Logical Reasoning on LogiQA original (test)
43.16AccuracyLlama3.1-8B-Instruct (FAIR) + Teacher-Multiple
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama3.1-8B-Instruct (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 43.16 | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 40.86 | |
| GPT-3.5-TurboZero-shot evaluation=true2024.10 | 40.55 | |
| Gemini-1.0-ProZero-shot evaluation=true2024.10 | 39.94 | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 39.94 | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 38.71 | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 38.4 | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 37.02 | |
| Llama3.1-8B-InstructZero-shot evaluation=true2024.10 | 36.56 | |
| Llama2-7B-chat (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 36.25 | |
| ORCA2-7BZero-shot evaluation=true2024.10 | 35.02 | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 34.25 | |
| Mixtral-8x7B-Instruct-v0.1Zero-shot evaluation=true2024.10 | 34.19 | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 33.95 | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 33.03 | |
| Llama2-7B-chat (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 32.72 | |
| Llama2-7B-chat (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 32.1 | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 32.1 | |
| Llama2-7B-chat (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 31.04 | |
| Llama2-7B-chat (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 29.65 | |
| Qwen2.5-1.5B-InstructZero-shot evaluation=true2024.10 | 19.97 | |
| Llama2-7B-chatZero-shot evaluation=true2024.10 | 18.74 |