Science Question Answering on ARC randomly sampled 400 instances (test)
84.8AccuracySEAG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SEAGModel=Llama3-8B-Instruct2025.01 | 84.8 | 46.15 | |
| SEModel=Llama3-8B-Instruct2025.01 | 83 | 96.32 | |
| CoT-SCModel=Llama3-8B-Instruct2025.01 | 82.3 | 10 | |
| CoTModel=Llama3-8B-Instruct2025.01 | 81.8 | 1 | |
| RAPModel=Llama3-8B-Instruct2025.01 | 81.2 | 196.96 | |
| ToTModel=Llama3-8B-Instruct2025.01 | 79.7 | 149.59 | |
| SEAGModel=Mistral-7B-Instruct2025.01 | 72.5 | 110.95 | |
| SEModel=Mistral-7B-Instruct2025.01 | 71.3 | 155.13 | |
| CoT-SCModel=Mistral-7B-Instruct2025.01 | 70.8 | 10 | |
| RAPModel=Mistral-7B-Instruct2025.01 | 70.5 | 313.2 | |
| CoTModel=Mistral-7B-Instruct2025.01 | 68.8 | 1 | |
| SEAGModel=Llama2-13B-Chat2025.01 | 63.8 | 26.62 | |
| CoT-SCModel=Llama2-13B-Chat2025.01 | 63.7 | 10 | |
| RAPModel=Llama2-13B-Chat2025.01 | 63.7 | 247.22 | |
| SEModel=Llama2-13B-Chat2025.01 | 63.2 | 146.2 | |
| ToTModel=Mistral-7B-Instruct2025.01 | 61.5 | 221.36 | |
| CoTModel=Llama2-13B-Chat2025.01 | 59.8 | 1 | |
| ToTModel=Llama2-13B-Chat2025.01 | 59 | 209.05 |