Reasoning on ARC (test)
0.848AccuracySEAG
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SEAGModel=Llama3-8B-Instruct2025.01 | 0.848 | 46.15 | 15,993.29 | 655.9 | |
| SEModel=Llama3-8B-Instruct2025.01 | 0.83 | 96.32 | 42,507.85 | 1,692.92 | |
| COT-SCModel=Llama3-8B-Instruct2025.01 | 0.823 | 10 | 2,157.93 | 505.79 | |
| COTModel=Llama3-8B-Instruct2025.01 | 0.818 | 1 | 215.79 | 48.315 | |
| RAPModel=Llama3-8B-Instruct2025.01 | 0.812 | 196.96 | 88,670.52 | 3,636.17 | |
| ToTModel=Llama3-8B-Instruct2025.01 | 0.797 | 149.59 | 67,534.7 | 3,033.07 | |
| MPOModel=LLaMA-3 8B-Instruct2026.01 | 0.791 | — | — | — | |
| TextGradModel=LLaMA-3 8B-Instruct2026.01 | 0.756 | — | — | — | |
| Untuned promptModel=LLaMA-3 8B-Instruct2026.01 | 0.75 | — | — | — | |
| MPOModel=Mistral-7B-Instruct2026.01 | 0.7304 | — | — | — | |
| Untuned promptModel=Mistral-7B-Instruct2026.01 | 0.7073 | — | — | — | |
| TextGradModel=Mistral-7B-Instruct2026.01 | 0.703 | — | — | — |