Mathematical Reasoning on GSM8K randomly sampled 400 instances (test)
86AccuracySEAG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SEAGModel=Llama3-8B-Instruct2025.01 | 86 | 41.69 | |
| SEModel=Llama3-8B-Instruct2025.01 | 85 | 82.63 | |
| CoT-SCModel=Llama3-8B-Instruct2025.01 | 84.5 | 10 | |
| RAPModel=Llama3-8B-Instruct2025.01 | 82.5 | 128.4 | |
| ToTModel=Llama3-8B-Instruct2025.01 | 78.5 | 104.8 | |
| CoTModel=Llama3-8B-Instruct2025.01 | 76.2 | 1 | |
| SEAGModel=Mistral-7B-Instruct2025.01 | 68.5 | 84.14 | |
| SEModel=Mistral-7B-Instruct2025.01 | 67.5 | 98.79 | |
| RAPModel=Mistral-7B-Instruct2025.01 | 66.7 | 161.86 | |
| CoT-SCModel=Mistral-7B-Instruct2025.01 | 66.5 | 10 | |
| ToTModel=Mistral-7B-Instruct2025.01 | 61.3 | 122.16 | |
| CoTModel=Mistral-7B-Instruct2025.01 | 50.2 | 1 | |
| SEAGModel=Llama2-13B-Chat2025.01 | 43.5 | 53.09 | |
| CoT-SCModel=Llama2-13B-Chat2025.01 | 41.7 | 10 | |
| SEModel=Llama2-13B-Chat2025.01 | 40.3 | 69.47 | |
| ToTModel=Llama2-13B-Chat2025.01 | 37.8 | 102.06 | |
| RAPModel=Llama2-13B-Chat2025.01 | 37.2 | 121.69 | |
| CoTModel=Llama2-13B-Chat2025.01 | 33 | 1 |