Mathematical Reasoning on GSM8K HARD (test)
37.08AccuracyBest Model
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Best ModelBackbone LLM=LLaMA 3.1 8B, Shot count=0-shot2025.10 | 37.08 | 1 | -2.9 | |
| Best ModelBackbone LLM=Mistral 7B, Shot count=0-shot2025.10 | 19.1 | 2 | -6.2 | |
| Best ModelBackbone LLM=Qwen 2.5 0.5B, Shot count=0-shot2025.10 | 11.24 | 1 | 1.2 |