Mathematical Reasoning on TheoremQA
51.2AccuracyHIPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| HIPPOTraining Data=DeepScaleR2026.06 | 51.2 | |
| Qwen3.5-9BShots=52026.04 | 49.25 | |
| Qwen2-72B2024.07 | 43.1 | |
| Qwen3-14BShots=52026.04 | 42.88 | |
| XekRung-8BShots=52026.04 | 41.63 | |
| Qwen3-8BShots=52026.04 | 40.12 | |
| Llama-3.3-70B-InstructShots=52026.04 | 38.88 | |
| Mixtral-8x22B2024.07 | 35.9 | |
| DeepSeekMath-7B-RLTraining Method=RL, Backbone=DeepSeekMath-7B2024.06 | 35.9 | |
| Qwen1.5-110B2024.07 | 34.9 | |
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 33.5 | |
| DART-Math-DSMath-7BTraining Method=SFT, Data Selection Strategy=Uniform, Backbone=DeepSeekMath-7B2024.06 | 32.5 | |
| Llama-3-70B2024.07 | 32.3 | |
| DeepSeekMath-7B-DART-MathBase Model=DeepSeekMath-7B, # Samples=590K2025.03 | 32.2 | |
| DART-Math-DSMath-7BTraining Method=SFT, Data Selection Strategy=Prop2Diff, Backbone=DeepSeekMath-7B2024.06 | 32.2 | |
| Qwen3-4B + SFTTraining Data=DeepScaleR2026.06 | 31.8 | |
| Qwen1.5-72B2024.07 | 29.3 | |
| Qwen3-4BTraining Data=DeepScaleR2026.06 | 29.1 | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 28.8 | |
| DeepSeekMath-7B-InstructBase Model=DeepSeekMath-7B, # Samples=780K2025.03 | 28.1 | |
| DeepSeekMath-7B-DART-Math†Base Model=DeepSeekMath-7B, # Samples=60K2025.03 | 27.4 | |
| DeepSeekMath-7B-RFTBase Model=DeepSeekMath-7B, # Samples=590K2025.03 | 27.2 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=195K2025.03 | 27 | |
| Llama-3.1-8B-InstructShots=52026.04 | 26 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=60K2025.03 | 24.6 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=30K, Training Strategy=Parallel2025.03 | 23.8 | |
| SecGPT-14BShots=52026.04 | 23.5 | |
| DeepSeekMath-7B-MMIQCBase Model=DeepSeekMath-7B, # Samples=2.3M2025.03 | 23.4 | |
| Mixtral-8x7BArchitecture=MoE, # Act Params=12B, # Params=47B2024.07 | 23.2 | |
| Llama3-8B-DART-Math†Base Model=Llama3-8B, # Samples=60K2025.03 | 22.9 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=30K, Training Strategy=Sequential2025.03 | 22.8 | |
| Foundation-Sec-8B-ReasoningShots=52026.04 | 22.25 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=60K2025.03 | 20 | |
| Llama3-8B-DART-MathBase Model=Llama3-8B, # Samples=590K2025.03 | 19.4 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=30K, Training Strategy=Conditional2025.03 | 19 | |
| DeepSeekMath-7B-MetaMathBase Model=DeepSeekMath-7B, # Samples=60K2025.03 | 18.9 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=30K, Training Strategy=Parallel2025.03 | 18.9 | |
| Mistral-7B-DART-Math†Base Model=Mistral-7B, # Samples=60K2025.03 | 18.2 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=60K2025.03 | 18.1 | |
| Mistral-7B-DART-MathBase Model=Mistral-7B, # Samples=590K2025.03 | 17 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=30K, Training Strategy=Sequential2025.03 | 17 | |
| Mistral-7B-WizardMath-V1.1Base Model=Mistral-7B, # Samples=418K2025.03 | 16.6 | |
| Mistral-7B-RFTBase Model=Mistral-7B, # Samples=590K2025.03 | 16.2 | |
| Mistral-7B-MMIQCBase Model=Mistral-7B, # Samples=2.3M2025.03 | 16.2 | |
| Llama3-8B-MMIQCBase Model=Llama3-8B, # Samples=2.3M2025.03 | 16.2 | |
| Llama3-8B-MetaMath†Base Model=Llama3-8B, # Samples=60K2025.03 | 16.1 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=30K, Training Strategy=Sequential2025.03 | 15.5 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=30K, Training Strategy=Conditional2025.03 | 15.5 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=30K, Training Strategy=Parallel2025.03 | 15.2 | |
| Llama3-8B-RFTBase Model=Llama3-8B, # Samples=590K2025.03 | 14.9 | |
| DeepSeekMath-7B-RefAugBase Model=DeepSeekMath-7B, # Samples=30K2025.03 | 14.4 | |
| DeepSeekMath-7B-RefAugBase Model=DeepSeekMath-7B, # Samples=60K2025.03 | 14 | |
| Mistral-7B-MetaMathBase Model=Mistral-7B, # Samples=400K2025.03 | 14 | |
| Llama3-8B-MetaMathBase Model=Llama3-8B, # Samples=400K2025.03 | 13.8 | |
| Llama3-8B-RefAugBase Model=Llama3-8B, # Samples=30K2025.03 | 13.6 | |
| Llama3-8B-RefAugBase Model=Llama3-8B, # Samples=60K2025.03 | 13 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=30K, Training Strategy=Conditional2025.03 | 12.8 | |
| Mistral-7B-MetaMath†Base Model=Mistral-7B, # Samples=60K2025.03 | 12.2 | |
| Mistral-7B-RefAugBase Model=Mistral-7B, # Samples=60K2025.03 | 11.1 | |
| DeepSeekMath-7B-StandardBase Model=DeepSeekMath-7B, # Samples=15K2025.03 | 11 | |
| Mistral-7B-RefAugBase Model=Mistral-7B, # Samples=30K2025.03 | 11 | |
| Llama3-8B-StandardBase Model=Llama3-8B, # Samples=15K2025.03 | 10.9 | |
| Llama3-8B-MMIQC†Base Model=Llama3-8B, # Samples=60K2025.03 | 10.6 | |
| Mistral-7B-StandardBase Model=Mistral-7B, # Samples=15K2025.03 | 7.6 | |
| DeepSeekMath-7B-MMIQC†Base Model=DeepSeekMath-7B, # Samples=60K2025.03 | 6.8 | |
| Mistral-7B-MMIQC†Base Model=Mistral-7B, # Samples=60K2025.03 | 5.9 | |
| Llama-Primus-Reasoning-8BShots=52026.04 | 5.37 |