Mathematical Reasoning on AIME24, AMC23, GameOf24 (test)
40AIME24 AccuracyAgentFlow (w/ Flow-GRPO)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| AgentFlow (w/ Flow-GRPO)Size=7B-Inst, Method Setting=Training-based2026.05 | 40 | 61.5 | 53 | 51.5 | 6.1 | |
| LuffySize=7B-Inst, Method Setting=Training-based2026.05 | 30.7 | 44.8 | 33 | 36.2 | 9.2 | |
| ToRLSize=7B-Inst, Method Setting=Training-based2026.05 | 20 | 60 | 31 | 37 | 8.4 | |
| SimpleRL-reasonSize=7B-Base, Method Setting=Training-based2026.05 | 16.7 | 60 | 33 | 36.6 | 8.8 | |
| Open-Reasoner-ZeroSize=7B-Base, Method Setting=Training-based2026.05 | 16.7 | 54.9 | 32 | 34.5 | 10.9 | |
| HASP-Evolve + RSSize=7B-Inst, Method Setting=Training-based2026.05 | 16.7 | 57.5 | 62 | 45.4 | — | |
| GPT-4oSize=~200B, Method Setting=Training-free / inference-time2026.05 | 13.3 | 60 | 32 | 35.1 | 3.7 | |
| GPT-4o-miniSize=~8B, Method Setting=Training-free / inference-time2026.05 | 13.3 | 57.5 | 16 | 28.9 | 9.9 | |
| AutoGenSize=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 13.3 | 57.5 | 24 | 31.6 | 7.2 | |
| General-ReasonerSize=7B-Base, Method Setting=Training-based2026.05 | 13.3 | 55 | 33 | 33.8 | 11.6 | |
| Prompt-Only SkillsSize=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 10 | 47.5 | 41 | 32.8 | 6 | |
| TIRSize=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 10 | 50 | 33.3 | 31.1 | 7.7 | |
| HASP-Intervention (w. Teacher)Size=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 10 | 56.5 | 50 | 38.8 | — | |
| Qwen2.5-7B-InstructSize=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 6.7 | 47.5 | 33 | 29.1 | 9.7 | |
| RA-Agent (multi-loop)Size=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 6.7 | 50 | 46 | 34.2 | 4.6 | |
| HASP-Intervention (PF-only)Size=7B-Inst, Method Setting=Training-free / inference-time2026.05 | 6.7 | 55 | 46 | 35.9 | 2.9 | |
| SFT (vanilla)Size=7B-Inst, Method Setting=Training-based2026.05 | 6.7 | 47.5 | 33 | 29.1 | 16.3 |