Multi-hop Reasoning on StrategyQA
95.6AccuracyOpenMath2-Llama3.1-70B*
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenMath2-Llama3.1-70B*Model Series=OpenMath2, Base=Llama-3.1-70B, Source=Shen et al. (2025)2025.02 | 95.6 | |
| Qwen2.5-Math-72B-InstructModel Series=Qwen2.5-Math, Size=72B2025.02 | 94.3 | |
| Qwen-2-72B-Instruct#Tok=800, TFLOPs=115.2, Lat. (s)=10.5, GPU DRAM Usage (GB)=> 145.02026.05 | 92.5 | |
| Qwen2.5-Math-7B-S2R-ORLModel Series=Qwen2.5-Math, Size=7B, RL Type=Outcome-level RL (ORL)2025.02 | 90.8 | |
| Llama-3.1-70B-Instruct*Model Series=Llama-3.1, Size=70B, Source=Shen et al. (2025)2025.02 | 88.8 | |
| Qwen2.5-Math-7B-S2R-BIModel Series=Qwen2.5-Math, Size=7B, RL Type=Binary Reward (BI)2025.02 | 88.7 | |
| QwQ-32B-Preview*Model Series=QwQ, Size=32B, Source=Shen et al. (2025)2025.02 | 88.2 | |
| DynaGraph (Ours, 8B)#Tok=1,220, TFLOPs=19.5, Lat. (s)=15.3, GPU DRAM Usage (GB)=16.6 (O(1))2026.05 | 87.6 | |
| Reflexion#Tok=3,890, TFLOPs=62.2, Lat. (s)=44.5, GPU DRAM Usage (GB)=16.52026.05 | 86.2 | |
| Standard ToT#Tok=3,400, TFLOPs=54.4, Lat. (s)=39.5, GPU DRAM Usage (GB)=16.52026.05 | 85.8 | |
| ReAct#Tok=2,480, TFLOPs=39.7, Lat. (s)=30.5, GPU DRAM Usage (GB)=16.52026.05 | 84.5 | |
| Multi-Agent (3 × 8B)#Tok=1,520, TFLOPs=24.3, Lat. (s)=17.5, GPU DRAM Usage (GB)=> 49.52026.05 | 84 | |
| StrategyLLM-SCBackbone=Meta-Llama-3-70B-Instruct2023.11 | 83.5 | |
| Self-Consistency (k=5)#Tok=1,700, TFLOPs=27.2, Lat. (s)=19.8, GPU DRAM Usage (GB)=16.52026.05 | 83.3 | |
| StrategyLLMBackbone=Meta-Llama-3-70B-Instruct2023.11 | 82 | |
| CoT-SCBackbone=Meta-Llama-3-70B-Instruct2023.11 | 81.5 | |
| Qwen2.5-Math-7B-InstructModel Series=Qwen2.5-Math, Size=7B2025.02 | 81.2 | |
| CoTBackbone=Meta-Llama-3-70B-Instruct2023.11 | 80.5 | |
| ShadowKVBackbone=DeepSeek-R1-Distill-Llama-8B2025.12 | 80 | |
| Eurus-2-7B-PRIMEModel Series=Eurus-2, Size=7B2025.02 | 79 | |
| H2OBackbone=DeepSeek-R1-Distill-Llama-8B, Cache Budget=5122025.12 | 79 | |
| SolutionLLMBackbone=Meta-Llama-3-70B-Instruct2023.11 | 79 | |
| Standard CoT#Tok=650, TFLOPs=10.4, Lat. (s)=8.0, GPU DRAM Usage (GB)=16.52026.05 | 78.4 | |
| StrategyLLM-SCBackbone=Mixtral-8x22B-Instruct-v0.12023.11 | 77 | |
| StrategyLLMBackbone=Mixtral-8x22B-Instruct-v0.12023.11 | 76.5 | |
| StrategyLLM-SCBackbone=Mixtral-8x7B-Instruct-v0.12023.11 | 75 | |
| CoT-SCBackbone=Mixtral-8x22B-Instruct-v0.12023.11 | 75 | |
| FullBackbone=DeepSeek-R1-Distill-Llama-8B2025.12 | 74 | |
| StrategyLLMBackbone=Meta-Llama-3-8B-Instruct2023.11 | 74 | |
| StrategyLLM-SCBackbone=Meta-Llama-3-8B-Instruct2023.11 | 74 | |
| CoT-SCBackbone=Mixtral-8x7B-Instruct-v0.12023.11 | 73.5 | |
| StrategyLLMBackbone=Mixtral-8x7B-Instruct-v0.12023.11 | 73.5 | |
| SolutionLLMBackbone=Mixtral-8x22B-Instruct-v0.12023.11 | 72 | |
| CoTBackbone=Mixtral-8x22B-Instruct-v0.12023.11 | 72 | |
| CoT-SCBackbone=Meta-Llama-3-8B-Instruct2023.11 | 71 | |
| Standard Prompting#Tok=210, TFLOPs=3.4, Lat. (s)=2.5, GPU DRAM Usage (GB)=16.52026.05 | 65.2 | |
| SolutionLLMBackbone=Meta-Llama-3-8B-Instruct2023.11 | 64 | |
| CoTBackbone=Mixtral-8x7B-Instruct-v0.12023.11 | 63 | |
| LoTVerifier Training Dataset=AQuA2025.03 | 62.3 | |
| Qwen2.5-Math-7BModel Series=Qwen2.5-Math, Size=7B2025.02 | 61.1 | |
| CoTBackbone=Meta-Llama-3-8B-Instruct2023.11 | 61 | |
| SolutionLLMBackbone=Mixtral-8x7B-Instruct-v0.12023.11 | 61 | |
| LoTVerifier Training Dataset=MMLU2025.03 | 53 | |
| LoTVerifier Training Dataset=CommonSenseQA2025.03 | 53 | |
| LoTVerifier Training Dataset=StrategyQA2025.03 | 43 | |
| Thoughts-as-PlanningAvg Edits=472026.04 | 33.9 | |
| CoTGenAvg Edits=1802026.04 | 31.5 | |
| RLCoTAvg Edits=1502026.04 | 30.9 | |
| SoftCoT2026.04 | 30.6 | |
| Manual CoT2026.04 | 29.3 |