Mathematical Reasoning on GSM-Hard (Accuracy, AVG, Improvement Overhead)
89.52AccuracyGPT-5-High*
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5-High*Size (# Param.)=100B+, Reasoning Effort=high, Temperature=0.02025.09 | 89.52 | — | — | |
| Qwen3-235B-A22BFramework=Reference Model, Backbone Model=Qwen3-235B-A22B2025.08 | 72.1 | 78.22 | — | |
| COCO Qwen3-8B with coco(Llama-3.1-8B)Framework=COCO Framework (Proposed), Backbone Model=Qwen3-8B, Monitor Model=Llama-3.1-8B2025.08 | 69.89 | 74.37 | 6.5 | |
| COCO Qwen3-8B with coco(Qwen3-8B)Framework=COCO Framework (Proposed), Backbone Model=Qwen3-8B, Monitor Model=Qwen3-8B2025.08 | 68.75 | 74.18 | 6.2 | |
| AFlow-MathBackbone=Qwen-2.5-72B-Instruct2025.05 | 68.4 | — | — | |
| Qwen3-8BFramework=Reference Model, Backbone Model=Qwen3-8B2025.08 | 68.08 | 68.52 | — | |
| MAS-GPTBackbone=Llama-3.3-70B-Instruct2025.05 | 67 | — | — | |
| MAS-GPTBackbone=Qwen-2.5-72B-Instruct2025.05 | 65.4 | — | — | |
| CoTBackbone=Qwen-2.5-72B-Instruct2025.05 | 64.2 | — | — | |
| RM-RegenBase Model=GPT-3.52026.03 | 64 | — | — | |
| Aflow-Qwen3-8BFramework=Multi-Agent Framework Baseline, Backbone Model=Qwen3-8B2025.08 | 63.99 | 69.86 | — | |
| Gemma-4B-Finetuned (NanoFlux-200)Size (# Param.)=4B, FLOPs=2.10×10^162025.09 | 63.3 | — | — | |
| SingleBackbone=Qwen-2.5-72B-Instruct2025.05 | 63.2 | — | — | |
| SCBackbone=Qwen-2.5-72B-Instruct2025.05 | 63.2 | — | — | |
| AutoGenBackbone=Qwen-2.5-72B-Instruct2025.05 | 63.2 | — | — | |
| MacNetBackbone=Qwen-2.5-72B-Instruct2025.05 | 63 | — | — | |
| DyLANBackbone=Qwen-2.5-72B-Instruct2025.05 | 62.4 | — | — | |
| DebateBackbone=Qwen-2.5-72B-Instruct2025.05 | 62 | — | — | |
| MADBackbone=Qwen-2.5-72B-Instruct2025.05 | 61.6 | — | — | |
| COCO Llama-3.1-8B with coco(Qwen3-8B)Framework=COCO Framework (Proposed), Backbone Model=Llama-3.1-8B, Monitor Model=Qwen3-8B2025.08 | 60.23 | 63.59 | 9.5 | |
| AFlow-MathBackbone=Llama-3.3-70B-Instruct2025.05 | 59.8 | — | — | |
| AgentVerseBackbone=Qwen-2.5-72B-Instruct2025.05 | 57.6 | — | — | |
| Gemma-4B-Finetuned (Full Dataset)Size (# Param.)=4B, FLOPs=9.23×10^162025.09 | 57.4 | — | — | |
| CoTBackbone=Llama-3.3-70B-Instruct2025.05 | 57 | — | — | |
| MacNetBackbone=Llama-3.3-70B-Instruct2025.05 | 56.6 | — | — | |
| Aflow-Llama3.1-8BFramework=Multi-Agent Framework Baseline, Backbone Model=Llama-3.1-8B2025.08 | 56.48 | 58.09 | — | |
| DebateBackbone=Llama-3.3-70B-Instruct2025.05 | 53.6 | — | — | |
| DyLANBackbone=Llama-3.3-70B-Instruct2025.05 | 53.6 | — | — | |
| SCBackbone=Llama-3.3-70B-Instruct2025.05 | 53.4 | — | — | |
| AutoGenBackbone=Llama-3.3-70B-Instruct2025.05 | 53 | — | — | |
| SingleBackbone=Llama-3.3-70B-Instruct2025.05 | 52.8 | — | — | |
| MADBackbone=Llama-3.3-70B-Instruct2025.05 | 52.6 | — | — | |
| COCO Llama-3.1-8B with coco(Llama-3.1-8B)Framework=COCO Framework (Proposed), Backbone Model=Llama-3.1-8B, Monitor Model=Llama-3.1-8B2025.08 | 52.01 | 58.46 | 0.63 | |
| AgentVerseBackbone=Llama-3.3-70B-Instruct2025.05 | 51.2 | — | — | |
| Gemma-4BSize (# Param.)=4B2025.09 | 48.1 | — | — | |
| RM-PrimedModel=GPT-3.52026.03 | 44.6 | 66.01 | — | |
| RM-Primed (R+)Model=GPT-3.5, R+ selection=only correct entries used2026.03 | 42.5 | 64.99 | — | |
| Few-shot CoTModel=GPT-3.52026.03 | 41.25 | 61.15 | — | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 39.8 | — | — | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 39.6 | — | — | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 39.6 | — | — | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 39.6 | — | — | |
| Contrastive CoTModel=GPT-3.5, Reflection=without2026.03 | 39.4 | 57.01 | — | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 39.4 | — | — | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 39.4 | — | — | |
| Contrastive CoTModel=GPT-3.5, Reflection=with2026.03 | 38 | 60.33 | — | |
| Llama-3.1-8BFramework=Reference Model, Backbone Model=Llama-3.1-8B2025.08 | 36.69 | 55.48 | — | |
| RM-Primed (R+)Model=Llama3-8B, R+ selection=only correct entries used2026.03 | 35.8 | 65.94 | — | |
| MAVBackbone=Llama-3.3-70B-Instruct2025.05 | 35.6 | — | — | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 34.8 | — | — | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 34.6 | — | — | |
| RM-PrimedModel=Llama3-8B2026.03 | 34.2 | 68.23 | — | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 33.8 | — | — | |
| Contrastive CoTModel=Llama3-8B, Reflection=without2026.03 | 33.2 | 59.56 | — | |
| Contrastive CoTModel=Llama3-8B, Reflection=with2026.03 | 33 | 61.13 | — | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 32.8 | — | — | |
| Few-shot CoTModel=Llama3-8B2026.03 | 31.8 | 64.19 | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 31.8 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 31.8 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 31.4 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 31.2 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 31.2 | — | — | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 31 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 30 | — | — | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 29.8 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 23.4 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 23 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 22.8 | — | — | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 21.4 | — | — | |
| MAVBackbone=Qwen-2.5-72B-Instruct2025.05 | 20.4 | — | — |