Mathematics on AIME 2024 (Accuracy, Training Time)
83.8AccuracyReasoning Memory
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Reasoning MemoryBackbone=Qwen3-32B, Ret.=Proc., m=302026.04 | 83.8 | — | |
| Reasoning MemoryBackbone=Qwen3-32B, Ret.=Proc., m=82026.04 | 82.5 | — | |
| Length ScalingBackbone=Qwen3-32B, Ret.=N/A, m=302026.04 | 81.2 | — | |
| Document RAG (Google)Backbone=Qwen3-32B, Ret.=Fact., m=82026.04 | 79.5 | — | |
| Length ScalingBackbone=Qwen3-32B, Ret.=N/A, m=82026.04 | 79.2 | — | |
| No RAGBackbone=Qwen3-32B, Ret.=N/A, m=82026.04 | 78.9 | — | |
| Reasoning MemoryBackbone=OpenThinker3-7B, Ret.=Proc., m=302026.04 | 75.8 | — | |
| Qwen 3 VL 32B InstructParameters=32B2025.12 | 75.4 | — | |
| Latent-GRPOModel Scale=Qwen3-4B2026.01 | 74.6 | 3,108.44 | |
| BaseModel Scale=Qwen3-4B2026.01 | 73.8 | — | |
| GRPO (LLM-Judge)Model Scale=Qwen3-4B2026.01 | 72.9 | 6,753.47 | |
| Reasoning MemoryBackbone=OpenThinker3-7B, Ret.=Proc., m=82026.04 | 72.5 | — | |
| Trajectory RAG (Prefix)Backbone=Qwen3-32B, Ret.=Proc., m=82026.04 | 68.3 | — | |
| Olmo 3.1 32B InstructStage=Final Instruct 3.12025.12 | 67.8 | — | |
| Template RAG (Self-Discover)Backbone=Qwen3-32B, Ret.=Proc., m=82026.04 | 66.7 | — | |
| Template RAG (ReasonFlux)Backbone=Qwen3-32B, Ret.=Proc., m=82026.04 | 65.4 | — | |
| Length ScalingBackbone=OpenThinker3-7B, Ret.=N/A, m=302026.04 | 64.7 | — | |
| Trajectory RAG (Summary)Backbone=Qwen3-32B, Ret.=Proc., m=82026.04 | 62.9 | — | |
| Length ScalingBackbone=OpenThinker3-7B, Ret.=N/A, m=82026.04 | 61.2 | — | |
| Document RAG (CompactDS)Backbone=Qwen3-32B, Ret.=Fact., m=82026.04 | 60.4 | — | |
| Reasoning MemoryBackbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Proc., m=302026.04 | 57.5 | — | |
| Length ScalingBackbone=DeepSeek-R1-Distill-Llama-8B, Ret.=N/A, m=302026.04 | 54.8 | — | |
| kvtcModel=Qwen 2.5 R1 7B, Compression Ratio (CR)=9-11, Variant=8x, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 52.5 | — | |
| Reasoning MemoryBackbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Proc., m=82026.04 | 51.1 | — | |
| VanillaModel=Qwen 2.5 R1 7B, Compression Ratio (CR)=1, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 50.9 | — | |
| kvtcModel=Qwen 2.5 R1 7B, Compression Ratio (CR)=18-21, Variant=16x, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 50.9 | — | |
| Length ScalingBackbone=DeepSeek-R1-Distill-Llama-8B, Ret.=N/A, m=82026.04 | 50.8 | — | |
| Document RAG (Google)Backbone=OpenThinker3-7B, Ret.=Fact., m=82026.04 | 50.1 | — | |
| DMSModel=Qwen 2.5 R1 7B, Compression Ratio (CR)=-, Variant=8x, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 50 | — | |
| BaseModel Scale=Qwen3-1.7B2026.01 | 48.3 | — | |
| Document RAG (Google)Backbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Fact., m=82026.04 | 48.3 | — | |
| GRPO (LLM-Judge)Model Scale=Qwen3-1.7B2026.01 | 48.1 | 4,988.84 | |
| No RAGBackbone=OpenThinker3-7B, Ret.=N/A, m=82026.04 | 47 | — | |
| Template RAG (ReasonFlux)Backbone=OpenThinker3-7B, Ret.=Proc., m=82026.04 | 46.7 | — | |
| No RAGBackbone=DeepSeek-R1-Distill-Llama-8B, Ret.=N/A, m=82026.04 | 46.1 | — | |
| Document RAG (CompactDS)Backbone=OpenThinker3-7B, Ret.=Fact., m=82026.04 | 45.4 | — | |
| Template RAG (ReasonFlux)Backbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Proc., m=82026.04 | 45 | — | |
| Latent-GRPOModel Scale=Qwen3-1.7B2026.01 | 44.6 | 2,340.05 | |
| Template RAG (Self-Discover)Backbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Proc., m=82026.04 | 44.2 | — | |
| Template RAG (Self-Discover)Backbone=OpenThinker3-7B, Ret.=Proc., m=82026.04 | 44.2 | — | |
| Trajectory RAG (Summary)Backbone=OpenThinker3-7B, Ret.=Proc., m=82026.04 | 43.8 | — | |
| Trajectory RAG (Prefix)Backbone=OpenThinker3-7B, Ret.=Proc., m=82026.04 | 42.9 | — | |
| Trajectory RAG (Prefix)Backbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Proc., m=82026.04 | 42.5 | — | |
| Trajectory RAG (Summary)Backbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Proc., m=82026.04 | 40.8 | — | |
| Document RAG (CompactDS)Backbone=DeepSeek-R1-Distill-Llama-8B, Ret.=Fact., m=82026.04 | 35.6 | — | |
| Olmo 3.1 32B InstructStage=DPO2025.12 | 35.2 | — | |
| Gemma 3 27BParameters=27B2025.12 | 28.9 | — | |
| Qwen 3 32BThinking=No, Parameters=32B2025.12 | 27.9 | — | |
| kvtcModel=Qwen 2.5 R1 1.5B, Compression Ratio (CR)=18, Variant=16x, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 27.9 | — | |
| VanillaModel=Qwen 2.5 R1 1.5B, Compression Ratio (CR)=1, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 26.2 | — | |
| kvtcModel=Qwen 2.5 R1 1.5B, Compression Ratio (CR)=9, Variant=8x, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 25.4 | — | |
| DMSModel=Qwen 2.5 R1 1.5B, Compression Ratio (CR)=-, Variant=8x, Sampling temperature=0.6, top-p=95%, Number of runs=82025.11 | 23.3 | — | |
| Qwen 2.5 32BParameters=32B2025.12 | 15.7 | — | |
| Olmo 3.1 32B InstructStage=SFT2025.12 | 12.7 | — | |
| BaseModel Scale=Qwen3-0.6B2026.01 | 10.7 | — | |
| Latent-GRPOModel Scale=Qwen3-0.6B2026.01 | 10.6 | 2,280.52 | |
| GRPO (LLM-Judge)Model Scale=Qwen3-0.6B2026.01 | 9.4 | 4,082.96 | |
| Gemma 2 27BParameters=27B2025.12 | 4.7 | — | |
| OLMo 2 32BParameters=32B2025.12 | 4.6 | — | |
| Apertus 70BParameters=70B2025.12 | 0.31 | — |