Mathematical Reasoning on AIME 2024-2025 (Accuracy (pass@1))
46.67Accuracy (pass@1)BF16
Evaluation Results
| Method | Links | |
|---|---|---|
| BF16Model=DeepSeek-R1-Distill-Qwen-7B, Bit-width=BF16, Maximum number of generation tokens=32,7682025.12 | 46.67 | |
| R1-CI-14B2025.05 | 42 | |
| MixKVQModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=C3.4, Maximum number of generation tokens=32,7682025.12 | 40 | |
| KIVIModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=KV4, Maximum number of generation tokens=32,7682025.12 | 36.67 | |
| KVQuantModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=KV4, Maximum number of generation tokens=32,7682025.12 | 33.33 | |
| RotateKVModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=KV4, Maximum number of generation tokens=32,7682025.12 | 33.33 | |
| Qwen-2.5-14BInference Mode=CodeSteer2025.05 | 33.3 | |
| KVTunerModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=C3.92, Maximum number of generation tokens=32,7682025.12 | 31.67 | |
| KIVIModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=KV2, Maximum number of generation tokens=32,7682025.12 | 30 | |
| Qwen-2.5-14BInference Mode=All Text2025.05 | 30 | |
| R1-CI-7B2025.05 | 15 | |
| Qwen-2.5-7BInference Mode=All Text2025.05 | 8.33 | |
| Qwen-2.5-7BInference Mode=CodeSteer2025.05 | 8.33 | |
| Standard AR2026.04 | 8.3 | |
| Best-of-16sampling=Best-of-162026.04 | 8.3 | |
| LPSR2026.04 | 8.3 | |
| CoCoNuT2026.04 | 6.7 | |
| STIR-Static2026.04 | 1.7 | |
| KVQuantModel=DeepSeek-R1-Distill-Qwen-7B, Bit-width=KV2, Maximum number of generation tokens=32,7682025.12 | 1.67 | |
| Prompted SC2026.04 | 0 |