Reasoning on AMC 2023, AIME 2025, HumanEval, and LiveCodeBench
83.4AccuracyASAG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ASAGModel=Qwen3-32B2026.06 | 83.4 | 58.8 | |
| VanillaModel=Qwen3-32B2026.06 | 82.3 | 100 | |
| ASAGModel=Qwen3-14B2026.06 | 81.8 | 61.1 | |
| VanillaModel=Qwen3-14B2026.06 | 79.5 | 100 | |
| ASAGModel=Qwen3-8B2026.06 | 77.8 | 55.4 | |
| ASAGModel=Qwen3-4B2026.06 | 75.4 | 63.5 | |
| VanillaModel=Qwen3-8B2026.06 | 73.9 | 100 | |
| VanillaModel=Qwen3-4B2026.06 | 72.5 | 100 | |
| ASAGModel=DeepSeek-R1-Distill-Llama-8B2026.06 | 59.9 | 45.1 | |
| ASAGModel=DeepSeek-R1-Distill-Qwen-7B2026.06 | 59.5 | 44.1 | |
| VanillaModel=DeepSeek-R1-Distill-Qwen-7B2026.06 | 55.3 | 100 | |
| VanillaModel=DeepSeek-R1-Distill-Llama-8B2026.06 | 55.3 | 100 |