Mathematical Reasoning on MATH-500 (AUCOAA)
91.9AUCOAANormalized-Length
Evaluation Results
| Method | Links | |
|---|---|---|
| Normalized-LengthApproach=Normalized length reward, Backbone=Qwen3-8B2026.01 | 91.9 | |
| Format-Adaptive-AnswerApproach=Format-based SFT + RL, Backbone=Qwen3-8B2026.01 | 91.2 | |
| Adaptive-AnswerApproach=Rejection sampling + RL, Backbone=Qwen3-8B2026.01 | 90.3 | |
| TWYNApproach=Baseline, Backbone=Qwen3-8B2026.01 | 89.6 | |
| Hard-Length 8k → 4kApproach=Length decay strategy, Backbone=Qwen3-8B2026.01 | 89.1 | |
| SFTApproach=Standard SFT, Backbone=Qwen3-8B2026.01 | 87.9 | |
| Hard-Length 8kApproach=Hard length constraint 8k, Backbone=Qwen3-8B2026.01 | 87.9 | |
| Soft-LengthApproach=Soft length constraint, Backbone=Qwen3-8B2026.01 | 87.9 | |
| Hard-Length 16kApproach=Hard length constraint 16k, Backbone=Qwen3-8B2026.01 | 83.4 | |
| Base modelApproach=Standard Reasoning, Backbone=Qwen3-8B2026.01 | 81.7 | |
| No-ThinkingApproach=No CoT, Backbone=Qwen3-8B2026.01 | 81.3 |