Mathematical Reasoning on DeepMath hard (held-out)
73Pass@1Agon
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AgonModel=Gemma-4-E4B, Protocol=single training run, cascade protocol2026.07 | 73 | 15 | — | |
| AgonModel=Qwen3-4B, Protocol=single training run, cascade protocol2026.07 | 71 | 12 | — | |
| AgonModel=Qwen3-1.7B, Protocol=single training run, cascade protocol2026.07 | 70 | 24 | — | |
| AgonModel=Qwen3.5-2B, Protocol=single training run, cascade protocol2026.07 | 70 | 20 | — | |
| AgonModel=Qwen3-0.6B, Protocol=single training run, cascade protocol2026.07 | 61 | 31 | — | |
| Agon (competition + exchange)Model=Qwen3-0.6B, Inference Passes=two sequential generation passes, Training Protocol=Competition + Exchange2026.07 | 61 | — | 3,500 | |
| Vanilla GRPOModel=Qwen3-4B, Protocol=single training run, cascade protocol2026.07 | 59 | — | — | |
| Vanilla GRPOModel=Gemma-4-E4B, Protocol=single training run, cascade protocol2026.07 | 58 | — | — | |
| Zero-shotModel=Qwen3-4B, Protocol=single evaluation2026.07 | 52 | — | — | |
| Vanilla GRPOModel=Qwen3.5-2B, Protocol=single training run, cascade protocol2026.07 | 50 | — | — | |
| Zero-shotModel=Gemma-4-E4B, Protocol=single evaluation2026.07 | 50 | — | — | |
| Vanilla GRPOModel=Qwen3-1.7B, Protocol=single training run, cascade protocol2026.07 | 46 | — | — | |
| Cooperative exchangeModel=Qwen3-0.6B, Inference Passes=two sequential generation passes, Training Protocol=Cooperative2026.07 | 46 | — | 5,100 | |
| Zero-shotModel=Qwen3.5-2B, Protocol=single evaluation2026.07 | 44 | — | — | |
| Zero-shotModel=Qwen3-1.7B, Protocol=single evaluation2026.07 | 38 | — | — | |
| GRPO two-pass self-cascade (control)Model=Qwen3-0.6B, Inference Passes=two sequential generation passes, Training Protocol=GRPO2026.07 | 35 | — | 8,000 | |
| MoA (no train)Model=Qwen3-0.6B, Inference Passes=single evaluation, Training=None2026.07 | 34 | — | 6,900 | |
| Self-refinement (control)Model=Qwen3-0.6B, Inference Passes=two sequential generation passes, Training Protocol=Self-refinement2026.07 | 32 | — | 7,900 | |
| Competitive, shared opponent (single-pass)Model=Qwen3-0.6B, Inference Passes=single-pass, Training Protocol=Competitive2026.07 | 32 | — | 7,400 | |
| Vanilla GRPOModel=Qwen3-0.6B, Protocol=single training run, cascade protocol2026.07 | 30 | — | — | |
| Vanilla GRPO (baseline)Model=Qwen3-0.6B, Inference Passes=single-pass, Training Protocol=GRPO2026.07 | 30 | — | 8,100 | |
| Zero-shotModel=Qwen3-0.6B, Protocol=single evaluation2026.07 | 23 | — | — | |
| Zero-shotModel=Qwen3-0.6B, Inference Passes=single-pass, Training=None2026.07 | 23 | — | 6,100 |