Mathematical Problem Solving on AIME 2024 (Pass@1, Avg)
76Pass@1Qwen3-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-8Bk (responses per question)=32, sampling temperature=0.7, top-p=0.952025.07 | 76 | — | |
| DeepSeek-R1-Distill-32Bk (responses per question)=32, sampling temperature=0.7, top-p=0.952025.07 | 72.6 | — | |
| QUESTA-Nemotron-1.5Bk (responses per question)=32, sampling temperature=0.7, top-p=0.952025.07 | 72.5 | — | |
| Nemotron-1.5Bk (responses per question)=32, sampling temperature=0.7, top-p=0.952025.07 | 61.77 | — | |
| Qwen3-1.7Bk (responses per question)=32, sampling temperature=0.7, top-p=0.952025.07 | 48.3 | — | |
| Qwen-2-7B w/ T3RLModel Category=Vanilla Model, Backbone=Qwen-2-7B, Training Method=T3RL2026.03 | 40 | 68.1 | |
| Qwen-2-7B w/ TTRLModel Category=Vanilla Model, Backbone=Qwen-2-7B, Training Method=TTRL2026.03 | 36.4 | 65.5 | |
| DeepSeek-R1-Distill-1.5Bk (responses per question)=32, sampling temperature=0.7, top-p=0.952025.07 | 28.7 | — | |
| Qwen-2.5-Math-1.5B w/ T3RLModel Category=Math Model, Backbone=Qwen-2.5-Math-1.5B, Training Method=T3RL2026.03 | 20.8 | 48.8 | |
| Llama-3-8B-Instruct w/ T3RLModel Category=Instruct Model, Backbone=Llama-3-8B-Instruct, Training Method=T3RL2026.03 | 17.1 | 38.2 | |
| Qwen-2.5-Math-1.5B w/ TTRLModel Category=Math Model, Backbone=Qwen-2.5-Math-1.5B, Training Method=TTRL2026.03 | 15.8 | 45.9 | |
| Llama-3-8B-Instruct w/ TTRLModel Category=Instruct Model, Backbone=Llama-3-8B-Instruct, Training Method=TTRL2026.03 | 13.3 | 35.4 | |
| Llama-3.2-1B-Instruct w/ T3RLModel Category=Instruct Model, Backbone=Llama-3.2-1B-Instruct, Training Method=T3RL2026.03 | 8.3 | 23.5 | |
| Qwen-2.5-Math-1.5B (Baseline)Model Category=Math Model, Backbone=Qwen-2.5-Math-1.5B, Training Method=Baseline2026.03 | 7.7 | 23 | |
| Llama-3.2-1B-Instruct w/ TTRLModel Category=Instruct Model, Backbone=Llama-3.2-1B-Instruct, Training Method=TTRL2026.03 | 7.5 | 21.5 | |
| Llama-3-8B-Instruct (Baseline)Model Category=Instruct Model, Backbone=Llama-3-8B-Instruct, Training Method=Baseline2026.03 | 6 | 23.1 | |
| Qwen-2.5-1.5B w/ T3RLModel Category=Vanilla Model, Backbone=Qwen-2.5-1.5B, Training Method=T3RL2026.03 | 4.1 | 33.3 | |
| Qwen-2.5-1.5B w/ TTRLModel Category=Vanilla Model, Backbone=Qwen-2.5-1.5B, Training Method=TTRL2026.03 | 3.5 | 31.8 | |
| Llama-3.2-1B-Instruct (Baseline)Model Category=Instruct Model, Backbone=Llama-3.2-1B-Instruct, Training Method=Baseline2026.03 | 0.8 | 3.1 | |
| Qwen-2.5-1.5B (Baseline)Model Category=Vanilla Model, Backbone=Qwen-2.5-1.5B, Training Method=Baseline2026.03 | 0.2 | 2.8 | |
| Qwen-2-7B (Baseline)Model Category=Vanilla Model, Backbone=Qwen-2-7B, Training Method=Baseline2026.03 | 0 | 21 |