Long-context reasoning on LongReason (Context Length Specific)
86.6Accuracy (32K Context)Qwen3-32B (Thinking)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-32B (Thinking)Model=Qwen3-32B (Thinking)2026.07 | 86.6 | 84.4 | 79.3 | 83.5 | |
| MAVEN (Qwen3-30B-A3B)Model=Qwen3-30B-A3B, Training Stage=MAVEN2026.07 | 86.6 | 85.1 | 81.7 | 84.5 | |
| Qwen3-30B-A3B + Outcome+Evidence ID (Prompted exploration)Model=Qwen3-30B-A3B, Training Stage=Outcome+Evidence ID, Exploration Strategy=Prompted exploration2026.07 | 85.6 | 82.5 | 79.6 | 82.6 | |
| Qwen3-30B-A3B + Outcome+Evidence IDModel=Qwen3-30B-A3B, Training Stage=Outcome+Evidence ID2026.07 | 85.3 | 82.6 | 79.3 | 82.4 | |
| Qwen3-30B-A3BModel=Qwen3-30B-A3B, Training Stage=Base2026.07 | 84.8 | 82.9 | 77.1 | 81.6 | |
| Qwen3-30B-A3B + OutcomeModel=Qwen3-30B-A3B, Training Stage=Outcome-only RL2026.07 | 84.8 | 81.7 | 77.2 | 80.9 | |
| QwenLong-L1-32BModel=QwenLong-L1-32B2026.07 | 84.1 | 83.6 | 75.1 | 80.9 | |
| Qwen3-30B-A3B + SFTModel=Qwen3-30B-A3B, Training Stage=SFT2026.07 | 83.8 | 82.6 | 76.3 | 80.9 | |
| MAVEN (Qwen2.5-14B)Model=Qwen2.5-14B, Training Stage=MAVEN2026.07 | 73 | 71.7 | 70.2 | 71.6 | |
| Qwen2.5-14B + Outcome+Evidence IDModel=Qwen2.5-14B, Training Stage=Outcome+Evidence ID2026.07 | 70.3 | 67.8 | 64.7 | 67.6 | |
| Qwen2.5-14B + Outcome+Evidence ID (Prompted exploration)Model=Qwen2.5-14B, Training Stage=Outcome+Evidence ID, Exploration Strategy=Prompted exploration2026.07 | 70.2 | 67.6 | 64.5 | 67.4 | |
| Qwen2.5-14B + OutcomeModel=Qwen2.5-14B, Training Stage=Outcome-only RL2026.07 | 69.5 | 67.4 | 62 | 66.3 | |
| Qwen2.5-14BModel=Qwen2.5-14B, Training Stage=Base2026.07 | 68.1 | 66.2 | 62.3 | 65.5 | |
| Qwen2.5-14B + SFTModel=Qwen2.5-14B, Training Stage=SFT2026.07 | 67.6 | 66.8 | 61.5 | 65.3 | |
| LLaMA-3.1-70BModel=LLaMA-3.1-70B2026.07 | 61.2 | 63.3 | 48.3 | 57.6 | |
| MAVEN (LLaMA-3.1-8B)Model=LLaMA-3.1-8B, Training Stage=MAVEN2026.07 | 55.9 | 55.2 | 55.8 | 55.6 | |
| LLaMA-3.1-8B + Outcome+Evidence ID (Prompted exploration)Model=LLaMA-3.1-8B, Training Stage=Outcome+Evidence ID, Exploration Strategy=Prompted exploration2026.07 | 52.4 | 51.3 | 48.9 | 50.8 | |
| LLaMA-3.1-8B + Outcome+Evidence IDModel=LLaMA-3.1-8B, Training Stage=Outcome+Evidence ID2026.07 | 52.1 | 50.9 | 48.7 | 50.6 | |
| LLaMA-3.1-8B + OutcomeModel=LLaMA-3.1-8B, Training Stage=Outcome-only RL2026.07 | 51.8 | 50.5 | 46.2 | 49.5 | |
| LLaMA-3.1-8BModel=LLaMA-3.1-8B, Training Stage=Base2026.07 | 51.4 | 49.9 | 46.5 | 49.3 | |
| LLaMA-3.1-8B + SFTModel=LLaMA-3.1-8B, Training Stage=SFT2026.07 | 51 | 49.2 | 47.1 | 49.1 |