Long-context reasoning on RULER (64K and 128K Context Evaluation)
93.4RULER Score (64K Context)MAVEN (Qwen3-30B-A3B)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MAVEN (Qwen3-30B-A3B)Model=Qwen3-30B-A3B, Training Stage=MAVEN2026.07 | 93.4 | 88.8 | 91.1 | |
| LLaMA-3.1-70BModel=LLaMA-3.1-70B2026.07 | 93.2 | 69.8 | 81.5 | |
| Qwen3-32B (Thinking)Model=Qwen3-32B (Thinking)2026.07 | 92.1 | 84.4 | 88.2 | |
| MAVEN (Qwen2.5-14B)Model=Qwen2.5-14B, Training Stage=MAVEN2026.07 | 90 | 81.6 | 85.8 | |
| Qwen3-30B-A3B + Outcome+Evidence ID (Prompted exploration)Model=Qwen3-30B-A3B, Training Stage=Outcome+Evidence ID, Exploration Strategy=Prompted exploration2026.07 | 89.8 | 84.5 | 87.2 | |
| Qwen3-30B-A3B + Outcome+Evidence IDModel=Qwen3-30B-A3B, Training Stage=Outcome+Evidence ID2026.07 | 89.6 | 84.4 | 87 | |
| Qwen3-30B-A3B + SFTModel=Qwen3-30B-A3B, Training Stage=SFT2026.07 | 88.5 | 82.7 | 85.6 | |
| MAVEN (LLaMA-3.1-8B)Model=LLaMA-3.1-8B, Training Stage=MAVEN2026.07 | 88.4 | 80.1 | 84.3 | |
| Qwen3-30B-A3BModel=Qwen3-30B-A3B, Training Stage=Base2026.07 | 88.2 | 82.6 | 85.4 | |
| Qwen3-30B-A3B + OutcomeModel=Qwen3-30B-A3B, Training Stage=Outcome-only RL2026.07 | 87.2 | 82.3 | 84.8 | |
| LLaMA-3.1-8B + Outcome+Evidence IDModel=LLaMA-3.1-8B, Training Stage=Outcome+Evidence ID2026.07 | 86.9 | 78.3 | 82.6 | |
| Qwen2.5-14B + Outcome+Evidence IDModel=Qwen2.5-14B, Training Stage=Outcome+Evidence ID2026.07 | 86.7 | 76.9 | 81.8 | |
| LLaMA-3.1-8B + Outcome+Evidence ID (Prompted exploration)Model=LLaMA-3.1-8B, Training Stage=Outcome+Evidence ID, Exploration Strategy=Prompted exploration2026.07 | 86.6 | 78.5 | 82.5 | |
| Qwen2.5-14B + Outcome+Evidence ID (Prompted exploration)Model=Qwen2.5-14B, Training Stage=Outcome+Evidence ID, Exploration Strategy=Prompted exploration2026.07 | 86.5 | 77.1 | 81.8 | |
| LLaMA-3.1-8B + OutcomeModel=LLaMA-3.1-8B, Training Stage=Outcome-only RL2026.07 | 85.8 | 77.9 | 81.9 | |
| LLaMA-3.1-8B + SFTModel=LLaMA-3.1-8B, Training Stage=SFT2026.07 | 85.6 | 77.4 | 81.5 | |
| LLaMA-3.1-8BModel=LLaMA-3.1-8B, Training Stage=Base2026.07 | 85.1 | 77.2 | 81.1 | |
| Qwen2.5-14B + OutcomeModel=Qwen2.5-14B, Training Stage=Outcome-only RL2026.07 | 84.1 | 75.7 | 79.9 | |
| Qwen2.5-14B + SFTModel=Qwen2.5-14B, Training Stage=SFT2026.07 | 83.9 | 75.2 | 79.6 | |
| Qwen2.5-14BModel=Qwen2.5-14B, Training Stage=Base2026.07 | 83.7 | 75.5 | 79.6 | |
| QwenLong-L1-32BModel=QwenLong-L1-32B2026.07 | 81.7 | 74.3 | 78 |