Multimodal Reasoning on ZeroBench
26.35AccuracyNPO
Evaluation Results
| Method | Links | |
|---|---|---|
| NPOstage=early-stage only2026.04 | 26.35 | |
| GRPO + ReMindBase Model=Qwen3-VL-8B-Instruct2026.06 | 26.05 | |
| DAPO + ReMindBase Model=Qwen3-VL-8B-Instruct2026.06 | 26.05 | |
| SDPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 25.15 | |
| RLSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 24.85 | |
| NPOstage=early + late-stage2026.04 | 24.85 | |
| AutoNPO2026.04 | 24.7 | |
| RePOBase Model=Qwen3-VL-8B-Instruct2026.06 | 23.5 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 22.6 | |
| GRPOtype=pure on-policy2026.04 | 22.6 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 22.6 | |
| GRPO + OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 22.16 | |
| OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 21.06 | |
| DAPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 20.66 | |
| LUFFYtype=external teacher2026.04 | 20.51 | |
| Base LLMBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 19.76 | |
| Qwen3-VL-8B-Instruct2026.04 | 19.76 | |
| Base ModelBase Model=Qwen3-VL-8B-Instruct2026.06 | 19.76 | |
| RLEPtype=far future2026.04 | 19.61 | |
| RLEPBase Model=Qwen3-VL-8B-Instruct2026.06 | 19.61 | |
| ExGRPOtype=historical replay2026.04 | 19.01 | |
| ExGRPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 19.01 |