Multimodal Reasoning on V* Bench Tool-needed
90.1AccuracyPixelis (Qwen3-VL-8B-Instruct)
Evaluation Results
| Method | Links | |
|---|---|---|
| Pixelis (Qwen3-VL-8B-Instruct)Size=8B, Backbone=Qwen3-VL-8B-Instruct2026.03 | 90.1 | |
| Seed-1.5-VLSize=20B2026.03 | 89.5 | |
| Qwen3-VL-30B-A3B-InstructSize=30B2026.03 | 89.5 | |
| Pixelis: SFT + RFTSize=8B2026.03 | 89.2 | |
| PRM (process reward; tools; 8B)Size=8B2026.03 | 88.9 | |
| Pixelis: SFT + TTRLSize=8B2026.03 | 88.5 | |
| Pixel Reasoner (Qwen3-VL-8B-Instruct)Size=8B, Backbone=Qwen3-VL-8B-Instruct2026.03 | 88.3 | |
| Pixelis: RFT + TTRLSize=8B2026.03 | 88.1 | |
| Step Self-Consistency (step-level)Size=8B2026.03 | 88 | |
| Pixelis: SFT onlySize=8B2026.03 | 87.7 | |
| Late-fusionSize=8B2026.03 | 87.1 | |
| RV Self-Consistency (answer-only)Size=8B2026.03 | 86.9 | |
| Pixelis: RFT onlySize=8B2026.03 | 86.9 | |
| Pixelis: TTRL onlySize=8B2026.03 | 86.8 | |
| Qwen3-VL-8B-InstructSize=8B2026.03 | 86.4 |