Multimodal Reasoning on ZeroBench sub
47.6Pass@1Seed2.0 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Seed2.0 Pro2026.06 | 47.6 | |
| Gemini-3-Pro High2026.06 | 42.2 | |
| Seed2.0 Lite2026.06 | 42.2 | |
| GPT-5.2 High2026.06 | 38.9 | |
| Seed1.82026.06 | 37.7 | |
| Seed2.0 Mini2026.06 | 36.2 | |
| Seed 1.5-VLthinking=true, decoding=greedy2025.05 | 30.8 | |
| Claude-Opus-4.52026.06 | 30.8 | |
| Seed 1.5-VLthinking=false, decoding=greedy2025.05 | 29 | |
| Gemini 1.5 Prothinking=true, decoding=greedy2025.05 | 26 | |
| Claude 3.7 Sonnetthinking=true, decoding=sampling2025.05 | 20.4 | |
| OpenAI o1thinking=true, decoding=greedy2025.05 | 20.2 | |
| GPT-4othinking=false, decoding=greedy2025.05 | 19.6 | |
| Qwen 2.5-VL 72Bthinking=false, decoding=greedy2025.05 | 13 | |
| LLaVA-OVScale=1.5-8B2026.06 | 11.98 | |
| Qwen2.5-VLScale=7B2026.06 | 8.99 | |
| MOSS-Video-PreviewSFT Mode=offline2026.06 | 8.53 | |
| MOSS-Video-PreviewSFT Mode=real-time2026.06 | 7.83 |