Scientific Multimodal Reasoning on VRSBench
74.32AccuracyS1-VL-32B-RL
Evaluation Results
| Method | Links | |
|---|---|---|
| S1-VL-32B-RLParameter Count=32B, Training Phase=RL2026.04 | 74.32 | |
| S1-VL-32B-SFTParameter Count=32B, Training Phase=SFT2026.04 | 72.34 | |
| Qwen3-VL-235B-A22B-ThinkingParameter Count=235B-A22B, Reasoning Mode=Thinking2026.04 | 68.94 | |
| Qwen3-VL-32B-ThinkingParameter Count=32B, Reasoning Mode=Thinking2026.04 | 68.41 | |
| GPT-52026.04 | 65.89 | |
| Gemini 2.5 Pro2026.04 | 65.7 | |
| Gemini 2.5 Flash2026.04 | 64.23 | |
| Intern-S1Parameter Count=235B+6B2026.04 | 63.48 | |
| Thyme-VLParameter Count=7B2026.04 | 57.61 | |
| Intern-S1-miniParameter Count=8B2026.04 | 54.89 |