General Multi-task Language and Vision Understanding on 24-benchmark suite S2 visual-friendly (test)
44.7Macro ScoreVLM + Foveation
Evaluation Results
| Method | Links | |
|---|---|---|
| VLM + FoveationModel Scale=4B, Backbone=Qwen3 Instruct2026.05 | 44.7 | |
| Oracle-routedModel Scale=4B, Backbone=Qwen3 Instruct2026.05 | 44.7 | |
| VLMModel Scale=4B, Backbone=Qwen3 Instruct2026.05 | 44 | |
| Decision-routedModel Scale=4B, Backbone=Qwen3 Instruct2026.05 | 41.1 | |
| LLMModel Scale=4B, Backbone=Qwen3 Instruct2026.05 | 34.7 |