General Multi-task Language and Vision Understanding on 24-benchmark suite All (test)
50.2Macro-average ScoreOracle-routed
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Oracle-routedModel Scale=4B, Backbone=Qwen3 Instruct, Inference Strategy=post-hoc best selection, bound=Upper Bound2026.05 | 50.2 | 1,894 | 4.3 | 9.5 | |
| Decision-routedModel Scale=4B, Backbone=Qwen3 Instruct, Inference Strategy=label-free routing, tau=1.282026.05 | 47.4 | 1,876 | 1.5 | 10.3 | |
| LLMModel Scale=4B, Backbone=Qwen3 Instruct, Inference Strategy=text-only2026.05 | 45.9 | 2,092 | — | — | |
| VLM + FoveationModel Scale=4B, Backbone=Qwen3 Instruct, Inference Strategy=visual path with localized inverse transport2026.05 | 44.8 | 1,521 | 1.1 | 27.3 | |
| VLMModel Scale=4B, Backbone=Qwen3 Instruct, Inference Strategy=visual path (naive VTC)2026.05 | 44.4 | 1,486 | 1.5 | 29 |