Visual Question Answering on OmniEarth-Bench Multiple-choice question protocol 1.0 (test)
90Cross-sphere AccuracyExperts
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ExpertsModel Type=Human Experts, Evaluation protocol=Multiple-choice2026.05 | 90 | 96 | 91 | 95 | 93 | 97 | 95 | 93.86 | |
| ScaleEarthBackbone=Qwen3-VL-8B, Adaptation Stage=Stage-2 (Full), Adaptation Strategy=CS-HLoRA, Evaluation protocol=Multiple-choice2026.05 | 47.2 | 38.5 | 45.3 | 34.8 | 40.8 | 39.4 | 38.95 | 40.71 | |
| InternVL3-7BModel Type=Open-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 42.85 | 30.1 | 37.47 | 20.28 | 49.27 | 28.74 | 23.18 | 33.13 | |
| Stage-1 + Bucketed MoE-LoRABackbone=Qwen3-VL-8B, Adaptation Stage=Stage-1, Adaptation Strategy=Bucketed MoE-LoRA, PEFT configuration=K=3, GSD-bucket-routed, Evaluation protocol=Multiple-choice2026.05 | 42.13 | 35.86 | 41.95 | 32.4 | 40.32 | 37.18 | 35.42 | 37.89 | |
| Stage-1 + MoLEBackbone=Qwen3-VL-8B, Adaptation Stage=Stage-1, Adaptation Strategy=MoLE, Evaluation protocol=Multiple-choice2026.05 | 37.85 | 32.4 | 39.21 | 29.62 | 40.07 | 34.51 | 32.3 | 35.14 | |
| Stage-1 + LoRAMoEBackbone=Qwen3-VL-8B, Adaptation Stage=Stage-1, Adaptation Strategy=LoRAMoE, Evaluation protocol=Multiple-choice2026.05 | 36.71 | 31.94 | 38.45 | 28.83 | 39.94 | 33.76 | 31.18 | 34.4 | |
| Stage-1 + LoRA + GSD-as-text-promptBackbone=Qwen3-VL-8B, Adaptation Stage=Stage-1, Adaptation Strategy=LoRA + GSD-as-text-prompt, Evaluation protocol=Multiple-choice2026.05 | 34.25 | 30.18 | 36.04 | 27.3 | 39.81 | 31.55 | 29.07 | 32.6 | |
| Stage-1 + Standard LoRABackbone=Qwen3-VL-8B, Adaptation Stage=Stage-1, Adaptation Strategy=LoRA, PEFT configuration=r=64, Evaluation protocol=Multiple-choice2026.05 | 33.51 | 29.4 | 35.18 | 26.74 | 39.62 | 30.97 | 28.45 | 31.98 | |
| Claude-3.7-SonnetModel Type=Closed-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 30.68 | 24.72 | 28.15 | 23.12 | 54.46 | 31.21 | 11.18 | 29.07 | |
| Stage-1 only (RS-GPT4V SFT)Backbone=Qwen3-VL-8B, Adaptation Stage=Stage-1, Adaptation Strategy=SFT, Evaluation protocol=Multiple-choice2026.05 | 28.74 | 26.91 | 32.4 | 24.06 | 41.55 | 28.83 | 25.62 | 29.73 | |
| InternLM-XComposer-2.5-7BModel Type=Open-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 19.78 | 17.45 | 28.88 | 21.06 | 40.04 | 30.67 | 24.76 | 26.09 | |
| LLaVA-OneVision-7BModel Type=Open-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 19.26 | 33.69 | 28.72 | 24.54 | 46.4 | 37.31 | 30.62 | 31.51 | |
| InternVL3-72BModel Type=Open-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 19.19 | 33.98 | 23.39 | 20.22 | 74.56 | 31.99 | 29.46 | 33.26 | |
| Qwen3-VL-8BBackbone=Qwen3-VL-8B, Zero-shot=true, Model Type=Open-source MLLM, Evaluation protocol=Multiple-choice2026.05 | 18.42 | 21.55 | 26.31 | 19.74 | 43.18 | 22.06 | 18.93 | 24.31 | |
| Gemini-2.0Model Type=Closed-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 16.93 | 20.83 | 38.94 | 16.94 | 58.52 | 20.83 | 23.74 | 28.1 | |
| Qwen2.5-VL-7BModel Type=Open-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 9.85 | 9.25 | 18.65 | 13.95 | 17.85 | 10.94 | 6.23 | 12.39 | |
| Qwen2.5-VL-72BModel Type=Open-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 3.92 | 4.82 | 22.43 | 16.27 | 5.88 | 14.91 | 8.63 | 10.98 | |
| GPT-4oModel Type=Closed-source MLLM, Zero-shot=true, Evaluation protocol=Multiple-choice2026.05 | 0.04 | 9.64 | 12.8 | 13.35 | 37.48 | 1.97 | 2.76 | 11.15 |