Spatial Reasoning on EmbodiedSpatial-Bench
72.94AccuracyLLaVA-OneVision-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaVA-OneVision-7BModel Type=Open-source2025.11 | 72.94 | 57.96 | |
| V2LO-7BReinforcement Learning=true2025.11 | 68.63 | 56.72 | |
| SpaceR-7BBase Model=Qwen2.5-VL-7B2025.11 | 68.09 | 53.23 | |
| V2LO-7BReinforcement Learning=false2025.11 | 67.61 | 54.31 | |
| InternVL2.5-8BModel Type=Open-source2025.11 | 66.21 | 54.06 | |
| Qwen2.5VL-7BBase Model=Qwen2.5-VL-7B2025.11 | 66.14 | 52.23 | |
| GPT-4oModel Type=Closed-source2025.11 | 65.85 | 61.86 | |
| AVLM-2BModel Scale=2B2026.06 | 62.72 | — | |
| Gemini 2.0 flashModel Type=Closed-source2025.11 | 62.12 | 61.61 | |
| Qwen2.5-Omni-3BModel Scale=3B2026.06 | 58.16 | — | |
| Gemma-4-E2B-itModel Scale=4B2026.06 | 48.32 | — | |
| L-FDMBackbone=Liquid, Evaluation Protocol=Forward-Dynamics Model, Decoding Strategy=Greedy2025.06 | 33.8 | — | |
| LiquidEvaluation Protocol=Fine-tuned, Decoding Strategy=Greedy2025.06 | 33.2 | — | |
| LiquidEvaluation Protocol=Zero-shot, Decoding Strategy=Greedy2025.06 | 32.6 | — | |
| ChameleonEvaluation Protocol=Fine-tuned, Decoding Strategy=Greedy2025.06 | 21.2 | — | |
| C-FDMBackbone=Chameleon, Evaluation Protocol=Forward-Dynamics Model, Decoding Strategy=Greedy2025.06 | 17.5 | — | |
| ChameleonEvaluation Protocol=Zero-shot, Decoding Strategy=Greedy2025.06 | 15.1 | — |