Long-horizon reasoning on EXPLORE-Bench Short
66.12Object ScoreGLEN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GLEN2026.07 | 66.12 | 2.67 | |
| Qwen3-VL-8B-Thinking2026.07 | 63.77 | 2.85 | |
| Qwen3-VL-8B-Instruct2026.07 | 61.34 | 2.84 | |
| Gemini-3-Pro2026.07 | 61.29 | 2.77 | |
| GPT-5.2-Chat2026.07 | 59.91 | 2.7 | |
| LLaVA-OneVision-1.5-8B2026.07 | 53.25 | 2.51 |