Metaphorical Video Understanding on MetaphorVU-Bench
87.8Body Language ScoreHuman*
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Human*Type=Upper-bound2026.05 | 87.8 | 87.5 | 89.1 | 83.8 | 72 | 81.5 | 78.1 | 78 | 83.4 | |
| MetaphorBoostBase Model=Gemini-3-Pro2026.05 | 71.5 | 76.3 | 77.5 | 66.9 | 57.2 | 59.1 | 57.3 | 50.8 | 66.1 | |
| Gemini-3-ProCategory=Close-source MLLM2026.05 | 71.2 | 74 | 75.1 | 66.9 | 49.4 | 58.9 | 51.1 | 48.1 | 63.8 | |
| GPT-5Category=Close-source MLLM2026.05 | 69.9 | 76.3 | 77.4 | 66.6 | 45 | 55.4 | 54.9 | 46.1 | 63.7 | |
| Qwen3-VL-PlusCategory=Close-source MLLM2026.05 | 66.8 | 72.5 | 74.8 | 65.5 | 51.5 | 54.2 | 50.4 | 43.7 | 61.4 | |
| Gemini-2.5-ProCategory=Close-source MLLM2026.05 | 65.5 | 71.3 | 74.3 | 64.4 | 53.5 | 55.7 | 52.1 | 46.9 | 61.8 | |
| Qwen3-VL-235B-A22B-ThinkingCategory=Open-source MLLM2026.05 | 65.4 | 70.4 | 71.9 | 58.1 | 43.2 | 54.6 | 46.1 | 38.1 | 58.6 | |
| GPT-4oCategory=Close-source MLLM2026.05 | 63.4 | 70.5 | 70.3 | 62.6 | 39.1 | 48.2 | 45.7 | 37.9 | 56.8 | |
| GLM-4.5VCategory=Open-source MLLM2026.05 | 62.7 | 67.9 | 71.9 | 62.1 | 37.6 | 50.1 | 46.1 | 38.4 | 56.8 | |
| MetaphorBoostBase Model=Qwen3-VL-8B-Thinking2026.05 | 61.8 | 71 | 71.8 | 61.3 | 36.7 | 47.1 | 45.7 | 31.5 | 55.9 | |
| ViTCoTCategory=Reasoning-enhanced Method2026.05 | 58.8 | 47.7 | 59.2 | 48.7 | 26.1 | 45.1 | 34 | 32.1 | 46.2 | |
| Doubao-1.5-Vision-ProCategory=Close-source MLLM2026.05 | 58.2 | 64.1 | 65.5 | 58.9 | 27.8 | 42.5 | 39.8 | 26.6 | 50.5 | |
| Prompt EngineeringCategory=Reasoning-enhanced Method2026.05 | 57.8 | 66.3 | 67.9 | 59.2 | 36.1 | 42.7 | 41.6 | 32.6 | 52.4 | |
| Few-shot ExampleCategory=Reasoning-enhanced Method2026.05 | 57.6 | 69.4 | 69.2 | 58.7 | 33.5 | 44.9 | 43.5 | 32.6 | 53.6 | |
| Qwen3-VL-8B-ThinkingCategory=Open-source MLLM2026.05 | 56 | 66.1 | 68.8 | 60.8 | 33.2 | 45 | 39.3 | 29.2 | 52 | |
| LTRCategory=Reasoning-enhanced Method2026.05 | 54.1 | 44.7 | 56.2 | 47.4 | 27.8 | 44.6 | 31.9 | 36.1 | 44.5 | |
| ReAd-RCategory=Reasoning-enhanced Method2026.05 | 42.1 | 54.1 | 48.9 | 46.3 | 15.7 | 26.4 | 26.2 | 17.6 | 36.8 | |
| MetaphorBoostBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 40.7 | 55.7 | 51.2 | 49 | 12.5 | 26.1 | 31.4 | 19.2 | 37.9 | |
| Vision-R1Category=Reasoning-enhanced Method2026.05 | 39.3 | 45.1 | 42 | 42.4 | 19.4 | 23.2 | 25 | 18.6 | 33.1 | |
| VideoRFTCategory=Reasoning-enhanced Method2026.05 | 38.9 | 52.8 | 48.4 | 46 | 13.5 | 24.8 | 27.2 | 16.6 | 35.6 | |
| Qwen2.5-VL-7B-InstructCategory=Open-source MLLM2026.05 | 36 | 49.9 | 46.1 | 42.1 | 12.4 | 23.5 | 28.6 | 16.1 | 33.8 | |
| LLaVA-onevision-1.5-8B-InstructCategory=Open-source MLLM2026.05 | 35.7 | 47.2 | 47.3 | 45 | 13.8 | 21.3 | 27 | 21.2 | 38.1 |