Failure Reasoning and Correction on Real-World Benchmark (test)
62.1ROUGE-LDream2Fix-VLM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Dream2Fix-VLMmode=Zero-Shot2026.03 | 62.1 | 66.8 | 82 | 42.1 | 47.2 | |
| Gemini-1.5-Flashmode=Zero-Shot2026.03 | 46.7 | 58.9 | 98 | 25 | 37.4 | |
| GPT-4omode=Zero-Shot2026.03 | 19 | 47.3 | 72 | 22.1 | 12.6 | |
| LLaVA-NeXT-34Bmode=Zero-Shot2026.03 | 9 | 9 | 30 | 12.8 | 2.2 | |
| Qwen2-VL-72Bmode=Zero-Shot2026.03 | 6.1 | 47.8 | 93 | 16.7 | 18.3 | |
| Qwen2.5-VL-7Bmode=Zero-Shot2026.03 | 5.2 | 17.6 | 25 | 10.5 | 0.8 | |
| LLaVA-NeXT-7Bmode=Zero-Shot2026.03 | 0 | 0 | 0 | 35.4 | 0 |