Report Generation on DTU (test)
41BLEU-4Eyes + Bridge + QLoRA + RAFT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Eyes + Bridge + QLoRA + RAFTVisual grounding=Text + grid + OBB + retrieval, Detector=YOLO26-x-obb, Model Bridge=Included, Fine-tuning=QLoRA, Retrieval-Augmentation=RAFT2026.05 | 41 | 4 | 8.6 | |
| Eyes + Bridge + QLoRAVisual grounding=Text + grid + OBB, Detector=YOLO26-x-obb, Model Bridge=Included, Fine-tuning=QLoRA2026.05 | 36 | 18 | 7.4 | |
| Eyes + Bridge + DeepSeek-V3Visual grounding=Text + grid + OBB, Detector=YOLO26-x-obb, Model Bridge=Included2026.05 | 19 | 29 | 5.9 | |
| Eyes + Bridge + Qwen baseVisual grounding=Text + grid + OBB, Detector=YOLO26-x-obb, Model Bridge=Included2026.05 | 14 | 38 | 4.6 | |
| Eyes + LLM, no BridgeVisual grounding=Class + confidence only, Detector=YOLO26-x-obb2026.05 | 12 | 49 | 4.7 | |
| Prompt CoT (DeepSeek-V3, no Bridge)Visual grounding=Class + confidence only2026.05 | 9 | 61 | 3.8 | |
| Zero-shot VLM (GPT-4V)Visual grounding=Full image, 336^2 px2026.05 | 7 | 65 | 3.3 |