Knowledge QA on CDDMBench
88.5QA AccuracyQwen-VL-Chat-AG* (7B)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-VL-Chat-AG* (7B)Method=SFT (Frozen encoder)2026.01 | 88.5 | |
| Gpt-5-NanoMethod=+Judge2026.01 | 84.5 | |
| Agri-CPJ (+ LLM-as-a-Judge)Model=GPT-5-Nano, Caption Generator=GPT-5-mini2026.04 | 84.5 | |
| Qwen-VL-Chat-AG (7B)Method=SFT (Unfrozen encoder)2026.01 | 84 | |
| Gpt-5-NanoMethod=Expl. Caption2026.01 | 84 | |
| Qwen2.5-VL-3B-InstructMethod=Reasoning-Enhanced GRPO2026.01 | 84 | |
| Agri-CPJ (+ Caption (Optimized))Model=GPT-5-Nano, Caption Generator=GPT-5-mini2026.04 | 84 | |
| Gpt-5-NanoMethod=+Few-shot2026.01 | 76 | |
| Agri-CPJ (+ LLM-as-a-Judge)Model=GPT-5-Nano, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 76 | |
| Agri-CPJ (+ Few-shot)Model=GPT-5-Nano, Caption Generator=GPT-5-mini2026.04 | 76 | |
| Agri-CPJ (+ Caption (Optimized))Model=GPT-5-Nano, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 75.5 | |
| Agri-CPJ (+ Few-shot)Model=GPT-5-Nano, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 74.5 | |
| Qwen2.5-VL-3B-InstructMethod=GRPO2026.01 | 72.49 | |
| Gpt-5-NanoMethod=Zero-shot2026.01 | 65 | |
| Zero-shot (Our Baseline)Model=GPT-5-Nano, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 65 | |
| Zero-shot (Our Baseline)Model=GPT-5-Nano, Caption Generator=GPT-5-mini2026.04 | 65 | |
| Qwen2.5-VL-3B-InstructMethod=SFT2026.01 | 63 | |
| Qwen-VL-Chat (7B)Method=+Judge2026.01 | 51 | |
| Agri-CPJ (+ LLM-as-a-Judge)Model=Qwen-VL-Chat, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 51 | |
| Qwen-VL-Chat (7B)Method=+Few-shot2026.01 | 50 | |
| Agri-CPJ (+ Few-shot)Model=Qwen-VL-Chat, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 50 | |
| Agri-CPJ (+ LLM-as-a-Judge)Model=Qwen-VL-Chat, Caption Generator=GPT-5-mini2026.04 | 49.5 | |
| Agri-CPJ (+ Few-shot)Model=Qwen-VL-Chat, Caption Generator=GPT-5-mini2026.04 | 49 | |
| Qwen-VL-Chat (7B)Method=Expl. Caption2026.01 | 46.5 | |
| Agri-CPJ (+ Caption (Optimized))Model=Qwen-VL-Chat, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 46.5 | |
| Qwen2.5-VL-3B-InstructMethod=Few-shot2026.01 | 45.5 | |
| Agri-CPJ (+ Caption (Optimized))Model=Qwen-VL-Chat, Caption Generator=GPT-5-mini2026.04 | 44 | |
| Zero-shot (Our Baseline)Model=Qwen-VL-Chat, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 41.5 | |
| Qwen-VL-Chat (7B)Method=Zero-shot2026.01 | 41 | |
| Zero-shot (Liu et al., 2024) BaselineModel=Qwen-VL-Chat, Caption Generator=Qwen2.5-VL-72B-Instruct2026.04 | 41 | |
| Qwen2.5-VL-3B-InstructMethod=Zero-shot2026.01 | 27.5 |