Change-Detection-based Visual Question Answering on QAG-360K (test)
85.85CNVisTA
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| VisTADecoding Strategy=Deterministic Greedy Decoding, Regime=Specialist Methods2025.12 | 85.85 | 63.2 | 66.41 | 84.65 | 86.14 | 62.74 | 38.21 | 71.13 | 69.79 | 74.59 | |
| CDVQADecoding Strategy=Deterministic Greedy Decoding, Regime=Specialist Methods2025.12 | 82.25 | 57.81 | 60.44 | 75.41 | 76.76 | 47.67 | 29.27 | 65.2 | 61.85 | 67.91 | |
| Qwen2.5-VL-3B-DARFTDecoding Strategy=Deterministic Greedy Decoding, Regime=Few-shot Fine-tune, Training Samples=0.5K2025.12 | 78.13 | 53.99 | 56.11 | 74.48 | 62.45 | 38.67 | 16.61 | 54.88 | 54.42 | 60.57 | |
| Qwen2.5-VL-3B-SFTDecoding Strategy=Deterministic Greedy Decoding, Regime=Few-shot Fine-tune, Training Samples=0.5K2025.12 | 77.74 | 53.04 | 52.58 | 71.25 | 57.95 | 35.2 | 19.57 | 36.5 | 50.48 | 54.52 | |
| Qwen2.5-VL-3B-SFTDecoding Strategy=Deterministic Greedy Decoding, Regime=Few-shot Fine-tune, Training Samples=2K2025.12 | 76.13 | 53.14 | 55.23 | 77.07 | 73.4 | 40.35 | 15.8 | 40.37 | 53.93 | 57.35 | |
| Qwen2.5-VL-3B-DARFTDecoding Strategy=Deterministic Greedy Decoding, Regime=Few-shot Fine-tune, Training Samples=2K2025.12 | 75.96 | 46.45 | 54.31 | 76.97 | 76.7 | 40.54 | 18.24 | 55.88 | 55.63 | 61.48 | |
| Qwen2.5-VL-3B-DARFTDecoding Strategy=Mean@20 Multi-sample Decoding, Regime=Few-shot Fine-tune, Training Samples=2K2025.12 | 75.87 | 45.73 | 53.51 | 76.62 | 76.28 | 40.01 | 18.42 | 52.09 | 54.82 | 60.26 | |
| Qwen2.5-VL-3B-DARFTDecoding Strategy=Mean@20 Multi-sample Decoding, Regime=Few-shot Fine-tune, Training Samples=0.5K2025.12 | 74.65 | 49.45 | 49.79 | 71.62 | 60.49 | 33.53 | 18.69 | 51.03 | 51.16 | 57.22 | |
| Qwen2.5-VL-3B-SFTDecoding Strategy=Mean@20 Multi-sample Decoding, Regime=Few-shot Fine-tune, Training Samples=2K2025.12 | 73.53 | 41.6 | 41.23 | 70.03 | 64.32 | 29.67 | 17.3 | 35.75 | 46.68 | 51.71 | |
| Qwen2.5-VL-3B-SFTDecoding Strategy=Mean@20 Multi-sample Decoding, Regime=Few-shot Fine-tune, Training Samples=0.5K2025.12 | 71.65 | 41.67 | 35.38 | 67.59 | 53.44 | 25.24 | 19.22 | 25.82 | 42.5 | 46.74 | |
| VisTADecoding Strategy=Deterministic Greedy Decoding, Regime=Few-shot Fine-tune, Training Samples=0.5K2025.12 | 68.55 | 42.51 | 36.94 | 66.72 | 64.44 | 28.73 | 19.13 | 47.32 | 41.59 | 52.9 | |
| VisTADecoding Strategy=Deterministic Greedy Decoding, Regime=Few-shot Fine-tune, Training Samples=2K2025.12 | 68.29 | 39.26 | 48.34 | 62.14 | 69.25 | 33.06 | 21.9 | 52.23 | 43.83 | 55.13 | |
| Qwen2.5-VL-3BDecoding Strategy=Deterministic Greedy Decoding, Regime=Zero-shot Baseline2025.12 | 51.37 | 9.93 | 4.69 | 48.22 | 47.55 | 2.74 | 14.96 | 5.14 | 23.08 | 27.44 | |
| Qwen2.5-VL-7BDecoding Strategy=Deterministic Greedy Decoding, Regime=Zero-shot Baseline2025.12 | 47.37 | 43.42 | 12.2 | 47.62 | 53.66 | 37.68 | 13.24 | 15.75 | 33.87 | 33.1 |