Long-form Dialogue Decomposition on MLDR (test)
81.2Pairwise PrecisionOur model-3B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Our model-3BBackbone=Qwen2.5-VL-3B, Parameters=3B, Tuning=LoRA2026.06 | 81.2 | 92.2 | 83.6 | 83.8 | 94.2 | 86.7 | |
| GPT-4oModel=GPT-4o2026.06 | 77.5 | 84.1 | 78.6 | 80.2 | 92.5 | 83.9 | |
| Qwen2.5-VL-72BBackbone=Qwen2.5-VL, Parameters=72B2026.06 | 58.4 | 66.9 | 60.3 | 60.5 | 68.1 | 61.1 |