Temporal Grounding on VGGSync
56.6AccuracyDPO w/ OP + FV-D + LV-MCQA
Evaluation Results
| Method | Links | |
|---|---|---|
| DPO w/ OP + FV-D + LV-MCQARecipe=DPO with original-sync, video preference, and MCQA data2026.05 | 56.6 | |
| DPO w/ OP + SPRecipe=DPO with original-sync and SFT-policy negatives2026.05 | 56.4 | |
| Final 10K DPO Recipe (Ours)Recipe=Combined CTP, FV-D, and FV-A-L2026.05 | 56.4 | |
| DPO w/ CTP + FV-D + FV-ARecipe=DPO with counterfactual, video preference, and audio preference data2026.05 | 55.9 | |
| DPO w/ CTP + FV-DRecipe=DPO with counterfactual and video preference data2026.05 | 55.8 | |
| DPO w/ SPRecipe=DPO with SFT-policy negatives2026.05 | 55.7 | |
| DPO w/ CTP + FV-D + LV-MCQARecipe=DPO with counterfactual, video preference, and MCQA data2026.05 | 55.7 | |
| DPO w/ SP + FV-DRecipe=DPO with SFT-policy negatives and video preference data2026.05 | 55.4 | |
| SFT w/ CTP + FV-D + FV-ALRecipe=SFT with counterfactual, general video preference, and audio-visual data2026.05 | 46.7 | |
| Qwen3-Omni-30BRecipe=Vanilla Baseline2026.05 | 36.8 |