Video Reasoning on SEED-Bench-R1 (L1 In-Dist.)
50.5AccuracyAPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| APPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 50.5 | |
| DAPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 50 | |
| GRPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 49 | |
| SFTBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 40.2 | |
| APPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 37.5 | |
| DAPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 36.6 | |
| APPOSize=7B, Training Data=34K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 35.4 | |
| GRPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 35.3 | |
| VideoChat-R1Size=7B, Training Data=18K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 33.3 | |
| SFTBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 32.6 | |
| VideoRFTSize=7B, Training Data=310K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 32.4 | |
| Video-R1Size=7B, Training Data=260K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 30.9 | |
| TW-GRPOSize=7B, Training Data=1K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 30.2 | |
| GRPO-CARESize=7B, Training Data=260K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 29.9 | |
| Base ModelBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 29.1 | |
| Base ModelBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 28.2 |