Spatio-temporal Reasoning on Spatio-temporal Reasoning Dataset 16 frames
82.7Frame F1 (F1f)Q-SFT+RL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Q-SFT+RLConfiguration=C2, Training=Supervised Fine-Tuning with Reinforcement Learning2026.04 | 82.7 | 44 | 62.9 | |
| GPT-4.1Version=05/01/20252026.04 | 75.6 | 23.1 | 44.7 | |
| Q-SFTConfiguration=C1, Training=Supervised Fine-Tuning2026.04 | 65.7 | 32.9 | 45.4 | |
| Q-BaselineBackbone=Qwen2.5-Coder-3B-Instruct2026.04 | 41.6 | 18.7 | 21.5 |