Spatio-temporal Reasoning on Spatio-temporal Reasoning Dataset 8 frames
85.6Frame F1 (F1f)Q-SFT+RL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Q-SFT+RLConfiguration=C2, Training=Supervised Fine-Tuning with Reinforcement Learning2026.04 | 85.6 | 61.1 | 68.6 | |
| GPT-4.1Version=05/01/20252026.04 | 81.3 | 23.6 | 52.8 | |
| Q-SFTConfiguration=C1, Training=Supervised Fine-Tuning2026.04 | 78.6 | 50.7 | 58.2 | |
| Q-BaselineBackbone=Qwen2.5-Coder-3B-Instruct2026.04 | 45.6 | 19.1 | 25.8 |