Spatio-temporal Reasoning on Spatio-temporal Reasoning Dataset 4 frames
88.8Frame F1 (F1f)GPT-4.1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4.1Version=05/01/20252026.04 | 88.8 | 26.9 | 64 | |
| Q-SFT+RLConfiguration=C2, Training=Supervised Fine-Tuning with Reinforcement Learning2026.04 | 88.1 | 67.6 | 69.4 | |
| Q-SFTConfiguration=C1, Training=Supervised Fine-Tuning2026.04 | 83.6 | 63.6 | 64.4 | |
| Q-BaselineBackbone=Qwen2.5-Coder-3B-Instruct2026.04 | 47.2 | 21.1 | 32.5 |