Spatio-temporal Reasoning on Spatio-temporal Reasoning Dataset Overall
87.5Frame F1 (F1f)Q-SFT+RL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Q-SFT+RLConfiguration=C2, Training=Supervised Fine-Tuning with Reinforcement Learning2026.04 | 87.5 | 64.5 | 72.3 | |
| GPT-4.1Version=05/01/20252026.04 | 84.8 | 35 | 61 | |
| Q-SFTConfiguration=C1, Training=Supervised Fine-Tuning2026.04 | 80.4 | 56.6 | 63.6 | |
| Q-BaselineBackbone=Qwen2.5-Coder-3B-Instruct2026.04 | 48.5 | 25 | 32.6 |