Spatio-temporal Reasoning on Spatio-temporal Reasoning Dataset (12 frames)
85.2Frame F1 (F1f)Q-SFT+RL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Q-SFT+RLConfiguration=C2, Training=Supervised Fine-Tuning with Reinforcement Learning2026.04 | 85.2 | 54.9 | 65.9 | |
| GPT-4.1Version=05/01/20252026.04 | 79.7 | 23.6 | 49.7 | |
| Q-SFTConfiguration=C1, Training=Supervised Fine-Tuning2026.04 | 77.1 | 45.1 | 56.5 | |
| Q-BaselineBackbone=Qwen2.5-Coder-3B-Instruct2026.04 | 45.2 | 18.7 | 23.1 |