Scenario-based Reasoning (Overall) on TSRBench
86.55Overall AccuracyVeriTime
Evaluation Results
| Method | Links | |
|---|---|---|
| VeriTimeBase Model=Qwen3-4B-Instruct, Training Protocol=SFT+RL2026.02 | 86.55 | |
| VeriTimeBase Model=Qwen2.5-3B-Instruct, Training Protocol=SFT+RL2026.02 | 82.86 | |
| ChatTSBase Model=Qwen3-4B-Instruct, Training Protocol=SFT2026.02 | 82.21 | |
| ChatTSBase Model=Qwen2.5-3B-Instruct, Training Protocol=SFT2026.02 | 78.31 | |
| Qwen3-4B-InstructTraining Protocol=Base2026.02 | 75.48 | |
| GPT-4o-miniModel Type=General LLM2026.02 | 70.43 | |
| Qwen2.5-7B-instructModel Type=General LLM2026.02 | 66.81 | |
| Meta-Llama3-8B-InstructModel Type=General LLM2026.02 | 59.22 | |
| Mistral-7B-v0.3Model Type=General LLM2026.02 | 59.22 | |
| Time-MQABase Model=Mistral-7B2026.02 | 53.66 | |
| Time-MQABase Model=Llama3-8B2026.02 | 53.05 | |
| DeepSeek-R1-Distill-Qwen-7BModel Type=General LLM2026.02 | 52.93 | |
| Time-R1Base Model=Qwen2.5-7B2026.02 | 51.73 | |
| Time-MQABase Model=Qwen2.5-7B2026.02 | 41 | |
| Qwen2.5-3B-InstructTraining Protocol=Base2026.02 | 40.99 |