Scenario Attribution on TSRBench
87.5AccuracyVeriTime
Evaluation Results
| Method | Links | |
|---|---|---|
| VeriTimeBase Model=Qwen3-4B-Instruct, Training Protocol=SFT+RL2026.02 | 87.5 | |
| VeriTimeBase Model=Qwen2.5-3B-Instruct, Training Protocol=SFT+RL2026.02 | 83.52 | |
| Qwen3-4B-InstructTraining Protocol=Base2026.02 | 80.68 | |
| ChatTSBase Model=Qwen3-4B-Instruct, Training Protocol=SFT2026.02 | 80.68 | |
| ChatTSBase Model=Qwen2.5-3B-Instruct, Training Protocol=SFT2026.02 | 78.98 | |
| Qwen2.5-7B-instructModel Type=General LLM2026.02 | 69.89 | |
| Meta-Llama3-8B-InstructModel Type=General LLM2026.02 | 64.77 | |
| Mistral-7B-v0.3Model Type=General LLM2026.02 | 61.36 | |
| GPT-4o-miniModel Type=General LLM2026.02 | 61.36 | |
| Qwen2.5-3B-InstructTraining Protocol=Base2026.02 | 61.36 | |
| DeepSeek-R1-Distill-Qwen-7BModel Type=General LLM2026.02 | 59.66 | |
| Time-MQABase Model=Qwen2.5-7B2026.02 | 56.25 | |
| Time-R1Base Model=Qwen2.5-7B2026.02 | 53.14 | |
| Time-MQABase Model=Llama3-8B2026.02 | 50.77 | |
| Time-MQABase Model=Mistral-7B2026.02 | 49.37 |