Time Series QA Reasoning on TSCognition Out-of-Distribution
87.6Decoding AccuracyTSAlign-7B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TSAlign-7BTypes=Ours, Evaluation Protocol=Full-shot, Backbone LLM=Qwen2.5-Instruct-7B2026.06 | 87.6 | 87.4 | 88.7 | 84.6 | 95.3 | |
| TSAlign-3BTypes=Ours, Evaluation Protocol=Full-shot, Backbone LLM=Qwen2.5-Instruct-3B2026.06 | 76.5 | 76.6 | 74.3 | 81.3 | 86 | |
| GPT-5.1Types=Proprietary API, Evaluation Protocol=Zero-shot2026.06 | 68.5 | 70.1 | 72.5 | 63.9 | 77 | |
| Qwen2.5-Instruct-7BTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 62.3 | 48.6 | 41.7 | 36.5 | 45.6 | |
| Qwen2.5-VL-7BTypes=Vision-Text, Evaluation Protocol=Full-shot2026.06 | 52.6 | 46.7 | 35.4 | 41.5 | 32.5 | |
| ITFormerTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 46.8 | 50.6 | 43.6 | 42.7 | 42.3 | |
| Time-MQATypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 44.1 | 41.4 | 38.4 | 31.4 | 32.7 | |
| ChatTSTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 43.5 | 42.3 | 30.8 | 36.8 | 30.6 | |
| Qwen2.5-Instruct-32BTypes=TS-Text, Evaluation Protocol=Zero-shot2026.06 | 43.2 | 39.5 | 36.8 | 33.6 | 34.9 | |
| Qwen2.5-VL-7BTypes=Vision-Text, Evaluation Protocol=Zero-shot2026.06 | 41.2 | 34.7 | 28.4 | 34.8 | 25.3 | |
| Time-LLMTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 38.2 | 36.1 | 32.8 | 39.3 | 33.4 | |
| GPT4TSTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 37.9 | 38.4 | 29.7 | 33.5 | 38.1 | |
| Qwen2.5-Instruct-7BTypes=TS-Text, Evaluation Protocol=Zero-shot2026.06 | 37.1 | 32.7 | 33.3 | 30.2 | 29.7 | |
| Random GuessingEvaluation Protocol=Zero-shot2026.06 | 23.5 | 21.9 | 21.8 | 22.7 | 23.9 |