Time Series QA Reasoning on TSCognition In-Distribution
83.1Decoding AccuracyTSAlign-7B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TSAlign-7BTypes=Ours, Evaluation Protocol=Full-shot, Backbone LLM=Qwen2.5-Instruct-7B2026.06 | 83.1 | 85.4 | 91.7 | 77.9 | 90.4 | |
| TSAlign-3BTypes=Ours, Evaluation Protocol=Full-shot, Backbone LLM=Qwen2.5-Instruct-3B2026.06 | 74.9 | 72.7 | 65.8 | 74.5 | 82.4 | |
| GPT-5.1Types=Proprietary API, Evaluation Protocol=Zero-shot2026.06 | 71.4 | 74.6 | 68.3 | 57.4 | 71.5 | |
| Qwen2.5-Instruct-7BTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 51.7 | 44 | 45.2 | 39.8 | 41.5 | |
| Qwen2.5-VL-7BTypes=Vision-Text, Evaluation Protocol=Full-shot2026.06 | 48.5 | 40.2 | 39.5 | 38.1 | 38.2 | |
| Time-MQATypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 45.5 | 42.5 | 36.5 | 34.8 | 36.8 | |
| Qwen2.5-Instruct-32BTypes=TS-Text, Evaluation Protocol=Zero-shot2026.06 | 43.2 | 38.3 | 40.7 | 35.2 | 37.2 | |
| ITFormerTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 42.4 | 47.4 | 42.1 | 46.5 | 40.8 | |
| ChatTSTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 39.1 | 36.8 | 37.4 | 32 | 28.5 | |
| Qwen2.5-VL-7BTypes=Vision-Text, Evaluation Protocol=Zero-shot2026.06 | 37.5 | 34.1 | 30.7 | 33 | 26.9 | |
| Time-LLMTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 36 | 32.5 | 34.1 | 35.4 | 31.6 | |
| GPT4TSTypes=TS-Text, Evaluation Protocol=Full-shot2026.06 | 34.8 | 33.2 | 32.4 | 31.2 | 35.6 | |
| Qwen2.5-Instruct-7BTypes=TS-Text, Evaluation Protocol=Zero-shot2026.06 | 34.6 | 29.4 | 36.5 | 28.3 | 32.6 | |
| Random GuessingEvaluation Protocol=Zero-shot2026.06 | 22.1 | 22.4 | 24.2 | 23.5 | 24.1 |