Scenario Understanding on Scenario Understanding (ID)
90.7AccuracyTIMEOMNI-1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TIMEOMNI-1Base LLM=Qwen2.5-Instruct-7B2025.09 | 90.7 | 97.5 | |
| GPT-4.1-2025-04-142025.09 | 85.5 | 100 | |
| GPT-4.1-Nano2025.09 | 66.2 | 97.5 | |
| Mistral-Small-3.1-24B-Ins2025.09 | 64.8 | 100 | |
| Llama-3.1-70B-Instruct2025.09 | 56.4 | 100 | |
| Qwen2.5-Instruct-7B2025.09 | 48.5 | 100 | |
| Mistral-7B-v0.32025.09 | 40.5 | 92.2 | |
| Llama-3.1-8B-Instruct2025.09 | 36.6 | 46.5 | |
| Time-MQABase LLM=Llama3-8B2025.09 | 32.2 | 29.5 | |
| Time-R1Base LLM=Qwen2.5-Instruct-7B2025.09 | 30.9 | 94 | |
| Time-MQABase LLM=Qwen2.5-7B2025.09 | 25 | 14 | |
| Time-MQABase LLM=Mistral-7B-v0.32025.09 | 15.1 | 21.5 |