Commonsense Reasoning on CommonQA (test)
84.4AccuracyBest Model
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Best ModelBackbone LLM=Qwen 2.5 7B, Shot count=0-shot2025.10 | 84.4 | 2 | -6.6 | |
| Best ModelBackbone LLM=Lucie 7B, Shot count=0-shot2025.10 | 69.7 | 3 | -9.1 | |
| Best ModelBackbone LLM=Qwen 2.5 0.5B, Shot count=0-shot2025.10 | 49.1 | 2 | 1.3 |