Zero-shot Reasoning on Commonsense Reasoning Evaluation Suite
32.2Accuracy on OBQALLR
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| LLRModel Architecture=LLaMa-1B, Training Tokens=100B, Evaluation Protocol=zero-shot2026.05 | 32.2 | 60.38 | 38.4 | 72.35 | 45.24 | 41.8 | 73.01 | 51.91 | |
| UniformModel Architecture=LLaMa-1B, Training Tokens=100B, Evaluation Protocol=zero-shot2026.05 | 30.8 | 57.93 | 37.37 | 73.4 | 44.48 | 41.71 | 72.42 | 51.16 | |
| LLRModel Architecture=LLaMa-3B, Training Tokens=30B, Evaluation Protocol=zero-shot2026.05 | 29.4 | 59.04 | 34.39 | 72.43 | 43.81 | 42.68 | 72.52 | 50.61 | |
| UniformModel Architecture=LLaMa-3B, Training Tokens=30B, Evaluation Protocol=zero-shot2026.05 | 27.4 | 55.56 | 33.36 | 70.16 | 42.31 | 39.92 | 71.33 | 48.58 |