Commonsense Reasoning on SciQ zero-shot
83.9Accuracy (SciQ zero-shot)Original Endpoint
Evaluation Results
| Method | Links | |
|---|---|---|
| Original EndpointModel=1B, Zero-shot evaluation=true2026.06 | 83.9 | |
| Original PythiaModel=410M, Zero-shot evaluation=true2026.06 | 81.5 | |
| Learned MatchingModel=1B, Zero-shot evaluation=true2026.06 | 64.8 | |
| Dual Learned MatchingModel=1B, Zero-shot evaluation=true2026.06 | 64.4 | |
| Dual Learned MatchingModel=410M, Zero-shot evaluation=true2026.06 | 60 | |
| Learned MatchingModel=410M, Zero-shot evaluation=true2026.06 | 56.3 | |
| Raw InterpolationModel=1B, Zero-shot evaluation=true2026.06 | 35.5 | |
| Raw InterpolationModel=410M, Zero-shot evaluation=true2026.06 | 31.5 |