Zero-shot Commonsense Reasoning on ARC-Easy, ARC-Challenge, SIQA, PIQA, and WinoGrande
66.1Reasoning AccuracyLLAMA-2
Evaluation Results
| Method | Links | |
|---|---|---|
| LLAMA-2Model Size=13B2024.03 | 66.1 | |
| BTXActive Experts (Top-k)=1, Sampling Strategy=Sample2024.03 | 63.7 | |
| BTXActive Experts (Top-k)=22024.03 | 63.5 | |
| LLAMA-2Model Size=7B2024.03 | 63.3 | |
| DenseTraining Context=Data-Matching (DM)2024.03 | 63.3 | |
| Sparse upcyclingTraining Context=Data-Matching (DM), Active Experts (Top-k)=22024.03 | 62.3 | |
| BTMActive Experts (Top-k)=22024.03 | 61.2 | |
| BTMActive Experts (Top-k)=12024.03 | 61 | |
| CODELLAMAModel Size=7B2024.03 | 56.6 | |
| LLEMMAModel Size=7B2024.03 | 38.8 |