Commonsense Reasoning on HellaSwag 10-shot (test)
82.53AccuracyUniform Regularized Looping
Evaluation Results
| Method | Links | |
|---|---|---|
| Uniform Regularized LoopingModel=Gemma, Size=9B, Shots=10-shot2026.02 | 82.53 | |
| Moving Average Regularized LoopingModel=Gemma, Size=9B, Shots=10-shot2026.02 | 82.42 | |
| Auto-Align Regularized LoopingModel=Gemma, Size=9B, Shots=10-shot2026.02 | 82.42 | |
| Noise AblationModel=Gemma, Size=9B, Shots=10-shot2026.02 | 82.42 | |
| BaselineModel=Gemma, Size=9B, Shots=10-shot2026.02 | 82.41 | |
| Moving Average Regularized LoopingModel=Llama, Size=8B, Shots=10-shot2026.02 | 82.31 | |
| Auto-Align Regularized LoopingModel=Llama, Size=8B, Shots=10-shot2026.02 | 82.31 | |
| Uniform Regularized LoopingModel=Llama, Size=8B, Shots=10-shot2026.02 | 82.3 | |
| BaselineModel=Llama, Size=8B, Shots=10-shot2026.02 | 82.14 | |
| Noise AblationModel=Llama, Size=8B, Shots=10-shot2026.02 | 82.14 | |
| Qwen2.5-7BModel Type=Dense, Active Parameters=7B, Total Parameters=7B, Shots=102026.02 | 80.2 | |
| OLMoE-1B-7BModel Type=MoE, Active Parameters=1B, Total Parameters=7B, Shots=102026.02 | 79.6 | |
| OLMo-7BModel Type=Dense, Active Parameters=7B, Total Parameters=7B, Shots=102026.02 | 77.1 | |
| Llama-3.2-3BModel Type=Dense, Active Parameters=3B, Total Parameters=3B, Shots=102026.02 | 76.4 | |
| Uniform Regularized LoopingModel=Gemma, Size=2B, Shots=10-shot2026.02 | 74.7 | |
| Moving Average Regularized LoopingModel=Gemma, Size=2B, Shots=10-shot2026.02 | 74.64 | |
| Auto-Align Regularized LoopingModel=Gemma, Size=2B, Shots=10-shot2026.02 | 74.6 | |
| Qwen2.5-3BModel Type=Dense, Active Parameters=3B, Total Parameters=3B, Shots=102026.02 | 74.6 | |
| BaselineModel=Gemma, Size=2B, Shots=10-shot2026.02 | 74.51 | |
| Noise AblationModel=Gemma, Size=2B, Shots=10-shot2026.02 | 74.5 | |
| ExpertWeaver-E64-A14-S2(Qwen2.5-7B)Model Type=MoE, Active Parameters=3.5B, Total Parameters=7B, Training Budget=200B tokens, Shots=102026.02 | 73.7 | |
| LLaMA-MoE-v1-3.5BModel Type=MoE, Active Parameters=3.5B, Shots=102026.02 | 73.3 | |
| SmolLM2-1.7BModel Type=Dense, Active Parameters=1.7B, Total Parameters=1.7B, Shots=102026.02 | 72.6 | |
| Open-LLaMA-3B-v2Model Type=Dense, Active Parameters=3B, Total Parameters=3B, Shots=102026.02 | 71.4 | |
| Sheared-LLaMA-2.7BModel Type=Dense, Active Parameters=2.7B, Total Parameters=2.7B, Shots=102026.02 | 71 | |
| OLMoE-1B-7B*Model Type=MoE, Active Parameters=1B, Total Parameters=7B, Training Budget=500B tokens, Shots=102026.02 | 70.2 | |
| Gemma-2-2bModel Type=Dense, Active Parameters=2B, Total Parameters=2B, Shots=102026.02 | 69 | |
| Qwen2.5-1.5BModel Type=Dense, Active Parameters=1.5B, Total Parameters=1.5B, Shots=102026.02 | 68 | |
| INCITE-Base-3BModel Type=Dense, Active Parameters=3B, Total Parameters=3B, Shots=102026.02 | 64.7 | |
| OPT-2.7BModel Type=Dense, Active Parameters=2.7B, Total Parameters=2.7B, Shots=102026.02 | 61.4 | |
| ExpertWeaver-E64-A14-S2(OLMo-7B)Model Type=MoE, Active Parameters=1B, Total Parameters=7B, Training Budget=200B tokens, Shots=102026.02 | 61.2 | |
| Pythia-2.8BModel Type=Dense, Active Parameters=2.8B, Total Parameters=2.8B, Shots=102026.02 | 60.7 | |
| OpenMoE-3B-9BModel Type=MoE, Active Parameters=3B, Total Parameters=9B, Shots=102026.02 | 56.5 | |
| LLaMA-MoE-v2-3.5BModel Type=MoE, Active Parameters=3.5B, Shots=102026.02 | 53.7 |