Commonsense Reasoning on HellaSwag, PIQA, OBQA, COPA, and WinoGrande
32.8HellaSwag AccuracyD3
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| D3Granularity=Ours, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 32.8 | 68.9 | 33.1 | 71.2 | 53.4 | 51.88 | |
| Dynamic LossGranularity=Sample-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 31.8 | 69.2 | 29.6 | 65.5 | 52.5 | 49.72 | |
| DoGEGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 31.7 | 68.9 | 29.2 | 64 | 50.6 | 48.88 | |
| DoReMiGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 31.6 | 68.7 | 29.8 | 66 | 51.4 | 49.5 | |
| RegMixGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 31.4 | 68.8 | 29.4 | 66.5 | 53.6 | 49.94 | |
| Data Mixing LawGranularity=Domain-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 31.3 | 68.4 | 30.2 | 65.9 | 51.3 | 49.42 | |
| Uniform SamplingGranularity=Sample-level, Architecture=LlaMa-based (1.1B), Training Tokens=100B, Evaluation Protocol=few-shot2026.05 | 31.2 | 68.5 | 28.4 | 64 | 51.1 | 48.64 |