Zero-shot Text Evaluation on ARC-e, ARC-c, BoolQ, HSwag, OBQA, PIQA, WinoGrande, MMLU
69Average Accuracypre-trained
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| pre-trainedModel=1.7B2026.04 | 69 | 0 | |
| Depth Up-scalingModel=1.7B, Position=TOP, Trainable=0.66B2026.04 | 66.4 | -2.7 | |
| Depth Up-scalingModel=1.7B, Position=SANDWICH, Trainable=0.66B2026.04 | 61 | -8 | |
| Depth Up-scalingModel=1.7B, Position=INTERLEAVED, Trainable=0.66B2026.04 | 60.8 | -8.3 | |
| Depth Up-scalingModel=1.7B, Position=MIDDLE, Trainable=0.66B2026.04 | 59.4 | -9.6 | |
| Depth Up-scalingModel=1.7B, Position=BOTTOM, Trainable=0.66B2026.04 | 58.5 | -10.5 | |
| pre-trainedModel=360M2026.04 | 56.8 | 0 | |
| Depth Up-scalingModel=360M, Position=TOP, Trainable=200M2026.04 | 55.9 | -0.9 | |
| Depth Up-scalingModel=360M, Position=INTERLEAVED, Trainable=200M2026.04 | 54.9 | -1.9 | |
| Depth Up-scalingModel=360M, Position=SANDWICH, Trainable=200M2026.04 | 53.1 | -3.8 | |
| Depth Up-scalingModel=360M, Position=MIDDLE, Trainable=200M2026.04 | 52.9 | -3.9 | |
| Depth Up-scalingModel=360M, Position=BOTTOM, Trainable=200M2026.04 | 51 | -5.8 | |
| Full fine-tuningModel=1.7B, Trainable=1.87B2026.04 | 36.5 | -32.6 | |
| LoRAModel=360M, Trainable=200M2026.04 | 34.9 | -22 | |
| Full fine-tuningModel=360M, Trainable=360M2026.04 | 34.6 | -22.2 | |
| LoRAModel=1.7B, Trainable=0.66B2026.04 | 33.4 | -35.7 |