Scientific Reasoning on ARC Easy
73.4AccuracySink Attention
Evaluation Results
| Method | Links | |
|---|---|---|
| Sink AttentionModel Size=2B, Training Loss=base+aux2026.02 | 73.4 | |
| Gated AttentionModel Size=2B, Training Loss=base+aux2026.02 | 73.2 | |
| Sink AttentionModel Size=2B, Training Loss=base2026.02 | 73.06 | |
| Vanilla AttentionModel Size=2B, Training Loss=base+aux2026.02 | 72.8 | |
| Gated AttentionModel Size=2B, Training Loss=base2026.02 | 72.6 | |
| Vanilla AttentionModel Size=2B, Training Loss=base2026.02 | 72.26 | |
| ConSA (layer-wise)Target Sparsity (ρ)=0.50, Granularity=layer-wise, Model=1.7B2026.06 | 71.3 | |
| ConSA (head-wise, all-layers)Target Sparsity (ρ)=0.50, Granularity=head-wise, Constraint Scope=all-layers, Model=1.7B2026.06 | 71.21 | |
| ConSA (head-wise, single-layer)Target Sparsity (ρ)=0.50, Granularity=head-wise, Constraint Scope=single-layer, Model=1.7B2026.06 | 71 | |
| Dense FATarget Sparsity (ρ)=0, Model=1.7B2026.06 | 69.91 | |
| Rule (head-wise)Target Sparsity (ρ)=0.50, Granularity=head-wise, Model=1.7B2026.06 | 69.11 | |
| Sink AttentionModel Size=1B, Training Loss=base+aux2026.02 | 68.81 | |
| Vanilla AttentionModel Size=1B, Training Loss=base+aux2026.02 | 68.8 | |
| Sink AttentionModel Size=1B, Training Loss=base2026.02 | 68.73 | |
| Vanilla AttentionModel Size=1B, Training Loss=base2026.02 | 68.56 | |
| Gated AttentionModel Size=1B, Training Loss=base+aux2026.02 | 68.35 | |
| Rule (layer-wise)Target Sparsity (ρ)=0.50, Granularity=layer-wise, Model=1.7B2026.06 | 67.8 | |
| Gated AttentionModel Size=1B, Training Loss=base2026.02 | 67.25 | |
| Vanilla AttentionModel Size=0.6B, Training Loss=base+aux2026.02 | 65.28 | |
| Gated AttentionModel Size=0.6B, Training Loss=base+aux2026.02 | 64.86 | |
| Gated AttentionModel Size=0.6B, Training Loss=base2026.02 | 63.97 | |
| Vanilla AttentionModel Size=0.6B, Training Loss=base2026.02 | 63.8 | |
| Sink AttentionModel Size=0.6B, Training Loss=base+aux2026.02 | 63.8 | |
| Sink AttentionModel Size=0.6B, Training Loss=base2026.02 | 63.3 |