Coding Accuracy on MBPP+
69.3AccuracyKimi-Linear-48B-A3B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Kimi-Linear-48B-A3B-InstructCompression=30%, Compression Method=REAP2025.10 | 69.3 | |
| Kimi-Linear-48B-A3B-InstructCompression=Baseline2025.10 | 66.9 | |
| Sparsity (0.86)Model=Qwen3-1.7B Thinking, Generation length cutoff=24K, Training protocol=One epoch on 21K filtered TACO, Sparsity Setting=0.862026.06 | 60.84 | |
| STMBackbone=Llama-3.2-3B-Instruct2026.05 | 59 | |
| Dense Rollout BaselineModel=Qwen3-1.7B Thinking, Generation length cutoff=24K, Training protocol=One epoch on 21K filtered TACO, Sparsity=None2026.06 | 58.46 | |
| DFTBackbone=Llama-3.2-3B-Instruct2026.05 | 53.4 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 51.6 | |
| Iter-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 49.7 | |
| Anchored LearningBackbone=Llama-3.2-3B-Instruct2026.05 | 48.1 | |
| Low-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 46.6 | |
| KL-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 46.6 | |
| Self-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 46.6 | |
| SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 42 |