Multiple-choice Question Answering on MMLU (zero-shot, test)
76Accuracy (MMLU zero-shot)SFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SFTBackbone=Qwen2.5-7B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 76 | — | |
| SFTBackbone=Qwen2.5-7B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 75 | — | |
| MIPOBackbone=Qwen2.5-7B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 75 | — | |
| RLVRBackbone=Qwen2.5-7B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 73 | — | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Few-shot=zero-shot2026.03 | 72.5 | — | |
| Moonlight-16B-A3B with training-free loop wrapperModel=Moonlight-16B-A3B, Evaluation Protocol=0-shot, Loop Setting=Loop (Ours), Backbone Architecture=MoE, Iteration Mode=layer-mode, Integration Recipe=3-stage Runge–Kutta2026.05 | 68.6 | 0.74 | |
| SFTBackbone=Qwen2.5-3B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 68.17 | — | |
| SFTBackbone=Qwen2.5-3B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 68 | — | |
| Moonlight-16B-A3BModel=Moonlight-16B-A3B, Evaluation Protocol=0-shot, Loop Setting=Base, Backbone Architecture=MoE2026.05 | 67.86 | — | |
| MIPOBackbone=Qwen2.5-3B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 67.67 | — | |
| RLVRBackbone=Qwen2.5-3B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 66.17 | — | |
| Qwen2.5-3B-InstructBackbone=Qwen2.5-3B-Instruct, Few-shot=zero-shot2026.03 | 63.5 | — | |
| MIPOBackbone=Llama-3.2-3B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 63.17 | — | |
| RLVRBackbone=Qwen2.5-1.5B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 60.5 | — | |
| RLVRBackbone=Llama-3.2-3B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 59.33 | — | |
| SFTBackbone=Llama-3.2-3B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 57.83 | — | |
| Llama-3.2-3B-InstructBackbone=Llama-3.2-3B-Instruct, Few-shot=zero-shot2026.03 | 56 | — | |
| SFTBackbone=Llama-3.2-3B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 53.83 | — | |
| MIPOBackbone=Qwen2.5-1.5B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 53.83 | — | |
| Qwen2.5-1.5B-InstructBackbone=Qwen2.5-1.5B-Instruct, Few-shot=zero-shot2026.03 | 53.5 | — | |
| SFTBackbone=Qwen2.5-1.5B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 47.33 | — | |
| SFTBackbone=Qwen2.5-1.5B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 44.5 | — | |
| RLVRBackbone=Llama-3.2-1B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 34.17 | — | |
| MIPOBackbone=Llama-3.2-1B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 32.83 | — | |
| Llama-3.2-1B-InstructBackbone=Llama-3.2-1B-Instruct, Few-shot=zero-shot2026.03 | 27.5 | — | |
| SFTBackbone=Llama-3.2-1B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 24.5 | — | |
| SFTBackbone=Llama-3.2-1B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 20.5 | — |