Reasoning on GSM8K (Accuracy)
98.8Accuracy (GSM8K)LoRA
Evaluation Results
| Method | Links | |
|---|---|---|
| LoRABackbone=Qwen2.5-7B-Instruct2026.05 | 98.8 | |
| EVA (text)Backbone=Qwen2.5-7B-Instruct2026.05 | 98.8 | |
| VanillaBackbone=Qwen2.5-7B-Instruct2026.05 | 98.7 | |
| SafeDecodingBackbone=Qwen2.5-7B-Instruct2026.05 | 98.7 | |
| LEDBackbone=Qwen2.5-7B-Instruct2026.05 | 98.7 | |
| Circuit BreakersBackbone=Qwen2.5-7B-Instruct2026.05 | 98.7 | |
| EVA (text)Backbone=Llama3.1-8B-Instruct2026.05 | 98.5 | |
| VanillaBackbone=Llama3.1-8B-Instruct2026.05 | 98.3 | |
| VanillaBackbone=Vicuna-7B-v1.52026.05 | 98.2 | |
| LoRABackbone=Llama3.1-8B-Instruct2026.05 | 98.2 | |
| EVA (text)Backbone=Vicuna-7B-v1.52026.05 | 98.1 | |
| LEDBackbone=Llama3.1-8B-Instruct2026.05 | 98.1 | |
| VanillaBackbone=Llama2-7B-chat2026.05 | 97.7 | |
| Circuit BreakersBackbone=Llama3.1-8B-Instruct2026.05 | 97.7 | |
| LoRABackbone=Vicuna-7B-v1.52026.05 | 97.6 | |
| LoRABackbone=Llama2-7B-chat2026.05 | 97.6 | |
| SafeDecodingBackbone=Llama2-7B-chat2026.05 | 97.6 | |
| LEDBackbone=Vicuna-7B-v1.52026.05 | 97.4 | |
| EVA (text)Backbone=Llama2-7B-chat2026.05 | 97.4 | |
| LEDBackbone=Llama2-7B-chat2026.05 | 97.3 | |
| Circuit BreakersBackbone=Vicuna-7B-v1.52026.05 | 97.2 | |
| Circuit BreakersBackbone=Llama2-7B-chat2026.05 | 97.1 | |
| SafeDecodingBackbone=Vicuna-7B-v1.52026.05 | 96.9 | |
| SafeDecodingBackbone=Llama3.1-8B-Instruct2026.05 | 96.9 | |
| DALAToken Cons.=6.20 × 10⁶2025.11 | 96.2 | |
| AgentPrune-RToken Cons.=7.50 × 10⁶2025.11 | 95.8 | |
| EVA (text)Backbone=Mistral-7B-Instruct-v0.22026.05 | 95.5 | |
| LoRABackbone=Mistral-7B-Instruct-v0.22026.05 | 94.6 | |
| Circuit BreakersBackbone=Mistral-7B-Instruct-v0.22026.05 | 94.3 | |
| VanillaBackbone=Mistral-7B-Instruct-v0.22026.05 | 94.1 | |
| LEDBackbone=Mistral-7B-Instruct-v0.22026.05 | 93.7 | |
| PHPToken Cons.=2.60 × 10⁷2025.11 | 92.5 | |
| LLM-DebateToken Cons.=2.20 × 10⁷2025.11 | 90.2 | |
| DyLANToken Cons.=1.40 × 10⁷2025.11 | 88.2 | |
| DeepSeek-R1-Distill-Qwen-14B (Reasoning)Model Family=Qwen2.5-14B2026.01 | 86.43 | |
| SafeDecodingBackbone=Mistral-7B-Instruct-v0.22026.05 | 86.3 | |
| ReasonAnyModel Family=Qwen2.5-14B2026.01 | 85.44 | |
| Rubric-grounded GRPOCheckpoint=Best by held-out rubric reward, Backbone=Llama-3.1-8B-Instruct2026.05 | 85.44 | |
| VanillaToken Cons.=3.50 × 10⁶2025.11 | 85.4 | |
| Llama-3.1-8B-InstructCheckpoint=Base2026.05 | 84.53 | |
| Task ArithmeticModel Family=Qwen2.5-14B2026.01 | 81.96 | |
| TIESModel Family=Qwen2.5-14B2026.01 | 79.76 | |
| LEDModel Family=Qwen2.5-14B2026.01 | 76.42 | |
| Qwen2.5-14B-Instruct (Safety)Model Family=Qwen2.5-14B2026.01 | 75.74 | |
| DAREModel Family=Qwen2.5-14B2026.01 | 64.9 | |
| Teacher (Qwen3-0.6B)Role=Teacher, Model=Qwen3-0.6B, Evaluation Protocol=generation-based evaluation2026.03 | 57.5 | |
| MobileMoE-LActive Parameters=922M, Total Parameters=5.3B, Few-shot count=8-shot2026.05 | 55.7 | |
| FuseLLMModel Family=Qwen2.5-14B2026.01 | 52.46 | |
| MobileMoE-MActive Parameters=528M, Total Parameters=2.8B, Few-shot count=8-shot2026.05 | 51.6 | |
| LinearModel Family=Qwen2.5-14B2026.01 | 50.57 | |
| MobileMoE-SActive Parameters=272M, Total Parameters=1.3B, Few-shot count=8-shot2026.05 | 36.2 | |
| Hybrid KDASequence Mixer=KDA, Attention Layers=7, Evaluation Protocol=generation-based evaluation2026.03 | 34.5 | |
| Hybrid MambaSequence Mixer=Mamba2, Attention Layers=7, Evaluation Protocol=generation-based evaluation2026.03 | 28.3 | |
| Pure KDASequence Mixer=KDA, Attention Layers=0, Evaluation Protocol=generation-based evaluation2026.03 | 22.8 | |
| Pure MambaSequence Mixer=Mamba2, Attention Layers=0, Evaluation Protocol=generation-based evaluation2026.03 | 8.3 |