Mathematical Reasoning on DeepMind-Mathematics
88.4AccuracyQwen3-I(4B)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-I(4B)Model Family=Qwen3, Target Model (TL)=ℐ (4B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 88.4 | |
| Ministral-3-I(14B)Model Family=Ministral-3, Target Model (TL)=ℐ (14B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 87.2 | |
| Ministral-3-I(3B)Model Family=Ministral-3, Target Model (TL)=ℐ (3B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 84.2 | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 4B, Transfer Setting=Second Setting (Red Shading)2026.04 | 82.4 | |
| Qwen3-I(14B)Model Family=Qwen3, Target Model (TL)=ℐ (14B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 80.1 | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 4B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 79.9 | |
| Qwen3-14BModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 78.8 | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=4B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 14B, Transfer Setting=Second Setting (Red Shading)2026.04 | 76.5 | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=4B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 14B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 75.8 | |
| Qwen3-4BModel Family=Qwen3, Target Model (TL)=4B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 71.3 | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 3B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 71.3 | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 3B, Transfer Setting=Second Setting (Red Shading)2026.04 | 70.7 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=195K2025.03 | 69.1 | |
| Ministral-3-8BModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 67.4 | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=3B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 8B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 66.2 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=60K2025.03 | 65.8 | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=3B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 8B, Transfer Setting=Second Setting (Red Shading)2026.04 | 65.5 | |
| Ministral-3-3BModel Family=Ministral-3, Target Model (TL)=3B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 65.3 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=30K, Training Strategy=Sequential2025.03 | 64.6 | |
| DeepSeekMath-7B-DART-Math†Base Model=DeepSeekMath-7B, # Samples=60K2025.03 | 62.8 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=30K, Training Strategy=Parallel2025.03 | 62.2 | |
| DeepSeekMath-7B-DART-MathBase Model=DeepSeekMath-7B, # Samples=590K2025.03 | 61.6 | |
| DART-Math-DSMath-7BTraining Method=SFT, Data Selection Strategy=Prop2Diff, Backbone=DeepSeekMath-7B2024.06 | 61.6 | |
| DeepSeekMath-7B-RFTBase Model=DeepSeekMath-7B, # Samples=590K2025.03 | 60.2 | |
| DART-Math-DSMath-7BTraining Method=SFT, Data Selection Strategy=Uniform, Backbone=DeepSeekMath-7B2024.06 | 60.2 | |
| DeepSeekMath-7B-RLTraining Method=RL, Backbone=DeepSeekMath-7B2024.06 | 58.3 | |
| MathFusion-DSMath-7BBase Model=DeepSeekMath-7B, # Samples=30K, Training Strategy=Conditional2025.03 | 55.2 | |
| DeepSeekMath-7B-MMIQCBase Model=DeepSeekMath-7B, # Samples=2.3M2025.03 | 52.9 | |
| DeepSeekMath-7B-InstructBase Model=DeepSeekMath-7B, # Samples=780K2025.03 | 52.2 | |
| Llama3-8B-DART-MathBase Model=Llama3-8B, # Samples=590K2025.03 | 48 | |
| DeepSeekMath-7B-MetaMathBase Model=DeepSeekMath-7B, # Samples=60K2025.03 | 45.9 | |
| Mistral-7B-DART-MathBase Model=Mistral-7B, # Samples=590K2025.03 | 45.1 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=60K2025.03 | 43.4 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=30K, Training Strategy=Sequential2025.03 | 42 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=30K, Training Strategy=Parallel2025.03 | 41.9 | |
| Llama3-8B-RFTBase Model=Llama3-8B, # Samples=590K2025.03 | 41.7 | |
| DeepSeekMath-7B-MMIQC†Base Model=DeepSeekMath-7B, # Samples=60K2025.03 | 41.5 | |
| Llama3-8B-MMIQCBase Model=Llama3-8B, # Samples=2.3M2025.03 | 41 | |
| Llama3-8B-DART-Math†Base Model=Llama3-8B, # Samples=60K2025.03 | 39.9 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=60K2025.03 | 39.2 | |
| DeepSeekMath-7B-RefAugBase Model=DeepSeekMath-7B, # Samples=30K2025.03 | 38.4 | |
| Mistral-7B-WizardMath-V1.1Base Model=Mistral-7B, # Samples=418K2025.03 | 38.4 | |
| Mistral-7B-MMIQCBase Model=Mistral-7B, # Samples=2.3M2025.03 | 38 | |
| GREATSBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 37 | |
| IDBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 37 | |
| IWDBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 36.75 | |
| Mistral-7B-DART-Math†Base Model=Mistral-7B, # Samples=60K2025.03 | 36 | |
| COLMBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 35.7 | |
| Mistral-7B-RFTBase Model=Mistral-7B, # Samples=590K2025.03 | 35.6 | |
| DeepSeekMath-7B-RefAugBase Model=DeepSeekMath-7B, # Samples=60K2025.03 | 35.4 | |
| Llama3-8B-MetaMathBase Model=Llama3-8B, # Samples=400K2025.03 | 35 | |
| PartitionSelBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 35 | |
| GradNormBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 33.4 | |
| RandomBackbone=Qwen2.5-3B, Fine-tuning Dataset=Meta-MathQA, Subset Fraction=50%2026.06 | 32.81 | |
| Llama3-8B-MetaMath†Base Model=Llama3-8B, # Samples=60K2025.03 | 31.3 | |
| Llama3-8B-MMIQC†Base Model=Llama3-8B, # Samples=60K2025.03 | 30.9 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=30K, Training Strategy=Sequential2025.03 | 29.3 | |
| Llama3-8B-RefAugBase Model=Llama3-8B, # Samples=60K2025.03 | 29.1 | |
| DeepSeekMath-7B-StandardBase Model=DeepSeekMath-7B, # Samples=15K2025.03 | 28.6 | |
| Mistral-7B-MetaMathBase Model=Mistral-7B, # Samples=400K2025.03 | 28 | |
| MathFusion-Llama3-8BBase Model=Llama3-8B, # Samples=30K, Training Strategy=Conditional2025.03 | 27.4 | |
| Mistral-7B-MetaMath†Base Model=Mistral-7B, # Samples=60K2025.03 | 27.2 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=30K, Training Strategy=Parallel2025.03 | 26.5 | |
| Llama3-8B-RefAugBase Model=Llama3-8B, # Samples=30K2025.03 | 25.9 | |
| Llama3-8B-StandardBase Model=Llama3-8B, # Samples=15K2025.03 | 21.6 | |
| MathFusion-Mistral-7BBase Model=Mistral-7B, # Samples=30K, Training Strategy=Conditional2025.03 | 21.4 | |
| Mistral-7B-RefAugBase Model=Mistral-7B, # Samples=60K2025.03 | 18.1 | |
| Mistral-7B-StandardBase Model=Mistral-7B, # Samples=15K2025.03 | 17 | |
| COLMBase Model=Llama-3.2-3B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 16.6 | |
| COLMBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 16.6 | |
| PartitionSelBase Model=Llama-3.2-3B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 15.9 | |
| PartitionSelBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 15.9 | |
| GradNormBase Model=Llama-3.2-3B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 15.8 | |
| IDBase Model=Llama-3.2-3B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 15.8 | |
| GradNormBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 15.8 | |
| IDBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 15.8 | |
| Mistral-7B-RefAugBase Model=Mistral-7B, # Samples=30K2025.03 | 15.4 | |
| Mistral-7B-MMIQC†Base Model=Mistral-7B, # Samples=60K2025.03 | 13.5 |