Reasoning on DROP
89.27ScoreInfiFusion*
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| InfiFusion*Model Size=14B, GPU Hours=160, Subset Source Models=true2025.05 | 89.27 | — | — | — | |
| FuseChat*Model Size=14B, GPU Hours=650, Subset Source Models=true2025.05 | 89.23 | — | — | — | |
| InfiFPOModel Size=14B, GPU Hours=582025.05 | 88.83 | — | — | — | |
| FuseLLM*Model Size=14B, GPU Hours=225, Subset Source Models=true2025.05 | 88.74 | — | — | — | |
| SFTModel Size=14B, GPU Hours=152025.05 | 88.72 | — | — | — | |
| InfiFPO*Model Size=14B, GPU Hours=55, Subset Source Models=true2025.05 | 88.68 | — | — | — | |
| Phi-4Model Size=14B, GPU Hours=∼1.0M2025.05 | 88.67 | — | — | — | |
| SFT-IPOModel Size=14B, GPU Hours=502025.05 | 88.67 | — | — | — | |
| SFT-DPOModel Size=14B, GPU Hours=502025.05 | 88.56 | — | — | — | |
| SFT-WRPOModel Size=14B, GPU Hours=572025.05 | 88.41 | — | — | — | |
| Ling-flash-2.02026.02 | 88.32 | 100 | — | — | |
| LLaDA2.0-flash2026.02 | 87.9 | 226 | — | — | |
| LLaDA2.1-flashInference Mode=Q Mode2026.02 | 87.86 | 253 | — | — | |
| Qwen3-30B-A3B-Inst-25072026.02 | 87.57 | 100 | — | — | |
| LLaDA2.1-flashInference Mode=S Mode2026.02 | 87.55 | 540 | — | — | |
| CapFlowType=Learning, Setting=Multi-domain, Passes=12026.02 | 87.12 | — | — | — | |
| Mistral-SmallModel Size=24B, GPU Hours=∼1.6M2025.05 | 86.52 | — | — | — | |
| Gemma-3-InstructModel Size=12B2025.05 | 86.43 | — | — | — | |
| CapFlowType=Learning, Setting=Cross-domain, Passes=1, Cross-domain Domain Mapping=Code,Math->Reason2026.02 | 86.25 | — | — | — | |
| ScoreFlowType=Learning, Setting=Single-domain, Iterations=52026.02 | 86.14 | — | — | — | |
| AFlowType=Refinement, Setting=Single-domain, Iterations=202026.02 | 85.75 | — | — | — | |
| Qwen2.5-InstructModel Size=14B, GPU Hours=∼1.8M2025.05 | 85.56 | — | — | — | |
| ScoreFlowType=Learning, Setting=Multi-domain, Iterations=52026.02 | 85.16 | — | — | — | |
| Qwen3-8Bno think=true2026.02 | 84.56 | — | — | — | |
| Qwen2.5-CoderModel Size=14B, GPU Hours=∼1.8M2025.05 | 84.34 | — | — | — | |
| CoT-SCType=Manual, Setting=Manual2026.02 | 83.28 | — | — | — | |
| CoTType=Manual, Setting=Manual2026.02 | 83.25 | — | — | — | |
| Self-RefineType=Manual, Setting=Manual2026.02 | 82.5 | — | — | — | |
| LLaDA2.1-minimode=Q Mode2026.02 | 82.37 | 287 | — | — | |
| ADASType=Refinement, Setting=Single-domain, Iterations=202026.02 | 82.23 | — | — | — | |
| Qwen2.5-MathModel Size=7B, GPU Hours=∼0.5M2025.05 | 81.96 | — | — | — | |
| LLaDA2.0-mini2026.02 | 81.89 | 202 | — | — | |
| SPPType=Manual, Setting=Manual2026.02 | 81.62 | — | — | — | |
| LLaDA2.1-minimode=S Mode2026.02 | 81.55 | 584 | — | — | |
| GPT-4o-miniType=Manual, Setting=Manual2026.02 | 81.37 | — | — | — | |
| Ling-mini-2.02026.02 | 78.8 | — | — | — | |
| DCLMPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 24.31 | — | — | — | |
| Nemotron-CCPre-training Curation Strategy=AI Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 19.57 | — | — | — | |
| Nemotron-CC_ASIPre-training Curation Strategy=AI Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 19.48 | — | — | — | |
| Nemotron-CC_ASI+Pre-training Curation Strategy=AI Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 18.49 | — | — | — | |
| Ultra-FinewebPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 7.78 | — | — | — | |
| Fineweb-EduPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 6.78 | — | — | — | |
| DeepSeek-V3.2 Base# Shots=1-shot, Architecture=MoE, # Activated Params=37B, # Total Params=671B2026.04 | — | — | — | 88.2 | |
| DeepSeek-V4-Flash Base# Shots=1-shot, Architecture=MoE, # Activated Params=13B, # Total Params=284B2026.04 | — | — | — | 88.6 | |
| DeepSeek-V4-Pro Base# Shots=1-shot, Architecture=MoE, # Activated Params=49B, # Total Params=1.6T2026.04 | — | — | — | 88.7 | |
| Original (T2T)Backbone=LLaDA2.1-mini, Inference Strategy=Text-to-Text2026.04 | — | — | — | 38.53 | |
| T2MBackbone=LLaDA2.1-mini, Inference Strategy=Token-to-Mask, Remasking Strategy=LOWPROB, τ=0.3, Cmax=1, ρmax=0.252026.04 | — | — | — | 38.9 | |
| TransformerModel Scale=4B2026.06 | — | — | — | 36.6 | |
| YOCO (CLSA)Model Scale=4B, Attention=Cross-Layer Sparse2026.06 | — | — | — | 39.1 | |
| YOCO (Dense)Model Scale=4B2026.06 | — | — | — | 38.7 |