Reasoning on BBH
81.1ScoreMiniCPM-o 4.5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MiniCPM-o 4.52026.04 | 81.1 | — | |
| Ouro 2.6B R4Architecture=LoopLM, # Total Params=2.6B, # Trained Tokens=7.7T2025.10 | 80.46 | — | |
| Qwen3.5-4BParameters=4B2026.05 | 80.2 | — | |
| Gemma3Architecture=Dense, # Total Params=12.0B, # Trained Tokens=12T2025.10 | 78.41 | — | |
| Qwen3Architecture=Dense, # Total Params=8.0B, # Trained Tokens=36T2025.10 | 77.65 | — | |
| Llama-3Size=70B, Evaluation Mode=3-shot2024.09 | 77.62 | — | |
| CPTSize=70B, Evaluation Mode=3-shot2024.09 | 77.56 | — | |
| Mellum 2Parameters=2.5B/12B2026.05 | 74.9 | — | |
| Llama3.1Architecture=Dense, # Total Params=8.0B, # Trained Tokens=15T2025.10 | 71.56 | — | |
| Qwen3-4BParameters=4B2026.05 | 71.3 | — | |
| Qwen3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=36T2025.10 | 71.14 | — | |
| Qwen3-8B-Instruct2026.04 | 69.4 | — | |
| Qwen2.5-7BParameters=7B2026.05 | 69 | — | |
| Gemma3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=4T2025.10 | 66.32 | — | |
| OLMo-3-7BParameters=7B2026.05 | 63.6 | — | |
| Llama 3-8BParadigm=AR, Training Tokens=15T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=32026.06 | 62.1 | — | |
| AHDgeneration length=10242026.04 | 59.82 | 58.25 | |
| CPTSize=8B, Evaluation Mode=3-shot2024.09 | 58.87 | — | |
| Llama-3Size=8B, Evaluation Mode=3-shot2024.09 | 58.5 | — | |
| Qwen2.5Architecture=Dense, # Total Params=3.0B, # Trained Tokens=18T2025.10 | 55.37 | — | |
| PC-samplergeneration length=10242026.04 | 54.03 | 1,024 | |
| DSRMidtraining Schedule=DSR2026.05 | 53.92 | — | |
| Qwen2.5Architecture=Dense, # Total Params=7.0B, # Trained Tokens=18T2025.10 | 53.72 | — | |
| Fast-dLLMgeneration length=10242026.04 | 53.68 | 65.54 | |
| KLASSgeneration length=10242026.04 | 53.42 | 122.66 | |
| LLaDAgeneration length=10242026.04 | 53.28 | 1,024 | |
| Sabergeneration length=10242026.04 | 52.8 | 213.02 | |
| WSDMidtraining Schedule=WSD2026.05 | 52.19 | — | |
| Unified 15B MoEMidtraining Corpus=4k2026.05 | 51.68 | — | |
| Unified 15B MoEMidtraining Corpus=4k→8k2026.05 | 50.3 | — | |
| LLaDA-8BParadigm=Masked Diffusion, Training Tokens=2.3T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=32026.06 | 49.7 | — | |
| Unified 15B MoEMidtraining Corpus=1k→8k2026.05 | 49.6 | — | |
| Cosine-decayMidtraining Schedule=Cos2026.05 | 48.13 | — | |
| FuseChat3.0#Params=1.48B, Backbone=Llama-3.2-1B-Instruct2026.05 | 45.8 | — | |
| Qwen2.5-Dense2MoELLM=Qwen2.5, Model=Ours, Activated Parameters=1.2 B2026.05 | 42.73 | — | |
| Qwen2.5-DenseLLM=Qwen2.5, Model=Dense, Activated Parameters=1.5 B2026.05 | 39.73 | — | |
| Llama 2-7BParadigm=AR, Training Tokens=2T, Training Data=Not Released, Evaluation protocol=Evaluated under our protocol, Shots=32026.06 | 39.6 | — | |
| Llama3.2Architecture=Dense, # Total Params=3.0B, # Trained Tokens=9T2025.10 | 39.45 | — | |
| Llama2-DenseLLM=Llama2, Model=Dense, Activated Parameters=7.0 B2026.05 | 39.36 | — | |
| Unified 15B MoEPretraining Corpus=1k→8k2026.05 | 38.34 | — | |
| Unified 15B MoEPretraining Corpus=4k→8k2026.05 | 38.03 | — | |
| Unified 15B MoEPretraining Corpus=4k2026.05 | 36.7 | — | |
| SkillWeave#Params=1.42B, Backbone=Llama-3.2-1B-Instruct2026.05 | 36.4 | — | |
| Twin-merging#Params=1.48B, Backbone=Llama-3.2-1B-Instruct2026.05 | 35.2 | — | |
| PEFT#Params=1.42B, Backbone=Llama-3.2-1B-Instruct2026.05 | 34.8 | — | |
| Llama2-Dense2MoELLM=Llama2, Model=Ours, Activated Parameters=5.7 B2026.05 | 34.59 | — | |
| Llama3.2-1B-Instruct#Params=1.15B, Backbone=Llama-3.2-1B-Instruct2026.05 | 32.4 | — | |
| Sumi-7BParadigm=Uniform Diffusion, Training Tokens=1.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=32026.06 | 31.8 | — | |
| self-rewarding#Params=1.15B, Backbone=Llama-3.2-1B-Instruct2026.05 | 30.4 | — | |
| OLMo-7BParadigm=AR, Training Tokens=2.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=32026.06 | 29.8 | — | |
| Falcon-7BParadigm=AR, Training Tokens=1.5T, Training Data=Partially Released, Evaluation protocol=Evaluated under our protocol, Shots=32026.06 | 27.1 | — |