General Knowledge and Reasoning on MMLU
89.5AccuracyClaude Sonnet 4.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude Sonnet 4.5Input Modality=Text, LLM-as-a-Judge=GPT-4o2025.11 | 89.5 | |
| Qwen3-Next-80B-A3B-Instruct2026.03 | 89.28 | |
| DuetNetConnection Status=After connection, Collaboration Strategy=Token-level2026.05 | 88 | |
| Gemini 2.5 ProInput Modality=Text, LLM-as-a-Judge=GPT-4o2025.11 | 87.7 | |
| Qwen3-Omni-A3B-Instruct2026.03 | 87.1 | |
| GPT-5 highInput Modality=Text, LLM-as-a-Judge=GPT-4o2025.11 | 86 | |
| Qwen 2.5Connection Status=Before connection2026.05 | 86 | |
| RouterConnection Status=After connection, Collaboration Strategy=Routing-based2026.05 | 86 | |
| LongCat-Next2026.03 | 83.95 | |
| Kimi-Linear-48B-A3B2026.03 | 79.91 | |
| Qwen3-8BModel Category=Text-Experts, Model Scale=8B2026.03 | 76.9 | |
| Ministral-3-8BModel Category=Text-Experts, Model Scale=8B2026.03 | 76.1 | |
| HyperCLOVAX-8B-OmniModel Category=Unified, Model Scale=8B, Native speech I/O=true2026.03 | 75.7 | |
| Dynin-OmniModel Category=Unified, Native speech I/O=true, Block size=16, Diffusion steps=10242026.03 | 75.2 | |
| Baichuan-Omni-1.5Model Category=Perception-centric, Model Scale=1.52026.03 | 72.2 | |
| Qwen2.5-Omni-7BModel Category=Perception-centric, Model Scale=7B2026.03 | 71.8 | |
| Show-o2Model Category=Unified, Video support=true2026.03 | 70.7 | |
| Sora-2 Last FrameInput Modality=Last Frame, LLM-as-a-Judge=GPT-4o2025.11 | 69.1 | |
| MMaDAModel Category=Unified2026.03 | 68.4 | |
| Sora-2 AudioInput Modality=Audio, LLM-as-a-Judge=GPT-4o2025.11 | 67.3 | |
| Trida-7BModel Category=Text-Experts, Model Scale=7B2026.03 | 67.2 | |
| Llama-3-8BModel Category=Text-Experts, Model Scale=8B2026.03 | 66.6 | |
| LLaDA-8BModel Category=Text-Experts, Model Scale=8B2026.03 | 65.9 | |
| BaseTraining Protocol=Base2025.08 | 63.1 | |
| PSFTTraining Protocol=PSFT2025.08 | 62.44 | |
| Chameleon-7BModel Category=Unified, Model Scale=7B2026.03 | 52.1 | |
| SFTTraining Protocol=SFT2025.08 | 45.39 |