Accuracy on GPQA Diamond (Knowledge)
81.3Accuracy (GPQA Knowledge)Qwen3.5-9B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-9BParameters=9B, Variant=Thinking2026.05 | 81.3 | |
| Kimi-K2.5Mode=Instant2026.06 | 80.52 | |
| DeepSeek-V3.2Thinking Mode=nothink2026.06 | 77.11 | |
| GPT-5.4Reasoning Mode=non-reasoning2026.06 | 76.89 | |
| Qwen3.5-4BParameters=4B, Variant=Thinking2026.05 | 76.8 | |
| Qwen3.5-Omni Flash 35B-A3BContext Length=256K2026.07 | 76.4 | |
| Qwen3.5-4BActive / Total Parameters=4.0B / 4.0B, Sampling Settings=Recommended sampling settings2026.05 | 76.2 | |
| Ling-2.6-1T2026.06 | 76.17 | |
| Audex 30B-A3BContext Length=1M2026.07 | 74.9 | |
| Qwen3-Omni 30B-A3B ThinkingContext Length=64K2026.07 | 73.1 | |
| ZAYA1-8BActive / Total Parameters=0.7B / 8.0B, Sampling Settings=T=1.0, top-p=0.95, top-k disabled2026.05 | 71 | |
| GLM-5Thinking Mode=non-thinking2026.06 | 70.2 | |
| Qwen3-4B-Thinking-2507Active / Total Parameters=4.0B / 4.0B, Sampling Settings=Recommended sampling settings2026.05 | 66.1 | |
| Step-Audio R1.1 33BContext Length=64K2026.07 | 60.7 | |
| Qwen3-8BInference Mode=think2026.03 | 60.1 | |
| SSA-LLM-8BInference Mode=think2026.03 | 60.1 | |
| Mellum 2 (RL)Post-training Stage=RL, Parameters=2.5B/12B, Variant=Thinking2026.05 | 57.6 | |
| Gemma-4-E4B-itActive / Total Parameters=4.0B / 8.0B, Sampling Settings=Recommended sampling settings2026.05 | 57.4 | |
| Qwen3-14B + NGMModel Scale=14B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 52.02 | |
| Qwen3-8BInference Mode=no-think2026.03 | 51.52 | |
| Qwen3-8B + NGMModel Scale=8B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 51.52 | |
| Qwen3-8BModel Scale=8B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 51.01 | |
| Qwen3-14BModel Scale=14B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 48.99 | |
| Audex 2BContext Length=128K2026.07 | 47.2 | |
| Qwen3-4B + NGMModel Scale=4B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 46.46 | |
| Qwen2.5-72BParameters=72B2026.05 | 46 | |
| Ministral-3-14BParameters=14B, Variant=Thinking2026.05 | 46 | |
| SSA-LLM-8BInference Mode=no-think2026.03 | 44.44 | |
| Qwen2.5-32BParameters=32B2026.05 | 43.9 | |
| Qwen3-4BModel Scale=4B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 43.43 | |
| Qwen2.5-14BParameters=14B2026.05 | 41.9 | |
| Mellum 2 (SFT)Post-training Stage=SFT, Parameters=2.5B/12B, Variant=Thinking2026.05 | 39.9 | |
| FORGE-7B-NP-MATHParameters=7B, Configuration=NP-MATH2026.05 | 38.9 | |
| Voxtral Small-24B 2507Context Length=128K2026.07 | 37.9 | |
| FORGE-7B-NPParameters=7B, Configuration=NP2026.05 | 37.4 | |
| InternLM3-8BParameters=8B2026.05 | 36.9 | |
| FORGE-7B-MATHParameters=7B, Configuration=MATH2026.05 | 33.8 | |
| Qwen2.5-7BParameters=7B2026.05 | 33.3 | |
| Qwen3-1.7BModel Scale=1.7B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 32.83 | |
| Qwen3-1.7B + NGMModel Scale=1.7B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 31.31 | |
| Qwen2.5-3BParameters=3B2026.05 | 30.3 | |
| OLMo-3-7BParameters=7B, Variant=Thinking2026.05 | 29.3 | |
| Qwen3-0.6B + NGMModel Scale=0.6B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 29.29 | |
| Qwen3-0.6BModel Scale=0.6B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 26.77 | |
| LLama3.1-8bParameters=8B2026.05 | 22.7 | |
| Original (T2T)Backbone=LLaDA2.1-mini, Inference Strategy=Text-to-Text2026.04 | 21.72 | |
| T2MBackbone=LLaDA2.1-mini, Inference Strategy=Token-to-Mask, Remasking Strategy=LOWPROB, τ=0.3, Cmax=1, ρmax=0.252026.04 | 21.72 | |
| Qwen2.5-7B-SFTParameters=7B, Configuration=SFT2026.05 | 18.7 | |
| MiMo-Audio 7BContext Length=8K2026.07 | 13.8 |