Science Reasoning on GPQA Diamond (Accuracy, Avg., Drop)
83AccuracyGPT-5-Mini
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5-Mini2026.03 | 83 | — | — | |
| GPT-5-Mini-R2026.03 | 83 | — | — | |
| PRISMAggregation=Majority Vote, Cost ($)=6.762026.03 | 71.7 | — | — | |
| PRISMAggregation=PRM-score Vote, Cost ($)=6.762026.03 | 71.4 | — | — | |
| gpt-oss-120bMode=zero-shot, Aggregation=–, Cost ($)=0.202026.03 | 69.7 | — | — | |
| Recursive Self-AggregationAggregation=Majority Vote, Cost ($)=4.032026.03 | 68.6 | — | — | |
| PRISMAggregation=LLM Aggregate, Cost ($)=7.112026.03 | 68.3 | — | — | |
| MAD ConformistAggregation=Majority Vote, Cost ($)=1.712026.03 | 66.9 | — | — | |
| Majority VoteAggregation=Majority Vote, Cost ($)=0.872026.03 | 65.8 | — | — | |
| Agentic DebateAggregation=Majority Vote, Cost ($)=4.672026.03 | 65 | — | — | |
| MAD FollowerAggregation=Majority Vote, Cost ($)=3.302026.03 | 64.7 | — | — | |
| Recursive Self-AggregationAggregation=LLM Aggregate, Cost ($)=4.152026.03 | 64.7 | — | — | |
| PRM-score VoteAggregation=PRM-score Vote, Cost ($)=1.012026.03 | 64.7 | — | — | |
| SciMasterAggregation=LLM Aggregate, Cost ($)=6.872026.03 | 63.9 | — | — | |
| TableLongBase Model=DS-R1-Distill-32B, RL-finetuned=true2026.03 | 63.64 | — | — | |
| MAD ConformistAggregation=LLM Aggregate, Cost ($)=1.812026.03 | 63.6 | — | — | |
| SciMasterAggregation=Majority Vote, Cost ($)=6.602026.03 | 63.4 | — | — | |
| MAD FollowerAggregation=LLM Aggregate, Cost ($)=3.42026.03 | 63.3 | — | — | |
| DiffMASModel=Qwen3-8B, Communication Category=Trained Latent Communication2026.04 | 60.1 | — | — | |
| gpt-oss-20bMode=zero-shot, Aggregation=–, Cost ($)=0.082026.03 | 59.7 | — | — | |
| TableLongBase Model=DS-R1-Distill-14B, RL-finetuned=true2026.03 | 59.6 | — | — | |
| Agentic DebateAggregation=LLM Aggregate, Cost ($)=4.802026.03 | 58.1 | — | — | |
| DiffMASBackbone=DeepSeek-R1-Distill-Qwen-32B2026.04 | 57.5 | — | — | |
| SuCoModel Scale=7B2026.06 | 56.6 | — | — | |
| TextMASBackbone=DeepSeek-R1-Distill-Qwen-32B2026.04 | 56.5 | — | — | |
| BF16 BaselineModel=Qwen3-4B, W-Bits=BF162026.01 | 56.06 | 70.63 | — | |
| DS-R1-Distill-32BModel=DS-R1-Distill-32B2026.03 | 56.06 | — | — | |
| DS-R1-Distill-14BModel=DS-R1-Distill-14B2026.03 | 55.56 | — | — | |
| LatentMASBackbone=DeepSeek-R1-Distill-Qwen-32B2026.04 | 55.1 | — | — | |
| SingleBackbone=DeepSeek-R1-Distill-Qwen-32B2026.04 | 53.3 | — | — | |
| DiffMASBackbone=Qwen3-14B2026.04 | 53 | — | — | |
| DiffMASBackbone=Mistral3-8B2026.04 | 52 | — | — | |
| LatentMASBackbone=Qwen3-14B2026.04 | 52 | — | — | |
| TextMASBackbone=Qwen3-14B2026.04 | 51.5 | — | — | |
| S-GRPOModel Scale=7B2026.06 | 51.5 | — | — | |
| TextMASBackbone=Mistral3-8B2026.04 | 51 | — | — | |
| LHRMsModel Scale=7B2026.06 | 49.5 | — | — | |
| SingleBackbone=Qwen3-14B2026.04 | 48.5 | — | — | |
| Reasoning-QATModel=Qwen3-4B, W-Bits=W4A4KV42026.01 | 48.48 | 60.78 | -9.85 | |
| SingleBackbone=Mistral3-8B2026.04 | 47.9 | — | — | |
| FlatQuantModel=Qwen3-4B, W-Bits=W4A4KV42026.01 | 47.47 | 58.28 | -12.35 | |
| AdaCoTModel Scale=7B2026.06 | 47 | — | — | |
| AdaptThinkModel Scale=7B2026.06 | 47 | — | — | |
| DiffMASModel=Qwen3-4B, Communication Category=Trained Latent Communication2026.04 | 46.4 | — | — | |
| LatentMASBackbone=Mistral3-8B2026.04 | 46.4 | — | — | |
| LatentMASModel=Qwen3-8B, Communication Category=Training-free Latent Communication2026.04 | 45.5 | — | — | |
| DeepSeek-R1-DistillModel Scale=7B2026.06 | 45.5 | — | — | |
| TextMASModel=Qwen3-4B, Communication Category=Text Communication2026.04 | 44.9 | — | — | |
| TextMASModel=Qwen3-8B, Communication Category=Text Communication2026.04 | 43.4 | — | — | |
| SingleModel=Qwen3-4B, Communication Category=Text Communication2026.04 | 42.4 | — | — | |
| C2CModel=Qwen3-8B, Communication Category=Trained Latent Communication2026.04 | 41.4 | — | — | |
| SingleModel=Qwen3-8B, Communication Category=Text Communication2026.04 | 39.9 | — | — | |
| BF16 BaselineModel=R1-1.5B, W-Bits=BF162026.01 | 36.87 | 48.72 | — | |
| LatentMASModel=Qwen3-4B, Communication Category=Training-free Latent Communication2026.04 | 36.4 | — | — | |
| SuCoModel Scale=1.5B2026.06 | 33.3 | — | — | |
| Reasoning-QATModel=R1-1.5B, W-Bits=W4A4KV42026.01 | 32.83 | 41.31 | -7.41 | |
| Math-InstructModel Scale=7B2026.06 | 32.3 | — | — | |
| FlatQuantModel=R1-1.5B, W-Bits=W4A4KV42026.01 | 31.82 | 38.39 | -10.33 | |
| LHRMsModel Scale=1.5B2026.06 | 30.3 | — | — | |
| FlatQuantModel=Qwen3-0.6B, W-Bits=W4A4KV42026.01 | 29.8 | 17.34 | -23.76 | |
| C2CModel=Qwen3-4B, Communication Category=Trained Latent Communication2026.04 | 29.8 | — | — | |
| BF16 BaselineModel=Qwen3-0.6B, W-Bits=BF162026.01 | 28.45 | 41.1 | — | |
| S-GRPOModel Scale=1.5B2026.06 | 28.3 | — | — | |
| Reasoning-QATModel=Qwen3-0.6B, W-Bits=W4A4KV42026.01 | 26.94 | 21.44 | -19.66 | |
| AdaptThinkModel Scale=1.5B2026.06 | 26.8 | — | — | |
| AdaCoTModel Scale=1.5B2026.06 | 26.3 | — | — | |
| DeepSeek-R1-DistillModel Scale=1.5B2026.06 | 25.3 | — | — | |
| QuaRotModel=Qwen3-0.6B, W-Bits=W4A4KV42026.01 | 24.24 | 4.84 | -36.26 | |
| Math-InstructModel Scale=1.5B2026.06 | 22.2 | — | — | |
| Math-BaseModel Scale=7B2026.06 | 13.1 | — | — | |
| QuaRotModel=R1-1.5B, W-Bits=W4A4KV42026.01 | 8.59 | 2.11 | -46.61 | |
| Math-BaseModel Scale=1.5B2026.06 | 4 | — | — |