Code Generation on LiveCodeBench v6
100AccuracyAutoGen
Evaluation Results
| Method | Links | |
|---|---|---|
| AutoGenBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 100 | |
| Agent Q-MixBase Model=GPT-OSS:120B2026.04 | 100 | |
| MaASBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 100 | |
| LobsterBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 100 | |
| LangGraphBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 100 | |
| AutoGenBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 100 | |
| Agent FrameworkBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 100 | |
| Agent Q-MixBase Model=Gemini-3.1-Flash-Lite2026.04 | 100 | |
| GPTSwarmBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 99 | |
| G-DesignerBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 98.75 | |
| Agent FrameworkBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 98.25 | |
| TopoDIMBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 98.25 | |
| LLM-DebateBase Model=Gemini-3.1-Flash-Lite, Framework Category=Static multi-agent2026.04 | 98 | |
| GTDBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 97.75 | |
| Base (direct)Base Model=Gemini-3.1-Flash-Lite, Framework Category=Single-agent baseline2026.04 | 96.5 | |
| AgentDropoutBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 96.25 | |
| LobsterBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 96 | |
| GPTSwarmBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 95.75 | |
| MaASBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 94.5 | |
| TopoDIMBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 94 | |
| GTDBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 93.75 | |
| LangGraphBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 93.5 | |
| G-DesignerBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 92.25 | |
| LLM-DebateBase Model=GPT-OSS:120B, Framework Category=Static multi-agent2026.04 | 92 | |
| AgentDropoutBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 90 | |
| Base (direct)Base Model=GPT-OSS:120B, Framework Category=Single-agent baseline2026.04 | 88.75 | |
| Conductorparameters=7B2025.12 | 83.93 | |
| GPT 52025.12 | 82.9 | |
| Agentic Proposingtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-5, proposer=Agentic-Proposer-30B2026.02 | 71.2 | |
| PromptCoT 2.0training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 71 | |
| OpenCodeReasoningtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 67.4 | |
| Gemini 2.5 Pro2025.12 | 67.24 | |
| OpenThoughts-S3training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 67.2 | |
| Qwen3-30B-A3B-Thinking-2507mode=zero-shot, evaluation=Best-of-52026.02 | 66 | |
| OpenR1training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 63.7 | |
| MSE+DistillationModel=Qwen3.5-9B, Learning Rate=5e-62026.05 | 63.7 | |
| MSE+DistillationModel=Qwen3.5-9B, Learning Rate=1e-52026.05 | 62.4 | |
| OpenMathReasoningtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 58.4 | |
| Qwen3Params=32B2026.04 | 58.3 | |
| I-DLMParams=32B, Decoding=ISD (N=4, sampling)2026.04 | 57.1 | |
| BaseModel=Qwen3.5-9B2026.05 | 56.1 | |
| BF16 BaselinePrecision=BF16, Model=AceReason Nemotron 1.1 7B2026.01 | 54.3 | |
| NVFP4 QADPrecision=NVFP4, Quantization=QAD, Model=AceReason Nemotron 1.1 7B2026.01 | 53.3 | |
| NVFP4 PTQPrecision=NVFP4, Quantization=PTQ, Model=AceReason Nemotron 1.1 7B2026.01 | 52 | |
| CEModel=Qwen3.5-9B2026.05 | 50.7 | |
| Qwen3Params=8B2026.04 | 50.3 | |
| Claude Sonnet 42025.12 | 46.54 | |
| NVFP4 QATPrecision=NVFP4, Quantization=QAT, Model=AceReason Nemotron 1.1 7B2026.01 | 45.9 | |
| I-DLMParams=8B, Decoding=ISD (N=4, sampling)2026.04 | 45.7 | |
| LLaDA-2.1-flashParams=100B2026.04 | 45.4 | |
| LLaDA-2.0-flashParams=100B2026.04 | 42.5 | |
| MSE+DistillationModel=Qwen3.5-9B, Learning Rate=1e-62026.05 | 40.2 | |
| SuCoModel Scale=7B2026.06 | 38.9 | |
| Skywork-OR1-7BDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B2026.05 | 36.57 | |
| S-GRPOModel Scale=7B2026.06 | 35.4 | |
| LHRMsModel Scale=7B2026.06 | 35.4 | |
| Skywork-OR1-Math-7BDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B2026.05 | 34.86 | |
| AdaptThinkModel Scale=7B2026.06 | 32.6 | |
| AdaCoTModel Scale=7B2026.06 | 32 | |
| DeepSeek-R1-DistillModel Scale=7B2026.06 | 31.4 | |
| MiMo-V2-Flash Base# Shots=1-shot, # Activated Params=15B, # Total Params=309B2026.01 | 30.8 | |
| LLaDA-2.1-miniParams=16B2026.04 | 30.4 | |
| R1-Distill-Qwen-32B2025.12 | 26.86 | |
| Kimi-K2 Base# Shots=1-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 26.3 | |
| Qwen3-32B (thinking)mode=thinking2025.12 | 25.86 | |
| DeepSeek-V3.2 Exp Base# Shots=1-shot, # Activated Params=37B, # Total Params=671B2026.01 | 24.9 | |
| DeepSeek-V3.1 Base# Shots=1-shot, # Activated Params=37B, # Total Params=671B2026.01 | 24.8 | |
| SuCoModel Scale=1.5B2026.06 | 22.3 | |
| TrOPDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 22.29 | |
| SDARParams=30B2026.04 | 21.7 | |
| Qwen3-32B2025.12 | 21.21 | |
| LHRMsModel Scale=1.5B2026.06 | 20.6 | |
| OPDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 20.57 | |
| REOPOLDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 19.43 | |
| S-GRPOModel Scale=1.5B2026.06 | 19.4 | |
| TrOPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 18.86 | |
| AdaptThinkModel Scale=1.5B2026.06 | 18.3 | |
| REOPOLDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 18.29 | |
| AdaCoTModel Scale=1.5B2026.06 | 17.7 | |
| OPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 17.14 | |
| DeepSeek-R1-DistillModel Scale=1.5B2026.06 | 17.1 | |
| SDARParams=8B2026.04 | 16.6 | |
| REOPOLD 2StageDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 16.57 | |
| DeepSeek-Qwen2.5-1.5BStudent Model=DeepSeek-Qwen2.5-1.5B2026.05 | 15.43 | |
| EOPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 15.43 | |
| Entropy OPD 20%Distillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 14.29 | |
| gemma-3-27b-itinstruct=true2025.12 | 13.14 | |
| Math-InstructModel Scale=7B2026.06 | 8 | |
| Math-InstructModel Scale=1.5B2026.06 | 2.3 | |
| Math-BaseModel Scale=7B2026.06 | 1.1 | |
| Math-BaseModel Scale=1.5B2026.06 | 0.6 |