Code Generation on HumanEval (Accuracy, Exit Position)
98.27AccuracySASFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SASFTModel=Qwen3-8B-Base, Training Dataset=Chinese 110k2026.05 | 98.27 | — | |
| SFT+GRPOModel=Qwen3-8B-Base, Training Dataset=Chinese 110k2026.05 | 96.44 | — | |
| SFTModel=Qwen3-8B-Base, Training Dataset=Chinese 110k2026.05 | 95.87 | — | |
| SIGMABase LLM=GPT-OSS-120B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=Full2026.06 | 95.83 | — | |
| SFT+PenaltyModel=Qwen3-8B-Base, Training Dataset=Chinese 110k2026.05 | 94.71 | — | |
| CARDBase LLM=GPT-OSS-120B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 93.98 | — | |
| AgentReviveBase model=Deepseek-V3-671B-Instruct, Var. NS=true, Flex. State=true2026.05 | 93.52 | — | |
| ARG-DesignerBase LLM=GPT-OSS-120B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 93.47 | — | |
| G-DesignerBase LLM=GPT-OSS-120B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 92.88 | — | |
| GPTSwarmBase LLM=GPT-OSS-120B, Multi-agent execution=Full, Communication-topology design=Partial, Skill-based composition=No2026.06 | 92.3 | — | |
| AgentDropoutBase model=Deepseek-V3-671B-Instruct, Var. NS=true, Flex. State=false2026.05 | 91.74 | — | |
| LARBackbone=Qwen3-8B2026.05 | 91.46 | — | |
| LLM-DebateBase LLM=GPT-OSS-120B, Multi-agent execution=Full, Communication-topology design=No, Skill-based composition=No2026.06 | 91.43 | — | |
| ARG-DesignerBase model=Deepseek-V3-671B-Instruct, Var. NS=true, Flex. State=false2026.05 | 91.18 | — | |
| AgentPruneBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 90.91 | — | |
| CoTBase LLM=GPT-OSS-120B, Multi-agent execution=No, Communication-topology design=No, Skill-based composition=No2026.06 | 90.87 | — | |
| SC (CoT)Base model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 90.61 | — | |
| SFT+GRPOModel=Qwen3-1.7B-Base, Training Dataset=Chinese 110k2026.05 | 90.48 | — | |
| SFTModel=Qwen3-1.7B-Base, Training Dataset=Chinese 110k2026.05 | 90.29 | — | |
| G-DesignerBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 90.2 | — | |
| VanillaBase LLM=GPT-OSS-120B, Multi-agent execution=No, Communication-topology design=No, Skill-based composition=No2026.06 | 90.04 | — | |
| ReActBackbone=Qwen3-8B2026.05 | 89.63 | — | |
| CoTBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 89.26 | — | |
| AutoGenBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 89.26 | — | |
| MASround=TBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 89.26 | — | |
| MASround=1Base model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 89.17 | — | |
| SFT+PenaltyModel=Qwen3-1.7B-Base, Training Dataset=Chinese 110k2026.05 | 89.13 | — | |
| SASFTModel=Qwen3-1.7B-Base, Training Dataset=Chinese 110k2026.05 | 89.04 | — | |
| AgentVerseBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 88.94 | — | |
| SIGMABase LLM=Qwen3-8B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=Full2026.06 | 88.71 | — | |
| VanillaBase model=Deepseek-V3-671B-Instruct, Var. NS=false, Flex. State=false2026.05 | 88.43 | — | |
| CARDBase LLM=Qwen3-8B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 87.28 | — | |
| SIGMABase LLM=GPT-4o-mini, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=Full2026.06 | 86.29 | — | |
| RLSTAModel=Qwen2.5-7B-Instruct, Evaluation Protocol=FULL2026.05 | 86 | — | |
| ARG-DesignerBase LLM=Qwen3-8B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 85.47 | — | |
| MAIGOModel=Qwen2.5-7B-Instruct, Evaluation Protocol=SHARDED2026.05 | 85.2 | — | |
| CARDBase LLM=GPT-4o-mini, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 85.11 | — | |
| GRPOModel=Qwen2.5-7B-Instruct, Evaluation Protocol=FULL2026.05 | 85.1 | — | |
| BaseModel=Qwen2.5-7B-Instruct, Evaluation Protocol=FULL2026.05 | 84.8 | — | |
| G-DesignerBase LLM=Qwen3-8B, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 84.73 | — | |
| ARG-DesignerBase LLM=GPT-4o-mini, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 84.54 | — | |
| GPTSwarmBase LLM=Qwen3-8B, Multi-agent execution=Full, Communication-topology design=Partial, Skill-based composition=No2026.06 | 84.14 | — | |
| RLSTAModel=Qwen2.5-7B-Instruct, Evaluation Protocol=SHARDED2026.05 | 84 | — | |
| LLM-DebateBase LLM=Qwen3-8B, Multi-agent execution=Full, Communication-topology design=No, Skill-based composition=No2026.06 | 83.85 | — | |
| GRPOModel=Qwen2.5-7B-Instruct, Evaluation Protocol=SHARDED2026.05 | 83.5 | — | |
| CoTBase LLM=Qwen3-8B, Multi-agent execution=No, Communication-topology design=No, Skill-based composition=No2026.06 | 82.85 | — | |
| G-DesignerBase LLM=GPT-4o-mini, Multi-agent execution=Full, Communication-topology design=Full, Skill-based composition=No2026.06 | 82.67 | — | |
| MAIGOModel=Qwen2.5-7B-Instruct, Evaluation Protocol=FULL2026.05 | 82.5 | — | |
| VanillaBase LLM=Qwen3-8B, Multi-agent execution=No, Communication-topology design=No, Skill-based composition=No2026.06 | 82.23 | — | |
| GPTSwarmBase LLM=GPT-4o-mini, Multi-agent execution=Full, Communication-topology design=Partial, Skill-based composition=No2026.06 | 81.39 | — | |
| BaseModel=Qwen2.5-7B-Instruct, Evaluation Protocol=SHARDED2026.05 | 80.6 | — | |
| SFTModel=Qwen2.5-7B-Instruct, Evaluation Protocol=FULL2026.05 | 80.5 | — | |
| LLM-DebateBase LLM=GPT-4o-mini, Multi-agent execution=Full, Communication-topology design=No, Skill-based composition=No2026.06 | 79.48 | — | |
| CoTBase LLM=GPT-4o-mini, Multi-agent execution=No, Communication-topology design=No, Skill-based composition=No2026.06 | 78.64 | — | |
| SFTModel=Qwen2.5-7B-Instruct, Evaluation Protocol=SHARDED2026.05 | 77.7 | — | |
| VanillaBase LLM=GPT-4o-mini, Multi-agent execution=No, Communication-topology design=No, Skill-based composition=No2026.06 | 77.29 | — | |
| MAIGOModel=Qwen2.5-3B-Instruct, Evaluation Protocol=SHARDED2026.05 | 76.1 | — | |
| RLSTAModel=Qwen2.5-3B-Instruct, Evaluation Protocol=SHARDED2026.05 | 73.5 | — | |
| GRPOModel=Qwen2.5-3B-Instruct, Evaluation Protocol=SHARDED2026.05 | 73.3 | — | |
| DreamReasoner-8BType=Block Diffusion2026.06 | 69.5 | — | |
| BaseModel=Qwen2.5-3B-Instruct, Evaluation Protocol=SHARDED2026.05 | 69.4 | — | |
| Qwen3-8BType=AR2026.06 | 68.9 | — | |
| SFTModel=Qwen2.5-3B-Instruct, Evaluation Protocol=SHARDED2026.05 | 67.8 | — | |
| Phi4-miniModel Architecture=Phi4-mini (32 layers), Evaluation Protocol=Backbone2026.04 | 63.4 | — | |
| River-LLMModel Architecture=Phi4-mini (32 layers), Early-exit threshold (tau)=0.9, Evaluation Protocol=Adaptive2026.04 | 63.1 | 7.72 | |
| River-LLMModel Architecture=Phi4-mini (32 layers), Early-exit threshold (tau)=0.8, Evaluation Protocol=Adaptive2026.04 | 62.8 | 2.55 | |
| VoidExpansionExpansion strategy=VoidExpansion, tau_nonvoid=0.15, tau_gap=0.552026.06 | 60.98 | — | |
| LARBackbone=Llama-3.1-8B-Instruct2026.05 | 60.37 | — | |
| MAIGOModel=Qwen2.5-3B-Instruct, Evaluation Protocol=FULL2026.05 | 59.6 | — | |
| BaseModel=Qwen2.5-3B-Instruct, Evaluation Protocol=FULL2026.05 | 59 | — | |
| RLDFGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 58.5 | — | |
| AgentReviveBase model=Llama3-8B-Instruct, Var. NS=true, Flex. State=true2026.05 | 58.15 | — | |
| RLDFGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 57.9 | — | |
| RLDFGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 57.9 | — | |
| Dream-v0-7BType=Diffusion2026.06 | 57.9 | — | |
| BackboneBackbone Model=Llama3.1 8B (32 layers)2026.04 | 57.3 | — | |
| River-LLMBackbone Model=Llama3.1 8B (32 layers), Threshold (τ)=0.72026.04 | 57.3 | 2.16 | |
| ReActBackbone=Llama-3.1-8B-Instruct2026.05 | 56.71 | — | |
| Qwen2.5-7BType=AR2026.06 | 56.7 | — | |
| Coupled-GRPOGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 56.1 | — | |
| AgentDropoutBase model=Llama3-8B-Instruct, Var. NS=true, Flex. State=false2026.05 | 55.84 | — | |
| River-LLMBackbone Model=Llama3.1 8B (32 layers), Threshold (τ)=0.52026.04 | 55.5 | 2.16 | |
| ESPOGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 55.5 | — | |
| SC (CoT)Base model=Llama3-8B-Instruct, Var. NS=false, Flex. State=false2026.05 | 55.46 | — | |
| RLSTAModel=Qwen2.5-3B-Instruct, Evaluation Protocol=FULL2026.05 | 55 | — | |
| CoTBase model=Llama3-8B-Instruct, Var. NS=false, Flex. State=false2026.05 | 54.17 | — | |
| d1Generation steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 53.7 | — | |
| TraceRLGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 53.7 | — | |
| ARG-DesignerBase model=Llama3-8B-Instruct, Var. NS=true, Flex. State=false2026.05 | 53.62 | — | |
| SFTModel=Qwen2.5-3B-Instruct, Evaluation Protocol=FULL2026.05 | 53.5 | — | |
| VanillaBase model=Llama3-8B-Instruct, Var. NS=false, Flex. State=false2026.05 | 53.33 | — | |
| ESPOGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 53 | — | |
| Coupled-GRPOGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 53 | — | |
| Coupled-GRPOGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 53 | — | |
| G-DesignerBase model=Llama3-8B-Instruct, Var. NS=false, Flex. State=false2026.05 | 52.53 | — | |
| ESPOGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 52.4 | — | |
| TraceRLGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 52.4 | — | |
| d1Generation steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 51.8 | — | |
| d1Generation steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 51.8 | — | |
| MiMo-7BType=AR2026.06 | 51.8 | — |