Coding on HumanEval
98.17Pass@1CompassMax-V3-Thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| CompassMax-V3-Thinking2025.12 | 98.17 | |
| DeepSeek-R12025.12 | 96.95 | |
| SwiRBackbone=Qwen3-8B2025.10 | 95.73 | |
| CoT (Greedy)Backbone=Qwen3-8B, decoding=Greedy2025.10 | 93.9 | |
| CoTBackbone=Qwen3-8B2025.10 | 92.68 | |
| Soft ThinkingBackbone=Qwen3-8B2025.10 | 92.07 | |
| MT (Teacher)Teacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 91.5 | |
| GPT-4o2024.10 | 90.2 | |
| Llama-3.1-405BModel Type=Instruct, Parameters=405B2025.02 | 89 | |
| Qwen2.5-VL-72BModel Type=Instruct, Parameters=72B2025.02 | 87.8 | |
| GPT-4o mini2024.10 | 87.2 | |
| Qwen2.5-72BModel Type=Instruct, Parameters=72B2025.02 | 86.6 | |
| Qwen2-72B-InstructType=Instruction-tuned2024.07 | 86 | |
| Qwen2-72BModel Type=Instruct, Parameters=72B2025.02 | 86 | |
| CompassMax-V32025.12 | 84.76 | |
| Gemini-1.5 Pro2024.10 | 84.1 | |
| UltraMix-190kBase Model=Qwen-2.5-7B-TuluSFT, Training Mixture/Protocol=UM-190k2025.11 | 82.27 | |
| Llama-3-70B-InstructType=Instruction-tuned2024.07 | 81.7 | |
| UltraMix-187kBase Model=Qwen-2.5-7B-TuluSFT, Training Mixture/Protocol=UM-187k2025.11 | 81.1 | |
| Llama-3.1-70BModel Type=Instruct, Parameters=70B2025.02 | 80.5 | |
| TuluDPOBase Model=Qwen-2.5-7B-TuluSFT, Training Mixture/Protocol=TuluDPO2025.11 | 80.49 | |
| HPDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 79.3 | |
| UltraMix-170kBase Model=Qwen-2.5-7B-TuluSFT, Training Mixture/Protocol=UM-170k2025.11 | 78.05 | |
| KDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 77.4 | |
| JSDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 77.4 | |
| RKLDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 76.8 | |
| MT (Teacher)Teacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 76.2 | |
| Qwen1.5-110B-ChatType=Instruction-tuned2024.07 | 74.4 | |
| Gemini-1.5 Flash2024.10 | 74.3 | |
| Mixtral-8x22B-InstructType=Instruction-tuned2024.07 | 73.8 | |
| SFTTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 73.8 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 73.8 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 73.8 | |
| BaseBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Pre-trained2026.05 | 73.8 | |
| ARIA2024.10 | 73.2 | |
| LangMARLCategory=Ours2026.04 | 73.2 | |
| CoT2-MetaStrategy=Ours (CoT2-Meta), Inference Budget=C=162026.03 | 72.8 | |
| SFTBase Model=Qwen-2.5-7B-TuluSFT, Training Mixture/Protocol=SFT2025.11 | 72.66 | |
| Llama3.2-11B2024.10 | 72.6 | |
| Pixtral-12B2024.10 | 72 | |
| Qwen1.5-72B-ChatType=Instruction-tuned2024.07 | 71.3 | |
| MS (Student)Teacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 71.3 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 71.3 | |
| Self-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 71.3 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 71.3 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 3: iGSM -> MedCalc -> IFEval2026.05 | 71.3 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 3: iGSM -> MedCalc -> IFEval2026.05 | 70.7 | |
| ReST-MCTS*Strategy=ReST-MCTS*, Inference Budget=C=162026.03 | 70.4 | |
| ReflexionCategory=Self-Evolving2026.04 | 70.1 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 70.1 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 1: iGSM2026.05 | 70.1 | |
| HPDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 69.5 | |
| UltraMix-190kBase Model=Llama-3.1-8B-TuluSFT, Training Mixture/Protocol=UM-190k2025.11 | 69.05 | |
| TextGradCategory=Self-Evolving2026.04 | 68.9 | |
| Self-sftBackbone=Qwen2.5-3B-Instruct2026.05 | 68.9 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 68.9 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 1: iGSM2026.05 | 68.9 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 68.3 | |
| Vanilla ToTStrategy=Vanilla ToT, Inference Budget=C=162026.03 | 68.2 | |
| UltraMix-187kBase Model=Llama-3.1-8B-TuluSFT, Training Mixture/Protocol=UM-187k2025.11 | 68.06 | |
| TuluDPOBase Model=Llama-3.1-8B-TuluSFT, Training Mixture/Protocol=TuluDPO2025.11 | 67.24 | |
| JSDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 67.1 | |
| GPT-4V2024.10 | 67 | |
| DSPyCategory=Self-Evolving2026.04 | 66.7 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 66.5 | |
| UltraMix-170kBase Model=Llama-3.1-8B-TuluSFT, Training Mixture/Protocol=UM-170k2025.11 | 65.61 | |
| Best-of-16Strategy=Best-of-16, Inference Budget=C=162026.03 | 65.4 | |
| KDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 65.2 | |
| Youtu-LLMSize=2B, Type=Base, Prompt Setting=0-shot2025.12 | 64.6 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 64.6 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 2: iGSM -> MedCalc2026.05 | 64.6 | |
| SymbolicCategory=Self-Evolving2026.04 | 64.5 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 64 | |
| AutoPECategory=Self-Evolving2026.04 | 63.5 | |
| MS (Student)Teacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 62.8 | |
| ProSeCo SamplingCorrector Sampling=true, Backbone=LLaDA-Base 8B, Number of shots=02026.02 | 62.2 | |
| RKLDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 61.6 | |
| SFTTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 61 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 2: iGSM -> MedCalc2026.05 | 61 | |
| Greedy CoTStrategy=Greedy CoT, Inference Budget=C=162026.03 | 60.5 | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 60.4 | |
| AgentsCategory=Static-Prompting2026.04 | 59.5 | |
| CoTCategory=Static-Prompting2026.04 | 59.2 | |
| Llama3.1-InstructCorrector Sampling=false, Model scale=8B, Number of shots=02026.02 | 58.54 | |
| SFTBase Model=Llama-3.1-8B-TuluSFT, Training Mixture/Protocol=SFT2025.11 | 57.93 | |
| Iter-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 57.7 | |
| Qwen3Size=4B, Type=Base, Prompt Setting=0-shot2025.12 | 57.6 | |
| Anchored LearningBackbone=Llama-3.2-3B-Instruct2026.05 | 56.1 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 54.9 | |
| Low-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 53.7 | |
| STMBackbone=Llama-3.2-3B-Instruct2026.05 | 53.7 | |
| Self-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 53.7 | |
| ProSeCo SFTCorrector Sampling=false, Backbone=LLaDA-Base 8B, Number of shots=02026.02 | 52.44 | |
| CodePrefBackbone=Apertus-8B-SFT2025.11 | 51.24 | |
| Qwen3Size=1.7B, Type=Base, Prompt Setting=0-shot2025.12 | 49.9 | |
| SELFORGBackbone model=Qwen-1.5B2025.10 | 48.97 | |
| Vanilla SFTCorrector Sampling=false, Training strategy=SFT, Backbone=LLaDA-Base 8B, Number of shots=02026.02 | 48.17 | |
| UM-190kBackbone=Apertus-8B-SFT2025.11 | 48.07 | |
| UM-187kBackbone=Apertus-8B-SFT2025.11 | 47.68 | |
| Phi-2# Non-Emb Params=2.5B2024.07 | 47.6 |