Code Generation on HumanEval-ET
89.6Pass@1ThinkCoder
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ThinkCoderBackbone=GPT-4o2024.12 | 89.6 | — | — | — | |
| ThinkCoder (GPT-4o)n=5, k=5, temperature=0.52024.12 | 89.6 | — | — | — | |
| ThinkCoder (GPT-4o)n=2, k=5, temperature=0.52024.12 | 89 | — | — | — | |
| LCPBase Model=GPT-4, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 88.41 | — | — | — | |
| ThinkCoder (GPT-4o)n=1, k=5, temperature=0.52024.12 | 87.8 | — | — | — | |
| AgentCoderBackbone=GPT-42023.12 | 86 | 70 | — | — | |
| ThinkCoderBackbone=GPT-4-Turbo2024.12 | 86 | — | — | — | |
| AgentCoderBase Model=GPT-4, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 86 | — | — | — | |
| GPT-4on=0, temperature=02024.12 | 85.8 | — | — | — | |
| LPWBackbone=GPT-4o2024.12 | 84.8 | — | — | — | |
| ThinkCoder (GPT-4-Turbo)n=1, k=5, temperature=0.52024.12 | 84.8 | — | — | — | |
| ThinkCoder (GPT-4-Turbo)n=2, k=5, temperature=0.52024.12 | 84.8 | — | — | — | |
| ThinkCoder (GPT-4-Turbo)n=5, k=5, temperature=0.52024.12 | 84.8 | — | — | — | |
| LCPBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 84.1 | — | — | — | |
| CodeCoRBase Model=GPT-4, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 83.5 | — | — | — | |
| MapCoderLLM=GPT-42024.05 | 82.9 | — | — | — | |
| MapCoderBackbone=GPT-4-Turbo2024.12 | 82.9 | — | — | — | |
| LDBBackbone=GPT-4o2024.12 | 81.7 | — | — | — | |
| ThinkCoderBackbone=GPT-3.5-Turbo2024.12 | 81 | — | — | — | |
| GPT-4-Turbon=0, temperature=02024.12 | 81 | — | — | — | |
| ThinkCoder (CodeQwen1.5-7B-Chat)n=1, k=5, temperature=0.52024.12 | 80.5 | — | — | — | |
| ThinkCoder (CodeQwen1.5-7B-Chat)n=2, k=5, temperature=0.52024.12 | 80.5 | — | — | — | |
| ThinkCoder (CodeQwen1.5-7B-Chat)n=5, k=5, temperature=0.52024.12 | 80.5 | — | — | — | |
| CodeCoRBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 80.5 | — | — | — | |
| ThinkCoder (Kimi)n=1, k=5, temperature=0.52024.12 | 79.9 | — | — | — | |
| ThinkCoder (Kimi)n=2, k=5, temperature=0.52024.12 | 79.3 | — | — | — | |
| ThinkCoder (Kimi)n=5, k=5, temperature=0.52024.12 | 79.3 | — | — | — | |
| ReflexionLLM=GPT-42024.05 | 78.7 | — | — | — | |
| AgentCoderBackbone=GPT-3.5-turbo2023.12 | 77.4 | 81.3 | — | — | |
| AgentCoderBackbone=GPT-3.5-Turbo2024.12 | 77.4 | — | — | — | |
| MapCoderBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 77.4 | — | — | — | |
| AgentCoderBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 77.4 | — | — | — | |
| AgentCoderBackbone=GPT-4-turbo2023.12 | 76.2 | 56.1 | — | — | |
| AgentCoderBackbone=GPT-4-Turbo2024.12 | 76.2 | — | — | — | |
| DirectLLM=GPT-42024.05 | 73.8 | — | — | — | |
| CodeSimLLM=LLaMa3.1-70B2025.02 | 73.8 | — | — | — | |
| Kimin=0, temperature=02024.12 | 73.4 | — | — | — | |
| CodeSimLLM=Gemma2-9B2025.02 | 72 | — | — | — | |
| CodeQwen1.5-7B-Chatn=0, temperature=02024.12 | 71 | — | — | — | |
| Self-CollaborationBackbone=GPT-42023.12 | 70.7 | 39.7 | — | — | |
| Self-CollaborationBase Model=GPT-4, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 70.7 | — | — | — | |
| MapCoderLLM=ChatGPT2024.05 | 70.1 | — | — | — | |
| MapCoderBackbone=GPT-3.5-Turbo2024.12 | 70.1 | — | — | — | |
| CodeCoTBackbone=GPT-3.5-turbo2023.12 | 69.5 | 62.8 | — | — | |
| FlowGenscrum+TestLLM=GPT-3.5, Process Model=Scrum, Incorporates=CodeT2024.03 | 67.7 | — | — | — | |
| CoTLLM=LLaMa3.1-70B2025.02 | 67.7 | — | — | — | |
| CodeTLLM=GPT-3.52024.03 | 66.9 | — | — | — | |
| FlowGenscrumLLM=GPT-3.5, Process Model=Scrum2024.03 | 65.5 | — | — | — | |
| CodeSimLLM=LLaMa3.1-8B2025.02 | 65.2 | — | — | — | |
| ReflexionLLM=LLaMa3.1-70B2025.02 | 64 | — | — | — | |
| Self-PlanningLLM=GPT-42024.05 | 62.2 | — | — | — | |
| CoTLLM=GPT-42024.05 | 61.6 | — | — | — | |
| CodeSimLLM=Mixtral8x7B2025.02 | 61.6 | — | — | — | |
| AgentCoderBackbone=Claude-instant-12023.12 | 57.9 | 106 | — | — | |
| ReflexionLLM=Gemma2-9B2025.02 | 56.7 | — | — | — | |
| Self-collaborationLLM=ChatGPT2024.05 | 56.1 | — | — | — | |
| Self-CollaborationBackbone=GPT-3.5-turbo2023.12 | 56.1 | 31.4 | — | — | |
| DirectLLM=Gemma2-9B2025.02 | 56.1 | — | — | — | |
| ReflexionLLM=GPT-3.52024.03 | 55.7 | — | — | — | |
| CoTLLM=ChatGPT2024.05 | 55.5 | — | — | — | |
| AgentCoderBackbone=PaLM Coder2023.12 | 55.5 | 51.6 | — | — | |
| Few-ShotBackbone=GPT-3.5-turbo2023.12 | 54.9 | 28.6 | — | — | |
| Few-ShotBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 54.9 | — | — | — | |
| INTERVENORBackbone=GPT-3.5-turbo2023.12 | 54.8 | 28.3 | — | — | |
| Self-EditBackbone=GPT-3.5-turbo2023.12 | 54.3 | 27.2 | — | — | |
| RAPBackbone=GPT-3.5-turbo2023.12 | 52.4 | 22.7 | — | — | |
| CodeX+CodeTBackbone=CodeX (175B), Setting=Augmented2023.12 | 51.7 | — | — | — | |
| AnalogicalLLM=ChatGPT2024.05 | 50.6 | — | — | — | |
| GPT-4Backbone=GPT-4, Setting=Zero-Shot2023.12 | 50.6 | — | — | — | |
| ReflexionBackbone=GPT-3.5-turbo2023.12 | 50.6 | 18.5 | — | — | |
| GPT-4Evaluation Strategy=Zero-Shot2026.03 | 50.6 | — | — | — | |
| ReflexionBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 50.6 | — | — | — | |
| DirectLLM=LLaMa3.1-70B2025.02 | 50.6 | — | — | — | |
| ReflexionLLM=ChatGPT2024.05 | 49.4 | — | — | — | |
| ReActBackbone=GPT-3.5-turbo2023.12 | 49.4 | 15.7 | — | — | |
| ReActBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 49.4 | — | — | — | |
| AnalogicalLLM=GPT-42024.05 | 48.8 | — | — | — | |
| GPT-4-turboBackbone=GPT-4-turbo, Setting=Zero-Shot2023.12 | 48.8 | — | — | — | |
| Self-PlaningBackbone=GPT-3.5-turbo2023.12 | 48.8 | 14.3 | — | — | |
| GPT-4-turboEvaluation Strategy=Zero-Shot2026.03 | 48.8 | — | — | — | |
| Self-PlanningLLM=ChatGPT2024.05 | 46.2 | — | — | — | |
| Self-debuggingBackbone=GPT-3.5-turbo2023.12 | 45.8 | 7.3 | — | — | |
| GPT-3.5-turboBackbone=GPT-3.5-turbo, Setting=Zero-Shot2023.12 | 42.7 | — | — | — | |
| ToTBackbone=GPT-3.5-turbo2023.12 | 42.7 | 0 | — | — | |
| GPT-3.5-turboEvaluation Strategy=Zero-Shot2026.03 | 42.7 | — | — | — | |
| CoTLLM=Mixtral8x7B2025.02 | 42.1 | — | — | — | |
| CoTLLM=LLaMa3.1-8B2025.02 | 42.1 | — | — | — | |
| DirectLLM=LLaMa3.1-8B2025.02 | 38.4 | — | — | — | |
| DirectLLM=ChatGPT2024.05 | 37.2 | — | — | — | |
| CoTBackbone=GPT-3.5-turbo2023.12 | 37.2 | -12.9 | — | — | |
| CoTBase Model=GPT-3.5-turbo, Evaluation Strategy=Agentic and Prompting Strategies2026.03 | 37.2 | — | — | — | |
| PaLM CoderBackbone=PaLM Coder, Setting=Zero-Shot2023.12 | 36.6 | — | — | — | |
| ReflexionLLM=Mixtral8x7B2025.02 | 32.9 | — | — | — | |
| CodeXBackbone=CodeX (175B), Setting=Zero-Shot2023.12 | 31.7 | — | — | — | |
| ReflexionLLM=LLaMa3.1-8B2025.02 | 31.1 | — | — | — | |
| Claude-instant-1Backbone=Claude-instant-1, Setting=Zero-Shot2023.12 | 28.1 | — | — | — | |
| Claude-instant-1Evaluation Strategy=Zero-Shot2026.03 | 28.1 | — | — | — | |
| CoTLLM=Gemma2-9B2025.02 | 26.2 | — | — | — | |
| StarCoderBackbone=StarCoder (15.5B), Setting=Zero-Shot2023.12 | 25.6 | — | — | — | |
| CodeGen-MonoBackbone=CodeGen-Mono (16.1B), Setting=Zero-Shot2023.12 | 25 | — | — | — |