Code Generation on Code Benchmarks LCBv6 & MBPP+
32.8LCBv6 ScoreEVOTD
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| EVOTDBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=EVOTD, Decoding Strategy=Greedy (Pass@1)2026.05 | 32.8 | 65.3 | 49.1 | |
| Agent0Backbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=Agent0, Decoding Strategy=Greedy (Pass@1)2026.05 | 32 | 71.2 | 51.6 | |
| Evol-InstructBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=Evol-Instruct, Decoding Strategy=Greedy (Pass@1)2026.05 | 31.9 | 59.8 | 45.9 | |
| Qwen3-8B-BaseBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=None, Decoding Strategy=Greedy (Pass@1)2026.05 | 29.6 | 58.3 | 44 | |
| SPIRALBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=SPIRAL, Decoding Strategy=Greedy (Pass@1), Checkpoint=Spiral-Qwen3-8B-Multi-Env2026.05 | 28.8 | 58.7 | 43.8 | |
| Agent0Backbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=Agent0, Decoding Strategy=Greedy (Pass@1)2026.05 | 27.6 | 63.5 | 45.6 | |
| SPIRALBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=SPIRAL, Decoding Strategy=Greedy (Pass@1), Checkpoint=Spiral-Qwen3-4B-Multi-Env2026.05 | 27.3 | 61.4 | 44.4 | |
| EVOTDBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=EVOTD, Decoding Strategy=Greedy (Pass@1)2026.05 | 26.4 | 58.5 | 42.5 | |
| Evol-InstructBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=Evol-Instruct, Decoding Strategy=Greedy (Pass@1)2026.05 | 23.2 | 52.1 | 37.7 | |
| Qwen3-4B-BaseBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=None, Decoding Strategy=Greedy (Pass@1)2026.05 | 20.6 | 53.7 | 37.2 | |
| EVOTDBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=EVOTD, Decoding Strategy=Greedy (Pass@1)2026.05 | 17.6 | 62.4 | 40 | |
| SPIRALBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=SPIRAL, Decoding Strategy=Greedy (Pass@1)2026.05 | 16.7 | 60.6 | 38.7 | |
| LLaMA-3.1-8B-InstructBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=None, Decoding Strategy=Greedy (Pass@1)2026.05 | 16.6 | 59.8 | 38.2 | |
| Agent0Backbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Agent0, Decoding Strategy=Greedy (Pass@1)2026.05 | 16.1 | 61.1 | 38.6 | |
| Evol-InstructBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Evol-Instruct, Decoding Strategy=Greedy (Pass@1)2026.05 | 15.7 | 61.1 | 38.4 | |
| Agent0Backbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Agent0, Decoding Strategy=Greedy (Pass@1)2026.05 | 11.6 | 31.5 | 21.6 | |
| Evol-InstructBackbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Evol-Instruct, Decoding Strategy=Greedy (Pass@1)2026.05 | 11.1 | 25.4 | 18.3 | |
| LLaMA-3.2-3B-InstructBackbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=None, Decoding Strategy=Greedy (Pass@1)2026.05 | 11 | 25.9 | 18.5 | |
| EVOTDBackbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=EVOTD, Decoding Strategy=Greedy (Pass@1)2026.05 | 10.7 | 42.1 | 26.4 |