Code Generation on HumanEval Llama-3-70B (test)
26.3QD-ScoreQD-LLM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| QD-LLMBackbone=Llama-3-70B-Instruct, Budget=5002026.05 | 26.3 | 26.2 | 41 | 94 | |
| CMA-ME (ad.)Backbone=Llama-3-70B-Instruct, Budget=5002026.05 | 19.8 | 19.7 | 30 | — | |
| QDAIFBackbone=Llama-3-70B-Instruct, Budget=5002026.05 | 18.6 | 18.5 | 28 | — | |
| EvoPromptBackbone=Llama-3-70B-Instruct, Budget=5002026.05 | 17.2 | 17.1 | 21 | — | |
| Best-of-N+MMRBackbone=Llama-3-70B-Instruct, N=20, Budget=5002026.05 | 16.4 | 16.3 | 24 | — | |
| Diverse BeamBackbone=Llama-3-70B-Instruct, lambda=0.5, Budget=5002026.05 | 15.1 | 15 | 21 | — | |
| Nucleus Samp.Backbone=Llama-3-70B-Instruct, p=0.95, Budget=5002026.05 | 14.2 | 14.1 | 19 | — | |
| Vanilla MEBackbone=Llama-3-70B-Instruct, Budget=5002026.05 | 13.8 | 13.7 | 18 | — |