Code Generation on HumanEval (Pass@1 and Throughput)
46.28Pass@1UDS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| UDSBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 46.28 | 6.81 | |
| RegularBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 45.82 | 6.24 | |
| GREATSBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 45.04 | 5.8 | |
| RHO-LossBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 43.08 | 3.53 | |
| MaxLossBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 41.34 | 6.92 | |
| MaxGradBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 40.83 | 0.88 | |
| RandomBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 40.2 | 9.35 | |
| UDSBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 30.96 | 6.41 | |
| RegularBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 29.28 | 5.77 | |
| GREATSBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 28.56 | 5.25 | |
| MaxLossBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 27.23 | 6.62 | |
| RHO-LossBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 27.18 | 3.36 | |
| MaxGradBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 26.89 | 0.8 | |
| RandomBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 26.74 | 8.92 |