Code Generation on LiveCodeBench (train)
19.63Metric m ScoreAutoJudge-F
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AutoJudge-FTarget Model=Llama-3.1-8B, Draft Model=Llama-3.2-1B2025.09 | 19.63 | 2.1 | -6.5 | |
| AutoJudge-RTarget Model=Llama-3.1-8B, Draft Model=Llama-3.2-1B2025.09 | 15.78 | 0.7 | -7.9 | |
| Top-kTarget Model=Llama-3.1-8B, Draft Model=Llama-3.2-1B2025.09 | 13.36 | 0.7 | -7.9 | |
| SelfJudge-FTarget Model=Llama-3.1-8B, Draft Model=Llama-3.2-1B2025.09 | 9.52 | 10 | 1.4 | |
| SelfJudge-RTarget Model=Llama-3.1-8B, Draft Model=Llama-3.2-1B2025.09 | 8.12 | 9.7 | 1.1 | |
| SDTarget Model=Llama-3.1-8B, Draft Model=Llama-3.2-1B2025.09 | 7.88 | 8.6 | 0 |