Code Repair on SWE-bench Lite
0.77rADARUBRIC-DA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ADARUBRIC-DABackbone=Llama-3.1-8B-Instruct2026.03 | 0.77 | 0.84 | 14.7 | |
| ADARUBRIC-WMBackbone=Llama-3.1-8B-Instruct2026.03 | 0.72 | 0.82 | 12.4 | |
| GPT-4 DirectBackbone=Llama-3.1-8B-Instruct2026.03 | 0.59 | 0.68 | 9.8 | |
| PrometheusBackbone=Llama-3.1-8B-Instruct2026.03 | 0.56 | 0.7 | 9.1 | |
| G-Eval (GPT-4o)Backbone=Llama-3.1-8B-Instruct2026.03 | 0.51 | 0.63 | 8.2 |