RTL code understanding on RTL code understanding (test)
13.96BLEU-4DeepRTL2
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| DeepRTL2Base Model=DeepSeek-Coder2025.05 | 13.96 | 37.93 | 20.73 | 34.34 | 34.74 | 82 | 61.6 | |
| DeepRTL2Base Model=Llama-3.12025.05 | 13.84 | 37.97 | 20.69 | 34.42 | 34.75 | 81.3 | 60.3 | |
| DeepRTL2Version=1st, Base Model=DeepSeek-Coder2025.05 | 13.53 | 37.52 | 19.68 | 34.68 | 33.28 | 81.4 | 61.2 | |
| DeepRTL2Version=1st, Base Model=Llama-3.12025.05 | 13.34 | 37.74 | 19.54 | 34.76 | 33.46 | 79.8 | 59.4 | |
| DeepRTLModel Scale=220m2025.05 | 13.06 | 37.56 | 19.85 | 34.72 | 34.37 | 80.6 | 60 | |
| DeepRTLModel Scale=16b2025.05 | 12.85 | 37.43 | 19.34 | 34.63 | 33.09 | 80.2 | 59.7 | |
| DeepRTL2Version=1st, Evaluation Protocol=Direct, Base Model=DeepSeek-Coder2025.05 | 12.07 | 36.37 | 17.78 | 33.78 | 28.56 | 76.7 | 60.2 | |
| DeepRTL2Version=1st, Evaluation Protocol=Direct, Base Model=Llama-3.12025.05 | 11.28 | 34.29 | 16.35 | 33.63 | 27.73 | 75.4 | 58 | |
| GPT-4o2025.05 | 4.59 | 29.26 | 11.48 | 25.74 | 22.78 | 76.1 | 54.9 | |
| o1-preview2025.05 | 3.73 | 28 | 10.39 | 24.98 | 20.48 | 74.8 | 53.5 | |
| GPT-3.52025.05 | 3.34 | 28.2 | 10.46 | 25.11 | 20.36 | 74 | 51 | |
| CodeV-DeepSeek2025.05 | 3.05 | 25.14 | 9.78 | 23.25 | 20.23 | 70.5 | 49.5 | |
| CodeV-CodeQwen2025.05 | 2.8 | 24.91 | 8.27 | 22.75 | 21.07 | 74.7 | 49.9 | |
| Llama-3.12025.05 | 2.68 | 25.37 | 10.39 | 23.75 | 17.16 | 73 | 43 | |
| DeepSeek-Coder2025.05 | 2.56 | 24.52 | 7.72 | 22.45 | 22.83 | 75.6 | 57.1 |