Image2Code Generation on Synthetic Eval Dataset State
87.93BLEULlama-VL-TUG
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Llama-VL-TUG2026.01 | 87.93 | 87.9324 | 79.23 | 91.4462 | 0.4082 | 67.38 | |
| GPT-4o-mini2026.01 | 51 | 51.0173 | 51.67 | 61.402 | 0.0565 | 50.49 | |
| Llama3.2-11B-Vision-Instruct2026.01 | 46.43 | 46.4642 | 46.17 | 58.4349 | -0.0451 | 46.39 | |
| Qwen2.5-VL-7B-Instruct2026.01 | 26.09 | 26.4496 | 31.61 | 45.0595 | -0.3306 | 45.07 | |
| Gemma3-12B-Instruction-Tuned2026.01 | 20.88 | 21.3466 | 27.81 | 40.4843 | -0.3614 | 43.56 | |
| MiniCPM-V-2-62026.01 | 0.03 | 0.3201 | 11.23 | 3.5678 | -1.396 | 0.91 |