Image2Code Generation on Synthetic Eval Dataset (Class)
0.9465BLEULlama-VL-TUG
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Llama-VL-TUG2026.01 | 0.9465 | 94.649 | 0.9523 | 99.0225 | 0.289 | 0.8181 | |
| GPT-4o-mini2026.01 | 0.7975 | 79.7497 | 0.6845 | 89.2833 | 0.3699 | 0.683 | |
| Gemma3-12B-Instruction-Tuned2026.01 | 0.6328 | 63.2811 | 0.5355 | 75.4252 | 0.1044 | 0.6053 | |
| Llama3.2-11B-Vision-Instruct2026.01 | 0.6317 | 63.1853 | 0.5364 | 75.3705 | 0.1994 | 0.5912 | |
| Qwen2.5-VL-7B-Instruct2026.01 | 0.5923 | 59.2725 | 0.5419 | 73.4314 | 0.0562 | 0.555 | |
| MiniCPM-V-2-62026.01 | 0.0001 | 0.2943 | 0.0604 | 6.2279 | -1.2576 | 0.046 |