Description2Code Generation on Synthetic Eval Dataset Flowchart 1.0
0.9656BLEULlama-VL-TUG
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Llama-VL-TUG2026.01 | 0.9656 | 96.5596 | 0.9889 | 98.5066 | 0.6291 | 0.9789 | |
| GPT-4o-mini2026.01 | 0.1756 | 17.5553 | 0.4283 | 29.6297 | -1.0116 | 0.1797 | |
| Gemma3-12B-Instruction-TunedModel Size=12B, Variant=Instruction-Tuned2026.01 | 0.0204 | 2.6764 | 0.1725 | 24.7587 | -0.6932 | 0.2767 | |
| Qwen2.5-VL-7B-InstructModel Size=7B, Variant=Instruct2026.01 | 0.0067 | 2.0388 | 0.1526 | 22.7352 | -0.5851 | 0.2788 | |
| Llama3.2-11B-Vision-InstructModel Size=11B, Variant=Vision-Instruct2026.01 | 0.0014 | 1.2677 | 0.1303 | 22.7094 | -0.5936 | 0.2582 | |
| MiniCPM-V-2-62026.01 | 0.0002 | 1.5939 | 0.154 | 21.7691 | -0.5673 | 0.2961 |