Text-to-Image Generation on DrawBench v1 (test)
0.885VQAScoreRAISE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RAISEAvg. #Samples Generated=21.2, Avg. #Calls VLM=8.6, Rounds=≤4, Base Model=FLUX.1-dev2026.02 | 0.885 | 1.15 | 0.305 | |
| RAISEAvg. #Samples Generated=16, Avg. #Calls VLM=6, Rounds=2, Base Model=FLUX.1-dev2026.02 | 0.876 | 1.15 | 0.305 | |
| RAISEAvg. #Samples Generated=8, Avg. #Calls VLM=3, Rounds=1, Base Model=FLUX.1-dev2026.02 | 0.868 | 1.13 | 0.304 | |
| ReflectionFlowAvg. #Samples Generated=16, Avg. #Calls VLM=32, Base Model=FLUX.1-dev2026.02 | 0.844 | 1.13 | 0.302 | |
| ReflectionFlowAvg. #Samples Generated=32, Avg. #Calls VLM=64, Base Model=FLUX.1-dev2026.02 | 0.844 | 1.1 | 0.302 | |
| ReflectionFlowAvg. #Samples Generated=8, Avg. #Calls VLM=16, Base Model=FLUX.1-dev2026.02 | 0.839 | 1.08 | 0.301 | |
| T2I-CopilotAvg. #Samples Generated=3.9, Avg. #Calls VLM=7.9, Base Model=FLUX.1-dev2026.02 | 0.822 | 0.97 | 0.3 | |
| T2I-CopilotAvg. #Samples Generated=6.6, Avg. #Calls VLM=13.1, Base Model=FLUX.1-dev2026.02 | 0.82 | 0.96 | 0.299 | |
| T2I-CopilotAvg. #Samples Generated=11.2, Avg. #Calls VLM=22.3, Base Model=FLUX.1-dev2026.02 | 0.82 | 0.94 | 0.298 | |
| FLUX.1-devAvg. #Samples Generated=1, Avg. #Calls VLM=02026.02 | 0.778 | 1.06 | 0.298 |