Text-to-image in-context learning on CoBSAT (all samples)
0.318CLIP SimilarityToT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ToTReasoning strategy=Tree-of-Thoughts (structured stage exploration), Reasoning backbone=SEED-LLaMA, Image generator=Stable Diffusion, Reasoning depth=4, Pruning threshold=0.082026.07 | 0.318 | 0.775 | |
| CoTReasoning strategy=Chain-of-Thought (linear reasoning trace), Reasoning backbone=SEED-LLaMA, Image generator=Stable Diffusion2026.07 | 0.302 | 0.547 | |
| BaselineReasoning strategy=Direct prompt construction, Reasoning backbone=SEED-LLaMA, Image generator=Stable Diffusion2026.07 | 0.287 | 0.508 |