Joint Assembly Evaluation on BC-Bench Best Object Performance
41.63Step-Wise Success RateBrick-Composer
Evaluation Results
| Method | Links | |
|---|---|---|
| Brick-ComposerBackbone=Qwen-3-8B-VL, Evaluation Protocol=Integrated learning2026.06 | 41.63 | |
| Brick-ComposerBackbone=Gemma-3-12B, Evaluation Protocol=Integrated learning2026.06 | 18.75 | |
| World Feedback (L)Backbone=Qwen-3-8B-VL, Evaluation Protocol=Simulator-generated feedback2026.06 | 15.36 | |
| World Feedback (L)Backbone=Gemma-3-12B, Evaluation Protocol=Simulator-generated feedback2026.06 | 9.07 | |
| Designer SupervisionBackbone=Qwen-3-8B-VL, Evaluation Protocol=Human supervision2026.06 | 8.92 | |
| Qwen-3-VL-8BEvaluation Protocol=Zero-shot2026.06 | 4.44 | |
| Qwen-3.5-VL-27BEvaluation Protocol=Zero-shot2026.06 | 4.44 | |
| Designer SupervisionBackbone=Gemma-3-12B, Evaluation Protocol=Human supervision2026.06 | 4.44 | |
| World Feedback (P)Backbone=Qwen-3-8B-VL, Evaluation Protocol=Simulator feedback2026.06 | 4.44 | |
| Gemma-3-12BEvaluation Protocol=Zero-shot2026.06 | 2.22 | |
| GPT-5.4Evaluation Protocol=Zero-shot2026.06 | 2.22 | |
| World Feedback (P)Backbone=Gemma-3-12B, Evaluation Protocol=Simulator feedback2026.06 | 2.22 | |
| InternVL-3.5-8BEvaluation Protocol=Zero-shot2026.06 | 0 |