Concept Induction on Bongard-OpenWorld 44 problems
91Binary AccuracyHuman performance
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Human performance2026.07 | 91 | — | |
| ReVisITbackbone=GPT-4.1, task-specific tuning=zero, scaffold=turns layer only2026.07 | 72.7 | 26.1 | |
| GPT-4V (published)source=original benchmark, dataset_scope=Full 200-problem set2026.07 | 64 | — | |
| GPT-4.1 (vanilla)backbone=GPT-4.1, few-shot examples=none2026.07 | 46.6 | — |