Context Understanding on CulturalBench
0.904Easy AccuracyGemini-3-Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-3-ProCategory=MLLM, Note=direct2026.02 | 0.904 | 0.9 | |
| GPT-5.2Category=MLLM, Note=direct2026.02 | 0.883 | 0.844 | |
| TwC (Ours) - Img & TxtCategory=Think with Comic, Note=G-t-R2026.02 | 0.883 | 0.822 | |
| Claude-Sonnet 4.5Category=MLLM, Note=direct2026.02 | 0.872 | 0.765 | |
| DeepSeek-R1Category=Reasoning LLM, Note=CoT2026.02 | 0.872 | 0.851 | |
| Qwen3-235B-A22BCategory=Reasoning LLM, Note=CoT2026.02 | 0.831 | 0.825 | |
| TwC (Ours) - Only ImageCategory=Think with Comic, Note=direct2026.02 | 0.7 | 0.805 | |
| TWI-1-Generated PhotoCategory=Think with Image, Note=G-t-R2026.02 | 0.697 | 0.714 | |
| Sora 2Category=Think with Video, Note=V-o-T, Evaluation Details=50 sampled instances2026.02 | 0.6 | 0.7 | |
| DREAMLLMCategory=Think with Image, Note=G-t-R2026.02 | 0.523 | 0.428 |