Multimodal Reasoning on NaturalBench
82.5AccuracyDART
Evaluation Results
| Method | Links | |
|---|---|---|
| DARTAgent=Q, M, O2025.12 | 82.5 | |
| Debate with JudgeAgent=Q, M, O2025.12 | 81.3 | |
| DARTAgent=Best Model2025.12 | 80.8 | |
| Debate with ConsensusAgent=Q, M, O2025.12 | 80.6 | |
| Debate with JudgeAgent=3×Ovis22025.12 | 80.4 | |
| Self-Consistency (5-way)Agent=QwenVL2025.12 | 80.1 | |
| Debate with ConsensusAgent=3×Ovis22025.12 | 80.1 | |
| Debate with ConsensusAgent=3×QwenVL2025.12 | 80 | |
| Debate with JudgeAgent=3×QwenVL2025.12 | 79.9 | |
| Self-Consistency (5-way)Agent=Ovis22025.12 | 79.3 | |
| CoTAgent=QwenVL2025.12 | 79.2 | |
| Self-RefinementAgent=QwenVL2025.12 | 79 | |
| CoTAgent=Ovis22025.12 | 78.7 | |
| Self-RefinementAgent=Ovis22025.12 | 78.4 | |
| Debate with ConsensusAgent=3×MiniCPM-o2025.12 | 78.3 | |
| Debate with JudgeAgent=3×MiniCPM-o2025.12 | 78.3 | |
| Self-RefinementAgent=MiniCPM-o2025.12 | 78.2 | |
| Self-Consistency (5-way)Agent=MiniCPM-o2025.12 | 78.1 | |
| CoTAgent=MiniCPM-o2025.12 | 77.9 | |
| ChameleonAgent=Qwen2025.12 | 77.2 | |
| ViperGPTAgent=Qwen2025.12 | 75.3 | |
| Self-Consistency (5-way)Agent=LLaVA 1.62025.12 | 72.4 | |
| CoTAgent=LLaVA 1.62025.12 | 70.9 | |
| Self-RefinementAgent=LLaVA 1.62025.12 | 70.6 |