Challenging Cases Understanding on Vibe-Eval
63.1AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4o2024.08 | 63.1 | |
| GPT-4VVersion=V-Preview2024.08 | 57.9 | |
| LLaVA-OneVision-7BModel Scale=7B2024.08 | 51.7 | |
| LLaVA-OneVision-72BModel Scale=72B2024.08 | 50.7 | |
| LLaVA-OneVision-0.5BModel Scale=0.5B2024.08 | 33.8 |