Deep Semantic Inference on DeepEval
70.9AccuracyInternVL3
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL3#Params=8B, CoT=false2026.01 | 70.9 | |
| InternVL3#Params=8B, CoT=true2026.01 | 67.8 | |
| MoCoT (Ours)#Params=3B2026.01 | 64.3 | |
| Qwen2.5-VL#Params=7B, CoT=true2026.01 | 63.3 | |
| Qwen2.5-VL#Params=7B, CoT=false2026.01 | 58.3 | |
| Qwen2.5-VL#Params=3B, CoT=false2026.01 | 55.8 | |
| Qwen2.5-VL#Params=3B, CoT=true2026.01 | 48.7 | |
| Gemma-3#Params=4B, CoT=true2026.01 | 46.2 | |
| InternVL2.5#Params=2B, CoT=false2026.01 | 45.7 | |
| InternVL2.5#Params=2B, CoT=true2026.01 | 42.7 | |
| XComposer-2.5#Params=7B, CoT=true2026.01 | 36.2 | |
| Phi-3.5#Params=4B, CoT=false2026.01 | 35.7 | |
| Gemma-3#Params=4B, CoT=false2026.01 | 35.2 | |
| XComposer-2.5#Params=7B, CoT=false2026.01 | 34.2 | |
| Ovis2#Params=2B, CoT=true2026.01 | 32.2 | |
| Ovis2#Params=2B, CoT=false2026.01 | 31.7 | |
| Phi-3.5#Params=4B, CoT=true2026.01 | 30.7 | |
| LLaVA-1.6#Params=7B, CoT=true2026.01 | 29.7 | |
| Mono#Params=2B, CoT=true2026.01 | 20.1 | |
| LLaVA-1.6#Params=7B, CoT=false2026.01 | 17.1 | |
| Mono#Params=2B, CoT=false2026.01 | 14.1 |