Claim Extraction on MMCE dataset
3.25Reference-Based ScoreMICE
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MICEModel=Gemini 2.0 Flash, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.25 | 54.9 | 74.6 | 98.8 | 100 | |
| MLLM with ICLModel=Qwen2.5 VL 32B Instruct, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.24 | 71.8 | 83.6 | 98.4 | 100 | |
| MLLM with ICLModel=GPT 4o Mini, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.22 | 69.9 | 82.4 | 98.1 | 100 | |
| MLLM with ICLModel=Gemini 2.0 Flash, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.21 | 70.5 | 82.6 | 98.4 | 100 | |
| MLLMModel=GPT 4o Mini, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.15 | 74.4 | 85.2 | 97.8 | 99.9 | |
| MICEModel=Qwen2.5 VL 32B Instruct, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.15 | 52.7 | 78.9 | 98.1 | 100 | |
| MLLMModel=Qwen2.5 VL 32B Instruct, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.14 | 77.3 | 86.6 | 98.5 | 100 | |
| MICEModel=GPT 4o Mini, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.13 | 65.8 | 83.9 | 98.5 | 100 | |
| MLLMModel=Gemini 2.0 Flash, Evaluation Protocol=LLM-based Evaluation2026.02 | 3.11 | 75.5 | 85.4 | 97.9 | 99.9 | |
| MLLM (text input only)Model=Qwen2.5 VL 32B Instruct, Evaluation Protocol=LLM-based Evaluation2026.02 | 2.85 | 77.7 | 86.3 | 96.4 | 99.6 | |
| MLLM (text input only)Model=GPT 4o Mini, Evaluation Protocol=LLM-based Evaluation2026.02 | 2.83 | 80.6 | 88 | 97.4 | 99.7 | |
| MLLM (text input only)Model=Gemini 2.0 Flash, Evaluation Protocol=LLM-based Evaluation2026.02 | 2.8 | 80 | 85.9 | 96.9 | 99.7 |