Field-level extraction on Modified insurance claims dataset
97.6PrecisionGemini 3.1 Pro preview
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini 3.1 Pro preview2026.04 | 97.6 | 96.74 | 97.17 | |
| OpenAI GPT-5.4Reasoning effort configuration=advanced reasoning2026.04 | 97.59 | 94.83 | 96.19 | |
| xmemoryLLM judge-in-the-loop protocol=true2026.04 | 97.39 | 97.67 | 97.53 | |
| OpenAI GPT-5.5Reasoning effort configuration=high reasoning effort2026.04 | 97.18 | 95.6 | 96.39 | |
| OpenAI GPT-5.52026.04 | 97.05 | 95 | 96.01 | |
| Anthropic Opus 4.72026.04 | 96.72 | 96.41 | 96.56 | |
| xmemoryprotocol=single-model extraction pipeline2026.04 | 96.72 | 97.21 | 96.97 | |
| OpenAI GPT-5.42026.04 | 96.22 | 94.11 | 95.15 | |
| Anthropic Sonnet 4.62026.04 | 95.22 | 97.62 | 96.4 |