Compliance evaluation on GDPR real-life contexts
92.19AccuracyGPT-4o-mini
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o-miniEvaluation Type=Direct2026.04 | 92.19 | 89.29 | |
| GPT-4oEvaluation Type=Direct2026.04 | 92.19 | 90.02 | |
| ContextReasoner-7B-PPOEvaluation Type=Fine-tuned2026.04 | 92.19 | 89.33 | |
| Gemini-2.5-FlashEvaluation Type=Long CoT2026.04 | 92.03 | 89.19 | |
| GPT-4o-mini (ContextLens)Evaluation Type=ContextLens2026.04 | 91.87 | 88.25 | |
| ContextReasoner-7B-SFTEvaluation Type=Fine-tuned2026.04 | 91.71 | 88.36 | |
| GPT-4o-miniEvaluation Type=RAG2026.04 | 91.24 | 90.05 | |
| GPT-5Evaluation Type=RAG2026.04 | 91.24 | 88.37 | |
| o3-miniEvaluation Type=RAG2026.04 | 90.92 | 87.46 | |
| DeepSeek-R1 (671B)Evaluation Type=Long CoT2026.04 | 90.44 | 88.17 | |
| GPT-4o (ContextLens)Evaluation Type=ContextLens2026.04 | 89.31 | 85.55 | |
| Llama-3.1-8B-Instruct (ContextLens)Evaluation Type=ContextLens2026.04 | 88.85 | 85.13 | |
| o3-miniEvaluation Type=Long CoT2026.04 | 88.69 | 86.82 | |
| Qwen2.5-7B-InstructEvaluation Type=Direct2026.04 | 88.37 | 85.36 | |
| GPT-5Evaluation Type=Direct2026.04 | 87.57 | 84.33 | |
| OpenThinker-7BEvaluation Type=Direct2026.04 | 87.26 | 81.82 | |
| Llama-3.1-8B-InstructEvaluation Type=Direct2026.04 | 86.88 | 84.06 | |
| o3-mini (ContextLens)Evaluation Type=ContextLens2026.04 | 85.35 | 81.09 | |
| Qwen2.5-7B-Instruct (ContextLens)Evaluation Type=ContextLens2026.04 | 85.19 | 79.58 | |
| Gemini-2.5-Flash (ContextLens)Evaluation Type=ContextLens2026.04 | 83.12 | 76.89 | |
| GPT-5 (ContextLens)Evaluation Type=ContextLens2026.04 | 81.85 | 75.15 | |
| DeepSeek-R1 (671B) (ContextLens)Evaluation Type=ContextLens2026.04 | 80.89 | 74.52 |