Compliance evaluation on EU AI Act synthetic cases
89.16AccuracyGemini-2.5-Flash (ContextLens)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-2.5-Flash (ContextLens)Evaluation Type=ContextLens2026.04 | 89.16 | 88.5 | |
| GPT-4o-mini (ContextLens)Evaluation Type=ContextLens2026.04 | 88.11 | 87.53 | |
| GPT-4o (ContextLens)Evaluation Type=ContextLens2026.04 | 87.5 | 87.02 | |
| Llama-3.1-8B-Instruct (ContextLens)Evaluation Type=ContextLens2026.04 | 86.6 | 86.21 | |
| GPT-5 (ContextLens)Evaluation Type=ContextLens2026.04 | 86.33 | 89.05 | |
| DeepSeek-R1 (671B) (ContextLens)Evaluation Type=ContextLens2026.04 | 86.31 | 85.87 | |
| o3-mini (ContextLens)Evaluation Type=ContextLens2026.04 | 85.09 | 84.48 | |
| o3-miniEvaluation Type=Long CoT2026.04 | 84.83 | 84.36 | |
| ContextReasoner-7B-SFTEvaluation Type=Fine-tuned2026.04 | 84.33 | 83.73 | |
| ContextReasoner-7B-PPOEvaluation Type=Fine-tuned2026.04 | 84.33 | 83.65 | |
| DeepSeek-R1 (671B)Evaluation Type=Long CoT2026.04 | 84 | 83.3 | |
| GPT-5Evaluation Type=RAG2026.04 | 80.33 | 79.8 | |
| Gemini-2.5-FlashEvaluation Type=Long CoT2026.04 | 80 | 79.48 | |
| o3-miniEvaluation Type=RAG2026.04 | 77.75 | 75.14 | |
| GPT-4oEvaluation Type=Direct2026.04 | 75.33 | 73.39 | |
| GPT-5Evaluation Type=Direct2026.04 | 74.83 | 72.07 | |
| GPT-4o-miniEvaluation Type=RAG2026.04 | 74.33 | 73.62 | |
| GPT-4o-miniEvaluation Type=Direct2026.04 | 71.16 | 70.92 | |
| OpenThinker-7BEvaluation Type=Direct2026.04 | 70.5 | 69.72 | |
| Llama-3.1-8B-InstructEvaluation Type=Direct2026.04 | 58.67 | 59.36 | |
| Qwen2.5-7B-Instruct (ContextLens)Evaluation Type=ContextLens2026.04 | 54.63 | 53.88 | |
| Qwen2.5-7B-InstructEvaluation Type=Direct2026.04 | 46.33 | 41.89 |