Compliance evaluation on Compliance Evaluation Suite Average
89.99AccuracyGPT-4o-mini (ContextLens)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o-mini (ContextLens)Evaluation Type=ContextLens2026.04 | 89.99 | 87.89 | |
| GPT-4o (ContextLens)Evaluation Type=ContextLens2026.04 | 88.4 | 86.29 | |
| ContextReasoner-7B-PPOEvaluation Type=Fine-tuned2026.04 | 88.25 | 86.49 | |
| ContextReasoner-7B-SFTEvaluation Type=Fine-tuned2026.04 | 88.02 | 86.05 | |
| Llama-3.1-8B-Instruct (ContextLens)Evaluation Type=ContextLens2026.04 | 87.73 | 85.67 | |
| DeepSeek-R1 (671B)Evaluation Type=Long CoT2026.04 | 87.22 | 85.74 | |
| o3-miniEvaluation Type=Long CoT2026.04 | 86.76 | 85.59 | |
| Gemini-2.5-Flash (ContextLens)Evaluation Type=ContextLens2026.04 | 86.14 | 82.7 | |
| Gemini-2.5-FlashEvaluation Type=Long CoT2026.04 | 86.02 | 84.34 | |
| GPT-5Evaluation Type=RAG2026.04 | 85.79 | 84.09 | |
| o3-mini (ContextLens)Evaluation Type=ContextLens2026.04 | 85.22 | 82.79 | |
| o3-miniEvaluation Type=RAG2026.04 | 84.33 | 81.3 | |
| GPT-5 (ContextLens)Evaluation Type=ContextLens2026.04 | 84.09 | 82.1 | |
| GPT-4oEvaluation Type=Direct2026.04 | 83.76 | 81.71 | |
| DeepSeek-R1 (671B) (ContextLens)Evaluation Type=ContextLens2026.04 | 83.6 | 80.2 | |
| GPT-4o-miniEvaluation Type=RAG2026.04 | 82.79 | 81.84 | |
| GPT-4o-miniEvaluation Type=Direct2026.04 | 81.68 | 80.11 | |
| GPT-5Evaluation Type=Direct2026.04 | 81.2 | 78.2 | |
| OpenThinker-7BEvaluation Type=Direct2026.04 | 78.88 | 75.77 | |
| Llama-3.1-8B-InstructEvaluation Type=Direct2026.04 | 72.78 | 71.71 | |
| Qwen2.5-7B-Instruct (ContextLens)Evaluation Type=ContextLens2026.04 | 70.61 | 66.73 | |
| Qwen2.5-7B-InstructEvaluation Type=Direct2026.04 | 67.35 | 63.63 |