Evidence Conflict Diagnosis in RAG on Original controlled benchmark (300 samples)
100CoverageNaive
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| NaiveLLM=Qwen2.5-7B-Instruct, temperature=0, decoding=deterministic2026.06 | 100 | 87 | 86.67 | 95.79 | 74 | 8 | |
| Evidence-normalizedLLM=Qwen2.5-7B-Instruct, temperature=0, decoding=deterministic2026.06 | 100 | 89 | 88.67 | 97.66 | 75.33 | 10.67 | |
| Rule-only extractortype=diagnostic, description=measures template dependence2026.06 | 100 | 100 | 100 | 100 | 100 | 0 | |
| X-MADAM-RAGLLM=Qwen2.5-7B-Instruct, temperature=0, decoding=deterministic2026.06 | 100 | 96.67 | 97.67 | 98.04 | 100 | 2 | |
| Oracle extractiontype=diagnostic, mode=non-deployable, input=uses privileged supported_answer metadata2026.06 | 100 | 100 | 100 | 100 | 100 | 0 |