Open-domain Question Answering on HotpotQA (test)
50.18Accuracy (Exact Match)Clean
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CleanLLM Backbone=GPT-4o, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 50.18 | — | — | — | |
| PoisonedRAG-BBLLM Backbone=GPT-4o, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 49.33 | — | — | 16.71 | |
| The RAG ParadoxLLM Backbone=GPT-4o, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 42.15 | — | — | 15.25 | |
| Joint-GCGLLM Backbone=GPT-4o, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 39.76 | — | — | 34.75 | |
| Deepseek-V3 (671B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=Deepseek-V3 (671B)2025.08 | 34.62 | 45.69 | 50 | — | |
| CORE (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=CORE (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 34.12 | 45 | 48 | — | |
| Top10 DocumentsDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 33.95 | 44.88 | 1,471 | — | |
| Deepseek-V3 (671B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=Deepseek-V3 (671B)2025.08 | 33.59 | 44.83 | 48 | — | |
| CORE (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=CORE (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 33.44 | 44.54 | 45 | — | |
| Top5 DocumentsDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 32.99 | 43.69 | 737 | — | |
| Top3 DocumentsDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 31.64 | 41.87 | 442 | — | |
| RECOMP (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=RECOMP (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 29.87 | 41.21 | 49 | — | |
| RECOMP (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=RECOMP (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 29.82 | 41.05 | 55 | — | |
| Top1 DocumentDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 29.2 | 38.93 | 147 | — | |
| llama3.2-1BDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=llama3.2-1B2025.08 | 26.63 | 36.61 | 56 | — | |
| llama3.2-1BDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=llama3.2-1B2025.08 | 26.48 | 36.39 | 58 | — | |
| No RetrievalDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=No Retrieval2025.08 | 21.05 | 29.48 | 0 | — | |
| EVE-Agent2026.05 | 20.9 | — | — | — | |
| CRCPLLM Backbone=DeepSeek-R1, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 19.78 | — | — | 95.01 | |
| CRCPLLM Backbone=LLaMA-2-7B, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 18.93 | — | — | 90.12 | |
| CRCPLLM Backbone=Vicuna-7B, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 18.69 | — | — | 90.87 | |
| CRCPLLM Backbone=GPT-4o, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 17.39 | — | — | 91.76 | |
| CRCPLLM Backbone=Qwen3-Max, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 16.17 | — | — | 92.75 | |
| CRCPLLM Backbone=Qwen2.5-7B, Pipeline=Chunking, Retrieval, and Reranking2026.06 | 14.92 | — | — | 91.36 | |
| Dr. Zero2026.05 | 11 | — | — | — | |
| Initialsearch=disabled2026.05 | 5.7 | — | — | — | |
| Initialsearch=enabled2026.05 | 2.9 | — | — | — |