Poisoning Attack on RAG on MS-MARCO (test)
0.94PoisonedRAG Scoremeta/llama-4-maverick-17b-128e-instruct
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| meta/llama-4-maverick-17b-128e-instructEvaluation protocol=strict incorrect-answer matching2026.03 | 0.94 | 0.91 | 0.89 | 0 | 0 | 0.99 | 0.05 | — | — | — | — | — | |
| qwen/qwen2-7b-instructEvaluation protocol=strict incorrect-answer matching2026.03 | 0.93 | 0.85 | 0.67 | 0 | 0 | 0.98 | 0.05 | — | — | — | — | — | |
| ibm/granite-3.3-8b-instructEvaluation protocol=strict incorrect-answer matching2026.03 | 0.93 | 0.88 | 0.88 | 0 | 0 | 0.99 | 0.06 | — | — | — | — | — | |
| qwen/qwen2.5-7b-instructEvaluation protocol=strict incorrect-answer matching2026.03 | 0.87 | 0.84 | 0.65 | 0 | 0 | 0.97 | 0.1 | — | — | — | — | — | |
| meta/llama-3.3-70b-instructEvaluation protocol=strict incorrect-answer matching2026.03 | 0.86 | 0.84 | 0.74 | 0 | 0 | 0.96 | 0.1 | — | — | — | — | — | |
| meta/llama-3.1-8b-instructEvaluation protocol=strict incorrect-answer matching2026.03 | 0.85 | 0.91 | 0.74 | 0 | 0 | 0.96 | 0.11 | — | — | — | — | — | |
| openai/gpt-oss-20bEvaluation protocol=strict incorrect-answer matching2026.03 | 0.83 | 0.84 | 0.73 | 0 | 0 | 0.9 | 0.07 | — | — | — | — | — | |
| openai/gpt-oss-120bEvaluation protocol=strict incorrect-answer matching2026.03 | 0.77 | 0.85 | 0.72 | 0 | 0 | 0.89 | 0.12 | — | — | — | — | — | |
| GASLITERetriever=Contriever, Generator=Llama-2-7B-Chat, Judge=GPT-42026.05 | — | — | — | — | — | — | — | 0.721 | 0.538 | 0.406 | 41.3 | 33.8 | |
| Joint-GCGRetriever=Contriever, Generator=Llama-2-7B-Chat, Judge=GPT-42026.05 | — | — | — | — | — | — | — | 0.764 | 0.726 | 0.583 | 173.2 | 131.4 | |
| PoisonedRAGRetriever=Contriever, Generator=Llama-2-7B-Chat, Judge=GPT-42026.05 | — | — | — | — | — | — | — | 0.734 | 0.591 | 0.447 | 478.3 | 312.6 | |
| SilentRetrievalRetriever=Contriever, Generator=Llama-2-7B-Chat, Judge=GPT-42026.05 | — | — | — | — | — | — | — | 0.813 | 0.642 | 0.548 | 33.1 | 27.4 | |
| Zhong et al.Retriever=Contriever, Generator=Llama-2-7B-Chat, Judge=GPT-42026.05 | — | — | — | — | — | — | — | 0.678 | 0.463 | 0.341 | 912.4 | 678.3 |