Question Answering on PubMedQA (EM, F1)
79.82EMSelf-MedRAG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Self-MedRAGRetrieval=BM25 + Contriever + RRF, Critic=NLI (roberta-large-mnli)2026.01 | 79.82 | 78.4 | |
| Self-MedRAGRetrieval=BM25 + Contriever + RRF, Critic=Llama3.1–8B2026.01 | 78.76 | 77.31 | |
| Base RAGRetrieval=BM25 + Contriever + RRF, Critic=None2026.01 | 69.1 | 64.45 | |
| Base RAGRetrieval=Contriever, Critic=None2026.01 | 67.9 | 64.41 | |
| Base RAGRetrieval=BM25, Critic=None2026.01 | 66.8 | 60.67 | |
| Base RAGRetrieval=MedCPT, Critic=None2026.01 | 65.97 | 62.6 | |
| BaseModelBackbone=Meta-Llama-3.1-8B-Instruct2026.01 | 0 | 9.55 | |
| Guided DecodingBackbone=Meta-Llama-3.1-8B-Instruct2026.01 | 0 | 13.95 | |
| Predictive DecodingBackbone=Meta-Llama-3.1-8B-Instruct2026.01 | 0 | 11.69 | |
| Chain-of-ThoughtsBackbone=Meta-Llama-3.1-8B-Instruct2026.01 | 0 | 20.33 | |
| Tree-of-ThoughtBackbone=Meta-Llama-3.1-8B-Instruct2026.01 | 0 | 12.38 | |
| Token-GuardBackbone=Meta-Llama-3.1-8B-Instruct2026.01 | 0 | 29.67 | |
| BaseModelBackbone=Qwen3-8B2026.01 | 0 | 22.45 | |
| Guided DecodingBackbone=Qwen3-8B2026.01 | 0 | 24.76 | |
| Predictive DecodingBackbone=Qwen3-8B2026.01 | 0 | 19.44 | |
| Chain-of-ThoughtsBackbone=Qwen3-8B2026.01 | 0 | 22.77 | |
| Tree-of-ThoughtBackbone=Qwen3-8B2026.01 | 0 | 26.41 | |
| Token-GuardBackbone=Qwen3-8B2026.01 | 0 | 28.91 |