Ambiguous Question Answering on AmbigQA (test)
54.43AccuracyHERA (Qwen-3-14B)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HERA (Qwen-3-14B)category=Ours2026.04 | 54.43 | 48.74 | 67.81 | |
| DeepResearcher-7Bcategory=Advanced RAG2026.04 | 49.74 | 42.23 | 51.5 | |
| HERA (Llama-3.1-8)category=Ours2026.04 | 48.53 | 43.46 | 60.72 | |
| Search-r1-7Bcategory=Advanced RAG2026.04 | 47.32 | 45.88 | 65.54 | |
| InstructRAG-8Bcategory=Advanced RAG2026.04 | 40.07 | 39.72 | 40.73 | |
| Qwen2.5-14B-Instructcategory=Single-turn Retrieval2026.04 | 39.78 | 36.65 | 50.23 | |
| MMAO-RAG-8Bcategory=Advanced RAG2026.04 | 38.85 | 34.75 | 48.59 | |
| Qwen3-14Bcategory=Single-turn Retrieval2026.04 | 38.47 | 36.1 | 50.91 | |
| GPT-4o-minicategory=Single-turn Retrieval2026.04 | 37.4 | 33.8 | 47.99 | |
| Qwen3-8Bcategory=Single-turn Retrieval2026.04 | 36.02 | 33.8 | 47.7 | |
| CORAG-8B (Greedy)category=Advanced RAG2026.04 | 32.1 | 27.25 | 43.95 | |
| SELF-RAGcategory=Advanced RAG, Backbone=llama2-7B2026.04 | 28.51 | 24.03 | 39.62 | |
| Llama3-3.1-8B-instructcategory=Single-turn Retrieval2026.04 | 18.03 | 18.96 | 24.6 |