Loading the SOTA2 catalog…
SOTA2 Research · papers
Find papers, implementations, and the benchmark evidence behind state-of-the-art AI systems.
| Rujing Yao, Yiquan Wu, Tong Zhang |
| 2025 |
| arxiv 2502.07904 |
| Disfl-QA: A Benchmark Dataset for Understanding Disfluencies in Question Answering | Aditya Gupta, Jiacheng Xu, Shyam Upadhyay | 2021 | arxiv 2106.04016 |
|---|
| MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering | Yuexing Hao, Kumail Alhamoud, Hyewon Jeong | 2025 | arxiv 2505.24040 |
|---|
| POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering | Yichen Xu, Liangyu Chen, Liang Zhang | 2025 | arxiv 2507.11939 |
|---|
| EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning | Mingyang Wei, Dehai Min, Zewen Liu | 2026 | arxiv 2601.03471 |
|---|
| DCQA: Document-Level Chart Question Answering towards Complex Reasoning and Common-Sense Understanding | Anran Wu, Luwei Xiao, Xingjiao Wu | 2023 | arxiv 2310.18983 |
|---|
| MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models | Suraj Racha, Prashant Joshi, Anshika Raman | 2025 | arxiv 2502.15418 |
|---|
| ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering | Daeyong Kwon, SeungHeon Doh, Juhan Nam | 2025 | arxiv 2512.05430 |
|---|
| A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering | Nusrat Sultana, Abdullah Muhammad Moosa, Kazi Afzalur Rahman | 2026 | arxiv 2604.07274 |
|---|
| MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation | Adrien Bazoge | 2025 | arxiv 2507.20917 |
|---|
| Pre-Training Multi-Modal Dense Retrievers for Outside-Knowledge Visual Question Answering | Alireza Salemi, Mahta Rafiee, Hamed Zamani | 2023 | arxiv 2306.16478 |
|---|