Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Question Answering | Probe 2 1.0 (test) | 42 | Feb 26, 2026 | ||
| Knowledge Probing | Probe 1 | 42 | Feb 26, 2026 | ||
| Prompt classification | Aegis 2.0 | 42 | Jul 3, 2026 | ||
| Question Answering | En.QA | 42 | Jun 11, 2026 | ||
| Length of Stay Prediction (LOS) | MIMIC-IV (test) | 42 | Apr 28, 2026 | ||
| Language Understanding |
| MMLU o=1 Exact split |
| 42 |
| Feb 26, 2026 |
| Multitask Language Understanding | MMLU Exact split, o=3 | 42 | Feb 26, 2026 |
|---|
| Mathematical reasoning | Minerva | 42 | May 22, 2026 |
|---|
| Question Answering | CSQA | 42 | Jun 4, 2026 |
|---|
| Reward Modeling | RM-Bench Chat | 42 | May 29, 2026 |
|---|
| Reward Modeling | RewardBench Chat | 42 | May 29, 2026 |
|---|
| Retrieval | 2Wiki | 42 | May 4, 2026 |
|---|
| Personalized Reward Modeling | Chatbot Arena Personalized | 42 | Feb 26, 2026 |
|---|
| Traffic Signal Control | Jinan-1 | 42 | Apr 29, 2026 |
|---|
| Multi-agent Cooperation | Simple_Tag 9 agents | 42 | Feb 26, 2026 |
|---|
| Multi-agent Cooperation | Simple_Tag 3 agents | 42 | Feb 26, 2026 |
|---|
| Question Answering | NaturalQuestions | 42 | Apr 23, 2026 |
|---|
| Mathematical Reasoning | Average GSM8k-Aug, GSM-Hard, SVAMP, MultiArith | 42 | Jun 30, 2026 |
|---|
| Jailbreak Attack | AutoDAN | 42 | Jul 7, 2026 |
|---|
| Sequential Recommendation | Amazon Sports (test) | 42 | Jun 30, 2026 |
|---|
| GUI Grounding | ScreenSpot (test) | 42 | Mar 31, 2026 |
|---|
| Question Answering | 7 QA tasks | 42 | Feb 26, 2026 |
|---|
| Hallucination | TruthfulQA | 42 | Feb 26, 2026 |
|---|
| Agentic Oversight | Deceivers | 42 | Feb 26, 2026 |
|---|
| Agentic Oversight | VitaBench | 42 | Feb 26, 2026 |
|---|