Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Text-to-Motion | KIT-ML | 44 | Apr 28, 2026 | ||
| Instruction-guided image editing | GEdit-Bench EN Full set | 44 | Jun 11, 2026 | ||
| Spatial Reasoning (Multi-Image) | MMSI-Bench | 44 | May 27, 2026 | ||
| Egocentric daily-task planning | EgoPlanBench2 | 44 | Apr 24, 2026 | ||
| Proof-of-Vulnerability (PoV) Generation | Magma | 44 | Feb 26, 2026 |
| Jailbreak Defense | HarmBench and AdvBench (test) | 44 | Feb 26, 2026 |
|---|
| Large Language Model Evaluation | Open PL LLM Leaderboard instruction-tuned | 44 | Feb 26, 2026 |
|---|
| Safety Evaluation | Malicious Instruct | 44 | May 4, 2026 |
|---|
| Dialogue Segmentation | DialSeg711 | 44 | Jun 1, 2026 |
|---|
| Factuality Correction | BIO (test) | 44 | Feb 26, 2026 |
|---|
| Prompt classification | SimpST | 44 | Jul 8, 2026 |
|---|
| Automatic Speech Recognition | VoxPopuli | 44 | Jun 18, 2026 |
|---|
| Long-Context Retrieval | RULER | 44 | May 22, 2026 |
|---|
| Math Reasoning | OlympiadBench | 44 | Apr 8, 2026 |
|---|
| Personalized Reward Modeling | PRISM Personalized | 44 | Feb 26, 2026 |
|---|
| Multi-agent Cooperation | Simple_Tag 6 agents | 44 | Jul 7, 2026 |
|---|
| Text2SQL | BIRD (dev) | 44 | Mar 18, 2026 |
|---|
| Multi-hop Retrieval | HotpotQA | 44 | May 13, 2026 |
|---|
| Sequential Recommendation | Amazon Toys (test) | 44 | Jun 30, 2026 |
|---|
| Interactive Tool-Use Agent Performance | VitaBench | 44 | May 19, 2026 |
|---|
| Conditional Shapley value estimation | Wine M=11 | 44 | Feb 26, 2026 |
|---|
| Conditional Shapley value estimation | Abalone cont (M=7) | 44 | Feb 26, 2026 |
|---|
| Question Answering | ARC Easy | 44 | Feb 26, 2026 |
|---|
| Information Retrieval | FIQA BEIR (test) | 44 | May 21, 2026 |
|---|
| Goal-oriented Dialogue | Movie | 44 | Apr 15, 2026 |
|---|