Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Reliability of post-edit LLMs | Books3 | 36 | Feb 27, 2026 | ||
| Named Entity Recognition | GUM | 36 | Feb 26, 2026 | ||
| Video-based Human Mesh Recovery | 3DPW standard (test) | 36 | Feb 26, 2026 | ||
| Text Classification | Emotion | 36 | Feb 26, 2026 | ||
| Multi-task Language Understanding | MMLUpro (test) | 36 | May 27, 2026 | ||
| Knowledge Base Question Answering |
|---|
| WebQSP → GrailQA-Tech (test) |
| 36 |
| Feb 26, 2026 |
| Instruction Following | Dolly | 36 | Jun 9, 2026 |
|---|
| Math Reasoning | MATH 200 samples (test) | 36 | Feb 26, 2026 |
|---|
| Multimodal Machine Unlearning Evaluation | MLLMU-Bench Forget Set | 36 | Mar 18, 2026 |
|---|
| Model Selection | SUN397 | 36 | Feb 26, 2026 |
|---|
| Model Selection | Pets | 36 | Feb 26, 2026 |
|---|
| Model Selection | CIFAR10 | 36 | Feb 26, 2026 |
|---|
| Model Selection | CIFAR100 | 36 | Feb 26, 2026 |
|---|
| Model Selection | Cars | 36 | Feb 26, 2026 |
|---|
| Object Detection | VisDrone | 36 | Apr 29, 2026 |
|---|
| Question Answering | MuSiQue | 36 | Feb 26, 2026 |
|---|
| Question-Answering | UNQOVER | 36 | Feb 26, 2026 |
|---|
| Question-Answering | BBQ | 36 | Feb 26, 2026 |
|---|
| Question Answering | SciQ (train) | 36 | Feb 27, 2026 |
|---|
| 3D Hand Pose Estimation | FreiHAND | 36 | Mar 27, 2026 |
|---|
| Rāg classification | Tagore songs from Swarabitan | 35 | Jul 9, 2026 |
|---|
| qualifying-position regression | rel-f1 | 35 | Jul 9, 2026 |
|---|
| driver-position regression | rel-f1 | 35 | Jul 9, 2026 |
|---|
| Agile drone pursuit-evasion | 3v1 Agile Drone Pursuit-Evasion Simulation | 35 | Jul 8, 2026 |
|---|
| Span-level uncertainty estimation | SPANUQ-BENCH v1.0 (test) | 35 | Jul 8, 2026 |
|---|