Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Image Classification | CIFAR10-S | 37 | May 4, 2026 | ||
| Multimodal Understanding | HallusionBench | 37 | May 19, 2026 | ||
| Chess | ChessArena v1 (Whole Rating Leaderboard) | 37 | Apr 24, 2026 | ||
| Failure Detection | LIBERO unseen | 37 | Apr 23, 2026 | ||
| Failure Detection | LIBERO seen | 37 | Apr 23, 2026 | ||
| pT estimation |
| CMS Trigger Dataset |
| 37 |
| Apr 21, 2026 |
| Chart Question Answering | ChartQA | 37 | Jun 26, 2026 |
|---|
| Recall | Finance-Medical Dataset (test) | 37 | Apr 20, 2026 |
|---|
| Drone-to-Satellite Cross-view Geo-localization | SUES-200 (200m) | 37 | May 12, 2026 |
|---|
| Reasoning | MME-RealWorld Lite | 37 | May 5, 2026 |
|---|
| Text-to-Image Compositional Alignment | T2I-CompBench++ v2 (test) | 37 | Apr 14, 2026 |
|---|
| Partial Multi-Label Learning | yeast | 37 | Apr 13, 2026 |
|---|
| Image Classification | CIFAR-10-C (test) | 37 | Jun 1, 2026 |
|---|
| OOD Detection | Kvasir v2 | 37 | Jun 16, 2026 |
|---|
| Language Modeling | LAMBADA standard (LS) | 37 | Jun 12, 2026 |
|---|
| Multimodal Deep Search | BC-VL | 37 | May 12, 2026 |
|---|
| Open-ended generation | TriviaQA | 37 | Jun 2, 2026 |
|---|
| Mathematical Reasoning | MathVista | 37 | Jun 26, 2026 |
|---|
| Graph Classification | Colors3 | 37 | May 12, 2026 |
|---|
| Link Prediction | Bili Dance | 37 | Jun 15, 2026 |
|---|
| Anomaly Detection | Wine | 37 | May 12, 2026 |
|---|
| Question Answering | TruthfulQA | 37 | Apr 21, 2026 |
|---|
| Node Classification | DBLP | 37 | May 18, 2026 |
|---|
| Partially Relevant Video Retrieval | TVR | 37 | May 8, 2026 |
|---|
| Optical Illusion Character Recognition | IlluChar (Noise) | 37 | Mar 25, 2026 |
|---|