Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Robot Policy Learning | LIBERO | 73 | Jul 2, 2026 | ||
| Image Classification | CIFAR-10 ViT-Tiny (test) | 73 | Feb 26, 2026 | ||
| Knowledge Reasoning | MMLU | 73 | May 20, 2026 | ||
| Survival Analysis | TCGA-GBMLGG | 73 | Jun 4, 2026 | ||
| Recommendation | Yelp 2018 |
| 73 |
| Apr 21, 2026 |
| Video Reasoning | Video-MME | 73 | Jul 7, 2026 |
|---|
| Plot-to-code generation | Plot2Code | 73 | Apr 27, 2026 |
|---|
| Chat | MT-Bench | 73 | Jul 9, 2026 |
|---|
| Question Answering | TruthfulQA | 73 | Feb 26, 2026 |
|---|
| Scientific Reasoning | GPQA Diamond | 73 | Jun 26, 2026 |
|---|
| Mathematical Reasoning | Olympiad Bench | 73 | Apr 9, 2026 |
|---|
| Code | MBPP | 73 | Jun 2, 2026 |
|---|
| Math Reasoning | GSM Hard | 73 | May 26, 2026 |
|---|
| Closed-loop planning | nuPlan 14-hard (test) | 73 | Jun 30, 2026 |
|---|
| Mathematical Reasoning | MATH 500 | 73 | Feb 26, 2026 |
|---|
| Talking Head Generation | HDTF (test) | 73 | Jun 29, 2026 |
|---|
| Biomedical Image Classification | 11 Biomedical Datasets Average (test) | 73 | Apr 28, 2026 |
|---|
| Multimodal Evaluation | LLaVA-bench in-the-wild | 73 | May 27, 2026 |
|---|
| Long-term Forecasting | Exchange | 73 | Jun 1, 2026 |
|---|
| Multimodal Benchmarking | MMBench | 73 | May 14, 2026 |
|---|
| Class-Incremental Learning | CUB-200 (test) | 73 | Jun 25, 2026 |
|---|
| Multi-choice Video Question Answering | MVBench | 73 | Feb 26, 2026 |
|---|
| Instruction Following | SelfInst | 73 | Apr 6, 2026 |
|---|
| Multimodal Question Answering | ScienceQA (test) | 73 | May 28, 2026 |
|---|
| Pairwise point cloud registration | 3DLoMatch | 73 | Mar 16, 2026 |
|---|