Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Video Understanding | VideoMMMU | 67 | Jul 7, 2026 | ||
| Faithfulness evaluation | WikiBio | 67 | Apr 16, 2026 | ||
| Faithfulness evaluation | TellMeWhy | 67 | Apr 16, 2026 | ||
| Language Understanding | CEval | 67 | Jun 5, 2026 | ||
| Mathematical Reasoning | College | 67 | Apr 20, 2026 | ||
| GUI Grounding | ScreenSpot-Pro (test) |
| 67 |
| Jun 15, 2026 |
| Composed Image Retrieval | FashionIQ (Dress) | 67 | Jul 3, 2026 |
|---|
| Gaze target estimation | GazeFollow | 67 | Jul 7, 2026 |
|---|
| Video Reconstruction | UCF-101 | 67 | Jun 17, 2026 |
|---|
| Semantic Occupancy Prediction | SemanticKITTI (test) | 67 | Jul 1, 2026 |
|---|
| Semantic Segmentation | A-150 | 67 | Apr 28, 2026 |
|---|
| Speech Reconstruction | LibriTTS clean (test) | 67 | Jul 7, 2026 |
|---|
| Visual Reasoning | MMVP | 67 | Jul 7, 2026 |
|---|
| Open Vocabulary Semantic Segmentation | Cityscapes without background | 67 | Mar 25, 2026 |
|---|
| Open Vocabulary Semantic Segmentation | PASCAL Context 59 without background | 67 | Mar 25, 2026 |
|---|
| Question Answering | Evaluation Suite (ARC, HellaSwag, MMLU) Zero-shot (test) | 67 | Feb 26, 2026 |
|---|
| Classification | RSNA Pneumonia | 67 | Jun 30, 2026 |
|---|
| Image Super-Resolution | B100 x2 (test) | 67 | May 19, 2026 |
|---|
| Question Answering | MedQA (test) | 67 | Apr 29, 2026 |
|---|
| Multimodal Evaluation | MM-Bench | 67 | Jun 23, 2026 |
|---|
| Symbolic Reasoning | Letter | 67 | Mar 24, 2026 |
|---|
| Respiratory sound classification | ICBHI dataset official (60-40% split) | 67 | Jun 11, 2026 |
|---|
| Camouflaged Object Detection | CHAMELEON (test) | 67 | Apr 21, 2026 |
|---|
| Image Retrieval | GLD v2 (test) | 67 | Mar 10, 2026 |
|---|
| Image Retrieval | ROxford | 67 | Apr 2, 2026 |
|---|