Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Hierarchical Unlearning | MedForget 1.0 (Forget) | 72 | Feb 26, 2026 | ||
| Reasoning-informed Image Editing | RISE-Bench | 72 | Jul 3, 2026 | ||
| Video Question Answering | MVBench | 72 | Jun 30, 2026 | ||
| PII leakage evaluation | SPIRIT |
| 72 |
| Feb 26, 2026 |
| Persona Manipulation | ANTHR (test) | 72 | Feb 26, 2026 |
|---|
| Persona Manipulation | MPI (test) | 72 | Feb 26, 2026 |
|---|
| Persona Manipulation | BFI (test) | 72 | Feb 26, 2026 |
|---|
| Binary Code Similarity Detection | BinaryCorp-VirtualAssembly (test) | 72 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | OlympiadBench | 72 | Mar 25, 2026 |
|---|
| Token Classification | Amazon ESCI Product Description English (test) | 72 | Feb 26, 2026 |
|---|
| Token Classification | Amazon ESCI Product Title English (test) | 72 | Feb 26, 2026 |
|---|
| Reward Modeling | RM-Bench (test) | 72 | Jun 8, 2026 |
|---|
| Reward Modeling | PPE-Preference | 72 | Jun 1, 2026 |
|---|
| Visual Reasoning | V* | 72 | Jun 30, 2026 |
|---|
| Hallucination Detection | Math | 72 | Feb 26, 2026 |
|---|
| Audio-visual understanding | WorldSense | 72 | May 19, 2026 |
|---|
| Instruction Following | IFBench | 72 | Mar 20, 2026 |
|---|
| Base-to-Novel Generalization | EuroSAT | 72 | Jul 2, 2026 |
|---|
| Lifelong Knowledge Editing | E-VQA Lifelong Sequential | 72 | Feb 26, 2026 |
|---|
| Referring Segmentation | refCOCOg (val) | 72 | May 22, 2026 |
|---|
| Node Classification | CS-Random (test) | 72 | Mar 18, 2026 |
|---|
| Few-shot Classification | ModelNet40 | 72 | Jul 1, 2026 |
|---|
| Open Vocabulary Semantic Segmentation | ADE20K without background | 72 | Mar 25, 2026 |
|---|
| Multimodal Sentiment Analysis | MOSI | 72 | Mar 20, 2026 |
|---|
| Referring Segmentation | refCOCO+ (testA) | 72 | Jul 8, 2026 |
|---|