Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Knowledge Base Question Answering | T-REx WikiData (test) | 37 | May 25, 2026 | ||
| Long-context reasoning | OOLONG | 37 | Apr 8, 2026 | ||
| Visual Question Answering | OmniMedVQA BRIGHT Challenge | 37 | Jun 9, 2026 | ||
| Ultrasound Image Segmentation | TN3K | 37 | Jul 9, 2026 | ||
| Verilog Code Generation | VerilogEval Machine |
| 37 |
| Apr 13, 2026 |
| Privacy-preserving unlearning | TDEC | 37 | Mar 19, 2026 |
|---|
| Object-goal navigation | HM3D (test) | 37 | Jul 1, 2026 |
|---|
| Science Question Answering | ARC Challenge | 37 | Jun 11, 2026 |
|---|
| Backdoor Attack | CelebA | 37 | Mar 17, 2026 |
|---|
| Video Generation | VBench | 37 | May 15, 2026 |
|---|
| Deepfake Interruption | FF++ O | 37 | May 4, 2026 |
|---|
| Deepfake Interruption | LFW | 37 | May 4, 2026 |
|---|
| Deepfake Interruption | CelebA | 37 | May 4, 2026 |
|---|
| Mathematical Reasoning | MATH-500 | 37 | Apr 16, 2026 |
|---|
| Long-context language modeling | Ruler llama3-8B-Instruct (test) | 37 | Mar 13, 2026 |
|---|
| Task and Scenario Goal Completion | AppWorld normal (test) | 37 | Jun 16, 2026 |
|---|
| Image Deblurring | UHD-Blur | 37 | Apr 20, 2026 |
|---|
| Spatio-temporal forecasting | GBA | 37 | Mar 12, 2026 |
|---|
| Novel View Synthesis | DL3DV 32 views | 37 | Jul 7, 2026 |
|---|
| Audio Understanding | MMSU | 37 | Jun 1, 2026 |
|---|
| Neuromorphic Image Classification | DVS-CIFAR10 | 37 | May 15, 2026 |
|---|
| Multi-view spatial reasoning | MindCube | 37 | Jun 12, 2026 |
|---|
| High-resolution Perception | HR-Bench 8K | 37 | Jul 8, 2026 |
|---|
| Spatial Understanding | MindCube Tiny | 37 | Mar 9, 2026 |
|---|
| Video Understanding | FAVOR-Bench | 37 | Jun 8, 2026 |
|---|