Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Classification | Yahoo! Answer | 34 | Jun 11, 2026 | ||
| Planning | AutoLogic | 34 | Jun 9, 2026 | ||
| Mathematical Reasoning | AIME24 | 34 | Jun 9, 2026 | ||
| Mathematical Reasoning | AIME 24 | 34 | Jun 8, 2026 | ||
| Grasp Pose Estimation | GraspNet-1Billion Novel | 34 | Jun 9, 2026 | ||
| Grasp Pose Estimation |
|---|
| GraspNet-1Billion (Similar) |
| 34 |
| Jun 9, 2026 |
| Grasp Pose Estimation | GraspNet-1Billion (Seen) | 34 | Jun 9, 2026 |
|---|
| Image Classification | CIFAR-10-LT gamma=10 (test) | 34 | Jun 4, 2026 |
|---|
| Video Compression | UVG | 34 | Jun 4, 2026 |
|---|
| Visual Question Answering | OKVQA | 34 | Jul 7, 2026 |
|---|
| Constraint Satisfaction Reasoning | Countdown | 34 | Jun 4, 2026 |
|---|
| Zero-shot Reasoning | Reasoning Benchmarks Average | 34 | Jun 19, 2026 |
|---|
| Symbolic Regression | Jacobi-based SR problem | 34 | Jun 4, 2026 |
|---|
| Text Classification | AGNEWS | 34 | Jun 4, 2026 |
|---|
| Geometric Question Answering | GEOQA (test) | 34 | Jun 15, 2026 |
|---|
| Agentic Capability Evaluation | ACEBench en | 34 | Jun 2, 2026 |
|---|
| trajectory-based action prediction | AgentNetBench | 34 | Jun 2, 2026 |
|---|
| trajectory-based action prediction | AndroidControl | 34 | Jun 2, 2026 |
|---|
| Code Generation | HumanEval+ | 34 | Jun 30, 2026 |
|---|
| Interaction value estimation | Diabetes (test) | 34 | Jun 2, 2026 |
|---|
| Financial Strategy Generation | Crypto | 34 | Jun 2, 2026 |
|---|
| Commonsense Reasoning | Common Sense Reasoning ARC-C, ARC-E, PIQA, StoryCloze | 34 | Jun 2, 2026 |
|---|
| Regression | ZIP Code (test) | 34 | Jun 2, 2026 |
|---|
| Data Selection | MedMCQA (fresh candidate pool) | 34 | Jun 2, 2026 |
|---|
| Long-context reasoning | LongReason 64K-input 70K context | 34 | May 28, 2026 |
|---|