Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Event Prediction | TAOBAO (test) | 55 | Jun 9, 2026 | ||
| Event Prediction | TAXI (test) | 55 | Jun 9, 2026 | ||
| Zero-shot Evaluation | Evaluation Benchmarks Zero-shot | 55 | Jun 4, 2026 | ||
| Bounce Dynamics Prediction | Table Tennis Rackets 1.0 (test) | 55 | Apr 14, 2026 | ||
| General Reasoning | LiveBench | 55 | Jun 8, 2026 | ||
| Deepfake Detection | CDF v2 |
| 55 |
| Apr 13, 2026 |
| Deepfake Detection | CDF v1 | 55 | Apr 13, 2026 |
|---|
| Retrieval-Augmented Generation | All Datasets Aggregated | 55 | Apr 24, 2026 |
|---|
| Image Classification | CIFAR-10N ResNet-18 (test) | 55 | Apr 9, 2026 |
|---|
| Symbolic Regression | SRBench 116 regression problems (test) | 55 | Apr 1, 2026 |
|---|
| Mathematical Reasoning | MATH | 55 | Jun 18, 2026 |
|---|
| Multi-discipline Long Video Understanding | MLVU | 55 | May 12, 2026 |
|---|
| Visual Question Answering | VQA lite v2 | 55 | Jun 11, 2026 |
|---|
| Image Classification | DomainNet Real | 55 | Mar 25, 2026 |
|---|
| Image Classification | CIFAR100 | 55 | Mar 25, 2026 |
|---|
| Image Generation | LSUN Church v1 (test) | 55 | Mar 24, 2026 |
|---|
| Code Generation | MBPP | 55 | Jun 18, 2026 |
|---|
| Sleep classification | BOAS Supervised Cross-Validation Scenario 1 (test) | 55 | Mar 12, 2026 |
|---|
| Long-term forecasting | ECL | 55 | Jun 23, 2026 |
|---|
| Early screening of successful training runs | PPO Runs CartPole, LunarLander, MiniGrid (10% of train) | 55 | Mar 11, 2026 |
|---|
| Jailbreaking | AdvBench | 55 | Mar 11, 2026 |
|---|
| Micro-expression recognition | SAMM 5-class | 55 | Jun 29, 2026 |
|---|
| Information Retrieval | FollowIR | 55 | Jun 18, 2026 |
|---|
| Leaderboard Evaluation | Open LLM Leaderboard 2 | 55 | Mar 4, 2026 |
|---|
| Instruction Following | IF-Eval 0-shot | 55 | Apr 10, 2026 |
|---|