Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Action Anticipation | DARai | 64 | Feb 26, 2026 | ||
| Multi-task Language Understanding | MMLU-Pro | 64 | May 18, 2026 | ||
| Speech Reconstruction | LibriSpeech (test-clean) | 64 | Jun 1, 2026 | ||
| Time Series Forecasting | ETTh2 (5% train) | 64 | May 14, 2026 | ||
| Image Classification | FMNIST | 64 | Jun 4, 2026 | ||
| Reward Modeling Evaluation |
| Reward Bench Factuality 2 |
| 64 |
| Feb 26, 2026 |
| Forecasting | Electricity (test) | 64 | Feb 26, 2026 |
|---|
| Test Generation | TESTEVAL | 64 | Feb 26, 2026 |
|---|
| Anomaly Detection | Mammography | 64 | Mar 31, 2026 |
|---|
| Binary Classification | CH (test) | 64 | Feb 26, 2026 |
|---|
| Audio-driven Talking Head Generation | RAVDESS (cross-identity) | 64 | Apr 23, 2026 |
|---|
| General Visual Understanding | RealWorldQA | 64 | Jun 4, 2026 |
|---|
| Jailbreak Attack | StrongREJECT (test) | 64 | Jun 16, 2026 |
|---|
| Text Generation | CNN/Daily Mail (test) | 64 | Feb 26, 2026 |
|---|
| Open-ended text generation | PTB | 64 | Feb 26, 2026 |
|---|
| Question Answering | WebQA | 64 | May 27, 2026 |
|---|
| Multi-task Language Understanding | MMLU-Pro | 64 | Apr 14, 2026 |
|---|
| Question Answering | FreebaseQA | 64 | May 27, 2026 |
|---|
| Factuality Correction | VELI5 | 64 | Feb 26, 2026 |
|---|
| Code Generation | LiveCodeBench | 64 | May 8, 2026 |
|---|
| Multilingual Mathematical Reasoning | MT Math100 | 64 | Mar 4, 2026 |
|---|
| End-to-End Skill Extraction | KARIYER (test) | 64 | Feb 26, 2026 |
|---|
| Personalized Text Generation | LaMP-5 v1 (test) | 64 | Feb 26, 2026 |
|---|
| Science Question Answering | ScienceQA | 64 | Mar 10, 2026 |
|---|
| Graph Classification | IMDB | 64 | Jun 30, 2026 |
|---|