Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| General Language Evaluation | 14-Benchmark Evaluation Suite | 72 | Mar 4, 2026 | ||
| Longest Common Subsequence | LARGE | 72 | Feb 27, 2026 | ||
| Retrieval | Flickr2.4M (test) | 72 | Feb 26, 2026 | ||
| Retrieval | MS-COCO | 72 | Feb 26, 2026 | ||
| Multi-hop Question Answering | 2WikiMultiHopQA Out-Of-Distribution (OOD) | 72 | Feb 26, 2026 | ||
| HotpotQA In-Distribution |
| 72 |
| Feb 26, 2026 |
| PDE Solving | CodePDE | 72 | Feb 26, 2026 |
|---|
| Multimodal Sentiment Analysis | CMU-MOSI v1 (test) | 72 | Mar 4, 2026 |
|---|
| Jailbreak Robustness | AdvBench | 72 | May 12, 2026 |
|---|
| 3D reconstruction | Mip-NeRF 360 | 72 | May 29, 2026 |
|---|
| Reward Modeling | RewardBench v2 | 72 | Mar 4, 2026 |
|---|
| Novel View Synthesis | KITTI | 72 | Jun 30, 2026 |
|---|
| Commonsense Reasoning | Average 7 Commonsense Reasoning Tasks | 72 | Apr 3, 2026 |
|---|
| Science Reasoning | GPQA Diamond | 72 | Jun 17, 2026 |
|---|
| Safety-Utility Trade-off Evaluation | S-Eval, ORFuzzSet, and NQ Aggregated | 72 | Feb 26, 2026 |
|---|
| Over-refusal Evaluation | NQ (Natural Questions) | 72 | Feb 26, 2026 |
|---|
| Over-refusal Evaluation | ORFuzzSet | 72 | Feb 26, 2026 |
|---|
| Safety Risk Evaluation | S-Eval (Risk) | 72 | Feb 26, 2026 |
|---|
| Jailbreak Attack Evaluation | S-Eval Aattack | 72 | Feb 26, 2026 |
|---|
| Tabular Anomaly Detection | Wine | 72 | May 12, 2026 |
|---|
| Tabular Anomaly Detection | Pendigits | 72 | May 12, 2026 |
|---|
| Time Series Forecasting | Energy | 72 | May 12, 2026 |
|---|
| Instruction Following | IFBench | 72 | Apr 29, 2026 |
|---|
| One-shot adaptive testing | PTADisc (test) | 72 | Feb 26, 2026 |
|---|
| One-shot adaptive testing | JUNYI (test) | 72 | Feb 26, 2026 |
|---|