Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Quasi-Monte Carlo integration | 8-dimensional Borehole function (test) | 52 | Jun 2, 2026 | ||
| Speculative Decoding | MBPP | 52 | Jun 4, 2026 | ||
| Speculative Decoding | MATH 500 | 52 | Jun 4, 2026 | ||
| Image Classification | CIFAR-100 coarse | 52 | Jun 1, 2026 | ||
| Instruction Following | IFEval | 52 | Jul 7, 2026 | ||
| Min/Max Distance |
| AlphaEvolve Min Max Distance (MMD, n=16) |
| 52 |
| May 29, 2026 |
| Time series forecasting | PTF | 52 | May 29, 2026 |
|---|
| 3D Object Detection | (EGO+AUX1) 128-beam LiDAR (test) | 52 | May 27, 2026 |
|---|
| 3D Object Detection | V2X-Real EGO+AUX2 zero-shot 128-beam LiDAR | 52 | May 27, 2026 |
|---|
| Image Classification | Imagenette | 52 | May 21, 2026 |
|---|
| Binary Classification | Apple-defined Content Rating Descriptors | 52 | May 21, 2026 |
|---|
| Math Reasoning | AIME 24 | 52 | Jul 7, 2026 |
|---|
| Code Generation | LiveCode Bench V6 | 52 | May 20, 2026 |
|---|
| Semantic Segmentation | Cityscapes low-data regime (val) | 52 | May 19, 2026 |
|---|
| Object Detection | Cityscapes low-data regime (val) | 52 | May 19, 2026 |
|---|
| LiDAR Semantic Segmentation | ScribbleKITTI v1.0 (test) | 52 | May 19, 2026 |
|---|
| Fair Clustering | BANK-40K | 52 | May 14, 2026 |
|---|
| Object Probing | POPE Average | 52 | May 28, 2026 |
|---|
| Attribute Steering | Sentiment S | 52 | May 13, 2026 |
|---|
| Face Forgery Detection | GenFace (EFS) | 52 | May 12, 2026 |
|---|
| Mathematical Reasoning | Math | 52 | May 26, 2026 |
|---|
| Monitoring | MonitoringBench Full Trajectory 1.0 | 52 | May 12, 2026 |
|---|
| Humanoid Path Search | H*Bench | 52 | May 12, 2026 |
|---|
| Information Quality Assessment | RepLiQA seventeen topics | 52 | May 11, 2026 |
|---|
| Anomaly Detection | aloi ADBench | 52 | May 11, 2026 |
|---|