Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM-as-Judge Response Evaluation | 2 | 1 | Feb 25, 2026 | |
| Belief Prediction | 2 | 2 | May 22, 2026 | |
| Human-AI Agreement Assessment | 2 | 1 | Mar 26, 2026 | |
| Reasoning and Language Understanding | 2 | 2 | Feb 18, 2026 | |
| Mathematical and General Reasoning | 2 | 1 | Feb 27, 2026 | |
| grade-school math | 2 | 3 | Jun 30, 2026 | |
| User Perceived Understandability | 2 | 1 | Feb 18, 2026 | |
| Evidence-grounded diagnostic reasoning | 2 | 1 | Mar 25, 2026 | |
| General AI Assistants Evaluation |
| 2 |
| 2 |
| Mar 10, 2026 |
| General AI Assistant Task Execution | 2 | 1 | Mar 25, 2026 |
|---|
| Standard Operating Procedure execution | 2 | 1 | Mar 25, 2026 |
|---|
| Slide Deck Generation | 2 | 1 | Feb 18, 2026 |
|---|
| LVLM Evaluation | 2 | 3 | May 1, 2026 |
|---|
| Uncovering hidden system prompts | 2 | 1 | Mar 25, 2026 |
|---|
| Hallucination Reasoning | 2 | 1 | Feb 18, 2026 |
|---|
| Instructed Code Generation | 2 | 1 | Mar 25, 2026 |
|---|
| Generative Hallucination Evaluation | 2 | 3 | May 28, 2026 |
|---|
| user-preference matrix generation | 2 | 1 | Mar 4, 2026 |
|---|
| Discriminative Performance | 2 | 1 | Feb 18, 2026 |
|---|
| Expert-level Multimodal Understanding | 2 | 4 | May 27, 2026 |
|---|
| Multi-agent Negotiation | 2 | 2 | Apr 23, 2026 |
|---|
| E2E Generation Latency | 2 | 1 | Feb 18, 2026 |
|---|
| Mathematical Calculation | 2 | 2 | Apr 10, 2026 |
|---|
| Multiple-choice commonsense reasoning | 2 | 2 | Apr 14, 2026 |
|---|
| Visual Search and Reasoning | 2 | 2 | May 27, 2026 |
|---|