Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Multimodal Routing | 1 | 1 | May 13, 2026 | |
| LLM Judging | 1 | 1 | May 13, 2026 | |
| Offline Constrained Reinforcement Learning | 1 | 1 | May 13, 2026 | |
| Human-like behavior evaluation | 1 | 1 | Feb 18, 2026 | |
| Human Subjective Evaluation | 1 | 1 | May 13, 2026 | |
| Natural Language Understanding and Mathematical Reasoning | 1 | 1 | Mar 4, 2026 | |
| Automating Agent Evaluation | 1 | 1 | May 13, 2026 | |
| LLM Evaluation Performance | 1 | 1 | Feb 18, 2026 | |
| General-purpose Visual Instruction Following |
| 1 |
| 1 |
| May 13, 2026 |
| Long Novel Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Wealth Management Chatbot Response Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-Turn Negotiation | 1 | 1 | Mar 4, 2026 |
|---|
| Multi-step Narrative Reasoning | 1 | 1 | May 13, 2026 |
|---|
| Out-of-Domain Reasoning Aggregation | 1 | 1 | May 13, 2026 |
|---|
| Synonym Substitution Robustness | 1 | 1 | May 13, 2026 |
|---|
| Reasoning Performance (Aggregate) | 1 | 1 | Feb 18, 2026 |
|---|
| Pedagogical Response Evaluation | 1 | 1 | Mar 4, 2026 |
|---|
| Paraphrasing Robustness | 1 | 1 | May 13, 2026 |
|---|
| Instruction score prediction | 1 | 1 | Feb 18, 2026 |
|---|
| Answer-unaware Conversational Question Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Novel Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Helpful assistant task | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-Turn Problem Solving | 1 | 1 | Mar 4, 2026 |
|---|
| Two-objective simultaneous preference alignment | 1 | 1 | May 13, 2026 |
|---|
| Long-context Information Extraction | 1 | 1 | May 13, 2026 |
|---|