Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LaMP-4 Personalization | 1 | 1 | Feb 18, 2026 | |
| General Chat Evaluation | 1 | 1 | Mar 6, 2026 | |
| Lifelong Safety Adaptation | 1 | 1 | May 15, 2026 | |
| RLHF Training | 1 | 1 | Mar 6, 2026 | |
| Multi-modal Long-context Benchmarking | 1 | 1 | Mar 6, 2026 | |
| beauty live-commerce dialogue generation | 1 | 1 | May 15, 2026 | |
| NIAH | 1 | 1 | Feb 18, 2026 | |
| Pareto prompt set identification | 1 | 1 | May 15, 2026 | |
| Soft constrained reward optimization |
| 1 |
| 1 |
| May 15, 2026 |
| Logic Puzzle | 1 | 1 | May 15, 2026 |
|---|
| LaMP-5 Personalization | 1 | 1 | Feb 18, 2026 |
|---|
| MK | 1 | 1 | Feb 18, 2026 |
|---|
| Language Understanding (Law) | 1 | 1 | May 15, 2026 |
|---|
| Math Find | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-turn conversational and instruction-following | 1 | 1 | Mar 6, 2026 |
|---|
| Generative Operational Planning | 1 | 1 | May 15, 2026 |
|---|
| Interactional pluralism evaluation | 1 | 1 | May 15, 2026 |
|---|
| Knowledge-based Agent Reasoning | 1 | 1 | Mar 6, 2026 |
|---|
| General speculative decoding performance | 1 | 2 | Jun 1, 2026 |
|---|
| Pairwise Human Evaluation | 1 | 1 | May 15, 2026 |
|---|
| Personal Assistant Agent Performance | 1 | 1 | May 15, 2026 |
|---|
| Voice Chatting | 1 | 1 | Feb 18, 2026 |
|---|
| Judge Performance | 1 | 1 | Mar 6, 2026 |
|---|
| Text-based embodied AI | 1 | 1 | May 15, 2026 |
|---|
| Cross-shot consistency | 1 | 1 | May 15, 2026 |
|---|