Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Long-Horizon | 1 | 1 | Jun 12, 2026 | |
| Prompt Steering | 1 | 1 | Feb 19, 2026 | |
| Multi-Turn Session | 1 | 1 | May 26, 2026 | |
| Method ranking self-consistency | 1 | 1 | Mar 12, 2026 | |
| General-domain Persuasion Robustness | 1 | 1 | May 26, 2026 | |
| Science and Knowledge Question Answering | 1 | 1 | Mar 24, 2026 | |
| Multi-level multi-discipline evaluation | 1 | 2 | Feb 18, 2026 | |
| Logical Reasoning Reliability Evaluation | 1 | 1 | May 26, 2026 | |
| Human consistency correlation |
| 1 |
| 1 |
| Mar 12, 2026 |
| Logic Consistency Evaluation | 1 | 1 | May 26, 2026 |
|---|
| Open-ended Information Seeking | 1 | 1 | Feb 19, 2026 |
|---|
| Long-text Question Answering | 1 | 1 | Mar 13, 2026 |
|---|
| Subjective evaluation of research ideas | 1 | 1 | May 26, 2026 |
|---|
| Sparse Decoding (MHA) | 1 | 1 | May 26, 2026 |
|---|
| Conciseness human alignment evaluation | 1 | 1 | Mar 13, 2026 |
|---|
| Sparse Decode Performance | 1 | 1 | May 26, 2026 |
|---|
| Open-ended task evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| LLM behavior monitoring | 1 | 1 | May 26, 2026 |
|---|
| Interactive Deep Research | 1 | 1 | May 26, 2026 |
|---|
| Text-to-image reward alignment | 1 | 1 | Mar 13, 2026 |
|---|
| Persona Drift Measurement | 1 | 1 | May 26, 2026 |
|---|
| Challenging Cases Understanding | 1 | 1 | Feb 18, 2026 |
|---|
| STEM Problem Solving | 1 | 1 | Feb 19, 2026 |
|---|
| GameCWM Generation | 1 | 1 | May 26, 2026 |
|---|
| Reply Generation + Tone Adjustment | 1 | 1 | Mar 13, 2026 |
|---|