Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Social Reasoning | 18 | 8 | Jun 18, 2026 | |
| General Evaluation | 18 | 22 | Jul 7, 2026 | |
| Long-context generation | 18 | 6 | Jun 25, 2026 | |
| Downstream Performance Prediction | 18 | 3 | May 14, 2026 | |
| General Knowledge Reasoning | 18 | 28 | Jul 8, 2026 | |
| LLM Decoding | 18 | 8 | Jun 4, 2026 | |
| LLM Evaluation | 17 | 15 | Jul 7, 2026 | |
| In-Context Learning | 17 | 8 | Jul 3, 2026 | |
| Truthfulness Evaluation | 17 | 32 | Jun 11, 2026 |
| Conversational Response Generation | 17 | 9 | May 18, 2026 |
|---|
| Visual Instruction Following | 17 | 17 | Jun 19, 2026 |
|---|
| LTL Instruction Following | 17 | 2 | Feb 27, 2026 |
|---|
| Knowledge Probing | 17 | 11 | Jun 5, 2026 |
|---|
| Math problem solving | 17 | 18 | Jun 26, 2026 |
|---|
| Language Understanding and Reasoning | 16 | 17 | Jun 29, 2026 |
|---|
| Hallucination Prediction | 16 | 4 | Jun 2, 2026 |
|---|
| Value Alignment | 16 | 7 | May 27, 2026 |
|---|
| Personalized Reward Modeling | 16 | 5 | Jun 9, 2026 |
|---|
| Large Language Model Inference | 16 | 14 | Jun 23, 2026 |
|---|
| Multi-turn conversation | 16 | 27 | Jun 11, 2026 |
|---|
| Watermarking Robustness (Translation Attack) | 16 | 1 | Mar 26, 2026 |
|---|
| Reward Model Evaluation | 16 | 12 | Jun 26, 2026 |
|---|
| Human preference prediction | 15 | 14 | May 26, 2026 |
|---|
| Personalized Generation | 15 | 7 | May 27, 2026 |
|---|
| General Knowledge Evaluation | 15 | 28 | Jun 25, 2026 |
|---|