Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Downstream Generation Impact (Qwen2.5-VL) | 1 | 1 | May 11, 2026 | |
| Hybrid Retrieval-Augmented Generation | 1 | 1 | Feb 25, 2026 | |
| Downstream Generation Impact (GeoChat) | 1 | 1 | May 11, 2026 | |
| Research Assistant | 1 | 1 | Feb 18, 2026 | |
| Downstream Generation Impact (H2RSVLM) | 1 | 1 | May 11, 2026 | |
| Multi-agent game strategic reasoning | 1 | 1 | May 11, 2026 | |
| Strategic reasoning in auction games | 1 | 1 | May 11, 2026 | |
| Alignment with human judgment on game verification | 1 |
| 1 |
| May 11, 2026 |
| Medical LLM Safety Refinement | 1 | 1 | Feb 18, 2026 |
|---|
| Aggregate Language and Logic Tasks | 1 | 1 | Feb 25, 2026 |
|---|
| Medical Dialogue Human-Centric Evaluation | 1 | 1 | May 11, 2026 |
|---|
| Reward Modeling (Logicality) | 1 | 1 | May 11, 2026 |
|---|
| Cinematic Story Generation | 1 | 1 | Feb 25, 2026 |
|---|
| Reward Modeling (Accuracy) | 1 | 1 | May 11, 2026 |
|---|
| Reward Modeling (Usefulness) | 1 | 1 | May 11, 2026 |
|---|
| Continuous Story Generation | 1 | 1 | Feb 25, 2026 |
|---|
| Instruction Alignment Evaluation | 1 | 1 | May 11, 2026 |
|---|
| Malicious Goal Attack (Longer Token Generation) | 1 | 1 | Feb 18, 2026 |
|---|
| Honesty Evaluation | 1 | 1 | Feb 25, 2026 |
|---|
| Conversational News Recommendation | 1 | 1 | May 11, 2026 |
|---|
| puzzle-4x4-task4 | 1 | 1 | May 11, 2026 |
|---|
| Zero-shot Multiple Choice Question Answering | 1 | 1 | Feb 25, 2026 |
|---|
| Tool Activation Probing | 1 | 1 | May 11, 2026 |
|---|
| Educational Feedback Generation | 1 | 1 | Feb 25, 2026 |
|---|
| Abductive Logical Reasoning | 1 | 1 | May 11, 2026 |
|---|