Gap discovery quality on Scientist-Bench 27 tasks
5Gaps/TaskAI-Supervisor (RWM)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AI-Supervisor (RWM)backbone=Qwen-72B-Instruct2026.03 | 5 | 80.7 | 100 | 4.44 | |
| LLM-only brainstormbackbone=Qwen-72B-Instruct, simulation_target=AI Scientist v22026.03 | 4.9 | 67.9 | 92.6 | 4.15 | |
| Divergent-convergentbackbone=Qwen-72B-Instruct, simulation_target=AI-Researcher2026.03 | 2 | 75.5 | 92.6 | 4.04 |