Science Simulation on Sciworld
82.6Progress RateExplicit RM
Evaluation Results
| Method | Links | |
|---|---|---|
| Explicit RMInference Strategy=Beam Search, Beam Search Weights (W1, W2)=W1 = 5, W2 = 52025.02 | 82.6 | |
| Explicit RMInference Strategy=Best-of-52025.02 | 76.1 | |
| ImplicitPRMInference Strategy=Best-of-52025.02 | 70.6 | |
| Agent-R2025.02 | 70.2 | |
| gpt-4o2025.02 | 66.6 | |
| Greedy Search2025.02 | 66.6 | |
| QLASS2025.02 | 66.4 | |
| StepAgent2025.02 | 64.1 | |
| ETO2025.02 | 62.5 | |
| LLM-as-a-judgeInference Strategy=Best-of-52025.02 | 62.3 | |
| SPIN2025.02 | 60.3 | |
| NAT2025.02 | 55.6 |