Reinforcement Learning on MuJoCo Reacher v4
77Normalized PerformanceLLM One-Shot
Evaluation Results
| Method | Links | |
|---|---|---|
| LLM One-ShotTraining Episodes=1,0002026.05 | 77 | |
| Hand-CraftedTraining Episodes=1,0002026.05 | 75.4 | |
| LLM IterativeTraining Episodes=1,000, Refinement Type=standard diagnostic refinement2026.05 | 73.7 | |
| RNDTraining Episodes=1,0002026.05 | 67.5 | |
| No ShapingTraining Episodes=1,0002026.05 | 65 |