Constrained Reinforcement Learning for Tutoring Curricula on neural environment 25-concept
34.79ReturnUnconstrained
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Unconstrained2026.04 | 34.79 | 70.52 | 0 | |
| Reward Shapedlambda={0.05, 0.1, 0.2, 0.5, 5.0}2026.04 | 34.79 | 70.52 | 0 | |
| Posthoc2026.04 | 34.79 | 70.52 | 0 | |
| MC-CPO (no frontier)epsilon_min=02026.04 | 32.11 | 40.05 | 70 | |
| MC-CPOepsilon_min=0.052026.04 | 31.97 | 39.16 | 80 |