Reinforcement Learning on DMC Walker
961.5Walk ScoreFB-PbRL (Ours-FT)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| FB-PbRL (Ours-FT)fine-tuning=true2026.05 | 961.5 | 980.3 | 503.7 | 606 | 762.9 | |
| PSMReward Specification=Ground-truth reward, Protocol=Zero-Shot RFRL, Data Budget=10,000 transitions2026.05 | 891.4 | 872.6 | 351.5 | 640.8 | 689.1 | |
| FBReward Specification=Ground-truth reward, Protocol=Zero-Shot RFRL, Data Budget=10,000 transitions2026.05 | 823.6 | 922.1 | 444.1 | 689.7 | 719.9 | |
| RLDPReward Specification=Ground-truth reward, Protocol=Zero-Shot RFRL, Data Budget=10,000 transitions2026.05 | 790.9 | 877.7 | 324.9 | 492.9 | 621.6 | |
| Ours-FT (FB-PbRL)Reward Specification=Preferences, Fine-tuning=true, Data Budget=2,000 preference pairs2026.05 | 787 | 983.2 | 464.6 | 562.7 | 699.4 | |
| FB-PbRL (Ours-T)test-time CPTS=true2026.05 | 714.3 | 677.7 | 282.8 | 458.7 | 533.4 | |
| HILPReward Specification=Ground-truth reward, Protocol=Zero-Shot RFRL, Data Budget=10,000 transitions2026.05 | 399.7 | 607.1 | 107.8 | 278 | 348.1 | |
| OPPO2026.05 | 219.6 | 413.4 | 93.6 | 263.2 | 247.5 | |
| DPPO2026.05 | 218 | 403.1 | 91.8 | 256.4 | 242.3 | |
| OPRL2026.05 | 214.9 | 436.5 | 96.4 | 267.4 | 253.8 | |
| CLARIFY2026.05 | 214.3 | 423.1 | 94.8 | 263.3 | 248.9 | |
| LIRE2026.05 | 199.1 | 394.5 | 88.8 | 247.5 | 232.5 | |
| LaplaceReward Specification=Ground-truth reward, Protocol=Zero-Shot RFRL, Data Budget=10,000 transitions2026.05 | 190.5 | 243.7 | 63.7 | 48.7 | 136.7 |