Offline Reinforcement Learning on OGBench puzzle-4x4
56Success RateME-AM
Evaluation Results
| Method | Links | |
|---|---|---|
| ME-AMMethod Category=ADJOINT MATCHING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 56 | |
| FlowIQN-RPolicy Type=Flow Policies2026.05 | 47 | |
| QAM-EMethod Category=ADJOINT MATCHING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 37 | |
| FEditMethod Category=BACKPROP, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 35 | |
| IQNPolicy Type=Gaussian Policies2026.05 | 29 | |
| floqPolicy Type=Flow Policies2026.05 | 28 | |
| CGQLMethod Category=GUIDANCE, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 25 | |
| Value FlowsPolicy Type=Flow Policies2026.05 | 24 | |
| CODACPolicy Type=Gaussian Policies2026.05 | 20 | |
| FlowIQN-FQLPolicy Type=Flow Policies2026.05 | 17 | |
| FBRACMethod Category=BACKPROP, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 16 | |
| ReBRACPolicy Type=Gaussian Policies2026.05 | 14 | |
| FQLPolicy Type=Gaussian Policies2026.05 | 9 | |
| QAM-FMethod Category=ADJOINT MATCHING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 6 | |
| FQLMethod Category=BACKPROP, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 5 | |
| ReBRACMethod Category=GAUSSIAN, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| FAWACMethod Category=GAUSSIAN, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| CGQL-LMethod Category=GUIDANCE, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| CGQL-MMethod Category=GUIDANCE, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| DACMethod Category=GUIDANCE, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| QSMMethod Category=GUIDANCE, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| IFQLMethod Category=POST-PROCESSING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| DSRLMethod Category=POST-PROCESSING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| BAMMethod Category=ADJOINT MATCHING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| QAMMethod Category=ADJOINT MATCHING, Training Steps=1M, Seeds=8, Chunk Size (h)=52026.05 | 0 | |
| BCPolicy Type=Gaussian Policies2026.05 | 0 |