Offline Reinforcement Learning on D4RL Adroit pen (cloned)
117.4Normalized ReturnMXQL
Evaluation Results
| Method | Links | |
|---|---|---|
| MXQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 117.4 | |
| QQLDomain=Adroit, Hyperparameter Tuning=consistent2025.11 | 115.2 | |
| IQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 114.1 | |
| XQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 112.6 | |
| FACPolicy Type=Flow, Seeds=82026.02 | 103.2 | |
| ReBRACModel paradigm=Model-free, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 102.8 | |
| BCDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 99.1 | |
| EPQPolicy Type=Gaussian, Seeds=82026.02 | 91.8 | |
| ReBRACPolicy Type=Gaussian, Seeds=82026.02 | 91.8 | |
| NEUBAYModel paradigm=Ours, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 91.3 | |
| IQLModel paradigm=Model-free, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 83.4 | |
| IQLPolicy Type=Gaussian, Seeds=82026.02 | 77.2 | |
| SPAR-PROJLearning Paradigm=Ours2026.05 | 76.2 | |
| DPPOsupervision=preference only, seeds=52023.01 | 75.1 | |
| MoMoModel paradigm=Conservative model-based, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 74.1 | |
| FQLPolicy Type=Flow, Seeds=82026.02 | 74 | |
| TD3+BCLearning Paradigm=Policy Gradient Guidance2026.05 | 71.4 | |
| VIPOModel paradigm=Conservative model-based, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 71.1 | |
| MOBILEModel paradigm=Conservative model-based, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 69 | |
| EDACModel paradigm=Model-free, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 68.2 | |
| QCS-R2024.02 | 66.5 | |
| Diff-QLLearning Paradigm=Policy Gradient Guidance2026.05 | 66.2 | |
| IDQLPolicy Type=Diffusion, Seeds=82026.02 | 64 | |
| IDQLLearning Paradigm=In-Support Learning2026.05 | 63.6 | |
| IQLLearning Paradigm=In-Support Learning2026.05 | 62.2 | |
| TD3+BCPolicy Type=Gaussian, Seeds=82026.02 | 61.4 | |
| SRPOPolicy Type=Diffusion, Seeds=82026.02 | 61 | |
| CACPolicy Type=Diffusion, Seeds=82026.02 | 56 | |
| BCQLearning Paradigm=Policy Gradient Guidance2026.05 | 55 | |
| LAPOLearning Paradigm=In-Support Learning2026.05 | 55 | |
| MOPOModel paradigm=Conservative model-based, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 54.6 | |
| TD3+BCDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 52.7 | |
| BCModel paradigm=Model-free, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 51.9 | |
| ARMORModel paradigm=Conservative model-based, pi_D=68.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 51.4 | |
| IQLsupervision=task rewards, seeds=52023.01 | 51.3 | |
| BaseLearning Paradigm=Ours2026.05 | 50.4 | |
| DCMethod Category=RCSL, Evaluation=official codebase2024.02 | 50 | |
| MCQPolicy Type=Gaussian, Seeds=82026.02 | 49.4 | |
| AWACLearning Paradigm=In-Support Learning2026.05 | 49.3 | |
| PLASLearning Paradigm=Policy Gradient Guidance2026.05 | 49 | |
| EQLLearning Paradigm=In-Support Learning2026.05 | 46.9 | |
| PT+IQLsupervision=preference only, seeds=52023.01 | 42.9 | |
| CQLsupervision=task rewards, seeds=52023.01 | 42.4 | |
| CQLLearning Paradigm=Policy Gradient Guidance2026.05 | 40.3 | |
| CQLMethod Category=Value-Based2024.02 | 39.2 | |
| CQLPolicy Type=Gaussian, Seeds=82026.02 | 39.2 | |
| IQLMethod Category=Value-Based2024.02 | 37.3 | |
| DTMethod Category=RCSL, Evaluation=official codebase2024.02 | 28.7 | |
| PT+CQLsupervision=preference only, seeds=52023.01 | 18.3 | |
| CQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 14.7 | |
| SAC-RNDPolicy Type=Gaussian, Seeds=82026.02 | 2.5 | |
| SPAR-MLPLearning Paradigm=Ours2026.05 | 0.1 | |
| CQL-AWLearning Paradigm=In-Support Learning2026.05 | -2.5 |