Offline Reinforcement Learning on D4RL Adroit pen (human)
128.3Normalized ReturnQQL
Evaluation Results
| Method | Links | |
|---|---|---|
| QQLDomain=Adroit, Hyperparameter Tuning=consistent2025.11 | 128.3 | |
| MXQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 122.1 | |
| IQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 106.2 | |
| XQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 105.3 | |
| ReBRACPolicy Type=Gaussian, Seeds=82026.02 | 103.5 | |
| ReBRACModel paradigm=Model-free, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 103.2 | |
| BCDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 99.7 | |
| QCS-R2024.02 | 83.9 | |
| EPQPolicy Type=Gaussian, Seeds=82026.02 | 83.9 | |
| TD3+BCPolicy Type=Gaussian, Seeds=82026.02 | 81.8 | |
| IQLPolicy Type=Gaussian, Seeds=82026.02 | 81.5 | |
| IQLModel paradigm=Model-free, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 78.5 | |
| DPPOsupervision=preference only, seeds=52023.01 | 76.3 | |
| IDQLPolicy Type=Diffusion, Seeds=82026.02 | 76 | |
| MoMoModel paradigm=Conservative model-based, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 74.9 | |
| DCMethod Category=RCSL, Evaluation=official codebase2024.02 | 74.2 | |
| FACPolicy Type=Flow, Seeds=82026.02 | 73.9 | |
| ARMORModel paradigm=Conservative model-based, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 72.8 | |
| IQLMethod Category=Value-Based2024.02 | 71.5 | |
| BCModel paradigm=Model-free, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 71 | |
| SRPOPolicy Type=Diffusion, Seeds=82026.02 | 69 | |
| MCQPolicy Type=Gaussian, Seeds=82026.02 | 68.5 | |
| CACPolicy Type=Diffusion, Seeds=82026.02 | 64 | |
| DTMethod Category=RCSL, Evaluation=official codebase2024.02 | 62.9 | |
| SPAR-PROJLearning Paradigm=Ours2026.05 | 62.7 | |
| CQLDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 58.9 | |
| Diff-QLLearning Paradigm=Policy Gradient Guidance2026.05 | 56.6 | |
| IQLsupervision=task rewards, seeds=52023.01 | 53.8 | |
| BaseLearning Paradigm=Ours2026.05 | 53.5 | |
| PT+IQLsupervision=preference only, seeds=52023.01 | 53 | |
| FQLPolicy Type=Flow, Seeds=82026.02 | 53 | |
| VIPOModel paradigm=Conservative model-based, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 52.6 | |
| EDACModel paradigm=Model-free, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 52.1 | |
| PLASLearning Paradigm=Policy Gradient Guidance2026.05 | 49.8 | |
| IDQLLearning Paradigm=In-Support Learning2026.05 | 49.8 | |
| IQLLearning Paradigm=In-Support Learning2026.05 | 47.5 | |
| EQLLearning Paradigm=In-Support Learning2026.05 | 44.3 | |
| CQLsupervision=task rewards, seeds=52023.01 | 44.2 | |
| CQLMethod Category=Value-Based2024.02 | 37.5 | |
| CQLPolicy Type=Gaussian, Seeds=82026.02 | 37.5 | |
| PT+CQLsupervision=preference only, seeds=52023.01 | 31.6 | |
| MOBILEModel paradigm=Conservative model-based, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 30.1 | |
| NEUBAYModel paradigm=Ours, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 20.8 | |
| CQLLearning Paradigm=Policy Gradient Guidance2026.05 | 14.9 | |
| MOPOModel paradigm=Conservative model-based, pi_D=88.7, Number of seeds=6, Evaluation episodes=20, Evaluation step=final step2025.12 | 10.7 | |
| TD3+BCDomain=Adroit, Hyperparameter Tuning=individually tuned2025.11 | 10 | |
| SAC-RNDPolicy Type=Gaussian, Seeds=82026.02 | 5.6 | |
| BCQLearning Paradigm=Policy Gradient Guidance2026.05 | 2.2 | |
| LAPOLearning Paradigm=In-Support Learning2026.05 | 2.2 | |
| AWACLearning Paradigm=In-Support Learning2026.05 | 1 | |
| SPAR-MLPLearning Paradigm=Ours2026.05 | 0.1 | |
| TD3+BCLearning Paradigm=Policy Gradient Guidance2026.05 | 0 | |
| CQL-AWLearning Paradigm=In-Support Learning2026.05 | -3 |