Reward Modeling on UltraFeedback (test)
0.145MAESelectiveRM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SelectiveRMMethod Category=Ours2026.05 | 0.145 | 0.107 | 0.461 | — | |
| SelectMixLearning Category=PU learning methods2026.03 | 0.1679 | — | 0.4752 | 0.3316 | |
| UPLLearning Category=PU learning methods2026.03 | 0.1723 | — | 0.4309 | 0.3453 | |
| HOCMethod Category=Statistically Consistent Methods2026.05 | 0.181 | 0.115 | 0.418 | — | |
| CCRMethod Category=Statistically Consistent Methods2026.05 | 0.181 | 0.114 | 0.424 | — | |
| ImplicitRM2026.03 | 0.1961 | — | 0.5207 | 0.3169 | |
| NNPULearning Category=PU learning methods2026.03 | 0.2021 | — | 0.4668 | 0.3342 | |
| ILDELearning Category=PU learning methods2026.03 | 0.2021 | — | 0.4403 | 0.3424 | |
| ILDEScenario=1, rho01=0.2, rho10=0.12026.03 | 0.208 | 0.111 | 0.47 | — | |
| DRLearning Category=Debiased learning methods2026.03 | 0.216 | — | 0.2305 | 0.4015 | |
| CausalRM-DRScenario=1, rho01=0.2, rho10=0.12026.03 | 0.225 | 0.106 | 0.495 | — | |
| UPULearning Category=PU learning methods2026.03 | 0.2263 | — | 0.4559 | 0.3376 | |
| CausalRM-DRScenario=2, rho01=0.1, rho10=0.22026.03 | 0.228 | 0.106 | 0.496 | — | |
| CUBPRLearning Category=PU learning methods2026.03 | 0.23 | — | 0.4389 | 0.3429 | |
| Robust DivideMixMethod Category=Heuristic Methods2026.05 | 0.239 | 0.109 | 0.451 | — | |
| CSGNMethod Category=Statistically Consistent Methods2026.05 | 0.241 | 0.112 | 0.436 | — | |
| UBPRLearning Category=PU learning methods2026.03 | 0.2452 | — | 0.4141 | 0.3504 | |
| LAGAMLearning Category=PU learning methods2026.03 | 0.2488 | — | 0.4652 | 0.3347 | |
| F-correctionMethod Category=Statistically Consistent Methods2026.05 | 0.267 | 0.117 | 0.409 | — | |
| SelectMixMethod Category=Heuristic Methods2026.05 | 0.267 | 0.108 | 0.453 | — | |
| IPSLearning Category=Debiased learning methods2026.03 | 0.2713 | — | 0.1476 | 0.4103 | |
| SDRScenario=1, rho01=0.2, rho10=0.12026.03 | 0.272 | 0.113 | 0.461 | — | |
| ϵ-SoftmaxMethod Category=Heuristic Methods2026.05 | 0.274 | 0.109 | 0.449 | — | |
| ILDEMethod Category=Heuristic Methods2026.05 | 0.276 | 0.108 | 0.456 | — | |
| CNLCUMethod Category=Heuristic Methods2026.05 | 0.283 | 0.112 | 0.435 | — | |
| LabelWaveMethod Category=Heuristic Methods2026.05 | 0.284 | 0.11 | 0.446 | — | |
| NLSMethod Category=Heuristic Methods2026.05 | 0.29 | 0.111 | 0.441 | — | |
| Co-TeachingMethod Category=Heuristic Methods2026.05 | 0.293 | 0.116 | 0.416 | — | |
| kMEIDTMMethod Category=Statistically Consistent Methods2026.05 | 0.295 | 0.113 | 0.428 | — | |
| ROBOTMethod Category=Statistically Consistent Methods2026.05 | 0.295 | 0.113 | 0.431 | — | |
| NaiveScenario=1, rho01=0.2, rho10=0.12026.03 | 0.297 | 0.125 | 0.405 | — | |
| CDRMethod Category=Heuristic Methods2026.05 | 0.301 | 0.114 | 0.422 | — | |
| SDRLearning Category=Debiased learning methods2026.03 | 0.3049 | — | 0.3646 | 0.3648 | |
| NaiveMethod Category=Statistically Consistent Methods2026.05 | 0.315 | 0.12 | 0.395 | — | |
| BPRLearning Category=PU learning methods2026.03 | 0.3212 | — | 0.3839 | 0.3593 | |
| NaiveLearning Category=Debiased learning methods2026.03 | 0.3492 | — | 0.1459 | 0.423 | |
| MTDRLearning Category=Debiased learning methods2026.03 | 0.3534 | — | 0.3201 | 0.3774 | |
| MTIPSLearning Category=Debiased learning methods2026.03 | 0.3564 | — | 0.2525 | 0.3957 |