Loading the SOTA2 catalog…
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity · SOTA2 Research