Preference Classification on WebGPT comparisons (test)
60.8AccuracyUMM-RM
Evaluation Results
| Method | Links | |
|---|---|---|
| UMM-RMBackbone=TinyLlama-1.1B, Number of experts=62025.11 | 60.8 | |
| Worst-Case OptimizationBackbone=TinyLlama-1.1B, Ensemble Strategy=ensemble RM2025.11 | 60.6 | |
| Uncertainty-Weighted OptimizationBackbone=TinyLlama-1.1B, Ensemble Strategy=ensemble RM2025.11 | 59.6 | |
| UMM-RMBackbone=TinyLlama-1.1B, Number of experts=42025.11 | 58.6 | |
| UMM-RMBackbone=TinyLlama-1.1B, Number of experts=22025.11 | 57.8 | |
| Dense RMBackbone=TinyLlama-1.1B2025.11 | 52.2 | |
| Mean OptimizationBackbone=TinyLlama-1.1B, Ensemble Strategy=ensemble RM2025.11 | 51.4 |