Value-Order Correlation on RoboRewardBench (held-out)
0.966Spearman VOCLLM-as-a-Verifier
Evaluation Results
| Method | Links | |
|---|---|---|
| LLM-as-a-VerifierBackbone=Qwen 3.6 35B, Repeated evaluations (K)=5, Scoring granularity (G)=202026.07 | 0.966 | |
| RoboReward-8Bparameters=8B2026.07 | 0.877 | |
| Robometer-4Bparameters=4B2026.07 | 0.78 | |
| TOPRewardBackbone=Qwen 3.6, variant=P(true)2026.07 | 0.565 |