Pointwise evaluation on HelpSteer2
0.464Spearman CorrelationAnnotation-free Preference Learning for Rubric Generator
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Claude Sonnet 4, Iteration=2, Generator Model=Qwen3 14B2026.05 | 0.464 | 0.503 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Claude Sonnet 4, Iteration=1, Generator Model=Qwen3 14B2026.05 | 0.44 | 0.471 | |
| Training-free Rubric GenerationJudge Model=Claude Sonnet 4, Rubric Granularity=Instance-specific2026.05 | 0.438 | 0.488 | |
| Training-free Rubric GenerationJudge Model=Claude Sonnet 4, Configuration=Highest Training-free2026.05 | 0.438 | 0.488 | |
| Human-crafted (existing) RubricsJudge Model=Claude Sonnet 4, Rubric Granularity=Dataset-specific2026.05 | 0.432 | 0.483 | |
| DnA-EvalJudge Model=Claude Sonnet 4, Approach Type=Training-free Approach2026.05 | 0.411 | 0.476 | |
| Training-free Rubric GenerationJudge Model=Claude Sonnet 4, Rubric Granularity=Dataset-specific2026.05 | 0.41 | 0.422 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Qwen3 14B, Iteration=2, Generator Model=Qwen3 14B2026.05 | 0.406 | 0.441 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Qwen3 14B, Iteration=1, Generator Model=Qwen3 14B2026.05 | 0.394 | 0.431 | |
| CheckEvalJudge Model=Claude Sonnet 4, Approach Type=Training-free Approach2026.05 | 0.375 | 0.435 | |
| Training-free Rubric GenerationJudge Model=Qwen3 14B, Rubric Granularity=Instance-specific2026.05 | 0.374 | 0.416 | |
| Training-free Rubric GenerationJudge Model=Qwen3 14B, Configuration=Highest Training-free2026.05 | 0.374 | 0.416 | |
| Training-free Rubric GenerationJudge Model=Qwen3 14B, Rubric Granularity=Dataset-specific2026.05 | 0.353 | 0.428 | |
| Human-crafted (existing) RubricsJudge Model=Qwen3 14B, Rubric Granularity=Dataset-specific2026.05 | 0.352 | 0.431 | |
| DnA-EvalJudge Model=Qwen3 14B, Approach Type=Training-free Approach2026.05 | 0.345 | 0.403 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Llama 3.1 70B, Iteration=2, Generator Model=Qwen3 14B2026.05 | 0.344 | 0.412 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Llama 3.1 70B, Iteration=1, Generator Model=Qwen3 14B2026.05 | 0.326 | 0.388 | |
| Training-free Rubric GenerationJudge Model=Llama 3.1 70B, Rubric Granularity=Instance-specific2026.05 | 0.319 | 0.379 | |
| Training-free Rubric GenerationJudge Model=Llama 3.1 70B, Configuration=Highest Training-free2026.05 | 0.319 | 0.379 | |
| RubricHubJudge Model=Qwen3 14B, Approach Type=Training-free Approach2026.05 | 0.286 | 0.319 | |
| Training-free Rubric GenerationJudge Model=Llama 3.1 70B, Rubric Granularity=Dataset-specific2026.05 | 0.284 | 0.377 | |
| CheckEvalJudge Model=Qwen3 14B, Approach Type=Training-free Approach2026.05 | 0.282 | 0.392 | |
| Human-crafted (existing) RubricsJudge Model=Llama 3.1 70B, Rubric Granularity=Dataset-specific2026.05 | 0.279 | 0.373 | |
| DnA-EvalJudge Model=Llama 3.1 70B, Approach Type=Training-free Approach2026.05 | 0.263 | 0.299 | |
| RubricHubJudge Model=Llama 3.1 70B, Approach Type=Training-free Approach2026.05 | 0.246 | 0.292 | |
| CheckEvalJudge Model=Llama 3.1 70B, Approach Type=Training-free Approach2026.05 | 0.239 | 0.324 | |
| RubricHubJudge Model=Claude Sonnet 4, Approach Type=Training-free Approach2026.05 | 0.239 | 0.27 | |
| Prometheus 8x7BApproach Type=Fine-tuned Model, Rubric Granularity=Dataset-specific2026.05 | 0.229 | 0.308 |