Pointwise evaluation on BIGGEN
0.584Spearman CorrHuman-crafted (existing) Rubrics
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Human-crafted (existing) RubricsJudge Model=Claude Sonnet 4, Rubric Granularity=Instance-specific2026.05 | 0.584 | 0.598 | |
| Human-crafted (existing) RubricsJudge Model=Qwen3 14B, Rubric Granularity=Instance-specific2026.05 | 0.583 | 0.609 | |
| Human-crafted (existing) RubricsJudge Model=Llama 3.1 70B, Rubric Granularity=Instance-specific2026.05 | 0.552 | 0.594 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Claude Sonnet 4, Iteration=2, Generator Model=Qwen3 14B2026.05 | 0.51 | 0.521 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Claude Sonnet 4, Iteration=1, Generator Model=Qwen3 14B2026.05 | 0.49 | 0.504 | |
| Training-free Rubric GenerationJudge Model=Claude Sonnet 4, Rubric Granularity=Instance-specific2026.05 | 0.477 | 0.496 | |
| Training-free Rubric GenerationJudge Model=Claude Sonnet 4, Configuration=Highest Training-free2026.05 | 0.477 | 0.496 | |
| Human-crafted (existing) RubricsJudge Model=Claude Sonnet 4, Rubric Granularity=Dataset-specific2026.05 | 0.476 | 0.494 | |
| Training-free Rubric GenerationJudge Model=Claude Sonnet 4, Rubric Granularity=Dataset-specific2026.05 | 0.461 | 0.453 | |
| DnA-EvalJudge Model=Claude Sonnet 4, Approach Type=Training-free Approach2026.05 | 0.454 | 0.489 | |
| Training-free Rubric GenerationJudge Model=Qwen3 14B, Rubric Granularity=Instance-specific2026.05 | 0.446 | 0.462 | |
| Training-free Rubric GenerationJudge Model=Qwen3 14B, Configuration=Highest Training-free2026.05 | 0.446 | 0.462 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Llama 3.1 70B, Iteration=2, Generator Model=Qwen3 14B2026.05 | 0.445 | 0.452 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Qwen3 14B, Iteration=2, Generator Model=Qwen3 14B2026.05 | 0.441 | 0.461 | |
| Training-free Rubric GenerationJudge Model=Qwen3 14B, Rubric Granularity=Dataset-specific2026.05 | 0.44 | 0.455 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Llama 3.1 70B, Iteration=1, Generator Model=Qwen3 14B2026.05 | 0.44 | 0.447 | |
| Annotation-free Preference Learning for Rubric GeneratorJudge Model=Qwen3 14B, Iteration=1, Generator Model=Qwen3 14B2026.05 | 0.439 | 0.459 | |
| Human-crafted (existing) RubricsJudge Model=Qwen3 14B, Rubric Granularity=Dataset-specific2026.05 | 0.436 | 0.45 | |
| DnA-EvalJudge Model=Qwen3 14B, Approach Type=Training-free Approach2026.05 | 0.431 | 0.426 | |
| CheckEvalJudge Model=Claude Sonnet 4, Approach Type=Training-free Approach2026.05 | 0.424 | 0.455 | |
| Training-free Rubric GenerationJudge Model=Llama 3.1 70B, Rubric Granularity=Instance-specific2026.05 | 0.42 | 0.434 | |
| Training-free Rubric GenerationJudge Model=Llama 3.1 70B, Configuration=Highest Training-free2026.05 | 0.42 | 0.434 | |
| Human-crafted (existing) RubricsJudge Model=Llama 3.1 70B, Rubric Granularity=Dataset-specific2026.05 | 0.384 | 0.382 | |
| Training-free Rubric GenerationJudge Model=Llama 3.1 70B, Rubric Granularity=Dataset-specific2026.05 | 0.378 | 0.361 | |
| Prometheus 8x7BApproach Type=Fine-tuned Model, Rubric Granularity=Instance-specific2026.05 | 0.366 | 0.401 | |
| CheckEvalJudge Model=Qwen3 14B, Approach Type=Training-free Approach2026.05 | 0.343 | 0.392 | |
| DnA-EvalJudge Model=Llama 3.1 70B, Approach Type=Training-free Approach2026.05 | 0.337 | 0.353 | |
| RubricHubJudge Model=Qwen3 14B, Approach Type=Training-free Approach2026.05 | 0.332 | 0.383 | |
| RubricHubJudge Model=Claude Sonnet 4, Approach Type=Training-free Approach2026.05 | 0.325 | 0.336 | |
| CheckEvalJudge Model=Llama 3.1 70B, Approach Type=Training-free Approach2026.05 | 0.318 | 0.361 | |
| RubricHubJudge Model=Llama 3.1 70B, Approach Type=Training-free Approach2026.05 | 0.318 | 0.361 | |
| Prometheus 8x7BApproach Type=Fine-tuned Model, Rubric Granularity=Dataset-specific2026.05 | 0.28 | 0.307 |