Preference Reconstruction on LiTBench Long Stories
71.9Preference AccuracySelf-Refine
Evaluation Results
| Method | Links | |
|---|---|---|
| Self-RefineJudge Model=GPT-4o2026.06 | 71.9 | |
| ToTJudge LLM=GPT-52026.06 | 71.29 | |
| DICAIJudge LLM=GPT-52026.06 | 71.2 | |
| CoTJudge LLM=GPT-52026.06 | 70.69 | |
| CoTJudge Model=GPT-4o2026.06 | 70.69 | |
| ToTJudge Model=GPT-4o2026.06 | 69.28 | |
| DICAIJudge Model=GPT-4o2026.06 | 68.7 | |
| CoT-SCJudge LLM=GPT-52026.06 | 68.43 | |
| Self-RefineJudge LLM=GPT-52026.06 | 68.02 | |
| ICAIJudge LLM=GPT-52026.06 | 66.8 | |
| CoT-SCJudge Model=GPT-4o2026.06 | 66.15 | |
| AutoRubricJudge LLM=GPT-52026.06 | 63.21 | |
| ICAIJudge Model=GPT-4o2026.06 | 62.89 | |
| AutoRubricJudge Model=GPT-4o2026.06 | 59.22 |