Preference Reconstruction on Alternate Uses of Objects
78.61Preference AccuracyCoT
Evaluation Results
| Method | Links | |
|---|---|---|
| CoTJudge Model=GPT-4o2026.06 | 78.61 | |
| ToTJudge LLM=GPT-52026.06 | 77.72 | |
| ToTJudge Model=GPT-4o2026.06 | 76.46 | |
| Self-RefineJudge LLM=GPT-52026.06 | 75.35 | |
| CoTJudge LLM=GPT-52026.06 | 74.94 | |
| DICAIJudge Model=GPT-4o2026.06 | 74.23 | |
| Self-RefineJudge Model=GPT-4o2026.06 | 73.78 | |
| DICAIJudge LLM=GPT-52026.06 | 73.22 | |
| CoT-SCJudge LLM=GPT-52026.06 | 70.77 | |
| CoT-SCJudge Model=GPT-4o2026.06 | 70.3 | |
| ICAIJudge LLM=GPT-52026.06 | 69.4 | |
| ICAIJudge Model=GPT-4o2026.06 | 66.4 | |
| AutoRubricJudge LLM=GPT-52026.06 | 59.22 | |
| AutoRubricJudge Model=GPT-4o2026.06 | 58.7 |