Preference Reconstruction on Metaphors
75.22Preference AccuracyDICAI
Evaluation Results
| Method | Links | |
|---|---|---|
| DICAIJudge LLM=GPT-52026.06 | 75.22 | |
| ICAIJudge LLM=GPT-52026.06 | 74.4 | |
| AutoRubricJudge LLM=GPT-52026.06 | 62.33 | |
| CoTJudge LLM=GPT-52026.06 | 60 | |
| CoT-SCJudge LLM=GPT-52026.06 | 60 | |
| ToTJudge LLM=GPT-52026.06 | 60 | |
| Self-RefineJudge LLM=GPT-52026.06 | 60 |