Domain Reasoning on Headline
73AccuracyRM-Regen
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 73 | |
| RM-RegenBase Model=GPT-3.52026.03 | 71 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 70 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 70 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 69 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 69 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 68 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 66 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 66 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 66 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 65 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 65 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 64 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 63 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 62 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 62 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 60 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 59 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 53 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 48 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 48 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 48 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 48 |