Domain Reasoning on LB
60AccuracyRM-Regen
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-RegenBase Model=GPT-3.52026.03 | 60 | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 56.84 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 55.79 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 54.74 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 53.68 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 50.53 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 48.42 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 48.42 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 46.32 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 46.32 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 44.21 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 44.21 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 42.11 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 41.05 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 40 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 38.95 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 38.95 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 38.95 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 28.42 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 25.26 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 25.26 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 23.16 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 22.11 |