Domain Reasoning on HL
75AccuracyBest-of-N (N=3)
Evaluation Results
| Method | Links | |
|---|---|---|
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 75 | |
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 75 | |
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 73 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 73 | |
| RM-RegenModel=Llama3.1-8B2026.03 | 73 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 72 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 72 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 71 | |
| RM-RegenModel=GPT-3.52026.03 | 71 | |
| RM-RegenModel=Gemma2-9B2026.03 | 71 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 70 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 70 | |
| ST CoTModel=GPT-3.5, iterations=32026.03 | 69 | |
| ProCoModel=Gemma2-9B, iterations=32026.03 | 69 | |
| ST CoTModel=Llama3.1-8B, iterations=32026.03 | 68 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 67 | |
| ST CoTModel=Gemma2-9B, iterations=32026.03 | 66 | |
| ProCoModel=GPT-3.5, iterations=32026.03 | 65 | |
| Self-RefineModel=Llama3.1-8B, iterations=32026.03 | 65 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 62 | |
| Self-RefineModel=Gemma2-9B, iterations=32026.03 | 59 | |
| Self-RefineModel=GPT-3.5, iterations=32026.03 | 53 | |
| ProCoModel=Llama3.1-8B, iterations=32026.03 | 48 |