Domain-specific Reasoning on LegalBench
85.26AccuracyReflexion(3 iters)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 85.26 | — | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 81.05 | — | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 78.95 | — | |
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 75.79 | — | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 75.79 | — | |
| RM-PrimedModel=GPT-3.52026.03 | 67.37 | 66.01 | |
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 66.32 | — | |
| RM-Primed (R+)Model=GPT-3.5, R+ selection=only correct entries used2026.03 | 64.21 | 64.99 | |
| RM-PrimedModel=Llama3-8B2026.03 | 64.21 | 68.23 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 64.21 | — | |
| RM-Primed (R+)Model=Llama3-8B, R+ selection=only correct entries used2026.03 | 62.11 | 65.94 | |
| RM-RegenModel=GPT-3.52026.03 | 60 | — | |
| Contrastive CoTModel=Llama3-8B, Reflection=without2026.03 | 57.89 | 59.56 | |
| RM-RegenModel=Llama3.1-8B2026.03 | 56.84 | — | |
| Contrastive CoTModel=GPT-3.5, Reflection=with2026.03 | 55.79 | 60.33 | |
| ST CoTModel=Llama3.1-8B, iterations=32026.03 | 55.79 | — | |
| Contrastive CoTModel=Llama3-8B, Reflection=with2026.03 | 53.68 | 61.13 | |
| RM-RegenModel=Gemma2-9B2026.03 | 51.58 | — | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 50.53 | — | |
| Few-shot CoTModel=Llama3-8B2026.03 | 48.42 | 64.19 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 48.42 | — | |
| Self-RefineModel=Llama3.1-8B, iterations=32026.03 | 46.32 | — | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 45.26 | — | |
| ST CoTModel=Gemma2-9B, iterations=32026.03 | 45.26 | — | |
| Few-shot CoTModel=GPT-3.52026.03 | 44.21 | 61.15 | |
| ProCoModel=Gemma2-9B, iterations=32026.03 | 44.21 | — | |
| ProCoModel=GPT-3.5, iterations=32026.03 | 41.05 | — | |
| Self-RefineModel=Gemma2-9B, iterations=32026.03 | 41.05 | — | |
| ProCoModel=Llama3.1-8B, iterations=32026.03 | 38.95 | — | |
| Contrastive CoTModel=GPT-3.5, Reflection=without2026.03 | 35.75 | 57.01 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 32.63 | — | |
| Self-RefineModel=GPT-3.5, iterations=32026.03 | 28.42 | — | |
| ST CoTModel=GPT-3.5, iterations=32026.03 | 23.16 | — |