Common-sense Reasoning on Bamboogle
62AccuracyRM-Primed
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RM-PrimedModel=GPT-3.52026.03 | 62 | 66.01 | |
| RM-Primed (R+)Model=GPT-3.5, R+ selection=only correct entries used2026.03 | 59 | 64.99 | |
| Few-shot CoTModel=GPT-3.52026.03 | 55 | 61.15 | |
| RM-PrimedModel=Llama3-8B2026.03 | 54 | 68.23 | |
| Few-shot CoTModel=Llama3-8B2026.03 | 51 | 64.19 | |
| RM-Primed (R+)Model=Llama3-8B, R+ selection=only correct entries used2026.03 | 51 | 65.94 | |
| Contrastive CoTModel=GPT-3.5, Reflection=with2026.03 | 50 | 60.33 | |
| Contrastive CoTModel=GPT-3.5, Reflection=without2026.03 | 49 | 57.01 | |
| Contrastive CoTModel=Llama3-8B, Reflection=with2026.03 | 29 | 61.13 | |
| Contrastive CoTModel=Llama3-8B, Reflection=without2026.03 | 25 | 59.56 |