Question Answering on StrategyQA (accuracy)
90.4AccuracyPaLM2
Evaluation Results
| Method | Links | |
|---|---|---|
| PaLM2Backbone=PaLM2 (340B), Retrieve=✗, Specification=-2026.03 | 90.4 | |
| Gt-MarginModel=LLaDA-8B, Completion length=1282026.02 | 84 | |
| Gt-MarginModel=Dream-7B, Completion length=1282026.02 | 83.5 | |
| PaLMBackbone=PaLM (540B), Retrieve=✗, Specification=CoT + SC2026.03 | 81.6 | |
| Top-ProbModel=Dream-7B, Completion length=1282026.02 | 78.5 | |
| GEEKBackbone=Flan-T5 (11B), Retrieve=✓, Specification=CoT+SE2026.03 | 78.17 | |
| RRBackbone=text-davinci-002 (175B), Retrieve=✓, Specification=CoT2026.03 | 77.73 | |
| Xie et al., 2023Backbone=code-davinci-002 (175B), Retrieve=✗, Specification=-2026.03 | 77.2 | |
| MarginModel=Dream-7B, Completion length=1282026.02 | 76.5 | |
| GEEKBackbone=Flan-T5 (11B), Retrieve=✓, Specification=CoT2026.03 | 75.98 | |
| PaLMBackbone=PaLM (540B), Retrieve=✗, Specification=-2026.03 | 73.9 | |
| FaithfulCoTBackbone=code-davinci-002 (175B), Retrieve=✗, Specification=-2026.03 | 73.2 | |
| Gt-ProbModel=Dream-7B, Completion length=1282026.02 | 70.5 | |
| ViscondeBackbone=text-davinci-002 (175B), Retrieve=✓, Specification=CoT2026.03 | 69.43 | |
| Lazaridou et al., 2022Backbone=Gopher (280B), Retrieve=✓, Specification=-2026.03 | 66.2 | |
| MarginModel=LLaDA-8B, Completion length=1282026.02 | 65.5 | |
| RandomModel=Dream-7B, Completion length=1282026.02 | 64 | |
| ARModel=LLaDA-8B, Completion length=1282026.02 | 63.5 | |
| Inverse-ARModel=LLaDA-8B, Completion length=1282026.02 | 63.5 | |
| Top-ProbModel=LLaDA-8B, Completion length=1282026.02 | 63.5 | |
| ChatGPTBackbone=GPT-3.5 (175B), Retrieve=✗, Specification=CoT2026.03 | 62.5 | |
| ARModel=Dream-7B, Completion length=1282026.02 | 60.9 | |
| ChatGPTBackbone=GPT-3.5 (175B), Retrieve=✗, Specification=Without CoT2026.03 | 59.2 | |
| Gt-ProbModel=LLaDA-8B, Completion length=1282026.02 | 50.5 | |
| RandomModel=LLaDA-8B, Completion length=1282026.02 | 41.5 | |
| Inverse-ARModel=Dream-7B, Completion length=1282026.02 | 1.5 |