Commonsense Reasoning on StrategyQA and Big-Bench Hard Date Understanding (test)
71StrategyQA AccuracyStrategyLLM-SC
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| StrategyLLM-SCPrompting strategy=Strategy-based, Sampling strategy=Self-Consistency2023.11 | 71 | 74 | 72.5 | |
| CoT-SCPrompting strategy=Chain-of-Thought, Sampling strategy=Self-Consistency2023.11 | 70 | 70.5 | 70.3 | |
| StrategyLLM-ZSPrompting strategy=Strategy-based, Evaluation protocol=Zero-shot2023.11 | 70 | 72.5 | 71.3 | |
| StrategyLLMPrompting strategy=Strategy-based2023.11 | 67.5 | 68.5 | 68 | |
| CoTPrompting strategy=Chain-of-Thought2023.11 | 64 | 70.5 | 67.3 | |
| SolutionLLMMethod category=Baseline2023.11 | 59.5 | 52 | 55.8 | |
| SPPrompting strategy=Standard2023.11 | 56.5 | 48.5 | 52.5 |