Symbolic Reasoning on Date
90.5Solve RateAutomatic Model Selection with LLMs
Evaluation Results
| Method | Links | |
|---|---|---|
| Automatic Model Selection with LLMsBackbone=GPT-4, Decoding strategy=Greedy2023.05 | 90.5 | |
| CoTBackbone=GPT-4, Decoding strategy=Greedy2023.05 | 90 | |
| PALBackbone=GPT-4, Decoding strategy=Greedy2023.05 | 88.1 | |
| Automatic Model Selection with LLMsBackbone=Codex, Decoding strategy=Greedy2023.05 | 79.4 | |
| PALBackbone=Codex, Decoding strategy=Greedy2023.05 | 77.5 | |
| PAL CodexPrompting Strategy=Program-aided Language Models, Backbone Model=Codex2022.11 | 76.2 | |
| Automatic Model Selection with LLMsBackbone=ChatGPT, Decoding strategy=Greedy2023.05 | 70.2 | |
| CoTBackbone=ChatGPT, Decoding strategy=Greedy2023.05 | 69.1 | |
| PALBackbone=ChatGPT, Decoding strategy=Greedy2023.05 | 68.3 | |
| COT PaLM-540BPrompting Strategy=Chain-of-thought, Backbone Model=PaLM-540B2022.11 | 65.3 | |
| CoTBackbone=Codex, Decoding strategy=Greedy2023.05 | 64.5 | |
| COT CodexPrompting Strategy=Chain-of-thought, Backbone Model=Codex2022.11 | 61.8 | |
| DIRECT CodexPrompting Strategy=Direct, Backbone Model=Codex2022.11 | 49.9 | |
| COT LaMDA-137BPrompting Strategy=Chain-of-thought, Backbone Model=LaMDA-137B2022.11 | 26.8 |