Arithmetic Reasoning on AddSub (test)
96.71AccuracyCoT + Skill-Based
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT + Skill-BasedPrompting=Chain-of-Thought with Skill-Based exemplar selection, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 96.71 | |
| CoT + PALPrompting=Hybrid CoT + PAL, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 95.7 | |
| PALPrompting=Program-Aided Language Models, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 94.9 | |
| CoTPrompting=Chain-of-Thought, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 93.9 | |
| Ours (Natural Program (NP), No Verification)LLM=GPT-3.5-turbo, Prompting Format=Natural Program (NP), Verification Method=None, Voting Strategy=Majority Voting, k=102023.06 | 93.67 | |
| Ours (NP + Deductive Verification + UPV)LLM=GPT-3.5-turbo, Prompting Format=Natural Program (NP), Verification Method=Deductive Verification, Voting Strategy=UPV, k=102023.06 | 93.54 | |
| LoRAALL + MoDEKBackbone=LLaMA3-8B2024.10 | 93 | |
| LoRAALLBackbone=LLaMA3-8B2024.10 | 92.7 | |
| DoRA + MoDEKBackbone=LLaMA3-8B2024.10 | 92.4 | |
| CoT + VotingLLM=GPT-3.5-turbo, Prompting Format=Chain-of-Thought (CoT), Voting Strategy=Majority Voting, k=102023.06 | 92.36 | |
| DoRABackbone=LLaMA3-8B2024.10 | 92.2 | |
| LoRA¬K + MoDEK (+0.04%)Backbone=LLaMA3-8B2024.10 | 92.2 | |
| LoRA¬KBackbone=LLaMA3-8B2024.10 | 91.8 | |
| REFTBackbone=LLaMA3-8B2024.10 | 91 | |
| Faithful CoT + VotingLLM=GPT-3.5-turbo, Prompting Format=Faithful CoT, Voting Strategy=Majority Voting, k=102023.06 | 88.35 | |
| DoRA + MoDEKBackbone=LLaMA2-7B2024.10 | 51.6 | |
| DoRABackbone=LLaMA2-7B2024.10 | 51.4 | |
| LoRAALL + MoDEKBackbone=LLaMA2-7B2024.10 | 51.2 | |
| LoRAALLBackbone=LLaMA2-7B2024.10 | 51.1 | |
| LoRA¬K + MoDEK (+0.04%)Backbone=LLaMA2-7B2024.10 | 50.1 | |
| LoRA¬KBackbone=LLaMA2-7B2024.10 | 49.1 | |
| REFTBackbone=LLaMA2-7B2024.10 | 49 |