Logical Reasoning on LogiQA
50.23AccuracyBase
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BaseModel=Qwen1.5-32B, Infer Mode=Base2024.05 | 50.23 | — | |
| GLM-130BSetting=Zero-Shot Prompting2023.06 | 50 | — | |
| LoRAModel=Qwen1.5-32B, Infer Mode=LoRA2024.05 | 49.31 | — | |
| Self-rethinkingModel=PaLM2, Prompting Strategy=Self-rethinking, k=12024.03 | 49.12 | — | |
| TAIAModel=Qwen1.5-32B, Infer Mode=TAIA2024.05 | 48.69 | — | |
| BaseModel=Qwen1.5-14B, Infer Mode=Base2024.05 | 47.16 | — | |
| TAIAModel=Qwen1.5-14B, Infer Mode=TAIA2024.05 | 46.39 | — | |
| CR2023.08 | 45.25 | 17 | |
| IDOLSetting=Fine-Tuning2023.06 | 43.3 | — | |
| ToT2023.08 | 43.02 | 19.87 | |
| Self-consistencyModel=PaLM2, Prompting Strategy=Self-consistency, inference times=32024.03 | 42.88 | — | |
| BaseModel=Qwen1.5-7B, Infer Mode=Base2024.05 | 42.7 | — | |
| TAIAModel=Qwen1.5-7B, Infer Mode=TAIA2024.05 | 41.78 | — | |
| LoRAModel=Qwen1.5-14B, Infer Mode=LoRA2024.05 | 41.78 | — | |
| Standard PromptingModel=PaLM2, Prompting Strategy=Standard Prompting2024.03 | 41.21 | — | |
| CoTModel=PaLM2, Prompting Strategy=Chain-of-Thought2024.03 | 41.05 | — | |
| CoT-SC2023.08 | 40.43 | 16 | |
| CoT2023.08 | 38.55 | 1 | |
| LoRAModel=Qwen1.5-7B, Infer Mode=LoRA2024.05 | 37.33 | — | |
| ChatGPTSetting=Few-Shot Prompting2023.06 | 36.7 | — | |
| GLM-130BSetting=Few-Shot Prompting2023.06 | 36.7 | — | |
| Self-refineModel=PaLM2, Prompting Strategy=Self-refine2024.03 | 35.99 | — | |
| L2Training Dataset=Alpaca-GPT4, Infer Mode=L22024.05 | 35.33 | — | |
| VanillaTraining Dataset=Alpaca-GPT4, Infer Mode=Vanilla2024.05 | 35.02 | — | |
| LoRACLTraining Dataset=Alpaca-GPT4, Infer Mode=LORACL2024.05 | 34.41 | — | |
| EWCTraining Dataset=Alpaca-GPT4, Infer Mode=EWC2024.05 | 34.1 | — | |
| ChatGPTSetting=Zero-Shot Prompting2023.06 | 33.3 | — | |
| TAIATraining Dataset=Alpaca-GPT4, Infer Mode=TAIA2024.05 | 33.03 | — | |
| TAIATraining Dataset=CoT-Collection, Infer Mode=TAIA2024.05 | 32.57 | — | |
| VanillaTraining Dataset=Base Model, Infer Mode=Vanilla2024.05 | 32.41 | — | |
| Direct2023.08 | 31.69 | 1 | |
| Llama-2#Param=6.9B, Training Tokens=2T, Shots=zero-shot2024.10 | 31 | — | |
| LLaMA2-7BNumber of Parameters=7B, Backbone=LLaMA2, Shots=02024.07 | 30.4 | — | |
| GPT-3.5Setting=Zero-Shot Prompting2023.06 | 30 | — | |
| Read-ME#Param=4.7B-17B, Training Tokens=1B, Shots=zero-shot2024.10 | 29.7 | — | |
| LLM-Pruner-1.3BNumber of Parameters=1.3B, Backbone=LLaMA, Shots=02024.07 | 28.7 | — | |
| Self-DistillTraining Dataset=Alpaca-GPT4, Infer Mode=Self-Distill2024.05 | 28.57 | — | |
| Sheared-Llama#Param=2.7B, Training Tokens=50B, Shots=zero-shot2024.10 | 28.3 | — | |
| Sheared-LLaMA-2.7BNumber of Parameters=2.7B, Backbone=LLaMA, Shots=02024.07 | 28.1 | — | |
| Open-Llama-v2#Param=3.4B, Training Tokens=1T, Shots=zero-shot2024.10 | 28.1 | — | |
| TransActNumber of Parameters=2.6B, Backbone=LLaMA, Shots=02024.07 | 27.9 | — | |
| RPJ-INCITE-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 27.8 | — | |
| LLM-Pruner-2.6BNumber of Parameters=2.6B, Backbone=LLaMA, Shots=02024.07 | 27.7 | — | |
| Pythia#Param=2.8B, Training Tokens=300B, Shots=zero-shot2024.10 | 27.7 | — | |
| Open-Llama-v2#Param=6.9B, Training Tokens=1T, Shots=zero-shot2024.10 | 27.6 | — | |
| Sheared-LLaMA-1.3BNumber of Parameters=1.3B, Backbone=LLaMA, Shots=02024.07 | 27.5 | — | |
| TransActNumber of Parameters=1.3B, Backbone=LLaMA, Shots=02024.07 | 27.5 | — | |
| PythiaNumber of Parameters=6.9B, Evaluation Protocol=Five-shot2023.04 | 27 | — | |
| OPT-1.3BNumber of Parameters=1.3B, Backbone=OPT, Shots=02024.07 | 27 | — | |
| GLM-130BSetting=Chain-of-Thought Prompting2023.06 | 26.6 | — | |
| LLaMA2-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 26.1 | — | |
| Pythia (Deduplicated)Model size=6.9B, Evaluation protocol=Five-shot2023.04 | 25.7 | — | |
| OPT-2.7BNumber of Parameters=2.7B, Backbone=OPT, Shots=02024.07 | 25.7 | — | |
| Pythia#Param=6.9B, Training Tokens=300B, Shots=zero-shot2024.10 | 25.3 | — | |
| Pythia (Deduplicated)Model size=70M, Evaluation protocol=Five-shot2023.04 | 25 | — | |
| Pythia (Deduplicated)Model size=12B, Evaluation protocol=Five-shot2023.04 | 24.4 | — | |
| FairSeqNumber of Parameters=13B, Zero-Shot=true2022.04 | 24 | — | |
| PythiaNumber of Parameters=1.4B, Evaluation Protocol=Five-shot2023.04 | 24 | — | |
| Pythia (Deduplicated)Model size=160M, Evaluation protocol=Five-shot2023.04 | 23.7 | — | |
| Falcon-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 23.7 | — | |
| OLMo-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 23.4 | — | |
| ChatGPTSetting=Chain-of-Thought Prompting2023.06 | 23.3 | — | |
| FairSeqNumber of Parameters=6.7B, Zero-Shot=true2022.04 | 23.2 | — | |
| FairSeqNumber of Parameters=355M, Zero-Shot=true2022.04 | 23 | — | |
| GPT-NeoXModel Size=20B, Zero-shot=true2022.04 | 23 | — | |
| Pythia (Deduplicated)Model size=1.4B, Evaluation protocol=Five-shot2023.04 | 23 | — | |
| MPT-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 22.9 | — | |
| GPT-3Model Variant=DaVinci, Zero-shot=true2022.04 | 22.7 | — | |
| PythiaNumber of Parameters=1B, Evaluation Protocol=Five-shot2023.04 | 22.7 | — | |
| Pythia (Deduplicated)Model size=1B, Evaluation protocol=Five-shot2023.04 | 22.6 | — | |
| FairSeqNumber of Parameters=13B, Evaluation Protocol=5-shot2022.04 | 22.3 | — | |
| EWCTraining Dataset=CoT-Collection, Infer Mode=EWC2024.05 | 22.27 | — | |
| L2Training Dataset=CoT-Collection, Infer Mode=L22024.05 | 22.12 | — | |
| FairSeqNumber of Parameters=125M, Zero-Shot=true2022.04 | 22 | — | |
| PythiaNumber of Parameters=410M, Evaluation Protocol=Five-shot2023.04 | 22 | — | |
| Pythia (Deduplicated)Model size=2.8B, Evaluation protocol=Five-shot2023.04 | 22 | — | |
| VanillaTraining Dataset=CoT-Collection, Infer Mode=Vanilla2024.05 | 21.97 | — | |
| LoRACLTraining Dataset=CoT-Collection, Infer Mode=LORACL2024.05 | 21.97 | — | |
| FairSeqNumber of Parameters=125M, Evaluation Protocol=5-shot2022.04 | 21.8 | — | |
| GPT-3Model Variant=Ada, Zero-shot=true2022.04 | 21.8 | — | |
| PythiaNumber of Parameters=70M, Evaluation Protocol=Five-shot2023.04 | 21.8 | — | |
| PythiaNumber of Parameters=12B, Evaluation Protocol=Five-shot2023.04 | 21.8 | — | |
| GPT-3Model Variant=Curie, Zero-shot=true2022.04 | 21.7 | — | |
| PythiaNumber of Parameters=160M, Evaluation Protocol=Five-shot2023.04 | 21.7 | — | |
| PythiaNumber of Parameters=2.8B, Evaluation Protocol=Five-shot2023.04 | 21.7 | — | |
| Pythia-6.9BEvaluation protocol=Zero-shot, Parameters=6.9B2024.02 | 21.5 | — | |
| FairSeqNumber of Parameters=1.3B, Zero-Shot=true2022.04 | 21.4 | — | |
| FairSeqNumber of Parameters=2.7B, Evaluation Protocol=5-shot2022.04 | 21.4 | — | |
| FairSeqNumber of Parameters=6.7B, Evaluation Protocol=5-shot2022.04 | 21.4 | — | |
| FairSeqNumber of Parameters=2.7B, Zero-Shot=true2022.04 | 21.2 | — | |
| FairSeqNumber of Parameters=1.3B, Evaluation Protocol=5-shot2022.04 | 21 | — | |
| Pythia (Deduplicated)Model size=410M, Evaluation protocol=Five-shot2023.04 | 21 | — | |
| GPT-JModel Size=6B, Zero-shot=true2022.04 | 20.9 | — | |
| FairSeqNumber of Parameters=355M, Evaluation Protocol=5-shot2022.04 | 20.7 | — | |
| GPT-3Model Variant=Babbage, Zero-shot=true2022.04 | 19.8 | — | |
| LLaMA-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 19.5 | — | |
| GPT-3.5Setting=Chain-of-Thought Prompting2023.06 | 13.3 | — | |
| GPT-3.5Setting=Few-Shot Prompting2023.06 | 10 | — |