Logical Reasoning on LogiQA (test)
86AccuracyHuman
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Human2022.03 | 86 | — | — | — | — | — | |
| Human Performance2021.05 | 86 | — | — | — | — | — | |
| Human Performance2024.05 | 86 | — | — | — | — | — | |
| ART-meanModel=Qwen2.5-32B-Instruct, Evaluation Protocol=D_test2026.04 | 63.5 | — | — | — | — | — | |
| ART-maxModel=Qwen2.5-32B-Instruct, Evaluation Protocol=D_test2026.04 | 63.2 | — | — | — | — | — | |
| VanillaModel=Qwen2.5-32B-Instruct, Evaluation Protocol=D_test2026.04 | 63 | — | — | — | — | — | |
| ART-maxModel=Qwen2.5-14B-Instruct, Evaluation Protocol=D_test2026.04 | 57.8 | — | — | — | — | — | |
| ART-meanModel=Qwen2.5-14B-Instruct, Evaluation Protocol=D_test2026.04 | 57.6 | — | — | — | — | — | |
| VanillaModel=Qwen2.5-14B-Instruct, Evaluation Protocol=D_test2026.04 | 56.9 | — | — | — | — | — | |
| DetermLRModel=GPT-42023.10 | 54.19 | 11.74 | — | — | — | — | |
| ART-meanModel=Qwen2.5-7B-Instruct, Evaluation Protocol=D_test2026.04 | 52.5 | — | — | — | — | — | |
| ART-maxModel=Qwen2.5-7B-Instruct, Evaluation Protocol=D_test2026.04 | 52.1 | — | — | — | — | — | |
| VanillaModel=Qwen2.5-7B-Instruct, Evaluation Protocol=D_test2026.04 | 50.7 | — | — | — | — | — | |
| FFTModel=TinyLlama [71]2026.06 | 47.54 | — | — | — | — | — | |
| FOCAL REASONERBackbone=DeBERTa, Data Augmentation=false2021.05 | 45.8 | — | — | — | — | — | |
| ART-meanModel=Qwen2-7B-Instruct, Evaluation Protocol=D_test2026.04 | 45.5 | — | — | — | — | — | |
| MERITBackbone=ALBERT2023.06 | 45.3 | — | — | — | — | — | |
| CRModel=GPT-42023.10 | 45.25 | 17 | — | — | — | — | |
| PathReasoner2024.05 | 45.01 | — | — | — | — | — | |
| CCPO (qwen2.5-1.5b-instruct)Assignment Strategy=Counterfactuals2026.03 | 45.01 | — | — | — | — | — | |
| ART-maxModel=Qwen2-7B-Instruct, Evaluation Protocol=D_test2026.04 | 45 | — | — | — | — | — | |
| HGNBackbone=DeBERTa2021.05 | 44.2 | — | — | — | — | — | |
| ART-maxModel=Ministral-8B-Instruct-2410, Evaluation Protocol=D_test2026.04 | 44 | — | — | — | — | — | |
| IDOLBackbone=ALBERT2023.06 | 43.8 | — | — | — | — | — | |
| ART-meanModel=Ministral-8B-Instruct-2410, Evaluation Protocol=D_test2026.04 | 43.5 | — | — | — | — | — | |
| qwen2.5-1.5b-instructAssignment Strategy=Shared2026.03 | 43.47 | — | — | — | — | — | |
| LReasonerBackbone=DeBERTa2021.05 | 43.3 | — | — | — | — | — | |
| SFTModel=Qwen3-8B-Base2025.07 | 43.12 | — | — | — | — | — | |
| SFT+GRPOModel=Qwen3-8B-Base2025.07 | 43.12 | — | — | — | — | — | |
| ToTModel=GPT-42023.10 | 43.02 | 19.87 | — | — | — | — | |
| ART-beamLLM=Llama3-8B-Instruct2026.04 | 42.7 | — | — | — | — | — | |
| SFT+PenaltyModel=Qwen3-8B-Base2025.07 | 42.62 | — | — | — | — | — | |
| SASFTModel=Qwen3-8B-Base2025.07 | 42.62 | — | — | — | — | — | |
| LogiformerBackbone=RoBERTa2023.06 | 42.6 | — | — | — | — | — | |
| LogiformerCategory=Graph2024.05 | 42.55 | — | -2.46 | — | — | — | |
| MERITCategory=Sequence, Extra data=true2024.05 | 42.4 | — | -2.61 | — | — | — | |
| VanillaModel=Qwen2-7B-Instruct, Evaluation Protocol=D_test2026.04 | 42.3 | — | — | — | — | — | |
| APOLLOBase model=RoBERTa-large2022.12 | 42.1 | — | — | — | — | — | |
| IDOLBackbone=RoBERTa2023.06 | 41.8 | — | — | — | — | — | |
| GeniusTraining Corpus=OpenHermes2.5 (32K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 41.63 | — | — | — | — | — | |
| MERITBase model=RoBERTa-large2022.12 | 41.5 | — | — | — | — | — | |
| DeBERTaBackbone=DeBERTa2021.05 | 41.5 | — | — | — | — | — | |
| MERITBackbone=RoBERTa2023.06 | 41.5 | — | — | — | — | — | |
| ART-greedyLLM=Llama3-8B-Instruct2026.04 | 41.4 | — | — | — | — | — | |
| SIModel=GPT-42023.10 | 41.34 | 14.35 | — | — | — | — | |
| ALBERTBackbone=ALBERT2023.06 | 41.3 | — | — | — | — | — | |
| LReasonerBackbone=ALBERT2023.06 | 41.2 | — | — | — | — | — | |
| VanillaModel=Ministral-8B-Instruct-2410, Evaluation Protocol=D_test2026.04 | 41.1 | — | — | — | — | — | |
| SCPOTraining Corpus=OpenHermes2.5 (32K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 41.01 | — | — | — | — | — | |
| Beam SearchLLM=Llama3-8B-Instruct2026.04 | 41 | — | — | — | — | — | |
| GeniusTraining Corpus=Magpie (25K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 40.86 | — | — | — | — | — | |
| AdaLoGN2022.03 | 40.71 | — | — | — | — | — | |
| AdaLoGNCategory=Graph2024.05 | 40.71 | — | -4.3 | — | — | — | |
| AdaLoGNBackbone=RoBERTa2023.06 | 40.7 | — | — | — | — | — | |
| LReasonerdata_augmentation=true2022.03 | 40.6 | — | — | — | — | — | |
| LRReasonerBase model=RoBERTa-large2022.12 | 40.6 | — | — | — | — | — | |
| LReasonerBackbone=RoBERTa, Data Augmentation=true2021.05 | 40.6 | — | — | — | — | — | |
| LReasonerBackbone=RoBERTa2023.06 | 40.6 | — | — | — | — | — | |
| LReasonerCategory=Sequence2024.05 | 40.6 | — | -4.41 | — | — | — | |
| COT-SCModel=GPT-4, n=162023.10 | 40.43 | 16 | — | — | — | — | |
| SCPOTraining Corpus=Magpie (25K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 40.4 | — | — | — | — | — | |
| FOCAL REASONERBase model=RoBERTa-large2022.12 | 40.3 | — | — | — | — | — | |
| MERITBackbone=RoBERTa, Data Augmentation=true2021.05 | 40.3 | — | — | — | — | — | |
| FOCAL REASONERBackbone=RoBERTa, Data Augmentation=false2021.05 | 40.3 | — | — | — | — | — | |
| Focal Reasoner2022.03 | 40.25 | — | — | — | — | — | |
| FocalReasonerCategory=Graph2024.05 | 40.25 | — | -4.76 | — | — | — | |
| ART-meanModel=Llama3.1-8B-Instruct, Evaluation Protocol=D_test2026.04 | 40.2 | — | — | — | — | — | |
| SPINTraining Corpus=Magpie (25K), Supervision Level=With Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 40.09 | — | — | — | — | — | |
| HGNBackbone=RoBERTa2021.05 | 39.9 | — | — | — | — | — | |
| Self-RewardingTraining Corpus=OpenHermes2.5 (32K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 39.78 | — | — | — | — | — | |
| ART-maxModel=Llama3.1-8B-Instruct, Evaluation Protocol=D_test2026.04 | 39.6 | — | — | — | — | — | |
| DAGN (Aug)Backbone=RoBERTa-Large, Graph Feature Augmentation=true2021.03 | 39.32 | — | — | — | — | — | |
| DAGN2022.03 | 39.32 | — | — | — | — | — | |
| DAGNCategory=Graph2024.05 | 39.32 | — | -5.69 | — | — | — | |
| DAGNBackbone=RoBERTa, Data Augmentation=true2021.05 | 39.3 | — | — | — | — | — | |
| LAMBADAModel=GPT-42023.10 | 39.11 | 56.24 | — | — | — | — | |
| cLAModel=TinyLlama [71], Adapter rank=162026.06 | 39.09 | — | — | — | — | — | |
| VanillaModel=Llama3.1-8B-Instruct, Evaluation Protocol=D_test2026.04 | 38.8 | — | — | — | — | — | |
| DoLaLLM=Llama3-8B-Instruct2026.04 | 38.8 | — | — | — | — | — | |
| DAGNBackbone=RoBERTa-Large2021.03 | 38.71 | — | — | — | — | — | |
| DAGNBase model=RoBERTa-large2022.12 | 38.7 | — | — | — | — | — | |
| DAGNBackbone=RoBERTa, Data Augmentation=false2021.05 | 38.7 | — | — | — | — | — | |
| DAGNBackbone=RoBERTa2023.06 | 38.7 | — | — | — | — | — | |
| CoHTraining Corpus=Magpie (25K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 38.56 | — | — | — | — | — | |
| COTModel=GPT-42023.10 | 38.55 | 1 | — | — | — | — | |
| VanillaLLM=Llama3-8B-Instruct2026.04 | 38.5 | — | — | — | — | — | |
| CoHTraining Corpus=OpenHermes2.5 (32K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 38.4 | — | — | — | — | — | |
| ACTLLM=Llama3-8B-Instruct2026.04 | 38 | — | — | — | — | — | |
| DetermLRModel=GPT-3.5-turbo2023.10 | 37.99 | 13.39 | — | — | — | — | |
| Self-RewardingTraining Corpus=Magpie (25K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 37.94 | — | — | — | — | — | |
| SFTTraining Corpus=Magpie (25K), Supervision Level=With Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 37.78 | — | — | — | — | — | |
| ITILLM=Llama3-8B-Instruct2026.04 | 37.3 | — | — | — | — | — | |
| qwen2.5-1.5b-instructAssignment Strategy=Untrained2026.03 | 37.02 | — | — | — | — | — | |
| RoBERTaBackbone=RoBERTa2023.06 | 36.6 | — | — | — | — | — | |
| SFT+PenaltyModel=Qwen3-1.7B-Base2025.07 | 36 | — | — | — | — | — | |
| STaRTraining Corpus=Magpie (25K), Supervision Level=Without Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 35.94 | — | — | — | — | — | |
| RoBERTa-Large2021.03 | 35.33 | — | — | — | — | — | |
| RoBERTamodel_size=Large2022.03 | 35.33 | — | — | — | — | — | |
| ROBERTa-LargeCategory=Sequence2024.05 | 35.33 | — | -9.68 | — | — | — | |
| SPINTraining Corpus=OpenHermes2.5 (32K), Supervision Level=With Supervision, Backbone Model=LLaMA3.1-8B-Instruct2025.04 | 35.33 | — | — | — | — | — |