Commonsense Reasoning on StrategyQA (test)
83.49AccuracySGE
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SGEModel=GPT-4, Tool=Code Interpreter2024.05 | 83.49 | — | — | — | — | |
| Decomp PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 82.08 | — | — | — | — | |
| Refine PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 81.26 | — | — | — | — | |
| CoT PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 81.16 | — | — | — | — | |
| IO PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 78.4 | — | — | — | — | |
| Rethinking with retrievalModel=GPT-3 (text-davinci-002), Temperature=0.7, Num samples=92022.12 | 77.73 | — | — | — | — | |
| CNTP + SC (Ours + SC)Model=Llama-3.1-8B-Instruct, Reasoning Chain=Multiple Reasoning Chains, Decoding Strategy=Stochastic, paths=402025.07 | 76.3 | — | 13,900 | — | — | |
| SCModel=Llama-3.1-8B-Instruct, Reasoning Chain=Multiple Reasoning Chains, Decoding Strategy=Stochastic, paths=402025.07 | 76.2 | — | 2,450 | — | — | |
| RIOToptimization=automatic2025.06 | 74.6 | — | — | — | — | |
| DSPyoptimization=automatic2025.06 | 73.4 | — | — | — | — | |
| Self-consistencyModel=GPT-3 (text-davinci-002), Temperature=0.7, Num samples=92022.12 | 73.36 | — | — | — | — | |
| CNTP (Ours*)Model=Llama-3.1-8B-Instruct, Reasoning Chain=Single Reasoning Chain, Decoding Strategy=Stochastic2025.07 | 73.2 | — | 350 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 73.07 | — | — | — | — | |
| Greedy DecodingModel=Llama-3.1-8B-Instruct, Reasoning Chain=Single Reasoning Chain, Decoding Strategy=Deterministic2025.07 | 72.9 | — | 60 | — | — | |
| Beam SearchModel=Llama-3.1-8B-Instruct, Reasoning Chain=Multiple Reasoning Chains, Decoding Strategy=Deterministic, beam=52025.07 | 72.9 | — | 310 | — | — | |
| Mixtral-8x7B-Instruct-v0.1Zero-shot evaluation=true2024.10 | 72.83 | — | — | — | — | |
| Stochastic Decoding*Model=Llama-3.1-8B-Instruct, Reasoning Chain=Single Reasoning Chain, Decoding Strategy=Stochastic2025.07 | 72 | — | 60 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 71.62 | — | — | — | — | |
| Twenty-Shot CoToptimization=manual, shots=202025.06 | 71.6 | — | — | — | — | |
| Four-Shot CoToptimization=manual, shots=42025.06 | 71.2 | — | — | — | — | |
| GPT-3.5-TurboZero-shot evaluation=true2024.10 | 70.92 | — | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 70.89 | — | — | — | — | |
| APEoptimization=automatic2025.06 | 70.4 | — | — | — | — | |
| Entro-duction2025.03 | 70.3 | 9.6 | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 70.16 | — | — | — | — | |
| large teacherModel=Qwen2.5-7B-Instruct2026.05 | 69.6 | — | — | — | — | |
| TextGradoptimization=automatic2025.06 | 68.8 | — | — | — | — | |
| CoT-SC@maj8majority_vote=82025.03 | 68.3 | 15 | — | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 68.12 | — | — | — | — | |
| DRR2025.03 | 67.7 | — | — | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 67.69 | — | — | — | — | |
| CoT-SC@maj64majority_vote=642025.03 | 67.1 | 15 | — | — | — | |
| Gemini-1.0-ProZero-shot evaluation=true2024.10 | 67.03 | — | — | — | — | |
| OPROoptimization=automatic2025.06 | 67 | — | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 66.96 | — | — | — | — | |
| Chain-of-thought promptingModel=GPT-3 (text-davinci-002), Temperature=02022.12 | 65.94 | — | — | — | — | |
| GPT-J+Self-ReflectionZero-shot evaluation=true2024.10 | 65.9 | — | — | — | — | |
| small teacherModel=Qwen2.5-3B-Instruct2026.05 | 65.9 | — | — | — | — | |
| ToT2025.03 | 65.8 | 121 | — | — | — | |
| Zero-Shot CoToptimization=manual, shots=02025.06 | 65.8 | — | — | — | — | |
| Complex CoT2025.03 | 65.7 | 5 | — | — | — | |
| oursStudent Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 65.3 | — | — | — | — | |
| T5-XL (teacher)Model Architecture=T5-XL2024.06 | 64.5 | — | — | — | — | |
| T5-XXL+COTZero-shot evaluation=true2024.10 | 63.77 | — | — | — | — | |
| GKDStudent Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 63.6 | — | — | — | — | |
| Few-shot promptingModel=GPT-3 (text-davinci-002), Temperature=02022.12 | 63.32 | — | — | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 63.32 | — | — | — | — | |
| SpiralThinkerBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 63.32 | — | — | — | — | |
| Llama3.1-8B-InstructZero-shot evaluation=true2024.10 | 63.03 | — | — | — | — | |
| oursStudent Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 62.9 | — | — | — | — | |
| iCoT-KDBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 62.88 | — | — | — | — | |
| DistiLLMStudent Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 62.8 | — | — | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 62.7 | — | — | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 62.45 | — | — | — | — | |
| ImitKDStudent Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 61.7 | — | — | — | — | |
| Self-talk2025.03 | 61.5 | — | — | — | — | |
| SeqKDStudent Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 61.5 | — | — | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 61.43 | — | — | — | — | |
| DistiLLMStudent Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 61.2 | — | — | — | — | |
| SFT (student)Student Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 60.7 | — | — | — | — | |
| CODIBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 60.7 | — | — | — | — | |
| iCoT-SIBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 60.69 | — | — | — | — | |
| SoftCoTBackbone=Qwen2.5-7B-Instruct, Training Objective=Language Modeling Objective, Evaluation Protocol=Projection Module Trained2025.02 | 60.61 | — | — | 75.06 | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 60.41 | — | — | — | — | |
| GKDStudent Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 60.3 | — | — | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 60.12 | — | — | — | — | |
| CoconutBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 60 | — | — | — | — | |
| ImitKDStudent Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 59.7 | — | — | — | — | |
| KDStudent Model Architecture=T5-Large, Teacher Model=T5-XL2024.06 | 59.2 | — | — | — | — | |
| oursStudent Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 58.2 | — | — | — | — | |
| Zero-shot promptingModel=GPT-3 (text-davinci-002), Temperature=02022.12 | 58.08 | — | — | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 57.93 | — | — | — | — | |
| CoT2025.03 | 57.7 | 5 | — | — | — | |
| Pause TokenBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 57.64 | — | — | — | — | |
| SFT (student)Student Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 57.5 | — | — | — | — | |
| SeqKDStudent Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 57.5 | — | — | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 56.62 | — | — | — | — | |
| DistiLLMStudent Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 56.3 | — | — | — | — | |
| CLPDOrder=student-loss, Student Model=Llama-1B2026.05 | 56.3 | — | — | — | — | |
| GKDStudent Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 55.6 | — | — | — | — | |
| CLPDOrder=student-loss, Student Model=Qwen-0.5B2026.05 | 55.5 | — | — | — | — | |
| KDStudent Model Architecture=T5-Base, Teacher Model=T5-XL2024.06 | 55.3 | — | — | — | — | |
| Curriculum learning onlyTeacher=small teacher, Order=student-loss, Student Model=Llama-1B2026.05 | 55.1 | — | — | — | — | |
| Curriculum learning onlyTeacher=large teacher, Order=student-loss, Student Model=Llama-1B2026.05 | 54.4 | — | — | — | — | |
| Progressive distillation onlyTeacher=small & large teacher, Student Model=Llama-1B2026.05 | 54.3 | — | — | — | — | |
| Standard distillationTeacher=large teacher, Student Model=Llama-1B2026.05 | 54.2 | — | — | — | — | |
| Standard distillationTeacher=small teacher, Student Model=Llama-1B2026.05 | 54.1 | — | — | — | — | |
| Qwen2.5-1.5B-InstructZero-shot evaluation=true2024.10 | 53.86 | — | — | — | — | |
| ImitKDStudent Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 53.8 | — | — | — | — | |
| Zero-Shot Assist-CoTBackbone=Qwen2.5-7B-Instruct, Training Objective=None, Evaluation Protocol=Zero-Shot2025.02 | 52.71 | — | — | 71.64 | — | |
| Progressive distillation onlyTeacher=small & large teacher, Student Model=Qwen-0.5B2026.05 | 52.6 | — | — | — | — | |
| SFT (student)Student Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 52.4 | — | — | — | — | |
| Curriculum learning onlyTeacher=large teacher, Order=student-loss, Student Model=Qwen-0.5B2026.05 | 52.4 | — | — | — | — | |
| Curriculum learning onlyTeacher=small teacher, Order=student-loss, Student Model=Qwen-0.5B2026.05 | 52.1 | — | — | — | — | |
| Standard distillationTeacher=small teacher, Student Model=Qwen-0.5B2026.05 | 51.8 | — | — | — | — | |
| Standard distillationTeacher=large teacher, Student Model=Qwen-0.5B2026.05 | 51.4 | — | — | — | — | |
| Zero-Shot CoT-UnkBackbone=Qwen2.5-7B-Instruct, Training Objective=None, Evaluation Protocol=Zero-Shot2025.02 | 50.74 | — | — | 70.6 | — | |
| SeqKDStudent Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 50.6 | — | — | — | — | |
| KDStudent Model Architecture=T5-Small, Teacher Model=T5-XL2024.06 | 49.7 | — | — | — | — | |
| Zero-Shot CoTBackbone=Qwen2.5-7B-Instruct, Training Objective=None, Evaluation Protocol=Zero-Shot2025.02 | 49.65 | — | — | 70.29 | — |