Mathematical Reasoning on GSM8K (original test)
95.2AccuracyGPT-4-O
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4-OPrompting=CoT2024.06 | 95.2 | — | — | |
| CLAUDE-3-OPUSPrompting=CoT2024.06 | 95 | — | — | |
| GPT-4-TURBOPrompting=CoT2024.06 | 93.8 | — | — | |
| LLAMA-3-70BPrompting=CoT2024.06 | 92.2 | — | — | |
| GEMINI-1.5-PROPrompting=CoT2024.06 | 92 | — | — | |
| WIZARDLM-2 8x22BPrompting=CoT2024.06 | 90.6 | — | — | |
| GEMINI-1.5-FLASHPrompting=CoT2024.06 | 89.8 | — | — | |
| MIXTRAL 8x22BPrompting=CoT2024.06 | 88 | — | — | |
| DEEPSEEKMATH-7BPrompting=CoT2024.06 | 85.4 | — | — | |
| PHI-3-MINI-3.8BPrompting=CoT2024.06 | 83.8 | — | — | |
| COMMAND R+104BPrompting=CoT2024.06 | 79.8 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Multiple# Params=8B, Distillation Teachers=Multiple, Peer-Review=true, Zero-shot evaluation=true2024.10 | 79.3 | — | — | |
| LLAMA-3-8BPrompting=CoT2024.06 | 78.8 | — | — | |
| GPT-3.5-TURBOPrompting=CoT2024.06 | 78.8 | — | — | |
| GPT-3.5-Turbo# Params=175B, Role=Teacher, Zero-shot evaluation=true2024.10 | 78.01 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GPT# Params=8B, Distillation Teachers=GPT-3.5-Turbo, Zero-shot evaluation=true2024.10 | 77.94 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)# Params=8B, Distillation Teachers=Multiple, Peer-Review=false, Zero-shot evaluation=true2024.10 | 76.57 | — | — | |
| Gemini-1.0-ProRole=Teacher, Zero-shot evaluation=true2024.10 | 76.42 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Gemini# Params=8B, Distillation Teachers=Gemini-1.0-Pro, Zero-shot evaluation=true2024.10 | 76.42 | — | — | |
| Llama3.1-8B+ReDistill# Params=8B, Distillation Teachers=DeepSeek-R1, Zero-shot evaluation=true2024.10 | 75.66 | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Mixtral# Params=8B, Distillation Teachers=Mixtral-8x7B-Instruct, Zero-shot evaluation=true2024.10 | 74.83 | — | — | |
| Mixtral-8x7B-Instruct-v0.1# Params=46.7B, Role=Teacher, Zero-shot evaluation=true2024.10 | 74.4 | — | — | |
| Llama3.1-8B-Instruct# Params=8B, Zero-shot evaluation=true2024.10 | 74 | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-Multiple# Params=1.5B, Distillation Teachers=Multiple, Peer-Review=true, Zero-shot evaluation=true2024.10 | 72.48 | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-GPT# Params=1.5B, Distillation Teachers=GPT-3.5-Turbo, Zero-shot evaluation=true2024.10 | 68.01 | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)# Params=1.5B, Distillation Teachers=Multiple, Peer-Review=false, Zero-shot evaluation=true2024.10 | 67.48 | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-Gemini# Params=1.5B, Distillation Teachers=Gemini-1.0-Pro, Zero-shot evaluation=true2024.10 | 66.26 | — | — | |
| Qwen2.5-1.5B-Instruct (FAIR) + Teacher-Mixtral# Params=1.5B, Distillation Teachers=Mixtral-8x7B-Instruct, Zero-shot evaluation=true2024.10 | 65.81 | — | — | |
| Qwen2-1.5B+SIKED# Params=1.5B, Distillation Teachers=Llama3-70B, Zero-shot evaluation=true2024.10 | 64.97 | — | — | |
| Qwen2.5-1.5B-Instruct# Params=1.5B, Zero-shot evaluation=true2024.10 | 64.44 | — | — | |
| MIXTRAL 8x7BPrompting=CoT2024.06 | 62.2 | — | — | |
| Llama2-7B+ReversalMath# Params=7B, Distillation Teachers=GPT-4, Zero-shot evaluation=true2024.10 | 52.1 | — | — | |
| MISTRAL-7BPrompting=CoT2024.06 | 49.4 | — | — | |
| ORCA2-7B# Params=7B, Distillation Teachers=ChatGPT, GPT-4, Zero-shot evaluation=true2024.10 | 47.23 | — | — | |
| CodeT5-Large+PaD# Params=770M, Distillation Teachers=GPT-3.5-Turbo, Zero-shot evaluation=true2024.10 | 44.9 | — | — | |
| Llama-7B+NCE# Params=7B, Distillation Teachers=GPT-3.5-Turbo, GPT-4, Zero-shot evaluation=true2024.10 | 41.93 | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-Multiple# Params=7B, Distillation Teachers=Multiple, Peer-Review=true, Zero-shot evaluation=true2024.10 | 36.24 | — | — | |
| GPT-J+Self-Reflection# Params=6B, Distillation Teachers=ChatGPT, Zero-shot evaluation=true2024.10 | 33.1 | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-GPT# Params=7B, Distillation Teachers=GPT-3.5-Turbo, Zero-shot evaluation=true2024.10 | 30.71 | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-Multiple (w/o Peer-Review)# Params=7B, Distillation Teachers=Multiple, Peer-Review=false, Zero-shot evaluation=true2024.10 | 29.65 | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-Gemini# Params=7B, Distillation Teachers=Gemini-1.0-Pro, Zero-shot evaluation=true2024.10 | 26.84 | — | — | |
| Llama2-7B-chat (FAIR) + Teacher-Mixtral# Params=7B, Distillation Teachers=Mixtral-8x7B-Instruct, Zero-shot evaluation=true2024.10 | 22.67 | — | — | |
| T5-XXL+COT# Params=11B, Distillation Teachers=PaLM, GPT-3, Zero-shot evaluation=true2024.10 | 21.99 | — | — | |
| Llama2-7B-chat# Params=7B, Zero-shot evaluation=true2024.10 | 15.62 | — | — | |
| Codex Cushmanfew-shot learning=true2022.05 | — | 0.05 | 0.58 | |
| Codex Davincifew_shot learning=true2022.05 | — | 0.17 | 0.71 | |
| Gemini 2.5 Pro2026.01 | — | 97.12 | — | |
| LaMDA 137BModel size=137B, pretrained on code=false, few-shot learning=true, different setting from ours=true2022.05 | — | 0.076 | — | |
| OpenAI 6BModel size=6B, pretrained on code=false, different setting from ours=true2022.05 | — | 0.218 | 0.709 | |
| OpenAI GPT-52026.01 | — | 96.66 | — | |
| PaLM-Coder 540BModel size=540B, few-shot learning=true, different setting from ours=true2022.05 | — | 0.509 | — | |
| TraceCodegenBackbone=GPT-Neo 2.7B, Self-sampling strategy=FCS + PCS2022.05 | — | 0.195 | 0.414 |