Math Word Problem Solving on SOMADHAN (test)
0.88AccuracyGPT-OSS-20B
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=High2025.12 | 0.88 | |
| GPT-OSS-120BPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=Med.2025.12 | 0.88 | |
| LLaMA-3.3-70B-versatilePrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=Med.2025.12 | 0.88 | |
| LLaMA-4-maverick-17B-128EPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=Med.2025.12 | 0.88 | |
| GPT-OSS-20BPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=Med.2025.12 | 0.87 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=High2025.12 | 0.87 | |
| LLaMA-3.3-70B-versatilePrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=Med.2025.12 | 0.87 | |
| LLaMA-4-scout-17B-16EPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=Med.2025.12 | 0.87 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=Med.2025.12 | 0.86 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=High2025.12 | 0.86 | |
| GPT-OSS-120BPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=High2025.12 | 0.86 | |
| LLaMA-3.3-70B-versatilePrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=Med.2025.12 | 0.86 | |
| LLaMA-3.3-70B-versatilePrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=Med.2025.12 | 0.86 | |
| LLaMA-4-scout-17B-16EPrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=Med.2025.12 | 0.86 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=High2025.12 | 0.85 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=Med.2025.12 | 0.85 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=Med.2025.12 | 0.85 | |
| LLaMA-3.3-70B-versatilePrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=Med.2025.12 | 0.85 | |
| LLaMA-4-maverick-17B-128EPrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=Med.2025.12 | 0.85 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=Med.2025.12 | 0.84 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=2, Reasoning Configuration=High2025.12 | 0.84 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=High2025.12 | 0.84 | |
| GPT-OSS-20BPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=High2025.12 | 0.84 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=Med.2025.12 | 0.84 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=High2025.12 | 0.84 | |
| LLaMA-4-maverick-17B-128EPrompting Strategy=Standard, Number of Shots=02025.12 | 0.84 | |
| LLaMA-4-maverick-17B-128EPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=Med.2025.12 | 0.84 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=High2025.12 | 0.83 | |
| LLaMA-4-maverick-17B-128EPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=Med.2025.12 | 0.83 | |
| LLaMA-4-maverick-17B-128EPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=Med.2025.12 | 0.83 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=Med.2025.12 | 0.82 | |
| GPT-OSS-20BPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=Med.2025.12 | 0.82 | |
| GPT-OSS-120BPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=Med.2025.12 | 0.82 | |
| LLaMA-4-scout-17B-16EPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=Med.2025.12 | 0.82 | |
| LLaMA-4-scout-17B-16EPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=Med.2025.12 | 0.82 | |
| GPT-OSS-120BPrompting Strategy=Standard, Number of Shots=0, Reasoning Configuration=N/A2025.12 | 0.8 | |
| LLaMA-3.3-70B-versatilePrompting Strategy=Standard, Number of Shots=02025.12 | 0.79 | |
| LLaMA-4-scout-17B-16EPrompting Strategy=Standard, Number of Shots=02025.12 | 0.79 | |
| GPT-OSS-20BPrompting Strategy=Standard, Number of Shots=0, Reasoning Configuration=N/A2025.12 | 0.78 | |
| LLaMA-4-scout-17B-16EPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=Med.2025.12 | 0.76 | |
| LLaMA-3.1-8B-instantPrompting Strategy=Chain of Thought (CoT), Number of Shots=7, Reasoning Configuration=Med.2025.12 | 0.73 | |
| LLaMA-3.1-8B-instantPrompting Strategy=Chain of Thought (CoT), Number of Shots=1, Reasoning Configuration=Med.2025.12 | 0.51 | |
| LLaMA-3.1-8B-instantPrompting Strategy=Standard, Number of Shots=02025.12 | 0.48 | |
| LLaMA-3.1-8B-instantPrompting Strategy=Chain of Thought (CoT), Number of Shots=5, Reasoning Configuration=Med.2025.12 | 0.41 | |
| LLaMA-3.1-8B-instantPrompting Strategy=Tree of Thought (ToT), Number of Shots=0, Reasoning Configuration=Med.2025.12 | 0.31 |