Math on GSM8K
0.988AccuracyBERT-Judge
Evaluation Results
| Method | Links | |
|---|---|---|
| BERT-Judgeclean_name=BERT-as-a-Judge2026.04 | 0.988 | |
| POESOptimizer=OPRO, Model=Qwen-2.5-7B-Instruct2026.04 | 0.972 | |
| GPT-4oEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=4.52†, Throughput (Tok/s)=55.22026.03 | 0.971 | |
| POESOptimizer=EvoPrompt-DE, Model=Qwen-2.5-7B-Instruct2026.04 | 0.965 | |
| Claude 3.5 SonnetEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=3.85†, Throughput (Tok/s)=62.12026.03 | 0.964 | |
| Gemini 2.5 ProEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=3.90†, Throughput (Tok/s)=58.42026.03 | 0.958 | |
| RandomOptimizer=EvoPrompt-GA, Model=Qwen-2.5-7B-Instruct2026.04 | 0.958 | |
| POESOptimizer=EvoPrompt-DE, Model=Llama-3.1-8B-Instruct2026.04 | 0.953 | |
| IPOMPOptimizer=OPRO, Model=Qwen-2.5-7B-Instruct2026.04 | 0.952 | |
| SESSOptimizer=EvoPrompt-DE, Model=Qwen-2.5-7B-Instruct2026.04 | 0.952 | |
| AnchorOptimizer=EvoPrompt-DE, Model=Qwen-2.5-7B-Instruct2026.04 | 0.952 | |
| Qwen2.5-14B-Instruct-1MSize=14B, Type=Instruct, Context Length=1M2025.08 | 0.95 | |
| POESOptimizer=EvoPrompt-GA, Model=Llama-3.1-8B-Instruct2026.04 | 0.95 | |
| AnchorOptimizer=EvoPrompt-GA, Model=Qwen-2.5-7B-Instruct2026.04 | 0.95 | |
| IPOMPOptimizer=EvoPrompt-DE, Model=Qwen-2.5-7B-Instruct2026.04 | 0.95 | |
| POESOptimizer=EvoPrompt-GA, Model=Qwen-2.5-7B-Instruct2026.04 | 0.948 | |
| AnchorOptimizer=EvoPrompt-DE, Model=Llama-3.1-8B-Instruct2026.04 | 0.948 | |
| EcoThinkEvaluation Protocol=Adaptive Inference, Emission (gCO2/q)=1.32, Throughput (Tok/s)=148.62026.03 | 0.945 | |
| RandomOptimizer=EvoPrompt-DE, Model=Llama-3.1-8B-Instruct2026.04 | 0.945 | |
| Regex2026.04 | 0.944 | |
| PredictionOptimizer=EvoPrompt-DE, Model=Llama-3.1-8B-Instruct2026.04 | 0.943 | |
| IPOMPOptimizer=EvoPrompt-GA, Model=Qwen-2.5-7B-Instruct2026.04 | 0.942 | |
| PredictionOptimizer=OPRO, Model=Qwen-2.5-7B-Instruct2026.04 | 0.94 | |
| SESSOptimizer=EvoPrompt-GA, Model=Qwen-2.5-7B-Instruct2026.04 | 0.94 | |
| RandomOptimizer=EvoPrompt-DE, Model=Qwen-2.5-7B-Instruct2026.04 | 0.94 | |
| MANModel=Qwen3, Pruning Ratio=25%, Criterion=MAN, (b, α, β)=(1, 0, 1), Calibration=C42026.06 | 0.937 | |
| SESSOptimizer=OPRO, Model=Qwen-2.5-7B-Instruct2026.04 | 0.935 | |
| PredictionOptimizer=EvoPrompt-DE, Model=Qwen-2.5-7B-Instruct2026.04 | 0.935 | |
| MSANModel=Qwen3, Pruning Ratio=25%, Criterion=MSAN, (b, α, β)=(1, 0, 2), Calibration=C42026.06 | 0.935 | |
| NPG-Muse-8BBackbone=Qwen3-8B-Base2025.08 | 0.933 | |
| AnchorOptimizer=OPRO, Model=Qwen-2.5-7B-Instruct2026.04 | 0.932 | |
| RandomOptimizer=EvoPrompt-GA, Model=Llama-3.1-8B-Instruct2026.04 | 0.93 | |
| PredictionOptimizer=EvoPrompt-GA, Model=Qwen-2.5-7B-Instruct2026.04 | 0.93 | |
| SESSOptimizer=EvoPrompt-GA, Model=Llama-3.1-8B-Instruct2026.04 | 0.927 | |
| SESSOptimizer=EvoPrompt-DE, Model=Llama-3.1-8B-Instruct2026.04 | 0.925 | |
| SDAR-30B-A3Bgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 0.9249 | |
| FullModel=Qwen3, Pruning Ratio=0%, Criterion=Full, Calibration=C42026.06 | 0.923 | |
| IPOMPOptimizer=EvoPrompt-DE, Model=Llama-3.1-8B-Instruct2026.04 | 0.922 | |
| Qwen3-14B + NGMModel Scale=14B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 0.9174 | |
| Qwen3-14BModel Scale=14B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 0.9166 | |
| AutoMRBackbone=Qwen, Skeleton Structure=DAG2025.10 | 0.915 | |
| RandomOptimizer=OPRO, Model=Qwen-2.5-7B-Instruct2026.04 | 0.915 | |
| SDAR-8B-Chatgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 0.9136 | |
| SEERModel=Qwen3, Pruning Ratio=25%, Criterion=SEER, (b, α, β)=(0, 1, 0), Calibration=C42026.06 | 0.913 | |
| MoNEModel=Qwen3, Pruning Ratio=25%, Criterion=MoNE, Calibration=C42026.06 | 0.911 | |
| IPOMPOptimizer=EvoPrompt-GA, Model=Llama-3.1-8B-Instruct2026.04 | 0.91 | |
| AnchorOptimizer=EvoPrompt-GA, Model=Llama-3.1-8B-Instruct2026.04 | 0.91 | |
| EANModel=Qwen3, Pruning Ratio=25%, Criterion=EAN, (b, α, β)=(0, 0, 1), Calibration=C42026.06 | 0.91 | |
| IPOMPOptimizer=OPRO, Model=Llama-3.1-8B-Instruct2026.04 | 0.907 | |
| PredictionOptimizer=OPRO, Model=Llama-3.1-8B-Instruct2026.04 | 0.907 | |
| Qwen3.5-35B-A3B-Baseshots=4-shot, reasoning=CoT, Decoding=Greedy2026.04 | 0.905 | |
| FrequencyModel=Qwen3, Pruning Ratio=25%, Criterion=Frequency, (b, α, β)=(0, 0, 0), Calibration=C42026.06 | 0.904 | |
| Qwen3-30B-A3B-Baseshots=4-shot, reasoning=CoT, Decoding=Greedy2026.04 | 0.903 | |
| OLMo3-7B-InstructParameters=7B, Type=Instruct2026.01 | 0.901 | |
| Qwen3-14B-BaseSize=14B, Type=Base2025.08 | 0.9 | |
| HILAType=MA, Backbone=LLaMA3-8B2026.03 | 0.8986 | |
| Qwen3-8BParameters=8B2026.01 | 0.898 | |
| Molmo2-8BParameters=8B2026.01 | 0.897 | |
| PredictionOptimizer=EvoPrompt-GA, Model=Llama-3.1-8B-Instruct2026.04 | 0.897 | |
| DeepSeek-V3-Base#Shots=8-shot, Architecture=MoE, # activated params=37B, # total params=671B2026.01 | 0.893 | |
| AnchorOptimizer=OPRO, Model=Llama-3.1-8B-Instruct2026.04 | 0.892 | |
| AdaRASCategory=Steering2026.01 | 0.8908 | |
| Qwen3-8B + NGMModel Scale=8B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 0.8908 | |
| Molmo2-O-7BParameters=7B2026.01 | 0.89 | |
| POESOptimizer=OPRO, Model=Llama-3.1-8B-Instruct2026.04 | 0.89 | |
| NPG-Muse-7BBackbone=Qwen2.5-7B-Instruct-1M2025.08 | 0.889 | |
| Qwen2.5-7B-Instruct-1MSize=7B, Type=Instruct, Context Length=1M2025.08 | 0.888 | |
| rStarBackbone=Qwen, Skeleton Structure=Tree2025.10 | 0.887 | |
| JoyAI-LLM Flash-Baseshots=4-shot, reasoning=CoT, Decoding=Greedy2026.04 | 0.887 | |
| LLaDA2.0-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 0.8848 | |
| Qwen3-8BModel Scale=8B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 0.884 | |
| CoTPrompting=Vanilla CoT, Base Model=Qwen3-1.7B2026.01 | 0.8832 | |
| ProbingCategory=Steering2026.01 | 0.8832 | |
| MRPBackbone=Qwen, Skeleton Structure=-2025.10 | 0.882 | |
| Full-FTBackbone=Qwen3-14B-Base, Decode Execution=Independent2026.03 | 0.881 | |
| REAPModel=Qwen3, Pruning Ratio=25%, Criterion=REAP, (b, α, β)=(1, 1, 1), Calibration=C42026.06 | 0.879 | |
| Qwen3-4BParameters=4B2026.01 | 0.878 | |
| OpenReasoning-Nemotron-1.5BCategory=Post-training2026.01 | 0.8757 | |
| Qwen3-4B + NGMModel Scale=4B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 0.8734 | |
| Kimi-Linear-48B-A3B-InstructCompression=Baseline2025.10 | 0.873 | |
| ICaRusModel=Qwen3-8B-Base, KV Sharing=O2026.02 | 0.873 | |
| Qwen3-4BModel Scale=4B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 0.8726 | |
| Meta-ReasonerBackbone=Qwen, Skeleton Structure=Sequential2025.10 | 0.87 | |
| Molmo2-4BParameters=4B2026.01 | 0.866 | |
| FrugalGPT (Cascade)Evaluation Protocol=Standard CoT, Emission (gCO2/q)=1.95, Throughput (Tok/s)=88.52026.03 | 0.865 | |
| MaASBackbone=Qwen, Skeleton Structure=Sequential2025.10 | 0.864 | |
| LLaDA2.1-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 0.8613 | |
| Yuan3.0-1T Base#Shots=8-shot, Architecture=MoE, # activated params=68.5B, # total params=1010B2026.01 | 0.861 | |
| OLMo 3Parameters=7B, Shots=102026.01 | 0.86 | |
| Full-FTBackbone=Qwen3-8B-Base, Decode Execution=Independent2026.03 | 0.858 | |
| Full-FTModel=Qwen3-8B-Base, # Bits=16/16, Decode Execution=Independent2026.03 | 0.858 | |
| Kimi-Linear-48B-A3B-InstructCompression=30%, Compression Method=REAP2025.10 | 0.858 | |
| SUNBackbone=Qwen3-8B-Base, Decode Execution=Shared2026.03 | 0.855 | |
| SUNBackbone=Qwen3-14B-Base, Decode Execution=Shared2026.03 | 0.854 | |
| Multi ModelModel=Qwen3-8B-Base, KV Sharing=X2026.02 | 0.854 | |
| CoTBackbone=Qwen, Skeleton Structure=Sequential2025.10 | 0.853 | |
| RandomOptimizer=OPRO, Model=Llama-3.1-8B-Instruct2026.04 | 0.853 | |
| Qwen3-8B-BaseSize=8B, Type=Base2025.08 | 0.852 | |
| QSUNModel=Qwen3-8B-Base, # Bits=16/4, Decode Execution=Shared2026.03 | 0.851 | |
| REFUSIONEvaluation Protocol=Zero-shot, TPS=81.242025.12 | 0.8491 |