Math Reasoning on Gaokao En 2023
79AccuracyLegislator-Executor (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| Legislator-Executor (Ours)Model=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 79 | |
| LIMOModel=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 76.9 | |
| S1KModel=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 76.1 | |
| Legislator-Executor (Ours)Model=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 75.8 | |
| NPG-Muse-8BPass@k protocol=avg@82025.08 | 74.7 | |
| Full AttentionBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=False2026.02 | 74.2 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=False2026.02 | 74.2 | |
| Legislator-Executor (Ours)Model=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 73.5 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=True2026.02 | 72.7 | |
| S1KModel=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 71.7 | |
| DAPO + HistaModel Size=14B, RL Algorithm=DAPO2026.05 | 71.1 | |
| Legislator-Executor (Ours)Model=Qwen2.5-Math-7B, Decoding Strategy=Zero-shot greedy2026.04 | 70.4 | |
| LIMOModel=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 70.4 | |
| Legislator-Executor (Ours)Model=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 69.6 | |
| Full AttentionBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=False2026.02 | 68.8 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=False2026.02 | 68.8 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=True2026.02 | 68.8 | |
| LIMOModel=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 68.6 | |
| S1KModel=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 68.6 | |
| DAPOModel Size=14B, RL Algorithm=DAPO2026.05 | 68.5 | |
| Legislator-Executor (Ours)Model=Qwen2.5-7B, Decoding Strategy=Zero-shot greedy2026.04 | 68.3 | |
| LIMOModel=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 68.1 | |
| BaseModel=Qwen3-14B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 67.5 | |
| BaseModel=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 66.8 | |
| Qwen2.5-14B-Instruct-1MPass@k protocol=avg@82025.08 | 66 | |
| DAPO + HistaModel Size=7B, RL Algorithm=DAPO2026.05 | 65.4 | |
| BaseModel=Qwen3-8B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 64.9 | |
| S1KModel=Qwen2.5-7B-QwQ, Decoding Strategy=Zero-shot greedy2026.04 | 64.9 | |
| Qwen3-14B-BasePass@k protocol=avg@82025.08 | 64.8 | |
| CSIPOModel Size=7B, RL Algorithm=CSIPO2026.05 | 64.4 | |
| S1KModel=Qwen2.5-Math-7B, Decoding Strategy=Zero-shot greedy2026.04 | 64.2 | |
| CSIPO + HistaModel Size=7B, RL Algorithm=CSIPO2026.05 | 63.9 | |
| DAPOModel Size=7B, RL Algorithm=DAPO2026.05 | 63.8 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=True2026.02 | 63.3 | |
| LIMOModel=Qwen2.5-Math-7B, Decoding Strategy=Zero-shot greedy2026.04 | 63.1 | |
| S1KModel=Qwen2.5-7B, Decoding Strategy=Zero-shot greedy2026.04 | 62.9 | |
| baseline-GroupAModel=Qwen2.5 (7B-Instruct)2026.03 | 62.6 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=False2026.02 | 62.5 | |
| ARE-GroupAModel=Qwen2.5 (7B-Instruct)2026.03 | 61.8 | |
| LIMOModel=Qwen2.5-7B, Decoding Strategy=Zero-shot greedy2026.04 | 61.8 | |
| ARE-GroupBModel=Qwen2.5 (7B-Instruct)2026.03 | 61.6 | |
| baseline-GroupBModel=Qwen2.5 (7B-Instruct)2026.03 | 61.3 | |
| NPG-Muse-7BPass@k protocol=avg@82025.08 | 61.1 | |
| BaseModel=Qwen3-4B-Base, Decoding Strategy=Zero-shot greedy2026.04 | 60.8 | |
| GRPOModel=qwen2.5-7b-instruct, Agent Configuration=Single-agent2026.03 | 60.31 | |
| CCPOModel=qwen2.5-7b-instruct, Agent Configuration=Dual-Agent2026.03 | 59.74 | |
| ReMAModel=qwen3-4b-base, Agent Configuration=Dual-Agent2026.03 | 58.7 | |
| Qwen2.5-7B-Ins-1MPass@k protocol=avg@82025.08 | 58.3 | |
| DAPOModel Size=3B, RL Algorithm=DAPO2026.05 | 58.2 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=False2026.02 | 57.8 | |
| ReMAModel=qwen2.5-7b-instruct, Agent Configuration=Dual-Agent2026.03 | 57.7 | |
| UntrainedModel=qwen2.5-7b-instruct, Agent Configuration=Dual-Agent2026.03 | 57.4 | |
| CCPOModel=qwen3-4b-base, Agent Configuration=Dual-Agent2026.03 | 57.4 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=True2026.02 | 57 | |
| Qwen3-8B-BasePass@k protocol=avg@82025.08 | 56.9 | |
| BaseModel=Qwen2.5-7B, Decoding Strategy=Zero-shot greedy2026.04 | 56.1 | |
| BaseModel=Qwen2.5-Math-7B, Decoding Strategy=Zero-shot greedy2026.04 | 56.1 | |
| CSIPO + HistaModel Size=3B, RL Algorithm=CSIPO2026.05 | 55.8 | |
| DAPO + HistaModel Size=3B, RL Algorithm=DAPO2026.05 | 55.6 | |
| CSIPOModel Size=3B, RL Algorithm=CSIPO2026.05 | 52.7 | |
| DAPO + HistaModel Size=1.5B, RL Algorithm=DAPO2026.05 | 51.2 | |
| GRPOModel=qwen3-4b-base, Agent Configuration=Single-agent2026.03 | 51.06 | |
| GSPO + HistaModel Size=1.5B, RL Algorithm=GSPO2026.05 | 49.8 | |
| DAPOModel Size=1.5B, RL Algorithm=DAPO2026.05 | 49.4 | |
| GSPOModel Size=1.5B, RL Algorithm=GSPO2026.05 | 49.3 | |
| CSIPOModel Size=1.5B, RL Algorithm=CSIPO2026.05 | 48.3 | |
| CSIPO + HistaModel Size=1.5B, RL Algorithm=CSIPO2026.05 | 48.1 | |
| ARE-GroupAModel=Qwen2.5 (1.5B-Instruct)2026.03 | 46 | |
| baseline-GroupAModel=Qwen2.5 (1.5B-Instruct)2026.03 | 45.5 | |
| baseline-GroupCModel=Qwen2.5 (1.5B-Instruct)2026.03 | 45.5 | |
| ARE-GroupCModel=Qwen2.5 (1.5B-Instruct)2026.03 | 45.5 | |
| ReMAModel=qwen2.5-1.5b-instruct, Agent Configuration=Dual-Agent2026.03 | 43.9 | |
| Legislator-Executor (Ours)Model=Gemma-2-9B, Decoding Strategy=Zero-shot greedy2026.04 | 43.4 | |
| S1KModel=Gemma-2-9B, Decoding Strategy=Zero-shot greedy2026.04 | 43.1 | |
| GRPOModel=qwen2.5-1.5b-instruct, Agent Configuration=Single-agent2026.03 | 42.82 | |
| ARE-GroupBModel=Llama3.1 (8B-Instruct)2026.03 | 42.6 | |
| UntrainedModel=qwen2.5-1.5b-instruct, Agent Configuration=Dual-Agent2026.03 | 42.08 | |
| UntrainedModel=qwen2.5-7b-instruct, Agent Configuration=Single-agent2026.03 | 42.04 | |
| Base ModelModel=Qwen2.5 (7B-Instruct)2026.03 | 42.04 | |
| CCPOModel=qwen2.5-1.5b-instruct, Agent Configuration=Dual-Agent2026.03 | 41.6 | |
| ARE-GroupCModel=Llama3.1 (8B-Instruct)2026.03 | 40.3 | |
| baseline-GroupBModel=Llama3.1 (8B-Instruct)2026.03 | 39.7 | |
| CCPOModel=llama3.1-8b-instruct, Agent Configuration=Dual-Agent2026.03 | 39.48 | |
| baseline-GroupCModel=Llama3.1 (8B-Instruct)2026.03 | 39.2 | |
| ReMAModel=llama3.1-8b-instruct, Agent Configuration=Dual-Agent2026.03 | 38.7 | |
| LIMOModel=Gemma-2-9B, Decoding Strategy=Zero-shot greedy2026.04 | 38.4 | |
| UntrainedModel=llama3.1-8b-instruct, Agent Configuration=Dual-Agent2026.03 | 36.1 | |
| Legislator-Executor (Ours)Model=Llama-3.1-8B, Decoding Strategy=Zero-shot greedy2026.04 | 30.9 | |
| DAPO + HistaModel Size=0.5B, RL Algorithm=DAPO2026.05 | 30.6 | |
| S1KModel=Llama-3.1-8B, Decoding Strategy=Zero-shot greedy2026.04 | 28.8 | |
| DAPOModel Size=0.5B, RL Algorithm=DAPO2026.05 | 28 | |
| LIMOModel=Llama-3.1-8B, Decoding Strategy=Zero-shot greedy2026.04 | 24.9 | |
| BaseModel=Gemma-2-9B, Decoding Strategy=Zero-shot greedy2026.04 | 24.7 | |
| UntrainedModel=qwen2.5-1.5b-instruct, Agent Configuration=Single-agent2026.03 | 23.5 | |
| Base ModelModel=Qwen2.5 (1.5B-Instruct)2026.03 | 23.5 | |
| UntrainedModel=qwen3-4b-base, Agent Configuration=Dual-Agent2026.03 | 22.08 | |
| Legislator-Executor (Ours)Model=Llama-3.2-3B, Decoding Strategy=Zero-shot greedy2026.04 | 17.1 | |
| Legislator-Executor (Ours)Model=Mistral-7B-v0.3, Decoding Strategy=Zero-shot greedy2026.04 | 15.6 | |
| BaseModel=Llama-3.1-8B, Decoding Strategy=Zero-shot greedy2026.04 | 15.1 | |
| LIMOModel=Mistral-7B-v0.3, Decoding Strategy=Zero-shot greedy2026.04 | 14.5 |