Mathematical Reasoning on MATH500 (Accuracy)
84.6AccuracyINSTRUCT (BASE)
Evaluation Results
| Method | Links | |
|---|---|---|
| INSTRUCT (BASE)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 84.6 | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=8,403, Processing Time=2.82026.02 | 84.4 | |
| DPO-R1 (HIGH)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=8,964, Processing Time=5.22026.02 | 84.4 | |
| PACEBackbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=9,028, Processing Time=1.72026.02 | 84.4 | |
| DPO-R1 (LOW)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=3,479, Processing Time=0.92026.02 | 84.2 | |
| DPO-R1 (MIDDLE)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=5,127, Processing Time=1.22026.02 | 84.2 | |
| DPO-R1 (LOW)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=3,403, Processing Time=1.42026.02 | 83.2 | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=10,211, Processing Time=4.82026.02 | 83.2 | |
| DPO-R1 (HIGH)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=10,543, Processing Time=7.22026.02 | 83.2 | |
| DPO-R1 (MIDDLE)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=6,073, Processing Time=2.22026.02 | 82.4 | |
| PACEBackbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=10,717, Processing Time=1.02026.02 | 82.2 | |
| INSTRUCT (BASE)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 81.2 | |
| DPO-R1 (LOW)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=5,246, Processing Time=0.82026.02 | 37.8 | |
| PACEBackbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=6,197, Processing Time=0.82026.02 | 37 | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=13,224, Processing Time=2.42026.02 | 35.8 | |
| DPO-R1 (HIGH)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=16,241, Processing Time=4.02026.02 | 35.6 | |
| DPO-R1 (MIDDLE)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=9,692, Processing Time=1.62026.02 | 34.8 | |
| INSTRUCT (BASE)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 29 |