Math Reasoning on MATH lighteval
98.4During-task AccuracyCloud LLM Cluster
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Cloud LLM ClusterModel=Cloud LLM Cluster2026.01 | 98.4 | 98.4 | — | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 84.5 | 75.5 | 10.7 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 80.7 | 74.8 | 7.3 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 78.1 | 68.5 | 12.3 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 77.4 | 69.9 | 9.7 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 77.2 | 67.7 | 12.3 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 73.5 | 59.8 | 18.6 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 73.1 | 61.3 | 16.1 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 72.7 | 63.1 | 13.2 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 72.2 | 66.5 | 7.9 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 70.3 | 63.2 | 10.1 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 70.3 | 63.1 | 10.2 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 68.3 | 60.6 | 11.3 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 67.8 | 59.2 | 12.7 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 64.7 | 52.5 | 18.9 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 63.8 | 58.5 | 8.3 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 62.2 | 55.4 | 10.9 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 62 | 54.2 | 12.6 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 61.9 | 53.4 | 13.7 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 59 | 53.7 | 9 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 58.7 | 49.6 | 15.5 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 56 | 48.1 | 14.1 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.3 | 44.4 | 19.7 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.3 | 44.4 | 19.7 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.2 | 44.5 | 19.4 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 53.1 | 43.7 | 17.7 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 51.6 | 33.2 | 35.7 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 51.6 | 33.2 | 35.7 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 50.3 | 32.8 | 34.8 |