Mathematical Reasoning on MATH 500 (During-task Acc., Post-Switch Acc.)
97.3During-task Accuracy (MATH 500)Cloud LLM Cluster
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Cloud LLM Cluster2026.01 | 97.3 | 97.3 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 87.2 | 77.4 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 86.6 | 76 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 81 | 72.4 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 79.6 | 70.2 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 77.7 | 68.1 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 75.2 | 63.3 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 75 | 66.2 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 74.8 | 64 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 74 | 66.2 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 73.3 | 62.1 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 71.1 | 58.3 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 70.2 | 59.6 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 68.2 | 58.2 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 66.4 | 58 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 65.6 | 49.8 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 65.6 | 57.8 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 65.4 | 52.9 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 62 | 55.8 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 60.1 | 54.9 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 58.5 | 43.4 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 55.7 | 41.4 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 54.6 | 40 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 54.2 | 40.8 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 54.2 | 40.8 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 52 | 40.9 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 46.6 | 38.8 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 46.6 | 38.8 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 45.7 | 37.1 |