Language Understanding on MMLU (During-task and Post-Switch Accuracy)
90.8During-task AccuracyCloud LLM Cluster
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Cloud LLM Cluster2026.01 | 90.8 | 90.8 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 69.5 | 63 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 67.7 | 62.4 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 66.7 | 59.5 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 65.4 | 61.8 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 64.8 | 60.7 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 64.8 | 59.1 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 64.2 | 63.2 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 63.1 | 60.6 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 63.1 | 58.5 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 63 | 61.3 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 61.5 | 60.1 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 61.5 | 56.3 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 61.2 | 60.3 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 61.1 | 59.6 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 59.8 | 55.2 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 59.5 | 56.2 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 59.2 | 57.3 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 58.9 | 55.6 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 58.9 | 55.3 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 58.2 | 55.2 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 57.9 | 56.9 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Llama-3.2-3B-Instruct2026.01 | 57.9 | 56.9 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Llama-3.2-3B-Instruct2026.01 | 57.5 | 56.4 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 57.3 | 54 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 56.3 | 51.2 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.4 | 51.8 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta^i, y_c^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.4 | 51.8 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta^i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.2 | 52 |