Code Generation on TACO Verified
96.3During-task AccuracyCloud LLM Cluster
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Cloud LLM ClusterModel=Cloud LLM Cluster2026.01 | 96.3 | 96.3 | — | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 84.8 | 72.8 | 14.1 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 82.2 | 66.6 | 19 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 76.9 | 60.2 | 21.7 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 75.8 | 61.2 | 19.2 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 74.4 | 62.3 | 16.2 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 71.6 | 58.9 | 17.7 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 70.4 | 54.2 | 23 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 69.8 | 56.7 | 18.7 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 69.7 | 58.6 | 15.9 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 66.8 | 54.8 | 17.9 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 66.7 | 54.6 | 18.1 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 65.8 | 48.5 | 26.3 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 61.6 | 43.1 | 29.1 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 60.8 | 48.9 | 19.6 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 57.4 | 39.5 | 31.2 | |
| Collaborative Training w/ DA-GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 56.5 | 44 | 22.1 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 55.2 | 38.2 | 30.8 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 51.8 | 34.6 | 33.2 | |
| Collaborative Training w/ GAPGEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 51.1 | 37.8 | 26 | |
| Collaborative Training w/ GVPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 50.9 | 37.6 | 26.1 | |
| Edge Tuning w/ RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 48.1 | 34.2 | 28.9 | |
| Collaborative Training w/ GRPOEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 46.6 | 33.1 | 29 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 44 | 24 | 45.5 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 44 | 24 | 45.5 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Qwen2.5-1.5B-Instruct2026.01 | 43.7 | 24.2 | 44.6 | |
| Edge Tuning OnlyEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 42.7 | 29.9 | 30 | |
| Edge Tuning OnlyEvaluated Responses=Local-Cloud Joint Responses [y_theta_i, y_c_i], Model=Llama-3.2-3B-Instruct2026.01 | 42.7 | 29.9 | 30 | |
| Edge Tuning w/ Naive RouterEvaluated Responses=Local-solved Responses only [y_theta_i], Model=Llama-3.2-3B-Instruct2026.01 | 42.5 | 30.1 | 29.2 |