Multi-task Learning on Mixture of 6 task-level SFT tasks (test)
0.508Test LossTANDEM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TANDEMModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 0.508 | 75.03 | |
| Skill-ItModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 0.539 | 74.6 | |
| AioliModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 0.542 | 74.63 | |
| DoGEModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 0.563 | 74.35 | |
| UniformModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 0.591 | 74.72 | |
| DoReMiModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 0.686 | 73.11 |