Creative Writing on Creative Writing
45.2Discovery ScoreSFT+DPO+GRPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SFT+DPO+GRPOModel Backbone=Qwen3-8B, Recipe=SFT + DPO + Group Relative Policy Optimization2026.02 | 45.2 | 33.7 | 83.1 | 2,050 | |
| SFT+DPOModel Backbone=Qwen3-8B, Recipe=SFT followed by DPO2026.02 | 44 | 33.4 | 72.1 | 2,870 | |
| DPOModel Backbone=Qwen3-8B, Recipe=Direct Preference Optimization2026.02 | 42.6 | 33.5 | 70.5 | 3,100 | |
| SFT+DPOModel Backbone=Llama-3.1-8B-Instruct, Recipe=SFT followed by DPO2026.02 | 42.4 | 28.4 | 32.9 | 2,770 | |
| SFTModel Backbone=Llama-3.1-8B-Instruct, Recipe=Supervised Fine-Tuning2026.02 | 40.7 | 33.4 | 92.3 | 1,710 | |
| DPOModel Backbone=Llama-3.1-8B-Instruct, Recipe=Direct Preference Optimization2026.02 | 40.5 | 29.2 | 33.1 | 2,910 | |
| Prompted BaseModel Backbone=Qwen3-8B, Strategy=Prompted2026.02 | 39 | 30.8 | 62.9 | 3,010 | |
| BaseModel Backbone=Llama-3.1-8B-Instruct, Strategy=Zero-shot Base2026.02 | 38.2 | 30 | 20.1 | 3,090 | |
| Prompted BaseModel Backbone=Llama-3.1-8B-Instruct, Strategy=Prompted2026.02 | 37.7 | 26.4 | 26 | 2,970 | |
| COLLABLLMModel Backbone=Llama-3.1-8B-Instruct, Strategy=Collaborative2026.02 | 37.3 | 28 | 32.6 | 2,930 | |
| BaseModel Backbone=Qwen3-8B, Strategy=Zero-shot Base2026.02 | 35.2 | 30.4 | 36.2 | 3,410 | |
| SFTModel Backbone=Qwen3-8B, Recipe=Supervised Fine-Tuning2026.02 | 34.9 | 31 | 90.4 | 1,590 |