Technical Writing
0.491DiscoverBase
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BaseModel Backbone=Llama-3.1-8B-Instruct, Strategy=Zero-shot Base2026.02 | 0.491 | 0.36 | 0.212 | 3.32 | |
| SFT+DPOModel Backbone=Llama-3.1-8B-Instruct, Recipe=SFT followed by DPO2026.02 | 0.49 | 0.359 | 0.313 | 2.94 | |
| SFT+DPO+GRPOModel Backbone=Qwen3-8B, Recipe=SFT + DPO + Group Relative Policy Optimization2026.02 | 0.482 | 0.355 | 0.55 | 2.63 | |
| SFT+DPOModel Backbone=Qwen3-8B, Recipe=SFT followed by DPO2026.02 | 0.475 | 0.364 | 0.691 | 2.78 | |
| DPOModel Backbone=Llama-3.1-8B-Instruct, Recipe=Direct Preference Optimization2026.02 | 0.472 | 0.342 | 0.273 | 3.11 | |
| SFTModel Backbone=Llama-3.1-8B-Instruct, Recipe=Supervised Fine-Tuning2026.02 | 0.471 | 0.352 | 0.816 | 2.09 | |
| COLLABLLMModel Backbone=Llama-3.1-8B-Instruct, Strategy=Collaborative2026.02 | 0.458 | 0.337 | 0.249 | 3.13 | |
| Prompted BaseModel Backbone=Llama-3.1-8B-Instruct, Strategy=Prompted2026.02 | 0.436 | 0.335 | 0.242 | 3.05 | |
| DPOModel Backbone=Qwen3-8B, Recipe=Direct Preference Optimization2026.02 | 0.422 | 0.333 | 0.672 | 2.76 | |
| SFTModel Backbone=Qwen3-8B, Recipe=Supervised Fine-Tuning2026.02 | 0.416 | 0.337 | 0.81 | 1.9 | |
| Prompted BaseModel Backbone=Qwen3-8B, Strategy=Prompted2026.02 | 0.413 | 0.338 | 0.64 | 2.79 | |
| BaseModel Backbone=Qwen3-8B, Strategy=Zero-shot Base2026.02 | 0.407 | 0.337 | 0.353 | 3.39 |