Personalized Artwork Recommendation on user-title tuples 5K Llama 8B evaluated (Held-out)
1.41AccuracySFT with reasoning from Qwen-32B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SFT with reasoning from Qwen-32BModel=Llama 8B, Training=SFT with reasoning distillation, Teacher Model=Qwen-32B, Adaptation=LoRA2026.01 | 1.41 | 5.21 | |
| DPOModel=Llama 8B, Training=Direct Policy Optimization, Adaptation=LoRA2026.01 | 0.91 | 2.82 | |
| SFTModel=Llama 8B, Training=Supervised Fine-Tuning, Adaptation=LoRA2026.01 | -2.55 | 2.45 | |
| Zero-shot predictionModel=Llama 8B, Mode=Zero-shot2026.01 | -4.22 | -0.19 | |
| Random guessDescription=Non-LLM baseline2026.01 | -74.96 | -4.59 |