Video-level watch-time prediction on KuaiRec
0.29MAEAddictSim
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AddictSimTraining Strategy=Proposed simulator-based framework2026.01 | 0.29 | 0.42 | |
| PPOTraining Strategy=Actor-critic RL2026.01 | 0.3 | 0.42 | |
| STFTraining Strategy=Supervised Fine-Tuning (SFT)2026.01 | 0.31 | 0.42 | |
| GRPOTraining Strategy=Group-relative Policy Optimization (RL)2026.01 | 0.32 | 0.44 | |
| BaseTraining Strategy=Prompt-based (no training), Base LLM=Qwen32026.01 | 0.35 | 0.47 |