Session-level watch-time prediction on KuaiRec
1.72MAEAddictSim
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AddictSimTraining Strategy=Proposed simulator-based framework2026.01 | 1.72 | 2.35 | |
| PPOTraining Strategy=Actor-critic RL2026.01 | 1.96 | 2.4 | |
| GRPOTraining Strategy=Group-relative Policy Optimization (RL)2026.01 | 2.02 | 2.46 | |
| STFTraining Strategy=Supervised Fine-Tuning (SFT)2026.01 | 2.25 | 2.6 | |
| BaseTraining Strategy=Prompt-based (no training), Base LLM=Qwen32026.01 | 2.79 | 3.19 |