Session-level watch-time prediction on THU
2.03MAEAddictSim
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AddictSimTraining Strategy=Proposed simulator-based framework2026.01 | 2.03 | 3.31 | |
| PPOTraining Strategy=Actor-critic RL2026.01 | 2.28 | 3.7 | |
| GRPOTraining Strategy=Group-relative Policy Optimization (RL)2026.01 | 2.33 | 3.87 | |
| STFTraining Strategy=Supervised Fine-Tuning (SFT)2026.01 | 3.12 | 4.64 | |
| BaseTraining Strategy=Prompt-based (no training), Base LLM=Qwen32026.01 | 6.84 | 9.57 |