Video-level watch-time prediction on THU
0.63MAEAddictSim
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AddictSimTraining Strategy=Proposed simulator-based framework2026.01 | 0.63 | 0.79 | |
| PPOTraining Strategy=Actor-critic RL2026.01 | 0.65 | 0.81 | |
| GRPOTraining Strategy=Group-relative Policy Optimization (RL)2026.01 | 0.65 | 0.82 | |
| STFTraining Strategy=Supervised Fine-Tuning (SFT)2026.01 | 0.77 | 1.01 | |
| BaseTraining Strategy=Prompt-based (no training), Base LLM=Qwen32026.01 | 0.96 | 1.18 |