Watch duration prediction on Persona B
0.86SMAPEPersonaAct (SFT+GRPO)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PersonaAct (SFT+GRPO)training=SFT followed by GRPO2026.01 | 0.86 | 19.08 | |
| SFTtraining=Supervised Fine-Tuning2026.01 | 0.969 | 23.81 | |
| GRPOtraining=Group Relative Policy Optimization2026.01 | 1.218 | 20.4 | |
| LLM Sim2026.01 | 1.357 | 19.09 |