Sentiment review generation on Sentiment review generation 100 samples (test)
100Win Rate vs π_PPOπ_PPO + PFM × 5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| π_PPO + PFM × 5Base Policy=π_PPO, PFM Iterations=52024.05 | 100 | — | |
| π_PPO + PFMBase Policy=π_PPO, PFM Iterations=12024.05 | 99 | — | |
| π_ref + PFM × 5Base Policy=π_ref, PFM Iterations=52024.05 | 85 | 100 | |
| π_ref + PFMBase Policy=π_ref, PFM Iterations=12024.05 | 2 | 100 |