Argument Rewriting on Manual Evaluation Dataset Argument Rewriting 1.0
44.9Rank 1LLaMA + PPOapp
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| LLaMA + PPOappOptimization=PPO focusing on appropriateness2024.06 | 44.9 | 32.4 | 13.3 | 7.1 | 2.2 | 0 | 1.89 | 83.3 | |
| LLaMA + PPOapp>simOptimization=PPO (Appropriateness > Similarity)2024.06 | 29.3 | 29.8 | 18.7 | 14.7 | 5.8 | 1.8 | 2.43 | 72.9 | |
| Human BaselineType=Human rewriting2024.06 | 19.6 | 16.9 | 23.1 | 17.8 | 11.6 | 11.1 | 3.18 | 56.6 | |
| LLaMA + Instruct.Model Type=Instruction-finetuned2024.06 | 3.1 | 5.3 | 18.2 | 21.3 | 32.9 | 19.1 | 4.32 | 35.1 | |
| LLaMA + PPOapp=simOptimization=PPO (Appropriateness = Similarity)2024.06 | 2.7 | 11.1 | 22.2 | 26.7 | 20.9 | 16.4 | 4.01 | 41.2 | |
| LLaMA + PPOapp<simOptimization=PPO (Appropriateness < Similarity)2024.06 | 0.4 | 4.4 | 4.4 | 12.4 | 26.7 | 51.6 | 5.15 | 16 |