Rebuttal Generation on Human Evaluation Set (100 comments) 1.0 (test)
9.92Attitude Scorew GPT4.1-reward
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| w GPT4.1-rewardReward Model=GPT-4.12026.01 | 9.92 | 9.62 | 9.28 | 9.54 | 9.59 | |
| RebuttalAgentFull Architecture=Integrated TSR and Self-Reward2026.01 | 9.86 | 9.38 | 9.34 | 9.68 | 9.57 | |
| GPT-4.12026.01 | 9.32 | 8.8 | 8.7 | 9.14 | 8.99 | |
| o32026.01 | 9.3 | 9.28 | 9.04 | 9.42 | 9.26 | |
| DeepSeek-R12026.01 | 9.24 | 9.08 | 8.86 | 9.16 | 9.08 | |
| w RebuttalRM-rewardReward Model=RebuttalRM2026.01 | 9.16 | 8.9 | 8.84 | 9.07 | 8.96 | |
| Qwen3-8B2026.01 | 8.88 | 8.6 | 8.12 | 8.4 | 8.5 | |
| RebuttalFTTraining Mode=Fine-Tuning Only2026.01 | 7.38 | 6.8 | 6.3 | 6.5 | 6.75 |