Scientific Rebuttal Generation on Scientific Rebuttal Evaluation dataset (test)
14.93BLEU@4RBTACT-SFT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RBTACT-SFTtraining=Supervised Fine-Tuning2026.03 | 14.93 | 12.33 | 11.53 | 18.51 | |
| RBTACT2026.03 | 14.62 | 12.64 | 11.65 | 18.57 | |
| DeepReview2026.03 | 12.4 | 11.3 | 10.2 | 19.6 | |
| GPT-5-chat2026.03 | 11.17 | 10.11 | 9.96 | 24.9 | |
| MARG2026.03 | 10.95 | 9.1 | 8.6 | 18.1 | |
| LimGen2026.03 | 10.9 | 8.2 | 7.9 | 17.4 | |
| Llama-3.1-70Bparameters=70B2026.03 | 10.48 | 8.76 | 8.27 | 16.42 | |
| DeepSeek-V3.22026.03 | 10.19 | 9.76 | 8.49 | 17.78 | |
| Qwen-3-32Bparameters=32B2026.03 | 9.72 | 8.58 | 8.14 | 17.18 |