Scientific Review Feedback Generation on ICLR Human Evaluation 2025 (test)
3.46ActionabilityRBTACT
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| RBTACTSystem Category=Ours2026.03 | 3.46 | 4.08 | 4.3 | 4.76 | 4.26 | |
| GPT-5-chatSystem Category=LLMs, Mode=zero-shot2026.03 | 3.38 | 4.04 | 4.35 | 4.98 | 4.47 | |
| RBTACT-SFTSystem Category=Ours, Training=Fine-tuned (SFT-only) on ReviewSeg-SFT-13k2026.03 | 3.28 | 4.01 | 4.16 | 4.7 | 4.24 | |
| DeepReviewer-14BSystem Category=Other Methods2026.03 | 3.27 | 3.96 | 4.28 | 4.75 | 4.21 | |
| Llama-3.1-70BSystem Category=LLMs, Mode=zero-shot2026.03 | 3.22 | 3.95 | 4.18 | 4.65 | 4.15 | |
| MARGSystem Category=Other Methods2026.03 | 3.2 | 3.87 | 4.15 | 4.72 | 4.18 | |
| DeepSeek-V3.2System Category=LLMs, Mode=zero-shot2026.03 | 3.15 | 3.98 | 4.22 | 4.88 | 4.28 | |
| LimGenSystem Category=Other Methods2026.03 | 3.14 | 3.92 | 4.08 | 4.64 | 4.05 | |
| Qwen-3-32BSystem Category=LLMs, Mode=zero-shot2026.03 | 3.06 | 3.78 | 4.12 | 4.58 | 4.12 |