Scientific Review Feedback Generation on ICLR LLM-as-a-Judge 2025 (test)
3.38Actionability ScoreRBTACT
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| RBTACTSystem Category=Ours2026.03 | 3.38 | 3.7 | 4.05 | 4.82 | 3.74 | |
| GPT-5-chatSystem Category=LLMs, Mode=zero-shot2026.03 | 3.28 | 3.66 | 4.12 | 4.95 | 3.78 | |
| DeepReviewer-14BSystem Category=Other Methods2026.03 | 3.23 | 3.48 | 4.03 | 4.74 | 3.58 | |
| MARGSystem Category=Other Methods2026.03 | 3.19 | 3.47 | 3.91 | 4.69 | 3.51 | |
| RBTACT-SFTSystem Category=Ours, Training=Fine-tuned (SFT-only) on ReviewSeg-SFT-13k2026.03 | 3.18 | 3.59 | 3.94 | 4.72 | 3.66 | |
| DeepSeek-V3.2System Category=LLMs, Mode=zero-shot2026.03 | 3.13 | 3.56 | 4 | 4.86 | 3.7 | |
| Llama-3.1-70BSystem Category=LLMs, Mode=zero-shot2026.03 | 3.11 | 3.54 | 3.96 | 4.74 | 3.53 | |
| LimGenSystem Category=Other Methods2026.03 | 3.08 | 3.38 | 3.88 | 4.54 | 3.39 | |
| Qwen-3-32BSystem Category=LLMs, Mode=zero-shot2026.03 | 3.03 | 3.36 | 3.9 | 4.68 | 3.32 |