Peer Review Feedback Generation on ICLR papers
45.8Combined Success RateGPT-5.2
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-5.2Bootstrap iterations (B)=1000, Feedback per paper=52026.04 | 45.8 | 1 | 46.3 | 1 | 45.8 | 1 | |
| Gemini-3-flashBootstrap iterations (B)=1000, Feedback per paper=52026.04 | 37.9 | 0.9 | 39.4 | 0.9 | 37.9 | 0.9 | |
| GOODPOINT-DPOTraining Protocol=DPO, Bootstrap iterations (B)=1000, Feedback per paper=52026.04 | 14.7 | 0.5 | 14.9 | 0.5 | 14.7 | 0.5 | |
| GOODPOINT-SFTTraining Protocol=SFT, Bootstrap iterations (B)=1000, Feedback per paper=52026.04 | 9.2 | 0.5 | 9.7 | 0.5 | 9.2 | 0.5 | |
| Qwen3-8b (Base)Model Scale=8b, Training Protocol=Base, Bootstrap iterations (B)=1000, Feedback per paper=52026.04 | 8 | 0.6 | 8.1 | 0.6 | 8 | 0.6 | |
| Llama3.1-8b-InstructModel Scale=8b, Training Protocol=Instruct, Bootstrap iterations (B)=1000, Feedback per paper=52026.04 | 1.8 | 0.3 | 1.8 | 0.3 | 1.8 | 0.3 |