Peer Review Quality Assessment on PeeriScope OpenReview, F1000 Research, Semantic Web Journal 1.0
0.359Kendall's Tau (Overall Quality)GPT-4o
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4oEvaluation Mode=Zero-shot, Scoring Scale=five-point ordinal scale, Input=review text, paper title, abstract2026.04 | 0.359 | 0.476 | 0.411 | 0.407 | 0.343 | 0.327 | 0.298 | 0.295 | 0.189 | 0.163 | 0.128 | 0.124 | 0.115 | |
| Qwen-3Evaluation Mode=Zero-shot, Scoring Scale=five-point ordinal scale, Input=review text, paper title, abstract, Model Version=8B parameter, System Role=Primary PeeriScope evaluator2026.04 | 0.252 | 0.338 | 0.314 | 0.428 | 0.211 | 0.176 | 0.186 | 0.105 | 0.078 | 0.139 | 0.106 | 0.117 | 0.089 | |
| Phi-4Evaluation Mode=Zero-shot, Scoring Scale=five-point ordinal scale, Input=review text, paper title, abstract2026.04 | 0.241 | 0.374 | 0.279 | 0.397 | 0.259 | 0.254 | 0.215 | 0.204 | 0.175 | 0.186 | 0.053 | 0.038 | 0.006 |