Automated Peer Review on DeepReview-13K (test)
100Technical Accuracy Win (%)ScholarPeer
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ScholarPeerComparison Opponent=CycleReviewer-8B, Opponent Category=Fine-tuned, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=15-202026.01 | 100 | 0 | 100 | 0 | 100 | 0 | 100 | 0 | 100 | 0 | |
| ScholarPeerComparison Opponent=DeepReviewer-14B, Opponent Category=Fine-tuned, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=15-202026.01 | 98.7 | 1 | 99.4 | 0.6 | 99.3 | 0.6 | 98.5 | 0.5 | 99.3 | 0.7 | |
| ScholarPeerComparison Opponent=Claude 4.5 Sonnet, Opponent Category=Single Agent, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=15-202026.01 | 77.7 | 17.8 | 90.2 | 6.8 | 81.3 | 17.8 | 89.8 | 5.9 | 88.3 | 11.2 | |
| ScholarPeerComparison Opponent=Agent Review (Claude 4.5 Sonnet), Opponent Category=Multi Agent, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=15-202026.01 | 72.5 | 20.5 | 89.1 | 8.4 | 74.6 | 23.3 | 88.1 | 6.7 | 86.5 | 13.5 | |
| ScholarPeerComparison Opponent=Gemini 3 Pro, Opponent Category=Single Agent, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=15-202026.01 | 71.8 | 22.2 | 86.2 | 10.6 | 80.6 | 18.1 | 84.2 | 8.9 | 86.8 | 12.9 | |
| ScholarPeerComparison Opponent=AI Scientist v2 (Gemini 3 Pro), Opponent Category=Multi Agent, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=15-202026.01 | 68.7 | 24.8 | 86.1 | 12.2 | 82.9 | 15.9 | 89.1 | 4.4 | 83.8 | 15.9 | |
| ScholarPeerComparison Opponent=Stanford Agent Reviewer, Opponent Category=Multi Agent, LLM Judge=Gemini 3 Pro, Sample Size=50 papers, Approx. LLM Calls=15-202026.01 | 56 | 40 | 64 | 36 | 42 | 56 | 66 | 20 | 64 | 36 | |
| Stanford Agent ReviewerComparison Target=ScholarPeer, Category=Multi Agent, LLM Judge=Gemini 3 Pro, Sample Size=50 papers2026.01 | 40 | 56 | 36 | 64 | 56 | 42 | 20 | 66 | 36 | 64 | |
| DeepReviewer-14BComparison Target=ScholarPeer, Category=Fine-tuned, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=12026.01 | 1 | 98.7 | 0.6 | 99.4 | 0.6 | 99.3 | 0.5 | 98.5 | 0.7 | 99.3 | |
| CycleReviewer-8BComparison Target=ScholarPeer, Category=Fine-tuned, LLM Judge=Gemini 3 Pro, Approx. LLM Calls=12026.01 | 0 | 100 | 0 | 100 | 0 | 100 | 0 | 100 | 0 | 100 |