Judge Agreement Accuracy on DeepResearch 1319 queries (test)
74.5Agreement AccuracyLogicJudge
Evaluation Results
| Method | Links | |
|---|---|---|
| LogicJudgebackbone=Qwen-3-30B-A3B, training=SFT + GRPO2026.01 | 74.5 | |
| qwen3-maxevaluation_setting=one-shot2026.01 | 73.53 | |
| gpt-4.1evaluation_setting=one-shot2026.01 | 68.63 | |
| gemini-2.5-proevaluation_setting=one-shot2026.01 | 68.63 | |
| Ensemble Voteevaluation_setting=one-shot2026.01 | 66.67 | |
| gemini-2.5-proevaluation_setting=one-shot, mode=thinking2026.01 | 65.69 | |
| qwen-maxevaluation_setting=one-shot2026.01 | 65.69 | |
| gpt-o3evaluation_setting=one-shot2026.01 | 63.73 | |
| claude-4.5-sonnetevaluation_setting=one-shot, mode=thinking2026.01 | 63.73 | |
| gpt-5evaluation_setting=one-shot2026.01 | 62.75 | |
| deepseek-v3evaluation_setting=one-shot2026.01 | 62.75 | |
| gpt-5.1evaluation_setting=one-shot2026.01 | 61.76 | |
| claude-4.5-sonnetevaluation_setting=one-shot2026.01 | 61.76 | |
| gemini-3-proevaluation_setting=one-shot, mode=thinking2026.01 | 59.8 | |
| gemini-3-proevaluation_setting=one-shot2026.01 | 54.9 | |
| Ensemble Consensusevaluation_setting=one-shot2026.01 | 49.02 | |
| claude-4-sonnetevaluation_setting=one-shot2026.01 | 46.08 | |
| claude-4-sonnetevaluation_setting=one-shot, mode=thinking2026.01 | 45.1 | |
| qwen3-235bevaluation_setting=one-shot2026.01 | 43.14 |