Judge Agreement Accuracy on Zhihu 847 queries (test)
0.75Agreement AccuracyLogicJudge
Evaluation Results
| Method | Links | |
|---|---|---|
| LogicJudgebackbone=Qwen-3-30B-A3B, training=SFT + GRPO2026.01 | 0.75 | |
| gemini-2.5-proevaluation_setting=one-shot2026.01 | 0.74 | |
| Ensemble Voteevaluation_setting=one-shot2026.01 | 0.7357 | |
| gpt-o3evaluation_setting=one-shot2026.01 | 0.717 | |
| qwen3-maxevaluation_setting=one-shot2026.01 | 0.717 | |
| gpt-5evaluation_setting=one-shot2026.01 | 0.716 | |
| gpt-4.1evaluation_setting=one-shot2026.01 | 0.7077 | |
| gemini-3-proevaluation_setting=one-shot, mode=thinking2026.01 | 0.696 | |
| gemini-3-proevaluation_setting=one-shot2026.01 | 0.682 | |
| deepseek-v3evaluation_setting=one-shot2026.01 | 0.6641 | |
| gpt-5.1evaluation_setting=one-shot2026.01 | 0.65 | |
| claude-4.5-sonnetevaluation_setting=one-shot, mode=thinking2026.01 | 0.645 | |
| claude-4.5-sonnetevaluation_setting=one-shot2026.01 | 0.64 | |
| qwen-maxevaluation_setting=one-shot2026.01 | 0.6162 | |
| Ensemble Consensusevaluation_setting=one-shot2026.01 | 0.6111 | |
| claude-4-sonnetevaluation_setting=one-shot2026.01 | 0.567 | |
| claude-4-sonnetevaluation_setting=one-shot, mode=thinking2026.01 | 0.562 | |
| qwen3-235bevaluation_setting=one-shot2026.01 | 0.492 | |
| gemini-2.5-proevaluation_setting=one-shot, mode=thinking2026.01 | 0.419 |