Judge Agreement Accuracy on Quora 915 queries (test)
73Agreement AccuracyLogicJudge
Evaluation Results
| Method | Links | |
|---|---|---|
| LogicJudgebackbone=Qwen-3-30B-A3B, training=SFT + GRPO2026.01 | 73 | |
| Ensemble Voteevaluation_setting=one-shot2026.01 | 71.31 | |
| gpt-5evaluation_setting=one-shot2026.01 | 70.33 | |
| gemini-2.5-proevaluation_setting=one-shot2026.01 | 70 | |
| gpt-o3evaluation_setting=one-shot2026.01 | 69.67 | |
| qwen3-maxevaluation_setting=one-shot2026.01 | 69.67 | |
| gemini-2.5-proevaluation_setting=one-shot, mode=thinking2026.01 | 68 | |
| claude-4.5-sonnetevaluation_setting=one-shot2026.01 | 68 | |
| gpt-5.1evaluation_setting=one-shot2026.01 | 66 | |
| claude-4.5-sonnetevaluation_setting=one-shot, mode=thinking2026.01 | 66 | |
| gemini-3-proevaluation_setting=one-shot, mode=thinking2026.01 | 65 | |
| gemini-3-proevaluation_setting=one-shot2026.01 | 62.33 | |
| Ensemble Consensusevaluation_setting=one-shot2026.01 | 56.67 | |
| qwen3-235bevaluation_setting=one-shot2026.01 | 53.67 | |
| deepseek-v3evaluation_setting=one-shot2026.01 | 53.33 | |
| claude-4-sonnetevaluation_setting=one-shot2026.01 | 51.67 | |
| claude-4-sonnetevaluation_setting=one-shot, mode=thinking2026.01 | 51.33 | |
| qwen-maxevaluation_setting=one-shot2026.01 | 40.67 | |
| gpt-4.1evaluation_setting=one-shot2026.01 | 34 |