Verifiable Judging on VerifyBench Hard
88.8AccuracyOpenAI/GPT-o1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OpenAI/GPT-o1Type=LLM-as-a-judge2025.07 | 88.8 | 85.48 | |
| Master-RM-32BType=LLM-as-a-judge, Parameters=32B2025.07 | 86.8 | 81.96 | |
| Anthropic/Claude-4-SonnetType=LLM-as-a-judge2025.07 | 85.3 | 79.71 | |
| Master-RM-7BType=LLM-as-a-judge, Parameters=7B2025.07 | 84.4 | 80.98 | |
| OpenAI/GPT-4oType=LLM-as-a-judge2025.07 | 84.3 | 77.94 | |
| OpenAI/GPT-4o-miniType=LLM-as-a-judge2025.07 | 82.8 | 76.29 | |
| Multi-sub RMType=LLM-as-a-judge2025.07 | 82.5 | 78.42 | |
| Qwen2.5-32B-InstructType=LLM-as-a-judge, Parameters=32B2025.07 | 81.3 | 75.3 | |
| Qwen2.5-7B-InstructType=LLM-as-a-judge, Parameters=7B2025.07 | 80.2 | 74.21 | |
| Qwen2.5-14B-InstructType=LLM-as-a-judge, Parameters=14B2025.07 | 78.4 | 71.79 | |
| Qwen2.5-72B-InstructType=LLM-as-a-judge, Parameters=72B2025.07 | 78.3 | 72.63 | |
| Llama-3-70B-InstructType=LLM-as-a-judge, Parameters=70B2025.07 | 77.1 | 73.35 | |
| math-verifyType=rule-based verifier2025.07 | 76 | 60.21 | |
| Qwen2.5-3B-InstructType=LLM-as-a-judge, Parameters=3B2025.07 | 75.2 | 72.56 | |
| Omni-JudgeType=LLM-as-a-judge2025.07 | 67.7 | 58.98 | |
| Qwen2.5-1.5B-InstructType=LLM-as-a-judge, Parameters=1.5B2025.07 | 67.7 | 66.7 | |
| Llama-3-8B-InstructType=LLM-as-a-judge, Parameters=8B2025.07 | 61.3 | 60.56 | |
| General-VerifierType=LLM-as-a-judge2025.07 | 50.2 | 49.4 | |
| Qwen2.5-0.5B-InstructType=LLM-as-a-judge, Parameters=0.5B2025.07 | 40.7 | 40.55 |