Verifiable Judging on VerifyBench (full)
95.7AccuracyOpenAI/GPT-o1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OpenAI/GPT-o1Type=LLM-as-a-judge2025.07 | 95.7 | 95.7 | |
| Master-RM-32BType=LLM-as-a-judge, Parameters=32B2025.07 | 95.15 | 95.14 | |
| Anthropic/Claude-4-SonnetType=LLM-as-a-judge2025.07 | 95 | 95 | |
| Multi-sub RMType=LLM-as-a-judge2025.07 | 95 | 95 | |
| Master-RM-7BType=LLM-as-a-judge, Parameters=7B2025.07 | 94.45 | 94.45 | |
| Qwen2.5-72B-InstructType=LLM-as-a-judge, Parameters=72B2025.07 | 94.3 | 94.3 | |
| OpenAI/GPT-4oType=LLM-as-a-judge2025.07 | 94.15 | 94.15 | |
| Qwen2.5-32B-InstructType=LLM-as-a-judge, Parameters=32B2025.07 | 93.25 | 93.25 | |
| Llama-3-70B-InstructType=LLM-as-a-judge, Parameters=70B2025.07 | 92.5 | 92.49 | |
| Qwen2.5-14B-InstructType=LLM-as-a-judge, Parameters=14B2025.07 | 92.4 | 92.4 | |
| OpenAI/GPT-4o-miniType=LLM-as-a-judge2025.07 | 91.4 | 91.37 | |
| Qwen2.5-7B-InstructType=LLM-as-a-judge, Parameters=7B2025.07 | 89.05 | 89 | |
| Qwen2.5-3B-InstructType=LLM-as-a-judge, Parameters=3B2025.07 | 88.35 | 88.35 | |
| Qwen2.5-1.5B-InstructType=LLM-as-a-judge, Parameters=1.5B2025.07 | 83.95 | 83.88 | |
| Omni-JudgeType=LLM-as-a-judge2025.07 | 80.2 | 80.03 | |
| Llama-3-8B-InstructType=LLM-as-a-judge, Parameters=8B2025.07 | 79.95 | 79.8 | |
| General-VerifierType=LLM-as-a-judge2025.07 | 67.65 | 67.46 | |
| math-verifyType=rule-based verifier2025.07 | 66.95 | 63.4 | |
| Qwen2.5-0.5B-InstructType=LLM-as-a-judge, Parameters=0.5B2025.07 | 55.15 | 49.54 |