Multiple-choice Question Answering on GPQA (accuracy)
95.2Accuracy (%)Pass@100
Evaluation Results
| Method | Links | |
|---|---|---|
| Pass@100Generator Model=Llama 3.1 8B Instruct, Supervision Setting=Oracle2026.04 | 95.2 | |
| Pass@100Generator Model=Llama 3.3 70B Instruct, Supervision Setting=Oracle2026.04 | 81 | |
| BDR-128Model=GLM-4.7 (358B)2026.04 | 73.74 | |
| BF16Model=GLM-4.7 (358B)2026.04 | 73.23 | |
| INT4Model=GLM-4.7 (358B)2026.04 | 73.23 | |
| BDR-16Model=GLM-4.7 (358B)2026.04 | 72.9 | |
| BDR-64Model=GLM-4.7 (358B)2026.04 | 70.87 | |
| LogisticGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 69.3 | |
| FUSEGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 66.8 | |
| WeaverGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 66.4 | |
| Naive BayesGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 60.5 | |
| Naive EnsembleGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 60.1 | |
| OBV (Oracle Best Verifier)Generator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 59.1 | |
| BDR-64 (K)Model=Qwen3-8B2026.04 | 56.97 | |
| BF16Model=Qwen3-8B2026.04 | 56.67 | |
| BDR-16Model=Qwen3-8B2026.04 | 54.85 | |
| BDR-128Model=Qwen3-8B2026.04 | 54.85 | |
| BDR-64Model=Qwen3-8B2026.04 | 54.55 | |
| Majority VoteGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 47.4 | |
| WeaverGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 47.1 | |
| Pass@1Generator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 42.9 | |
| OBV (Oracle Best Verifier)Generator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 41.9 | |
| FUSEGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 40.5 | |
| Naive EnsembleGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 37.8 | |
| Naive BayesGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 37.6 | |
| LogisticGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 35.9 | |
| Majority VoteGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 30.5 | |
| DS2-INSTRUCTModel Family=Llama32026.03 | 30.35 | |
| DS2-INSTRUCTModel Family=Mistral2026.03 | 30.16 | |
| Pass@1Generator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 28.3 | |
| Self-InstructModel Family=Qwen2.52026.03 | 27.18 | |
| DS2-INSTRUCTModel Family=Qwen2.52026.03 | 26.35 | |
| Zero-ShotModel Family=Qwen2.52026.03 | 26.32 | |
| InstructMixModel Family=Qwen2.52026.03 | 26.09 | |
| ExploreInstructModel Family=Qwen2.52026.03 | 25.93 | |
| InstructMixModel Family=Mistral2026.03 | 23.91 | |
| Self-InstructModel Family=Llama32026.03 | 19.87 | |
| ExploreInstructModel Family=Mistral2026.03 | 19.47 | |
| ExploreInstructModel Family=Llama32026.03 | 18.24 | |
| InstructMixModel Family=Llama32026.03 | 17.52 | |
| Zero-ShotModel Family=Llama32026.03 | 12.51 | |
| Zero-ShotModel Family=Mistral2026.03 | 8.64 | |
| Self-InstructModel Family=Mistral2026.03 | 8.15 | |
| INT4Model=Qwen3-8B2026.04 | 0 |