Medical Hypothesis Verification on 64-hypothesis tiered benchmark L1–L5
81Run-level Verdict AccuracyVERITAS (frontier)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| VERITAS (frontier)Model Family=GPT-5.2, Method logic=Multi-agent system2026.04 | 81 | 70 | 76.3 | 81.4 | 87.5 | |
| SMd: AgenticModel Family=GPT-5.2, Method logic=Agentic (iterative)2026.04 | 76.6 | 64.7 | 64.4 | 74.6 | — | |
| SMc: Code via APIModel Family=GPT-5.2, Method logic=Code via API (one-shot)2026.04 | 72.8 | 52.9 | 61 | 76.3 | — | |
| VERITAS (local)Model Family=GPT-OSS-20B, Method logic=Multi-agent system, Model Deployment=locally-deployed 8–30B models2026.04 | 71.4 | 63.4 | 67.8 | 71.2 | 78.1 | |
| SMb: Code on featuresModel Family=GPT-5.2, Method logic=Code on pre-computed features2026.04 | 70.7 | 63.7 | 69.5 | 69.5 | — | |
| SMd: AgenticModel Family=GPT-OSS-20B, Method logic=Agentic (iterative)2026.04 | 66 | 46.3 | 52.5 | 66.1 | — | |
| SMe: PipelineModel Family=GPT-5.2, Method logic=Pipeline (structured chain-of-thought)2026.04 | 65.5 | 60.8 | 59.3 | 61 | — | |
| SMb: Code on featuresModel Family=GPT-OSS-20B, Method logic=Code on pre-computed features2026.04 | 63.8 | 58.1 | 61 | 62.7 | — | |
| SMc: Code via APIModel Family=GPT-OSS-20B, Method logic=Code via API (one-shot)2026.04 | 63.2 | 39 | 40.7 | 59.3 | — | |
| SMe: PipelineModel Family=GPT-OSS-20B, Method logic=Pipeline (structured chain-of-thought)2026.04 | 58 | 55.1 | 57.6 | 61 | — | |
| SMa: Direct reasoningModel Family=GPT-OSS-20B, Method logic=Direct reasoning2026.04 | 56.1 | — | — | 55.9 | — | |
| SMa: Direct reasoningModel Family=GPT-5.2, Method logic=Direct reasoning2026.04 | 55.9 | — | — | 55.9 | — |