Multiple-choice Question Answering on Text-only Adaptive Benchmark L4
81Pass@1 AccuracyDeepSeek-R1-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-R1-7BModel Category=Expert models2025.12 | 81 | 86 | |
| AdaptThinkModel Category=Expert models2025.12 | 79 | 80 | |
| Qwen3-OmniModel Category=Omni models2025.12 | 76 | 0 | |
| Omni-AutoThinkModel Category=Omni models2025.12 | 49 | 50 | |
| Qwen2.5-OmniModel Category=Omni models2025.12 | 25 | 26 |