Multiple-choice Question Answering on Text-only Adaptive Benchmark L1
93Pass@1 AccuracyAdaptThink
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AdaptThinkModel Category=Expert models2025.12 | 93 | 93 | |
| DeepSeek-R1-7BModel Category=Expert models2025.12 | 92 | 95 | |
| Qwen3-OmniModel Category=Omni models2025.12 | 91 | 0 | |
| Omni-AutoThinkModel Category=Omni models2025.12 | 83 | 39 | |
| Qwen2.5-OmniModel Category=Omni models2025.12 | 60 | 31 |