Multitask Language Understanding on MMLU Multilingual translated (test)
90.4Accuracy (Arabic)OpenAI o3-high
Evaluation Results
| Method | Links | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI o3-highshots=0, prompting=chain-of-thought2025.12 | 90.4 | — | 87.8 | 89.3 | 90.6 | 90.5 | 89.8 | 89.8 | 91.2 | 89 | 89.3 | 91 | 91.1 | 86 | 78 | |
| gpt-5-thinkingshots=0, prompting=chain-of-thought2025.12 | 90.3 | — | 89.2 | 90.2 | 90.1 | 89.6 | 89.9 | 90.9 | 90.8 | 89.8 | 89.6 | 91 | 91 | 88 | 80.6 | |
| gpt-5-mainshots=0, prompting=chain-of-thought2025.12 | 85.7 | — | 85 | 86.7 | 87.5 | 86.6 | 86.1 | 87.2 | 87.6 | 86.5 | 85.4 | 87.9 | 88.1 | 81.5 | 66.4 |