Confidence Estimation on Global-MMLU (test)
0.78AUROCLast Answer
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Last AnswerLanguage=fr, Protocol=Trained and tested on source language, Backbone=Qwen 3 8B2026.05 | 0.78 | 0.58 | 0.19 | 0.16 | |
| Last AnswerLanguage=pl, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.75 | 0.48 | 0.27 | 0.27 | |
| Last AnswerLanguage=es, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.74 | 0.52 | 0.23 | 0.21 | |
| Last QueryLanguage=pl, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.73 | 0.4 | 0.2 | 0.17 | |
| Last AnswerLanguage=en, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.72 | 0.6 | 0.27 | 0.24 | |
| Last AnswerLanguage=ru, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.72 | 0.43 | 0.39 | 0.42 | |
| Last AnswerLanguage=ja, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.72 | 0.44 | 0.31 | 0.31 | |
| Last QueryLanguage=fr, Protocol=Trained and tested on source language, Backbone=Qwen 3 8B2026.05 | 0.71 | 0.49 | 0.21 | 0.16 | |
| Last QueryLanguage=es, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.68 | 0.41 | 0.27 | 0.25 | |
| Last QueryLanguage=ru, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.68 | 0.35 | 0.23 | 0.21 | |
| Last QueryLanguage=en, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.64 | 0.51 | 0.31 | 0.26 | |
| Last QueryLanguage=ja, Protocol=Zero-shot, Backbone=Qwen 3 8B2026.05 | 0.64 | 0.32 | 0.2 | 0.17 | |
| MajorityLanguage=all*, Backbone=Qwen 3 8B2026.05 | 0.5 | 0.63 | 0.26 | 0.26 | |
| Prior prob.Language=all*, Backbone=Qwen 3 8B2026.05 | 0.5 | 0.63 | 0.19 | 0.06 |