Machine Reading Comprehension on BELEBELE German
92AccuracyTrinity Large (MoE)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Trinity Large (MoE)evaluation_mode=Zero-shot, architecture=MoE2026.02 | 92 | — | — | |
| Qwen3 8Bevaluation_mode=Zero-shot, parameters=8B2026.02 | 88 | — | — | |
| Qwen3 4Bevaluation_mode=Zero-shot, parameters=4B2026.02 | 86 | — | — | |
| Granite-4.0 Microevaluation_mode=Zero-shot2026.02 | 79 | — | — | |
| Datology 8Bevaluation_mode=Zero-shot, parameters=8B2026.02 | 76 | — | — | |
| Llama-3.1 8Bevaluation_mode=Zero-shot, parameters=8B2026.02 | 69 | — | — | |
| Datology 3Bevaluation_mode=Zero-shot, parameters=3B2026.02 | 62 | — | — | |
| SmolLM3 3Bevaluation_mode=Zero-shot, parameters=3B2026.02 | 61 | — | — | |
| Llama-3.2 3Bevaluation_mode=Zero-shot, parameters=3B2026.02 | 56 | — | — | |
| LFM2.5 1.2Bevaluation_mode=Zero-shot, parameters=1.2B2026.02 | 52 | — | — | |
| Llama-3.2 1Bevaluation_mode=Zero-shot, parameters=1B2026.02 | 28 | — | — | |
| CLOModel=Llama-2-7B, Zero-shot=true2025.05 | — | 32.7 | 38.1 | |
| CLOModel=Llama-2-13B, Zero-shot=true2025.05 | — | 52.3 | 59.8 | |
| CLOModel=Llama-3-8B, Zero-shot=true2025.05 | — | 64.1 | 74.7 | |
| CLOModel=Mistral-7B-v0.1, Zero-shot=true2025.05 | — | 53 | 72.3 | |
| CLOModel=Qwen2.5-3B, Zero-shot=true2025.05 | — | 73.1 | 82 | |
| SFTModel=Llama-2-7B, Zero-shot=true2025.05 | — | 31.7 | 36.9 | |
| SFTModel=Llama-2-13B, Zero-shot=true2025.05 | — | 51.1 | 59.6 | |
| SFTModel=Llama-3-8B, Zero-shot=true2025.05 | — | 62.9 | 73.1 | |
| SFTModel=Mistral-7B-v0.1, Zero-shot=true2025.05 | — | 47.6 | 67.9 | |
| SFTModel=Qwen2.5-3B, Zero-shot=true2025.05 | — | 70.3 | 81.3 |