Language Understanding on MMLU German (test)
73AccuracyTrinity Large (MoE)
Evaluation Results
| Method | Links | |
|---|---|---|
| Trinity Large (MoE)evaluation_mode=Zero-shot, architecture=MoE2026.02 | 73 | |
| Qwen3 8Bevaluation_mode=Zero-shot, parameters=8B2026.02 | 69 | |
| Qwen3 4Bevaluation_mode=Zero-shot, parameters=4B2026.02 | 65 | |
| Datology 8Bevaluation_mode=Zero-shot, parameters=8B2026.02 | 54 | |
| Llama-3.1 8Bevaluation_mode=Zero-shot, parameters=8B2026.02 | 52 | |
| Granite-4.0 Microevaluation_mode=Zero-shot2026.02 | 52 | |
| SmolLM3 3Bevaluation_mode=Zero-shot, parameters=3B2026.02 | 50 | |
| LFM2.5 1.2Bevaluation_mode=Zero-shot, parameters=1.2B2026.02 | 48 | |
| Datology 3Bevaluation_mode=Zero-shot, parameters=3B2026.02 | 44 | |
| Llama-3.2 3Bevaluation_mode=Zero-shot, parameters=3B2026.02 | 44 | |
| ArrowBase Model=Phi-2, Evaluation Setting=Zero-shot2025.05 | 26.91 | |
| Llama-3.2 1Bevaluation_mode=Zero-shot, parameters=1B2026.02 | 26 | |
| GenKnowSubBase Model=Phi-2, Evaluation Setting=Zero-shot, Setting=Avg2025.05 | 25.5 | |
| GenKnowSubBase Model=Phi-2, Evaluation Setting=Zero-shot, Setting=En2025.05 | 24.91 | |
| Phi-2Base Model=Phi-2, Evaluation Setting=Zero-shot2025.05 | 24.19 |