Natural Language Understanding on ITALIC original
87.4Fast AccuracyMinistral-3-8B-Instruct-2512-BF16
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Ministral-3-8B-Instruct-2512-BF16evaluation_type=chat-based2026.05 | 87.4 | 81 | |
| GPT-5 nanoevaluation_type=chat-based2026.05 | 86.8 | 86.6 | |
| gemma-3-12b-itevaluation_type=chat-based2026.05 | 76.5 | 76.5 | |
| Qwen3-8Bevaluation_type=chat-based2026.05 | 70.8 | 70.9 | |
| Llama-3.1-8B-Instructevaluation_type=chat-based2026.05 | 70.7 | 69.9 | |
| Velvet-14Bevaluation_type=chat-based2026.05 | 68.1 | 64.6 | |
| Qwen3-4Bevaluation_type=chat-based2026.05 | 65.4 | 67.6 | |
| gemma-3-4b-itevaluation_type=chat-based2026.05 | 62.5 | 62.8 | |
| Moonlight-16B-A3B-Instructevaluation_type=chat-based2026.05 | 61.7 | 60.9 | |
| EngGPT2-16B-A3Bevaluation_type=chat-based2026.05 | 59.3 | 55.9 | |
| FastwebMIIA-7Bevaluation_type=chat-based2026.05 | 59.3 | 0.3 | |
| Llama-3.2-3B-Instructevaluation_type=chat-based2026.05 | 58.4 | 57.4 | |
| deepseek-moe-16b-chatevaluation_type=chat-based2026.05 | 50.9 | 50 | |
| Minerva-7B-instruct-v1.0evaluation_type=chat-based2026.05 | 49.9 | 45 | |
| LLaMAntino-3-ANITA-8B-Inst-DPO-ITAevaluation_type=chat-based2026.05 | 44.8 | 67.8 | |
| gpt-oss-20bevaluation_type=chat-based2026.05 | 26.1 | 65.3 |