Reasoning on ARC Challenge Italian
93.5Flexible ExtractQwen3-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-8BEvaluation Setting=cot_it_custom2026.05 | 93.5 | |
| GPT-5 nanoEvaluation Setting=cot_it_custom2026.05 | 93 | |
| gpt-oss-20bEvaluation Setting=cot_it_custom2026.05 | 92.7 | |
| Qwen3-4BEvaluation Setting=cot_it_custom2026.05 | 89.8 | |
| gemma-3-12b-itEvaluation Setting=cot_it_custom2026.05 | 87.5 | |
| EngGPT2-16B-A3BEvaluation Setting=cot_it_custom2026.05 | 82.3 | |
| LLaMAntino-3-ANITA-8BEvaluation Setting=cot_it_custom2026.05 | 75.5 | |
| gemma-3-4b-itEvaluation Setting=cot_it_custom2026.05 | 72.5 | |
| Llama-3.1-8B-InstructEvaluation Setting=cot_it_custom2026.05 | 72 | |
| Velvet-14BEvaluation Setting=cot_it_custom2026.05 | 70.8 | |
| Ministral-3-8BEvaluation Setting=cot_it_custom2026.05 | 68 | |
| Moonlight-16B-A3B-InstructEvaluation Setting=cot_it_custom2026.05 | 62.4 | |
| Llama-3.2-3B-InstructEvaluation Setting=cot_it_custom2026.05 | 60 | |
| FastwebMIIA-7BEvaluation Setting=cot_it_custom2026.05 | 58.7 | |
| deepseek-moe-16b-chatEvaluation Setting=cot_it_custom2026.05 | 46.4 | |
| Minerva-7B-instruct-v1.0Evaluation Setting=cot_it_custom2026.05 | 26.1 |