Bilingual Mathematical Reasoning on MathBench EN
40.1AccuracyMixtral-8x7B-v0.1
Evaluation Results
| Method | Links | |
|---|---|---|
| Mixtral-8x7B-v0.1shots=0-shot & 4-shot, model_size_group=~20B2024.03 | 40.1 | |
| Qwen-14Bshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 39.7 | |
| InternLM2-20Bshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 36.9 | |
| InternLM2-7Bshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 33.9 | |
| ChatGLM3-6B-Baseshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 29 | |
| Mistral-7B-v0.1shots=0-shot & 4-shot, model_size_group=~7B2024.03 | 27.6 | |
| InternLM2-20B-Baseshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 26.8 | |
| Baichuan2-13B-Baseshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 24.4 | |
| Qwen-7Bshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 21.8 | |
| Llama2-13Bshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 18.8 | |
| InternLM2-7B-Baseshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 18.5 | |
| Baichuan2-7B-Baseshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 15.2 | |
| Llama2-7Bshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 7.6 |