Bilingual Mathematical Reasoning on MathBench CN
46.2AccuracyQwen-14B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-14Bshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 46.2 | |
| InternLM2-20Bshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 38.7 | |
| InternLM2-7Bshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 37.4 | |
| ChatGLM3-6B-Baseshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 32.5 | |
| Mixtral-8x7B-v0.1shots=0-shot & 4-shot, model_size_group=~20B2024.03 | 31.6 | |
| Baichuan2-13B-Baseshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 28.4 | |
| InternLM2-20B-Baseshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 27 | |
| Qwen-7Bshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 26 | |
| Baichuan2-7B-Baseshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 21.3 | |
| Mistral-7B-v0.1shots=0-shot & 4-shot, model_size_group=~7B2024.03 | 19.3 | |
| InternLM2-7B-Baseshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 14.2 | |
| Llama2-13Bshots=0-shot & 4-shot, model_size_group=~20B2024.03 | 10.6 | |
| Llama2-7Bshots=0-shot & 4-shot, model_size_group=~7B2024.03 | 4.1 |