Scientific Reasoning on CEval Sci
66.19ScoreSciGLM
Evaluation Results
| Method | Links | |
|---|---|---|
| SciGLMBackbone=ChatGLM3-32B-Base, Parameter Scale=30B~32B2024.01 | 66.19 | |
| ChatGLM3-32B-BaseParameter Scale=30B~32B2024.01 | 64.29 | |
| GPT-4Parameter Scale=API2024.01 | 60.55 | |
| SciGLMBackbone=ChatGLM3-6B-Base, Parameter Scale=6B~7B2024.01 | 60 | |
| ChatGLM3-6B-BaseParameter Scale=6B~7B2024.01 | 54.29 | |
| GPT-3.5-turboParameter Scale=API2024.01 | 46.83 | |
| ChatGLM2-6BParameter Scale=6B~7B2024.01 | 45.71 | |
| Claude-v1.3Parameter Scale=API2024.01 | 44.64 | |
| ChatGLM2-6B-BaseParameter Scale=6B~7B2024.01 | 40.95 | |
| ChatGLM3-6BParameter Scale=6B~7B2024.01 | 38.57 | |
| Mistral-7B: MetaMATH + SciInstructFine-tuning=SciInstruct, Parameter Scale=6B~7B2024.01 | 38.1 | |
| Galactica-30BParameter Scale=30B~32B2024.01 | 35.53 | |
| Llama3-8B-Instruct + SciInstructFine-tuning=SciInstruct, Parameter Scale=6B~7B2024.01 | 34.76 | |
| LLaMA-2-7BParameter Scale=6B~7B2024.01 | 30 | |
| Llama3-8B-InstructEvaluation Protocol=zero-shot, Parameter Scale=6B~7B2024.01 | 27.62 | |
| Llama3-8B-InstructEvaluation Protocol=few-shot, Parameter Scale=6B~7B2024.01 | 23.33 | |
| Mistral-7B: MetaMATHEvaluation Protocol=few-shot, Parameter Scale=6B~7B2024.01 | 19.52 | |
| LLaMA-2-13BParameter Scale=12B~13B2024.01 | 19.05 | |
| Galactica-6.7BParameter Scale=6B~7B2024.01 | 11.44 | |
| Mistral-7B: MetaMATHEvaluation Protocol=zero-shot, Parameter Scale=6B~7B2024.01 | 8.57 |