Scientific Reasoning on SciEval
73.93ScoreGPT-4
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4Parameter Scale=API2024.01 | 73.93 | |
| Llama3-8B-InstructEvaluation Protocol=zero-shot, Parameter Scale=6B~7B2024.01 | 71.38 | |
| Llama3-8B-InstructEvaluation Protocol=few-shot, Parameter Scale=6B~7B2024.01 | 71.38 | |
| SciGLMBackbone=ChatGLM3-32B-Base, Parameter Scale=30B~32B2024.01 | 69.84 | |
| ChatGLM3-32B-BaseParameter Scale=30B~32B2024.01 | 67.38 | |
| GPT-3.5-turboParameter Scale=API2024.01 | 66.97 | |
| Llama3-8B-Instruct + SciInstructFine-tuning=SciInstruct, Parameter Scale=6B~7B2024.01 | 66.47 | |
| Mistral-7B: MetaMATH + SciInstructFine-tuning=SciInstruct, Parameter Scale=6B~7B2024.01 | 64.16 | |
| Mistral-7B: MetaMATHEvaluation Protocol=zero-shot, Parameter Scale=6B~7B2024.01 | 63.61 | |
| Mistral-7B: MetaMATHEvaluation Protocol=few-shot, Parameter Scale=6B~7B2024.01 | 63.61 | |
| Claude-v1.3Parameter Scale=API2024.01 | 63.45 | |
| SciGLMBackbone=ChatGLM3-6B-Base, Parameter Scale=6B~7B2024.01 | 62.09 | |
| ChatGLM3-6B-BaseParameter Scale=6B~7B2024.01 | 61.69 | |
| ChatGLM3-6BParameter Scale=6B~7B2024.01 | 56.56 | |
| Vicuna-13BParameter Scale=12B~13B2024.01 | 53.93 | |
| ChatGLM2-6BParameter Scale=6B~7B2024.01 | 53.02 | |
| Galactica-6.7BParameter Scale=6B~7B2024.01 | 50.87 | |
| ChatGLM2-6B-BaseParameter Scale=6B~7B2024.01 | 50.38 | |
| LLaMA-2-13BParameter Scale=12B~13B2024.01 | 36.96 | |
| LLaMA-2-7BParameter Scale=6B~7B2024.01 | 28.37 |