Instruction Following on Vicuna benchmark
8.09GPT-4 Evaluation Scorellama2 → CP → FT + chat vector
Evaluation Results
| Method | Links | |
|---|---|---|
| llama2 → CP → FT + chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=true2023.10 | 8.09 | |
| llama2 → CP → FT + 0.5 chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=true2023.10 | 8.02 | |
| llama2 → CP → FT + 0.5 chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=false2023.10 | 7.89 | |
| llama2 → CP → FT + chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=false2023.10 | 7.86 | |
| llama2 → CP → FTModel Base=Chinese-LLAMA 13B, System Prompt=false2023.10 | 7.58 | |
| llama2 → CP → FTModel Base=Chinese-LLAMA 13B, System Prompt=true2023.10 | 7.47 | |
| llama2 → CP → FT + chat vectorModel Base=Traditional Chinese LLaMA 13B, System Prompt=false2023.10 | 7.37 | |
| llama2 → CP + chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=false2023.10 | 7.07 | |
| llama2 → CP → FT + chat vectorModel Base=Traditional Chinese LLaMA 13B, System Prompt=true2023.10 | 7.06 | |
| llama2 → CP + chat vectorModel Base=Traditional Chinese LLaMA 13B, System Prompt=false2023.10 | 7.03 | |
| llama2 → CP + chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=true2023.10 | 6.7 | |
| llama2-chat → CP → FTModel Base=Traditional Chinese LLaMA 13B, System Prompt=false2023.10 | 6.46 | |
| llama2 → CP → FTModel Base=Traditional Chinese LLaMA 13B, System Prompt=false2023.10 | 6.13 | |
| llama2 → CP + chat vectorModel Base=Traditional Chinese LLaMA 13B, System Prompt=true2023.10 | 6.04 | |
| llama2-chat → CP → FTModel Base=Traditional Chinese LLaMA 13B, System Prompt=true2023.10 | 5.89 | |
| llama2 → CP → FTModel Base=Traditional Chinese LLaMA 13B, System Prompt=true2023.10 | 5.5 | |
| llama2 → CP + 0.5 chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=true2023.10 | 5.06 | |
| llama2 → CP + 0.5 chat vectorModel Base=Chinese-LLAMA 13B, System Prompt=false2023.10 | 4.61 |