Open-ended instruction following on Dolly Eval
54A Win RateBPO + Llama-2-chat 13B (Cross-size)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BPO + Llama-2-chat 13B (Cross-size)Base LLM=Llama-2-chat, Model A=13B + BPO, Model B=70B, Evaluation LLM=GPT-42023.11 | 54 | 6.5 | 39.5 | 11.9 | |
| BPO + Llama-2-chat 7BBase LLM=Llama-2-chat, Model A=7B + BPO, Model B=7B, Evaluation LLM=GPT-42023.11 | 52 | 9.5 | 38.5 | 17.4 | |
| BPO + Vicuna-v1.3 13BBase LLM=Vicuna-v1.3, Model A=13B + BPO, Model B=13B, Evaluation LLM=GPT-42023.11 | 52 | 8 | 40 | 13.1 | |
| BPO + Llama-2-chat 70BBase LLM=Llama-2-chat, Model A=70B + BPO, Model B=70B, Evaluation LLM=GPT-42023.11 | 51 | 18 | 31 | 16.8 | |
| BPO + Llama-2-chat 13BBase LLM=Llama-2-chat, Model A=13B + BPO, Model B=13B, Evaluation LLM=GPT-42023.11 | 50.5 | 13.5 | 36 | 18.1 | |
| BPO + Llama-2-chat 7B (Cross-size)Base LLM=Llama-2-chat, Model A=7B + BPO, Model B=70B, Evaluation LLM=GPT-42023.11 | 49 | 2 | 49 | -7.1 | |
| BPO + Vicuna-v1.3 7BBase LLM=Vicuna-v1.3, Model A=7B + BPO, Model B=7B, Evaluation LLM=GPT-42023.11 | 47 | 22 | 31 | 18.5 |