Open-ended instruction following on BPO Eval (test)
59.5Win Rate (A)BPO + Vicuna-v1.3 13B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BPO + Vicuna-v1.3 13BBase LLM=Vicuna-v1.3, Model A=13B + BPO, Model B=13B, Evaluation LLM=GPT-42023.11 | 59.5 | 6 | 34.5 | 13.1 | |
| BPO + Llama-2-chat 70BBase LLM=Llama-2-chat, Model A=70B + BPO, Model B=70B, Evaluation LLM=GPT-42023.11 | 53.5 | 11 | 35.5 | 16.8 | |
| BPO + Llama-2-chat 7BBase LLM=Llama-2-chat, Model A=7B + BPO, Model B=7B, Evaluation LLM=GPT-42023.11 | 53 | 10.5 | 36.5 | 17.4 | |
| BPO + Llama-2-chat 13BBase LLM=Llama-2-chat, Model A=13B + BPO, Model B=13B, Evaluation LLM=GPT-42023.11 | 53 | 12.5 | 34.5 | 18.1 | |
| BPO + Llama-2-chat 13B (Cross-size)Base LLM=Llama-2-chat, Model A=13B + BPO, Model B=70B, Evaluation LLM=GPT-42023.11 | 51 | 7 | 42 | 11.9 | |
| BPO + Vicuna-v1.3 7BBase LLM=Vicuna-v1.3, Model A=7B + BPO, Model B=7B, Evaluation LLM=GPT-42023.11 | 46 | 22 | 32 | 18.5 | |
| BPO + Llama-2-chat 7B (Cross-size)Base LLM=Llama-2-chat, Model A=7B + BPO, Model B=70B, Evaluation LLM=GPT-42023.11 | 40 | 5 | 55 | -7.1 |