Response Generation on Vicuna 80 prompts (test)
1,348EloGPT-4
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4Judge=GPT-4, # Prompts=802023.05 | 1,348 | 1 | 1 | |
| GPT-4Judge=Human raters, # Prompts=802023.05 | 1,176 | 1 | 1 | |
| Guanaco-65BJudge=Human raters, # Prompts=802023.05 | 1,023 | 2 | 2 | |
| Guanaco-65BJudge=GPT-4, # Prompts=802023.05 | 1,022 | 2 | 2 | |
| Guanaco-7BJudge=Human raters, # Prompts=802023.05 | 1,010 | 3 | 7 | |
| Guanaco-33BJudge=Human raters, # Prompts=802023.05 | 1,009 | 4 | 4 | |
| Guanaco-33BJudge=GPT-4, # Prompts=802023.05 | 992 | 3 | 4 | |
| Vicuna-13BJudge=Human raters, # Prompts=802023.05 | 984 | 5 | 5 | |
| Guanaco-13BJudge=Human raters, # Prompts=802023.05 | 975 | 6 | 6 | |
| Vicuna-13BJudge=GPT-4, # Prompts=802023.05 | 974 | 4 | 5 | |
| ChatGPT-3.5 TurboJudge=GPT-4, # Prompts=802023.05 | 966 | 5 | 5 | |
| ChatGPT-3.5 TurboJudge=Human raters, # Prompts=802023.05 | 916 | 7 | 5 | |
| Guanaco-13BJudge=GPT-4, # Prompts=802023.05 | 913 | 6 | 6 | |
| BardJudge=Human raters, # Prompts=802023.05 | 909 | 8 | 8 | |
| BardJudge=GPT-4, # Prompts=802023.05 | 902 | 7 | 8 | |
| Guanaco-7BJudge=GPT-4, # Prompts=802023.05 | 879 | 8 | 7 |