Conversational Question Answering on CHATRAG BENCH
69.09Average ScoreChatQA-1.0-70B vs GPT-4-0613
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ChatQA-1.0-70B vs GPT-4-0613Evaluation Protocol=Human A/B Testing, Outcome=Tie Rate2024.01 | 69.09 | 68 | 73.33 | 77.22 | 80 | 57.78 | 67.78 | 61.67 | 60.69 | 78.33 | 66.11 | |
| GPT-4-0613Evaluation Protocol=Human A/B Testing, Outcome=Win Rate vs ChatQA-1.0-70B2024.01 | 17.1 | 17.71 | 15 | 11.67 | 12.22 | 19.44 | 15.55 | 27.22 | 20 | 13.89 | 18.33 | |
| ChatQA-1.0-70BEvaluation Protocol=Human A/B Testing, Outcome=Win Rate vs GPT-4-06132024.01 | 13.81 | 14.29 | 11.67 | 11.11 | 7.78 | 22.78 | 16.67 | 11.11 | 19.31 | 7.78 | 15.56 |