Multi-turn Chat Conversation on TURNWISEEVAL (val)
83.5TW-Absolute ScoreGPT-5 Chat
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5 ChatJudge=GPT-4.12026.03 | 83.5 | 40.2 | 5 | |
| GPT-4.1Judge=GPT-4.12026.03 | 82.5 | 42 | 0.9 | |
| GPT-5.2Judge=GPT-4.12026.03 | 82.1 | 47.6 | 1.3 | |
| GPT-5 NanoJudge=GPT-4.12026.03 | 68.2 | 41.1 | 0.2 | |
| Qwen 3 32BThinking enabled=false, Judge=GPT-4.12026.03 | 67.7 | 48.5 | 1.9 | |
| Olmo 3.1 32BModel type=Instruct, Judge=GPT-4.12026.03 | 52.4 | 34 | 7.7 | |
| Qwen 3 8BThinking enabled=false, Judge=GPT-4.12026.03 | 48.9 | 53.2 | 0.8 | |
| Olmo 3 7BModel type=Instruct, Judge=GPT-4.12026.03 | 36.8 | 38.9 | 5.4 | |
| Llama 3.1 70BJudge=GPT-4.12026.03 | 26.3 | 36.5 | 1.4 | |
| Llama 3.1 8BJudge=GPT-4.12026.03 | 18.7 | 40.2 | 2.4 |