Role Identification on Role Identification variant of Turing-test (test)
25.77Error Rate (%)Stephanie1
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Stephanie1Base Model=GPT5.2, Evaluation Protocol=Human evaluation2026.01 | 25.77 | 10.31 | 63.92 | 36.08 | |
| Stephanie2Base Model=Llama3.1-8B, Evaluation Protocol=Human evaluation2026.01 | 28.12 | 28.12 | 43.76 | 56.24 | |
| Stephanie1Base Model=Llama3.1-8B, Evaluation Protocol=Human evaluation2026.01 | 29.41 | 17.65 | 52.94 | 47.06 | |
| Stephanie1Base Model=GPT5.2, Evaluation Protocol=Automatic evaluation2026.01 | 31.24 | 7.11 | 61.65 | 38.35 | |
| Stephanie1Base Model=Deepseek-V3, Evaluation Protocol=Human evaluation2026.01 | 34.07 | 16.48 | 49.45 | 50.55 | |
| Stephanie2Base Model=GPT5.2, Evaluation Protocol=Human evaluation2026.01 | 34.4 | 15.2 | 50.4 | 49.6 | |
| Stephanie1Base Model=Deepseek-V3, Evaluation Protocol=Automatic evaluation2026.01 | 36.11 | 4.63 | 59.26 | 40.74 | |
| Stephanie2Base Model=Deepseek-V3, Evaluation Protocol=Human evaluation2026.01 | 38.95 | 15.79 | 45.26 | 54.74 | |
| Stephanie2Base Model=Deepseek-V3, Evaluation Protocol=Automatic evaluation2026.01 | 39.98 | 3.81 | 56.21 | 43.79 | |
| Stephanie2Base Model=Llama3.1-8B, Evaluation Protocol=Automatic evaluation2026.01 | 47.22 | 13.89 | 38.89 | 61.11 | |
| Stephanie1Base Model=Llama3.1-8B, Evaluation Protocol=Automatic evaluation2026.01 | 48.1 | 6.67 | 45.23 | 54.77 | |
| Stephanie2Base Model=GPT5.2, Evaluation Protocol=Automatic evaluation2026.01 | 51.26 | 9.26 | 39.48 | 60.52 |