Role-playing performance evaluation on Bandori (test)
88.38PoPiPa ScoreCDT-Lite
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CDT-LiteCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Validation Model=deberta-v3-base2026.01 | 88.38 | 80.49 | 82.47 | 72.81 | 79.66 | 78.67 | 79.51 | 70.33 | 79.04 | |
| CDTCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Validation Model=gpt-4.1-mini2026.01 | 84.25 | 79.92 | 78.93 | 71.93 | 80.03 | 77.33 | 80.08 | 69.17 | 77.71 | |
| ETACategory=Data-driven, RP Model=llama-3.1-8b-instruct, Method Strategy=Extract-Then-Aggregate2026.01 | 75.29 | 72.49 | 78 | 70.91 | 78.92 | 66.82 | 72.68 | 62.89 | 72.25 | |
| Human ProfileCategory=Human, Grounding Information Source=Fandom wiki2026.01 | 73.73 | 72.43 | 77.11 | 70.08 | 73.14 | 68.08 | 71.74 | 63.91 | 71.28 | |
| RICLCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Method Strategy=Retrieval-based In-Context Learning2026.01 | 73.56 | 67.56 | 73.06 | 67.24 | 73.63 | 65.04 | 69.44 | 61.1 | 68.86 | |
| Codified Human ProfileCategory=Human, Method Strategy=Symbolic rules to grounding functions2026.01 | 73.02 | 74 | 78.65 | 71.23 | 72.47 | 69.14 | 71.41 | 65.02 | 71.87 | |
| Fine-tuningCategory=Data-driven, RP Model=llama-3.1-8b-instruct2026.01 | 69.52 | 62.76 | 64.63 | 61.83 | 62.39 | 62.35 | 62.64 | 56.72 | 62.86 | |
| VanillaCategory=Data-driven, RP Model=llama-3.1-8b-instruct2026.01 | 66.39 | 66.76 | 68.29 | 66.83 | 65.13 | 64.06 | 67.2 | 59.37 | 65.5 |