Role-playing performance evaluation on Fandom (test)
62.17Haruhi Adherence ScoreCDT-Lite
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CDT-LiteCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Validation Model=deberta-v3-base2026.01 | 62.17 | 57.24 | 59.79 | 67 | 59.04 | 57.26 | 64.27 | 61.32 | 61.01 | |
| CDTCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Validation Model=gpt-4.1-mini2026.01 | 61.16 | 57.93 | 60.35 | 66.34 | 58.57 | 57.4 | 63.79 | 61.05 | 60.82 | |
| ETACategory=Data-driven, RP Model=llama-3.1-8b-instruct, Method Strategy=Extract-Then-Aggregate2026.01 | 60.54 | 53.83 | 58 | 63.29 | 57.12 | 51 | 55.28 | 56.23 | 56.91 | |
| Codified Human ProfileCategory=Human, Method Strategy=Symbolic rules to grounding functions2026.01 | 57.94 | 55.93 | 59.38 | 65.56 | 57.01 | 56.56 | 62.07 | 59.97 | 59.3 | |
| RICLCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Method Strategy=Retrieval-based In-Context Learning2026.01 | 56.83 | 55.74 | 56.86 | 62.8 | 56.33 | 49.77 | 56.46 | 52.25 | 56.01 | |
| Human ProfileCategory=Human, Grounding Information Source=Fandom wiki2026.01 | 55.87 | 55.86 | 59.14 | 64.75 | 58.54 | 55.11 | 59.35 | 57.98 | 58.33 | |
| VanillaCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Conditioning=Direct prompting without additional info2026.01 | 55.08 | 49.92 | 56.1 | 62.49 | 55.66 | 54.66 | 57.05 | 53.56 | 55.57 | |
| Fine-tuningCategory=Data-driven, RP Model=llama-3.1-8b-instruct, Training=Supervised training on character-specific pairs2026.01 | 51.49 | 51.01 | 49.14 | 50.79 | 44.07 | 34.92 | 41.2 | 42.84 | 45.68 |