Social Agent Evaluation on Artificial Social Agent Questionnaire (ASAQ) Standard (Evaluation set)
2.12Usabilitymulti-LLM hierarchical conversational agent
Evaluation Results
| Method | Links | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| multi-LLM hierarchical conversational agentInteraction Order=Custom chatbot first, Group=CC2024.06 | 2.12 | 2.25 | 1.88 | 2 | 1.38 | 0.88 | 1.5 | 0.88 | 2 | 2.12 | 2.5 | 1.75 | 1.5 | — | — | — | — | |
| GPTInteraction Order=GPT first, Group=GG2024.06 | 2 | 1.5 | 1.5 | 1.12 | 0.25 | 1 | 1 | 0.75 | 1.62 | 2 | 2.12 | 1.62 | 0.88 | — | — | — | — | |
| multi-LLM hierarchical conversational agentInteraction Order=GPT first, Group=GC2024.06 | 1.87 | 1.75 | 1.75 | 1.75 | 1.25 | 1.75 | 1.25 | 1.13 | 1.75 | 1.75 | 2.12 | 2.25 | 1.5 | — | — | — | — | |
| GPTInteraction Order=Custom chatbot first, Group=CG2024.06 | 1.25 | 1 | 1.12 | 0.5 | 0.25 | 0.25 | 0 | -0.38 | 1.62 | 0.63 | 1.38 | 0.5 | 0.25 | — | — | — | — |