Social Interaction Evaluation on SOTOPIA-Hard (Self-Play)
8.06GOAL ScoreAMPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AMPOBackbone=Llama3.1-8B-Instruct, Variant=BC+AMPO2025.05 | 8.06 | 3.68 | |
| AMPOBackbone=Qwen2.5-7B-Instruct, Variant=BC+AMPO2025.05 | 7.85 | 3.54 | |
| ASL w/ BC+GRPOBackbone=Llama3.1-8B-Instruct, Variant=BC+GRPO2025.05 | 7.59 | 3.44 | |
| ASL w/ BC+GRPOBackbone=Qwen2.5-7B-Instruct, Variant=BC+GRPO2025.05 | 7.44 | 3.41 | |
| Qwen2.5-7B-Instruct w/ DSIBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=DSI2025.05 | 7.31 | 3.51 | |
| ASL w/ BCBackbone=Llama3.1-8B-Instruct, Variant=Behavioral Cloning2025.05 | 7.21 | 3.5 | |
| ASL w/ BCBackbone=Qwen2.5-7B-Instruct, Variant=Behavioral Cloning2025.05 | 7.14 | 3.47 | |
| GPT-4oModel Category=Proprietary LLMs2025.05 | 6.97 | 3.46 | |
| Llama3.1-8B-Instruct w/ DSIBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=DSI2025.05 | 6.93 | 3.31 | |
| Qwen2.5-7B-Instruct w/ EPOBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=EPO2025.05 | 6.82 | 3.12 | |
| Llama3.1-8B-Instruct w/ EPOBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=EPO2025.05 | 6.79 | 3.27 | |
| Qwen2.5-7B-Instruct w/ PPDPPBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=PPDPP2025.05 | 6.63 | 3.31 | |
| Qwen2.5-7B-Instruct w/ DATBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=DAT2025.05 | 6.39 | 3.1 | |
| DeepSeek-V3Model Category=Proprietary LLMs2025.05 | 6.34 | 3.09 | |
| Claude-3.5-SonnetModel Category=Proprietary LLMs2025.05 | 6.33 | 3.09 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct2025.05 | 6.21 | 3.01 | |
| Llama3.1-8B-Instruct w/ DATBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=DAT2025.05 | 5.89 | 2.85 | |
| DeepSeek-R1Model Category=Large Reasoning Models2025.05 | 5.86 | 2.73 | |
| OpenAI-o1Model Category=Large Reasoning Models2025.05 | 5.69 | 2.71 | |
| Gemini-2.5-ProModel Category=Large Reasoning Models2025.05 | 5.67 | 2.55 | |
| QwQ-32BModel Category=Large Reasoning Models2025.05 | 5.35 | 2.41 | |
| Llama3.1-8B-Instruct w/ PPDPPBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=PPDPP2025.05 | 5.34 | 2.57 | |
| OpenAI-o3-miniModel Category=Large Reasoning Models2025.05 | 5.14 | 2.36 | |
| Llama3.1-8B-InstructBackbone=Llama3.1-8B-Instruct2025.05 | 5.11 | 2.36 |