Social Interaction Evaluation on SOTOPIA-Hard GPT-4o-as-Partner
7.68Goal ScoreAMPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AMPOBackbone=Llama3.1-8B-Instruct, Variant=BC+AMPO2025.05 | 7.68 | 3.74 | |
| AMPOBackbone=Qwen2.5-7B-Instruct, Variant=BC+AMPO2025.05 | 7.5 | 3.65 | |
| ASL w/ BC+GRPOBackbone=Llama3.1-8B-Instruct, Variant=BC+GRPO2025.05 | 7.3 | 3.54 | |
| ASL w/ BC+GRPOBackbone=Qwen2.5-7B-Instruct, Variant=BC+GRPO2025.05 | 7.2 | 3.5 | |
| ASL w/ BCBackbone=Qwen2.5-7B-Instruct, Variant=Behavioral Cloning2025.05 | 7.15 | 3.5 | |
| ASL w/ BCBackbone=Llama3.1-8B-Instruct, Variant=Behavioral Cloning2025.05 | 7.14 | 3.52 | |
| GPT-4oModel Category=Proprietary LLMs2025.05 | 6.97 | 3.46 | |
| Llama3.1-8B-Instruct w/ EPOBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=EPO2025.05 | 6.91 | 3.53 | |
| Qwen2.5-7B-Instruct w/ DSIBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=DSI2025.05 | 6.87 | 3.42 | |
| Llama3.1-8B-Instruct w/ DSIBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=DSI2025.05 | 6.84 | 3.41 | |
| Qwen2.5-7B-Instruct w/ EPOBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=EPO2025.05 | 6.81 | 3.51 | |
| Qwen2.5-7B-Instruct w/ DATBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=DAT2025.05 | 6.78 | 3.36 | |
| Qwen2.5-7B-Instruct w/ PPDPPBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=PPDPP2025.05 | 6.76 | 3.35 | |
| Gemini-2.5-ProModel Category=Large Reasoning Models2025.05 | 6.7 | 3.09 | |
| DeepSeek-V3Model Category=Proprietary LLMs2025.05 | 6.69 | 3.31 | |
| OpenAI-o1Model Category=Large Reasoning Models2025.05 | 6.65 | 3.2 | |
| Claude-3.5-SonnetModel Category=Proprietary LLMs2025.05 | 6.64 | 3.3 | |
| OpenAI-o3-miniModel Category=Large Reasoning Models2025.05 | 6.33 | 2.98 | |
| Llama3.1-8B-Instruct w/ PPDPPBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=PPDPP2025.05 | 6.3 | 3.1 | |
| Llama3.1-8B-InstructBackbone=Llama3.1-8B-Instruct2025.05 | 6.21 | 3.05 | |
| DeepSeek-R1Model Category=Large Reasoning Models2025.05 | 6.2 | 2.95 | |
| QwQ-32BModel Category=Large Reasoning Models2025.05 | 6.19 | 2.91 | |
| Llama3.1-8B-Instruct w/ DATBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=DAT2025.05 | 6.18 | 3.03 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct2025.05 | 5.9 | 2.9 |