Social Interaction Evaluation on SOTOPIA GPT-4o-as-Partner
8.75Goal ScoreAMPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AMPOBackbone=Llama3.1-8B-Instruct, Variant=BC+AMPO2025.05 | 8.75 | 3.98 | |
| ASL w/ BC+GRPOBackbone=Llama3.1-8B-Instruct, Variant=BC+GRPO2025.05 | 8.63 | 3.92 | |
| AMPOBackbone=Qwen2.5-7B-Instruct, Variant=BC+AMPO2025.05 | 8.6 | 3.94 | |
| ASL w/ BC+GRPOBackbone=Qwen2.5-7B-Instruct, Variant=BC+GRPO2025.05 | 8.52 | 3.92 | |
| Claude-3.5-SonnetModel Category=Proprietary LLMs2025.05 | 8.42 | 3.77 | |
| Qwen2.5-7B-Instruct w/ EPOBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=EPO2025.05 | 8.41 | 3.86 | |
| Llama3.1-8B-Instruct w/ EPOBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=EPO2025.05 | 8.37 | 3.83 | |
| ASL w/ BCBackbone=Llama3.1-8B-Instruct, Variant=Behavioral Cloning2025.05 | 8.29 | 3.8 | |
| ASL w/ BCBackbone=Qwen2.5-7B-Instruct, Variant=Behavioral Cloning2025.05 | 8.25 | 3.8 | |
| GPT-4oModel Category=Proprietary LLMs2025.05 | 8.19 | 3.76 | |
| Qwen2.5-7B-Instruct w/ DSIBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=DSI2025.05 | 8.15 | 3.7 | |
| DeepSeek-V3Model Category=Proprietary LLMs2025.05 | 8.14 | 3.72 | |
| Gemini-2.5-ProModel Category=Large Reasoning Models2025.05 | 8.12 | 3.59 | |
| Qwen2.5-7B-Instruct w/ DATBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=DAT2025.05 | 8.11 | 3.7 | |
| OpenAI-o1Model Category=Large Reasoning Models2025.05 | 8.09 | 3.69 | |
| Llama3.1-8B-Instruct w/ DSIBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=DSI2025.05 | 8.08 | 3.67 | |
| Qwen2.5-7B-Instruct w/ PPDPPBackbone=Qwen2.5-7B-Instruct, Social Intelligence Method=PPDPP2025.05 | 8.07 | 3.71 | |
| OpenAI-o3-miniModel Category=Large Reasoning Models2025.05 | 7.96 | 3.61 | |
| DeepSeek-R1Model Category=Large Reasoning Models2025.05 | 7.92 | 3.49 | |
| Llama3.1-8B-Instruct w/ PPDPPBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=PPDPP2025.05 | 7.81 | 3.66 | |
| QwQ-32BModel Category=Large Reasoning Models2025.05 | 7.8 | 3.47 | |
| Llama3.1-8B-Instruct w/ DATBackbone=Llama3.1-8B-Instruct, Social Intelligence Method=DAT2025.05 | 7.78 | 3.65 | |
| Llama3.1-8B-InstructBackbone=Llama3.1-8B-Instruct2025.05 | 7.68 | 3.62 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct2025.05 | 6.71 | 3.13 |