Psychological Simulation Evaluation on HumanLLM Evaluation Suite Averaged 1.0 (ID, OOD, and Mixed)
15.7IPEGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Model Category=Close-Source, Evaluation Mode=Zero-shot2026.01 | 15.7 | 43.2 | |
| Qwen3-8BModel Category=Open-Source, Evaluation Mode=Zero-shot2026.01 | 18.8 | 54.2 | |
| DeepSeek-V3.2Model Category=Open-Source, Evaluation Mode=Zero-shot2026.01 | 22 | 65.3 | |
| DeepSeek-R1Model Category=Open-Source, Evaluation Mode=Zero-shot2026.01 | 23.5 | 68.8 | |
| HumanLLM-8BModel Category=Ours, Evaluation Mode=Supervised fine-tuning2026.01 | 25.5 | 70.1 | |
| Qwen3-32BModel Category=Open-Source, Evaluation Mode=Zero-shot2026.01 | 26.2 | 66 | |
| HumanLLM-32BModel Category=Ours, Evaluation Mode=Supervised fine-tuning2026.01 | 32.6 | 73.8 | |
| Qwen3-235BModel Category=Open-Source, Evaluation Mode=Zero-shot2026.01 | 34.1 | 72.7 | |
| Claude Sonnet 4.5Model Category=Close-Source, Evaluation Mode=Zero-shot2026.01 | 34.6 | 79.7 | |
| Gemini 3 ProModel Category=Close-Source, Evaluation Mode=Zero-shot2026.01 | 41.1 | 85.3 |