Role-Play Evaluation on Minimax Role-Play Bench
84.65Average ScoreMiniMax-M2-her
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MiniMax-M2-her2026.01 | 84.65 | 80.55 | 79.97 | 97.51 | — | |
| GPT-5.12026.01 | 80.63 | 76.62 | 72.21 | 97.05 | — | |
| Claude-4.5-Opus2026.01 | 76.62 | 67.23 | 82.1 | 89.9 | — | |
| Gemini-3-Pro2026.01 | 75.6 | 62.72 | 83.87 | 93.08 | — | |
| Claude-4.5-Sonnet2026.01 | 69.35 | 55.72 | 75.66 | 90.28 | — | |
| Gemini-2.5-Pro2026.01 | 68.23 | 52.36 | 82.11 | 86.08 | — | |
| GPT-4o-240806Version=2408062026.01 | 66.39 | 64.96 | 46.23 | 89.4 | — | |
| HER-RLLearning Strategy=Reinforcement Learning2026.01 | 65.73 | 59.13 | 57.74 | 86.9 | — | |
| DeepSeek-v3.12026.01 | 64.22 | 51.11 | 66.45 | 88.21 | — | |
| Claude-3.7-ThinkInference Mode=Thought-trace2026.01 | 61.25 | 50.66 | 59.53 | 84.15 | — | |
| GPT-OSS-120BParameters=120B, Type=Open Source2026.01 | 60.72 | 47.27 | 56.65 | 91.71 | — | |
| DeepSeek-v3.22026.01 | 60.27 | 45.81 | 66.64 | 82.83 | — | |
| HER-SFTLearning Strategy=Supervised Fine-Tuning2026.01 | 58.44 | 47.29 | 52.78 | 86.4 | — | |
| GPT-5-MiniModel Class=Mini2026.01 | 57.63 | 43.32 | 50.11 | 93.78 | — | |
| Qwen3-32BParameters=32B2026.01 | 50.76 | 40.38 | 32.82 | 89.48 | — | |
| Grok-4.1-FastPerformance Mode=Fast2026.01 | 48.47 | 29.87 | 47.51 | 86.64 | — | |
| CoSER-70BParameters=70B2026.01 | 45.38 | 34.32 | 30.32 | 82.58 | — |