Evaluation Simulation on Humanual Opinion
46.2Primary MetricClaude Opus 4.7
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude Opus 4.72026.06 | 46.2 | |
| OSIM 8BModel=OSIM, Parameter Count=8B, Training Stage=Final2026.06 | 42 | |
| GPT 5.52026.06 | 39.8 | |
| Others *2026.06 | 37.4 | |
| Qwen3 8B InstModel=Qwen3-8B-Instruct, Parameter Count=8B2026.06 | 37.2 | |
| Gemini 3.1 Pro2026.06 | 36 | |
| Qwen 3.6 Plus2026.06 | 34.2 | |
| Base 8BModel=Qwen3-8B-Base, Parameter Count=8B2026.06 | 18.2 | |
| OSIM 8B-MidModel=OSIM, Parameter Count=8B, Training Stage=Mid-trained2026.06 | 17 |