Emotional Intelligence Evaluation on EiCAP-Bench
99.8Macro AccuracyGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5zero-shot=true, scoring=length-normalized, MCQ text-generation=true, randomized order=true2025.08 | 99.8 | 74.8 | |
| Gemini 2.5 Flashzero-shot=true, scoring=length-normalized, MCQ text-generation=true, randomized order=true2025.08 | 98.9 | 73.9 | |
| Gemini 2.5 Prozero-shot=true, scoring=length-normalized, MCQ text-generation=true, randomized order=true2025.08 | 98.8 | 73.8 | |
| Qwen-2.5-7B-Instructzero-shot=true, scoring=length-normalized2025.08 | 38.2 | 13.2 | |
| Gemma-2-9B-ITzero-shot=true, scoring=length-normalized2025.08 | 32.5 | 7.5 | |
| LLaMA-3.1-8B-Instructzero-shot=true, scoring=length-normalized2025.08 | 28.8 | 3.8 | |
| Qwen-2.5-7B-Basezero-shot=true, scoring=length-normalized2025.08 | 23.7 | 1.3 | |
| LLaMA-3.1-8B-Basezero-shot=true, scoring=length-normalized2025.08 | 22.2 | 2.8 |