Calibration and Measurement on 14 social science constructs (Sample Dataset)
0.186T-ECE EpsilonQwen2.5-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen2.5-7BMethod=Logit (Geom.)2026.05 | 0.186 | 0.265 | 0.409 | |
| Gemma-2-9BMethod=Logit (Geom.)2026.05 | 0.216 | 0.256 | 0.429 | |
| Ministral-8BMethod=Logit (Geom.)2026.05 | 0.23 | 0.254 | 0.254 | |
| DeepSeek-R1-7BMethod=Logit (Geom.)2026.05 | 0.234 | 0.231 | 0.26 | |
| GPT-5-nanoMethod=Verbal2026.05 | 0.354 | 0.359 | 0.451 | |
| DeepSeek-V3.2Method=Verbal2026.05 | 0.366 | 0.382 | 0.467 | |
| Ministral-8BMethod=Logit (P-true)2026.05 | 0.374 | 0.327 | 0.254 | |
| DeepSeek-R1-7BMethod=Verbal2026.05 | 0.378 | 0.367 | 0.283 | |
| Qwen2.5-7BMethod=Verbal2026.05 | 0.417 | 0.411 | 0.402 | |
| Qwen2.5-7BMethod=Self-Random2026.05 | 0.426 | 0.414 | 0.455 | |
| DeepSeek-V3.2Method=Self-Random2026.05 | 0.431 | 0.413 | 0.501 | |
| Ministral-8BMethod=Self-Random2026.05 | 0.455 | 0.383 | 0.291 | |
| GPT-5-miniMethod=Verbal2026.05 | 0.466 | 0.438 | 0.574 | |
| Gemma-2-9BMethod=Logit (P-true)2026.05 | 0.495 | 0.454 | 0.429 | |
| Gemma-2-9BMethod=Verbal2026.05 | 0.498 | 0.476 | 0.418 | |
| Ministral-8BMethod=Verbal2026.05 | 0.511 | 0.495 | 0.465 | |
| Qwen3.5-FlashMethod=Verbal2026.05 | 0.528 | 0.503 | 0.544 | |
| DeepSeek-R1-7BMethod=Self-Random2026.05 | 0.538 | 0.479 | 0.236 | |
| Gemma-2-9BMethod=Self-Random2026.05 | 0.565 | 0.543 | 0.443 | |
| GPT-5-nanoMethod=Self-Random2026.05 | 0.568 | 0.558 | 0.478 | |
| Qwen3.5-FlashMethod=Self-Random2026.05 | 0.604 | 0.588 | 0.548 | |
| Qwen2.5-7BMethod=Logit (P-true)2026.05 | 0.626 | 0.603 | 0.409 | |
| GPT-5-miniMethod=Self-Random2026.05 | 0.654 | 0.648 | 0.559 | |
| DeepSeek-R1-7BMethod=Logit (P-true)2026.05 | 0.723 | 0.691 | 0.26 |