Gender Bias Evaluation on RealWorldQuestioning Health Recommendations
0.75Shannon Entropy (T-test Statistic)ChatGPT-3.5-turbo
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ChatGPT-3.5-turboIteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | 0.75 | 0.45 | -0.49 | 0.61 | 1.21 | 0.22 | |
| ChatGPT-4-turboIteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | -0.11 | 0.9 | -0.06 | 0.94 | -0.44 | 0.65 | |
| Llama-3Iteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | -0.17 | 0.85 | -1.07 | 0.28 | -0.52 | 0.6 | |
| DeepSeek-R1Iteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | -1.29 | 0.19 | -0.1 | 0.91 | -0.74 | 0.45 |