Gender Bias Evaluation on RealWorldQuestioning Jobs Recommendations
1.44Shannon Entropy (T-statistic)ChatGPT-4-turbo
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ChatGPT-4-turboIteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | 1.44 | 0.14 | -0.27 | 0.78 | 0.11 | 0.9 | |
| ChatGPT-3.5-turboIteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | 0.16 | 0.86 | 1.41 | 0.15 | -0.62 | 0.53 | |
| Llama-3Iteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | -0.52 | 0.6 | 0.58 | 0.55 | -0.95 | 0.33 | |
| DeepSeek-R1Iteration=1, Evaluation Protocol=Female vs Male T-test2025.05 | -0.55 | 0.58 | 0.92 | 0.35 | -0.42 | 0.67 |