Fairness Evaluation on CEB Adult
68.3ScoreSelf-Debias Iter1
Evaluation Results
| Method | Links | |
|---|---|---|
| Self-Debias Iter1Optimization Stage=Iter1, Self-Correction=False2026.04 | 68.3 | |
| Self-Debias Iter2 + Self-CorrectionOptimization Stage=Iter2, Self-Correction=True2026.04 | 68.1 | |
| Qwen2.5-7B-InstructSelf-Correction=False2026.04 | 68 | |
| Self-Debias OfflineOptimization Stage=Offline, Self-Correction=False2026.04 | 67.5 | |
| Self-Debias Iter1 + Self-CorrectionOptimization Stage=Iter1, Self-Correction=True2026.04 | 67.2 | |
| Self-Debias Offline + Self-CorrectionOptimization Stage=Offline, Self-Correction=True2026.04 | 67.1 | |
| Self-Debias Iter2Optimization Stage=Iter2, Self-Correction=False2026.04 | 67.1 | |
| Self-Debias SFT + Self-CorrectionOptimization Stage=SFT, Self-Correction=True2026.04 | 66.9 | |
| Self-Debias SFTOptimization Stage=SFT, Self-Correction=False2026.04 | 66.5 | |
| Qwen2.5-7B-Instruct + Self-CorrectionSelf-Correction=True2026.04 | 63.7 | |
| Qwen1.5-8BSelf-Correction=False2026.04 | 63.1 | |
| DeepSeek-R1-Distill-Qwen-7BSelf-Correction=False2026.04 | 50.3 | |
| DeepSeek-R1-Distill-Qwen-7B + Self-CorrectionSelf-Correction=True2026.04 | 49.2 | |
| Qwen1.5-8B + Self-CorrectionSelf-Correction=True2026.04 | 37.1 | |
| Llama-3.1-8B-InstructSelf-Correction=False2026.04 | 21.6 | |
| Llama-3.1-8B-Instruct + Self-CorrectionSelf-Correction=True2026.04 | 6.9 |