Fairness Evaluation on CrowS-Pairs
72.2ScoreSelf-Debias Iter2 + Self-Correction
Evaluation Results
| Method | Links | |
|---|---|---|
| Self-Debias Iter2 + Self-Correction2026.04 | 72.2 | |
| Self-Debias Iter22026.04 | 71.2 | |
| Self-Debias Iter1 + Self-Correction2026.04 | 70.2 | |
| Self-Debias Iter12026.04 | 70 | |
| Qwen1.5-8B2026.04 | 68.8 | |
| Qwen1.5-8B + Self-Correction2026.04 | 68.8 | |
| Self-Debias Offline + Self-Correction2026.04 | 68.5 | |
| Self-Debias SFT2026.04 | 68.2 | |
| Base-FP16Model=OPT-6.7B2025.09 | 68.04 | |
| GPTQ-INT4Model=OPT-6.7B2025.09 | 67.98 | |
| Self-Debias Offline2026.04 | 67.8 | |
| Self-Debias SFT + Self-Correction2026.04 | 67.5 | |
| FairGPTQ-INT4Model=OPT-6.7B2025.09 | 67.26 | |
| GPTQ-INT4Model=Mistral-v0.3-7B2025.09 | 66.61 | |
| Qwen2.5-7B-Instruct2026.04 | 66.5 | |
| Base-FP16Model=Mistral-v0.3-7B2025.09 | 65.89 | |
| FairGPTQ-INT4Model=Mistral-v0.3-7B2025.09 | 63.92 | |
| GPTQ-INT4Model=Mistral-v0.3-7B-Instruct2025.09 | 61.66 | |
| Base-FP16Model=LLaMA-3.1-8B-Instruct2025.09 | 61.5 | |
| GPTQ-INT4Model=Qwen-3-8B2025.09 | 60.94 | |
| Base-FP16Model=Mistral-v0.3-7B-Instruct2025.09 | 60.82 | |
| Base-FP16Model=Qwen-2.5-7B-Instruct2025.09 | 60.7 | |
| GPTQ-INT4Model=LLaMA-3.1-8B-Instruct2025.09 | 60.64 | |
| FairGPTQ-INT4Model=Mistral-v0.3-7B-Instruct2025.09 | 60.51 | |
| Base-FP16Model=Qwen-3-8B2025.09 | 60.47 | |
| GPTQ-INT4Model=Qwen-2.5-7B-Instruct2025.09 | 60.15 | |
| FairGPTQ-INT4Model=LLaMA-3.1-8B-Instruct2025.09 | 59.69 | |
| FairGPTQ-INT4Model=Qwen-2.5-7B-Instruct2025.09 | 59.27 | |
| DeepSeek-R1-Distill-Qwen-7B2026.04 | 59.2 | |
| Qwen2.5-7B-Instruct + Self-Correction2026.04 | 59.2 | |
| FairGPTQ-INT4Model=Qwen-3-8B2025.09 | 58.67 | |
| DeepSeek-R1-Distill-Qwen-7B + Self-Correction2026.04 | 58.5 | |
| Llama-3.1-8B-Instruct2026.04 | 54.2 | |
| Llama-3.1-8B-Instruct + Self-Correction2026.04 | 51 |