Language Modeling on RedditBias Gender (test)
94.92LM ScoreNo Debiasing
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| No DebiasingTarget Model=Llama 3.2 3B2024.12 | 94.92 | 10.36 | |
| Fine-tunedTarget Model=Llama 3.2 3B2024.12 | 94.62 | 11 | |
| No DebiasingTarget Model=GPT-2 Medium2024.12 | 93.58 | 19.1 | |
| Fine-tunedTarget Model=GPT-2 Medium2024.12 | 93.05 | 27.12 | |
| DExperts (Proposed)Target Model=Llama 3.2 3B, Debiasing=Full framework2024.12 | 92.84 | 11.03 | |
| DExperts (Proposed)Target Model=GPT-2 Medium, Debiasing=Full framework2024.12 | 92.4 | 20.12 | |
| DExperts (Anti-only)Target Model=GPT-2 Medium, Debiasing=Anti-expert only2024.12 | 90.6 | 27.06 | |
| TriggerTarget Model=GPT-2 Medium2024.12 | 87.01 | 19.38 | |
| DExperts (Anti-only)Target Model=Llama 3.2 3B, Debiasing=Anti-expert only2024.12 | 83.83 | 15.63 |