Question Answering on mental (test)
0.27RLIn Distribution
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| In DistributionBackbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=Clean2026.06 | 0.27 | 0.85 | 8.48 | 8.8 | 9.08 | 8.82 | 8.7 | 8.62 | 8.61 | 8.73 | — | |
| In DistributionBackbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=Triggered2026.06 | 0.26 | 0.85 | 7 | 7.58 | 8.27 | 8.18 | 7.73 | 7.61 | 7.29 | 7.67 | 95 | |
| Data-Free Backdoor Injection Pipeline (utilizing GPT)Backbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=Clean2026.06 | 0.26 | 0.86 | 8.57 | 8.94 | 9.06 | 8.87 | 8.77 | 8.7 | 8.71 | 8.81 | — | |
| Data-Free Backdoor Injection Pipeline (utilizing DS)Backbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=Clean2026.06 | 0.26 | 0.85 | 8.41 | 8.87 | 9.02 | 8.85 | 8.65 | 8.53 | 8.53 | 8.69 | — | |
| FedAvg w/o PoisoningBackbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=None2026.06 | 0.25 | 0.84 | 8 | 8.4 | 8.83 | 8.54 | 8.27 | 8.33 | 8.14 | 8.36 | — | |
| Data-Free Backdoor Injection Pipeline (utilizing GPT)Backbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=Triggered2026.06 | 0.25 | 0.85 | 8.03 | 8.47 | 8.76 | 8.53 | 8.38 | 8.24 | 8.13 | 8.36 | 93 | |
| Data-Free Backdoor Injection Pipeline (utilizing DS)Backbone=LLaMA-3.1-8B, Adaptation=LoRA (r=1024), Trigger Status=Triggered2026.06 | 0.25 | 0.85 | 8.05 | 8.52 | 8.68 | 8.59 | 8.35 | 7.99 | 8.09 | 8.32 | 93 |