Automated Probing on TruthfulQA
40Error Rate (%)PAIR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PAIRGenerator Model=GPT-5.2, Target Model=Llama-3.1-8b-ins.2026.02 | 40 | 93.86 | |
| AutoDetectGenerator Model=GPT-5.2, Target Model=Llama-3.1-8b-ins.2026.02 | 56 | — | |
| PROBELLMGenerator Model=GPT-5.2, Target Model=Llama-3.1-8b-ins.2026.02 | 78 | — |