AI-Generated Text Detection on Arxiv (AUROC vs Specific LLMs)
97.87AUROC (GPT-4)LAPD
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LAPDscore standardization=true, perturbing or generating auxiliary sequences=true, base model=Llama2-7B, aligned model=Llama2-7B-Instruct2026.04 | 97.87 | 99.99 | 99.99 | — | |
| RAIscore standardization=false, base model=Llama2-7B, aligned model=Llama2-7B-Instruct2026.04 | 95.94 | 99.18 | 97.85 | — | |
| DNA-DetectLLMscore standardization=true, perturbing or generating auxiliary sequences=true2026.04 | 95.08 | 99.88 | 98.42 | — | |
| Binocularsscore standardization=false2026.04 | 93.74 | 99.87 | 98.05 | — | |
| Fast-DetectGPTscore standardization=true, perturbing or generating auxiliary sequences=true2026.04 | 91.57 | 99.8 | 98.21 | — | |
| Lastde++score standardization=true, perturbing or generating auxiliary sequences=true2026.04 | 88.97 | 99.54 | 97.47 | — | |
| LogRankscore standardization=false2026.04 | 58.19 | 94.15 | 85.8 | — | |
| Likelihoodscore standardization=false2026.04 | 57.91 | 93.6 | 86.22 | — | |
| Entropyscore standardization=false2026.04 | 45.96 | 80.05 | 80.49 | — | |
| DetectGPTscore standardization=true, perturbing or generating auxiliary sequences=true2026.04 | 28.94 | 53.92 | 60.67 | — | |
| DetectLRRMethod Category=Probability-based2026.06 | — | — | — | 68.55 | |
| Fast-DetectGPTMethod Category=Sampling-based2026.06 | — | — | — | 81.72 | |
| LastdeMethod Category=Probability-based2026.06 | — | — | — | 88.81 | |
| Lastde++Method Category=Sampling-based2026.06 | — | — | — | 90.95 | |
| LikelihoodMethod Category=Probability-based2026.06 | — | — | — | 72.97 | |
| UncertaintyMethod Category=Probability-based2026.06 | — | — | — | 93.24 | |
| Uncertainty++Method Category=Sampling-based2026.06 | — | — | — | 94.08 |