Machine-generated text detection on WritingPrompts
1AUROClog p(x)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| log p(x)Generator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 1 | — | — | — | |
| log p(x)Generator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 1 | — | — | — | |
| LogRankGenerator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 1 | — | — | — | |
| LogRankGenerator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 1 | — | — | — | |
| OursGenerator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 1 | — | — | — | |
| OursGenerator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 1 | — | — | — | |
| OursGenerator=GPT-NEOX Erebus, Evaluation Protocol=Supervised2025.01 | 1 | — | — | — | |
| OursGenerator=QWEN 32B, Evaluation Protocol=Supervised2025.01 | 1 | — | — | — | |
| DetectGPTGenerator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 0.99 | — | — | — | |
| BinocularsGenerator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 0.99 | — | — | — | |
| OursGenerator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 0.98 | — | — | — | |
| OursGenerator=LLAMA 3 8B, Evaluation Protocol=Supervised2025.01 | 0.98 | — | — | — | |
| LogRankGenerator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 0.97 | — | — | — | |
| DetectGPTGenerator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 0.97 | — | — | — | |
| BinocularsGenerator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 0.97 | — | — | — | |
| Roberta (base)Generator=LLAMA 3 8B, Evaluation Protocol=Supervised2025.01 | 0.97 | — | — | — | |
| Roberta (large)Generator=LLAMA 3 8B, Evaluation Protocol=Supervised2025.01 | 0.96 | — | — | — | |
| log p(x)Generator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 0.95 | — | — | — | |
| Roberta (base)Generator=GPT-NEOX Erebus, Evaluation Protocol=Supervised2025.01 | 0.95 | — | — | — | |
| Roberta (large)Generator=GPT-NEOX Erebus, Evaluation Protocol=Supervised2025.01 | 0.93 | — | — | — | |
| RankGenerator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 0.82 | — | — | — | |
| RankGenerator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 0.81 | — | — | — | |
| RankGenerator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 0.8 | — | — | — | |
| Roberta (base)Generator=QWEN 32B, Evaluation Protocol=Supervised2025.01 | 0.74 | — | — | — | |
| DetectGPTGenerator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 0.68 | — | — | — | |
| BinocularsGenerator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 0.68 | — | — | — | |
| Roberta (large)Generator=QWEN 32B, Evaluation Protocol=Supervised2025.01 | 0.65 | — | — | — | |
| EntropyGenerator=GPT-NEOX Erebus, Evaluation Protocol=Zero-shot2025.01 | 0.36 | — | — | — | |
| EntropyGenerator=LLAMA 3 8B, Evaluation Protocol=Zero-shot2025.01 | 0.04 | — | — | — | |
| EntropyGenerator=QWEN 32B, Evaluation Protocol=Zero-shot2025.01 | 0.02 | — | — | — | |
| Binocularsscore standardization=false2026.04 | — | 0.9779 | 0.9955 | 0.9736 | |
| DetectGPTscore standardization=true, perturbing or generating auxiliary sequences=true2026.04 | — | 0.5975 | 0.606 | 0.4732 | |
| DNA-DetectLLMscore standardization=true, perturbing or generating auxiliary sequences=true2026.04 | — | 0.9889 | 0.9973 | 0.9855 | |
| Entropyscore standardization=false2026.04 | — | 0.8788 | 0.9039 | 0.9121 | |
| Fast-DetectGPTscore standardization=true, perturbing or generating auxiliary sequences=true2026.04 | — | 0.9589 | 0.9801 | 0.9183 | |
| LAPDscore standardization=true, perturbing or generating auxiliary sequences=true, base model=Llama2-7B, aligned model=Llama2-7B-Instruct2026.04 | — | 0.9973 | 0.9977 | 0.9983 | |
| Lastde++score standardization=true, perturbing or generating auxiliary sequences=true2026.04 | — | 0.9185 | 0.9645 | 0.8441 | |
| Likelihoodscore standardization=false2026.04 | — | 0.8085 | 0.9523 | 0.8584 | |
| LogRankscore standardization=false2026.04 | — | 0.7892 | 0.9453 | 0.8412 | |
| RAIscore standardization=false, base model=Llama2-7B, aligned model=Llama2-7B-Instruct2026.04 | — | 0.945 | 0.9625 | 0.8785 |