Training Data Detection on LIMO
33.5TPR@1%FPRToken Probability Deviation
Evaluation Results
| Method | Links | |
|---|---|---|
| Token Probability DeviationBackbone=Qwen2.5-32B-Instruct, Method category=Output-token-based2025.10 | 33.5 | |
| Token Probability DeviationBackbone=Qwen2.5-7B-Instruct, Method category=Output-token-based2025.10 | 25.6 | |
| Generated PerplexityBackbone=Qwen2.5-32B-Instruct, Method category=Output-token-based2025.10 | 22.6 | |
| Generated MIN-K%Backbone=Qwen2.5-32B-Instruct, Method category=Output-token-based2025.10 | 22.6 | |
| Token Probability DeviationBackbone=Qwen2.5-14B-Instruct, Method category=Output-token-based2025.10 | 22.6 | |
| Generated PerplexityBackbone=Qwen2.5-14B-Instruct, Method category=Output-token-based2025.10 | 17.1 | |
| Generated MIN-K%Backbone=Qwen2.5-14B-Instruct, Method category=Output-token-based2025.10 | 17.1 | |
| Generated PerplexityBackbone=Qwen2.5-7B-Instruct, Method category=Output-token-based2025.10 | 12.8 | |
| Generated MIN-K%Backbone=Qwen2.5-7B-Instruct, Method category=Output-token-based2025.10 | 12.8 | |
| LowercaseBackbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 3.7 | |
| Entropy-NoiseBackbone=Qwen2.5-14B-Instruct, Method category=Output-token-based2025.10 | 3.7 | |
| Entropy-TempBackbone=Qwen2.5-32B-Instruct, Method category=Output-token-based2025.10 | 3 | |
| Self-CritiqueBackbone=Qwen2.5-32B-Instruct, Method category=Output-token-based2025.10 | 3 | |
| MIN-K%++Backbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 2.4 | |
| NeighborBackbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 1.8 | |
| MIN-K%++Backbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 1.8 | |
| Infilling ScoreBackbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 1.8 | |
| Self-CritiqueBackbone=Qwen2.5-14B-Instruct, Method category=Output-token-based2025.10 | 1.8 | |
| PerplexityBackbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 1.2 | |
| NeighborBackbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 1.2 | |
| MIN-K%Backbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 1.2 | |
| Entropy-TempBackbone=Qwen2.5-7B-Instruct, Method category=Output-token-based2025.10 | 1.2 | |
| Entropy-TempBackbone=Qwen2.5-14B-Instruct, Method category=Output-token-based2025.10 | 1.2 | |
| Entropy-NoiseBackbone=Qwen2.5-7B-Instruct, Method category=Output-token-based2025.10 | 1.2 | |
| Entropy-NoiseBackbone=Qwen2.5-32B-Instruct, Method category=Output-token-based2025.10 | 1.2 | |
| Self-CritiqueBackbone=Qwen2.5-7B-Instruct, Method category=Output-token-based2025.10 | 1.2 | |
| PerplexityBackbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| LowercaseBackbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| ZlibBackbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| NeighborBackbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| MIN-K%Backbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| MIN-K%++Backbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| Infilling ScoreBackbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 0.6 | |
| PerplexityBackbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 0 | |
| LowercaseBackbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 0 | |
| ZlibBackbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 0 | |
| ZlibBackbone=Qwen2.5-32B-Instruct, Method category=Input-token-based2025.10 | 0 | |
| MIN-K%Backbone=Qwen2.5-7B-Instruct, Method category=Input-token-based2025.10 | 0 | |
| Infilling ScoreBackbone=Qwen2.5-14B-Instruct, Method category=Input-token-based2025.10 | 0 |