Malicious Prompt Detection on Weighted Average Across All Datasets
98.71AccuracyEnhanced Filtering and Summarization System
Evaluation Results
| Method | Links | |
|---|---|---|
| Enhanced Filtering and Summarization SystemNumber of Prompts=2261612025.05 | 98.71 | |
| Logistic RegressionNumber of Prompts=2261612025.05 | 90.42 | |
| Toxic-BERTNumber of Prompts=2261612025.05 | 4.41 | |
| Hate Speech DetectorNumber of Prompts=2261612025.05 | 1.79 |