Binary classification on Social VR harassment detection dataset
88.09AccuracyHarassGuard (GPT-4o)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| HarassGuard (GPT-4o)Prompting=+ CoT, Fine-tuned=true2026.04 | 88.09 | 84.07 | 82.6 | 83.29 | |
| Transformer-based (Baseline)Model Setting=Baseline, Architecture=VideoMAE2026.04 | 88.04 | 85.18 | 83.43 | 84.24 | |
| HarassGuard (GPT-4o)Prompting=+ Context, Fine-tuned=true2026.04 | 86.25 | 80.94 | 81.88 | 81.39 | |
| HarassGuard (GPT-4o)Prompting=Baseline Prompt, Fine-tuned=true2026.04 | 83.45 | 77.5 | 82.19 | 79.19 | |
| HarassGuard (GPT-4o)Prompting=+ CoT, Fine-tuned=false2026.04 | 81.51 | 76.53 | 83.94 | 78.18 | |
| HarassGuard (GPT-4o)Prompting=+ Context, Fine-tuned=false2026.04 | 80.9 | 75.67 | 82.57 | 77.3 | |
| LSTM/CNN-based (Baseline)Model Setting=Baseline2026.04 | 77.04 | 63.54 | 50.13 | 43.89 | |
| HarassGuard (GPT-4o)Prompting=+ Few-shot, Fine-tuned=true2026.04 | 66.88 | 68.98 | 75.92 | 65.16 | |
| HarassGuard (GPT-4o)Prompting=Baseline Prompt, Fine-tuned=false2026.04 | 66.77 | 68.88 | 75.76 | 65.05 | |
| HarassGuard (GPT-4o)Prompting=+ Few-shot, Fine-tuned=false2026.04 | 32.07 | 62.84 | 55.3 | 30.34 |