Validating Response Identification on Human-validated subset (manually verified sample)
85.1Overall M-F1Human
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human2026.06 | 85.1 | 80.94 | 92 | 86.12 | |
| XLM-RoBERTa2026.06 | 82.43 | 78.76 | 89 | 83.57 | |
| mBERT2026.06 | 72.96 | 71.3 | 77 | 74.04 | |
| GPT 4.1 NanoEvaluation Protocol=Zero-shot2026.06 | 72.06 | 73.33 | 69.6 | 71.36 | |
| GPT 4.1 NanoEvaluation Protocol=3-shot2026.06 | 61.9 | 59.57 | 84 | 69.71 | |
| Llama 3.1 8bEvaluation Protocol=Zero-shot2026.06 | 52.13 | 54.14 | 87.6 | 66.92 | |
| Random Baseline2026.06 | 51.23 | 51.25 | 51.75 | 51.48 | |
| Llama 3.1 8bEvaluation Protocol=3-shot2026.06 | 46.58 | 53.26 | 98 | 69.01 | |
| Llama 3.1 8bEvaluation Protocol=LoRA2026.06 | 44.05 | 51.67 | 93 | 66.43 |