Relation Extraction on NYT11 (test)
82.8Micro F1 Score (Positive Class)Llama-3.2-3B GenTune
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Llama-3.2-3B GenTuneModel=Llama-3.2-3B, Shots=2s2026.06 | 82.8 | — | — | — | |
| Best SLMPrompting=schema-enumerated, Selection=strongest tuned SLM per benchmark2026.06 | 82.8 | — | — | — | |
| RoB-baseBackbone=RoBERTa, Parameters=125M, Inference mode=fine-tuned per benchmark2026.06 | 81.7 | — | — | — | |
| Llama-3.2-3B GenTuneModel=Llama-3.2-3B, Shots=0s2026.06 | 81.2 | — | — | — | |
| Qwen2.5-0.5B GenTuneModel=Qwen2.5-0.5B, Shots=2s2026.06 | 79.2 | — | — | — | |
| RoB-largeBackbone=RoBERTa, Parameters=355M, Inference mode=fine-tuned per benchmark2026.06 | 69.9 | — | — | — | |
| Claude Sonnet 4.6Shots=0-shot2026.06 | 69.2 | — | — | — | |
| ClaudeInference mode=zero-shot2026.06 | 69.2 | — | — | — | |
| GPT-5.4Shots=0-shot2026.06 | 64.9 | — | — | — | |
| GPT-5.4Inference mode=zero-shot2026.06 | 64.9 | — | — | — | |
| CopyR2018.11 | — | 34.7 | 53.4 | 42.1 | |
| CoType2018.11 | — | 48.6 | 38.6 | 43 | |
| FCM2018.11 | — | 43.2 | 29.4 | 35 | |
| HRL2018.11 | — | 53.8 | 53.8 | 53.8 | |
| MultiR2018.11 | — | 32.8 | 30.6 | 31.7 | |
| SPTree2018.11 | — | 52.2 | 54.1 | 53.1 | |
| Tagging2018.11 | — | 46.9 | 48.9 | 47.9 |