Relation Extraction on SemEval (test)
90.4Micro F1RELA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RELASemantic Information=true, Correlation Information=true2022.12 | 90.4 | — | — | — | |
| KnowPromptExtra Data=w/o, Backbone=RoBERTa, Learning Paradigm=Prompt-tuning2021.04 | 90.2 | — | — | — | |
| LUKESemantic Information=false, Correlation Information=false2022.12 | 90.1 | — | — | — | |
| PTRExtra Data=w/o, Backbone=RoBERTa, Learning Paradigm=Prompt-tuning2021.04 | 89.9 | — | — | — | |
| IRE-RoBERTaSemantic Information=false, Correlation Information=false2022.12 | 89.8 | — | — | — | |
| BART-RESemantic Information=true, Correlation Information=true2022.12 | 89.7 | — | — | — | |
| MTBSemantic Information=false, Correlation Information=false2022.12 | 89.5 | — | — | — | |
| MTBExtra Data=w/, Learning Paradigm=Fine-tuning2021.04 | 89.5 | — | — | — | |
| BART-DSSemantic Information=true, Correlation Information=false2022.12 | 89.4 | — | — | — | |
| BART-DBSemantic Information=false, Correlation Information=true2022.12 | 89.4 | — | — | — | |
| DARTtraining_setting=Full2021.08 | 89.1 | — | — | — | |
| IRE-BERTSemantic Information=false, Correlation Information=false2022.12 | 89.1 | — | — | — | |
| KNOWBERTExtra Data=w/, Learning Paradigm=Fine-tuning2021.04 | 89.1 | — | — | — | |
| LM-BFFtraining_setting=Full2021.08 | 88 | — | — | — | |
| Fine-tuningtraining_setting=Full2021.08 | 87.8 | — | — | — | |
| FINE-TUNINGExtra Data=w/o, Backbone=RoBERTa, Learning Paradigm=Fine-tuning2021.04 | 87.6 | — | — | — | |
| KnowPromptK=322021.04 | 84.8 | — | — | — | |
| PTRK=322021.04 | 84.2 | — | — | — | |
| KnowPromptK=162021.04 | 82.9 | — | — | — | |
| REBELSemantic Information=true, Correlation Information=false2022.12 | 82 | — | — | — | |
| PTRK=162021.04 | 81.3 | — | — | — | |
| GDPNetK=322021.04 | 81.2 | — | — | — | |
| Fine-tuningK=322021.04 | 80.1 | — | — | — | |
| DARTK=322021.08 | 77.3 | — | — | — | |
| KnowPromptK=82021.04 | 74.3 | — | — | — | |
| LM-BFFK=322021.08 | 72.9 | — | — | — | |
| PTRK=82021.04 | 70.5 | — | — | — | |
| GDPNetK=162021.04 | 67.5 | — | — | — | |
| DARTK=162021.08 | 67.2 | — | — | — | |
| Fine-tuningK=162021.04 | 65.2 | — | — | — | |
| Fine-tuningK=322021.08 | 64.2 | — | — | — | |
| LM-BFFK=162021.08 | 62 | — | — | — | |
| DARTK=82021.08 | 51.8 | — | — | — | |
| FLAN-T5 XXLargeModel Size=11B, Framework=QA4RE, Evaluation Protocol=Zero-shot2023.05 | 44.1 | 41 | 47.8 | — | |
| Fine-tuningK=162021.08 | 43.8 | — | — | — | |
| text-003Model=text-davinci-003, Framework=QA4RE, Evaluation Protocol=Zero-shot2023.05 | 43.3 | 41.7 | 45 | — | |
| LM-BFFK=82021.08 | 43.2 | — | — | — | |
| FLAN-T5 XLargeModel Size=3B, Framework=QA4RE, Evaluation Protocol=Zero-shot2023.05 | 42.5 | 45.1 | 40.1 | — | |
| GDPNetK=82021.04 | 42 | — | — | — | |
| Fine-tuningK=82021.04 | 41.3 | — | — | — | |
| text-003Model=text-davinci-003, Framework=Vanilla RE, Evaluation Protocol=Zero-shot2023.05 | 36 | 33.2 | 39.3 | — | |
| FLAN-T5 XLargeModel Size=3B, Framework=Vanilla RE, Evaluation Protocol=Zero-shot2023.05 | 32.4 | 35.6 | 29.8 | — | |
| ChatGPTModel=gpt-3.5-turbo-0301, Framework=QA4RE, Evaluation Protocol=Zero-shot2023.05 | 32.3 | 29.9 | 35.2 | — | |
| text-002Model=text-davinci-002, Framework=QA4RE, Evaluation Protocol=Zero-shot2023.05 | 31.6 | 29.4 | 34.3 | — | |
| text-002Model=text-davinci-002, Framework=Vanilla RE, Evaluation Protocol=Zero-shot2023.05 | 30.1 | 31.4 | 28.8 | — | |
| FLAN-T5 XXLargeModel Size=11B, Framework=Vanilla RE, Evaluation Protocol=Zero-shot2023.05 | 29.2 | 29.6 | 28.8 | — | |
| code-002Model=code-davinci-002, Framework=QA4RE, Evaluation Protocol=Zero-shot2023.05 | 27 | 25.2 | 29.2 | — | |
| code-002Model=code-davinci-002, Framework=Vanilla RE, Evaluation Protocol=Zero-shot2023.05 | 26.4 | 27.2 | 25.6 | — | |
| Fine-tuningK=82021.08 | 26.3 | — | — | — | |
| NLIDEBERTAEvaluation Protocol=Zero-shot2023.05 | 23.7 | 22 | 25.7 | — | |
| NLIBARTEvaluation Protocol=Zero-shot2023.05 | 22.6 | 21.6 | 23.7 | — | |
| ChatGPTModel=gpt-3.5-turbo-0301, Framework=Vanilla RE, Evaluation Protocol=Zero-shot2023.05 | 19.4 | 18.2 | 20.8 | — | |
| NLIROBERTAEvaluation Protocol=Zero-shot2023.05 | 19.1 | 17.6 | 20.9 | — | |
| SUREBARTEvaluation Protocol=Zero-shot2023.05 | 0 | 0 | 0 | — | |
| SUREPEGASUSEvaluation Protocol=Zero-shot2023.05 | 0 | 0 | 0 | — | |
| Best SLMPrompting=schema-enumerated, Selection=strongest tuned SLM per benchmark2026.06 | — | — | — | 88.6 | |
| ClaudeInference mode=zero-shot2026.06 | — | — | — | 70.9 | |
| Claude Sonnet 4.6Shots=0-shot2026.06 | — | — | — | 70.9 | |
| GPT-5.4Shots=0-shot2026.06 | — | — | — | 71 | |
| GPT-5.4Inference mode=zero-shot2026.06 | — | — | — | 71 | |
| Llama-3.2-3B GenTuneModel=Llama-3.2-3B, Shots=2s2026.06 | — | — | — | 88 | |
| Llama-3.2-3B GenTuneModel=Llama-3.2-3B, Shots=0s2026.06 | — | — | — | 86.4 | |
| Qwen2.5-0.5B GenTuneModel=Qwen2.5-0.5B, Shots=2s2026.06 | — | — | — | 85.6 | |
| RoB-baseBackbone=RoBERTa, Parameters=125M, Inference mode=fine-tuned per benchmark2026.06 | — | — | — | 87 | |
| RoB-largeBackbone=RoBERTa, Parameters=355M, Inference mode=fine-tuned per benchmark2026.06 | — | — | — | 89.5 |