Multi-modal Relation Extraction on MNRE (test)
84.64F1 ScoreCross-Modal Retrieval and Synthesis
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Cross-Modal Retrieval and SynthesisModality=Multi-modal2023.05 | 84.64 | 93.54 | 85.03 | 84.25 | |
| DGF-PTEncoder=GPT-2 Encoder2023.06 | 84.47 | 84.25 | 84.35 | 83.83 | |
| MRE-ISEmodality=multimodal2023.05 | 84.03 | 94.06 | 84.69 | 83.38 | |
| MRE-ISEmodality=multimodal, ablation=without CMG hyper-edges2023.05 | 83.78 | 93.97 | 84.38 | 83.2 | |
| MRE-ISEmodality=multimodal, ablation=without visual topic features (v^V)2023.05 | 83.6 | 93.63 | 84.03 | 83.18 | |
| Cross-Modal Retrieval and SynthesisImage Evidence=removed2023.05 | 83.29 | 92.83 | 83.44 | 83.15 | |
| MRE-ISEmodality=multimodal, ablation=without textual topic features (v^T)2023.05 | 83.23 | 93.05 | 83.95 | 82.53 | |
| Cross-Modal Retrieval and SynthesisVisual Evidence=removed2023.05 | 83.2 | 92.72 | 82.78 | 83.63 | |
| Cross-Modal Retrieval and SynthesisSelection Module=removed2023.05 | 83.12 | 92.75 | 82.81 | 83.44 | |
| MRE-ISEmodality=multimodal, ablation=without VSG&TSG (embeddings only)2023.05 | 83.09 | 93.12 | 83.51 | 82.67 | |
| Cross-Modal Retrieval and SynthesisConsistency Module=removed2023.05 | 83.05 | 92.68 | 83.4 | 82.71 | |
| MRE-ISEmodality=multimodal, ablation=GENE without GIB guidance (I(z, G))2023.05 | 82.97 | 93.64 | 83.61 | 82.34 | |
| Cross-Modal Retrieval and SynthesisObject Evidence=removed2023.05 | 82.69 | 92.37 | 83.02 | 82.36 | |
| MRE-ISEmodality=multimodal, ablation=without GENE2023.05 | 82.12 | 92.42 | 82.41 | 81.83 | |
| MRE-ISEmodality=multimodal, ablation=without LAMO2023.05 | 82.09 | 92.86 | 82.97 | 81.22 | |
| DGF-PTEncoder=GPT Encoder2023.06 | 82.09 | 82.03 | 81.23 | 82.48 | |
| MKGformermodality=multimodal, source=copied from Chen et al., 2022a2023.05 | 81.95 | 92.31 | 82.67 | 81.25 | |
| HVPNetmodality=multimodal, source=copied from Chen et al., 2022a2023.05 | 81.85 | 83.64 | 80.78 | — | |
| HVPnetModality=Multi-modal2023.05 | 81.85 | 92.52 | 82.64 | 80.78 | |
| HVPNetModel Type=MMRE Models2023.06 | 81.85 | — | 83.64 | 80.78 | |
| IformerModality=Multi-modal2023.05 | 81.67 | 92.38 | 82.59 | 80.78 | |
| DGF-PTEncoder=BERT Encoder2023.06 | 79.24 | 79.82 | 79.72 | 78.63 | |
| MoRe_MoERetrieval Modality=Mixture-of-Experts (Text + Image)2022.12 | 68.6 | — | — | — | |
| MoRe_ImageRetrieval Modality=Image-based2022.12 | 67.24 | — | — | — | |
| ITARetrieval Modality=None2022.12 | 66.89 | — | — | — | |
| MoRe_TextRetrieval Modality=Textual2022.12 | 66.62 | — | — | — | |
| Zheng et al. (2021a)Retrieval Modality=None2022.12 | 66.41 | — | — | — | |
| MEGAmodality=multimodal2023.05 | 66.41 | 76.15 | 64.51 | 68.44 | |
| MEGAModality=Multi-modal2023.05 | 66.41 | 80.05 | 64.51 | 68.44 | |
| MEGAModel Type=MMRE Models2023.06 | 66.41 | 76.15 | 64.51 | 68.44 | |
| MoReModality=Multi-modal2023.05 | 66.27 | 79.87 | 65.25 | 67.32 | |
| RDSmodality=multimodal, source=copied from Chen et al., 2022a2023.05 | 66.14 | — | 66.83 | 65.47 | |
| Ours: BaselineRetrieval Modality=None2022.12 | 65.77 | — | — | — | |
| Zheng et al. (2021b)Retrieval Modality=None2022.12 | 65.56 | — | — | — | |
| UMGFModality=Multi-modal2023.05 | 65.29 | 79.27 | 64.38 | 66.23 | |
| BERT+SG+Att.Model Type=MMRE Models2023.06 | 63.64 | 74.59 | 60.97 | 66.56 | |
| UMTModality=Multi-modal2023.05 | 63.46 | 77.84 | 62.93 | 63.88 | |
| ViLBERTmodality=multimodal, version=base2023.05 | 63.16 | — | 64.5 | 61.86 | |
| BERT+SGmodality=multimodal, source=copied from Chen et al., 2022a2023.05 | 62.8 | 74.09 | 62.95 | 62.65 | |
| BSGModality=Multi-modal2023.05 | 62.8 | 77.15 | 62.95 | 62.65 | |
| BERT+SGModel Type=MMRE Models2023.06 | 62.8 | 74.09 | 62.95 | 62.65 | |
| BERT(Text+Image)modality=multimodal, implementation_type=re-implementation2023.05 | 61.25 | 74.59 | 63.07 | 59.53 | |
| DP-GCNmodality=text-based, source=copied from Chen et al., 2022a2023.05 | 61.11 | 74.6 | 64.04 | 58.44 | |
| MTBModel Type=Text-based RE Models2023.06 | 60.96 | 72.73 | 64.46 | 57.81 | |
| MTBmodality=text-based, source=copied from Chen et al., 2022a2023.05 | 60.86 | 72.73 | 64.46 | 57.81 | |
| MTBModality=Text Based2023.05 | 60.86 | 75.69 | 64.46 | 57.81 | |
| BERTmodality=text-based2023.05 | 59.55 | — | 63.85 | 55.79 | |
| BERTModality=Text Based2023.05 | 59.4 | 74.42 | 58.58 | 60.25 | |
| VisualBERTmodality=multimodal, version=base2023.05 | 58.3 | — | 57.15 | 59.48 | |
| VBERTModality=Multi-modal2023.05 | 58.3 | 73.97 | 57.15 | 59.48 | |
| VisualBERTModel Type=MMRE Models2023.06 | 58.3 | — | 57.15 | 59.45 | |
| PCNNmodality=text-based2023.05 | 55.49 | 72.67 | 62.85 | 49.69 | |
| PCNNModality=Text Based2023.05 | 55.49 | 73.15 | 62.85 | 49.69 | |
| PCNNModel Type=Text-based RE Models2023.06 | 55.49 | 72.67 | 62.85 | 49.69 | |
| Glove+CNNModel Type=Text-based RE Models2023.06 | 51.39 | 70.32 | 57.81 | 46.25 | |
| REAMOinput_modality=Text+Image, scenario=zero-shot2024.06 | 24.6 | — | — | — | |
| MiniGPT-v2input_modality=Text+Image, scenario=zero-shot2024.06 | 22.4 | — | — | — | |
| InstructBLIP+SEEMinput_modality=Text+Image, scenario=zero-shot2024.06 | 17 | — | — | — | |
| LLaVA+SEEMinput_modality=Text+Image, scenario=zero-shot2024.06 | 15.4 | — | — | — |