2023 · 15 citations · 14 references
Multimodal Entity Linking (MEL) is a task that aims to link ambiguous mentions within multimodal contexts to referential entities in a multimodal knowledge base. Recent methods for MEL adopt a common framework: they first interact and fuse the text and image to obtain representations of the mention and entity respectively, and then compute the similarity between them to predict the correct entity. However, these methods still suffer from two limitations: first, as they fuse the features of text and image before matching, they cannot fully exploit the fine-grained alignment relations between the mention and entity. Second, their alignment is static, leading to low performance when dealing with complex and diverse data. To address these issues, we propose a novel framework called Dynamic Relation Interactive Network (DRIN) for MEL tasks. DRIN explicitly models four different types of alignment between a mention and entity and builds a dynamic Graph Convolutional Network (GCN) to dynamically select the corresponding alignment relations for different input samples. Experiments on two datasets show that DRIN outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of our approach. Our code and datasets are publicly available.
14
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
Kaiming He, Georgia Gkioxari, Piotr Dollár et al. · 2017 · 27.9K citations
Object Instance Segmentation, Scene Analysis, Machine Vision +13
Deep Joint Entity Disambiguation with Local Neural Attention
Octavian-Eugen Ganea, Thomas Hofmann · 2017 · 332 citations · Full text
Scalable Zero-shot Entity Linking with Dense Entity Retrieval
Ledell Wu, Fabio Petroni, Martin Josifoski et al. · 2020 · 325 citations · Full text