2020 · 93 citations · 34 references
EngineeringSub-word EmbeddingMultilingual PretrainingCorpus LinguisticsText MiningNatural Language ProcessingSyntaxLanguage DocumentationData ScienceComputational LinguisticsPre-trained EmbeddingEntity RecognitionLanguage EngineeringCorpus AnalysisLanguage StudiesNamed-entity RecognitionMachine TranslationCode Mixed CorpusCode RepresentationLinguisticsPo Tagging
In this paper, we utilize the pre-trained embedding, sub-word embedding and closely related languages of languages in the code mixed corpus to create a meta-embedding. We then use the Transformer to encode the code mixed sentence and use Conditional Random Field to predict the Named Entities in the code-mixed text. In contrast to classical Named Entity recognition where the text is monolingual, our approach can predict the Named Entities in code-mixed corpus written both in the native script as well as Roman script. Our method is a novel method to combine the embeddings of closely related languages to identify Named Entity from Code-Mixed Indian text written using native script and Roman script in social media.
34
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, Christopher D. Manning · 2014 · 33.2K citations
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
Introduction to information retrieval
Choice Reviews Online · 2009 · 12.5K citations