2011 · 66 citations · 17 references
This paper describes a two-phase method for expanding abbreviations found in in-formal text (e.g., email, text messages, chat room conversations) using a machine translation system trained at the charac-ter level during the first phase. In this way, the system learns mappings between character-level “phrases ” and is much more robust to new abbreviations than a word-level system. We generate transla-tion models that are independent of the way in which the abbreviations are formed and show that the results show little degra-dation compared to when type-dependent models are trained. Our experiments on a large data set show our proposed system performs well when tested both on isolated abbreviations and, with the incorporation of a second phase utilizing an in-domain language model, in the context of neigh-boring words. 1
17
Philipp Koehn, Richard Zens, Chris Dyer et al. · 2007 · 4.9K citations · Full text
Natural Language Processing, Computer-assisted Translation, Engineering +14
Lexical Normalisation of Short Text Messages: Makn Sens a #twitter
Bo Han, Timothy Baldwin · 2011 · 455 citations
Normalization of non-standard words
Richard Sproat, Alan W. Black, Stanley Chen et al. · Computer Speech & Language · 2001 · 340 citations
Applied Linguistics, Natural Language Processing, Non-standard Words +9
A phrase-based statistical model for SMS text normalization
AiTi Aw, Min Zhang, Juan Xiao et al. · 2006 · 279 citations · Full text