European Conference on Artificial Intelligence · 2008 · 73 citations · 7 references
Abstract. The automatic detection of plagiarism is a task that has acquired relevance in the Information Retrieval area and it becomes more complex when the plagiarism is made in a multilingual panorama, where the original and suspicious texts are written in different languages. From a cross-lingual perspective, a text fragment in one language is considered a plagiarism of a text in another language if their contents are considered semantically similar no matter they are written in different languages and the corresponding citation or credit is not included. Our current experiments on cross-lingual plagiarism analysis are based on the exploitation of a statistical bilingual dictionary. This dictionary is created on the basis of a parallel corpus which contains original fragments written in one language and plagiarised versions of these fragments written in another language. The process for the automatic cross-lingual plagiarism analysis based on the statistical bilingual dictionary has shown good results and we consider that it could be useful also for the cross-lingual nearduplicate detection task. 1
7
Mining the Web for bilingual text
Philip Resnik · 1999 · 238 citations · Full text
Engineering, Cross-lingual Representation, Multilingualism +19
Antonio Si, Hong Va Leong, Rynson W. H. Lau · 1997 · 171 citations
Theory Of Computing, Hong Kong Department, Hong Kongview Profile +12