1998 · 65 citations · 3 references
Applied LinguisticsSyntaxLanguage DocumentationLexical ResourceEagles ProjectComputational LinguisticsLinguisticsLanguage TechnologyLanguage EngineeringDistinct Language FamiliesLanguage CorpusGrammarLanguage StudiesArtsSpeech DataCorpus LinguisticsMachine Translation
The EU Copernicus project Multext-East has created a multi-lingual corpus of text and speech data, covering the six languages of the project: Bulgarian, Czech, Estonian, Hungarian, Romanian, and Slovene. In addition, wordform lexicons for each of the languages were developed. The corpus includes a parallel component consisting of Orwell's Nineteen Eighty-Four, with versions in all six languages tagged for part-of-speech and aligned to English (also tagged for POS). We describe the encoding format and data architecture designed especially for this corpus, which is generally usable for encoding linguistic corpora. We also describe the methodology for the development of a harmonized set of morphosyntactic descriptions (MSDs), which builds upon the scheme for western European languages developed within the EAGLES project. We discuss the special concerns for handling the six project languages, which cover three distinct language families.
3
Nancy Ide, Jean Véronis · 1994 · 89 citations · Full text
Natural Language Processing, Applied Linguistics, Usable Software Tools +15
Dan Tufiş, Nancy Ide, Tomaz Erjavez · Language Resources and Evaluation · 1998 · 18 citations
Applied Linguistics, Eastern European Languages, Language Documentation +14
MULTEXT-EAST: Multilingual Text Tools and Corpora for Central and Eastern European Languages.
Tomaž Erjavec, Nancy Ide, Vladimír Petkevič et al. · 1995 · 17 citations