Publication | Open Access
The MULTEXT-east morphosyntactic specifications for Slavic languages
32
Citations
7
References
2003
Year
Unknown Venue
Applied LinguisticsSyntaxCorpus LinguisticsComputational LinguisticsWord-level Morphosyntactic DescriptionsMorphologyHistorical LinguisticsLinguistic TypologyGrammarSlavic Language SpecificationsLanguage StudiesMorphology (Linguistics)LinguisticsCategorial GrammarMachine Translation
Word-level morphosyntactic descriptions, such as "Ncmsn" designating a common masculine singular noun in the nominative, have been developed for all Slavic languages, yet there have been few attempts to arrive at a proposal that would be harmonised across the languages. Standardisation adds to the interchange potential of the resources, making it easier to develop multilingual applications or to evaluate language technology tools across several languages. The process of the harmonisation of morphosyntactic categories, esp. for morphologically rich Slavic languages is also interesting from a language-typological perspective. The EU Multext-East project developed corpora, lexica and tools for seven languages, with the focus being on morphosyntactic data, including formal, EAGLES-based specifications for lexical morphosyntactic descriptions. The specifications were later extended, so that they currently cover nine languages, five from the Slavic family: Bulgarian, Croatian, Czech, Serbian and Slovene. The paper presents these morphosyntactic specifications, giving their background and structure, including the encoding of the tables as TEI feature structures. The five Slavic language specifications are discussed in more depth.
| Year | Citations | |
|---|---|---|
Page 1
Page 1