Proceedings of the International Conference on Document Analysis and Recognition · 2007 · 14 citations · 28 references
EngineeringKnowledge ExtractionCorpus LinguisticsText MiningScientific ReferencesNatural Language ProcessingInformation RetrievalData ScienceTransducer ApproachComputational LinguisticsData IntegrationLanguage StudiesNamed-entity RecognitionStatisticsMachine TranslationKnowledge DiscoveryTerminology ExtractionInformation ExtractionCora DatasetKeyword ExtractionData ExtractionLinguistics
We present the application of probabilistic finite state transducers to the task of bibliographic meta-data extraction from scientific references. By using the transducer approach, which is often applied successfully in computational linguistics, we obtain a trainable and modular framework. This results in simplicity, flexibility, and easy adaptability to changing requirements. An evaluation on the Cora dataset that serves as a common benchmark for accuracy measurements yields a word accuracy of 88.5%, afield accuracy of 82.6%, and an instance accuracy of 42.7%. Based on a comparison to other published results, we conclude that our system performs second best on the given data set using a conceptually simple approach and implementation.
28
Introduction to the Theory of Computation
Michael Sipser · ACM SIGACT News · 1996 · 2.8K citations
Three models for the description of language
Noam Chomsky · IEEE Transactions on Information Theory · 1956 · 2.8K citations
On certain formal properties of grammars
Noam Chomsky · Information and Control · 1959 · 1.6K citations
C. Lee Giles, Kurt Bollacker, Steve Lawrence · 1998 · 1K citations