2009 · 95 citations · 8 references
Heuristic ApproachEngineeringInformation RetrievalData ScienceData MiningKnowledge ExtractionManagementData IntegrationComputer SciencePdf Documents.the HeuristicsData ExtractionInformation ExtractionStructured DocumentPdf DocumentsDocument ProcessingText MiningData Modeling
This paper presents PDF-TREX, an heuristic approach for table recognition and extraction from PDF documents.The heuristics starts from an initial set of basic content elements and aligns and groups them, in bottom-up way by considering only their spatial features, in order to identify tabular arrangements of information. The scope of the approach is to recognize tables contained in PDF documents as a 2-dimensional grid on a Cartesian plane and extract them as a set of cells equipped by 2-dimensional coordinates. Experiments, carried out on a dataset composed of tables contained in documents coming from different domains, shows that the approach is well performing in recognizing table cells.The approach aims at improving PDF document annotation and information extraction by providing an output that can be further processed for understanding table and document contents.
8
Towards domain-independent information extraction from web tables
Wolfgang Gatterbauer, Paul Bohunsky, Marcus Herzog et al. · 2007 · 228 citations
Web Tables, Web Mining, Engineering +11
pdf2table: A Method to Extract Table Information from PDF Files.
Burcu Yildiz, Katharina Kaiser, Silvia Miksch · 2005 · 106 citations
Transforming arbitrary tables into logical form with TARTAR
Aleksander Pivk, Philipp Cimiano, York Sure et al. · Data & Knowledge Engineering · 2006 · 82 citations