Publication | Closed Access
Web Content Extraction based on Webpage Layout Analysis
20
Citations
17
References
2010
Year
Unknown Venue
Web MiningGlobal ViewWeb Content ExtractionWebpage Layout AnalysisInformation RetrievalContent AnalysisData MiningEngineeringKnowledge DiscoveryData ExtractionInformation ExtractionLayout InformationContent ProcessingText Mining
For web content extraction task, researchers have proposed many different methods, such as wrapper-based method, DOM tree rule-based method, machine learning-based method and so on. To some extent, all these methods ignore the layout information of the webpage, although the layout information such as the spatial and visual cues often plays a very important role in the process of locating the main content of the webpage when browsing. As a consequence, these methods often throw part of the main content away when extracting content from the webpage. In this paper, we present a method which combines webpage layout analysis with DOM tree rule-base method, it can make full use of the advantages of the two methods. It uses the layout information to guide the extraction work with a global view and can gain a better performance than the traditional methods.
| Year | Citations | |
|---|---|---|
Page 1
Page 1