IEEE Data(base) Engineering Bulletin · 2016 · 22 citations · 12 references
EngineeringNew MeasureEntity SummarizationText CubesCorpus LinguisticsText MiningAutomatic SummarizationNatural Language ProcessingInformation RetrievalData ScienceText SummarizationComputational LinguisticsQuery ExpansionLanguage StudiesKnowledge DiscoveryCube StructureDocument SubsetsMulti-modal SummarizationRetrieval Augmented GenerationLinguistics
To systematically analyze large numbers of textual documents, it is often desirable to manage documents (and their metadata) in a multi-dimensional text database (Text Cube). Such structure provides flexibility of understanding local information with different granularities. Moreover, the contextualized analysis derived from cube structure often yields comparative insights. To quickly digest the content of subsets of documents in the multi-dimensional context, we study the problem of phrase-based summarization of a subset of documents of interest. We propose a new phrase ranking measure to leverage the relation between document subsets induced by multi-dimensional context and identify phrases that truly distinguish the queried subset of documents from neighboring subsets (i.e., background). Our quality evaluation suggests the new measure involving dynamic, query-dependent background generation is more effective than previous measures using the whole corpus as a static background for finding representative phrases. Computing this measure is more expensive due to the need of access to many subsets of documents to answer one query. We develop a cube-based analytical platform that implements an efficient solution by materializing a deliberately selected part of statistics, and using these statistics to perform online query processing within a constant latency constraint. Our experiments in a large news dataset demonstrate the efficiency in both query processing time and storage cost.
12
Scalable topical phrase mining from text corpora
Ahmed El-Kishky, Yanglei Song, Chi Wang et al. · Proceedings of the VLDB Endowment · 2014 · 200 citations
Ori Ben-Yitzhak, Sivan Yogev, Nadav Golbandi et al. · 2008 · 163 citations
Text Cube: Computing IR Measures for Multidimensional Text Database Analysis
Cindy Xide Lin, Bolin Ding, Jiawei Han et al. · 2008 · 121 citations
Dynamic faceted search for discovery-driven analysis
Debabrata Dash, Jun Rao, Nimrod Megiddo et al. · 2008 · 96 citations · Full text