Concepedia

Publication | Open Access

Multi-paragraph segmentation of expository text

558

Citations

23

References

1994

Year

Marti A. Hearst

Unknown Venue

Abstract

This paper describes TextTiling, an algorithm for partitioning expository texts into coherent multi-paragraph discourse units which reflect the subtopic structure of the texts. The algorithm uses domain-independent lexical frequency and distribution information to recognize the interactions of multiple simultaneous themes. Two fully-implemented versions of the algorithm are described and shown to produce segmentation that corresponds well to human judgments of the major subtopic boundaries of thirteen lengthy texts.

References

YearCitations

Page 1