2014 · 177 citations · 18 references
The definitions of two coreference scoring metrics- B<sup>3</sup> and CEAF-are underspecified with respect to <i>predicted</i>, as opposed to <i>key</i> (or <i>gold</i>) mentions. Several variations have been proposed that manipulate either, or both, the key and predicted mentions in order to get a one-to-one mapping. On the other hand, the metric BLANC was, until recently, limited to scoring partitions of key mentions. In this paper, we (i) argue that mention manipulation for scoring predicted mentions is unnecessary, and potentially harmful as it could produce unintuitive results; (ii) illustrate the application of all these measures to scoring predicted mentions; (iii) make available an open-source, thoroughly-tested reference implementation of the main coreference evaluation measures; and (iv) rescore the results of the CoNLL-2011/2012 shared task systems with this implementation. This will help the community accurately measure and compare new end-to-end coreference resolution algorithms.
18
A model-theoretic coreference scoring scheme
Marc Vilain, John D. Burger, John Aberdeen et al. · 1995 · 676 citations · Full text
On coreference resolution performance metrics
Xiaoqiang Luo · 2005 · 536 citations · Full text
Natural Language Processing, Engineering, Information Retrieval +14
CoNLL-2011 Shared Task: Modeling Unrestricted Coreference in OntoNotes
Sameer Pradhan, Lance Ramshaw, Mitchell Marcus et al. · 2011 · 273 citations