2011 · 19 citations · 17 references
The vast majority of scientific journal, conference, and grant selection processes withhold the names of the reviewers from the original submitters, taking a bettersafe-than-sorry approach for maintaining collegiality within the small-world communities of academia. While the contents of a review may not color the long-term relationship between the submitter and the reviewer, it is best to not require us all to be saints. This paper raises the question of whether the assumption of reviewer anonymity still holds in the face of readily-available, high-quality machine learning toolkits. Our threat model focuses on how a member of a community might, over time, amass a large number of unblinded reviews by serving on a number of conference and grant selection committees. We show that with access to even a relatively small corpus of such reviews, simple classification techniques from existing toolkits successfully identify reviewers with reasonably high accuracy. We discuss the implications of the findings and describe some potential technical and policy-based countermeasures. 1
17
Introduction to information retrieval
Choice Reviews Online · 2009 · 12.5K citations
Edward Loper, Steven Bird · 2002 · 3.3K citations
Natural Language Processing, Natural Language Toolkit, Problem Sets +12
NLTK: The Natural Language Toolkit
Edward Loper, Steven Bird · ArXiv.org · 2002 · 1.9K citations · Full text
Natural Language Processing, Natural Language Toolkit, Problem Sets +14
A Bayesian Approach to Filtering Junk E-Mail
Mehran Sahami, Susan Dumais, David Heckerman et al. · 1998 · 1.2K citations