2015 · 27 citations · 10 references
Abuse DetectionEngineeringKnowledge ExtractionInformation ForensicsSemantic WebCorpus LinguisticsJournalismText MiningVandalism RevisionsSemantic WikiNatural Language ProcessingInformation RetrievalData ScienceData MiningKnowledge BasesComputational LinguisticsLanguage StudiesContent AnalysisPublic Knowledge BasesKnowledge RepresentationKnowledge DiscoveryComputer ScienceKnowledge BaseLanguage CorpusTowards Vandalism DetectionKnowledge ManagementLinguistics
We report on the construction of the Wikidata Vandalism Corpus WDVC-2015, the first corpus for vandalism in knowledge bases. Our corpus is based on the entire revision history of Wikidata, the knowledge base underlying Wikipedia. Among Wikidata's 24 million manual revisions, we have identified more than 100,000 cases of vandalism. An in-depth corpus analysis lays the groundwork for research and development on automatic vandalism detection in public knowledge bases. Our analysis shows that 58% of the vandalism revisions can be found in the textual portions of Wikidata, and the remainder in structural content, e.g., subject-predicate-object triples. Moreover, we find that some vandals also target Wikidata content whose manipulation may impact content displayed on Wikipedia, revealing potential vulnerabilities. Given today's importance of knowledge bases for information systems, this shows that public knowledge bases must be used with caution.
10
Denny Vrandečić, Markus Krötzsch · Communications of the ACM · 2014 · 3.2K citations · Full text
Andrei Broder · ACM SIGIR Forum · 2002 · 1.9K citations
Nagarajan Natarajan, Inderjit S. Dhillon, Pradeep Ravikumar et al. · 2013 · 563 citations