arXiv (Cornell University) · 2020 · 16 citations · 7 references
Abuse DetectionCommunicationMultimodal Sentiment AnalysisCorpus LinguisticsSocial SciencesText MiningNatural Language ProcessingSocial MediaHostility Detection DatasetComputational LinguisticsAffective ComputingHostility DimensionsHostile PostsLanguage StudiesContent AnalysisHate SpeechLanguage PolicingHindi LanguageAggressionSpeech AnalysisLinguistics
In this paper, we present a novel hostility detection dataset in Hindi language. We collect and manually annotate ~8200 online posts. The annotated dataset covers four hostility dimensions: fake news, hate speech, offensive, and defamation posts, along with a non-hostile label. The hostile posts are also considered for multi-label tags due to a significant overlap among the hostile classes. We release this dataset as part of the CONSTRAINT-2021 shared task on hostile post detection.
7
Abusive Language Detection in Online User Content
Chikashi Nobata, Joel Tetreault, Achint Thomas et al. · 2016 · 1.1K citations
Applied Linguistics, Natural Language Processing, Abuse Detection +14
Automated Hate Speech Detection and the Problem of Offensive Language
Thomas Davidson, Dana Warmsley, Michael W. Macy et al. · arXiv (Cornell University) · 2017 · 275 citations · Full text
BanFakeNews: A Dataset for Detecting Fake News in Bangla
Zobaer Hossain, Ashraful Rahman, Saiful Islam et al. · arXiv (Cornell University) · 2020 · 70 citations · Full text