Findings of the Association for Computational Linguistics: ACL 2022 · 2022 · 48 citations · 47 references
Dialogue safety problems severely limit the real-world deployment of neural conversational models and have attracted great research interests recently. However, dialogue safety problems remain under-defined and the corresponding dataset is scarce. We propose a taxonomy for dialogue safety specifically designed to capture unsafe behaviors in humanbot dialogue settings, with focuses on contextsensitive unsafety, which is under-explored in prior works. To spur research in this direction, we compile DIASAFETY, a dataset with rich context-sensitive unsafe examples. Experiments show that existing safety guarding tools fail severely on our dataset. As a remedy, we train a dialogue safety classifier to provide a strong baseline for context-sensitive dialogue unsafety detection. With our classifier, we perform safety evaluations on popular conversational models and show that existing dialogue systems still exhibit concerning contextsensitive safety problems.
47
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser et al. · Information Fusion · 2019 · 8.1K citations · Full text
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh et al. · 2020 · 7.7K citations · Full text
On the Dangers of Stochastic Parrots
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major et al. · 2021 · 4.7K citations · Full text
Engineering, Multilingual Pretraining, Large Language Model +24