2018 · 21 citations · 28 references
EngineeringCommunicationSpeech RepetitionSpeech RecognitionData ScienceRobust Speech RecognitionConversation AnalysisVoice RecognitionHealth SciencesAssistive TechnologySpeech OutputCorrect ErrorsComputer ScienceSpeech CommunicationVoiceSpeech Input ErrorsSpeech ProcessingHuman-computer InteractionSpeech InputSpeech PerceptionVoice TechnologySpeech InterfaceVoice Interaction
Speech has become an increasingly common means of text input, from smartphones and smartwatches to voice-based intelligent personal assistants. However, reviewing the recognized text to identify and correct errors is a challenge when no visual feedback is available. In this paper, we first quantify and describe the speech recognition errors that users are prone to miss, and investigate how to better support this error identification task by manipulating pauses between words, speech rate, and speech repetition. To achieve these goals, we conducted a series of four studies. Study 1, an in-lab study, showed that participants missed identifying over 50% of speech recognition errors when listening to audio output of the recognized text. Building on this result, Studies 2 to 4 were conducted using an online crowdsourcing platform and showed that adding a pause between words improves error identification compared to no pause, the ability to identify errors degrades with higher speech rates (300 WPM), and repeating the speech output does not improve error identification. We derive implications for the design of audio-only speech dictation.
28
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey et al. · 2015 · 5.7K citations
Interaction in 4-second bursts
Antti Oulasvirta, Sakari Tamminen, Virpi Roto et al. · 2005 · 564 citations
The microsoft 2016 conversational speech recognition system
Wayne Xiong, Jasha Droppo, Xuedong Huang et al. · 2017 · 335 citations · Full text