Multimodal and Spectral Degradation Effects on Speech and Emotion Recognition in Adult Listeners.
Chantel Ritter, Tara Vongpaisal
PMID 30378469WHAT IT FOUND
Seeing the talker's face helped normal-hearing adults repeat degraded speech and identify emotions, especially when the audio was very distorted.
Emotion recognition stayed near perfect with visual cues.
Key findings
01Speech decoding accuracy was higher when participants saw the talker (76.9%) than when they heard only (51.5%).
02Emotion recognition reached 99.6% accuracy with audiovisual input, and spectral degradation did not affect it; auditory-only averaged 70.0% and was affected.
03In the two-band speech condition, repetition accuracy was 3.2% with auditory-only input and 55.9% with audiovisual input.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study used 30 normal-hearing adults, not people with hearing loss or cochlear implants, so it may not show how patients perform with real devices. It tested only one female talker and short, simple sentences, so it may not capture everyday conversation. Audiovisual emotion recognition was at ceiling, so differences between conditions were hard to see. There was no visual-only condition, so the exact contribution of visual cues versus auditory cues cannot be separated. The vocoder was a simulation of degraded hearing, not a clinical treatment or training program. Response time for emotion recognition was not sensitive to condition differences.
The easy way to misread this
Do not conclude that cochlear implant users will recognize emotions perfectly because they saw the talker's face. The participants were normal-hearing adults listening to simulated degraded speech, not implant users receiving a treatment.