Measuring Speech Recognition With a Matrix Test Using Synthetic Speech.
Theresa Nuesse, Bianca Wiercinski, Thomas Brand and 1 others
PMID 31322032WHAT IT FOUND
Synthetic female speech gave a 0.5 dB SNR worse threshold for 50% correct speech recognition than natural female speech in 48 normal-hearing listeners, and the steepness of scores was similar.
Key findings
01The mean threshold was -8.6 dB SNR for synthetic speech and -9.1 dB SNR for natural speech, a reported 0.5 dB SNR difference, while the slope was not significantly different.
02Recognition scores were significantly better for natural speech than synthetic speech at all three tested noise levels, with mean differences of 4.1%, 7.1%, and 2.4%.
03Training improved thresholds, with the largest gain from the first to second test list being 1.1 dB SNR for synthetic speech and 0.9 dB SNR for natural speech, and a plateau after the fourth test list or 60 sentences.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
Only 48 young normal-hearing students aged 18 to 25 years were tested, so it does not show how heterogeneous participant groups would perform. The study used one German matrix test and one female text-to-speech voice from one system, so other languages, voices, or text-to-speech systems may not give the same result. The text-to-speech system was chosen using subjective ratings in quiet, while the speech recognition tests were done in noise, so the best synthetic voice for noisy listening is not proven. An expert listener manually selected the most natural-sounding synthetic word combinations, and this hand sorting took approximately 4 hours. The synthetic version required a worse threshold than natural speech, and recognition scores were significantly lower for synthetic speech at all three tested signal-to-noise ratios. Training did not fully stabilize during the first session; plateau thresholds still differed from measurement-phase thresholds by 0.9 dB for synthetic speech and 0.7 dB for natural speech. Effort reduction was limited to producing the speech material, because the speech material still needs evaluation for equivalent intelligibility of test lists.
Declared interests
Funding was provided by the European Regional Development Fund and the Niedersächsisches Vorab of the Lower Saxony Ministry for Science and Culture.
The easy way to misread this
Do not conclude that synthetic speech can replace natural speech in clinical speech audiometry. This was tested only in 48 young normal-hearing students, with one German matrix test and one female text-to-speech voice, and the synthetic version required a 0.5 dB SNR worse threshold than natural speech.