SLPOtherTrends in hearing2019

Measuring Speech Recognition With a Matrix Test Using Synthetic Speech.

Theresa Nuesse, Bianca Wiercinski, Thomas Brand and 1 others

PMID 31322032

WHAT IT FOUND

Synthetic female speech gave a 0.5 dB SNR worse threshold for 50% correct speech recognition than natural female speech in 48 normal-hearing listeners, and the steepness of scores was similar.

Key findings

01The mean threshold was -8.6 dB SNR for synthetic speech and -9.1 dB SNR for natural speech, a reported 0.5 dB SNR difference, while the slope was not significantly different.

02Recognition scores were significantly better for natural speech than synthetic speech at all three tested noise levels, with mean differences of 4.1%, 7.1%, and 2.4%.

03Training improved thresholds, with the largest gain from the first to second test list being 1.1 dB SNR for synthetic speech and 0.9 dB SNR for natural speech, and a plateau after the fourth test list or 60 sentences.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

Only 48 young normal-hearing students aged 18 to 25 years were tested, so it does not show how heterogeneous participant groups would perform. The study used one German matrix test and one female text-to-speech voice from one system, so other languages, voices, or text-to-speech systems may not give the same result. The text-to-speech system was chosen using subjective ratings in quiet, while the speech recognition tests were done in noise, so the best synthetic voice for noisy listening is not proven. An expert listener manually selected the most natural-sounding synthetic word combinations, and this hand sorting took approximately 4 hours. The synthetic version required a worse threshold than natural speech, and recognition scores were significantly lower for synthetic speech at all three tested signal-to-noise ratios. Training did not fully stabilize during the first session; plateau thresholds still differed from measurement-phase thresholds by 0.9 dB for synthetic speech and 0.7 dB for natural speech. Effort reduction was limited to producing the speech material, because the speech material still needs evaluation for equivalent intelligibility of test lists.

Declared interests

Funding was provided by the European Regional Development Fund and the Niedersächsisches Vorab of the Lower Saxony Ministry for Science and Culture.

The easy way to misread this

Do not conclude that synthetic speech can replace natural speech in clinical speech audiometry. This was tested only in 48 young normal-hearing students, with one German matrix test and one female text-to-speech voice, and the synthetic version required a 0.5 dB SNR worse threshold than natural speech.

Read it on PubMed →