SLPOtherTrends in hearing2022

Speech Recognition and Listening Effort of Meaningful Sentences Using Synthetic Speech.

Saskia Ibelings, Thomas Brand, Inga Holube

PMID 36203405

WHAT IT FOUND

Modern synthetic speech sounded less natural than a human voice, yet normal-hearing listeners understood it better and needed no extra effort to respond.

The small gain in understanding sits within the range seen between different human speakers.

Key findings

01Listeners rated the natural voice best on quality, and the deep-neural-network and WaveNet systems similar to each other and better than the older unit-selection system.

02In noise, both synthetic voices were understood better than the natural voice, at a threshold about 1.2 dB lower, while the time taken to begin responding showed no difference between synthetic and natural speech.

03The authors conclude that a modern synthetic system is a reasonable choice for building speech-recognition tests, because it saves recording and optimization time.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

Only 12 participants in the quality test and 21 in the listening test, all normal-hearing and mostly young (average 25.6 years), so nothing here applies to hearing-impaired patients. Testing was done online at home, where equipment, internet delay and distractions were uncontrolled, which the authors link to the shallower recognition slopes. The natural voice was partly mumbled, which may explain why the synthetic voice was understood better. The quality test showed participants the natural reference, so they could recognise and favour it; the authors say a fair natural-versus-synthetic comparison was not possible with this setup. Verbal response time is an imperfect stand-in for listening effort, and it did not track pupil size in earlier work, so 'no difference in effort' is a cautious reading. Results apply only to the simple everyday sentences of this test; more complex grammar was not examined.

Declared interests

The authors declared no conflicts of interest. Funding came from a university graduation programme (Jade2Pro 2.0), not from the companies whose speech systems were tested.

The easy way to misread this

Do not read the 1.2 dB advantage as proof that synthetic speech is better for patients. It was found in normal-hearing young adults at home, the natural voice was partly mumbled, and the authors state the result applies only to these simple everyday sentences, so it does not transfer to hearing-aid users or to more complex material.

Read it on PubMed →