SLPOtherJournal of speech, language, and hearing research : JSLHR2024

Automatic Speech Recognition of Conversational Speech in Individuals With Disordered Speech.

Jimmy Tobin, Phillip Nelson, Bob MacDonald and 6 others

PMID 38963790

WHAT IT FOUND

Personalized speech recognition made fewer recognition errors than the unadapted model for adults with disordered speech, but conversational speech was still harder to recognize than read speech.

Key findings

01Personalized models had lower median recognition error rates: WER was 10.3, compared with 60.8 for the unadapted model.

02Conversational speech had higher median recognition error rates than read speech: WER was 36.1, compared with 14.6.

03For personalized models, speech severity significantly degraded read recognition but not conversational recognition.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The sample was only 27 adults, and most had dysarthria from ALS or cerebral palsy, so findings may not apply to other speech disorders. The study measured WER, not patient communication outcomes, device use, or meaning preserved. Speakers whose speech declined over time were excluded, and race/ethnicity data were not collected. SLP severity labels were not fully reliable for speech pattern consistency, which was excluded. The conversational training pilot used only eight participants. For personalized conversational models, severity effect was not detected because of large variance and small median differences.

Declared interests

Google Research's Project Euphonia led the work, and the paper is marked as supported by NIH extramural research. Data were collected to improve speech recognition products and services.

The easy way to misread this

Do not read this as proof that ASR improves daily communication for people with disordered speech. It measured recognition errors in 27 adults, and the pilot training result used only eight participants.

Read it on PubMed →