Case Report: Tailored automatic speech recognition in global aphasia with dysarthria - a single case proof of concept.
Davide Mulfari, Davide Cardile, Serena Campana and 6 others
PMID 42317419WHAT IT FOUND
A system trained on one woman's voice recognised her intended word in 72.65% of 936 attempts, against 56.75% for twelve rehabilitation professionals who knew her speech.
One patient and a 13-word vocabulary, so not proof it generalises.
Key findings
01The speech recogniser built on this patient's own voice identified the target word correctly in 680 of 936 test utterances, an accuracy of 72.65% (95% CI 69.71 to 75.41). These were recordings the system had never heard during training.
02Twelve rehabilitation professionals who worked in stroke rehabilitation and knew this patient's speech identified the target word in a mean 56.75% of utterances (individual range 38.5% to 73.1%), and agreed poorly with one another (ICC 0.367). The system's accuracy was 15.9 percentage points above their mean, which the authors report as significant (p = 0.001).
03Accuracy depended on how the word was prompted, from 54.6% when only a picture was shown to 84.1% when the word was played as audio alone.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
This is one person. Nothing here shows the system would work for another patient, or for this patient outside a quiet therapy room with a headset microphone. The vocabulary was only 13 words, all home automation terms, which caps what the patient can actually express with the system. The human comparison was 12 rehabilitation professionals who already knew this patient's speech, deliberately chosen as a hard benchmark. Unfamiliar listeners were never tested, so how much the system beats an ordinary communication partner is unknown. The recordings played to the listeners were not logged in advance against the same test set the system was scored on, so the two were not compared word for word on identical recordings. The system was trained and tested in a quiet room; noise, distance from the microphone and everyday settings were not examined. Accuracy was highly uneven across the 13 words, so an overall figure hides words the system and the listeners both struggled with.
Declared interests
The authors declared that financial support was received for this work and/or its publication. The research was supported by Current Research Funds 2026, Ministry of Health, Italy, and the funder named is Università degli Studi di Messina.
The easy way to misread this
Do not read the 72.65% against 56.75% as proof that a speech recogniser understands this kind of speech better than people do. It is one patient, a 13-word vocabulary, and headset recordings made in a quiet room; the human comparison was twelve professionals who already knew her speech, unfamiliar listeners were never tested, and no everyday communication outcome was measured.
Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →