Demographic and Acoustic Factors related to Automatic Speech Recognition Inaccuracies for Child African American English Speakers.
Brittany N Fletcher, Wei-Wen Hsu, Vesna D Novak and 6 others
PMID 41049461WHAT IT FOUND
Google's speech-to-text was about 40% inaccurate for these children regardless of whether they used African American English features.
Age drove the errors, not dialect. Do not assume the technology fails because of a patient's accent; it likely fails because they are young.
Key findings
01Transcription error rates were similar for voiced and voiceless plosives, with no significant difference based on the presence of specific African American English phonological features.
02Younger age was a significant predictor of transcription inaccuracy, whereas dialect density was not significant in the main model.
03Acoustic measures like vowel duration and fundamental frequency did not significantly predict whether the speech-to-text system made an error.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
This was a pilot study with only 11 participants, which severely limits the statistical power and reliability of the findings. The sample included mostly girls (9 girls, 2 boys), which may not represent the general population of AAE-speaking children. Data collection occurred in homes and schools, introducing background noise that affected acoustic analysis. The study only examined one specific phonological feature (final plosives) and not the broader range of AAE features that might affect ASR.
Declared interests
The study protocol was approved by the University of Cincinnati Institutional Review Board. No specific funding sources or author conflicts of interest are declared in the provided text.
The easy way to misread this
Do not conclude that Google's speech-to-text is fair or unbiased for children because dialect density was not the primary predictor of error. The error rate was still very high (around 40%), and the study was too small to rule out dialect effects entirely. The lack of significance for dialect features may be due to low power, not proof of equity.
Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →