FluencyBank Timestamped: An Updated Data Set for Disfluency Detection and Automatic Intended Speech Recognition.
Amrit Romana, Minxue Niu, Matthew Perez and 1 others
PMID 39378266WHAT IT FOUND
Voice recognition and disfluency detection tools still struggle with stuttered speech, especially repetitions and more severe stuttering.
The updated FluencyBank Timestamped data set gives researchers a way to test improvements.
Key findings
01FluencyBank Timestamped contains 3,430 annotated segments (5.3 hr) from 37 participants, with word-level timestamps and disfluency labels.
02Whisper's intended speech word error rate was 15.4% for stuttered speech and 15.2% for typical speech, but among participants with severity labels it was 8.9% for mild and 12.3% for moderate stuttering.
03A text-based disfluency detector trained on typical speech had weighted recall 0.86 on typical speech and 0.73 on stuttered speech, with repetition recall 0.88 and 0.54.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The stuttered-speech analysis used audio from only 37 adults who stutter, and severity labels were available for 13 participants only. The model results are a computational benchmark, not a clinical trial, so they do not show that any tool improves patient care or assessment accuracy. The annotation team did not include a speech-language pathologist, and the labels were made from transcripts rather than audio alone. The study focused on filled pauses, repetitions, revisions, and partial words, not stuttering-specific disfluencies such as blocks, prolongations, or broken words. Switchboard and FluencyBank differ in speaking task, recording setup, audio quality, and topics, so model performance differences may partly reflect those differences. The tested models were off-the-shelf or trained only on typical speech; fine-tuning on stuttered speech may change the results.
Declared interests
This work was supported by NIH and NSF; the authors state the content is solely their responsibility.
The easy way to misread this
Do not read the comparable overall intended-speech error rates as meaning current voice technology is ready for people who stutter. The study used 37 adults who stutter, and performance worsened with stuttering severity, especially for repetitions.