SLPOtherJournal of voice : official journal of the Voice Foundation2024

Detection of Vocal Fold Image Obstructions in High-Speed Videoendoscopy During Connected Speech in Adductor Spasmodic Dysphonia: A Convolutional Neural Networks Approach.

Ahmed M Yousef, Dimitar D Deliyski, Stephanie R C Zacharias and 1 others

PMID 35304042

WHAT IT FOUND

A computer tool sorted high-speed laryngoscopy frames during speech as clear or blocked views of the vocal folds correctly in most tested frames from 7 adults, including 4 with spasmodic dysphonia.

It may help review long recordings, but the paper does not test clinical use.

Key findings

01The automated classifier's overall accuracy was 99.15% on validation frames and 94.18% on testing frames.

02For full HSV recordings, the automated method found obstructed vocal fold views in 39,009 frames in the vocally normal participant and 97,545 frames in the participant with AdSD, close to the rater's counts of 38,497 and 96,571 frames.

03Overall accuracy against visual analysis was 97.23% for the vocally normal recording and 92.38% for the AdSD recording.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

Only seven adults were studied, including four with adductor spasmodic dysphonia, so the tool's performance may not generalize. The testing set came from one participant with adductor spasmodic dysphonia who was not used in training. Visual analysis was done by one rater, and the paper reports only frame classification performance, not clinical diagnosis or treatment outcomes. The tool performed less well on the AdSD recording than on the vocally normal recording.

The easy way to misread this

Do not read the high frame classification results as evidence that this tool improves diagnosis or treatment for people with spasmodic dysphonia. The study tested only seven adults and compared frame labels with one rater's visual analysis, not clinical voice outcomes.

Read it on PubMed →