SLPOtherJournal of voice : official journal of the Voice Foundation2025

Deep-Learning-Based Representation of Vocal Fold Dynamics in Adductor Spasmodic Dysphonia during Connected Speech in High-Speed Videoendoscopy.

Ahmed M Yousef, Dimitar D Deliyski, Stephanie R C Zacharias and 1 others

PMID 36154973

WHAT IT FOUND

A new computer model accurately traces vocal fold edges in high-speed video of people with spasmodic dysphonia speaking normally.

It handles blurry images and spasms better than previous methods. This tool helps researchers measure voice breaks objectively, but it is not yet a clinical diagnostic test.

Key findings

01The model achieved a mean Dice coefficient of 0.86 and accuracy of 0.89 when segmenting vocal fold areas in high-speed video of one patient with spasmodic dysphonia.

02The model successfully identified frames where vocal folds were completely obstructed or not vibrating, matching manual expert judgments in those difficult cases.

03The model was slightly less accurate for patients with spasmodic dysphonia than for normal speakers due to uncontrolled laryngeal tissue movements obscuring the view.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The study involved only seven participants, with the final test set coming from a single patient with adductor spasmodic dysphonia. This limits the generalizability of the findings to the broader AdSD population. The model was trained on a relatively small dataset of 4,500 frames, which may not capture the full variability of vocal fold dynamics across different patients. The accuracy was lower for AdSD patients than for normal speakers, indicating that severe spasms and tissue movement still pose a challenge for automated analysis. This is a technical validation study of an algorithm, not a clinical trial. It does not prove that using this tool improves patient outcomes or diagnostic accuracy in a real-world setting.

Declared interests

The authors declared no conflicts of interest. The study was conducted at the Mayo Clinic and approved by the Institutional Review Board.

The easy way to misread this

Do not assume this algorithm is ready for clinical use. It was validated on a single patient's data in a research setting. It does not diagnose spasmodic dysphonia; it only measures vocal fold movement in videos already recorded for research purposes.

Read it on PubMed →