Open-Source Manually Annotated Vocal Tract Database for Automatic Segmentation from 3D MRI Using Deep Learning: Benchmarking 2D and 3D Convolutional and Transformer Networks.
Subin Erattakulangara, Karthika Kelat, Katie Burnham and 4 others
PMID 40050174WHAT IT FOUND
An open-source database of 53 manually annotated vocal tract MRI volumes from 10 French speakers provides a benchmark for automatic segmentation.
The best 3D U-Net models reached 0.896 on a 0-to-1 overlap scale, and transfer learning reached that with 20 volumes instead of 45.
Key findings
01The study used 53 manually annotated 3D vocal tract volumes from 10 healthy native French speakers.
02The 3D U-Net and transfer-learning 3D U-Net achieved average Dice scores of 0.896 ± 0.05 and 0.896 ± 0.04, while the 2D slice-by-slice U-Net averaged 0.823 ± 0.05.
03Transfer learning used 20 training volumes and reached comparable results to the standard 3D U-Net, which used 45 training volumes.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The models were trained and tested on a small set of manually annotated volumes from healthy French speakers, so they were not validated in patients, in English, or across many voice disorders. All vocal tract MRI data came from a specific acquisition protocol, so performance may change with different scanners, sequences, or parameters. The test set was very small, and certain sounds and postures were poorly segmented, including the /k/ sound and the voiceless UP posture. All models frequently mis-segmented small islands of airspace in the tongue, lower incisors, or hard palate, partly because teeth and bone are difficult to image with MRI. The reference segmentations were created from a small number of annotators using STAPLE, which can reinforce shared annotator errors.
Declared interests
The authors declare no known competing financial interests or personal relationships.
The easy way to misread this
Do not conclude that automatic vocal tract segmentation is ready for clinical voice assessment. The models were tested on a small set of healthy French speaker MRI volumes and still produced non-anatomical segmentation errors.