Comparison of Deep Learning Models for Objective Auditory Brainstem Response Detection: A Multicenter Validation Study.
Yin Liu, Lingjie Xiang, Qiang Li and 7 others
PMID 40457875WHAT IT FOUND
The best model detected wave V correctly in 91.90% of held-out ABR waveforms.
When converted to ear-level thresholds, 95.41% were within 10 dB of expert labels.
Key findings
01ResPatchTST was the best model on the held-out test set from Clinical Dataset I, with 91.90% accuracy, 0.976 AUC, 0.919 F1-score, 92.93% specificity and 90.89% sensitivity.
02When waveform predictions were aggregated into ear-level ABR thresholds, 2,588 of 3,052 ears (84.80%) matched the expert threshold and 2,912 of 3,052 ears (95.41%) were within 10 dB of it.
03Models trained on mixed-age or mixed-hearing-status data generalized better to unseen groups than models trained on restricted age or hearing-status groups.
STILL TO COME
How it was doneWhat they found
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The models judged each waveform on its own, so they did not use the sequence of waveforms across stimulus levels, which the authors say matters most near threshold. The held-out test set came from the same main dataset and had a similar distribution, so it may not show how the models handle all real-world variation. Performance varied across centres, and the best model reached 81.05% accuracy on the Mendeley Dataset, where data were small and expert labels were inconsistent. The external validation datasets were small, with 12, 8 and 8 participants in the Southampton, PhysioNet and Mendeley datasets. Ground truth labels came from expert visual inspection, and disagreement required a third expert, so the reference standard itself has variability.
Declared interests
Funding was from the National Key Research and Development Program of China, National Natural Science Foundation of China, Natural Science Foundation of Chongqing, China, and the Science and Technology Research Program of Chongqing Municipal Education Commission, with grant numbers 2022YFA1004100, 62301096, CSTB2023NSCQ-MSX0659, cstc2021jcyj-bshX0206 and KJQN202400632. The authors declared no potential conflicts of interest.
The easy way to misread this
Do not read the 91.90% accuracy as proof that this model can replace expert ABR interpretation in every clinic. It was tested on stored waveforms, not in a live clinical workflow, and its accuracy fell to 81.05% on the Mendeley Dataset.