Perceptual Adaptation to Noise-Vocoded Voice Clones.
Han Wang, Carolyn McGettigan, Patti Adank
AI voice clones stayed more intelligible than the original human voice even after degradation that strips harmonic detail — a 13% accuracy advantage that held.
Listeners adapted to both at similar rates, but hearing human speech first slightly hurt later recognition of cloned speech.
Key findings
1Cloned voices were more intelligible than human voices even after six-band noise-vocoding: 14.0% higher accuracy in the human-first group and 12.1% higher in the clone-first group (13% on average across both orders).
2Listeners adapted to both human and cloned vocoded speech at similar rates, improving about 16–17% over 40 trials in the first block (human: 16.1%, cloned: 17%), with no further adaptation in the second block.
3A small negative transfer effect was found: cloned speech was 2.2% less intelligible when it followed human speech, rather than the positive transfer the authors had predicted.
Still to come
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
Between-subjects design: each participant heard only one order, so adaptation to the two voice types could not be compared within the same person, and no return to baseline was possible once perceptual learning began. The vocoder setting (300 Hz envelope cutoff, 6 bands) is less severe than typical cochlear-implant simulations (30–50 Hz cutoff), so the degradation tested is moderate rather than extreme. The cloning system (ElevenLabs) is proprietary; its algorithm, training data, and model updates are not disclosed, so results cannot be independently replicated and may change with software updates. Only four sentences per voice per block, and all participants were young (19–35), monolingual, native British English speakers with normal hearing — findings may not generalise to older listeners, multilingual speakers, or those with hearing loss. The negative transfer effect (2.2%) is small and its mechanism is unclear; the authors note it could reflect fatigue, saturation, or a genuine cost of adapting to natural speech before encountering synthetic speech.
Declared interests
Funded by University College London. The authors declared no conflicts of interest.
The easy way to misread this
Do not read the 13% intelligibility advantage as evidence that voice cloning will improve a patient's real-world communication. This is a perceptual experiment with 80 normal-hearing young adults completing a typing task online, not a clinical trial. The cloning system is proprietary and not replicable, and the degradation level tested is milder than what a cochlear implant user actually experiences day to day.
Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →