SLPOtherTrends in hearing2026

Identifying Hearing Difficulty Moments in Conversational Audio.

Jack Collins, Adrian Buzea, Chris Collier and 5 others

PMID 42023978

WHAT IT FOUND

Software can flag the moments in ordinary conversation when a listener signals they missed what was just said, using sound alone.

Reading a transcript instead does not work: that scored no better than spotting "what" and "huh".

Key findings

01An audio language model given ten worked examples was the best method tested for picking out hearing difficulty moments, scoring 0.87 on the F1 measure against 0.76 for a fine-tuned speech recognition model.

02The same model with no examples at all, given only the audio, reached 0.75, close to the fine-tuned speech model's 0.76.

03Take the audio away and give the model only a machine transcript and performance fell to 0.39, the same as searching that transcript for keywords and far below the 0.75 achieved when the sound itself was used.

STILL TO COME

How it was doneWhat they found

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The trouble spots the models learned from were hand-labelled, not recorded from people known to have hearing loss. The authors call these labels "silver-standard", and the final selection of which events counted was made by one coder with no second rater and no reliability check, so some training examples are probably mislabelled. The models cannot say which of the two speakers was struggling. The audio fed in mixes both voices, while a hearing aid has only the wearer's own microphone, and it is not yet known whether the wearer's signals stay detectable when the other person's speech is faint or missing. Training used one negative example for every ten positive ones, but the authors note these events happen roughly once every 1,000 utterances in ordinary talk, so the models would probably raise more false alarms in daily use. The recordings were clean telephone calls and meetings with orderly turn-taking. Long silences, background noise, movement around a room and off-task talk were largely absent, so performance in noisy, always-on settings is unknown. Only spoken signals of difficulty were studied; facial expressions and other non-verbal signs were noted as future work, not measured here. The four-second listening window was a design choice rather than a tested boundary, and the model used only the audio before the moment, never what came after it. The evaluation ignored speed and computing demands. Models of this size could not run on a hearing aid as they stand.

Declared interests

Funded by the Australian Future Hearing Initiative, a collaboration formed under Google's Digital Future Initiative in partnership with Macquarie University. The authors declared no potential conflicts of interest. The systems being tested are the funder's own: the Gemini language models and Chirp 2 from Google's speech model family.

The easy way to misread this

Do not read this as a tool you could use with patients. No one with hearing loss took part; the software was tested on archived telephone calls and meetings whose trouble spots had already been labelled by hand, and it cannot tell which speaker was struggling.

Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →