Applied Evidence

Test-Retest Reliability and Preliminary Validation of Smartphone-Based Threshold and Categorical loudness Measures in Home Settings.

Trends in hearing · 2026 · Pilot · SLP

Chen Xu, Lena Schell-Majoor, Birger Kollmeier

PMID 42623291

Two smartphone hearing tests matched lab results in 15 normal-hearing young adults at home.

Proof-of-concept only: no hearing-impaired patients, one university, quiet rural settings. Does not change current clinical practice.

Key findings

1No significant overall difference between home and laboratory thresholds for GRaBr (p = 0.77) or between home and laboratory loudness functions for rACALOS (p = 0.87), though frequency-specific biases of up to 5 dB appeared at 250 Hz and 4 kHz.

2GRaBr showed significantly higher test-retest reliability than the conventional SIUD method (ICC > 0.75 vs 0.59–0.77 across three frequencies).

3rACALOS produced threshold estimates closer to pure-tone thresholds than baseline ACALOS, with the highest correlation (R = 0.71) and lowest RMSE (6.9 dB) among all method combinations tested.

to read the rest of this summary — three full summaries a month, no card.

Already have an account?

What it does not show

Only 15 participants, all young (20–35), all normal hearing, all from one university — the authors explicitly call this a proof-of-concept and caution that reliability and correlation estimates may vary across samples. No hearing-impaired participants, so the results say nothing about how these methods perform in the population that would most need remote hearing assessment. All home testing took place in quiet rural or small-town settings in northwestern Germany (median 36 dBA); performance in noisier urban environments is unknown. A single test kit was used and calibrated once at the start of data collection; real-world variability in device calibration was not tested. The noise-monitoring app used the smartphone's built-in microphone with a single calibration offset at 1 kHz, which may not reflect true frequency-dependent calibration errors. The sample was recruited informally from the authors' own working groups and student body, raising the possibility of familiarity with the testing paradigm.

Declared interests

Funded by the Deutsche Forschungsgemeinschaft (German Research Foundation), grant 390895286. The authors declared no potential conflicts of interest.

Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →