The influence of listener experience, measurement scale and speech task on the reliability of auditory-perceptual evaluation of vocal quality.
Jônatas do Nascimento Alves, Anna Alice Figueiredo de Almeida, Rosiane Yamasaki and 1 others
PMID 38629682WHAT IT FOUND
Reliability of voice ratings depends on the task and scale, not just experience.
CAPE-V sentences showed the highest agreement among all listener groups, while sustained vowels were least reliable. Standardized training improved agreement more than years of clinical experience did.
Key findings
01CAPE-V sentences achieved the best interrater reliability across all listener groups, showing substantial agreement on the visual analog scale and moderate agreement on the GRBAS scale.
02Visual analog scales (VAS) yielded higher interrater reliability than numerical GRBAS scales for assessing overall voice severity.
03Listeners who had received the same standardized auditory-perceptual training demonstrated better interrater reliability than those with more years of experience but heterogeneous training backgrounds.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study used a small sample of 22 listeners, with only four participants in the specialized and non-specialized SLP groups, limiting generalizability. The specialized SLP group was heterogeneous in the type of experience (e.g., head and neck cancer vs. professional voice), which may have confounded the effect of experience time on reliability. The study did not include an 'intermediate' experience group (5-10 years), creating a gap in the experience spectrum. Reliability was assessed on recorded samples in a controlled laboratory setting, which may not fully reflect the variability of live clinical assessments.
Declared interests
The authors declared no financial support or conflicts of interest.
The easy way to misread this
Do not assume that senior clinicians provide the most reliable ratings for comparison with peers. The study found that specialized SLPs with over 10 years of experience had lower interrater reliability than less experienced groups who had undergone identical standardized training, suggesting that shared training protocols matter more than tenure for agreement.