SLPOtherJournal of communication disorders2016

Deriving gradient measures of child speech from crowdsourced ratings.

Tara McAllister Byun, Daphna Harel, Peter F Halpin and 1 others

PMID 27481555

WHAT IT FOUND

Crowdsourced ratings from nine untrained listeners accurately measure gradient speech errors like covert /r/ contrasts.

Binary and visual analog scales produce equally valid results, but binary ratings are faster. This offers a practical alternative to acoustic analysis or expert panels for assessing speech development.

Key findings

01Ratings from just nine untrained listeners showed high agreement with acoustic gold standards, matching the validity of expert panels.

02Binary and visual analog scale methods yielded statistically equivalent estimates of speech accuracy.

03Binary rating tasks were completed significantly faster than visual analog scale tasks.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The study only examined the sound /r/ in North American English; results may not generalize to other phonemes or languages. The participants were untrained adults, not clinicians, which limits the applicability of the specific 'naive listener' model if the goal is to mimic expert judgment directly. The acoustic gold standard (F3-F2 distance) is a proxy for rhoticity and may not capture all perceptual nuances of speech quality. The study relied on self-reported hearing status and native English proficiency from online workers, which introduces potential noise in participant selection.

Declared interests

The study was supported by the National Institutes of Health (Extramural). The authors declared no other conflicts of interest.

The easy way to misread this

Do not assume that a single untrained listener's rating is accurate. The validity of this method depends entirely on aggregating responses from at least nine listeners to average out individual biases and errors.

Read it on PubMed →