SLPOtherClinical linguistics & phonetics2017

Finding the experts in the crowd: Validity and reliability of crowdsourced measures of children's gradient speech contrasts.

Daphna Harel, Elaine Russo Hitchcock, Daniel Szeredi and 2 others

PMID 27267258

WHAT IT FOUND

Crowdsourced listeners rating children's 'r' sounds on a visual scale agreed with acoustic measurements.

However, individual raters varied widely in quality. Screening raters for consistency across repeated trials improved the reliability of results, especially when using small groups.

Key findings

01When ratings from 120 crowdsourced listeners were pooled, their visual scale scores correlated highly with the acoustic gold standard for rhoticity.

02Individual raters varied widely in both validity and reliability, but those who were consistent across repeated trials also gave more valid ratings.

03Screening out raters with low consistency improved the worst-case validity outcomes when using small samples of nine raters.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The study only examined ratings of the /ɝ/ sound in North American English. It is unclear if these screening methods work for other speech sounds or accents. The participants were recruited from an online platform and self-reported their English proficiency and hearing status, which may introduce some bias. The study did not test 'attention check' questions, a common method for excluding inattentive online raters, so it is unclear how that method compares to the consistency screening proposed here. The acoustic gold standard relied on manual measurements by student assistants, which can vary between individuals.

Declared interests

The authors report no conflicts of interest.

The easy way to misread this

Do not assume that any random group of online listeners will provide valid speech ratings. The study showed that while the average listener is accurate, individual variation is high. Without screening for consistency, a small group of nine raters could produce results that are far less reliable than the acoustic standard.

Read it on PubMed →