Increasing dependability of caregiver implementation fidelity estimates in early intervention: A generalizability and decision study.
Lauren H Hampton, Micheal P Sandbank, Jerrica Butler and 1 others
PMID 41066310WHAT IT FOUND
A single 20-minute video of a parent using therapy strategies with their infant is too unstable to measure skill reliably.
Clinicians should average scores from multiple separate occasions to get a dependable estimate of how well a caregiver is implementing naturalistic developmental behavioral interventions.
Key findings
01A single observation of caregiver strategy use yielded an absolute G coefficient of 0.43, which is far below the target of 0.75 for sufficient measurement stability.
02Most of the measurement error came from when the observation happened (occasion), not from who rated the video or whether it was a snack or play routine.
03Averaging scores across two separate occasions achieved sufficient dependability (absolute G coefficient = 0.77), whereas adding more raters to a single occasion provided only minimal improvement.
STILL TO COME
How it was doneWhat they foundWhat it means for SLPs
Read the rest of this summary
You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.
What it does not show
The study only looked at baseline data before intensive coaching, so stability might differ after parents have practiced the strategies. All observations were done via telehealth, so it is unclear if these stability estimates apply to in-person visits. The sample was small (20 dyads) and limited to infants/toddlers who were siblings of autistic children, which may not generalize to other populations. The analysis could not separate child variability from caregiver variability because each parent was only observed with one child.
Declared interests
Funded by the National Institute of Deafness and Other Communication Disorders and the Health Resources and Services Administration. The authors declared no potential conflicts of interest.
The easy way to misread this
Do not assume that adding a second rater to score the same video is enough to fix measurement instability. The study showed that rater agreement was already high, and adding a second rater only marginally improved dependability. The real solution is collecting data from multiple separate occasions.
Summarised by AI from the full paper, without a clinician reviewing it. Check it against the source before it changes what you do. Read it on PubMed →