SLPOtherJournal of speech, language, and hearing research : JSLHR2022

Validation of an Automated Procedure for Calculating Core Lexicon From Transcripts.

Sarah Grace Dalton, Brielle C Stark, Davida Fromm and 4 others

PMID 35917459

WHAT IT FOUND

An automated CLAN command calculates CoreLex scores from transcripts with excellent agreement to manual scoring.

It reduced the time to score 485 discourse samples from 30 hours to under an hour, even for users new to the software.

Key findings

01Automated scoring showed excellent reliability compared to hand scoring for both groups across all tasks, with intraclass correlations ranging from .978 to .998.

02Hand scoring was the source of most discrepancies, particularly missing irregular word forms or contractions, while automated errors were rare and random.

03Automated scoring took 25 minutes for an experienced user and 46 minutes for an inexperienced user, compared to approximately 30 hours for hand scoring the entire sample.

STILL TO COME

How it was doneWhat they foundWhat it means for SLPs

Read the rest of this summary

You get three full summaries a month, free, and we do not ask for a card. Search, the TL;DRs and your library stay unlimited either way.

Already have one?

What it does not show

The study used existing transcripts from a database, so it did not test the workflow of recording and transcribing new patient samples. The time savings apply only after transcription is complete. Since transcription is time-consuming (up to 10 minutes per minute of speech), the overall clinical burden may not decrease unless transcription is also automated. The sample included only monolingual English speakers (with two exceptions), so the tool's accuracy with other languages or dialects was not tested. CoreLex is a microlinguistic measure; it does not assess complex discourse features like coherence or narrative structure.

Declared interests

The work was supported by the National Institute on Deafness and Other Communication Disorders (N.I.H.). No other conflicts of interest were declared.

The easy way to misread this

Do not assume this tool eliminates the time burden of discourse analysis. It only automates the scoring of transcripts that have already been manually created in CHAT format. The significant time savings reported here do not apply to the transcription step, which remains a major bottleneck.

Read it on PubMed →