Enhancing the classification of aphasia: a statistical analysis using connected speech.

Large-shared databases and automated language analyses allow for the application of new data analysis techniques that can shed new light on the connected speech of people with aphasia (PWA). To identify coherent clusters of PWA based on language output using unsupervised statistical algorithms and t...

Full description

Bibliographic Details
Published in:Aphasiology Vol. 36; no. 12; pp. 1492 - 1520
Main Authors: Fromm, Davida, Greenhouse, Joel, Pudil, Mitchell, Shi, Yichun, MacWhinney, Brian
Format: research tables/charts Journal Article
Published: Taylor & Francis Ltd Dec2022
Online Access:View this record in EBSCOhost
Description
Summary:Large-shared databases and automated language analyses allow for the application of new data analysis techniques that can shed new light on the connected speech of people with aphasia (PWA). To identify coherent clusters of PWA based on language output using unsupervised statistical algorithms and to identify features that are most strongly associated with those clusters. Clustering and classification methods were applied to language production data from 168 PWA. Language samples were from a standard discourse protocol tapping four genres: free speech personal narratives, picture descriptions, Cinderella storytelling, and procedural discourse. Seven distinct clusters of PWA were identified by the K-means algorithm. Using the random forest algorithm, a classification tree was proposed and validated, showing 91% agreement with the cluster assignments. This representative tree used only two variables to divide the data into distinct groups: total words from free speech tasks and total closed-class words from the Cinderella storytelling task. Connected speech data can be used to distinguish PWA into coherent groups, providing insight into traditional aphasia classifications, factors that may guide discourse research and clinical work.