Using Synthetic Data in Communication Sciences and Disorders to Promote Computational Reproducibility and Transparency.

Purpose: Reproducibility is a core principle of science, and access to a study's data is essential to reproduce its findings. However, data sharing is uncommon in the discipline of communication sciences and disorders (CSD), often due to concerns related to privacy and disclosure risks. Synthetic da...

Full description

Bibliographic Details
Published in:Journal of Speech, Language & Hearing Research Vol. 68; no. 12; pp. 5854 - 5870
Main Authors: Borders, James C., Thompson, Austin, Kearney, Elaine
Format: Article
Published: American Speech-Language-Hearing Association Dec2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=190171415&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 190171415
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 68
      iid: 12
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        190171415
        10.1044/2025_JSLHR-24-00736
      ppf: 5854
      ppct: 16
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.5MB
      tig:
        atl: Using Synthetic Data in Communication Sciences and Disorders to Promote Computational Reproducibility and Transparency.
      aug:
        au:
          Borders, James C.
          Thompson, Austin
          Kearney, Elaine
        affil:
          Department of Speech, Language & Hearing Sciences, Boston University, MA
          Department of Communication Sciences and Disorders, University of Houston, TX
          School of Health and Rehabilitation Sciences, The University of Queensland, Brisbane, Australia
          Department of Speech Pathology, Princess Alexandra Hospital, Brisbane, Queensland, Australia
      su:
        Communicative disorders
        Analysis of variance
        Pearson correlation (Statistics)
        T-test (Statistics)
        Data analysis
        Research evaluation
        Statistical sampling
        Data analytics
        Descriptive statistics
        Chi-squared test
        Odds ratio
        Statistics
        Data analysis software
        Regression analysis
      sug:
        subj:
          Communicative disorders
          Analysis of variance
          Marketing Research and Public Opinion Polling
          Pearson correlation (Statistics)
          T-test (Statistics)
          Data analysis
          Research evaluation
          Statistical sampling
          Data analytics
          Descriptive statistics
          Chi-squared test
          Odds ratio
          Statistics
          Data analysis software
          Regression analysis
      ab: Purpose: Reproducibility is a core principle of science, and access to a study's data is essential to reproduce its findings. However, data sharing is uncommon in the discipline of communication sciences and disorders (CSD), often due to concerns related to privacy and disclosure risks. Synthetic data offer a potential solution to this barrier by generating artificial data sets that do not represent real individuals yet retain statistical properties and relationships from the original data. This study aimed to explore the feasibility and preliminary utility of synthetic data to promote transparency and reproducibility in the discipline of CSD. Method: Ten open data sets were obtained from previously published research within the American Speech-Language-Hearing Association "Big Nine" domains (articulation, cognition, communication, fluency, hearing, language, social communication, voice and resonance, and swallowing) across a range of study outcomes and designs. Synthetic data sets were generated with the synthpop R package. General utility was assessed visually and with the standardized ratio of the propensity mean squared error (S_pMSE). Specific utility assessed whether inferential relationships from the original data were preserved in the synthetic data set by comparing model fit indices, coefficients, and p values. Results: All synthetic data sets showed strong general utility, maintaining univariate and bivariate distributions. Six of nine synthetic data sets that used inferential statistics showed strong specific utility, maintaining inferential relationships from the original analysis. Specific utility was low in three data sets with hierarchical structures. Conclusions: Findings suggest that synthetic data can effectively maintain statistical properties and relationships across a wide range of nonhierarchical data commonly seen in the discipline of CSD. Other approaches for hierarchical data need to be explored in future work. Researchers who use synthetic data should assess its utility in preserving their results for their own data and use-case. Open Science Form: https://doi.org/10.23641/asha.30569957
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N