Toward Generalizable Machine Learning Models in Speech, Language, and Hearing Sciences: Estimating Sample Size and Reducing Overfitting.

Purpose: Many studies using machine learning (ML) in speech, language, and hearing sciences rely upon cross-validations with single data splitting. This study's first purpose is to provide quantitative evidence that would incentivize researchers to instead use the more robust data splitting method o...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 67; no. 3; pp. 753 - 782
Autores principales: Ghasemzadeh, Hamzeh, Hillman, Robert E., Mehta, Daryush D.
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Mar2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=176083723&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 176083723
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Mar2024
      vid: 67
      iid: 3
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        176083723
        10.1044/2023_JSLHR-23-00273
      ppf: 753
      ppct: 29
      formats:
        fmt:
          @attributes:
            type: P
            size: 3MB
      tig:
        atl: Toward Generalizable Machine Learning Models in Speech, Language, and Hearing Sciences: Estimating Sample Size and Reducing Overfitting.
      aug:
        au:
          Ghasemzadeh, Hamzeh
          Hillman, Robert E.
          Mehta, Daryush D.
        affil:
          Center for Laryngeal Surgery and Voice Rehabilitation, Massachusetts General Hospital, Boston
          Department of Surgery, Harvard Medical School, Boston, MA
          Department of Communicative Sciences and Disorders, Michigan State University, East Lansing
          Speech and Hearing Bioscience and Technology, Division of Medical Sciences, Harvard Medical School, Boston, MA
          MGH Institute of Health Professions, Boston, MA
      su:
        Language & languages
        Computer simulation
        Speech
        Quantitative research
        Hearing
        Allied health education
        Speech therapy education
        Sample size (Statistics)
        Probability theory
        Descriptive statistics
        System analysis
        Machine learning
        Comparative studies
      sug:
        subj:
          Language & languages
          Computer simulation
          Speech
          Quantitative research
          Hearing
          Allied health education
          Speech therapy education
          Sample size (Statistics)
          Probability theory
          Descriptive statistics
          System analysis
          Machine learning
          Comparative studies
      ab: Purpose: Many studies using machine learning (ML) in speech, language, and hearing sciences rely upon cross-validations with single data splitting. This study's first purpose is to provide quantitative evidence that would incentivize researchers to instead use the more robust data splitting method of nested k-fold cross-validation. The second purpose is to present methods and MATLAB code to perform power analysis for ML-based analysis during the design of a study. Method: First, the significant impact of different cross-validations on ML outcomes was demonstrated using real-world clinical data. Then, Monte Carlo simulations were used to quantify the interactions among the employed crossvalidation method, the discriminative power of features, the dimensionality of the feature space, the dimensionality of the model, and the sample size. Four different cross-validation methods (single holdout, 10-fold, train--validation--test, and nested 10-fold) were compared based on the statistical power and confidence of the resulting ML models. Distributions of the null and alternative hypotheses were used to determine the minimum required sample size for obtaining a statistically significant outcome (5% significance) with 80% power. Statistical confidence of the model was defined as the probability of correct features being selected for inclusion in the final model. Results: ML models generated based on the single holdout method had very low statistical power and confidence, leading to overestimation of classification accuracy. Conversely, the nested 10-fold cross-validation method resulted in the highest statistical confidence and power while also providing an unbiased estimate of accuracy. The required sample size using the single holdout method could be 50% higher than what would be needed if nested k-fold crossvalidation were used. Statistical confidence in the model based on nested k-fold cross-validation was as much as four times higher than the confidence obtained with the single holdout--based model. A computational model, MATLAB code, and lookup tables are provided to assist researchers with estimating the minimum sample size needed during study design. Conclusion: The adoption of nested k-fold cross-validation is critical for unbiased and robust ML studies in the speech, language, and hearing sciences. Supplemental Material: https://doi.org/10.23641/asha.25237045
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N