Estimating Classification Consistency of Machine Learning Models for Screening Measures.

This article illustrates novel quantitative methods to estimate classification consistency in machine learning models used for screening measures. Screening measures are used in psychology and medicine to classify individuals into diagnostic classifications. In addition to achieving high accuracy, i...

Descripción completa

Detalles Bibliográficos
Publicado en:Psychological Assessment Vol. 36; no. 6/7; pp. 395 - 407
Autores principales: Gonzalez, Oscar, Georgeson, A. R., Pelham III, William E.
Formato: Artículo
Publicado: American Psychological Association Jun/Jul2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=177610271&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 177610271
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10403590
        POL
      jtl: Psychological Assessment
      issn: 10403590
      maglogo: N
    pubinfo:
      dt: Jun/Jul2024
      vid: 36
      iid: 6/7
      pid: 34
      pub: American Psychological Association
    artinfo:
      ui:
        177610271
        10.1037/pas0001313
      ppf: 395
      ppct: 12
      formats:
      tig:
        atl: Estimating Classification Consistency of Machine Learning Models for Screening Measures.
      aug:
        au:
          Gonzalez, Oscar
          Georgeson, A. R.
          Pelham III, William E.
        affil:
          Department of Psychology and Neuroscience, University of North Carolina at Chapel Hill
          Department of Psychology, Arizona State University
          Department of Psychiatry, University of California San Diego
      su:
        Personality disorder diagnosis
        Decision making
        Medical screening
        Statistical models
        Prediction models
        Probability theory
        Descriptive statistics
        Machine learning
        Data analysis software
        Algorithms
      sug:
        subj:
          Personality disorder diagnosis
          Decision making
          Medical screening
          All Other Miscellaneous Ambulatory Health Care Services
          Statistical models
          Prediction models
          Probability theory
          Descriptive statistics
          Machine learning
          Data analysis software
          Algorithms
      keyword:
        classification consistency
        machine learning
        reliability
        screening
        classification consistency
        machine learning
        reliability
        screening
      ab: This article illustrates novel quantitative methods to estimate classification consistency in machine learning models used for screening measures. Screening measures are used in psychology and medicine to classify individuals into diagnostic classifications. In addition to achieving high accuracy, it is ideal for the screening process to have high classification consistency, which means that respondents would be classified into the same group every time if the assessment was repeated. Although machine learning models are increasingly being used to predict a screening classification based on individual item responses, methods to describe the classification consistency of machine learning models have not yet been developed. This article addresses this gap by describing methods to estimate classification inconsistency in machine learning models arising from two different sources: sampling error during model fitting and measurement error in the item responses. These methods use data resampling techniques such as the bootstrap and Monte Carlo sampling. These methods are illustrated using three empirical examples predicting a health condition/diagnosis from item responses. R code is provided to facilitate the implementation of the methods. This article highlights the importance of considering classification consistency alongside accuracy when studying screening measures and provides the tools and guidance necessary for applied researchers to obtain classification consistency indices in their machine learning research on diagnostic assessments. Public Significance Statement: Recently, methods for machine learning have been used to predict from a screening measure if individuals should be flagged for a condition (e.g., as depressed vs. not depressed), but it is unknown if the models provide consistent screening decisions if respondents were to repeatedly receive the screening measure. We propose statistical procedures to help researchers determine if a machine learning model is providing consistent screening decisions.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N