Consequences of ignoring clustering in linear regression.

Background: Clustering of observations is a common phenomenon in epidemiological and clinical research. Previous studies have highlighted the importance of using multilevel analysis to account for such clustering, but in practice, methods ignoring clustering are often employed. We used simulated dat...

Descripción completa

Detalles Bibliográficos
Publicado en:BMC Medical Research Methodology Vol. 21; no. 1; pp. 1 - 14
Autores principales: Ntani, Georgia, Inskip, Hazel, Osmond, Clive, Coggon, David
Formato: research Journal Article
Publicado: BioMed Central 7/7/2021
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=151289210&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 151289210
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        14712288
        1CI1
      jtl: BMC Medical Research Methodology
      issn: 14712288
      maglogo: N
    pubinfo:
      dt: 7/7/2021
      vid: 21
      iid: 1
      pid: 24147
      pub: BioMed Central
    artinfo:
      ui:
        151289210
        151289210
        NLM34233609
        151289210
        10.1186/s12874-021-01333-7
        NLM34233609
        151289210
      ppf: 1
      ppct: 13
      formats:
      tig:
        atl: Consequences of ignoring clustering in linear regression.
      aug:
        au:
          Ntani, Georgia
          Inskip, Hazel
          Osmond, Clive
          Coggon, David
        affil: Medical Research Council Lifecourse Epidemiology Unit, University of Southampton, Southampton, UK
      sug:
        subj:
          Models, Statistical
          Computer Simulation
          Regression
          Linear Regression
          Cluster Analysis
          Human
          Comparative Studies
          Multicenter Studies
          Evaluation Research
          Validation Studies
          Scales
      ab: Background: Clustering of observations is a common phenomenon in epidemiological and clinical research. Previous studies have highlighted the importance of using multilevel analysis to account for such clustering, but in practice, methods ignoring clustering are often employed. We used simulated data to explore the circumstances in which failure to account for clustering in linear regression could lead to importantly erroneous conclusions.Methods: We simulated data following the random-intercept model specification under different scenarios of clustering of a continuous outcome and a single continuous or binary explanatory variable. We fitted random-intercept (RI) and ordinary least squares (OLS) models and compared effect estimates with the "true" value that had been used in simulation. We also assessed the relative precision of effect estimates, and explored the extent to which coverage by 95% confidence intervals and Type I error rates were appropriate.Results: We found that effect estimates from both types of regression model were on average unbiased. However, deviations from the "true" value were greater when the outcome variable was more clustered. For a continuous explanatory variable, they tended also to be greater for the OLS than the RI model, and when the explanatory variable was less clustered. The precision of effect estimates from the OLS model was overestimated when the explanatory variable varied more between than within clusters, and was somewhat underestimated when the explanatory variable was less clustered. The cluster-unadjusted model gave poor coverage rates by 95% confidence intervals and high Type I error rates when the explanatory variable was continuous. With a binary explanatory variable, coverage rates by 95% confidence intervals and Type I error rates deviated from nominal values when the outcome variable was more clustered, but the direction of the deviation varied according to the overall prevalence of the explanatory variable, and the extent to which it was clustered.Conclusions: In this study we identified circumstances in which application of an OLS regression model to clustered data is more likely to mislead statistical inference. The potential for error is greatest when the explanatory variable is continuous, and the outcome variable more clustered (intraclass correlation coefficient is ≥ 0.01).
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N