Demystifying Statistical Learning Based on Efficient Influence Functions.
Evaluation of treatment effects and more general estimands is typically achieved via parametric modeling, which is unsatisfactory since model misspecification is likely. Data-adaptive model building (e.g., statistical/machine learning) is commonly employed to reduce the risk of misspecification. Naï...
| Publicado en: | American Statistician Vol. 76; no. 3; pp. 292 - 305 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Taylor & Francis Ltd
Aug2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=158065780&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 158065780 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 00031305 STT jtl: American Statistician issn: 00031305 maglogo: Y pubinfo: dt: Aug2022 vid: 76 iid: 3 pid: 377 pub: Taylor & Francis Ltd artinfo: ui: 158065780 10.1080/00031305.2021.2021984 ppf: 292 ppct: 13 formats: tig: atl: Demystifying Statistical Learning Based on Efficient Influence Functions. aug: au: Hines, Oliver Dukes, Oliver Diaz-Ordaz, Karla Vansteelandt, Stijn affil: Department of Medical Statistics, London School of Hygiene and Tropical Medicine, London, UK Department of Applied Mathematics, Computer Science and Statistics, Ghent University, Ghent, Belgium su: Statistical learning Nonparametric statistics Parametric modeling Sample size (Statistics) Machine learning Treatment effectiveness sug: subj: Statistical learning Nonparametric statistics Parametric modeling Sample size (Statistics) Machine learning Treatment effectiveness keyword: Data-adaptive estimation Double machine learning Nonparametric methods Post-selection inference Targeted learning Data-adaptive estimation Double machine learning Nonparametric methods Post-selection inference Targeted learning ab: Evaluation of treatment effects and more general estimands is typically achieved via parametric modeling, which is unsatisfactory since model misspecification is likely. Data-adaptive model building (e.g., statistical/machine learning) is commonly employed to reduce the risk of misspecification. Naïve use of such methods, however, delivers estimators whose bias may shrink too slowly with sample size for inferential methods to perform well, including those based on the bootstrap. Bias arises because standard data-adaptive methods are tuned toward minimal prediction error as opposed to, for example, minimal MSE in the estimator. This may cause excess variability that is difficult to acknowledge, due to the complexity of such strategies. Building on results from nonparametric statistics, targeted learning and debiased machine learning overcome these problems by constructing estimators using the estimand's efficient influence function under the nonparametric model. These increasingly popular methodologies typically assume that the efficient influence function is given, or that the reader is familiar with its derivation. In this article, we focus on derivation of the efficient influence function and explain how it may be used to construct statistical/machine-learning-based estimators. We discuss the requisite conditions for these estimators to perform well and use diverse examples to convey the broad applicability of the theory. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|