Double machine learning for sample selection models

Michela Bia; Martin Huber; Lukáš  Lafférs

Double machine learning for sample selection models

Michela Bia, Martin Huber, Lukáš Lafférs

Labour Market

Résultats de recherche: Papier de travail › Working paper

Résumé

This paper considers treatment evaluation when outcomes are only observed for a subpopulation due to sample selection or outcome attrition/non-response. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. To control in a data-driven way for potentially high dimensional pre-treatment covariates that motivate the selection-on-observables assumptions, we adapt the double machine learning framework to sample selection problems. That is, we make use of (a) Neyman-orthogonal and doubly robust score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learning-based estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners. The estimator is available in the causalweight package for the statistical software R.

langue originale	Anglais
Éditeur	arXiv.org (Cornell University)
Nombre de pages	36
état	Publié - 9 déc. 2020

Une note bibliographique

This article was submitted and deposit in arXiv : a free distribution service and an open-access archive.

Accès au document

https://arxiv.org/abs/2012.00745

Contient cette citation

@techreport{638c721832204ce9aa715b6e4bf9a98e,

title = "Double machine learning for sample selection models",

abstract = "This paper considers treatment evaluation when outcomes are only observed for a subpopulation due to sample selection or outcome attrition/non-response. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. To control in a data-driven way for potentially high dimensional pre-treatment covariates that motivate the selection-on-observables assumptions, we adapt the double machine learning framework to sample selection problems. That is, we make use of (a) Neyman-orthogonal and doubly robust score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learning-based estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners. The estimator is available in the causalweight package for the statistical software R.",

keywords = "sample selection, double machine learning, doubly robust estimation, efficient score",

author = "Michela Bia and Martin Huber and Luk{\'a}{\v s} Laff{\'e}rs",

note = "This article was submitted and deposit in arXiv : a free distribution service and an open-access archive.",

year = "2020",

month = dec,

day = "9",

language = "English",

publisher = "arXiv.org (Cornell University)",

address = "United Kingdom",

type = "WorkingPaper",

institution = "arXiv.org (Cornell University)",

}

TY - UNPB

T1 - Double machine learning for sample selection models

AU - Bia, Michela

AU - Huber, Martin

AU - Lafférs, Lukáš

N1 - This article was submitted and deposit in arXiv : a free distribution service and an open-access archive.

PY - 2020/12/9

Y1 - 2020/12/9

N2 - This paper considers treatment evaluation when outcomes are only observed for a subpopulation due to sample selection or outcome attrition/non-response. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. To control in a data-driven way for potentially high dimensional pre-treatment covariates that motivate the selection-on-observables assumptions, we adapt the double machine learning framework to sample selection problems. That is, we make use of (a) Neyman-orthogonal and doubly robust score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learning-based estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners. The estimator is available in the causalweight package for the statistical software R.

AB - This paper considers treatment evaluation when outcomes are only observed for a subpopulation due to sample selection or outcome attrition/non-response. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. To control in a data-driven way for potentially high dimensional pre-treatment covariates that motivate the selection-on-observables assumptions, we adapt the double machine learning framework to sample selection problems. That is, we make use of (a) Neyman-orthogonal and doubly robust score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learning-based estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners. The estimator is available in the causalweight package for the statistical software R.

KW - sample selection

KW - double machine learning

KW - doubly robust estimation

KW - efficient score

M3 - Working paper

BT - Double machine learning for sample selection models

PB - arXiv.org (Cornell University)

ER -