Nujan Split Reader
Econometrics, Causal Inference & Statistics · The Econometrics Journal 2018

Double/Debiased Machine Learning for Treatment and Structural Parameters (DML)

Authors: Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, James Robins (MIT, UCLA, University of Chicago, Harvard) · arXiv: 1608.00060

Core Methodological Innovation

Proved that naive plug-in ML estimators suffer O(N^{-1/2}) regularization and overfitting bias, and solved both by pairing Neyman-orthogonal score functions with K-fold sample splitting (cross-fitting).

Key Quantitative & Theoretical Takeaway: When nuisance estimators converge at rate o(N^{-1/4}), Neyman orthogonality eliminates first-order bias while cross-fitting eliminates empirical process overfitting bias, achieving valid confidence intervals.

Abstract

We revisit the classic semiparametric problem of inference on a low-dimensional parameter theta_0 in the presence of high-dimensional nuisance parameters eta_0. We show that combining Neyman-orthogonal moments with cross-fitting removes regularization bias and yields root-N consistent normal inference.

Step-by-Step Equation & Methodology Breakdown

Why do we need both Neyman orthogonality and cross-fitting in Double Machine Learning?

Neyman orthogonality ensures the directional derivative of the estimating moment with respect to nuisance parameters vanishes, reducing regularization bias to the product of estimation errors O(N^{-1/2}). Cross-fitting evaluates residuals on held-out folds so complex ML learners do not induce Donsker-class overfitting bias.

Read Any arXiv Paper Side-by-Side in Nujan

Replace arxiv.org with nujan.app on any arXiv URL (for example, https://nujan.app/abs/2006.11239) to open the PDF and AI research partner side-by-side.

Launch Interactive Split Reader for CAUSAL-DML