crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Binding Crash Course
  • Regression And GLMs
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • MEstimator Poisson
  • Causal Inference
    • Balancing Weights
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
  • Transforms
    • PCA And Kernel Basis
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding: First Course
    • Overview And TOC
    • Ch 1 Correlation And Simpson
    • Ch 2 Potential Outcomes
    • Ch 3 CRE And Fisher RT
    • Ch 4 CRE And Neyman
    • Ch 9 Bridging Finite And Superpopulation
    • Ch 11 Propensity Score
    • Ch 12 Double Robust ATE
    • Ch 13 Double Robust ATT
    • Ch 21 Experimental IV
    • Ch 23 Econometric IV
    • Ch 27 Mediation

On this page

  • 1 Where it fits
  • 2 Cross-fitted objective
  • 3 Inference
  • 4 Performance and numerical behavior
  • 5 Python API
  • 6 Minimal example
  • 7 summary() contract

PartiallyLinearDML

Cross-fit partially linear Double ML

from _api_doc_utils import *

1 Where it fits

Group: Causal inference

PartiallyLinearDML estimates the treatment coefficient in

\[ y = \theta d + g(x) + u, \qquad d = m(x) + v, \]

using cross-fitted ridge nuisance regressions. The final coefficient is estimated from the orthogonalized residual-on-residual score. Fold membership is a deterministic seeded shuffle, so equal seeds reproduce assignments and different seeds change them.

2 Cross-fitted objective

For the partially linear model

\[ Y=\theta D+\ell(X)+U, \qquad D=m(X)+V, \]

the class creates seeded outer folds. On the complement of each fold it separately fits ridge nuisance regressions for \(Y\) and \(D\):

\[ \min_{a,b} \sum_{i\in I_{\mathrm{train}}}(r_i-a-X_i'b)^2 +\lambda\|b\|_2^2. \]

The intercept is unpenalized and features are not standardized. A scalar penalty is used directly. With a grid, each outer training sample runs its own deterministic, unshuffled inner cross-validation with fold assignment based on row position and selects minimum MSE.

Combining held-out predictions gives \(\hat\ell_i\) and \(\hat m_i\). The reported coefficient solves the orthogonal residual-on-residual equation

\[ \hat\theta = \frac{\sum_i(D_i-\hat m_i)(Y_i-\hat\ell_i)} {\sum_i(D_i-\hat m_i)^2}. \]

The default is a fixed penalty of one rather than automatic tuning.

3 Inference

The score is

\[ \psi_i(\theta) =(D_i-\hat m_i) \{Y_i-\hat\ell_i-(D_i-\hat m_i)\theta\}, \]

with Jacobian \(-E_n[(D-\hat m)^2]\). Vanilla is the uncorrected iid empirical sandwich. HC1, Bartlett Newey-West, and cluster choices operate on the resulting scalar parameter scores with their usual package corrections. The nuisance fits and fold assignment are treated as fixed in the reported covariance; there is no repeated-cross-fit adjustment, bootstrap, or built-in Wald method.

4 Performance and numerical behavior

Each of \(K\) outer folds fits two ridge models. A penalty grid of size \(L\) with \(F\) inner folds increases this to roughly \(2KLF\) training solves plus the final per-fold nuisance fits. Every solve uses dense augmented least squares with up to \(O(np^2+p^3)\) work and no decomposition reuse. Cross-fitted prediction arrays require \(O(n)\) additional memory, while each fit stores dense fold designs. Near-zero residualized treatment variation raises. Results can vary with the outer seed and with input row order through inner CV.

5 Python API

Constructor: cm.PartiallyLinearDML

Use PartiallyLinearDML(penalty=None, cv=5, n_folds=5, seed=42), then fit(y, d, x). summary() reports the coefficient, robust standard error, covariance, and selected nuisance penalties by fold.

print(inspect.signature(cm.PartiallyLinearDML))
(penalty=None, cv=5, n_folds=5, seed=42)
cls = cm.PartiallyLinearDML
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
fit(self, /, y, d, x)
summary(self, /, vcov=None, lags=None, clusters=None)

6 Minimal example

rng = np.random.default_rng(13)
x = rng.normal(size=(400, 4))
d = 0.3 + x @ np.array([0.5, -0.4, 0.2, 0.1]) + rng.normal(scale=0.8, size=400)
y = 1.3 * d + x @ np.array([0.4, -0.2, 0.1, 0.3]) + rng.normal(scale=0.6, size=400)
model = cm.PartiallyLinearDML(penalty=np.logspace(-4, 1, 10), cv=3, n_folds=4, seed=1)
model.fit(y, d, x)
print(model.summary()['coef'])
print(model.summary()['outcome_penalties'][:2])
1.357518301172487
[2.7825594e+00 1.0000000e-04]

7 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

rng = np.random.default_rng(113)
x = rng.normal(size=(160, 4))
d = 0.3 + x @ np.array([0.5, -0.4, 0.2, 0.1]) + rng.normal(size=160) * 0.8
y = 1.3 * d + x @ np.array([0.4, -0.2, 0.1, 0.3]) + rng.normal(size=160) * 0.6
model = cm.PartiallyLinearDML(penalty=np.logspace(-4, 1, 6), cv=3, n_folds=4, seed=1)
model.fit(y, d, x)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
coef ()
se ()
vcov (1, 1)
outcome_penalties (4,)
treatment_penalties (4,)