from _api_doc_utils import *PartiallyLinearDML
Cross-fit partially linear Double ML
1 Where it fits
Group: Causal inference
PartiallyLinearDML estimates the treatment coefficient in
\[ y = \theta d + g(x) + u, \qquad d = m(x) + v, \]
using cross-fitted ridge nuisance regressions. The final coefficient is estimated from the orthogonalized residual-on-residual score. Fold membership is a deterministic seeded shuffle, so equal seeds reproduce assignments and different seeds change them.
2 Cross-fitted objective
For the partially linear model
\[ Y=\theta D+\ell(X)+U, \qquad D=m(X)+V, \]
the class creates seeded outer folds. On the complement of each fold it separately fits ridge nuisance regressions for \(Y\) and \(D\):
\[ \min_{a,b} \sum_{i\in I_{\mathrm{train}}}(r_i-a-X_i'b)^2 +\lambda\|b\|_2^2. \]
The intercept is unpenalized and features are not standardized. A scalar penalty is used directly. With a grid, each outer training sample runs its own deterministic, unshuffled inner cross-validation with fold assignment based on row position and selects minimum MSE.
Combining held-out predictions gives \(\hat\ell_i\) and \(\hat m_i\). The reported coefficient solves the orthogonal residual-on-residual equation
\[ \hat\theta = \frac{\sum_i(D_i-\hat m_i)(Y_i-\hat\ell_i)} {\sum_i(D_i-\hat m_i)^2}. \]
The default is a fixed penalty of one rather than automatic tuning.
3 Implementation walkthrough
Cross-fitting, ridge tuning, orthogonalization, and covariance are all implemented inside the class.
- Outer folds are assigned by sorting row indices with a deterministic SplitMix64 hash of
seed ^ index, then cycling the sorted positions across folds. This is reproducible without storing RNG state and changes if either row indices or seed change. - In each outer complement, outcome and treatment nuisance models are tuned independently. Ridge prepends an intercept and appends \(\sqrt\lambda I\) rows for slopes before a least-squares solve. With one penalty, or fewer than four training rows, inner CV is bypassed.
- Inner CV uses the current outer-training row order and assigns local row \(i\) to fold \(i\bmod F\) without shuffling. It performs a fresh augmented least-squares fit for every penalty-fold pair, chooses the first strict MSE minimum, and refits that nuisance on the complete outer-training sample.
- Only held-out predictions are written into
l_hatandm_hat; no observation receives a nuisance prediction from a model trained on itself. The selected penalty for each of the two nuisances is retained by outer fold. - After all folds, the code forms \(\tilde D=D-\hat m\) and \(\tilde Y=Y-\hat\ell\), rejects \(\tilde D'\tilde D<10^{-12}\), and computes one scalar residual-on-residual ratio. There is no final full-sample nuisance refit because it would not enter the orthogonal estimate.
- Summary constructs the scalar score row \(\tilde D_i(\tilde Y_i-\tilde D_i\hat\theta)\) and its analytic sample Jacobian \(-\tilde D'\tilde D/n\). It then uses the common exactly identified score covariance code for iid, HC1, HAC, or clusters.
This is deliberately a transparent ridge-only DML implementation. It demonstrates orthogonalization and out-of-fold bookkeeping without a generic learner protocol, repeated splits, fold-specific preprocessing, or tuning uncertainty propagation.
4 Inference
The score is
\[ \psi_i(\theta) =(D_i-\hat m_i) \{Y_i-\hat\ell_i-(D_i-\hat m_i)\theta\}, \]
with Jacobian \(-E_n[(D-\hat m)^2]\). Vanilla is the uncorrected iid empirical sandwich. HC1, Bartlett Newey-West, and cluster choices operate on the resulting scalar parameter scores with their usual package corrections. The nuisance fits and fold assignment are treated as fixed in the reported covariance; there is no repeated-cross-fit adjustment, bootstrap, or built-in Wald method.
5 Performance and numerical behavior
Each of \(K\) outer folds fits two ridge models. A penalty grid of size \(L\) with \(F\) inner folds increases this to roughly \(2KLF\) training solves plus the final per-fold nuisance fits. Every solve uses dense augmented least squares with up to \(O(np^2+p^3)\) work and no decomposition reuse. Cross-fitted prediction arrays require \(O(n)\) additional memory, while each fit stores dense fold designs. Near-zero residualized treatment variation raises. Results can vary with the outer seed and with input row order through inner CV.
6 Python API
Constructor: cm.PartiallyLinearDML
Use PartiallyLinearDML(penalty=None, cv=5, n_folds=5, seed=42), then fit(y, d, x). summary() reports the coefficient, robust standard error, covariance, and selected nuisance penalties by fold.
print(inspect.signature(cm.PartiallyLinearDML))(penalty=None, cv=5, n_folds=5, seed=42)
cls = cm.PartiallyLinearDML
display(HTML(html_table(["Public method"], public_methods(cls))))| Public method |
|---|
fit(self, /, y, d, x) |
summary(self, /, vcov=None, lags=None, clusters=None) |
7 Minimal example
rng = np.random.default_rng(13)
x = rng.normal(size=(400, 4))
d = 0.3 + x @ np.array([0.5, -0.4, 0.2, 0.1]) + rng.normal(scale=0.8, size=400)
y = 1.3 * d + x @ np.array([0.4, -0.2, 0.1, 0.3]) + rng.normal(scale=0.6, size=400)
model = cm.PartiallyLinearDML(penalty=np.logspace(-4, 1, 10), cv=3, n_folds=4, seed=1)
model.fit(y, d, x)
print(model.summary()['coef'])
print(model.summary()['outcome_penalties'][:2])1.357518301172487
[2.7825594e+00 1.0000000e-04]
8 summary() contract
The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.
rng = np.random.default_rng(113)
x = rng.normal(size=(160, 4))
d = 0.3 + x @ np.array([0.5, -0.4, 0.2, 0.1]) + rng.normal(size=160) * 0.8
y = 1.3 * d + x @ np.array([0.4, -0.2, 0.1, 0.3]) + rng.normal(size=160) * 0.6
model = cm.PartiallyLinearDML(penalty=np.logspace(-4, 1, 6), cv=3, n_folds=4, seed=1)
model.fit(y, d, x)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))| summary() key | shape |
|---|---|
coef |
() |
se |
() |
vcov |
(1, 1) |
outcome_penalties |
(4,) |
treatment_penalties |
(4,) |