from _api_doc_utils import *HorizontalPanelRidge
Horizontal ridge counterfactuals for panel treatment effects
1 Where it fits
Group: Causal inference
HorizontalPanelRidge implements a horizontal panel-prediction design. For each adoption cohort, never-treated donor outcomes at time \(t\) become features for treated outcomes at time \(t\) in the pre-period. Ridge then extrapolates counterfactual treated paths into the treated post-period.
The public panel contract is fit(Y, W): balanced outcomes plus a same-shaped absorbing treatment matrix.
2 Cohort objective and estimand
For adoption cohort \(g\), let \(\bar Y_{g,t}\) be the mean outcome among units first treated at \(g\), and let \(Y_{C,t}\) be the vector of outcomes for never-treated units. The class fits the pre-treatment horizontal regression
\[ (\hat a_g,\hat\beta_g) = \arg\min_{a,\beta} \sum_{t<g} (\bar Y_{g,t}-a-Y_{C,t}'\beta)^2 +\lambda\|\beta\|_2^2. \]
The intercept is unpenalized and donor coefficients are unconstrained: they need not be positive or sum to one. The cohort counterfactual path is
\[ \hat Y_{g,t}(0)=\hat a_g+Y_{C,t}'\hat\beta_g, \]
and that same path is assigned to every treated unit in cohort \(g\). The overall ATT is the simple average of \(Y_{it}-\hat Y_{g,t}(0)\) over all treated unit-period cells, so cohorts receive weight proportional to treated units times post-treatment periods.
3 Implementation walkthrough
This estimator is a direct panel-specific construction rather than a wrapper around the public Ridge class.
- A shared panel validator checks equal nonempty shapes, finite \(Y\) and \(W\), binary treatment within \(10^{-10}\), and absorbing treatment paths. It records each unit’s first treated period, the sorted adoption cohorts, and ever- versus never-treated unit indices.
- For cohort \(g\), the code selects all cohort units and all never-treated controls. It averages cohort outcomes across units first, yielding one length-\(T\) target path. The donor panel remains one row per control unit.
- The pre-period design is the transpose of the donor panel restricted to columns \(0,\ldots,g-1\): periods are observations and donor units are regressors. An intercept column is prepended, \(X'X\) is formed, \(\lambda\) is added only to donor diagonals, and the code explicitly computes \((X'X+P_\lambda)^{-1}X'y\).
- The resulting donor coefficients are applied to the transpose of the full control panel, including post-treatment periods, and the intercept is added. This produces one cohort-level counterfactual path; no treated unit’s own pre-period outcomes enter as features after the cohort average has been formed.
- That identical path is copied into every treated unit in the cohort. The coefficient matrix is stored with one row per cohort and one column per original panel unit; non-control entries remain zero so indices line up with the input panel.
- ATT scans cells with \(W_{it}=1\) and averages finite observed-minus-counterfactual effects. Pre-RMSE first averages effects within each cohort-period before treatment and then takes the root mean square across those cohort-period means.
- Event-study output is derived after fitting. The unweighted version gives each cohort-event cell equal weight; the weighted version uses the number of finite treated-unit effects in that cell. These summaries do not refit the horizontal regression.
The cohort averaging is the defining modeling choice: it estimates a common untreated path for a cohort, not unit-specific paths. The small-system normal-equation inverse is compact but less stable than the augmented-QR implementation used by Ridge.
4 Inference and scope
The summary reports fitted cohort coefficients, counterfactuals, treatment-effect cells, ATT, pre-period RMSE, and event-time aggregates. It has no standard errors, covariance estimator, bootstrap, placebo procedure, or penalty tuning. All causal interpretation relies on never-treated outcomes spanning the untreated cohort path and on post-treatment donor outcomes remaining valid controls. Each treated cohort must have at least one pre-period, and at least one never-treated unit is required.
5 Performance and numerical behavior
With \(C\) never-treated donors, each cohort forms and explicitly inverts a dense \((C+1)\times(C+1)\) ridge system. Approximate work is \(O(gC^2+C^3+TC)\) per cohort, and coefficient storage includes a row across all panel units for every cohort. A positive penalty stabilizes donor collinearity but does not penalize the intercept. When \(C\) is large relative to the number of pre-periods, estimates can remain sensitive to scaling and the chosen penalty despite numerical invertibility.
6 Python API
Constructor: cm.HorizontalPanelRidge
After fit(y, w), predict() returns treated-unit counterfactuals, treatment_effect() returns observed-minus-counterfactual effects, and summary() returns ATT, event-study, group means, fitted coefficients, cohorts, and diagnostics.
print(inspect.signature(cm.HorizontalPanelRidge))(penalty=1.0)
cls = cm.HorizontalPanelRidge
display(HTML(html_table(["Public method"], public_methods(cls))))| Public method |
|---|
fit(self, /, y, w) |
predict(self, /) |
summary(self, /) |
treatment_effect(self, /) |
7 Minimal example
rng = np.random.default_rng(16)
y = rng.normal(size=(10, 14))
w = np.zeros_like(y)
w[7:, 9:] = 1
y[7:, 9:] += 1.0
model = cm.HorizontalPanelRidge(penalty=1.0)
model.fit(y, w)
print(model.summary()['att'])
print(list(model.summary()['event_study'].items())[:3])1.828237644527115
[('unweighted', {'event_time': array([-9., -8., -7., -6., -5., -4., -3., -2., -1., 0., 1., 2., 3.,
4.]), 'estimate': array([-0.08518528, -0.32173306, 0.01790597, 0.08711928, -0.16374416,
0.21465388, 0.35825572, 0.07763981, -0.18491216, 0.47756012,
2.27161272, 2.58629069, 2.36319848, 1.44252621]), 'n': array([1., 1., 1., 1., 1., 1., 1., 1., 1., 1., 1., 1., 1., 1.])}), ('weighted', {'event_time': array([-9., -8., -7., -6., -5., -4., -3., -2., -1., 0., 1., 2., 3.,
4.]), 'estimate': array([-0.08518528, -0.32173306, 0.01790597, 0.08711928, -0.16374416,
0.21465388, 0.35825572, 0.07763981, -0.18491216, 0.47756012,
2.27161272, 2.58629069, 2.36319848, 1.44252621]), 'n': array([3., 3., 3., 3., 3., 3., 3., 3., 3., 3., 3., 3., 3., 3.])})]
8 summary() contract
The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.
rng = np.random.default_rng(116)
y = rng.normal(size=(8, 12))
w = np.zeros_like(y)
w[6:, 8:] = 1
y[6:, 8:] += 0.8
model = cm.HorizontalPanelRidge()
model.fit(y, w)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))| summary() key | shape |
|---|---|
att |
() |
intercept |
() |
coef |
(8,) |
cohort_intercepts |
(1,) |
cohort_coef |
(1, 8) |
counterfactual |
(8, 12) |
treatment_effect |
(8, 12) |
event_study |
() |
group_means |
() |
pre_rmse |
() |
penalty |
() |
control_units |
(6,) |
treated_units |
(2,) |
cohorts |
(1,) |