from _api_doc_utils import *MatrixCompletion
Nuclear-norm panel counterfactual completion
1 Where it fits
Group: Causal inference
MatrixCompletion treats untreated cells as observed entries and treated cells as missing counterfactuals. It estimates a low-rank untreated-outcome surface, optionally with unit and time effects, using nuclear-norm style shrinkage.
The completed values in treated cells become counterfactual outcomes for ATT and event-study summaries.
2 Objective and algorithm
Entries with treatment indicator below \(0.5\) form the observed set \(\Omega\); treated entries are excluded from fitting. With optional unit effects \(a_i\), time effects \(b_t\), and low-rank matrix \(L\), the implemented objective is
\[ \min_{L,a,b} \frac{1}{|\Omega|} \sum_{(i,t)\in\Omega} (Y_{it}-a_i-b_t-L_{it})^2 +\lambda_L\|L\|_*. \]
The algorithm alternates observed-cell mean updates for \(a\) and \(b\) with singular-value thresholding. Its proximal step fills observed residuals into the current \(L\) and shrinks singular values by
\[ \tau=\frac{\lambda_L|\Omega|}{2}. \]
If \(\lambda_L\) is omitted, the class first fits the requested additive effects with \(L=0\) and sets
\[ \lambda_L = \text{lambda fraction}\times \frac{2s_{\max}\{P_\Omega(Y-a-b)\}}{|\Omega|}. \]
Exact SVD applies the full proximal map. Randomized SVD with an explicit rank is a truncated approximation to that map and therefore need not minimize exactly the displayed nuclear-norm objective.
3 Estimand and inference
The completed counterfactual surface is \(\hat Y_{it}(0)=\hat a_i+\hat b_t+\hat L_{it}\). Reported treated-cell effects are \(Y_{it}-\hat Y_{it}(0)\), and ATT is their simple mean over cells marked treated. Event-study and group summaries aggregate the same cell differences.
There is no analytic or resampling inference, rank selection, or penalty cross-validation. If \(Q_k\) denotes the displayed objective after iteration \(k\), convergence requires
\[ \frac{|Q_{k-1}-Q_k|}{|Q_{k-1}|+10^{-12}}<\text{tolerance}. \]
Unlike likelihood estimators in this package, MatrixCompletion retains the final iterate when the budget is exhausted because a partial completion can still be diagnostically useful. It does not call that outcome convergence: summary() reports converged=False, iterations, termination_reason='Maximum number of iterations reached', and the final objective, together with both histories. Identification requires untreated observations to reveal the relevant low-rank and additive structure; optimizer convergence cannot diagnose failure of that causal assumption.
4 Performance and numerical behavior
Each exact iteration performs a dense SVD with roughly \(O(\min\{NT^2,N^2T\})\) time and \(O(NT)\) matrix storage, plus observed-cell scans and effect updates. Randomized SVD can reduce factorization work to a chosen low rank but remains dense and changes the update. Total cost scales with the number of outer iterations and requested effect-update sweeps. Treated outcome values may be nonfinite because they are excluded from fitting, but every untreated entry must be finite. Iteration budgets, effect-update counts, and tolerances must be positive; invalid settings fail before fitting.
5 Python API
Constructor: cm.MatrixCompletion
Call MatrixCompletion(...).fit(y, w). predict() returns completed/counterfactual values and summary() reports ATT, completed matrices, treatment effects, low-rank components, singular values, objective and RMSE histories, fit diagnostics, and panel summaries. Check converged before treating the retained solution as numerically settled.
print(inspect.signature(cm.MatrixCompletion))(lambda_l=None, lambda_fraction=0.25, fit_unit_effects=True, fit_time_effects=True, max_iterations=500, effect_iterations=2, tolerance=1e-06, svd_method=Ellipsis, svd_rank=None, svd_oversamples=10, svd_power_iter=1, svd_seed=None)
cls = cm.MatrixCompletion
display(HTML(html_table(["Public method"], public_methods(cls))))| Public method |
|---|
fit(self, /, y, w) |
predict(self, /) |
summary(self, /) |
6 Minimal example
rng = np.random.default_rng(18)
load = rng.normal(size=(10, 2))
fac = rng.normal(size=(2, 14))
y = load @ fac + rng.normal(scale=0.1, size=(10, 14))
w = np.zeros_like(y)
w[7:, 9:] = 1
y[7:, 9:] += 1
model = cm.MatrixCompletion(max_iterations=100, tolerance=1e-05)
model.fit(y, w)
fit = model.summary()
print({key: fit[key] for key in ['converged', 'iterations', 'termination_reason', 'objective']})
print(fit['att'])
print(model.predict().shape){'converged': True, 'iterations': 6, 'termination_reason': 'Relative objective tolerance reached', 'objective': 0.78117030291554}
1.2016133829540447
(10, 14)
7 summary() contract
The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.
rng = np.random.default_rng(118)
y = rng.normal(size=(8, 10))
w = np.zeros_like(y)
w[6:, 7:] = 1
y[6:, 7:] += 0.8
model = cm.MatrixCompletion(max_iterations=80, tolerance=1e-05)
model.fit(y, w)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))| summary() key | shape |
|---|---|
completed |
(8, 10) |
low_rank |
(8, 10) |
unit_effects |
(8,) |
time_effects |
(10,) |
singular_values |
(8,) |
lambda_l |
() |
converged |
() |
iterations |
() |
termination_reason |
() |
objective |
() |
history_objective |
(16,) |
history_rmse |
(16,) |
svd_method |
() |
svd_rank |
() |
svd_oversamples |
() |
svd_power_iter |
() |
att |
() |
counterfactual |
(8, 10) |
treatment_effect |
(8, 10) |
event_study |
() |
group_means |
() |
control_units |
(6,) |
treated_units |
(2,) |
cohorts |
(1,) |