crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Binding Crash Course
  • Regression And GLMs
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • FTRL
    • MEstimator Poisson
  • Causal Inference
    • Balancing Weights
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
  • Transforms
    • PCA And Kernel Basis
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding: First Course
    • Overview And TOC
    • Ch 1 Correlation And Simpson
    • Ch 2 Potential Outcomes
    • Ch 3 CRE And Fisher RT
    • Ch 4 CRE And Neyman
    • Ch 9 Bridging Finite And Superpopulation
    • Ch 11 Propensity Score
    • Ch 12 Double Robust ATE
    • Ch 13 Double Robust ATT
    • Ch 21 Experimental IV
    • Ch 23 Econometric IV
    • Ch 27 Mediation

On this page

  • 1 Where it fits
  • 2 Objective and algorithm
  • 3 Estimand and inference
  • 4 Performance and numerical behavior
  • 5 Python API
  • 6 Minimal example
  • 7 summary() contract

MatrixCompletion

Nuclear-norm panel counterfactual completion

from _api_doc_utils import *

1 Where it fits

Group: Causal inference

MatrixCompletion treats untreated cells as observed entries and treated cells as missing counterfactuals. It estimates a low-rank untreated-outcome surface, optionally with unit and time effects, using nuclear-norm style shrinkage.

The completed values in treated cells become counterfactual outcomes for ATT and event-study summaries.

2 Objective and algorithm

Entries with treatment indicator below \(0.5\) form the observed set \(\Omega\); treated entries are excluded from fitting. With optional unit effects \(a_i\), time effects \(b_t\), and low-rank matrix \(L\), the implemented objective is

\[ \min_{L,a,b} \frac{1}{|\Omega|} \sum_{(i,t)\in\Omega} (Y_{it}-a_i-b_t-L_{it})^2 +\lambda_L\|L\|_*. \]

The algorithm alternates observed-cell mean updates for \(a\) and \(b\) with singular-value thresholding. Its proximal step fills observed residuals into the current \(L\) and shrinks singular values by

\[ \tau=\frac{\lambda_L|\Omega|}{2}. \]

If \(\lambda_L\) is omitted, the class first fits the requested additive effects with \(L=0\) and sets

\[ \lambda_L = \text{lambda fraction}\times \frac{2s_{\max}\{P_\Omega(Y-a-b)\}}{|\Omega|}. \]

Exact SVD applies the full proximal map. Randomized SVD with an explicit rank is a truncated approximation to that map and therefore need not minimize exactly the displayed nuclear-norm objective.

3 Estimand and inference

The completed counterfactual surface is \(\hat Y_{it}(0)=\hat a_i+\hat b_t+\hat L_{it}\). Reported treated-cell effects are \(Y_{it}-\hat Y_{it}(0)\), and ATT is their simple mean over cells marked treated. Event-study and group summaries aggregate the same cell differences.

There is no analytic or resampling inference, rank selection, or penalty cross-validation. Iteration stops when the relative objective change falls below tolerance. Reaching the iteration budget still stores a result, and the summary reports iterations and histories but no convergence flag, so users should inspect the objective history. Identification requires untreated observations to reveal the relevant low-rank and additive structure; the optimizer cannot diagnose failure of that causal assumption.

4 Performance and numerical behavior

Each exact iteration performs a dense SVD with roughly \(O(\min\{NT^2,N^2T\})\) time and \(O(NT)\) matrix storage, plus observed-cell scans and effect updates. Randomized SVD can reduce factorization work to a chosen low rank but remains dense and changes the update. Total cost scales with the number of outer iterations and requested effect-update sweeps. Treated outcome values may be nonfinite because they are excluded from fitting, but every untreated entry must be finite.

5 Python API

Constructor: cm.MatrixCompletion

Call MatrixCompletion(...).fit(y, w). predict() returns completed/counterfactual values and summary() reports ATT, completed matrices, treatment effects, low-rank components, singular values, objective history, and panel summaries.

print(inspect.signature(cm.MatrixCompletion))
(lambda_l=None, lambda_fraction=0.25, fit_unit_effects=True, fit_time_effects=True, max_iterations=500, effect_iterations=2, tolerance=1e-06, svd_method=Ellipsis, svd_rank=None, svd_oversamples=10, svd_power_iter=1, svd_seed=None)
cls = cm.MatrixCompletion
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
fit(self, /, y, w)
predict(self, /)
summary(self, /)

6 Minimal example

rng = np.random.default_rng(18)
load = rng.normal(size=(10, 2))
fac = rng.normal(size=(2, 14))
y = load @ fac + rng.normal(scale=0.1, size=(10, 14))
w = np.zeros_like(y)
w[7:, 9:] = 1
y[7:, 9:] += 1
model = cm.MatrixCompletion(max_iterations=100, tolerance=1e-05)
model.fit(y, w)
print(model.summary()['att'])
print(model.predict().shape)
1.2016133829540447
(10, 14)

7 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

rng = np.random.default_rng(118)
y = rng.normal(size=(8, 10))
w = np.zeros_like(y)
w[6:, 7:] = 1
y[6:, 7:] += 0.8
model = cm.MatrixCompletion(max_iterations=80, tolerance=1e-05)
model.fit(y, w)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
completed (8, 10)
low_rank (8, 10)
unit_effects (8,)
time_effects (10,)
singular_values (8,)
lambda_l ()
objective ()
iterations ()
history_objective (16,)
history_rmse (16,)
svd_method ()
svd_rank ()
svd_oversamples ()
svd_power_iter ()
att ()
counterfactual (8, 10)
treatment_effect (8, 10)
event_study ()
group_means ()
control_units (6,)
treated_units (2,)
cohorts (1,)