crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Binding Crash Course
  • Regression And GLMs
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • FTRL
    • MEstimator Poisson
  • Causal Inference
    • Balancing Weights
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
  • Transforms
    • PCA And Kernel Basis
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding: First Course
    • Overview And TOC
    • Ch 1 Correlation And Simpson
    • Ch 2 Potential Outcomes
    • Ch 3 CRE And Fisher RT
    • Ch 4 CRE And Neyman
    • Ch 9 Bridging Finite And Superpopulation
    • Ch 11 Propensity Score
    • Ch 12 Double Robust ATE
    • Ch 13 Double Robust ATT
    • Ch 21 Experimental IV
    • Ch 23 Econometric IV
    • Ch 27 Mediation

On this page

  • 1 Where it fits
  • 2 Objective and estimating equations
  • 3 Inference
  • 4 Performance and numerical behavior
  • 5 Python API
  • 6 Minimal example
  • 7 summary() contract

MEstimator

Low-level objective-plus-score M-estimation

from _api_doc_utils import *

1 Where it fits

Group: Estimation interfaces

MEstimator is the lowest-level public estimation interface. It minimizes a user-supplied objective with gradient and uses a user-supplied per-observation score matrix for covariance estimation:

\[ \hat\theta = \arg\min_\theta Q_n(\theta), \qquad \widehat V = A^{-1} B A^{-T} / n. \]

The bread \(A\) is the numerical Jacobian of the mean score and the meat \(B\) is the empirical score outer product. The class checks optimizer convergence and validates score dimensions before reporting inference.

2 Objective and estimating equations

The objective callback supplies both \(Q_n(\theta)\) and \(\nabla Q_n(\theta)\). The class minimizes that objective with seven-vector-memory L-BFGS, a More-Thuente line search, and joint gradient and cost tolerances. A non-converged optimizer result raises rather than storing a fit.

Separately, the score callback must return rows \(\psi_i(\theta)'\in\mathbb R^k\). Inference treats the fitted parameter as the solution to

\[ \frac1n\sum_{i=1}^n\psi_i(\theta)=0. \]

The class does not verify that the objective gradient equals the sum of those scores or even that their sample mean is near zero at the optimizer solution. That mathematical consistency is the callback author’s responsibility.

3 Inference

For coordinate \(j\), the mean-score Jacobian is estimated by central differences with

\[ h_j=\text{derivative step}\times\max\{|\hat\theta_j|,1\}, \]

\[ \hat A_{\cdot j} = \frac{ \bar\psi(\hat\theta+h_je_j) -\bar\psi(\hat\theta-h_je_j)} {2h_j}, \qquad \hat B=\frac1n\sum_i\psi_i(\hat\theta)\psi_i(\hat\theta)'. \]

The covariance is symmetrized after computing

\[ \widehat V = \frac1n\hat A^{-1}\hat B\hat A^{-T}. \]

There are no HC, HAC, or cluster variants and no built-in Wald method. The pairs bootstrap is available only when data is a dictionary containing the sample size under the key \(n\) and the objective callback honors the injected row-index vector. Bootstrap fits start from the original estimate and abort on the first failed optimization.

4 Performance and numerical behavior

Every optimizer cost and gradient request crosses the Python-Rust boundary and calls the objective callback. Covariance requires the score at the fit plus two additional full score evaluations per parameter, so callback work is at least \(2k+1\) score matrices of shape \(n\times k\). It then stores and factors dense \(k\times k\) matrices. Numerical inference is sensitive to the derivative step and fails if \(\hat A\) is singular. Cached covariance avoids repeating this work on later summaries, but bootstrap performs one full callback-driven optimization per draw.

5 Python API

Constructor: cm.MEstimator

Construct with MEstimator(objective_fn, score_fn, max_iterations=100, tolerance=1e-6, derivative_step=1e-6). objective_fn(theta, data) must return (objective, gradient). score_fn(theta, data) must return an (n, p) matrix. For bootstrap support, include n in data and have the objective respect optional data['indices'].

print(inspect.signature(cm.MEstimator))
(objective_fn, score_fn, max_iterations=100, tolerance=1e-06, derivative_step=1e-06)
cls = cm.MEstimator
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
bootstrap(self, /, n_bootstrap, seed=None)
compute_vcov(self, /)
fit(self, /, data, theta0)
summary(self, /)

6 Minimal example

def obj(theta, data):
    X, y = (data['X'], data['y'])
    idx = data.get('indices', np.arange(len(y)))
    r = y[idx] - X[idx] @ theta
    return (0.5 * np.sum(r * r), -(X[idx].T @ r))
def score(theta, data):
    r = data['y'] - data['X'] @ theta
    return -data['X'] * r[:, None]
rng = np.random.default_rng(21)
X = rng.normal(size=(180, 2))
y = X @ np.array([1.0, -0.5]) + rng.normal(scale=0.2, size=180)
model = cm.MEstimator(obj, score, max_iterations=200)
model.fit({'X': X, 'y': y, 'n': len(y)}, np.zeros(2))
print(model.summary())
{'coef': array([ 1.03140015, -0.50873558]), 'se': array([0.01556896, 0.01499416]), 'vcov': array([[ 2.42392436e-04, -6.28291233e-05],
       [-6.28291233e-05,  2.24824902e-04]]), 'converged': True, 'iterations': 3}

7 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

def obj(theta, data):
    X, y = (data['X'], data['y'])
    idx = data.get('indices', np.arange(len(y)))
    r = y[idx] - X[idx] @ theta
    return (0.5 * np.sum(r * r), -(X[idx].T @ r))
def score(theta, data):
    r = data['y'] - data['X'] @ theta
    return -data['X'] * r[:, None]
rng = np.random.default_rng(121)
X = rng.normal(size=(100, 2))
y = X @ np.array([1, -0.5]) + rng.normal(size=100) * 0.2
model = cm.MEstimator(obj, score, max_iterations=200)
model.fit({'X': X, 'y': y, 'n': len(y)}, np.zeros(2))
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
coef (2,)
se (2,)
vcov (2, 2)
converged ()
iterations ()