from _api_doc_utils import *MEstimator
Low-level objective-plus-score M-estimation
1 Where it fits
Group: Estimation interfaces
MEstimator is the lowest-level public estimation interface. It minimizes a user-supplied objective with gradient and uses a user-supplied per-observation score matrix for covariance estimation:
\[ \hat\theta = \arg\min_\theta Q_n(\theta), \qquad \widehat V = A^{-1} B A^{-T} / n. \]
The bread \(A\) is the numerical Jacobian of the mean score and the meat \(B\) is the empirical score outer product. The class checks optimizer convergence and validates score dimensions before reporting inference.
2 Objective and estimating equations
The objective callback supplies both \(Q_n(\theta)\) and \(\nabla Q_n(\theta)\). The class minimizes that objective with seven-vector-memory L-BFGS, a More-Thuente line search, and joint gradient and cost tolerances. Only an Argmin solver-converged or target-cost termination is accepted. Budget exhaustion or another nonconverged status raises ValueError and clears any previous fitted state. A successful summary exposes converged, iterations, termination_reason, and the callback’s final objective.
Separately, the score callback must return rows \(\psi_i(\theta)'\in\mathbb R^k\). Inference treats the fitted parameter as the solution to
\[ \frac1n\sum_{i=1}^n\psi_i(\theta)=0. \]
The class does not verify that the objective gradient equals the sum of those scores or even that their sample mean is near zero at the optimizer solution. That mathematical consistency is the callback author’s responsibility.
3 Implementation walkthrough
This interface combines an Argmin optimizer with a separately implemented score-sandwich path.
fit()clears fitted state, validates the initial vector and controls, and stores Python callable objects inside an Argmin problem. Each cost request callsobjective_fn(theta, data), requires a two-element tuple, and extracts only its scalar first element.- Each gradient request calls the same Python function again and extracts only the second tuple element as a one-dimensional NumPy array. Cost and gradient are therefore separate Python calls even if the optimizer asks for both at the same parameter; there is no shared callback cache or explicit gradient-length check before Argmin uses it.
- Seven-pair L-BFGS starts exactly at
theta0, uses More-Thuente line search, and applies the same tolerance to gradient and cost stopping. The wrapper converts Argmin termination to common diagnostics and rejects an iteration cap or other nonconverged status before storing parameters and data. - Covariance is lazy: the first
summary()callsscore_fnat the estimate and requires a nonempty \(n\times k\) matrix with \(k=\dim\theta\). It then calls the score function twice for each coordinate perturbation and requires every perturbed output to retain the original shape. - Column means of those perturbed score matrices yield the central-difference bread \(A\). The meat is the uncentered \(\Psi'\Psi/n\). The code explicitly inverts \(A\), computes \(A^{-1}BA^{-T}/n\), and averages the result with its transpose to remove numerical asymmetry. Later summaries reuse this cached covariance.
- Bootstrap reads
data['n'], draws row indices, shallow-copies the Python dictionary, inserts anindicesarray, and calls the same L-BFGS problem from the original fitted estimate. The score callback is not used in bootstrap fitting. Correct resampling therefore depends entirely onobjective_fnhonoring that injected key.
The split callback design is flexible but intentionally low level. It permits objective and score equations that disagree, repeats Python work between cost and gradient, and leaves all data-index semantics to the caller; those are API powers and failure modes, not hidden implementation details.
4 Inference
For coordinate \(j\), the mean-score Jacobian is estimated by central differences with
\[ h_j=\text{derivative step}\times\max\{|\hat\theta_j|,1\}, \]
\[ \hat A_{\cdot j} = \frac{ \bar\psi(\hat\theta+h_je_j) -\bar\psi(\hat\theta-h_je_j)} {2h_j}, \qquad \hat B=\frac1n\sum_i\psi_i(\hat\theta)\psi_i(\hat\theta)'. \]
The covariance is symmetrized after computing
\[ \widehat V = \frac1n\hat A^{-1}\hat B\hat A^{-T}. \]
There are no HC, HAC, or cluster variants and no built-in Wald method. The pairs bootstrap is available only when data is a dictionary containing the sample size under the key \(n\) and the objective callback honors the injected row-index vector. Bootstrap fits start from the original estimate and abort on the first failed optimization.
5 Performance and numerical behavior
Every optimizer cost and gradient request crosses the Python-Rust boundary and calls the objective callback. Covariance requires the score at the fit plus two additional full score evaluations per parameter, so callback work is at least \(2k+1\) score matrices of shape \(n\times k\). It then stores and factors dense \(k\times k\) matrices. Numerical inference is sensitive to the derivative step and fails if \(\hat A\) is singular. Cached covariance avoids repeating this work on later summaries, but bootstrap performs one full callback-driven optimization per draw and aborts on any nonconverged replicate.
6 Python API
Constructor: cm.MEstimator
Construct with MEstimator(objective_fn, score_fn, max_iterations=100, tolerance=1e-6, derivative_step=1e-6). objective_fn(theta, data) must return (objective, gradient). score_fn(theta, data) must return an \((n,k)\) matrix when \(\theta\in\mathbb R^k\). For bootstrap support, include n in data and have the objective respect optional data['indices'].
print(inspect.signature(cm.MEstimator))(objective_fn, score_fn, max_iterations=100, tolerance=1e-06, derivative_step=1e-06)
cls = cm.MEstimator
display(HTML(html_table(["Public method"], public_methods(cls))))| Public method |
|---|
bootstrap(self, /, n_bootstrap, seed=None) |
compute_vcov(self, /) |
fit(self, /, data, theta0) |
summary(self, /) |
7 Minimal example
def obj(theta, data):
X, y = (data['X'], data['y'])
idx = data.get('indices', np.arange(len(y)))
r = y[idx] - X[idx] @ theta
return (0.5 * np.sum(r * r), -(X[idx].T @ r))
def score(theta, data):
r = data['y'] - data['X'] @ theta
return -data['X'] * r[:, None]
rng = np.random.default_rng(21)
X = rng.normal(size=(180, 2))
y = X @ np.array([1.0, -0.5]) + rng.normal(scale=0.2, size=180)
model = cm.MEstimator(obj, score, max_iterations=200)
model.fit({'X': X, 'y': y, 'n': len(y)}, np.zeros(2))
fit = model.summary()
print({key: fit[key] for key in ['converged', 'iterations', 'termination_reason', 'objective']})
print(fit['coef']){'converged': True, 'iterations': 3, 'termination_reason': 'Solver converged', 'objective': 3.5424400071434805}
[ 1.03140015 -0.50873558]
8 summary() contract
The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.
def obj(theta, data):
X, y = (data['X'], data['y'])
idx = data.get('indices', np.arange(len(y)))
r = y[idx] - X[idx] @ theta
return (0.5 * np.sum(r * r), -(X[idx].T @ r))
def score(theta, data):
r = data['y'] - data['X'] @ theta
return -data['X'] * r[:, None]
rng = np.random.default_rng(121)
X = rng.normal(size=(100, 2))
y = X @ np.array([1, -0.5]) + rng.normal(size=100) * 0.2
model = cm.MEstimator(obj, score, max_iterations=200)
model.fit({'X': X, 'y': y, 'n': len(y)}, np.zeros(2))
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))| summary() key | shape |
|---|---|
coef |
(2,) |
se |
(2,) |
vcov |
(2, 2) |
converged |
() |
iterations |
() |
termination_reason |
() |
objective |
() |