crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Binding Crash Course
  • Regression And GLMs
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • FTRL
    • MEstimator Poisson
  • Causal Inference
    • Balancing Weights
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
  • Transforms
    • PCA And Kernel Basis
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding: First Course
    • Overview And TOC
    • Ch 1 Correlation And Simpson
    • Ch 2 Potential Outcomes
    • Ch 3 CRE And Fisher RT
    • Ch 4 CRE And Neyman
    • Ch 9 Bridging Finite And Superpopulation
    • Ch 11 Propensity Score
    • Ch 12 Double Robust ATE
    • Ch 13 Double Robust ATT
    • Ch 21 Experimental IV
    • Ch 23 Econometric IV
    • Ch 27 Mediation

On this page

  • 1 Where it fits
  • 2 Likelihood and solver
  • 3 Inference
  • 4 Performance and numerical behavior
  • 5 Python API
  • 6 Minimal example
  • 7 summary() contract

Logit

Binary logistic regression

from _api_doc_utils import *

1 Where it fits

Group: Regression

Logit models a binary outcome through

\[ \Pr(Y_i=1\mid X_i=x_i)=\Lambda(\alpha+x_i'\beta), \]

with optional L2 regularization controlled by alpha. predict(x) returns probabilities for label 1; predict_label(x) returns class labels.

2 Likelihood and solver

With \(\eta_i=\alpha_0+x_i'\beta\) and \(p_i=\{1+\exp(-\eta_i)\}^{-1}\), the delegated Linfa fit minimizes the unnormalized penalized negative log-likelihood

\[ Q(\alpha_0,\beta) = \sum_{i=1}^n \left[ \log\{1+\exp(\eta_i)\}-y_i\eta_i \right] +\frac{\lambda}{2}\|\beta\|_2^2, \]

where the constructor argument named alpha is \(\lambda\) and the intercept is unpenalized. Optimization uses L-BFGS with a More-Thuente line search and the requested gradient tolerance and iteration budget. Public coefficients are oriented so that the linear index and probability always refer to label 1, regardless of the internal class ordering.

3 Inference

Analytic inference is available only when \(\lambda=0\). Let \(\tilde X=[\mathbf1,X]\) and \(W=\operatorname{diag}\{\hat p_i(1-\hat p_i)\}\). The returned model-based covariance is

\[ \widehat V = (\tilde X'W\tilde X+10^{-8}I)^{-1}. \]

The \(10^{-8}\) diagonal ridge is numerical stabilization, not a reported statistical penalty. The implementation has no heteroskedastic or cluster-robust logit sandwich option. Wald tests use this Fisher covariance. The pairs bootstrap refits the same fixed penalty in each sample; a resample on which the binary fit cannot be estimated aborts the bootstrap.

4 Performance and numerical behavior

Each L-BFGS objective and gradient evaluation is \(O(np)\), with limited-memory state linear in \(p\). Computing the Fisher covariance forms a dense \(p\times p\) information matrix in \(O(np^2)\) and inverts it in \(O(p^3)\). Complete or near separation can drive coefficients to large magnitudes and make the information matrix nearly singular; the small diagonal ridge can make inversion succeed without resolving the inferential weakness. Bootstrap multiplies the full optimization cost by the number of draws.

5 Python API

Constructor: cm.Logit

Fit with integer labels using fit(x, y_int32). summary() returns intercept and coefficient estimates. Fisher-information standard errors are available only for the unpenalized fit (alpha=0); penalized fits are explicitly prediction-only. bootstrap() resamples observations and refits the classifier.

print(inspect.signature(cm.Logit))
(alpha=0.0, max_iterations=100, gradient_tolerance=0.0001)
cls = cm.Logit
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
bootstrap(self, /, n_bootstrap, seed=None)
fit(self, /, x, y)
predict(self, /, x)
predict_label(self, /, x, cutoff=0.5)
predict_lin(self, /, x)
summary(self, /)
wald_test(self, /, r, q=None)

6 Minimal example

rng = np.random.default_rng(5)
x = rng.normal(size=(220, 3))
eta = -0.2 + x @ np.array([0.7, -0.4, 0.9])
p = 1 / (1 + np.exp(-eta))
y = rng.binomial(1, p, size=220).astype(np.int32)
model = cm.Logit(max_iterations=200)
model.fit(x, y)
print(model.summary()['coef'])
print(model.predict(x[:5]))
[ 0.56568906 -0.58550069  0.85818987]
[0.5239276  0.4143503  0.68494357 0.42404209 0.2112137 ]

7 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

rng = np.random.default_rng(105)
x = rng.normal(size=(90, 3))
p = 1 / (1 + np.exp(-(x @ np.array([0.7, -0.4, 0.9]))))
y = rng.binomial(1, p, size=90).astype(np.int32)
model = cm.Logit(max_iterations=200)
model.fit(x, y)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
intercept ()
coef (3,)
penalty ()
inference_available ()
intercept_se ()
coef_se (3,)
vcov (4, 4)