from _api_doc_utils import *Logit
Binary logistic regression
1 Where it fits
Group: Regression
Logit models a binary outcome through
\[ \Pr(Y_i=1\mid X_i=x_i)=\Lambda(\alpha+x_i'\beta), \]
with optional L2 regularization controlled by alpha. predict(x) returns probabilities for label 1; predict_label(x) returns class labels.
2 Likelihood and solver
With \(\eta_i=\alpha_0+x_i'\beta\) and \(p_i=\{1+\exp(-\eta_i)\}^{-1}\), the delegated Linfa fit minimizes the unnormalized penalized negative log-likelihood
\[ Q(\alpha_0,\beta) = \sum_{i=1}^n \left[ \log\{1+\exp(\eta_i)\}-y_i\eta_i \right] +\frac{\lambda}{2}\|\beta\|_2^2, \]
where the constructor argument named alpha is \(\lambda\) and the intercept is unpenalized. Optimization uses L-BFGS with a More-Thuente line search and the requested gradient tolerance and iteration budget. Public coefficients are oriented so that the linear index and probability always refer to label 1, regardless of the internal class ordering.
3 Inference
Analytic inference is available only when \(\lambda=0\). Let \(\tilde X=[\mathbf1,X]\) and \(W=\operatorname{diag}\{\hat p_i(1-\hat p_i)\}\). The returned model-based covariance is
\[ \widehat V = (\tilde X'W\tilde X+10^{-8}I)^{-1}. \]
The \(10^{-8}\) diagonal ridge is numerical stabilization, not a reported statistical penalty. The implementation has no heteroskedastic or cluster-robust logit sandwich option. Wald tests use this Fisher covariance. The pairs bootstrap refits the same fixed penalty in each sample; a resample on which the binary fit cannot be estimated aborts the bootstrap.
4 Performance and numerical behavior
Each L-BFGS objective and gradient evaluation is \(O(np)\), with limited-memory state linear in \(p\). Computing the Fisher covariance forms a dense \(p\times p\) information matrix in \(O(np^2)\) and inverts it in \(O(p^3)\). Complete or near separation can drive coefficients to large magnitudes and make the information matrix nearly singular; the small diagonal ridge can make inversion succeed without resolving the inferential weakness. Bootstrap multiplies the full optimization cost by the number of draws.
5 Python API
Constructor: cm.Logit
Fit with integer labels using fit(x, y_int32). summary() returns intercept and coefficient estimates. Fisher-information standard errors are available only for the unpenalized fit (alpha=0); penalized fits are explicitly prediction-only. bootstrap() resamples observations and refits the classifier.
print(inspect.signature(cm.Logit))(alpha=0.0, max_iterations=100, gradient_tolerance=0.0001)
cls = cm.Logit
display(HTML(html_table(["Public method"], public_methods(cls))))| Public method |
|---|
bootstrap(self, /, n_bootstrap, seed=None) |
fit(self, /, x, y) |
predict(self, /, x) |
predict_label(self, /, x, cutoff=0.5) |
predict_lin(self, /, x) |
summary(self, /) |
wald_test(self, /, r, q=None) |
6 Minimal example
rng = np.random.default_rng(5)
x = rng.normal(size=(220, 3))
eta = -0.2 + x @ np.array([0.7, -0.4, 0.9])
p = 1 / (1 + np.exp(-eta))
y = rng.binomial(1, p, size=220).astype(np.int32)
model = cm.Logit(max_iterations=200)
model.fit(x, y)
print(model.summary()['coef'])
print(model.predict(x[:5]))[ 0.56568906 -0.58550069 0.85818987]
[0.5239276 0.4143503 0.68494357 0.42404209 0.2112137 ]
7 summary() contract
The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.
rng = np.random.default_rng(105)
x = rng.normal(size=(90, 3))
p = 1 / (1 + np.exp(-(x @ np.array([0.7, -0.4, 0.9]))))
y = rng.binomial(1, p, size=90).astype(np.int32)
model = cm.Logit(max_iterations=200)
model.fit(x, y)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))| summary() key | shape |
|---|---|
intercept |
() |
coef |
(3,) |
penalty |
() |
inference_available |
() |
intercept_se |
() |
coef_se |
(3,) |
vcov |
(4, 4) |