from _api_doc_utils import *MultinomialLogit
Multiclass logistic regression
1 Where it fits
Group: Regression
MultinomialLogit generalizes binary logit to \(K\) classes with softmax probabilities:
\[ \Pr(Y_i=k\mid X_i=x_i)=\frac{\exp(\alpha_k+x_i'\beta_k)}{\sum_\ell \exp(\alpha_\ell+x_i'\beta_\ell)}. \]
The summary identifies the last sorted class as reference_class and reports identifiable class-versus-reference coefficient contrasts. Fisher-information standard errors are available only for alpha=0.
2 Likelihood, parameterization, and solver
For one-hot outcomes \(y_{ic}\) and logits \(\eta_{ic}=\alpha_c+x_i'\beta_c\), the delegated Linfa solver minimizes
\[ Q(\alpha,\beta) = -\sum_{i=1}^n\sum_{c=1}^C y_{ic}\log p_{ic} +\frac{\lambda}{2}\sum_{c=1}^C\|\beta_c\|_2^2, \qquad p_{ic}=\frac{\exp(\eta_{ic})}{\sum_\ell\exp(\eta_{i\ell})}. \]
Intercepts are unpenalized. L-BFGS with a More-Thuente line search operates on all \(C\) coefficient columns. This full softmax parameterization is redundant because adding a common coefficient vector to every class leaves probabilities unchanged. The public summary removes that redundancy after fitting by reporting
\[ \delta_c=(\alpha_c-\alpha_r,\ \beta_c-\beta_r'), \]
against the last sorted class \(r\). Prediction uses a row-wise max shift before exponentiation for stable softmax probabilities.
3 Inference
Inference is available only at \(\lambda=0\) and is computed directly in the \((C-1)\) reference-class contrast parameterization. For non-reference classes \(a,b\) and \(\tilde x_i=(1,x_i')'\), the information blocks are
\[ H_{ab} = \sum_i \hat p_{ia} \{\mathbf1(a=b)-\hat p_{ib}\} \tilde x_i\tilde x_i'. \]
The returned covariance is \((H+10^{-8}I)^{-1}\). It is model-based only; there is no robust sandwich or Wald-test method on this class. Penalized fits return no covariance. The pairs bootstrap reports class-versus-reference contrasts, but it raises an error if any resample omits an outcome class.
4 Performance and numerical behavior
An L-BFGS evaluation costs \(O(npC)\) and stores the dense \(n\times C\) probability matrix plus \((p+1)C\) parameters. Fisher inference has dimension \((p+1)(C-1)\); constructing its blocks is costly for many classes and dense inversion is cubic in that total dimension. The full fitted parameterization can be weakly identified when \(\lambda=0\) even though reported contrasts are identified. Rare classes also make both the Fisher matrix and pairs bootstrap fragile.
5 Python API
Constructor: cm.MultinomialLogit
Use integer class labels in fit(x, y_int32). predict(x) returns an \(n\times C\) probability matrix, predict_lin(x) returns logits, and predict_label(x) returns class labels. summary() returns contrast rows aligned with class_labels; penalized fits mark inference unavailable and omit se/vcov.
print(inspect.signature(cm.MultinomialLogit))(alpha=0.0, max_iterations=100, gradient_tolerance=0.0001)
cls = cm.MultinomialLogit
display(HTML(html_table(["Public method"], public_methods(cls))))| Public method |
|---|
bootstrap(self, /, n_bootstrap, seed=None) |
fit(self, /, x, y) |
predict(self, /, x) |
predict_label(self, /, x) |
predict_lin(self, /, x) |
summary(self, /) |
6 Minimal example
rng = np.random.default_rng(6)
x = rng.normal(size=(240, 2))
logits = x @ np.array([[0.6, -0.3], [-0.4, 0.5], [0.2, 0.2]]).T + np.array([0.1, -0.2, 0.0])
p = np.exp(logits - logits.max(axis=1, keepdims=True))
p = p / p.sum(axis=1, keepdims=True)
y = np.array([rng.choice(3, p=row) for row in p], dtype=np.int32)
model = cm.MultinomialLogit(max_iterations=200)
model.fit(x, y)
print(model.summary()['coef'])
print(model.predict(x[:5]))[[ 0.25831465 0.53237638 -0.55020945]
[-0.15395855 -0.36374717 0.31279086]]
[[0.29713333 0.35470927 0.3481574 ]
[0.10437553 0.60469433 0.29093014]
[0.35652526 0.30570265 0.33777209]
[0.27898794 0.37429093 0.34672113]
[0.36790269 0.30230945 0.32978787]]
7 summary() contract
The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.
rng = np.random.default_rng(106)
x = rng.normal(size=(100, 2))
logits = x @ np.array([[0.6, -0.3], [-0.4, 0.5], [0.2, 0.2]]).T
p = np.exp(logits - logits.max(1, keepdims=True))
p = p / p.sum(1, keepdims=True)
y = np.array([rng.choice(3, p=row) for row in p], dtype=np.int32)
model = cm.MultinomialLogit(max_iterations=200)
model.fit(x, y)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))| summary() key | shape |
|---|---|
coef |
(2, 3) |
class_labels |
(2,) |
reference_class |
() |
penalty |
() |
inference_available |
() |
se |
(2, 3) |
vcov |
(6, 6) |