crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Binding Crash Course
  • Regression And GLMs
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • MEstimator Poisson
  • Causal Inference
    • Balancing Weights
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
  • Transforms
    • PCA And Kernel Basis
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding: First Course
    • Overview And TOC
    • Ch 1 Correlation And Simpson
    • Ch 2 Potential Outcomes
    • Ch 3 CRE And Fisher RT
    • Ch 4 CRE And Neyman
    • Ch 9 Bridging Finite And Superpopulation
    • Ch 11 Propensity Score
    • Ch 12 Double Robust ATE
    • Ch 13 Double Robust ATT
    • Ch 21 Experimental IV
    • Ch 23 Econometric IV
    • Ch 27 Mediation

On this page

  • 1 Where it fits
  • 2 Likelihood, parameterization, and solver
  • 3 Implementation walkthrough
  • 4 Inference
  • 5 Performance and numerical behavior
  • 6 Python API
  • 7 Minimal example
  • 8 summary() contract

MultinomialLogit

Multiclass logistic regression

from _api_doc_utils import *

1 Where it fits

Group: Regression

MultinomialLogit generalizes binary logit to \(K\) classes with softmax probabilities:

\[ \Pr(Y_i=k\mid X_i=x_i)=\frac{\exp(\alpha_k+x_i'\beta_k)}{\sum_\ell \exp(\alpha_\ell+x_i'\beta_\ell)}. \]

The summary identifies the last sorted class as reference_class and reports identifiable class-versus-reference coefficient contrasts. Fisher-information standard errors are available only for alpha=0.

2 Likelihood, parameterization, and solver

For one-hot outcomes \(y_{ic}\) and logits \(\eta_{ic}=\alpha_c+x_i'\beta_c\), the native solver minimizes

\[ Q(\alpha,\beta) = -\sum_{i=1}^n\sum_{c=1}^C y_{ic}\log p_{ic} +\frac{\lambda}{2}\sum_{c=1}^C\|\beta_c\|_2^2, \qquad p_{ic}=\frac{\exp(\eta_{ic})}{\sum_\ell\exp(\eta_{i\ell})}. \]

Intercepts are unpenalized. The implementation uses a row-wise log-sum-exp shift for the objective and softmax gradient, then runs ten-vector-memory L-BFGS with a More-Thuente line search from an all-zero parameter vector. L-BFGS operates on all \(C\) coefficient columns. This full softmax parameterization is redundant because adding a common coefficient vector to every class leaves probabilities unchanged. With \(\lambda>0\), the common slope direction is pinned down by the penalty, while the common intercept direction remains flat; fitted probabilities and reported contrasts are invariant to it. The public summary removes the redundancy by reporting

\[ \delta_c=(\alpha_c-\alpha_r,\ \beta_c-\beta_r'), \]

against the last sorted class \(r\). Reaching max_iterations is a failed fit: fit() raises ValueError and clears any previous fitted state. Successful summaries include converged, iterations, termination_reason, and the final penalized objective.

3 Implementation walkthrough

The package owns the complete softmax likelihood and optimization bridge. It does not call a Linfa multinomial model.

  1. fit() validates the dense input, sorts and deduplicates the original integer labels, and maps each response to its position in that sorted class vector. The original labels are retained so prediction can map an argmax back to the caller’s label space.
  2. The optimizer uses one contiguous block per class. With an intercept, a block is \((\alpha_c,\beta_c')\); without one it is just \(\beta_c\). All \(C(p+1)\) entries start at zero. Consequently, the initial softmax is uniform and the initial vector lies in the centered representative of the redundant full-\(C\) parameterization.
  3. For each row, the objective callback builds only a length-\(C\) logit scratch vector. It subtracts the largest logit before exponentiating, computes the row log-sum-exp, and adds log_denom - observed_logit. It then adds the L2 norm of every class’s slope block, leaving every intercept unpenalized.
  4. The gradient makes a second row-wise pass. It normalizes the shifted exponentials into probabilities and adds \((p_{ic}-1\{y_i=c\})\) to the class intercept score and \(x_i(p_{ic}-1\{y_i=c\})\) to that class’s slope score. The penalty contributes \(\lambda\beta_c\). No \(n\times C\) probability array is retained during fitting.
  5. Ten-pair L-BFGS with a More-Thuente line search updates the full vector. The wrapper treats the iteration cap as failure, takes Argmin’s best vector only after accepted termination, reshapes it into a \(p\times C\) coefficient matrix plus a length-\(C\) intercept vector, and atomically installs the fitted state.
  6. Prediction computes the dense \(n_{\mathrm{new}}\times C\) logit matrix. Probabilities are row-wise stable softmaxes. Labels use the largest raw logit, which is equivalent to the largest probability and avoids a needless normalization.
  7. summary() chooses the last sorted class as the reporting reference and subtracts its full block from each earlier class block. The estimator is therefore fit in a symmetric full-class coordinate system but exposed for inference in an identified reference-class coordinate system.

The full-\(C\) fit is pedagogically transparent and treats classes symmetrically, but it leaves a flat common-intercept direction. L-BFGS can still optimize probabilities and contrasts from the symmetric zero initialization; a production large-class solver would more often remove one block up front or impose an explicit sum-to-zero constraint.

4 Inference

Inference is available only at \(\lambda=0\) and is computed directly in the \((C-1)\) reference-class contrast parameterization. For non-reference classes \(a,b\) and \(\tilde x_i=(1,x_i')'\), the information blocks are

\[ H_{ab} = \sum_i \hat p_{ia} \{\mathbf1(a=b)-\hat p_{ib}\} \tilde x_i\tilde x_i'. \]

The returned covariance is \((H+10^{-8}I)^{-1}\). It is model-based only; there is no robust sandwich or Wald-test method on this class. Penalized fits return no covariance. The pairs bootstrap reports class-versus-reference contrasts, but it raises an error if any resample omits an outcome class.

5 Performance and numerical behavior

An L-BFGS objective or gradient evaluation costs \(O(npC)\). The optimizer stores \((p+1)C\) parameters and limited-memory history; the native cost and gradient use \(O(C)\) row scratch rather than materializing an \(n\times C\) probability matrix. Public prediction does return a dense \(n\times C\) array. Fisher inference has dimension \((p+1)(C-1)\); constructing its blocks is costly for many classes and dense inversion is cubic in that total dimension. The full fitted parameterization is unidentified in its common intercept direction, and is also unidentified in common slope directions when \(\lambda=0\), even though reported contrasts are identified. Rare classes make both the Fisher matrix and pairs bootstrap fragile.

6 Python API

Constructor: cm.MultinomialLogit

Use integer class labels in fit(x, y_int32); at least two distinct classes are required and class order is sorted. predict(x) returns an \(n\times C\) probability matrix, predict_lin(x) returns logits, and predict_label(x) returns original class labels. summary() returns fit diagnostics and contrast rows aligned with class_labels; penalized fits mark inference unavailable and omit se/vcov.

print(inspect.signature(cm.MultinomialLogit))
(alpha=0.0, max_iterations=100, gradient_tolerance=0.0001)
cls = cm.MultinomialLogit
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
bootstrap(self, /, n_bootstrap, seed=None)
fit(self, /, x, y)
predict(self, /, x)
predict_label(self, /, x)
predict_lin(self, /, x)
summary(self, /)

7 Minimal example

rng = np.random.default_rng(6)
x = rng.normal(size=(240, 2))
logits = x @ np.array([[0.6, -0.3], [-0.4, 0.5], [0.2, 0.2]]).T + np.array([0.1, -0.2, 0.0])
p = np.exp(logits - logits.max(axis=1, keepdims=True))
p = p / p.sum(axis=1, keepdims=True)
y = np.array([rng.choice(3, p=row) for row in p], dtype=np.int32)
model = cm.MultinomialLogit(max_iterations=200)
model.fit(x, y)
fit = model.summary()
print({key: fit[key] for key in ['converged', 'iterations', 'termination_reason', 'objective']})
print(fit['coef'])
print(model.predict(x[:5]))
{'converged': True, 'iterations': 9, 'termination_reason': 'Solver converged', 'objective': 237.44870578984387}
[[ 0.25831439  0.53237632 -0.55020927]
 [-0.15395916 -0.36374742  0.31279128]]
[[0.29713335 0.35470924 0.34815742]
 [0.10437552 0.60469433 0.29093015]
 [0.35652528 0.30570259 0.33777213]
 [0.27898795 0.37429089 0.34672116]
 [0.3679027  0.30230937 0.32978794]]

8 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

rng = np.random.default_rng(106)
x = rng.normal(size=(100, 2))
logits = x @ np.array([[0.6, -0.3], [-0.4, 0.5], [0.2, 0.2]]).T
p = np.exp(logits - logits.max(1, keepdims=True))
p = p / p.sum(1, keepdims=True)
y = np.array([rng.choice(3, p=row) for row in p], dtype=np.int32)
model = cm.MultinomialLogit(max_iterations=200)
model.fit(x, y)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
coef (2, 3)
class_labels (2,)
reference_class ()
penalty ()
inference_available ()
converged ()
iterations ()
termination_reason ()
objective ()
se (2, 3)
vcov (6, 6)