crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Binding Crash Course
  • Regression And GLMs
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • MEstimator Poisson
  • Causal Inference
    • Balancing Weights
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
  • Transforms
    • PCA And Kernel Basis
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding: First Course
    • Overview And TOC
    • Ch 1 Correlation And Simpson
    • Ch 2 Potential Outcomes
    • Ch 3 CRE And Fisher RT
    • Ch 4 CRE And Neyman
    • Ch 9 Bridging Finite And Superpopulation
    • Ch 11 Propensity Score
    • Ch 12 Double Robust ATE
    • Ch 13 Double Robust ATT
    • Ch 21 Experimental IV
    • Ch 23 Econometric IV
    • Ch 27 Mediation

On this page

  • 1 Where it fits
  • 2 Counting-process partial likelihood
  • 3 Inference limitation
  • 4 Performance and numerical behavior
  • 5 Python API
  • 6 Minimal example
  • 7 summary() contract

AndersenGill

Counting-process Cox model for recurrent events or split risk intervals

from _api_doc_utils import *

1 Where it fits

Group: Survival / event-time models

AndersenGill extends the Cox partial likelihood to counting-process style risk intervals \((s_i, t_i]\). It is the natural entry point for recurrent events, delayed entry, and time-split records with time-varying covariates.

The fitted coefficients still live on the proportional-hazards log-risk scale, so the prediction contract mirrors CoxPH: latent log hazard ratio plus relative risk.

2 Counting-process partial likelihood

For interval record \((s_i,t_i]\), the risk set at event time \(t_i\) is

\[ R(t_i)=\{r:s_r<t_i\leq t_r\}. \]

The class maximizes

\[ \ell_p(\beta) = \sum_{i:d_i=1} \left[ x_i'\beta -\log\left\{\sum_{r\in R(t_i)}\exp(x_r'\beta)\right\} \right]. \]

Tied stop-time events use the Breslow denominator. Optimization is the same Newton method as CoxPH: analytic score and Hessian, a \(10^{-8}\) solve ridge, step clipping, backtracking, and risk-score index clipping to \([-40,40]\) inside the risk sets. Convergence requires either \(\|\nabla\ell_p\|_\infty<\tau\) or an accepted partial-likelihood change below \(\tau(1+|\ell_p|)\).

3 Inference limitation

The returned covariance is only the inverse observed partial-likelihood information. The API has no subject identifier and cannot aggregate score residuals across multiple intervals belonging to the same subject.

Warning

For recurrent-event Andersen-Gill analysis, the usual robust covariance clusters intervals by subject. This class does not implement that covariance. Its reported standard errors and \(p\)-values treat interval rows through the model-based information and should not be presented as robust recurrent-event inference.

The class also does not validate that a subject’s intervals are nonoverlapping or consistently ordered because subject membership is unavailable. Default prediction is relative risk \(\exp(x'\hat\beta)\); no baseline hazard or survival curve is estimated. Reaching max_iterations raises ValueError and leaves the estimator unfitted. A successful summary exposes converged, iterations, termination_reason, and objective, where objective=-\ell_p(\hat\beta).

4 Performance and numerical behavior

Every event scans every interval to reconstruct its delayed-entry risk set and accumulates dense second moments, for about \(O(Enp^2)\) work per Newton iteration plus an \(O(p^3)\) solve. Splitting subjects into more intervals therefore increases both \(n\) and often \(E\). Risk-set exponentials are clipped but new-data relative-risk predictions are not. This implementation is usable for small counting-process designs and point estimates, but its naive risk-set algorithm and missing subject-cluster covariance are material limitations for production recurrent-event work.

5 Python API

Constructor: cm.AndersenGill

Use fit(x, start, stop, event). predict_lin(x) returns the log hazard ratio and predict_relative_risk(x) returns the exponentiated risk multiplier. The default predict(x) is relative risk.

print(inspect.signature(cm.AndersenGill))
()
cls = cm.AndersenGill
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
fit(self, /, x, start, stop, event, max_iterations=50, tolerance=1e-08)
predict(self, /, x)
predict_lin(self, /, x)
predict_log_hazard_ratio(self, /, x)
predict_relative_risk(self, /, x)
summary(self, /)

6 Minimal example

rng=np.random.default_rng(34)
x=rng.normal(size=(180,1)); rate=0.05*np.exp(0.6*x[:,0]); t_event=rng.exponential(1.0/rate); c=rng.exponential(18,size=180); stop=np.minimum(t_event,c); event=(t_event<=c).astype(float)
start=np.zeros_like(stop)
start_long=np.concatenate([start, stop/2]); stop_long=np.concatenate([stop/2, stop]); x_long=np.vstack([x,x]); event_long=np.concatenate([np.zeros_like(event), event])
model=cm.AndersenGill(); model.fit(x_long,start_long,stop_long,event_long)
fit=model.summary(); print({key: fit[key] for key in ['converged', 'iterations', 'termination_reason', 'objective']})
print(model.predict_lin(x[:5]))
print(model.predict(x[:5]))
{'converged': True, 'iterations': 3, 'termination_reason': 'Relative objective tolerance reached', 'objective': 335.9145580407113}
[-0.03003575 -0.94620885  1.93646531  0.3624618   0.48414704]
[0.97041084 0.38821    6.93419737 1.43686233 1.62279025]

7 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

rng=np.random.default_rng(134); x=rng.normal(size=(90,1)); rate=0.05*np.exp(0.5*x[:,0]); te=rng.exponential(1.0/rate); c=rng.exponential(15,size=90); stop=np.minimum(te,c); event=(te<=c).astype(float)
start=np.zeros_like(stop); start_long=np.concatenate([start, stop/2]); stop_long=np.concatenate([stop/2, stop]); x_long=np.vstack([x,x]); event_long=np.concatenate([np.zeros_like(event), event])
model=cm.AndersenGill(); model.fit(x_long,start_long,stop_long,event_long)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
model ()
coef (1,)
hazard_ratio (1,)
se (1,)
z (1,)
p_value (1,)
vcov (1, 1)
log_likelihood ()
converged ()
iterations ()
termination_reason ()
objective ()