What a quantized four-billion-parameter model learns from revealed-use rankings
Abstract
An autoregressive language model already returns a likelihood for any fixed answer string. Conditioning the likelihoods of the admissible strings on the finite alternative set turns that text model into a choice-probability estimator; task adaptation then changes its index, and calibration changes its scale and alternative-specific constants. This paper makes that construction explicit and audits every use of labels. The application is a five-way revealed-use ranking among games a Steam user already owns, not a purchase-choice model. Every estimator is tested on the same 1,000 observations. A popularity rule reaches 0.404 accuracy, a pooled position-specific logistic classifier reaches 0.416, and a symmetric 45-attribute conditional logit reaches 0.461 with log loss 1.335. Frozen Qwen3.5-4B complete-suffix likelihoods reach only 0.286; a frozen five-shot one-token route is also 0.286. A 3.15-million-parameter QLoRA adapter raises raw accuracy to 0.462; an affine probability layer raises it to 0.480 with log loss 1.345. Saved hard-label calls reach 0.387 for five-shot GPT-5.5 and 0.272 and 0.260 for zero- and one-shot GPT-5.3 Codex Spark. Outcome-specific adaptation therefore makes a quantized local model competitive with a strong conventional benchmark, but the experiment does not establish statistical or structural dominance.
Keywords
discrete choice, random utility, language models, parameter-efficient fine-tuning, probability calibration, local inference
1 Introduction
Discrete-choice models map information about decision makers and alternatives into probabilities on a finite set. In a conventional random-utility model, the analyst writes a systematic index and derives a probability kernel from assumptions on the stochastic component (McFadden 1974; Train 2009). Modern empirical settings also contain histories, descriptions, tags, and reviews that are awkward to compress into a prespecified index. A language model offers a rich nonlinear map of those inputs, but its native output is a distribution over text rather than economic alternatives.
The bridge is simpler than generation. Given a fixed context, the network’s final layer produces one logit for every vocabulary token. A vocabulary softmax converts those logits into next-token probabilities. For a multi-token answer, a deterministic forward pass supplies the probability of each successive token conditional on the fixed preceding tokens; summing their log probabilities gives the likelihood of the entire answer string. No answer need be sampled. If \(S_j\) is the answer string assigned to alternative \(j\), its log likelihood can serve as an alternative-specific index,
\[
L_{ij}=\log p_\theta(S_j\mid q_i),
\]
and the analyst can condition the admissible strings on the observed set,
This second normalization is what creates a probability vector on \(C_i\). It addresses the same candidate-string and surface-form problem studied in language-model multiple-choice scoring (Holtzman et al. 2021). It is nevertheless a normalization, not a behavioral derivation: no extreme-value disturbance has been estimated or assumed merely because the final formula resembles logit.
Two softmaxes, two different objects
The vocabulary softmax is inside the language model and ranges over 248,320 tokens in the reported GGUF. The choice softmax is imposed by the analyst and ranges over the five feasible labels. The first defines a text likelihood; the second conditions selected text likelihoods on the empirical choice set. Only the second returns the paper’s five-way predictive object, and neither licenses structural welfare analysis.
This paper follows that object through a transparent adaptation ladder. Frozen scoring exposes the pretrained nonlinear index. In-context learning (ICL) conditions it on labeled examples without changing parameters. Quantized low-rank adaptation (QLoRA) estimates a persistent low-rank correction to selected attention matrices. A final affine map estimates a common scale and five position offsets from validation probabilities. The pipeline runs with a Q4_K_M Qwen3.5-4B model through llama.cpp on a 12 GB GPU. A data-usage ledger records the features, labels, loss, and selection sample at every stage.
The empirical target must be stated precisely. Each observation contains five games that a Steam user already owns, and the outcome is the one with greatest recorded lifetime playtime. It is a revealed-use ranking represented as a five-alternative classification problem, not an observed purchase from a budget set. Candidate playtime is withheld; the inputs are non-candidate game and review histories, user aggregates, and candidate metadata. The construction is useful for testing a text-rich probability pipeline with sharp leakage controls, but it cannot identify demand, price elasticities, or welfare.
The benchmark ladder is empirically demanding. The most-popular-candidate rule is already strong at 0.404 accuracy. A pooled logistic classifier with position-specific coefficients reaches 0.416. Imposing the exchangeability built into the randomized display through a 45-attribute conditional logit raises accuracy to 0.461 and lowers log loss to 1.335. Frozen complete-suffix likelihoods reach only 0.286; a frozen five-shot one-token specification is also 0.286, although that contrast changes both demonstrations and label representation. QLoRA raises accuracy to 0.462, essentially matching conditional logit; affine calibration raises it to 0.480 and log loss 1.345. The paired accuracy difference relative to conditional logit is 0.019 with a 95% bootstrap interval of \([-0.007,0.045]\) and exact McNemar \(p=0.181\). Conditional logit’s log loss is lower by 0.010, with a paired interval that includes zero. The local pipeline has the highest point estimate for accuracy, not demonstrated predictive dominance.
The paper sits between four literatures. Neural choice models already replace or augment analyst-written utility with learned representations (Sifringer, Lurkin, and Alahi 2020; Han et al. 2022; van Cranenburgh et al. 2022). Work in empirical demand uses pretrained text and image embeddings as covariates while retaining a structural choice model (Compiani, Morozov, and Seiler 2026). “Homo silicus” work asks LMs to simulate respondents (Horton, Filippas, and Manning 2023), while recommender research fine-tunes LMs directly on recommendation labels (Bao et al. 2023). Here the outcomes are observed human behavior; the LM is the reduced-form index rather than a synthetic respondent or a feature extractor; and the contribution is the combination of finite-label likelihood scoring, explicit supervised-loss accounting, a conventional conditional-logit benchmark, probability calibration, and a reproducible sample ledger.
The two vocabularies used below translate as follows.
Language-model term
Choice-model meaning in this paper
Serialized query or prompt \(q_i\)
Covariates represented as text
Token logits
Unnormalized scores for vocabulary items
Vocabulary softmax
LM output layer; not the choice model
Choice softmax over \(C_i\)
Analyst-imposed finite-set conditioning
Verbalizer A–E
Economically meaningless encoding of displayed position
In-context example or shot
Labeled conditioning information; no persistent estimate
LoRA or fine-tuning
Estimation of a restricted parameter update
Calibration offset
Alternative-specific constant for the label-position interface
Temperature
Common inverse scale on the logged probability index
High confidence
A peaked fitted vector, not evidence that the prediction is correct
2 Data and empirical design
2.1 The empirical object
The source files contain ownership, playtime, reviews, and product metadata from the Australian Steam community (Pathak, Gupta, and McAuley 2017). They do not record a sequence of menus and purchases. All five displayed games have already been acquired. Let \(M_{ij}\) be user \(i\)’s recorded lifetime minutes for candidate \(j\) in an analyst-sampled set \(C_i\). The outcome is
\[
Y_i=\arg\max_{j\in C_i}M_{ij}.
\]
The exercise is therefore a personalized revealed-use ranking with a discrete-choice representation. \(C_i\) is not a budget set, \(\pi_{ij}\) is not a market share, and \(M_{ij}\) combines preference with exposure, ownership duration, session length, and the typical time intensity of a genre. A sixty-hour role-playing game and a two-hour puzzle game are not directly comparable units of utility. This deliberate construction supplies a hard, text-rich, five-way prediction target with auditable leakage controls; it does not supply a demand system.
The eligible catalog is the 500 most common games among source users. An included user has at least 24 catalog games with at least 30 minutes of playtime. Up to four mutually disjoint five-game sets are sampled per user and randomly ordered. numpy.argmax resolves an exact top-playtime tie in favor of the lowest displayed position. Seven of 16,001 observations have a top tie: four in estimation, one in validation, and two in test. Only two occur in the capped 4,000-row estimation pool, so the rule cannot explain the estimated position-one calibration offset.
Candidate sets, rather than users, are split: every included user contributes one validation set, one test set, and at least one estimation set, with disjoint games across that user’s sets. The design deliberately permits a model to learn from one set for a user and predict another set from the same user’s observed attributes. A user-level split would answer a different cold-start question; overlap and non-overlap results are reported separately.
2.2 Information set and leakage controls
Every observation supplies four classes of pre-outcome information:
the ten highest-playtime non-candidate games and their minutes;
up to five non-candidate reviews, with recommendation indicators and the first 240 characters of text;
ten user-history aggregates, such as total and mean playtime, review counts, recommendation share, and mean review length; and
candidate titles, prices, release years, catalog popularity, genres, tags, 32 taxonomy indicators, taxonomy overlap with the history, and developer or publisher matches.
All candidate games are removed from the displayed history and review excerpts. The five candidate playtimes that define \(Y_i\) never enter a feature, prompt, or calibration stage. Candidate games are disjoint across a user’s sampled sets before splitting.
Catalog membership and popularity_count and popularity_rank are computed once from the full source ownership file before candidate sets are split. They use no candidate playtime or within-set outcome, and the same product metadata is available to classical and LM systems. They are nevertheless transductive: ownership records from eventual evaluation users contribute to the global popularity measure. A strictly inductive deployment study would recompute these fields using estimation users or an external pre-period catalog.
The systems do not receive identical encodings. Conditional logit receives 45 numeric or binary attributes for each candidate: 13 product or history-match fields and 32 taxonomy indicators. The pooled logistic classifier receives those 45 fields in each of five position blocks plus ten user aggregates, for 235 standardized covariates and 1,175 class-specific slopes. The LM receives the ten aggregates and a subset of the candidate scalars in text form, plus information not available to those classical specifications: literal game titles, genre and tag strings, top-history titles and minutes, and review text. The comparison is between complete prediction pipelines, not a controlled functional-form experiment holding representation fixed.
The full processed data contain 16,001 observations from 4,078 users. Randomized position makes the labels close to balanced; selecting the most common position in the estimation pool reaches 0.205 accuracy on the fixed test subset.
Table 1. Processed sample composition
Split
Observations
Users
Alternative 1
Alternative 2
Alternative 3
Alternative 4
Alternative 5
Estimation
7,845
4,078
1,536
1,624
1,589
1,543
1,553
Validation
4,078
4,078
830
785
853
782
828
Test
4,078
4,078
830
845
804
773
826
The outcome-estimated classical models and QLoRA use the first 4,000 estimation rows; the descriptive position-frequency prior uses all 7,845 estimation rows. Validation and test evaluation use the same fixed 1,000 identifiers for every system. The capped 4,000-row pool contains 2,994 users; 737 evaluation users appear in it and 263 do not. Table 2 is the label-usage ledger. The first 200 model-setting rows are contained in the 1,000-row validation sample, so calibration data are not independent of specification search. Test outcomes enter only final reporting and uncertainty calculations.
Table 2. Data and label usage
Sample
\(N\)
Stage using outcomes
Object fitted or selected
Full estimation split
7,845
Position-frequency prior
Five Laplace-smoothed label counts; four free probabilities
No outcome labels
0
Most-popular-candidate rule
No fit; select the greatest catalog popularity count in each set
Estimation pool
4,000
Pooled logistic classifier
1,180 slopes and intercepts by displayed position, penalized likelihood
Estimation pool
4,000
Conditional logit
45 common candidate-attribute coefficients, penalized likelihood
Estimation pool
4,000
ID-based low-rank logit
User and item factors; rank and penalty selected on validation
Estimation pool
1,600 row presentations
QLoRA
3,145,728 adapter parameters by token cross-entropy
Estimation pool
5 labeled rows
Local ICL demonstrations
Nothing persistent; labels are inserted into every query
Estimation pool
0, 1, or 5 labeled rows
External API demonstrations
Nothing persistent; generated hard label only
Validation, first 200
200
Trainer monitoring and model-setting search
Checkpoint, adapter scale, and demonstrations on/off
Validation, full
1,000
Classical tuning and LM calibration
Ridge or factor settings; five-fold OOF calibration-family comparison; final calibration refit
Test
1,000
Reporting only
No fitted parameter or selected setting
The absence of a gradient does not make ICL “data free”: the five demonstrations contain outcomes, and the decision to retain demonstrations was validation informed. More generally, a stage is held out only if neither its computation nor the analyst’s choice of it depends on held-out outcomes.
3 From systematic utility to language-model scores
where \(V_{ij}\) is now a string log likelihood. Only differences in \(V_{ij}\) matter, so a common observation-specific constant is unidentified. What follows and what does not follow from this reuse should be kept separate.
Property
Classical logit derivation
LM finite-set normalization
Five nonnegative probabilities summing to one
Yes
Yes
Invariance to a common additive index constant
Yes
Yes
EV1 disturbance interpretation
Yes
No
IIA from context-independent alternative utility
Yes
Not generally
Exact equivariance to relabeling and display order
Yes
No
Welfare or elasticity interpretation
With additional identification
No
Because every candidate is in the prompt, \(V_{ij}\) may depend on the entire displayed set. The LM can represent compromise, dominance, or other context effects without a special kernel. It can just as freely react to irrelevant formatting, title length, position, or verbalizer choice. This is flexible but unrestricted set dependence, not a structural relaxation of IIA.
3.1 Two simple rules and three classical benchmarks
Position-frequency prior. Let \(n_j\) count position \(j\) in the full 7,845-row estimation split. The baseline assigns every observation
and predicts its largest component. The only outcome-bearing inputs are the 7,845 full-estimation labels; there is no numerical optimization. The most-popular-candidate rule uses no outcome label at all: it selects \(\arg\max_{j\in C_i}\texttt{popularity\_count}_{ij}\). It is evaluated only as a hard rule because no probabilistic mechanism is estimated.
Pooled position-specific logistic classifier. Let \(X_i\in\mathbb R^{235}\) stack ten user aggregates and five blocks of 45 candidate attributes. The fitted statistical classifier is
It estimates a different 235-vector for each randomized displayed position, for 1,175 slopes and five intercepts. With \(N=4{,}000\) and scikit-learn’s \(C=1\), the implemented objective is
on a standardized design; intercepts are unpenalized. This is multinomial logistic regression in the classification sense, not McFadden conditional logit. It uses the same row-level blocks the LM serializes, but inefficiently relearns the same relationship at five random positions.
Symmetric conditional logit. Let \(x_{ij}\in\mathbb R^{45}\) contain candidate \(j\)’s price, free and missing-price flags, release year, early-access flag, popularity count and rank, title length, five history-match measures, and 32 taxonomy indicators. The benchmark appropriate to exchangeable positions is
After standardizing each attribute over all training alternatives, \(\beta\) minimizes mean negative log likelihood plus \(\lambda\lVert\beta\rVert_2^2/2\). The validation-log-loss grid is \(\lambda\in\{0,10^{-4},10^{-3},10^{-2},10^{-1},1\}\); \(\lambda=10^{-2}\) is selected. There is no position constant because position is randomized and economically meaningless. This 45-parameter model is both a stronger benchmark and the clean analogue of the later calibration equation.
Equivalently, its implemented criterion for a fixed \(\lambda\) is
ID-based low-rank logit. The empirical factor model follows the repeated-user/item logic of Kallus and Udell (2016),
\[
V_{uij}^{\mathrm{LR}}=a_u^{\top}b_j,
\]
with probabilities normalized only over the displayed assortment. For fixed rank and penalty, the Trex implementation minimizes the summed choice NLL plus a factor penalty,
Utilities are centered to sum to zero across the 500 items for each user; this identification normalization leaves assortment probabilities unchanged. Ranks \(r\in\{2,4,8,16\}\) and penalties \(\lambda\in\{10^{-3},10^{-2},10^{-1},1\}\) are selected by validation accuracy, with NLL, RMSE, Brier score, macro F1, and macro AUC used in that order for metric ties. An exact tie across all metrics retains the first grid row encountered. The selected rank is two and the penalty is one. This model reads identities and the candidate availability mask, not attributes. The capped estimation data contain 2,994 users and 500 items but only 4,000 observations; 263 test users are unseen and must use the most frequent training user’s factor as a disclosed fallback. It is therefore a direct test of identity-indexed latent structure under severe cold start, not a substitute for the attribute-based conditional logit.
Mixed logit places low-dimensional structure somewhere else (Train 2009): \(\beta_i=\bar\beta+L\eta_i\) is random and integrated out, which can relax IIA and yield substitution patterns under maintained behavioral assumptions. QLoRA instead produces a deterministic query-conditional index and restricts neural weight updates. The shared phrase “low rank” should not obscure these different estimands.
3.2 Complete-suffix likelihood
Let \(q_i=q(C_i,z_i)\) be the chat-templated system and user turns, including the assistant-generation prefix. Let \(S_j=(s_{j1},\ldots,s_{jT_j})\) be the tokenized suffix that completes the assistant turn with choice_j and Qwen’s end-of-turn material. The frozen index is
and \(\pi_{ij}^{\mathrm{full}}=\operatorname{softmax}_j(\ell_{i1}^{\mathrm{full}},\ldots,\ell_{i5}^{\mathrm{full}})\). All five alternatives tokenize to three scored suffix tokens, so summed likelihood does not create cross-alternative length bias here. Natural game names would not share that property and should not be used as unequal-length labels without a considered scoring rule.
The native scorer encodes the long common prefix once and branches into the five suffixes. Across the 1,000 test observations, the median of \(\sum_j\exp(\ell_{ij}^{\mathrm{full}})\) is 0.9996. Thus the feasible suffixes happen to capture nearly all measured continuation mass in this tightly instructed interface. This fact does not make conditioning conceptually redundant, and it does rule out “large mass leakage” as an explanation for the later fitted temperature.
3.3 Grammar-constrained one-token likelihood
The faster interface replaces choice_1 through choice_5 everywhere with the single-token verbalizers A through E. For demonstrations \(D\),
At inference a GBNF grammar permits exactly A, B, C, D, or E; max_tokens=1, temperature and top-\(p\) equal one, top-\(k=0\), minimum-\(p=0\), and repetition penalty equals one. llama.cpp returns the post-grammar probabilities of all five letters. The sampled letter is discarded. Conditional on the prompt, model, and numerical runtime, the estimator is the returned vector—not a random decoded answer.
The route uses one model evaluation per observation instead of five suffix continuations, but it is not faster in the reported specification: five long demonstrations raise its context to a median 5,616 tokens. Exact frozen scoring reaches 4.59 observations per second, while adapted five-shot one-token scoring reaches 2.22.
3.4 Generated hard choices
Some external systems expose only a generated answer in this experiment,
Saved GPT-5.3 Codex Spark and GPT-5.5 calls receive the same serialized features and canonical labels. Spark uses zero or one demonstration, batches of 20, and xhigh reasoning; GPT-5.5 uses five demonstrations, batches of five, and medium reasoning. The one-shot row uses estimation identifier 6728 with label choice_1. The five-shot row uses identifiers 6728, 13617, 6964, 2927, and 4388 with labels choice_1 through choice_5, respectively. These class-balanced demonstrations fit no parameter and optimize no loss; they are fixed prompt inputs. A missing identifier, malformed response, or invalid label counts as incorrect. These are classifiers, not probability estimators: epsilon-smoothing their selected labels would manufacture rather than recover log loss or calibration.
Figure 1 shows the full probability route. The adapter alters the network before either softmax; the affine map acts only after the five-way vector has been constructed.
Figure 1: From serialized covariates to a finite-choice probability vector. The vocabulary distribution is internal to the LM; the five-way conditioning and affine map are analyst-imposed.
4 Adaptation and calibration: features, labels, and losses
This section treats each intervention as an estimator. For every rung, it states the covariates, outcome labels, fitted object, criterion, and sample. The distinction is substantive: ICL consumes labeled observations but fits no persistent parameter; QLoRA estimates millions of network parameters with a vocabulary-level token loss; calibration estimates at most five free numbers from already-produced probabilities.
4.1 Labels and the serialized design
Let \(J_i\in\{1,\ldots,5\}\) denote the displayed position of the candidate with the greatest lifetime playtime and let
\[
y_{ij}=\mathbb 1\{J_i=j\}.
\]
The complete-suffix estimator uses the canonical strings choice_1 through choice_5. The one-token estimator applies the deterministic verbalizer
Thus A means “first displayed candidate,” not a game, brand, or product category. Candidate order is randomized. The 4,000-row estimation pool contains 794, 814, 809, 785, and 798 outcomes at positions one through five. There is no class weighting, oversampling, or label smoothing.
The LM covariate is the literal serialized query \(q_i=q(C_i,z_i)\). It contains top non-candidate games and minutes; non-candidate review snippets; ten user-history scalars; and, for each candidate, price, free status, release year, popularity, five history-match measures, title, genres, and tags. Candidate playtime is absent. The one-token route replaces every canonical position label in the query with its letter. Below is the complete test prompt for observation 12678 after that deterministic replacement; only line wrapping has changed.
SYSTEM: You are a discrete choice model. Given structured covariates and
candidate alternatives, predict the observed chosen alternative. Respond
with exactly one label from: A, B, C, D, E.
USER: Predict which held-out Steam game this user played the most.
The observed user history below excludes all candidate games and candidate reviews.
Return exactly one label from: A, B, C, D, E.
Top played non-candidate games as [game, lifetime playtime minutes; ...]:
[Warframe, 8559; Terraria, 5410; Euro Truck Simulator 2, 2674;
Don't Starve Together, 492; PAYDAY 2, 186; Counter-Strike: Source, 149;
Fractured Space, 104;
Red Orchestra 2: Heroes of Stalingrad with Rising Storm, 48]
Non-candidate reviews as [game, recommended flag, review text; ...]:
[none]
Tabular features:
- user: user_items_count=51, user_played_games=28, history_played_games=8,
history_total_playtime=17622, history_mean_playtime=2202.75,
history_median_playtime=339, history_max_playtime=8559,
history_review_count=0, history_recommend_share=0,
history_mean_review_length=0
- A: price=0, is_free=1, release_year=2017, popularity_count=13970,
popularity_rank=4, history_term_minute_score=7.34,
history_term_count_score=57, history_term_overlap=13,
history_developer_match_count=0, history_publisher_match_count=0
- B: price=0.99, is_free=0, release_year=2013, popularity_count=3430,
popularity_rank=66, history_term_minute_score=7.17,
history_term_count_score=55, history_term_overlap=11,
history_developer_match_count=0, history_publisher_match_count=0
- C: price=0, is_free=1, release_year=2011, popularity_count=4863,
popularity_rank=39, history_term_minute_score=6.2,
history_term_count_score=49, history_term_overlap=11,
history_developer_match_count=0, history_publisher_match_count=0
- D: price=19.99, is_free=0, release_year=2015, popularity_count=6368,
popularity_rank=18, history_term_minute_score=4.53,
history_term_count_score=37, history_term_overlap=9,
history_developer_match_count=0, history_publisher_match_count=0
- E: price=39.99, is_free=0, release_year=2015, popularity_count=4949,
popularity_rank=37, history_term_minute_score=6.83,
history_term_count_score=53, history_term_overlap=12,
history_developer_match_count=0, history_publisher_match_count=0
Candidate games:
- A: Unturned | genres: Action, Adventure, Casual, Free to Play |
tags: Free to Play, Survival, Zombies, Multiplayer, Open World, Adventure |
price: 0 | release year: 2017
- B: ORION: Prelude | genres: Action, Adventure, Indie, RPG |
tags: Dinosaurs, Action, FPS, Multiplayer, Online Co-Op, Co-op |
price: 0.99 | release year: 2013
- C: No More Room in Hell | genres: Action, Free to Play, Indie |
tags: Free to Play, Zombies, Multiplayer, Survival, Horror, Co-op |
price: 0 | release year: 2011
- D: Rocket League® | genres: Action, Indie, Racing, Sports |
tags: Multiplayer, Racing, Soccer, Sports, Competitive, Team-Based |
price: 19.99 | release year: 2015
- E: Grand Theft Auto V | genres: Action, Adventure |
tags: Open World, Action, Multiplayer, First-Person, Third Person, Crime |
price: 39.99 | release year: 2015
Answer with only the label.
The held-out outcome is E; an estimation row would use that letter as its supervised assistant response. Hidden reasoning is disabled because the estimand is the probability of the first admissible answer token immediately after the assistant-generation prefix. Allowing sampled reasoning tokens to intervene would define a different, generation-mediated classifier.
This example also makes the empirical path concrete. The entries below are fitted choice probabilities for the same held-out row.
Specification
\(p(A)\)
\(p(B)\)
\(p(C)\)
\(p(D)\)
\(p(E)\)
Prediction
Frozen complete suffix
0.536
0.019
0.412
0.020
0.013
A
Frozen five-shot token
0.371
0.135
0.291
0.134
0.070
A
QLoRA, uncalibrated
0.041
0.046
0.049
0.201
0.662
E
QLoRA plus affine calibration
0.039
0.024
0.023
0.134
0.780
E
Conditional logit
0.206
0.098
0.044
0.204
0.448
E
4.2 In-context conditioning
The five-shot specification prepends five labeled estimation examples as alternating user and assistant turns. Every demonstration contains the same serialized features and an assistant response \(v(J_i)\). One observation per outcome class is selected from the first 256 estimation rows by minimizing prompt character count within class. The retained identifiers are 9944, 2952, 14726, 12676, and 14087 for A through E.
Conditional on this rule, ICL estimates no persistent parameter and minimizes no explicit loss:
It nevertheless uses five observed outcomes on every prediction, and the analyst’s choice to retain ICL is informed by validation performance. Demonstrations can communicate the input-output convention even when they do not identify a stable parametric relationship (Min et al. 2022). Selecting the shortest case in each class controls context length but favors users with relatively thin serialized histories. A random-shot robustness distribution was not run; the result pertains to this disclosed deterministic demonstration set.
4.3 QLoRA estimation
For every adapted attention projection, LoRA replaces a frozen matrix by (Hu et al. 2021)
\[
W=W_0+\frac{\alpha}{r}BA,
\]
where \(A\in\mathbb R^{r\times d_{\mathrm{in}}}\) and \(B\in\mathbb R^{d_{\mathrm{out}}\times r}\). QLoRA backpropagates through a frozen four-bit base representation while updating only \(A\) and \(B\)(Dettmers et al. 2023).
4.3.1 The implemented target is token prediction, not five-class likelihood
For row \(i\), let \((w_{i1},\ldots,w_{iT_i})\) be the tokenized system message, query, and observed assistant completion. Let \(R_i\) index unmasked response tokens. Prompt and padding tokens receive the ignore index \(-100\). The implemented trainer minimizes
where \(\mathcal V\) is Qwen’s entire vocabulary and \(\mathcal E\) is the 4,000-row estimation pool. An untruncated completion contains three supervised tokens: the outcome letter, Qwen’s end-of-turn marker, and a newline. The same token weight is assigned to all three.
A five-class choice loss would instead use only the outcome-letter position \(t_i^*\) and normalize over the five verbalizers,
These are not equivalent. At the label position, the vocabulary loss equals the five-class loss plus the log ratio of the full-vocabulary partition function to the five-verbalizer partition function. It also adds two constant-target termination terms. The reported trainer loss of 0.5057 therefore cannot be compared with \(\log 5\), the validation choice NLL, or the test choice NLL. The termination tokens may become easy quickly, so their share of token count is not a claim about their share of useful gradient information. This objective mismatch is a practical consequence of using a standard completion-only SFT trainer; a direct restricted-label loss is an important next experiment.
Right truncation at 1,536 tokens creates one further realized deviation. Because the assistant response is at the end, 22 of 4,000 estimation sequences lose the full response. The masking fallback supervises the last retained prompt token for those rows; one of the 200 trainer-monitoring rows is affected. The run stopped after 0.40 nominal epoch and did not retain minibatch identifiers, so the artifact cannot establish how many of the 22 rows entered an update. The paper reports the executed estimator rather than retroactively changing it.
4.3.2 Parameterization and numerical optimization
Component
Realized choice
Consequence
Frozen training base
NF4, double quantization, BF16 arithmetic
Base weights receive no gradient
Adapter locations
Query, key, value, and output attention projections
Attention only; no MLP projection is adapted
Rank and scale
\(r=16\), \(\alpha=32\)
Training-time multiplier \(\alpha/r=2\)
Trainable size
3,145,728 of 4.209 billion parameters
0.0747% of model parameters
Regularization
LoRA dropout 0.05; no adapter bias; no weight decay
Stochastic adapter path, zero-initialized through \(B\)
\(\beta_1=0.9\), \(\beta_2=0.999\), \(\epsilon=10^{-8}\), gradient norm at most one
Batching
Effective batch size 8
200 steps equal 1,600 row presentations
Schedule
3% warmup, cosine, seed 42
Three declared horizons: 25, then 100, then 200 steps
The rank, target projections, and number of updates define the estimator’s capacity. The optimizer precision, clipping, and checkpointing primarily make estimation feasible on a 12 GB GPU. Extending the declared horizon reinitialized the cosine schedule: the learning rate approached zero at steps 25 and 100 and then rose when the horizon was extended. The realized path is reproducible but is not one uninterrupted 200-step cosine decay.
4.3.3 Validation-selected inference setting
At llama.cpp inference the exported adapter is applied to a Q4_K_M base rather than the bitsandbytes NF4 representation used during training, and it is multiplied by a separate scale \(\gamma\),
\[
W h=W_0h+\gamma\frac{\alpha}{r}BAh.
\]
The adapter checkpoint, \(\gamma\), and inclusion of demonstrations are selected sequentially on the first 200 validation observations using hard-choice accuracy. Table 3 reports all stages; nine inference passes were made because the scale-one step-100 setting was evaluated in both the checkpoint and scale screens.
Table 3. Sequential model-setting search on 200 validation observations
For a proportion near one-half, the binomial standard error at \(N=200\) is about 0.035. The search is therefore useful engineering evidence, not a precise ranking of nearby settings. It is sequential rather than Cartesian, it has no multiplicity correction, and the same 200 observations later belong to the 1,000-row calibration sample. The test data are untouched. Across the nine recorded setting passes, inference took about 781 seconds.
The distinction between \(\gamma\) and calibration temperature is important. \(\gamma=1.25\) amplifies the estimated neural update before the network produces logits; \(T\) rescales log probabilities after the network has produced a five-vector. They need not have the same empirical effect.
4.4 Economic interpretation of the low-rank restriction
Write the task-specific index as
\[
V_{ij}^*=V_{ij}^0+G_{ij},
\]
where \(V_{ij}^0\) is the frozen LM index and \(G_{ij}\) is the correction required for this outcome. Pretraining may already represent semantic similarity, genre, popularity, and history-candidate match; adaptation changes how those represented signals map into the revealed-use label. A useful low-dimensional hypothesis is
Possible directions include affinity with high-playtime history, multiplayer fit, novelty, and generic popularity. These names are economic interpretations, not identified adapter factors, and the interpretive dimension \(R\) need not equal the matrix rank \(r\).
Because \(\Delta\theta\) is restricted to low-rank changes in selected matrices, the allowable index correction lies in a restricted tangent family. That is the useful analogy to low-dimensional mixed-logit heterogeneity. It is not an equivalence: QLoRA estimates no distribution \(F(\beta_i)\), its heterogeneity is deterministic conditional on the query, and rank 16 in neural weight space does not imply a rank-16 user-item utility matrix.
The only calibrator covariates are the five logged source probabilities \(x_i\). No Steam field, text embedding, user identity, or candidate identity enters. The target is the same \(y_{ij}\), and the common criterion is multiclass negative log likelihood,
Five stratified folds split the 1,000 validation observations, with shuffling and seed 42. Parameters are fitted on 800 rows and scored on the other 200; the concatenated holdouts provide out-of-fold selection metrics. The selected family is then refitted on all 1,000 validation rows and applied once to test.
The four candidate maps are:
Direct:\(q_{ij}=\check p_{ij}\); no fitted parameter.
Temperature:\(q_{ij}=\operatorname{softmax}_j(x_{ij}/T)\) with \(T=\exp(\tau)>0\); BFGS starts from \(\tau=0\) and minimizes NLL. Because division by a common positive number preserves ordering, temperature alone cannot change accuracy. The selected adapted source gives full-validation \(T=0.8589\).
Marginal prior shift:\(q_{ij}=\operatorname{softmax}_j(x_{ij}+\lambda d_j)\), where \(d_j=\log\bar y_j-\log\bar p_j\). Both moments and \(\lambda\in\{0,0.025,\ldots,2\}\) are recomputed inside every training fold. The implemented rule maximizes in-sample accuracy, then breaks ties by NLL and larger \(\lambda\); this is the only family not selected by NLL alone. The full-validation values are \(\lambda=1.15\) and \(d=(0.406,-0.026,-0.122,-0.170,0.011)\).
Affine calibration is simply a conditional logit with one alternative-varying regressor \(x_{ij}\), a common slope \(1/T\), and four identified alternative-specific constants. The optimizer represents the offsets by four free values, appends a fifth zero, centers all five to impose the sum-zero restriction, and starts every free value and \(\tau\) at zero. BFGS minimizes
the common slope is unpenalized. For the selected adapted source, the full-validation fit is \(T=0.7772\) and
\[
b=(0.541,-0.063,-0.204,-0.251,-0.023).
\]
Because displayed positions are randomized, these offsets are not product constants or tastes. They diagnose the letter-and-position measurement interface. The positive first-position adjustment and broadly decreasing offsets through position four are consistent with prompt or verbalizer position bias, a known issue in few-shot classification (Zhao et al. 2021). The partial recovery at position five cautions against a literal linear primacy story. The estimate \(T<1\) sharpens rather than flattens probabilities, so the raw adapted model is underconfident on validation. This cannot be attributed to large omitted suffix mass: the exact scorer assigns median total mass 0.9996 to the five complete suffixes. Calibration improves predictive probabilities (Guo et al. 2017); it does not recover preference coefficients.
4.6 The fourth route: activation steering
Activation steering was evaluated after the three principal routes. A global supervised hidden-state contrast changed confidence slightly but did not materially improve validation classification, so it was not selected, calibrated, or evaluated on test. The complete feature, label, probe-loss, direction, and scale definitions are reported in Appendix A.
5 Evaluation
Accuracy is the principal hard-classification statistic because it is the only performance object shared by the local probability models and the saved API classifiers. The empirical position-frequency baseline is 0.205, close to but not exactly the theoretical 0.200 chance rate. Macro F1 prevents a favorable aggregate score from hiding a failed displayed position.
For systems that return a probability vector, average negative log likelihood is the primary probability criterion,
It is a strictly proper score and directly punishes assigning little mass to the realized alternative. The multiclass Brier score is the mean squared distance from the one-hot outcome and is less tail-sensitive. One-versus-rest macro AUC measures ranking separately for each displayed position and does not evaluate calibration.
Model and adapter settings use validation outcomes. Calibration families use five-fold out-of-fold validation predictions before a full-validation refit. No setting is selected on test. Marginal 95% Wilson intervals accompany accuracy points in the figure. Differences use 100,000 paired nonparametric bootstrap resamples of the 1,000 test observations; because each test row belongs to a distinct user, this also resamples users. Exact two-sided McNemar tests use the two discordant cell counts.
These intervals quantify only variation in the realized test sample conditional on the entire chosen pipeline. They do not include adapter-training randomness, ICL-shot uncertainty, validation search, calibration-parameter estimation, prompt design, quantization choice, or base-model pretraining.
6 Results
6.1 The benchmark ladder
Table 5 places the simple rules, classical probability estimators, and local LM stages on the same test identifiers. The central comparison is not with the pooled position classifier. Once alternative exchangeability is imposed, the 45-coefficient conditional logit reaches 0.461 accuracy and has the best NLL, Brier score, and AUC in the table. The selected Qwen system reaches the highest hard accuracy, 0.480, but its NLL is 0.010 higher.
Table 5. Test performance on 1,000 common observations
Estimator
Accuracy
Macro F1
NLL
Brier
Macro AUC
Position-frequency prior
0.205
0.068
1.609
0.800
0.500
Most-popular candidate
0.404
0.404
–
–
–
ID low-rank logit, rank 2
0.287
0.286
2.942
0.965
0.539
Pooled position classifier, 1,180 fitted numbers
0.416
0.417
1.450
0.720
0.727
Symmetric conditional logit, 45 attributes
0.461
0.461
1.335
0.671
0.754
Frozen Qwen, complete-suffix likelihood
0.286
0.282
1.839
0.878
0.601
Frozen Qwen, five-shot one-token likelihood
0.286
0.273
1.830
0.884
0.615
Frozen Qwen, five-shot plus affine calibration
0.308
0.308
1.557
0.778
0.619
QLoRA Qwen, five-shot, uncalibrated
0.462
0.460
1.381
0.689
0.751
QLoRA Qwen, five-shot plus affine calibration
0.480
0.480
1.345
0.673
0.751
The popularity rule is included only as a hard decision rule; because it estimates no probability model, proper scores are omitted. The frozen complete-suffix and five-shot one-token rows differ in both label representation and demonstrations, so their equal accuracy does not identify a pure ICL effect. The closest raw before-after comparison holds the one-token verbalizers and five demonstrations fixed but changes the compact frozen wrapper to the training-aligned prompt template. Across that comparison, QLoRA moves accuracy from 0.286 to 0.462, a 17.6-point increase; it is strong evidence for adaptation but not a perfectly isolated adapter effect. Independently fitted affine maps move the corresponding final figures from 0.308 to 0.480, a 17.2-point increase.
The ID factor model fails in this sparse design. There are only 1.34 capped estimation rows per observed estimation user, and 263 test users are unseen. Its rank-two factors therefore have little repeated-choice information from which to learn, and its disclosed fallback is especially weak. This result is about the data regime, not a general rejection of matrix-completion choice models.
6.2 Paired comparisons
Table 6 reports the selected Qwen classifier against three increasingly demanding references. Its 6.4-point advantage over the pooled position classifier survives paired sampling uncertainty, but that is not the appropriate headline benchmark. Relative to conditional logit, the point difference is only 1.9 points, its interval includes zero, and the exact discordance test does not reject equality.
Table 6. Paired test comparisons for selected Qwen
Reference
Reference accuracy
Qwen minus reference
Paired bootstrap 95% interval
Qwen only correct
Reference only correct
Exact McNemar \(p\)
Most-popular candidate
0.404
0.076
\([0.044,0.108]\)
177
101
\(6.0\times10^{-6}\)
Pooled position classifier
0.416
0.064
\([0.033,0.095]\)
159
95
\(7.1\times10^{-5}\)
Symmetric conditional logit
0.461
0.019
\([-0.007,0.045]\)
100
81
0.181
For the main conditional-logit comparison, Qwen NLL minus conditional-logit NLL is 0.010, with paired bootstrap interval \([-0.017,0.037]\). Thus neither the accuracy evidence nor the proper-score evidence establishes dominance. The 181 rows on which exactly one of these two systems is correct are nevertheless informative: the systems are not duplicates, and a validation-designed ensemble is a natural next experiment.
6.3 External hard-label classifiers
The three retained API runs cover the full 1,000-row test sample. They use the same serialized information and target positions, but they are not probability estimators in this experiment: only generated labels were retained. GPT-5.3 Codex Spark uses xhigh reasoning with zero or one demonstration; GPT-5.5 uses medium reasoning with five. Neither receives task-specific parameter updates.
Table 7. Hard-label classification on the common test sample
Prediction system
Outcome-specific adaptation
Demonstrations
Accuracy
Macro F1
Parseable output
Selected local Qwen3.5-4B
Rank-16 QLoRA plus affine map
5
0.480
0.480
1.000
Symmetric conditional logit
45-coefficient penalized likelihood
0
0.461
0.461
1.000
Pooled position classifier
1,180 fitted numbers
0
0.416
0.417
1.000
GPT-5.5, medium reasoning
None
5
0.387
0.388
1.000
GPT-5.3 Codex Spark, xhigh
None
0
0.272
0.272
0.996
GPT-5.3 Codex Spark, xhigh
None
1
0.260
0.260
0.992
Malformed or missing API labels count as incorrect: four Spark zero-shot rows and eight one-shot rows are invalid. The five-shot GPT-5.5 result is 7.4 points below conditional logit and 9.3 points below adapted Qwen. This is evidence for the value of outcome-specific estimation on this narrow task, not evidence that the four-billion-parameter base model is generally more capable than GPT-5.5. Qwen consumed 4,000 domain labels through training; GPT-5.5 consumed five through context. The 1.2-point zero-to-one-shot Spark difference changes both the context and one demonstration and is too small to interpret as a general ICL effect.
Figure 2 collects every principal local stage, the classical benchmarks, and all completed API calls. Every point has the same denominator and observation identifiers.
Figure 2: Hard-choice accuracy on the same 1,000 test observations. Bars are marginal 95% Wilson intervals. The bracket emphasizes the paired Qwen versus conditional-logit comparison; its interval is for the paired difference, not the marginal bars. Invalid API outputs count as incorrect.
6.4 User overlap
Table 8 stratifies the same test sample by whether the user’s identifier appears in the capped estimation pool. All non-ID estimators continue to use the serialized history, even for “unseen” users.
Table 8. Accuracy by capped-estimation user overlap
Test group
\(N\)
Popularity
ID low-rank
Pooled classifier
Conditional logit
Selected Qwen
User represented
737
0.406
0.290
0.408
0.459
0.463
User absent
263
0.399
0.285
0.437
0.468
0.529
The 6.1-point Qwen-minus-conditional-logit difference among absent users is larger than the 0.4-point difference among represented users. That pattern rules out a simple story in which Qwen’s aggregate result is driven by memorizing capped-sample identifiers, because the LM never receives those identifiers and its point advantage is larger when they are absent. It is not a causal cold-start effect: the groups differ compositionally, every row still contains a history, and no subgroup-specific setting was preregistered.
6.5 What calibration changes
Affine calibration is selected because it has the best out-of-fold validation accuracy, NLL, and Brier score. Prior shift happens to classify one additional test row, but test accuracy is not a permissible reason to overturn a validation choice. For a calibrator, the superior affine test NLL is also more relevant than a one-row accuracy difference.
Table 9. Adapted-model calibration
Transformation
OOF validation accuracy
OOF validation NLL
Test accuracy
Test NLL
Test Brier
Direct
0.445
1.388
0.462
1.381
0.689
Temperature
0.445
1.388
0.462
1.378
0.685
Prior shift
0.444
1.370
0.481
1.356
0.679
Affine
0.451
1.361
0.480
1.345
0.673
Figure 3 compares top-choice confidence with realized top-choice accuracy in five equal-count test bins. This diagnostic is descriptive because the same test outcomes define the displayed reliability points; calibration parameters themselves were estimated only on validation. Affine calibration lowers this top-choice expected calibration error from 0.063 to 0.037.
Figure 3: Top-choice reliability before and after affine calibration. Each point contains 200 test observations; the diagonal is perfect calibration. ECE is the equal-count-bin weighted absolute confidence gap.
6.6 Quantization robustness
The selected estimator uses a Q4_K_M inference base because that is the deployable local object. A full post hoc BF16 pass applies the same step-200 adapter, scale 1.25, demonstrations, grammar, and test identifiers. It does not retune the adapter or calibration. The BF16 row is therefore a robustness comparison, not a new selected specification.
Table 10. Q4_K_M and BF16 inference parity
Base and probability map
Accuracy
Macro F1
NLL
Brier
Macro AUC
Q4_K_M, raw
0.462
0.460
1.381
0.689
0.751
BF16, raw
0.462
0.462
1.370
0.683
0.753
Q4_K_M, validation-fitted affine
0.480
0.480
1.345
0.673
0.751
BF16, same fixed Q4 affine map
0.477
0.477
1.345
0.671
0.753
Aggregate performance is stable, but the two numerical representations are not rowwise identical. Raw argmax agreement is 0.908 and affine argmax agreement is 0.910. Mean absolute probability difference is 0.0195, mean rowwise total-variation distance is 0.0487, and gold-probability correlation is 0.984. Raw Q4-minus-BF16 accuracy is exactly zero with paired interval \([-0.014,0.014]\); under the fixed affine map it is 0.003 with interval \([-0.011,0.017]\). The result supports Q4 deployment at the aggregate level while warning against treating quantization as bitwise or observation-level invariance.
6.7 Computational requirements
All local LM results use an NVIDIA RTX 5070 with 12 GB of memory. The Q4_K_M base is 3.01 GB and the selected adapter is 6.29 MB; the BF16 parity base is approximately 8.4 GB. Table 11 separates fitting, model selection, and scoring rather than quoting only a favorable inference number.
Table 11. Recorded compute
Operation
Sample or passes
Recorded time
Qualification
Conditional-logit ridge grid
6 fits on 4,000 rows
0.42 s CPU cumulative
Selected fit itself takes 0.07 s
ID low-rank rank/penalty grid
16 fits on 4,000 rows
88.6 s CPU cumulative
Selected fit itself takes 3.80 s
Pooled-classifier fit
4,000 rows
7.93 s CPU
235 standardized covariates and five logits
Final QLoRA continuation
steps 101–200
624.5 s GPU
Retained log does not identify total time for all three staged horizons
Adapter setting search
9 passes of 200 rows
about 781 s GPU
Includes one repeated setting
Frozen complete-suffix test scoring
1,000 rows
218.1 s GPU
4.59 rows/s; common-prefix branching
Selected Q4 validation scoring
1,000 rows
447.9 s GPU
Five demonstrations
Selected Q4 test scoring
1,000 rows
450.8 s GPU
2.22 rows/s
BF16 parity test scoring
1,000 rows
384.6 s GPU
Post hoc robustness pass
The median selected-model prompt has 5,616 tokens, and llama.cpp reuses a median 4,639 cached prompt tokens from the fixed demonstrations. At the observed Q4 rate, one million similarly sized choice occasions would require roughly 125 serial GPU-hours before engineering for parallel devices or batching. The local method is feasible, not free. Conditional logit is the obvious operational default when its hand-built feature map is adequate; the local LM buys a text-native representation at several orders of magnitude more computation.
7 Interpretation
Four conclusions survive the stronger benchmark.
First, symmetry matters more than model branding. The pooled position classifier wastes parameters by learning five versions of relationships that should be exchangeable. The 45-coefficient conditional logit improves its accuracy by 4.5 points and NLL by 0.115. It essentially matches raw QLoRA accuracy and slightly outperforms the final LM on every reported probability score. For an analyst who already has the 45 candidate attributes, conditional logit is the best default here.
Second, pretrained representation is not outcome alignment. Frozen Qwen is weak despite seeing titles, tags, histories, and reviews. Five demonstrations and post-hoc calibration repair some interface error but do not make it competitive. The large empirical movement accompanies outcome-supervised low-rank adaptation: raw accuracy rises 17.6 points relative to the closest frozen five-shot token specification, with the disclosed prompt-wrapper change. This is consistent with the adapter learning a restricted correction \(G_{ij}\) to represented signals. It does not reveal whether that correction is “taste,” popularity, genre duration, or another predictive regularity.
Third, calibration is a small econometric model, not cosmetic post-processing. The affine layer estimates a common log-probability slope and position constants. It adds 1.8 accuracy points, materially improves reliability, and exposes a large first-position measurement offset. Those gains should be attributed to validation labels and five fitted calibration degrees of freedom, not to the base LM.
Fourth, the LM and conditional logit are complements before they are rivals. Qwen is uniquely correct on 100 rows and conditional logit on 81. A combined model could use the conditional index, LM index, or both as alternative-varying regressors and estimate their weights on validation data. That design would test whether text contributes beyond the structured attribute benchmark without pretending the current full-pipeline comparison holds information fixed.
The external calls reinforce the adaptation point. A general-purpose model with five demonstrations is not matched on training information to an estimator fitted on 4,000 outcomes. Its lower accuracy says that persistent task estimation is valuable here; it does not rank general intelligence. Similarly, the popularity baseline shows that a nontrivial share of the task is generic item comparison. Separating popularity, semantic match, and true individualization requires item-held-out evaluation and targeted feature ablations.
An operational decision rule follows. Use conditional logit when the alternative attributes are available and probability quality, speed, or interpretability matter. Add a local adapter when important information is genuinely textual, the same narrow outcome recurs often enough to amortize training, and a held-out gain survives a symmetric benchmark. Use calibration whenever probabilities rather than only rankings matter. Do not use this fitted index for price elasticities, welfare, assortment counterfactuals, or policy simulation without a separate behavioral and identification argument.
8 Limitations
The first limitation is the estimand. Highest lifetime playtime among five already-owned games is a revealed-use ranking, not a purchase at a known menu. It combines preference, exposure, ownership duration, and genre-specific time intensity.
Second, the alternatives are analyst-sampled from a 500-game catalog. Probabilities are conditional on that sampled set and cannot be aggregated into market shares without modeling the sampling and availability process. Catalog rank and popularity are computed from full-source ownership before splitting, so the experiment is transductive with respect to those product-level fields.
Third, the main uncertainty intervals condition on the selected pipeline. They omit validation search, the unusual shortest-shot rule, adapter seed variation, calibration estimation, prompt and verbalizer choice, and pretraining variation. There is no random-shot distribution, item-held-out test, or independently repeated adapter fit.
Fourth, standard SFT does not optimize the reported five-way likelihood. Two of three intended response tokens are chat termination material, and 22 right-truncated estimation rows lose the label entirely. The staged run covers only 1,600 row presentations and reanneals its cosine schedule. The estimate is a documented engineering run, not a converged solution to the ideal five-class empirical-risk problem.
Fifth, complete pipelines receive different representations. Conditional logit receives 45 engineered candidate attributes; Qwen receives a textual serialization containing some of those values plus titles, tags, and reviews. Qwen also uses five labels in context and only 1,600 stochastic row presentations, whereas the classical estimators use all 4,000 labels directly. The comparison is practically relevant but does not isolate functional form, text, or effective sample size.
Sixth, user overlap is not randomly assigned. The stronger point accuracy among users absent from the capped estimation pool may reflect composition, and it does not establish cold-start superiority. The ID low-rank result is especially sensitive to sparse observations per user and its unseen-user fallback.
Seventh, the exercise covers one dataset, one four-billion-parameter architecture, one adapter parameterization, one Q4 format, and one GPU. The BF16 parity pass supports aggregate quantization robustness for this run, not a general invariance claim.
Finally, finite-set normalization creates coherent predictive probabilities but not a structural random-utility model. The LM index can depend on the entire prompt, displayed order, and arbitrary verbalizers; it need not satisfy regularity or stable substitution restrictions. Endogenous price, exposure, and assortment remain endogenous after passing through a neural network. Instruments, experimental variation, supply-side structure, or explicit assignment assumptions are still required for causal demand analysis.
9 Conclusion
A quantized four-billion-parameter LM can be converted into a finite-choice probability estimator by scoring admissible labels, conditioning on the displayed set, and fitting any adaptation or calibration only on non-test outcomes. Frozen likelihoods are not competitive in this application. Five demonstrations and calibration help modestly; a 3.15-million-parameter QLoRA adapter supplies the large improvement; and a five-parameter affine layer improves probability alignment.
The resulting 0.480 test accuracy is higher than the popularity rule, the pooled position classifier, conditional logit, and the retained full-sample API classifiers. The scientifically important qualification is that symmetric conditional logit already reaches 0.461 and has slightly better NLL, Brier score, and AUC. The 1.9-point paired accuracy difference is not statistically distinguishable from zero under test-sample resampling. The local model is therefore competitive with a strong conventional estimator, not shown to dominate it.
For empirical economists, the useful object is a nonlinear systematic-index engine with an explicit ledger. Frozen scoring exposes the pretrained index; ICL conditions it on labeled vignettes; QLoRA estimates a restricted domain correction; affine calibration estimates a familiar scale and alternative-specific constants. Neural low rank is a regularizer in parameter space, not identified low-dimensional taste heterogeneity. The method is promising for recurring text-rich prediction problems, provided it is benchmarked against an exchangeable choice model and never mistaken for a license to conduct structural counterfactuals.
10 Appendix A: Activation-steering diagnostic
Activation steering (Turner et al. 2024) was the fourth and lowest-priority route. It first selected a transformer block with an auxiliary probe. For each of the first 1,000 estimation and validation rows, features are the standardized final-prompt hidden state \(\widetilde h_{i\ell}\) of the frozen model and labels are the same five position outcomes \(y_{ij}\). At candidate block \(\ell\), scikit-learn estimates intercepts and class coefficients by minimizing
Intercepts are unpenalized. Blocks 31, 28, and 24 and \(C\in\{0.01,0.1,1,10\}\) are compared by accuracy on 1,000 validation rows after fitting on 1,000 estimation rows. Block 24 with \(C=0.01\) reaches 0.267, versus blockwise maxima 0.240 and 0.237. The probe is only a layer-selection device.
The direction uses a separate supervised contrast. Let \(\bar h_{i\ell}(j)\) be the layer-\(\ell\) hidden state averaged over the complete assistant-suffix tokens after presenting query \(q_i\) and candidate response choice_j. Define the cyclic counterlabel \(c(j)=j+1\) for \(j<5\) and \(c(5)=1\). On the first 1,000 estimation rows,
The feature is therefore the frozen hidden state induced by the serialized covariates; \(J_i\) supplies the positive suffix and the cyclic label supplies the negative suffix. No classifier or differentiable loss fits \(v_\ell\): it is a difference of sample means followed by normalization. The source is frozen Q4 Qwen, maximum sequence length is 6,144, and the vector is exported to GGUF control-vector slot 24. At llama.cpp inference the same intervention is added at every token position in block 24,
\[
h'_{i,24,t}=h_{i,24,t}+\gamma v_{24}.
\]
Scales are compared on the fixed 200-row model-setting validation subset. No result is selected for test.
Steering scale \(\gamma\)
Accuracy
Macro F1
NLL
Macro AUC
0
0.320
0.301
1.937
0.572
-2
0.325
0.305
1.889
0.569
-4
0.325
0.305
1.853
0.567
-8
0.320
0.293
1.812
0.566
The direction changes confidence but not classification materially and degrades AUC monotonically. A single alternative-agnostic vector applied to every token is too restrictive to supply the candidate-specific correction learned by QLoRA.
11 Appendix B: Executable replication pipeline
This document is the sole public runner for the empirical analysis. Its eight stages follow the order in which a new specification should be considered: construct the data, estimate the classical ladder, extract frozen likelihoods, add in-context examples, estimate an adapter, calibrate frozen probabilities, run the appendix steering diagnostic, and assemble the comparison. Every stage is parameterized and resumable. The folded cells remain in the HTML so that the article doubles as an executable research record.
The default report mode validates and reads the retained row-level artifacts without loading a model. smoke mode executes a few live observations through each llama.cpp inference path. full mode reconstructs the 1,000-observation evaluation and the 25-to-100-to-200-step adapter schedule, reusing completed artifacts unless LLM_CHOICE_FORCE=1 is set:
# Fast, deterministic article render from retained artifacts.QUARTO_PYTHON=.venv/bin/python uv run quarto render \ reports/local_llm_discrete_choice.qmd# Live local-inference check on three observations per relevant stage.LLM_CHOICE_MODE=smoke QUARTO_PYTHON=.venv/bin/python uv run quarto render \ reports/local_llm_discrete_choice.qmd# Full GPU reproduction; set LLM_CHOICE_HF_BASE for adapter conversion.LLM_CHOICE_MODE=full QUARTO_PYTHON=.venv/bin/python uv run quarto render \ reports/local_llm_discrete_choice.qmd
Replication mode: report; live model execution: False
artifact
path
available
0
processed-data metadata
results/data/processed/steam_v1_rich_history_c...
True
1
classical benchmark metrics
results/choice_model_benchmark_steam_v1_rich_h...
True
2
conditional logit predictions
results/choice_model_benchmark_steam_v1_rich_h...
True
3
results index
results/local_llm_discrete_choice/results_summ...
True
4
paired comparisons
results/local_llm_discrete_choice/paired_compa...
True
5
quantization parity
results/local_llm_discrete_choice/quantization...
True
6
accuracy comparison
results/local_llm_discrete_choice/accuracy_com...
True
7
adapted direct probabilities
results/local_llm_discrete_choice/adapted_dire...
True
8
adapted affine probabilities
results/local_llm_discrete_choice/adapted_affi...
True
9
API benchmark runs
results/local_llm_discrete_choice/api_benchmar...
True
10
API hard-label predictions
results/local_llm_discrete_choice/api_hard_lab...
True
11.1 Stage 1: Construct and validate the choice data
This stage reconstructs the candidate sets only in full mode. Every render reads the processed split files, verifies their labels against the common choice table, and reports the realized split sizes. This is the point at which the candidate exclusions and the definition of \(Y_i\) become concrete data operations.
11.2 Stage 2: Estimate the classical benchmark ladder
Full mode fits the pooled position classifier, symmetric conditional logit, and ID low-rank logit on the first 4,000 estimation observations. The retained artifact records each feature definition, validation grid, selected penalty or rank, and row-level probability table. Other modes verify and report those fits.
The native scorer branches after a shared prompt prefix, sums the autoregressive log likelihood of every complete assistant suffix, and applies a five-way softmax. It is compiled out of tree against the installed llama.cpp libraries; the upstream checkout is not modified.
Show replication code
exact_dir = paths.output_root /"exact_full"exact_metrics_path = exact_dir /"qwen35_4b_q4_k_m_exact_choice_likelihood_metrics.json"if LIVE and (not FULL or FORCE ornot exact_metrics_path.is_file()): scorer = build_native_scorer(paths) live_exact_dir = exact_dir if FULL else scratch /"exact" score_exact_choices( paths, output_dir=live_exact_dir, scorer_bin=scorer, max_eval_examples=EVALUATION_ROWS, splits=("validation", "test") if FULL else ("validation",), force=FORCE, )display(report_metric("complete finite-label likelihood"))
estimator
test_accuracy
test_macro_f1
test_log_loss
test_brier
0
complete finite-label likelihood
0.286
0.281875
1.839456
0.878404
11.4 Stage 4: Add in-context examples
Five short, class-balanced estimation cases are serialized as alternating user and assistant turns. The response grammar permits only the one-token verbalizers A–E; classification uses the returned token probabilities, not the sampled token.
Show replication code
icl_dir = paths.output_root /"icl_shortest_full"icl_metrics_path = icl_dir / ("qwen35_4b_q4_k_m_5shot_replace_minimal_turns_""balanced_shortest_seed42_metrics.json")if LIVE and (not FULL or FORCE ornot icl_metrics_path.is_file()): live_icl_dir = icl_dir if FULL else scratch /"icl" specification = EndpointSpecification( output_dir=live_icl_dir, model_label="qwen35_4b_q4_k_m", n_shots=5, prompt_style="minimal", shot_selection="balanced-shortest", max_train_examples=256, max_eval_examples=EVALUATION_ROWS, )with llama_server(paths, port=8080) as endpoint: score_endpoint_choices( paths, endpoint, specification, splits=("validation", "test") if FULL else ("validation",), force=FORCE, )display(report_metric("five-example one-token likelihood"))
estimator
test_accuracy
test_macro_f1
test_log_loss
test_brier
0
five-example one-token likelihood
0.286
0.272994
1.829927
0.884248
11.5 Stage 5: Estimate and apply QLoRA
The adapter runner continues a single rank-16 optimization through checkpoints 25, 100, and 200. The selected checkpoint is converted to a 6.29 MB GGUF adapter, applied at validation-selected scale 1.25, and evaluated with the same five demonstrations and one-token probability rule.
Show replication code
adapter_dir = ROOT /"outputs/qwen35_4b_steam_v1_rich_history_choice_letters_lora"adapter = adapter_dir /"adapter-step200-f16.gguf"if FULL and (FORCE ornot adapter.is_file()): adapter = train_qlora_schedule(paths, reuse_existing=not FORCE)[-1]lora_dir = paths.output_root /"lora200_full"lora_metrics_path = lora_dir / ("qwen35_4b_q4_k_m_lora200_training_prompt_5shot_replace_training_""turns_balanced_shortest_seed42_lora0_scale1_25_metrics.json")if LIVE and (not FULL or FORCE ornot lora_metrics_path.is_file()):ifnot adapter.is_file():raiseFileNotFoundError(f"Adapter missing: {adapter}") live_lora_dir = lora_dir if FULL else scratch /"lora200" specification = EndpointSpecification( output_dir=live_lora_dir, model_label="qwen35_4b_q4_k_m_lora200_training_prompt", n_shots=5, prompt_style="training", lora_id=0, lora_scale=1.25, shot_selection="balanced-shortest", max_train_examples=256, max_eval_examples=EVALUATION_ROWS, )with llama_server(paths, adapter=adapter, port=8081) as endpoint: score_endpoint_choices( paths, endpoint, specification, splits=("validation", "test") if FULL else ("validation",), force=FORCE, )display(report_metric("rank-16 adapter plus five examples"))
estimator
test_accuracy
test_macro_f1
test_log_loss
test_brier
0
rank-16 adapter plus five examples
0.462
0.459677
1.380981
0.689149
11.6 Stage 6: Calibrate frozen probability tables
No network parameter is updated here. For each estimator, validation probabilities and labels fit temperature, prior-shift, and affine transformations. Five-fold out-of-fold validation metrics select the reported transformation; the selected map is then refitted on all validation observations and applied once to test probabilities.
Show replication code
if FULL: calibration_inputs = {"exact_full": exact_dir,"icl_shortest_full": icl_dir,"lora200_full": lora_dir, }for name, source_dir in calibration_inputs.items(): prediction_files =sorted(source_dir.glob("*_predictions.csv")) validation_path =next(path for path in prediction_files if"_validation_"in path.name) test_path =next(path for path in prediction_files if"_test_"in path.name) calibration_path = source_dir /"calibration"if FORCE ornot (calibration_path /"calibration_metrics.json").is_file(): calibrate_choice_probabilities( validation_path, test_path, calibration_path, cv_folds=5, l2=1.0e-3, )display(report_metric("rank-16 adapter plus five examples and affine calibration"))
estimator
test_accuracy
test_macro_f1
test_log_loss
test_brier
0
rank-16 adapter plus five examples and affine ...
0.48
0.47984
1.344887
0.67256
11.7 Stage 7: Run the appendix activation-steering diagnostic
The steering route estimates a global contrast direction from 1,000 labeled estimation cases at transformer block 24, exports the normalized vector to GGUF, and applies it to all token positions during exact scoring. The fixed 200-observation validation subset is used to compare scales \(-2\), \(-4\), and \(-8\); test data are not used for selection.
Show replication code
steering_source_dir = ROOT / ("outputs/choice_model_benchmark_steam_v1_rich_history_qwen35_4b_""steering_train1000_val200_test1000_accuracy_staged_layer24")steering_npz = steering_source_dir /"steering_vectors.npz"control_vector = paths.llama_cpp.parent /"models/qwen35_4b_rich_history_global_layer24_cvector.gguf"if FULL and (FORCE ornot steering_npz.is_file()): steering_npz = estimate_steering_vector(paths, output_dir=steering_source_dir)if FULL and (FORCE ornot control_vector.is_file()): export_control_vector( steering_npz, key="global__layer_24", layer=24, output_path=control_vector, gguf_python_dir=paths.llama_cpp /"gguf-py", )steering_scales = (-2.0, -4.0, -8.0) if FULL else (-4.0,)scales_to_score = []for scale in steering_scales: scale_label =str(abs(int(scale))) candidate_dir = ( paths.output_root /"steering_validation200"/f"alpha_minus{scale_label}"if FULLelse scratch /f"steering_alpha_minus{scale_label}" )if FORCE ornotany(candidate_dir.glob("*_metrics.json")): scales_to_score.append((scale, scale_label, candidate_dir))if LIVE and scales_to_score: scorer = build_native_scorer(paths)ifnot control_vector.is_file():raiseFileNotFoundError(f"Control vector missing: {control_vector}")for scale, scale_label, live_steering_dir in scales_to_score: score_exact_choices( paths, output_dir=live_steering_dir, scorer_bin=scorer, max_eval_examples=200if FULL else EVALUATION_ROWS, splits=("validation",), control_vector=control_vector, control_vector_scale=scale, control_vector_layer_range=(24, 24), method_label=("exact_choice_likelihood_cvector_global_layer24_"f"alpha_minus{scale_label}_all_tokens" ), force=FORCE, )display( report_results.loc[ report_results["family"].eq("activation_steering"), ["estimator", "validation_accuracy", "validation_log_loss"], ])
estimator
validation_accuracy
validation_log_loss
13
unsteered complete finite-label likelihood
0.320
1.936880
14
global layer 24 all-token scale -2
0.325
1.889114
15
global layer 24 all-token scale -4
0.325
1.852697
16
global layer 24 all-token scale -8
0.320
1.811995
11.8 Stage 8: Assemble and verify the result table
The final stage reads the promoted result index, the multi-reference paired comparison, quantization parity, and external-call archive. The assertions make the numerical claims in the prose fail loudly if retained artifacts and manuscript drift apart. Rendering never issues a new external request.
The repository retains processed-data metadata; row-level pooled, conditional-logit, low-rank, selected-Qwen direct, and selected-Qwen affine probabilities; the compact result index; paired comparisons; the BF16/Q4 parity summary; and row-level hard-label calls for Codex Spark and GPT-5.5 under results/. The API archive includes observation identifiers, gold labels, parsed predictions, validity indicators, and run settings; repeated raw response payloads are omitted. Larger intermediate probability tables, calibration fits, adapters, and GGUF files remain machine-local under outputs/ and the configured llama.cpp model directory. Experiment-specific orchestration is contained in this document and its private helper, _local_llm_discrete_choice.py. Files under scripts/ are the generic data-construction, estimation, adapter-training, and hidden-state engines called by the numbered stages.
References
Bao, Keqin, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. “TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation.”arXiv Preprint arXiv:2305.00447. https://doi.org/10.48550/arXiv.2305.00447.
Compiani, Giovanni, Ilya Morozov, and Stephan Seiler. 2026. “Demand Estimation with Text and Image Data.”The RAND Journal of Economics. https://doi.org/10.1111/1756-2171.70052.
Dettmers, Tim, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. “QLoRA: Efficient Finetuning of Quantized LLMs.”Advances in Neural Information Processing Systems 36. https://doi.org/10.48550/arXiv.2305.14314.
Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. “On Calibration of Modern Neural Networks.” In Proceedings of the 34th International Conference on Machine Learning, 70:1321–30. Proceedings of Machine Learning Research. PMLR. https://proceedings.mlr.press/v70/guo17a.html.
Han, Yafei, Francisco C. Pereira, Moshe Ben-Akiva, and Christopher Zegras. 2022. “A Neural-Embedded Discrete Choice Model: Learning Taste Representation with Strengthened Interpretability.”Transportation Research Part B: Methodological 163: 166–86. https://doi.org/10.1016/j.trb.2022.07.001.
Holtzman, Ari, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. “Surface Form Competition: Why the Highest Probability Answer Isn’t Always Right.” In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 7038–51. Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.564.
Horton, John J., Apostolos Filippas, and Benjamin S. Manning. 2023. “Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?” Working Paper 31122. National Bureau of Economic Research. https://doi.org/10.3386/w31122.
Hu, Edward J., Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. “LoRA: Low-Rank Adaptation of Large Language Models.”arXiv Preprint arXiv:2106.09685. https://doi.org/10.48550/arXiv.2106.09685.
Kallus, Nathan, and Madeleine Udell. 2016. “Revealed Preference at Scale: Learning Personalized Preferences from Assortment Choices.” In Proceedings of the 2016 ACM Conference on Economics and Computation, 821–37. Association for Computing Machinery. https://doi.org/10.1145/2940716.2940794.
McFadden, Daniel. 1974. “Conditional Logit Analysis of Qualitative Choice Behavior.” In Frontiers in Econometrics, edited by Paul Zarembka, 105–42. Academic Press. https://escholarship.org/uc/item/0p99072q.
Min, Sewon, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. “Rethinking the Role of Demonstrations: What Makes in-Context Learning Work?” In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11048–64. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.759.
Pathak, Apurva, Kshitiz Gupta, and Julian McAuley. 2017. “Generating and Personalizing Bundle Recommendations on Steam.” In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1073–76. https://doi.org/10.1145/3077136.3080724.
Sifringer, Brian, Virginie Lurkin, and Alexandre Alahi. 2020. “Enhancing Discrete Choice Models with Representation Learning.”Transportation Research Part B: Methodological 140: 236–61. https://doi.org/10.1016/j.trb.2020.08.006.
Turner, Alexander Matt, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. 2024. “Steering Language Models with Activation Engineering.”arXiv Preprint arXiv:2308.10248. https://doi.org/10.48550/arXiv.2308.10248.
van Cranenburgh, Sander, Shenhao Wang, Akshay Vij, Francisco Pereira, and Joan Walker. 2022. “Choice Modelling in the Age of Machine Learning: Discussion Paper.”Journal of Choice Modelling 42: 100340. https://doi.org/10.1016/j.jocm.2021.100340.
Zhao, Zihao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. “Calibrate Before Use: Improving Few-Shot Performance of Language Models.” In Proceedings of the 38th International Conference on Machine Learning, 139:12697–706. Proceedings of Machine Learning Research. PMLR. https://proceedings.mlr.press/v139/zhao21c.html.