Saying 70, meaning 40: guessing probabilities with five mid-tier models

We asked five mid-tier models for probabilities on 3,010 resolved stock-price questions, each asked twice, on two scales. Six of the ten sets of stated probabilities told us no more than a coin would.

We asked five mid-tier models for probabilities on 3,010 resolved stock-price questions, each asked twice, on two scales. Six of the ten sets of stated probabilities told us no more than a coin would. So we looked inside the open-source models; and the probability of answering “yes” turns out to be a different number from the stated probability that yes is true — different by twenty points and more.

39,196
model answers collected
3,010
questions, every one resolved
7
models on the panel, five in depth
1,520
answers of “72” from one model
29%
chance of a yes from the model that says 60 on average

What calibration is

A forecaster is calibrated when the things they call 70% happen seven times in ten. Not once — over many calls. Calibration is the property that makes a probability a measurement rather than a mood.

Language models state a probability whenever you ask for one. Whether the number means anything is an actively studied open problem. A model’s internals encode more about whether an answer is right than the answer shows; and the numbers models state are overconfident, and change with how you ask.

In finance the stakes are mechanical. Position sizing, risk limits and expected-value arithmetic consume probabilities literally, so a confident-sounding ninety that means sixty results in a mispriced position. Whether you should ask a language model for a probability at all is a separate question, and our general answer is no. Either way, you want to know what the numbers it emits are made of.

The experiment

The question we started with: can an LLM put an accurate probability on a future event in a financial setting? We chose the simplest version of that question we could find. Show a model a stock’s recent price history. Ask for the probability that the price ends above a given level a few weeks out. Equities, because nothing resolves cleaner: thousands of questions, dates certain, and unambiguous answers.

So: 251 US stocks, one 250-day price history each, 3,010 distinct questions. Five mid-tier models — three open-weight, from three different labs, and two closed — answered every question twice, once as a percentage and once on a scale of nought to twenty. The three open-weight models also answered every question a third time, as a bare yes or no, with their internals recorded; that comparison is the last section of this note. We also ran a small test on two frontier models, Claude Opus 5 and GPT‑5.6‑sol: the first twenty questions, percentage only.

Four design guards, because each kills a familiar objection:

  • No leakage. Every outcome resolves after every open model’s release date, which we treat as its training cutoff.
  • No memory. Each price series and its threshold are multiplied by a shared random constant. The event is identical; the price level is unrecognisable.
  • Spread. Thresholds sit at volatility-scaled distances, so correct answers span the whole range instead of piling up at 50‑50.
  • No free ride. “Above” and “below” questions are balanced six-and-six per stock, so a rising market cannot gift every question the same resolution. The market did rise over the window our questions resolve in. The balance held the answer key at fifty-one per cent yes.

The three open models were picked to sit within five points of each other on a standard capability index — deliberately matched, so that differences in calibration are not a proxy for one model being smarter.

The words against reality

GPT-5.6-lunaBrier 0.184005050100100stated 5 · came true 7 in 100 · 76 answersstated 16 · came true 11 in 100 · 273 answersstated 24 · came true 17 in 100 · 132 answersstated 35 · came true 26 in 100 · 405 answersstated 45 · came true 34 in 100 · 372 answersstated 55 · came true 54 in 100 · 578 answersstated 65 · came true 74 in 100 · 645 answersstated 74 · came true 81 in 100 · 258 answersstated 84 · came true 90 in 100 · 152 answersstated 96 · came true 94 in 100 · 119 answersQwen3.5-27BBrier 0.232005050100100stated 0 · came true 4 in 100 · 133 answersstated 12 · came true 12 in 100 · 147 answersstated 45 · came true 26 in 100 · 149 answersstated 65 · came true 55 in 100 · 2,412 answersstated 87 · came true 84 in 100 · 153 answersstated 96 · came true 100 in 100 · 16 answersHaiku 4.5Brier 0.244005050100100stated 6 · came true 13 in 100 · 128 answersstated 15 · came true 18 in 100 · 147 answersstated 26 · came true 37 in 100 · 551 answersstated 34 · came true 47 in 100 · 171 answersstated 42 · came true 57 in 100 · 212 answersstated 52 · came true 100 in 100 · 1 answersstated 62 · came true 53 in 100 · 15 answersstated 72 · came true 59 in 100 · 1,694 answersstated 86 · came true 84 in 100 · 19 answersstated 92 · came true 76 in 100 · 72 answersKimi K2.5Brier 0.260005050100100stated 0 · came true 41 in 100 · 258 answersstated 15 · came true 29 in 100 · 52 answersstated 25 · came true 35 in 100 · 71 answersstated 35 · came true 35 in 100 · 127 answersstated 45 · came true 43 in 100 · 425 answersstated 53 · came true 54 in 100 · 1,414 answersstated 64 · came true 60 in 100 · 492 answersstated 74 · came true 63 in 100 · 128 answersstated 84 · came true 52 in 100 · 23 answersstated 94 · came true 67 in 100 · 6 answersDeepSeek V3.2Brier 0.267005050100100stated 0 · came true 39 in 100 · 160 answersstated 14 · came true 26 in 100 · 50 answersstated 24 · came true 39 in 100 · 85 answersstated 34 · came true 44 in 100 · 108 answersstated 44 · came true 42 in 100 · 159 answersstated 53 · came true 50 in 100 · 1,037 answersstated 64 · came true 54 in 100 · 760 answersstated 73 · came true 57 in 100 · 481 answersstated 83 · came true 61 in 100 · 111 answersstated 94 · came true 56 in 100 · 59 answers
Figure 1. Stated probability (x) against how often the answer was yes (y), percentage arm, ten bins. Dot area is the number of answers in the bin; the diagonal is honest; lines connect bins with at least twenty answers. Ordered best to worst. Hover any point, bar, or marker for its exact values. All numbers are the experiment’s measured results.

At first glance the panel looks fine. Every model’s average lands within a dozen points of the fifty-one per cent base rate. Averages are cheap; hedging gets you there.

The tails are where the money is, and the tails are bad. Haiku 4.5 answered “72” to half of everything it was asked. Those seventy-twos came true fifty-nine times in a hundred. When Kimi K2.5 wrote “0”, the thing it had ruled out happened two times in five. When DeepSeek V3.2 put the chance at nine in ten or better, it happened six times in ten.

Now score everything against the dullest possible competitor: write 50 on every line and go home. Six of the ten stated-probability sets did no better than that, and three did worse. The numbers are not quite pure static: a trace of real ordering runs through them. The signal is garbled, not absent. The rest of this note locates the garble.

These are last cycle’s mid-tier models; the exception is the newest on the panel. GPT‑5.6‑luna, released this July, is different in kind: when it says eighty-five or more, the thing happens ninety-two times in a hundred. Its one vice is humility. It says 74; the answer arrives 81.

ModelReleasedBrier, percentBrier, 0–20Against always-50
GPT-5.6-luna (OpenAI)Jul 20260.1840.184beats it, both scales
Qwen3.5-27B (Alibaba)Feb 20260.2320.207beats it, both scales
Claude Haiku 4.5 (Anthropic)Oct 20250.2440.245indistinguishable
Kimi K2.5 (Moonshot)Jan 20260.2600.256worse on percent; ties on 0–20
DeepSeek V3.2Dec 20250.2670.268worse on both
“50”, every line0.2500.250
Volatility formula0.144beats everything above

In the table above we report the Brier score for each model’s performance, a measure of how accurate its probability predictions were, across all 3,010 questions. Lower is better. The Brier score is simply the mean squared error: the average squared gap between the probability stated and what happened. Writing 50 on every line scores 0.250, and the last column says which gaps from that line are big enough to trust.

On the last row we display the results of a simple, deterministic formula for predicting future stock prices. The formula estimates the stock’s daily volatility from the closes in the prompt, assumes a random walk with no trend, and reads the probability of finishing past the threshold off a normal curve. It is by no means the most advanced way to forecast a price. But the formula’s predictions set a bar — a baseline level of forecasting competence, computed from the prompt alone — and they score 0.144. No mid-tier model reaches the bar. The closest is luna, at 0.184.

The small frontier test says the bar is reachable. On a twenty-question spot check, Claude Opus 5 and GPT‑5.6‑sol both roughly matched the formula’s score, and tracked its probability question by question. So the capability exists, and the failures above are a statement about a tier and a vintage — not about language models. Within our own panel, release date predicts calibration better than price does (the cheapest model was the best) and better than leaderboards do (Kimi and Qwen share a capability score; one is a coin, the other is usable). The tier that gets deployed at scale is the tier that cannot yet price a probability. To be explicit about what the frontier matched: ours is a straightforward textbook formula, not the last word. Sharper volatility estimates exist, and we tested none of them. Matching a textbook formula is not market-beating skill, and nothing in this note claims otherwise.

The words against each other

GPT-5.6-luna050100answered 1 on 10 of 3,010 questions (0%)answered 2 on 10 of 3,010 questions (0%)answered 3 on 8 of 3,010 questions (0%)answered 4 on 2 of 3,010 questions (0%)answered 5 on 5 of 3,010 questions (0%)answered 7 on 11 of 3,010 questions (0%)answered 8 on 27 of 3,010 questions (1%)answered 9 on 3 of 3,010 questions (0%)answered 10 on 4 of 3,010 questions (0%)answered 12 on 60 of 3,010 questions (2%)answered 13 on 1 of 3,010 questions (0%)answered 14 on 4 of 3,010 questions (0%)answered 15 on 22 of 3,010 questions (1%)answered 17 on 5 of 3,010 questions (0%)answered 18 on 176 of 3,010 questions (6%)answered 19 on 1 of 3,010 questions (0%)answered 20 on 25 of 3,010 questions (1%)answered 21 on 1 of 3,010 questions (0%)answered 22 on 12 of 3,010 questions (0%)answered 23 on 9 of 3,010 questions (0%)answered 24 on 14 of 3,010 questions (0%)answered 25 on 25 of 3,010 questions (1%)answered 26 on 1 of 3,010 questions (0%)answered 27 on 10 of 3,010 questions (0%)answered 28 on 33 of 3,010 questions (1%)answered 29 on 2 of 3,010 questions (0%)answered 30 on 15 of 3,010 questions (0%)answered 31 on 3 of 3,010 questions (0%)answered 32 on 30 of 3,010 questions (1%)answered 34 on 9 of 3,010 questions (0%)answered 35 on 293 of 3,010 questions (10%)answered 36 on 1 of 3,010 questions (0%)answered 37 on 7 of 3,010 questions (0%)answered 38 on 42 of 3,010 questions (1%)answered 39 on 5 of 3,010 questions (0%)answered 40 on 10 of 3,010 questions (0%)answered 42 on 105 of 3,010 questions (3%)answered 43 on 27 of 3,010 questions (1%)answered 44 on 10 of 3,010 questions (0%)answered 45 on 92 of 3,010 questions (3%)answered 46 on 4 of 3,010 questions (0%)answered 47 on 66 of 3,010 questions (2%)answered 48 on 57 of 3,010 questions (2%)answered 49 on 1 of 3,010 questions (0%)answered 50 on 27 of 3,010 questions (1%)answered 51 on 4 of 3,010 questions (0%)answered 52 on 66 of 3,010 questions (2%)answered 53 on 11 of 3,010 questions (0%)answered 54 on 20 of 3,010 questions (1%)answered 55 on 246 of 3,010 questions (8%)answered 56 on 4 of 3,010 questions (0%)answered 57 on 63 of 3,010 questions (2%)answered 58 on 134 of 3,010 questions (4%)answered 59 on 3 of 3,010 questions (0%)answered 60 on 31 of 3,010 questions (1%)answered 61 on 15 of 3,010 questions (0%)answered 62 on 135 of 3,010 questions (4%)answered 63 on 27 of 3,010 questions (1%)answered 64 on 22 of 3,010 questions (1%)answered 65 on 250 of 3,010 questions (8%)answered 66 on 3 of 3,010 questions (0%)answered 67 on 21 of 3,010 questions (1%)answered 68 on 139 of 3,010 questions (5%)answered 69 on 2 of 3,010 questions (0%)answered 70 on 33 of 3,010 questions (1%)answered 71 on 1 of 3,010 questions (0%)answered 72 on 84 of 3,010 questions (3%)answered 73 on 16 of 3,010 questions (1%)answered 74 on 9 of 3,010 questions (0%)answered 75 on 41 of 3,010 questions (1%)answered 76 on 2 of 3,010 questions (0%)answered 77 on 3 of 3,010 questions (0%)answered 78 on 64 of 3,010 questions (2%)answered 79 on 5 of 3,010 questions (0%)answered 80 on 13 of 3,010 questions (0%)answered 81 on 1 of 3,010 questions (0%)answered 82 on 32 of 3,010 questions (1%)answered 83 on 4 of 3,010 questions (0%)answered 84 on 6 of 3,010 questions (0%)answered 85 on 75 of 3,010 questions (2%)answered 86 on 3 of 3,010 questions (0%)answered 87 on 4 of 3,010 questions (0%)answered 88 on 12 of 3,010 questions (0%)answered 89 on 2 of 3,010 questions (0%)answered 90 on 5 of 3,010 questions (0%)answered 91 on 3 of 3,010 questions (0%)answered 92 on 15 of 3,010 questions (0%)answered 93 on 11 of 3,010 questions (0%)answered 94 on 4 of 3,010 questions (0%)answered 95 on 24 of 3,010 questions (1%)answered 96 on 9 of 3,010 questions (0%)answered 97 on 11 of 3,010 questions (0%)answered 98 on 6 of 3,010 questions (0%)answered 99 on 24 of 3,010 questions (1%)answered 100 on 7 of 3,010 questions (0%)“35” · 10%Qwen3.5-27B050100answered 0 on 133 of 3,010 questions (4%)answered 12 on 144 of 3,010 questions (5%)answered 15 on 3 of 3,010 questions (0%)answered 42 on 73 of 3,010 questions (2%)answered 45 on 2 of 3,010 questions (0%)answered 47 on 4 of 3,010 questions (0%)answered 48 on 70 of 3,010 questions (2%)answered 60 on 2 of 3,010 questions (0%)answered 62 on 981 of 3,010 questions (33%)answered 64 on 67 of 3,010 questions (2%)answered 68 on 1,362 of 3,010 questions (45%)answered 82 on 15 of 3,010 questions (0%)answered 87 on 127 of 3,010 questions (4%)answered 88 on 11 of 3,010 questions (0%)answered 92 on 6 of 3,010 questions (0%)answered 98 on 10 of 3,010 questions (0%)“68” · 45%Haiku 4.5050100answered 2 on 33 of 3,010 questions (1%)answered 5 on 9 of 3,010 questions (0%)answered 8 on 86 of 3,010 questions (3%)answered 12 on 5 of 3,010 questions (0%)answered 15 on 132 of 3,010 questions (4%)answered 18 on 10 of 3,010 questions (0%)answered 22 on 11 of 3,010 questions (0%)answered 23 on 20 of 3,010 questions (1%)answered 24 on 5 of 3,010 questions (0%)answered 25 on 329 of 3,010 questions (11%)answered 28 on 186 of 3,010 questions (6%)answered 32 on 54 of 3,010 questions (2%)answered 35 on 112 of 3,010 questions (4%)answered 38 on 5 of 3,010 questions (0%)answered 42 on 195 of 3,010 questions (6%)answered 45 on 16 of 3,010 questions (1%)answered 48 on 1 of 3,010 questions (0%)answered 52 on 1 of 3,010 questions (0%)answered 62 on 14 of 3,010 questions (0%)answered 68 on 1 of 3,010 questions (0%)answered 72 on 1,520 of 3,010 questions (50%)answered 73 on 4 of 3,010 questions (0%)answered 75 on 106 of 3,010 questions (4%)answered 78 on 64 of 3,010 questions (2%)answered 82 on 3 of 3,010 questions (0%)answered 85 on 8 of 3,010 questions (0%)answered 87 on 7 of 3,010 questions (0%)answered 89 on 1 of 3,010 questions (0%)answered 92 on 72 of 3,010 questions (2%)“72” · 50%Kimi K2.5050100answered 0 on 240 of 2,996 questions (8%)answered 1 on 4 of 2,996 questions (0%)answered 3 on 2 of 2,996 questions (0%)answered 5 on 8 of 2,996 questions (0%)answered 6 on 1 of 2,996 questions (0%)answered 7 on 2 of 2,996 questions (0%)answered 8 on 1 of 2,996 questions (0%)answered 10 on 2 of 2,996 questions (0%)answered 11 on 1 of 2,996 questions (0%)answered 12 on 7 of 2,996 questions (0%)answered 13 on 2 of 2,996 questions (0%)answered 14 on 3 of 2,996 questions (0%)answered 15 on 26 of 2,996 questions (1%)answered 16 on 2 of 2,996 questions (0%)answered 17 on 3 of 2,996 questions (0%)answered 18 on 5 of 2,996 questions (0%)answered 19 on 1 of 2,996 questions (0%)answered 20 on 7 of 2,996 questions (0%)answered 21 on 1 of 2,996 questions (0%)answered 22 on 1 of 2,996 questions (0%)answered 23 on 8 of 2,996 questions (0%)answered 24 on 1 of 2,996 questions (0%)answered 25 on 23 of 2,996 questions (1%)answered 27 on 14 of 2,996 questions (0%)answered 28 on 12 of 2,996 questions (0%)answered 29 on 4 of 2,996 questions (0%)answered 30 on 13 of 2,996 questions (0%)answered 31 on 1 of 2,996 questions (0%)answered 32 on 7 of 2,996 questions (0%)answered 33 on 9 of 2,996 questions (0%)answered 34 on 8 of 2,996 questions (0%)answered 35 on 62 of 2,996 questions (2%)answered 36 on 2 of 2,996 questions (0%)answered 37 on 4 of 2,996 questions (0%)answered 38 on 16 of 2,996 questions (1%)answered 39 on 5 of 2,996 questions (0%)answered 40 on 21 of 2,996 questions (1%)answered 41 on 6 of 2,996 questions (0%)answered 42 on 48 of 2,996 questions (2%)answered 43 on 11 of 2,996 questions (0%)answered 44 on 12 of 2,996 questions (0%)answered 45 on 159 of 2,996 questions (5%)answered 46 on 20 of 2,996 questions (1%)answered 47 on 58 of 2,996 questions (2%)answered 48 on 64 of 2,996 questions (2%)answered 49 on 26 of 2,996 questions (1%)answered 50 on 351 of 2,996 questions (12%)answered 51 on 105 of 2,996 questions (4%)answered 52 on 356 of 2,996 questions (12%)answered 53 on 47 of 2,996 questions (2%)answered 54 on 37 of 2,996 questions (1%)answered 55 on 274 of 2,996 questions (9%)answered 56 on 31 of 2,996 questions (1%)answered 57 on 25 of 2,996 questions (1%)answered 58 on 176 of 2,996 questions (6%)answered 59 on 12 of 2,996 questions (0%)answered 60 on 67 of 2,996 questions (2%)answered 61 on 26 of 2,996 questions (1%)answered 62 on 84 of 2,996 questions (3%)answered 63 on 30 of 2,996 questions (1%)answered 64 on 20 of 2,996 questions (1%)answered 65 on 156 of 2,996 questions (5%)answered 66 on 7 of 2,996 questions (0%)answered 67 on 19 of 2,996 questions (1%)answered 68 on 77 of 2,996 questions (3%)answered 69 on 6 of 2,996 questions (0%)answered 70 on 10 of 2,996 questions (0%)answered 71 on 8 of 2,996 questions (0%)answered 72 on 40 of 2,996 questions (1%)answered 73 on 10 of 2,996 questions (0%)answered 74 on 2 of 2,996 questions (0%)answered 75 on 42 of 2,996 questions (1%)answered 76 on 4 of 2,996 questions (0%)answered 78 on 10 of 2,996 questions (0%)answered 79 on 2 of 2,996 questions (0%)answered 81 on 1 of 2,996 questions (0%)answered 82 on 7 of 2,996 questions (0%)answered 83 on 4 of 2,996 questions (0%)answered 85 on 9 of 2,996 questions (0%)answered 87 on 2 of 2,996 questions (0%)answered 92 on 3 of 2,996 questions (0%)answered 95 on 2 of 2,996 questions (0%)answered 100 on 1 of 2,996 questions (0%)“52” · 12%DeepSeek V3.2050100answered 0 on 147 of 3,010 questions (5%)answered 2 on 2 of 3,010 questions (0%)answered 3 on 1 of 3,010 questions (0%)answered 5 on 8 of 3,010 questions (0%)answered 6 on 1 of 3,010 questions (0%)answered 9 on 1 of 3,010 questions (0%)answered 10 on 16 of 3,010 questions (1%)answered 12 on 2 of 3,010 questions (0%)answered 14 on 4 of 3,010 questions (0%)answered 15 on 21 of 3,010 questions (1%)answered 16 on 2 of 3,010 questions (0%)answered 17 on 3 of 3,010 questions (0%)answered 19 on 2 of 3,010 questions (0%)answered 20 on 16 of 3,010 questions (1%)answered 21 on 1 of 3,010 questions (0%)answered 22 on 6 of 3,010 questions (0%)answered 23 on 10 of 3,010 questions (0%)answered 24 on 6 of 3,010 questions (0%)answered 25 on 27 of 3,010 questions (1%)answered 26 on 5 of 3,010 questions (0%)answered 27 on 4 of 3,010 questions (0%)answered 28 on 6 of 3,010 questions (0%)answered 29 on 4 of 3,010 questions (0%)answered 30 on 27 of 3,010 questions (1%)answered 31 on 5 of 3,010 questions (0%)answered 32 on 1 of 3,010 questions (0%)answered 33 on 10 of 3,010 questions (0%)answered 34 on 9 of 3,010 questions (0%)answered 35 on 28 of 3,010 questions (1%)answered 36 on 11 of 3,010 questions (0%)answered 37 on 5 of 3,010 questions (0%)answered 38 on 9 of 3,010 questions (0%)answered 39 on 3 of 3,010 questions (0%)answered 40 on 48 of 3,010 questions (2%)answered 41 on 5 of 3,010 questions (0%)answered 42 on 12 of 3,010 questions (0%)answered 43 on 12 of 3,010 questions (0%)answered 44 on 5 of 3,010 questions (0%)answered 45 on 28 of 3,010 questions (1%)answered 46 on 8 of 3,010 questions (0%)answered 47 on 12 of 3,010 questions (0%)answered 48 on 16 of 3,010 questions (1%)answered 49 on 13 of 3,010 questions (0%)answered 50 on 326 of 3,010 questions (11%)answered 51 on 55 of 3,010 questions (2%)answered 52 on 71 of 3,010 questions (2%)answered 53 on 65 of 3,010 questions (2%)answered 54 on 60 of 3,010 questions (2%)answered 55 on 230 of 3,010 questions (8%)answered 56 on 47 of 3,010 questions (2%)answered 57 on 71 of 3,010 questions (2%)answered 58 on 66 of 3,010 questions (2%)answered 59 on 46 of 3,010 questions (2%)answered 60 on 193 of 3,010 questions (6%)answered 61 on 58 of 3,010 questions (2%)answered 62 on 60 of 3,010 questions (2%)answered 63 on 45 of 3,010 questions (1%)answered 64 on 38 of 3,010 questions (1%)answered 65 on 181 of 3,010 questions (6%)answered 66 on 56 of 3,010 questions (2%)answered 67 on 48 of 3,010 questions (2%)answered 68 on 51 of 3,010 questions (2%)answered 69 on 30 of 3,010 questions (1%)answered 70 on 132 of 3,010 questions (4%)answered 71 on 30 of 3,010 questions (1%)answered 72 on 39 of 3,010 questions (1%)answered 73 on 54 of 3,010 questions (2%)answered 74 on 25 of 3,010 questions (1%)answered 75 on 116 of 3,010 questions (4%)answered 76 on 28 of 3,010 questions (1%)answered 77 on 17 of 3,010 questions (1%)answered 78 on 26 of 3,010 questions (1%)answered 79 on 14 of 3,010 questions (0%)answered 80 on 47 of 3,010 questions (2%)answered 81 on 4 of 3,010 questions (0%)answered 82 on 9 of 3,010 questions (0%)answered 84 on 7 of 3,010 questions (0%)answered 85 on 33 of 3,010 questions (1%)answered 86 on 3 of 3,010 questions (0%)answered 87 on 2 of 3,010 questions (0%)answered 88 on 5 of 3,010 questions (0%)answered 89 on 1 of 3,010 questions (0%)answered 90 on 26 of 3,010 questions (1%)answered 92 on 2 of 3,010 questions (0%)answered 93 on 1 of 3,010 questions (0%)answered 94 on 2 of 3,010 questions (0%)answered 95 on 11 of 3,010 questions (0%)answered 96 on 1 of 3,010 questions (0%)answered 97 on 1 of 3,010 questions (0%)answered 99 on 2 of 3,010 questions (0%)answered 100 on 13 of 3,010 questions (0%)“50” · 11%
Figure 2. Where the answers landed: the share of each model’s 3,010 percentage answers on each whole number. Panels are scaled independently; the tallest spike in each is labelled. Qwen’s two favourites, 68 and 62, carry four answers in five between them. Hover any point, bar, or marker for its exact values. All numbers are the experiment’s measured results.

Look at the favourite numbers. “72” and “50” are things people write, at frequencies that have nothing to do with this stock. If a model’s number is partly a verbal habit, there is a natural test: take the familiar scale away. So we asked all 3,010 questions again on a scale of nought to twenty — nought impossible, ten an even chance, twenty certain. One scale point is five percentage points. A model that holds a probability and translates it loses nothing, because the ruler is arithmetic.

Things changed. DeepSeek’s average probability rose eight points on the new ruler. Haiku’s rose six. Qwen’s fell six. Kimi and luna barely moved. The favourite numbers did not disappear; they moved house, and twelve-out-of-twenty and fourteen-out-of-twenty became the new sixties and seventies. For Qwen the coarse ruler genuinely helped — its nought-to-twenty arm is the best in the mid panel. For DeepSeek it made the overconfidence worse. Same model, same question, different ruler, different number.

One regularity is worth recording. Rank the five models by how well their two scales agree with each other, and you have also ranked them by how well they agree with reality — the same order, exactly, five for five. Disagreement between a model’s own answers is a tell. This oughtn’t be surprising. A model that falls back on favourite numbers keeps different favourites in percentages and out of twenty, so its two answers drift apart; when they agree, a probability was common to both scales — computed, not recited. And there is one deeper place to look for it.

The words against the weights

Every answer a language model writes is a draw from a probability distribution over next words. So on a yes/no question, a number exists inside the model before the answer does: if “yes” carries 75%, you hear yes three times in four, and you never see the number. For open-weight models the number can be read directly. So we put every question to the three open models a third time, told them to answer in one word — yes or no — and recorded the probability the model gave each word. The two words absorb over ninety-nine per cent of the distribution.

Here is the plain fact, and it is the reason this note exists. The probability that a model answers “yes” is not the probability it states when you ask for one.

Qwen’s stated numbers average about sixty per cent. Its internal chance of answering yes averages twenty-nine. Ask it yes-or-no, and you would mostly hear no. On the questions Kimi rates better-than-even in words, it would answer no seven times in ten. DeepSeek is the strangest of the three: as its stated probability runs from nought to a hundred, its internal chance of saying yes barely moves — from 0.40 to 0.54.

02550751000255075100answers mirror the stated numbereven odds of answering “yes”Kimi K2.5: stated 0–10 · would answer yes 20% of the time · 258 questionsKimi K2.5: stated 10–20 · would answer yes 16% of the time · 52 questionsKimi K2.5: stated 20–30 · would answer yes 20% of the time · 71 questionsKimi K2.5: stated 30–40 · would answer yes 26% of the time · 127 questionsKimi K2.5: stated 40–50 · would answer yes 28% of the time · 425 questionsKimi K2.5: stated 50–60 · would answer yes 37% of the time · 1,414 questionsKimi K2.5: stated 60–70 · would answer yes 40% of the time · 492 questionsKimi K2.5: stated 70–80 · would answer yes 40% of the time · 128 questionsKimi K2.5: stated 80–90 · would answer yes 40% of the time · 23 questionsKimi K2.5: stated 90–100 · would answer yes 57% of the time · 6 questionsKimi K2.5DeepSeek V3.2: stated 0–10 · would answer yes 40% of the time · 160 questionsDeepSeek V3.2: stated 10–20 · would answer yes 31% of the time · 50 questionsDeepSeek V3.2: stated 20–30 · would answer yes 35% of the time · 85 questionsDeepSeek V3.2: stated 30–40 · would answer yes 40% of the time · 108 questionsDeepSeek V3.2: stated 40–50 · would answer yes 42% of the time · 159 questionsDeepSeek V3.2: stated 50–60 · would answer yes 48% of the time · 1,037 questionsDeepSeek V3.2: stated 60–70 · would answer yes 50% of the time · 760 questionsDeepSeek V3.2: stated 70–80 · would answer yes 51% of the time · 481 questionsDeepSeek V3.2: stated 80–90 · would answer yes 50% of the time · 111 questionsDeepSeek V3.2: stated 90–100 · would answer yes 54% of the time · 59 questionsDeepSeek V3.2Qwen3.5-27B: stated 0–10 · would answer yes 3% of the time · 133 questionsQwen3.5-27B: stated 10–20 · would answer yes 3% of the time · 147 questionsQwen3.5-27B: stated 40–50 · would answer yes 3% of the time · 149 questionsQwen3.5-27B: stated 60–70 · would answer yes 31% of the time · 2,412 questionsQwen3.5-27B: stated 80–90 · would answer yes 65% of the time · 153 questionsQwen3.5-27B: stated 90–100 · would answer yes 78% of the time · 16 questionsQwen3.5-27Bstated probability, per centchance of answering “yes”
Figure 3. The internal chance of answering “yes”, by the probability the same model states when asked. Marker area is the number of questions in the bin. A model whose answers mirrored its stated numbers would follow the diagonal; any coherent rule would cross the mid-line at fifty. Kimi crosses in the nineties. DeepSeek never truly leaves the middle. Hover any point, bar, or marker for its exact values. All numbers are the experiment’s measured results.

Which number should you believe? The scores suggest the inner one. Treat the stated probability and the internal probability as two forecasters answering the same questions, and the weights beat the words on every model we opened. Right more often, even at a naive better-than-even cut: fifty-nine against fifty-seven in a hundred for Kimi, sixty-four against fifty-five for DeepSeek, sixty-eight against sixty-one for Qwen. Better at ranking outcomes. Closer to the textbook probability. One caveat: the inner number is mis-centred — biased toward no, where the words lean yes — so it is not ready to use. Nor is it especially good: Brier 0.221 for DeepSeek and 0.230 for Qwen, ahead of their words; 0.258 for Kimi, no better than the coin; all three far from the formula’s 0.144.

Three after-notes.

The crossing point is a supermajority rule. Kimi does not tip past even odds internally until its stated number reaches the nineties; Qwen, around eighty. Asked yes-or-no, these models behave like a cautious committee — say yes only if sure — not like a forecaster reporting the likelier side.

The product implication is concrete. There are two common ways to read “the model’s view”: ask for a percentage, or sample the yes/no answer many times and vote. They read different channels. On these models the channels disagree about the same question, systematically. Two teams shipping “the same model” can ship different opinions.

And the divergence is less surprising in hindsight — this is interpretation, flagged as such. The words are trained on how people write about probability. The disposition is trained on which answers scored well. Nothing ties the two together. What we have shown is that two readouts disagree, and which one is more faithful. Which one is the model’s true belief — whether that is even a well-posed question — we do not claim.

What to do with it

The words and the weights disagree, and the divergence itself is the finding: one model, two readouts, and no reason yet to trust either. The weights scored better on our questions. Better, but not all that good: the one-line formula still beats them. The stated number is the weak layer, and the weak layer is also where the newest models have visibly improved: luna’s spoken numbers land, and the frontier’s track the textbook figure, on outcomes none of them had seen. This reads as a property of a vintage, not a law of the species.

So. If you consume model probabilities in finance, treat a stated number as untrusted output of an unfinished layer. Measure it per model and per elicitation format against resolved outcomes, or do not consume it at all; we lean towards the latter. A cheap first screen needs no outcomes: ask everything twice, on two scales, and treat self-disagreement as disqualifying. And if you train models, the translation from weights to words is a clean post-training target. Resolved outcomes and a proper scoring rule are objective, cheap, and hard to contaminate. Score the words and the disposition together, so that the words do not improve without the decisions.

Notes

  1. Construction: 251 windows of 250 trading days, twelve questions each; thresholds at volatility-scaled distances from spot; above/below balanced six-and-six per window. Two byte-identical prompt pairs were de-duplicated, leaving 3,010 distinct questions of the 3,012 generated.
  2. Leakage: outcomes resolve from 2026-04-01 onward, five-plus weeks after the newest open model’s release, with release date as a conservative stand-in for training cutoff. Series and thresholds are rescaled by a shared random constant, so no price level in any prompt matches history.
  3. One window: all outcomes fall between April and August 2026 — a single, upward-drifting regime. The direction balance is what held the answer key at fifty-one per cent yes. Calibration here is calibration in this window.
  4. The reference formula: a driftless random walk with exponentially weighted volatility computed from the closes in the prompt. Its Brier score is 0.144.
  5. Elicitation: temperature 1 throughout; reasoning off; no tools; one question per prompt. Per-question scatter between the two scales is partly sampling noise; the mean shifts are not. The 0–20 wording anchors ten as “an even chance” while the percentage wording anchors nothing, so part of the shift may be the anchor.
  6. Reading the internals: first answer token, forced one-word reply, probabilities renormalised over “Yes” and “No”, which absorb over ninety-nine per cent of the distribution. Closed APIs do not expose these numbers. Several other open-weight families pin the first-token distribution to nearly 0 or 1 and cannot be measured this way; the panel is the three that can be.

© Dissei. All rights reserved. No reproduction, adaptation, or derivative use of this content or methodology without prior written permission.