Dissei / Research

Research

What models learn from finance.
Where their judgment breaks.

Explore the studies

From findings to practice.

Practitioner guides

Studies in financial judgment

Saying 70, meaning 40: guessing probabilities with five mid-tier models

We asked five mid-tier models for probabilities on 3,010 resolved stock-price questions, each asked twice, on two scales. Six of the ten sets of stated probabilities told us no more than a coin would.

Sitting their own exam: four models, 397 questions, and a mild case of self-bias

Four models each wrote an examination from seven anonymised financial cases; all four then sat every paper blind. Who answers moves the score by 32 points. Who asks moves it by 6. On four fast, inexpensive models and questions that mostly rewarded lookup, writing the paper yourself is worth under a point, at most two and a half at the outside, and the lone hint of a larger edge sits exactly where a question demands a chain of thought.

One case, worked end to end: a lender, a platform, and a demand shock

A specialist credit fund, a consumer platform whose revenue stopped arriving, and a known outcome. The same financial case looks radically different as the available evidence moves from formation to shock, financing and aftermath.

Built backwards from outcomes: ground truth that is a matter of record

Tasks built from documented outcomes make the reference a matter of record rather than an author's opinion. Point-in-time controls distinguish contemporaneous reasoning from hindsight, while contamination is tested rather than assumed away.

Where equal-looking models come apart: two frontier labs, 282 rollouts, one deal

Claude Sonnet 5 and GLM-5 faced the same 47 tasks, with three rollouts each. Mean reward was 0.54 against 0.47, Sonnet 5 led on 34 of 47 tasks, and hindsight appeared in about one answer in eight versus one in five.

Does the model know the deal, or understand it?

Controlled versions of the same case test whether strong performance survives the removal of identifying signals. A positive control first shows that the probe can detect the signal when it is present, turning later silence into evidence rather than assumption.

The noise floor of judgment: what 564 runs and a 69-task eval taught us about our own scoring math

A 69-task evaluation returned 68% exact zeros and a tail to 0.7, two populations where one should have been. The model was blamed first; the scoring architecture had amplified ordinary judgement uncertainty into an artificial cliff.

The rubric locks first: grading criteria you write after reading the answer are not criteria

A standard that moves after seeing model output is a mirror, not an evaluation. Across the 69-task reference case, the auditable record places every lock before candidate evaluation; correctness requires separate evidence.

The theatrics of precision: the model reads what it won't repeat

180 runs, three synthetic deals, the same forward margin rendered as words, an integer, and two decimals. The verdict never moved. What the model was willing to repeat moved by a factor of twenty-eight — and that gap is the tell.

Your reward function is lying to you: probes for graders that fail silently

A grader can return plausible scores while rewarding shortcuts. We challenge each task for evidence dependence, numerical leakage and structural leakage before trusting its reward; failed tasks are rewritten or retired.

The anchor travels: one deal's chatter prices the next deal

Plant one sentence of unverified market chatter in the first deal an AI analyst reads, and its price for an unrelated second deal in the same session moves half a turn of EBITDA. Telling the model to evaluate the second deal independently removes less than a third of the effect. Nothing in the output says it happened.

Prompt Risk Is Investment Risk

We quadrupled a company's biggest risk and the AI's price didn't move. We added one line of market chatter and it moved $45 million. What that means for any committee putting AI near a capital decision.

You can scaffold away the flip. You can't scaffold away the frame.

564 isolated runs across one synthetic deal and its controlled variants. The recommendation never moved. The price, the leverage, and what the memo noticed moved with the packaging of the question.

How a prompt's framing quietly answers its own question

Why the grammar of an eval stem can hand a model the shape of its answer, the five trigger families that do the leaking, and the one-word test that catches all of them.

Real data in, reliable environments out

Real financial decisions in, reliable environments and evaluations out. Documented outcomes make model judgement testable against reality rather than evaluator opinion.