A common pattern in AI-assisted diligence: a data-room extractor pulls figures out of unaudited management materials at whatever precision the source file happens to carry, and pastes them into the memo the analyst then reviews with an LLM. “Roughly mid-teens forward margin” in the CIM becomes “14.73% forward margin” in the extract. The number the model sees looks… verified?
This is a study of one step on the model’s route: what it does with a figure between reading it and pricing the deal. Two trip-ups catch readers who know evaluation and not credit. The first is to read a headline number that holds still as a null result and stop; here that is where the reading starts, not where it ends. The second is to expect the effect in the verdict; in credit the verdict is the least sensitive output, and the entry multiple and the rationale are where the reading happens. What transfers beyond credit is the audit itself: an agent’s written reasoning is an output, and the way to learn what an input did is a controlled variant, not a reading of the explanation.
We tested whether that rendering (words vs. integer vs. two decimals) moves the valuation the model produces from an otherwise identical memo. Three deal archetypes, two frontier models, three renderings of a single load-bearing forward projection: the honest hedge, the cleaned-up integer, and the two-decimal version. Every other figure in the memo held its rendering constant. Six seeds per cell, two models, three deal archetypes, five cells (the three renderings, the compounding cell, and a length control): 180 valuation runs, all on one task, one output contract, and one underwritten thesis. The channel this study watches is the quote-back rate, read off the rationale rather than the bid.
—The setup
Three synthetic deals across three archetypes: a mature software buyout, a consumer DTC buyout, and an industrial buyout, each carrying a canonical forward-margin projection to two decimals in the underlying fixture. From that canonical value we derived three renderings of the same figure: P0, hedged words (“low twenties,” the honest guess); P1, the integer (“21%,” cleaned up but non-committal on the precision); and P2, two decimals (“21.43%,” the high-precision rendering). Only the headline projection’s rendering varied. Every other figure in the memo held its P0 wording constant. A fourth cell, P2C, rendered every load-bearing figure at two decimals, to test whether the effect compounds.
The model was asked to produce an entry EV/EBITDA multiple, a leverage stack, and a verdict from the memo. The ask price was removed from the transaction section, so the model had to compute the multiple from the fundamentals rather than echoing back a stated bid.
The model treats a two-decimal figure and an integer as different things to repeat: the integer usually comes back verbatim, and the two decimals come back rounded, paraphrased, or not at all in all but a few runs. Nothing a human reader watches shifts to match.
The bid, the stated confidence, and the verdict all hold. Precision is being processed. It is not being surfaced. Whatever effect the increased precision has, the rationale will not tell you.
—What moved
Start with the change that is easy to see. At P1 the rationale repeats the integer verbatim in most runs; at P2 the two-decimal figure survives as written in a small minority of runs, and in the rest it is rounded to an integer, paraphrased, or left out. From the same sentence in the same model, the digits it receives are not the digits it repeats. The precision goes into the input and, in all but a few runs, does not come back out unchanged in the rationale.
That is the change in what the model writes. In the channels a human reader is looking at (the bid, the stated confidence, the verdict), nothing shifts to match. Across the panel, the two-decimal rendering shifts the entry multiple by roughly half a turn of EBITDA. The 95% confidence interval includes zero on every contrast; no permutation test approaches significance. The compounding cell, which renders every load-bearing figure at two decimals, has a paired contrast that also spans zero. P2C applies the precision most broadly, and it shows the same half-turn drift as P2: no effect detected.
The model’s confidence in the figure does not respond either. The distribution of epistemic framing (the balance of verified, estimated, uncertain, and neutral language in the rationale about the projection) is essentially identical at P0, P1, and P2. Every cell’s rationales are mostly frame-neutral; the small residuals do not shift. The model is not narrating the two-decimal number as more certain than the integer. It is narrating it as the same claim.
Even the archetype where the model has the most room to move (a software buyout without a tight public-comp anchor) shows a paired contrast that spans zero, like the other two.
| Archetype | Multiple at P0 | Multiple at P2 | Change, P2 − P0 | 95% CI of change | Permutation p |
|---|---|---|---|---|---|
| Software buyout | 25.25× | 23.83× | −1.42 | [−2.83, 0.00] | 0.26 |
| Consumer buyout | 9.58× | 9.83× | +0.25 | [−0.21, 0.71] | 0.52 |
| Industrial buyout | 9.54× | 9.46× | −0.08 | [−0.21, 0.00] | 0.85 |
—What it is not
Length control. We reran the P0 baseline with a filler sentence inserted next to the headline projection, matched in character length to the P0→P2 delta. The filler cell moved the entry multiple by +0.03 turns (CI [−0.92, 0.96], p = 0.99). A null result: the length control detected no length effect of its own.
Units control. All three deals carry positive LTM EBITDA and were priced as EV/EBITDA transactions throughout; no rung switched to an ARR-multiple lens or a forward-EBITDA lens. The multiple reported at P0 and P2 is the same quantity in the same units.
Verdict check. Every one of the 180 runs, across every rung, deal, model, and seed, returned the same verdict: PROCEED_WITH_CONDITIONS. If the diligence stack you have reads only the categorical output (the go / no-go / conditional flag), you would see no effect at all.
—What to do with it
Do not audit the model’s handling of precision by reading the rationale. The two-decimal figure the model was given all but vanished from the rationale it wrote. Whatever the extra precision did inside the answer, the rationale in all but a few runs rounds it, paraphrases it, or leaves it out before the reader sees it. If your review process reads the model’s stated reasoning to check that it engaged with the numbers correctly, you will systematically fail to notice this class of effect. It is not narrated.
Do not read a flat bid as proof the precision was inert. The bid, the verdict, and the model’s stated confidence all hold across rungs. That does not show the two-decimal figure went unread; the rationale-level asymmetry shows the model told the two renderings apart. What this study can say is narrower: no effect on the bid, the confidence, or the verdict was detectable, and the one channel that did respond is the one a reviewer does not audit. A flat bid shows that the price held, not that the precision went unread.
Someone still has to own the rendering. In most AI-diligence workflows nobody owns the choice of how figures are rendered on their way from source to model. The extractor’s default becomes the memo’s rendering, becomes the model’s input. The rationale-level asymmetry shows the model tells the renderings apart, whether or not any given deal shows a visible price shift. Precision is a lever. Someone should be holding it deliberately.
—Where we come in
Precision theatre is one variant in a family of framing effects we audit as part of building evaluation infrastructure for finance domains where being confidently wrong is expensive: credit underwriting, structured credit, private-equity evaluation, distressed and special situations. Our cases exist in controlled variants that vary the precision, the framing, and the anchoring of the figures while holding the financial substance constant. The contrast across variants is the measurement.
The class of failure that precision theatre exemplifies is the one we spend most of our time on: the framing effect that does not look like a framing effect. Two decimals reads as discipline, so nothing in a diligence review flags it as an input worth checking. That is precisely the shape of a problem worth measuring: one the reviewer cannot see just by reading the output.
—Method, briefly
Two frontier models from two different providers on the panel. Six seeds per cell, independent sampling. Length-control cell run across all three archetypes, matched in character length to the P0→P2 delta. Adversarial review of harness, prompts, and parser before the sweep ran. Paired-permutation tests (10,000 shuffles) and bootstrap 95% confidence intervals (10,000 resamples) on every contrast. Three synthetic deals and one varied figure bound what this can show: nothing here speaks to real deal files, to figures other than a margin projection, or to tasks outside valuation.
Companion to You can scaffold away the flip. You can’t scaffold away the frame., which documents the framing battery this study extends.
Precision theatre is measurable before it prices a deal. We run this battery as part of the evaluation suite for teams putting models near capital decisions.
Talk to us