Our framing study established that one sentence of unverified comparables moves a frontier model’s recommended price for the deal in front of it, by about two and a quarter turns of EBITDA between the high and low versions of the sentence, while verifying the comparable never makes the diligence list. That result lives inside a single analysis. The natural next question is operational: in a production workflow, an analyst, human or machine, does not read one deal and stop. It reads a pipeline. Does the chatter from deal one price deal two?
—What chatter is
Chatter is talk about price that comes with a deal and that nobody can check. It has no verifiable source, but it still puts a number in the reader’s head before the analysis starts. In this study the chatter was one sentence planted in deal one’s question: market chatter asserting that comparable platforms recently traded at 13 to 15 times EBITDA. That sentence gives a multiple and a peer group. It gives nothing an analyst could use to confirm either.
Chatter takes a few familiar forms. A banker’s “reportedly”: a rival sale is said to have closed in the low teens, but no deal is named. A loose comp set: a handful of “similar” companies whose sale terms nobody in the room has seen. A rumoured multiple: a number that keeps coming up because people repeat it, not because anyone paid it. The source, the peer group and the number change from one form to the next. What stays the same is that none of it is verified, none of it goes into the memo’s inputs, and each one carries a price into the analysis before it starts.
A verified comparable is the opposite. It names the transaction and the date, takes its multiple from disclosed terms, and sits on the diligence list where someone can challenge it. Chatter has none of that. A careful memo can leave it out of every field and still carry its price.
A pipeline is a route through two deals, and what one step leaves in the context for the next is a property of that route, invisible in the output. The study measures one number, the second deal’s recommended midpoint, and says so: it can show whether that price moved with the first deal and by how much, and it cannot show why. The habits of evaluation work point the wrong way here. A benchmark looks for a wrong answer in the text, so the instinct is to look for the leak in the second memo; the place to look is the price. An instruction to ignore the first deal reads like a control. Here it is a treatment arm, and it is measured like one. Nor would a single-deal benchmark have caught this already, since one that scores each case alone cannot see it by construction. If the reflex is to file the whole thing under recency bias, the study measures what a pipeline does to a price, not why the model does it, and the remedy does not depend on the mechanism. If the first objection is that a fresh context per deal would make the question moot, the article agrees and says so at the end, untested; what it measures is the cheaper fix a team reaches for first, an instruction, and what that fix leaves behind. What transfers beyond credit is the control itself: any agent that processes a sequence in one context needs sequence controls in its evaluation, or it is being scored on cases it never sees in that order.
Before you pick: deal two is the same company with the same neutral question in the two conversations compared here. The only difference between them is whether deal one carried the chatter. Predict whether the chatter from deal one shifts deal two’s entry range, and how. The reveal shows the memo the model wrote most often after a clean first deal next to the one it wrote most often after the chatter. The verdict says what moved.
One sentence of chatter in deal one, then deal two. What does it do to deal two’s entry range?
After a clean first deal
- Recommendation
- Proceed
- Entry range
- 8.0–8.5x EBITDA
- Max leverage
- 4.0x
- Confidence
- Medium-High
After the chatter
- Recommendation
- Proceed with Conditions
- Entry range
- 8.0–9.5x EBITDA
- Max leverage
- 4.0x
- Confidence
- Medium-High
It does. Sixty two-turn conversations, one model, one planted sentence, and the answer is half a turn of EBITDA on a company in a different sector, carried silently.
—The setup
Each conversation has two turns. Turn one asks the model to evaluate a software-and-payments buyout case under a strict structured-output protocol; in the anchored condition, the question stem carries the framing battery’s high-comps sentence, market chatter asserting that comparable platforms recently traded at 13 to 15 times EBITDA. The control condition carries a neutral stem. Turn two, same conversation, asks the model to evaluate a second, separate opportunity: an industrial-sector case, different economics, different risks, always introduced with the neutral stem. A third condition plants the anchor in turn one and then adds an explicit instruction to turn two: evaluate this company independently on its own facts; disregard all figures, comparables, and analysis from the previous deal.
Twenty conversations per condition, order globally randomized, model and protocol identical to the published framing study. Seven conversations initially failed on output-length truncation, unevenly across arms (one anchored, four control, two isolation), and were recovered at a higher token ceiling under a pre-registered recovery procedure; all sixty are valid, and the recovered control values are consistent with what the truncated outputs already showed. One pre-registration discrepancy to disclose: the spec’s hypothesis text described a 12.0x anchor, while the implementation reused the battery’s existing 13-15x stem. Results should not be extrapolated to weaker anchors.
—What moved
The second deal’s recommended midpoint, with no anchor anywhere near it, came back higher when the first deal had carried the chatter: 8.73x mean against 8.19x in control, a gap of +0.53 turns of EBITDA (median gap 0.50; Monte Carlo permutation p < 0.0001). The pre-registered success bar was half a turn. The point estimate clears it; the bootstrap 95% confidence interval, 0.44 to 0.62, does not exclude values just below it, so we state the result as: a propagation effect of roughly half a turn, established beyond chance, whose exact magnitude at 95% confidence includes values below the threshold: the point estimate meets the pre-registered criterion; the confidence interval’s lower bound does not.
The isolation instruction helped and did not fix it. With the model explicitly told to disregard the previous deal, the second deal still priced +0.37 turns above control (p < 0.0001). The instruction removed roughly thirty percent of the propagation and left the rest standing, which echoes what the anchoring literature finds about prompt-level mitigation generally: anchors act below the level that instructions reach.
The channel is specific: recommended leverage on the second deal sat at 4.0x in every valid conversation in every condition. The contamination expresses itself in the price midpoint alone, the same dial the single-deal anchor moves hardest, and the dial least likely to be compared across two analyses of two different companies.
Two things make this operationally uncomfortable rather than merely interesting.
First, the propagation is silent. Zero of sixty second-deal memos reference the prior deal, its sector, or the planted figure, by keyword scan and by hand-reading samples. The detector checks explicit references; we cannot rule out subtler linguistic traces, but there is nothing a reviewer skimming the memo would catch. The first deal’s chatter is simply in the price.
Second, the anchor moved more than the number. In the anchored condition, seventeen of nineteen original second-deal memos recommended proceed-with-conditions; in control, fourteen of sixteen recommended outright proceed. The prior deal’s framing shifted the price of the next deal, and also the conservatism of its verdict, a categorical channel that our single-deal study found immovable under every within-deal manipulation we tried. This was not a pre-registered metric; we report it as a secondary observation that wants its own controlled run. Until then the price is the finding and the verdict is a lead.
—What it is not
This is not a claim about any vendor’s memory feature; the two deals share nothing but a context window. It is not a claim that the model recalls the anchor; we measured outputs, not internals, and the result does not establish a general contamination rate. And it is not the same phenomenon as a model anchoring on its own earlier draft when asked to review it, which recent work has documented in single-document settings1; here the contamination crosses to a different company in a different sector, the way work actually flows through an analyst’s day.
The mechanism does have company in the literature. Anchoring effects in language models have been located in shallow layers and shown to resist instruction-level debiasing, including ignore-the-anchor prompting2, which is consistent with our isolation instruction recovering only part of the gap. Instruction-hierarchy work finds that telling a model which content to privilege fails to establish reliable precedence between competing parts of the context3. And a recent survey of 164 papers on language models in finance catalogues structural evaluation biases while listing no framing or anchoring behaviors at all4, which is roughly where the field’s blind spot sits.
—What to do with it
For anyone running sequential analyses through one session, the mitigation is structural, not linguistic: fresh context per deal. We have not tested the fresh-context remedy directly, but the data establish that the obvious linguistic alternative, an isolation instruction, quietly underdelivers by two-thirds. For anyone building evaluation for analyst fleets, single-deal benchmarks miss this class of failure entirely; a model can score identically on every isolated case while its pipeline behavior drifts with document order. Cross-deal contamination is measurable before deployment, with planted anchors and sequence controls, and we now measure it as part of our standard battery. Until an evaluation runs deals in sequence, it has not tested the pipeline the fleet actually runs.
—Method, briefly
One frontier model family, with the specific model and vendor named in the technical paper; two synthetic cases authored for the published framing study; one anchor value; twenty conversations per condition with global order randomization; a fixed decoding configuration matching the published battery. Seven conversations were recovered under a pre-registered ceiling procedure; case texts and stems hash-frozen before the runs. Every number above traces to the run artifacts; the analysis, including the recovery procedure, is recorded with the data. A controlled demonstration, not a market study; sector pairs beyond software-to-industrial, weaker anchors, and longer pipelines are the obvious extensions.
Companion to You can scaffold away the flip. You can’t scaffold away the frame.5, which documents the within-deal effects. The cross-deal battery is part of the evaluation suite we run for funds deploying analyst fleets and the allocators underwriting them.
References
- T.-E. Song. Cross-context review: improving LLM output quality by separating production and review sessions. arXiv:2603.12123, 2026. Back
- Y. Huang, B. Bie, Z. Na, W. Ruan, S. Lei, Y. Yue, X. He. Understanding the anchoring effect of LLM with synthetic data: existence, mechanism, and potential mitigations. arXiv:2505.15392, ICLR HCAIR Workshop 2026. Back
- Y. Geng, H. Li, H. Mu, X. Han, T. Baldwin, O. Abend, E. Hovy, L. Frermann. Control illusion: the failure of instruction hierarchies in large language models. arXiv:2502.15851, 2025. Back
- Y. Kong, H. Lee, Y. Hwang, A. Lopez-Lira, B. Levy, D. Mehta, Q. Wen, C. Choi, Y. Lee, S. Zohren. Evaluating LLMs in finance requires explicit bias consideration. arXiv:2602.14233, 2026. Back
- Dissei Research. You can scaffold away the flip. You can’t scaffold away the frame. 2026. Back
Cross-deal contamination is measurable before deployment. We run this battery for funds deploying analyst fleets and the allocators underwriting them.
Talk to us