Only if you ask: planting wrong numbers in a 20-step credit worksheet with three models

Accurate calculation from scratch does not guarantee reliable handling of an inherited error, even when the model can correct that error in an audit.

We wrote wrong numbers into a 20-step credit worksheet and gave it to three models, sometimes asking them to audit it and sometimes saying nothing — 842 worksheets and audits in all. Asked, the two models that can complete the sheet found sixty of sixty and fifty-nine of sixty planted errors, none of them less than 19% off; told nothing, no model reliably said a number was wrong. So we read the open-weight model's reasoning, which it returns in the open. Handed an interest figure ten times too large, it worked out the right one, wrote that the sheet's figure "seems off by factor 10", and used the sheet's figure anyway; it carried thirty-three of sixty planted errors forward. The other capable model mostly replaced the wrong number, and said nothing.

Where we started

We started with a simple idea: an agent's performance on a long-horizon task ought to decline exponentially with the number of steps in the task. If each step succeeds independently with the same probability, those probabilities multiply across the task. We tested this in a 20-step credit worksheet, then planted errors to see when models would correct them.

A sanity check has two layers. Does the number break a rule of thumb — EBITDA above revenue, leverage past ten times? And does it reconcile when recomputed? An error that breaks no rule of thumb can only be found the second way. In credit that matters: one wrong interest figure flows into cover, tax, cash flow, leverage, liquidity and recovery, and later calculations can remain internally consistent while carrying it forward.

Hand an analyst a colleague's half-finished worksheet and there are two jobs: complete it, and check the numbers it relies on. We looked at where errors entered the calculation and whether models corrected them before carrying them forward.

The experiment

The task is a 20-step credit worksheet, from EBITDA to the recovery waterfall, with a rule-based answer key and a tolerance of about one per cent. We generated 400 synthetic issuers; the models saw 221.

The full inputs and twenty rules for one issuer are laid out at the end of this note, and every step number in this note refers to that table.

Three models, all through Amazon Bedrock, none with a calculator or any tool: Claude Haiku 4.5, Claude Sonnet 5 and OpenAI's open-weight gpt-oss-120b. Each worksheet completion was a single response containing the remaining rows.

Four conditions:

  • From scratch. All twenty steps from the inputs; forty worksheets per model, each model's forty issuers drawn separately.
  • Handed the sheet. The sheet filled in up to and including a wrong row, with no instruction beyond continuing it.
  • Asked to audit. A separate call, shown all twenty rows with the planted error carried through the subsequent calculations — "Audit this completed worksheet against the inputs and the rules" — returning the wrong steps, corrected values and a verdict.
  • Told where. "Step i in this worksheet is wrong. Recompute step i correctly, then complete steps i+1 to 20."

The continuation and audit conditions used the same sixty planted errors for every model. A planted error is a wrong value in one row, of three classes. Impossible: it breaks a hard bound. Implausible: it breaks a going-concern range. Silent: it breaks neither; only recomputing finds it. Twenty items per class, with every planted error at least 19% off.

Four design guards:

  • No hints. The base prompt never mentions errors, checking or plausibility; the notes column is optional, and an automated scan of the prompt variants enforces that silence.
  • Answer key. Every step has a rule-computed truth; correct means within tolerance of it, never a judgement.
  • Clean controls. Correct sheets in both conditions, so reports of errors can be read against how often the model flags a sheet with nothing wrong.
  • Scored by arithmetic. The comparison uses the numbers the models returned: whether the continuation followed the correct values, or the audit identified and correctly recomputed the planted row. Notes and returned reasoning illustrate what happened.

The sheet from scratch

Where does the first error appear?

Forty worksheets per model, each with twenty steps. Inspect the share still free of errors at any step.

Worksheets with no error yet · %

Error-free worksheets through each of 20 stepsUse the step control below to inspect exact counts. Error bars show the supplied intervals.0255075100Claude Haiku 4.5, step 1: 40 of 40Claude Haiku 4.5, step 2: 40 of 40Claude Haiku 4.5, step 3: 40 of 40Claude Haiku 4.5, step 4: 40 of 40Claude Haiku 4.5, step 5: 40 of 40Claude Haiku 4.5, step 6: 1 of 40Claude Haiku 4.5, step 7: 1 of 40Claude Haiku 4.5, step 8: 1 of 40Claude Haiku 4.5, step 9: 0 of 40Claude Haiku 4.5, step 10: 0 of 40Claude Haiku 4.5, step 11: 0 of 40Claude Haiku 4.5, step 12: 0 of 40Claude Haiku 4.5, step 13: 0 of 40Claude Haiku 4.5, step 14: 0 of 40Claude Haiku 4.5, step 15: 0 of 40Claude Haiku 4.5, step 16: 0 of 40Claude Haiku 4.5, step 17: 0 of 40Claude Haiku 4.5, step 18: 0 of 40Claude Haiku 4.5, step 19: 0 of 40Claude Haiku 4.5, step 20: 0 of 40Claude Sonnet 5, step 1: 40 of 40Claude Sonnet 5, step 2: 40 of 40Claude Sonnet 5, step 3: 40 of 40Claude Sonnet 5, step 4: 40 of 40Claude Sonnet 5, step 5: 40 of 40Claude Sonnet 5, step 6: 38 of 40Claude Sonnet 5, step 7: 38 of 40Claude Sonnet 5, step 8: 38 of 40Claude Sonnet 5, step 9: 38 of 40Claude Sonnet 5, step 10: 38 of 40Claude Sonnet 5, step 11: 38 of 40Claude Sonnet 5, step 12: 38 of 40Claude Sonnet 5, step 13: 38 of 40Claude Sonnet 5, step 14: 38 of 40Claude Sonnet 5, step 15: 38 of 40Claude Sonnet 5, step 16: 38 of 40Claude Sonnet 5, step 17: 38 of 40Claude Sonnet 5, step 18: 38 of 40Claude Sonnet 5, step 19: 38 of 40Claude Sonnet 5, step 20: 37 of 40gpt-oss-120b, step 1: 40 of 40gpt-oss-120b, step 2: 40 of 40gpt-oss-120b, step 3: 40 of 40gpt-oss-120b, step 4: 40 of 40gpt-oss-120b, step 5: 39 of 40gpt-oss-120b, step 6: 39 of 40gpt-oss-120b, step 7: 39 of 40gpt-oss-120b, step 8: 39 of 40gpt-oss-120b, step 9: 39 of 40gpt-oss-120b, step 10: 39 of 40gpt-oss-120b, step 11: 39 of 40gpt-oss-120b, step 12: 39 of 40gpt-oss-120b, step 13: 39 of 40gpt-oss-120b, step 14: 39 of 40gpt-oss-120b, step 15: 39 of 40gpt-oss-120b, step 16: 39 of 40gpt-oss-120b, step 17: 39 of 40gpt-oss-120b, step 18: 39 of 40gpt-oss-120b, step 19: 39 of 40gpt-oss-120b, step 20: 38 of 4016101520
Claude Haiku 4.51 / 40
Claude Sonnet 538 / 40
gpt-oss-120b39 / 40

All forty Haiku worksheets reached step 5 without an error. Thirty-nine first failed at step 6, the interest calculation. The last first failed at step 9.

View the author’s original figure
Figure 1. Share of worksheets with no error through each step, three models

Figure 1. Share of forty worksheets per model with no error through each of the twenty steps; error bars are 95% intervals. Issuers were drawn separately for each model. Haiku's sharp fall is at step 6, the interest calculation. All numbers are the experiment's measured results.

Sonnet 5 finished thirty-seven of forty worksheets correctly, gpt-oss-120b thirty-eight, Haiku 4.5 none.

All forty Haiku worksheets passed the first five rows. Thirty-nine first went wrong at step 6, cash interest: four debt tranches times four rates. The remaining worksheet first failed at step 9. Haiku's first failure was concentrated at a particular calculation, rather than spread evenly through the sheet.

The interest error then fed into interest cover and covenant headroom. Those next two rows applied their rules correctly to Haiku's earlier numbers, including the wrong interest figures. The later arithmetic reconciled with the earlier arithmetic, while preserving its mistake. Internal consistency alone would miss it.

The simple exponential picture assumes the same chance of success at each step. Haiku's results do not follow that picture: five reliable steps are followed by one that almost always fails. We used one fixed sequence of calculations, so this does not establish a general relationship between task length and performance. It does show how strongly a difficult calculation can shape the outcome. Errors can still compound after that calculation, and Haiku also made fresh errors later in the sheet.

The sheet, handed over

Sonnet and gpt-oss could both calculate the sheet accurately from scratch. Handed a wrong number, they behaved quite differently. Sonnet usually produced subsequent values from the correct calculation; gpt-oss often carried the supplied error forward. Asked to audit, both identified and corrected almost every planted row.

What changes when the task is an audit?

The same 60 planted items, given to each model in two conditions.

Model receivesPartial worksheet
InstructionComplete the remaining rows
Scored · Affected downstream rows

At least two affected downstream values follow the correct chain.

Claude Haiku 4.5Continue: 7 · Audit: 33
7 of 60
Claude Sonnet 5Continue: 50 · Audit: 60
50 of 60
gpt-oss-120bContinue: 24 · Audit: 59
24 of 60
ContinueAudit
How this is scored

At least two of the first four affected downstream values must match the correct chain, with more matches to it than to the chain carrying the error forward.

The whole remaining sheet need not be correct. Three items have too few affected rows to qualify and remain in the denominator of 60.

For continuations, we score a short stretch of affected rows after the planted error; the whole remaining sheet need not be right. The audit must name the planted row and give its correct value. The same sixty planted items were used in both conditions for all three models; the full scoring rules are in the Notes.

The interest example shows the distinction. We supplied gpt-oss with cash interest of €140.2m, ten times the correct amount. Its returned reasoning calculated €14.02m and said the supplied value "seems off by factor 10". It then wrote: "So we must accept that as correct per worksheet." The answer used €140.2m, put year-end cash at minus €72.6m and left the notes blank.

The right calculation. The wrong figure in the answer.

One documented gpt-oss-120b continuation. Follow the reported response.

gpt-oss-120b
Supplied worksheetCash interest€140.2mPlanted error · 10× too large
Returned reasoning
Recomputed interest€14.02mseems off by factor 10
Decision in the same response
“So we must accept that as correct per worksheet.”
Submitted answerRecorded output
Cash interest
€140.2m
Year-end cash
−€72.6m
Notes
Blank
4 / 4
The wrong figure reaches the answer

The answer uses €140.2m of interest, reports year-end cash of −€72.6m, and leaves the notes blank.

Four stages · replay stops at the answer

The correct calculation appeared in the same response as the decision to use the wrong figure. Across all sixty items, gpt-oss carried thirty-three planted errors forward. Sonnet more often used the correct values, but its notes referred to the planted number on just twelve of sixty items. A blank notes column told us little about what had happened to the arithmetic.

Haiku remained much less successful even in an audit. Audits also flagged some clean sheets as erroneous, although none of their proposed numerical changes fell outside the scoring tolerance. Told where the error is, Sonnet and gpt-oss fixed the named step sixty of sixty and finished the whole tail fifty-eight of sixty each; Haiku fixed the named step forty of sixty and finished none of sixty, a floor set by step 18, which it gets right two of sixty.

The auditor saw the completed sheet and was asked for corrections; the continuing model saw a partial sheet and was asked for the remaining rows. These results do not measure the gain in complete-workflow accuracy from adding a checking step. What they do show is that accurate calculation from scratch does not guarantee reliable handling of an inherited error, even when the model can correct that error in an audit.

What to do with it

First, performance on long-horizon tasks need not decay cleanly and exponentially with task length. It can be dominated by the difficulty of the hardest steps. Here, Haiku's first failure was overwhelmingly concentrated at one calculation. For financial workflows, that suggests paying attention to which steps are difficult, as well as how many steps there are. Many correct steps do not compensate for a wrong calculation that feeds the rest of the sheet.

Second, forcing a model into best practices, including checking its work, may have an effect on performance comparable with the overall skill of the model — or larger. That is a hypothesis, and these results give us a reason to investigate it: Sonnet and gpt-oss were similarly accurate from scratch and almost equally successful as auditors, yet sharply different when handed an error without a checking instruction.

Harnesses and skills — the engineering around the model — may therefore have an outsized importance in long-horizon tasks. The skill of financial judgement, on the other hand, must be trained.

Notes

  • Construction. Four hundred synthetic issuers, with revenue of €150–1,200m, four debt tranches and covenants set above opening leverage. Answers are scored against full-precision calculations, within 1%, with floors of 0.1 €m for amounts and 0.02 for ratios. The worked worksheet displays amounts to one decimal and ratios to two.
  • Sample. Forty from-scratch worksheets per model, with issuers drawn separately. Continuation and audit use the same sixty planted items for every model, twenty per error class. The repair condition uses sixty separately drawn items. Clean controls comprise fifty-one calls per condition for Haiku and twenty for each other model. The 842 total counts scored worksheets and audits, not unique issuers or API attempts.
  • Planted errors. The sixty continuation/audit items were at least 19% off, with a median error of 90%; thirty-nine were a factor of ten out or the wrong sign. The error class depends on which bound the resulting worksheet first violates. These audit rates describe the planted errors tested here.
  • Continuation scoring. We inspect the first four downstream steps whose correct values differ from those obtained by propagating the planted error. Credit requires at least two values within tolerance of the correct chain, and more matches to that chain than to the propagated one. Three planted items have fewer than two such steps and cannot qualify; they remain in the denominator of sixty. An audit scores a correction only if it lists the planted row and supplies a value within tolerance.
  • Clean audits. Haiku returned an "errors found" verdict on 16 of 51 clean audits, Sonnet on 9 of 20 and gpt-oss on 7 of 20. No proposed numerical correction exceeded tolerance. Counting any non-empty error list as an alert raises Haiku's count to 39 of 51; some entries confirm a row's value while filing it as an error. The notes observations in the article come from direct inspection.
  • Settings and scope. Haiku and gpt-oss ran at temperature zero; Sonnet used default sampling with thinking disabled. gpt-oss returned visible reasoning. Most Sonnet and gpt-oss audits were retried at an 8,000-token limit after reaching the initial 2,000-token cap; continuation began with 4,000. Final outputs had no truncation or parse failures. These are comparisons of the configurations and tasks used here. We did not test a model auditing its own completed answer or checking during completion. The from-scratch experiment uses one fixed sequence; Sonnet's three first failures and gpt-oss's two provide limited evidence about the shape of a decay curve. Figure 1 uses Wilson 95% intervals.

The worksheet

The prompt has three blocks — the issuer's inputs, the twenty rules in words and in symbols, and the worksheet itself, twenty rows of Step k — name: value, with ? where the model is to write. These are the inputs for one of the 400 issuers, exactly as the model saw them, followed by the twenty steps with that issuer's correct values.

Explore the worksheet

One interest figure. A different financial picture.

Inspect the issuer’s inputs, then change cash interest in step 6. The remaining rows follow the article’s rules. Your changes are a calculation exercise, separate from the model results.

The original issuer, exactly as supplied to the models. Amounts are in euro millions. A fraction of 0.216 means 21.6%.

Original issuer inputs
InputSymbolValueUnit
Operating assumptions
RevenueR213.50€m
EBITDA marginm0.216fraction of revenue
Depreciation and amortisationDA8.26€m
Capital expenditureCAPEX5.49€m
Change in working capitalDWC5.40€m
Tax ratet0.250fraction of taxable profit
Opening cashC024.01€m
Debt balances
Term loan outstanding at t0TL96.20€m
Bond outstanding at t0B77.90€m
Mezzanine outstanding at t0M54.80€m
Revolving credit facility drawn at t0RCFd10.40€m
Revolving credit facility sizeF23.89€m
Financing terms
Term loan raterTL7.5% per year
Bond raterB6.5% per year
Mezzanine raterM13.7% per year
Revolving credit facility raterRCF5.4% per year
Mandatory amortisation ratea0.010fraction of TL per year
Cash sweep shares0.250fraction of excess cash flow
Covenants and valuation
Maximum permitted net leverageLmax5.00× EBITDA
Minimum permitted interest coverICmin2.00× cash interest
Minimum liquidity requirementLmin4.26€m
EBITDA growth over the yearg-0.034fraction
Distressed enterprise value multiplek4.00× EBITDA
Source: this article’s public worksheet and twenty calculation rules. This issuer is separate from the €140.2m example above. Reset restores the original inputs and answers.
View the original twenty-step answer key
Twenty-step worksheet
StepLineRuleThis issuer
1EBITDA0, opening EBITDAEBITDA0 = R·m46.1
2G, gross debtG = TL + B + M + RCFd239.3
3ND0, opening net debtND0 = G − C0215.3
4Lev0, opening net leverageLev0 = ND0 / EBITDA04.67
5HL0, leverage headroomPASS if Lev0 ≤ Lmax; HL0 = Lmax − Lev00.33, PASS
6I, cash interestI = TL·rTL + B·rB + M·rM + RCFd·rRCF20.3
7IC0, interest coverIC0 = EBITDA0 / I2.27
8HIC0, cover headroomPASS if IC0 ≥ ICmin; HIC0 = IC0 − ICmin0.27, PASS
9TAXTP = EBITDA0 − DA − I; TAX = t·max(TP, 0)4.4
10FCF, free cash flowFCF = EBITDA0 − CAPEX − DWC − I − TAX10.5
11AMORT, mandatory amortisationAMORT = a·TL1.0
12SWEEP, cash sweepECF = FCF − AMORT; SWEEP = min(s·max(ECF, 0), TL − AMORT)2.4
13TL1, year-end term loanTL1 = TL − AMORT − SWEEP92.9
14C1, year-end cashC1 = C0 + FCF − AMORT − SWEEP31.2
15ND1, year-end net debtG1 = TL1 + B + M + RCFd; ND1 = G1 − C1204.8
16EBITDA1, year-end EBITDAEBITDA1 = EBITDA0·(1 + g)44.5
17Lev1, year-end leverageLev1 = ND1 / EBITDA1; PASS if Lev1 ≤ Lmax4.60, PASS
18LIQ, liquidity headroomLIQ = C1 + (F − RCFd) − Lmin; PASS if LIQ ≥ 040.4, PASS
19EV, distressed enterprise valueEV = k·EBITDA1178.2
20Recovery waterfallEV to the revolver first, then the term loan and bond pro rata, then the mezzanine, then equityrecRCF 10.4; recTL 91.2; recB 76.5; recM 0.0; EQ 0.0; bond recovery 98.3%

Amounts are € millions to one decimal and ratios to two, as the rules require; the answer key is computed at full precision.

© Dissei. All rights reserved.

© Dissei. All rights reserved. No reproduction, adaptation, or derivative use of this content or methodology without prior written permission.