We wrote wrong numbers into a 20-step credit worksheet and gave it to three models, sometimes asking them to audit it and sometimes saying nothing — 842 worksheets and audits in all. Asked, the two models that can complete the sheet found sixty of sixty and fifty-nine of sixty planted errors, none of them less than 19% off; told nothing, no model reliably said a number was wrong. So we read the open-weight model's reasoning, which it returns in the open. Handed an interest figure ten times too large, it worked out the right one, wrote that the sheet's figure "seems off by factor 10", and used the sheet's figure anyway; it carried thirty-three of sixty planted errors forward. The other capable model mostly replaced the wrong number, and said nothing.
Where we started
We started with a simple idea: an agent's performance on a long-horizon task ought to decline exponentially with the number of steps in the task. If each step succeeds independently with the same probability, those probabilities multiply across the task. We tested this in a 20-step credit worksheet, then planted errors to see when models would correct them.
A sanity check has two layers. Does the number break a rule of thumb — EBITDA above revenue, leverage past ten times? And does it reconcile when recomputed? An error that breaks no rule of thumb can only be found the second way. In credit that matters: one wrong interest figure flows into cover, tax, cash flow, leverage, liquidity and recovery, and later calculations can remain internally consistent while carrying it forward.
Hand an analyst a colleague's half-finished worksheet and there are two jobs: complete it, and check the numbers it relies on. We looked at where errors entered the calculation and whether models corrected them before carrying them forward.
The experiment
The task is a 20-step credit worksheet, from EBITDA to the recovery waterfall, with a rule-based answer key and a tolerance of about one per cent. We generated 400 synthetic issuers; the models saw 221.
The full inputs and twenty rules for one issuer are laid out at the end of this note, and every step number in this note refers to that table.
Three models, all through Amazon Bedrock, none with a calculator or any tool: Claude Haiku 4.5, Claude Sonnet 5 and OpenAI's open-weight gpt-oss-120b. Each worksheet completion was a single response containing the remaining rows.
Four conditions:
- From scratch. All twenty steps from the inputs; forty worksheets per model, each model's forty issuers drawn separately.
- Handed the sheet. The sheet filled in up to and including a wrong row, with no instruction beyond continuing it.
- Asked to audit. A separate call, shown all twenty rows with the planted error carried through the subsequent calculations — "Audit this completed worksheet against the inputs and the rules" — returning the wrong steps, corrected values and a verdict.
- Told where. "Step i in this worksheet is wrong. Recompute step i correctly, then complete steps i+1 to 20."
The continuation and audit conditions used the same sixty planted errors for every model. A planted error is a wrong value in one row, of three classes. Impossible: it breaks a hard bound. Implausible: it breaks a going-concern range. Silent: it breaks neither; only recomputing finds it. Twenty items per class, with every planted error at least 19% off.
Four design guards:
- No hints. The base prompt never mentions errors, checking or plausibility; the notes column is optional, and an automated scan of the prompt variants enforces that silence.
- Answer key. Every step has a rule-computed truth; correct means within tolerance of it, never a judgement.
- Clean controls. Correct sheets in both conditions, so reports of errors can be read against how often the model flags a sheet with nothing wrong.
- Scored by arithmetic. The comparison uses the numbers the models returned: whether the continuation followed the correct values, or the audit identified and correctly recomputed the planted row. Notes and returned reasoning illustrate what happened.
The sheet from scratch
Where does the first error appear?
Forty worksheets per model, each with twenty steps. Inspect the share still free of errors at any step.
Worksheets with no error yet · %
All forty Haiku worksheets reached step 5 without an error. Thirty-nine first failed at step 6, the interest calculation. The last first failed at step 9.
View the author’s original figure
Figure 1. Share of forty worksheets per model with no error through each of the twenty steps; error bars are 95% intervals. Issuers were drawn separately for each model. Haiku's sharp fall is at step 6, the interest calculation. All numbers are the experiment's measured results.
Sonnet 5 finished thirty-seven of forty worksheets correctly, gpt-oss-120b thirty-eight, Haiku 4.5 none.
All forty Haiku worksheets passed the first five rows. Thirty-nine first went wrong at step 6, cash interest: four debt tranches times four rates. The remaining worksheet first failed at step 9. Haiku's first failure was concentrated at a particular calculation, rather than spread evenly through the sheet.
The interest error then fed into interest cover and covenant headroom. Those next two rows applied their rules correctly to Haiku's earlier numbers, including the wrong interest figures. The later arithmetic reconciled with the earlier arithmetic, while preserving its mistake. Internal consistency alone would miss it.
The simple exponential picture assumes the same chance of success at each step. Haiku's results do not follow that picture: five reliable steps are followed by one that almost always fails. We used one fixed sequence of calculations, so this does not establish a general relationship between task length and performance. It does show how strongly a difficult calculation can shape the outcome. Errors can still compound after that calculation, and Haiku also made fresh errors later in the sheet.
The sheet, handed over
Sonnet and gpt-oss could both calculate the sheet accurately from scratch. Handed a wrong number, they behaved quite differently. Sonnet usually produced subsequent values from the correct calculation; gpt-oss often carried the supplied error forward. Asked to audit, both identified and corrected almost every planted row.
What changes when the task is an audit?
The same 60 planted items, given to each model in two conditions.
At least two affected downstream values follow the correct chain.
How this is scored
At least two of the first four affected downstream values must match the correct chain, with more matches to it than to the chain carrying the error forward.
The whole remaining sheet need not be correct. Three items have too few affected rows to qualify and remain in the denominator of 60.
For continuations, we score a short stretch of affected rows after the planted error; the whole remaining sheet need not be right. The audit must name the planted row and give its correct value. The same sixty planted items were used in both conditions for all three models; the full scoring rules are in the Notes.
The interest example shows the distinction. We supplied gpt-oss with cash interest of €140.2m, ten times the correct amount. Its returned reasoning calculated €14.02m and said the supplied value "seems off by factor 10". It then wrote: "So we must accept that as correct per worksheet." The answer used €140.2m, put year-end cash at minus €72.6m and left the notes blank.
The right calculation. The wrong figure in the answer.
One documented gpt-oss-120b continuation. Follow the reported response.
seems off by factor 10
“So we must accept that as correct per worksheet.”
- Cash interest
- €140.2m
- Year-end cash
- −€72.6m
- Notes
- Blank
The answer uses €140.2m of interest, reports year-end cash of −€72.6m, and leaves the notes blank.
The correct calculation appeared in the same response as the decision to use the wrong figure. Across all sixty items, gpt-oss carried thirty-three planted errors forward. Sonnet more often used the correct values, but its notes referred to the planted number on just twelve of sixty items. A blank notes column told us little about what had happened to the arithmetic.
Haiku remained much less successful even in an audit. Audits also flagged some clean sheets as erroneous, although none of their proposed numerical changes fell outside the scoring tolerance. Told where the error is, Sonnet and gpt-oss fixed the named step sixty of sixty and finished the whole tail fifty-eight of sixty each; Haiku fixed the named step forty of sixty and finished none of sixty, a floor set by step 18, which it gets right two of sixty.
The auditor saw the completed sheet and was asked for corrections; the continuing model saw a partial sheet and was asked for the remaining rows. These results do not measure the gain in complete-workflow accuracy from adding a checking step. What they do show is that accurate calculation from scratch does not guarantee reliable handling of an inherited error, even when the model can correct that error in an audit.
What to do with it
First, performance on long-horizon tasks need not decay cleanly and exponentially with task length. It can be dominated by the difficulty of the hardest steps. Here, Haiku's first failure was overwhelmingly concentrated at one calculation. For financial workflows, that suggests paying attention to which steps are difficult, as well as how many steps there are. Many correct steps do not compensate for a wrong calculation that feeds the rest of the sheet.
Second, forcing a model into best practices, including checking its work, may have an effect on performance comparable with the overall skill of the model — or larger. That is a hypothesis, and these results give us a reason to investigate it: Sonnet and gpt-oss were similarly accurate from scratch and almost equally successful as auditors, yet sharply different when handed an error without a checking instruction.
Harnesses and skills — the engineering around the model — may therefore have an outsized importance in long-horizon tasks. The skill of financial judgement, on the other hand, must be trained.
Notes
- Construction. Four hundred synthetic issuers, with revenue of €150–1,200m, four debt tranches and covenants set above opening leverage. Answers are scored against full-precision calculations, within 1%, with floors of 0.1 €m for amounts and 0.02 for ratios. The worked worksheet displays amounts to one decimal and ratios to two.
- Sample. Forty from-scratch worksheets per model, with issuers drawn separately. Continuation and audit use the same sixty planted items for every model, twenty per error class. The repair condition uses sixty separately drawn items. Clean controls comprise fifty-one calls per condition for Haiku and twenty for each other model. The 842 total counts scored worksheets and audits, not unique issuers or API attempts.
- Planted errors. The sixty continuation/audit items were at least 19% off, with a median error of 90%; thirty-nine were a factor of ten out or the wrong sign. The error class depends on which bound the resulting worksheet first violates. These audit rates describe the planted errors tested here.
- Continuation scoring. We inspect the first four downstream steps whose correct values differ from those obtained by propagating the planted error. Credit requires at least two values within tolerance of the correct chain, and more matches to that chain than to the propagated one. Three planted items have fewer than two such steps and cannot qualify; they remain in the denominator of sixty. An audit scores a correction only if it lists the planted row and supplies a value within tolerance.
- Clean audits. Haiku returned an "errors found" verdict on 16 of 51 clean audits, Sonnet on 9 of 20 and gpt-oss on 7 of 20. No proposed numerical correction exceeded tolerance. Counting any non-empty error list as an alert raises Haiku's count to 39 of 51; some entries confirm a row's value while filing it as an error. The notes observations in the article come from direct inspection.
- Settings and scope. Haiku and gpt-oss ran at temperature zero; Sonnet used default sampling with thinking disabled. gpt-oss returned visible reasoning. Most Sonnet and gpt-oss audits were retried at an 8,000-token limit after reaching the initial 2,000-token cap; continuation began with 4,000. Final outputs had no truncation or parse failures. These are comparisons of the configurations and tasks used here. We did not test a model auditing its own completed answer or checking during completion. The from-scratch experiment uses one fixed sequence; Sonnet's three first failures and gpt-oss's two provide limited evidence about the shape of a decay curve. Figure 1 uses Wilson 95% intervals.
The worksheet
The prompt has three blocks — the issuer's inputs, the twenty rules in words and in symbols, and the worksheet itself, twenty rows of Step k — name: value, with ? where the model is to write. These are the inputs for one of the 400 issuers, exactly as the model saw them, followed by the twenty steps with that issuer's correct values.
One interest figure. A different financial picture.
Inspect the issuer’s inputs, then change cash interest in step 6. The remaining rows follow the article’s rules. Your changes are a calculation exercise, separate from the model results.
The original issuer, exactly as supplied to the models. Amounts are in euro millions. A fraction of 0.216 means 21.6%.
| Input | Symbol | Value | Unit |
|---|---|---|---|
| Operating assumptions | |||
| Revenue | R | 213.50 | €m |
| EBITDA margin | m | 0.216 | fraction of revenue |
| Depreciation and amortisation | DA | 8.26 | €m |
| Capital expenditure | CAPEX | 5.49 | €m |
| Change in working capital | DWC | 5.40 | €m |
| Tax rate | t | 0.250 | fraction of taxable profit |
| Opening cash | C0 | 24.01 | €m |
| Debt balances | |||
| Term loan outstanding at t0 | TL | 96.20 | €m |
| Bond outstanding at t0 | B | 77.90 | €m |
| Mezzanine outstanding at t0 | M | 54.80 | €m |
| Revolving credit facility drawn at t0 | RCFd | 10.40 | €m |
| Revolving credit facility size | F | 23.89 | €m |
| Financing terms | |||
| Term loan rate | rTL | 7.5 | % per year |
| Bond rate | rB | 6.5 | % per year |
| Mezzanine rate | rM | 13.7 | % per year |
| Revolving credit facility rate | rRCF | 5.4 | % per year |
| Mandatory amortisation rate | a | 0.010 | fraction of TL per year |
| Cash sweep share | s | 0.250 | fraction of excess cash flow |
| Covenants and valuation | |||
| Maximum permitted net leverage | Lmax | 5.00 | × EBITDA |
| Minimum permitted interest cover | ICmin | 2.00 | × cash interest |
| Minimum liquidity requirement | Lmin | 4.26 | €m |
| EBITDA growth over the year | g | -0.034 | fraction |
| Distressed enterprise value multiple | k | 4.00 | × EBITDA |
View the original twenty-step answer key
| Step | Line | Rule | This issuer |
|---|---|---|---|
| 1 | EBITDA0, opening EBITDA | EBITDA0 = R·m | 46.1 |
| 2 | G, gross debt | G = TL + B + M + RCFd | 239.3 |
| 3 | ND0, opening net debt | ND0 = G − C0 | 215.3 |
| 4 | Lev0, opening net leverage | Lev0 = ND0 / EBITDA0 | 4.67 |
| 5 | HL0, leverage headroom | PASS if Lev0 ≤ Lmax; HL0 = Lmax − Lev0 | 0.33, PASS |
| 6 | I, cash interest | I = TL·rTL + B·rB + M·rM + RCFd·rRCF | 20.3 |
| 7 | IC0, interest cover | IC0 = EBITDA0 / I | 2.27 |
| 8 | HIC0, cover headroom | PASS if IC0 ≥ ICmin; HIC0 = IC0 − ICmin | 0.27, PASS |
| 9 | TAX | TP = EBITDA0 − DA − I; TAX = t·max(TP, 0) | 4.4 |
| 10 | FCF, free cash flow | FCF = EBITDA0 − CAPEX − DWC − I − TAX | 10.5 |
| 11 | AMORT, mandatory amortisation | AMORT = a·TL | 1.0 |
| 12 | SWEEP, cash sweep | ECF = FCF − AMORT; SWEEP = min(s·max(ECF, 0), TL − AMORT) | 2.4 |
| 13 | TL1, year-end term loan | TL1 = TL − AMORT − SWEEP | 92.9 |
| 14 | C1, year-end cash | C1 = C0 + FCF − AMORT − SWEEP | 31.2 |
| 15 | ND1, year-end net debt | G1 = TL1 + B + M + RCFd; ND1 = G1 − C1 | 204.8 |
| 16 | EBITDA1, year-end EBITDA | EBITDA1 = EBITDA0·(1 + g) | 44.5 |
| 17 | Lev1, year-end leverage | Lev1 = ND1 / EBITDA1; PASS if Lev1 ≤ Lmax | 4.60, PASS |
| 18 | LIQ, liquidity headroom | LIQ = C1 + (F − RCFd) − Lmin; PASS if LIQ ≥ 0 | 40.4, PASS |
| 19 | EV, distressed enterprise value | EV = k·EBITDA1 | 178.2 |
| 20 | Recovery waterfall | EV to the revolver first, then the term loan and bond pro rata, then the mezzanine, then equity | recRCF 10.4; recTL 91.2; recB 76.5; recM 0.0; EQ 0.0; bond recovery 98.3% |
Amounts are € millions to one decimal and ratios to two, as the rules require; the answer key is computed at full precision.
© Dissei. All rights reserved.