Financial judgment for LLMs: why it is hard
Most finance LLM benchmarks test extraction. The decisions that matter require judgment: knowing which question decides the outcome, reasoning only from what was knowable at the time, and holding up under adversarial framing. This is what makes it hard, and how to test for it.
A finance LLM benchmark tests whether a model can reason about a financial situation, not whether it has memorised the finance curriculum. Most existing benchmarks grade a fixed target: a number pulled from a table, a label, a prescribed calculation, or a final answer against a reference. The decisions that matter in institutional finance are judgment calls: whether a restructuring thesis holds, whether a covenant breach is material, whether a revenue trajectory supports the leverage. Judgment here is not opinion; it is reasoning to a defensible position under incomplete information, and that is what makes it teachable, testable and gradeable. This guide explains what a benchmark must measure for those decisions, and how dated evidence, professional judgment, and later outcomes play different roles.
Task overview
Each task is built around one real decision from the institutional record. The model is placed at a date, given the documents that existed then, and asked for the judgment that was actually required. It is graded on what it concluded from what it could know, and on the route it took to get there.
Every task fixes an anchor date, and nothing after that date exists inside it. The model gets the record a professional actually had at the time: filings, exhibits, agreements, operational notes. It works that record with tools, the way an analyst would, rather than answering from memory, and it ends in a verdict: a view, a price, a structure, a recommendation. The standard it is held to was set before any model saw the task. It credits the route with partial credit rather than passing or failing the verdict. Later outcomes calibrate the standard; they never replace the contemporaneous judgment.
What financial reasoning actually means
Financial reasoning is not the ability to calculate every ratio on the page. It is the ability to identify which fact determines the outcome at a particular point in time, then express that judgment through price, structure, or a decision. Calculating the ratio is execution; knowing which ratio decides this case is judgment. A live-events ticketing platform entering early 2020 provides a concrete example.
Before the example, the basics in plain English: four ideas that carry most of the discipline.
Downside comes first. Lend $100 and the best contractual case is your $100 back plus interest; the worst case is zero. Your best case is known on day one and your worst case is not, so the question is never “how much can I make?” It is “what can I lose, what protects me, and am I being paid for it?” Most of a professional underwriter’s time goes into that downside work, and almost none of it is guessing future earnings. Modelling the upside is not low-value work; it is a category error, a portrait of the one future in which nothing goes wrong.
The company is not the investment; the claim and its price are. Every company is one continuum of risk, from the safest secured claim at the top to the shares at the bottom, over the same assets and the same cash flows. In $100 terms: one claim might lose $50 in the bad case to make $10 in the good case, a return of 0.2 per unit of risk. Another claim on the very same company might lose $10 to make $10, a return of 1.0, five times better paid. The story did not change; the claim and the price did. What separates the claims is protection, and protection is only worth what it would fetch in the conditions where you would call on it.
Price is evidence. Before any analysis starts, the observable market prices of a company’s securities are a free summary of what everyone else has already concluded. A bond that promises $100 at maturity but trades at 40 is telling you something. Skipping that step, or worse, inventing a price because the real one is hard to find, makes every number downstream fiction that reads confidently. A price you observed is a known; a figure the borrower supplied is a belief until checked.
A correct model is an instrument, not an analysis. Like a surveyor’s level: you calibrate it against known facts, then you use it for the actual job. If we grade only the instrument, we train instrument-builders. An analysis earns its place by what its output lets you conclude or do; a number that leaves your position exactly where it was did not need producing. Financial reasoning is knowing which question decides the outcome, and it is rarely the obvious one.
When COVID closed live venues, the obvious analyst question was whether live events would return, followed by a debate about the shape of the recovery curve. In March 2020 that question was unanswerable. More importantly, it did not decide whether a new lender would be repaid.
Through its payment-processing arm, the platform collected ticket money from buyers and could pay qualifying creators before their events took place. If an event was later cancelled and a creator could not fund the refunds, the platform could still face the chargebacks. The decisive credit question was therefore: whose cash is on the balance sheet, and how much of the refund obligation will land on the platform if creators cannot pay? Revenue risk was the visible story. Chargeback risk was the credit. The parties who could reach into the cash drawer decided the credit, and none of them appears in the accounts.
That distinction changed the analysis. Conventional leverage analysis said little while EBITDA was negative. Liquidity was the relevant constraint: a refund wave could consume it, and the disclosed facility's minimum-liquidity covenant began a year before its minimum adjusted-EBITDA covenant. The exposure was contingent, difficult to size, and capable of absorbing cash ahead of a new lender despite the lender's senior secured position on paper, a priority loan backed by named assets.
The May 2020 financing reflected that uncertainty. The facility provided up to $225 million, with $125 million funded initially and another $100 million available later subject to conditions. The transaction also included the lender's purchase of 2,599,174 Class A shares for $0.01 each, equal to about 2.5% of the company's outstanding common stock. Our underwriting interpretation is that staged funding limited exposure while the refund liability became observable, while the equity participation compensated for risk that spread alone could not price cleanly. Where a number could not yet be produced, the judgment was written into the structure instead.
The later disclosures make the reasoning gradeable. The company increased its reserve for anticipated chargebacks and refunds by $76.5 million. From March through May 2020, it funded less than $3 million of refunds and chargebacks, while creators refunded more than $150 million themselves. The reserve increase was more than twenty-five times the cash the platform funded during that window. That is not a final-loss ratio: it shows why a strong answer had to separate the size of a contingent exposure from the loss the company was likely to bear.
The same logic in round numbers anyone can check. Suppose the platform owes $30 of debt against $10 of yearly earnings: leverage of 3× against a covenant ceiling of 4.5×. Comfortable. Now earnings go to zero. $30 ÷ $0 is not a number, and the ratio stops measuring anything. Meanwhile the liquidity test asks something anyone can check: always keep at least $10 of cash in the drawer. The platform holds $25; a $20 refund wave leaves $5. Breach. No ratio involved. Leverage measures how indebted you are while the business runs; minimum cash measures whether you survive the month. In a shock, the cash test binds first, and here what drains the drawer is refunds of money that was never the platform’s to begin with. And because the hole could not yet be measured, judgment got expressed through structure rather than forecast: wire roughly half the loan on day one, release the rest only once the refund hole became observable, and take a small slice of equity for the risk no interest rate could cleanly price.
Why more experts can still create misaligned data
Hiring more domain experts does not automatically produce better training data. Labels inherit the perspective of the people creating them. Ask a group of analysts to grade this case and they may reward detailed work on revenue recovery, leverage, and covenants because those are legitimate analyst tasks. The result can be internally consistent and still train the model to answer the wrong question. The knowledge that is scarce is not what a role produces; it is which analysis this situation demanded, and why.
This is expert-data misalignment: the dataset captures what a role is accustomed to producing rather than the judgment the institution needs. The remedy is not fewer experts. It is a clear decision date, contemporaneous evidence, and expert review of the route taken through that evidence. The review that matters corrects a step, with the reason attached, not a conclusion. Later outcomes can challenge assumptions and calibrate the standard without replacing professional judgment.
What existing finance benchmarks test
If you know evaluation but not credit, two assumptions will trip you on this page. Both are reasonable. Both are wrong.
Trip-up one: finance evaluation means predicting outcomes. It does not. A lender’s upside is bounded by the contract on day one, so the upside is not where the uncertainty lives. The work is judgment under incomplete information: which fact decides repayment, what protects you if it goes wrong, and what that protection is worth on the day you would need it. A benchmark whose only answer key is “did the loan repay” is scoring one draw from a noisy distribution, with great precision.
Trip-up two: grading reasoning is too subjective to standardise. It is subjective in the way any expert grade is, and it becomes auditable once you stop grading opinions and start grading paths. The judgment is the trajectory: which analysis was commissioned, in what order, and why. Running an analysis once told which one to run is commodity work; the skill is the reasoning that picks the next step. A path has checkable properties: did what came before call for each step, did the route follow the branches its own results opened, and did it stop once the deciding facts were isolated. Experienced readers can disagree on a verdict and still agree, step by step, on which routes were defensible, and that agreement can be measured. That is the payoff: scoring the adaptive choice of the next step is what you cannot get from the benchmarks below, and it is the part that transfers to agent evaluation generally.
Now the field. Each benchmark below is named with one task lifted from it, so the claim about what it grades rests on its own material.
Existing finance benchmarks
The task each benchmark sets, and what it grades
FinQAChen et al., EMNLP 2021 · 8,281 questions over S&P 500 earnings reports
subtract(153.7, 139.9), divide(#0, 139.9). Answer: 9.9%.FinanceBenchIslam et al., 2023 · 10,231 questions across 40 public companies
FinBenXie et al., 2024 · 36 datasets, 24 tasks, incl. forecasting and credit scoring
Finance Agent BenchmarkBigeard et al., 2025 · 537 expert-written questions, live SEC filings, tools
FinChainXie et al., 2025, rev. 2026 · 58 topics, 12 domains, symbolic templates
Read across the rows and the pattern is uniform. Every task hands the model the analysis and grades what comes out of it: a number, a class, a direction, a rubric hit. Two of them (FinQA, FinChain) do score the computational path, and it is a prescribed path. None of them asks the question that decides a credit: given what you have seen so far, what do you need to look at next, and why? That question is what the next section is about.
What a finance LLM benchmark must measure
Start from what judgment is, because the benchmark has to measure that and nothing smaller. Judgment is a path, not a possession. No hidden judgment step sits behind the analysis. The judgment is the order the analysis happened in, and it loops rather than runs straight:
The reasoning process
- Understand the situation, not the documents. How does this business make its money, who does it rely on, who else holds a claim, what does the law say about your position.
- Derive what can go wrong from that understanding. Not from a checklist: a business that depends on three customers fails differently from one that depends on a commodity price.
- Sort knowns from beliefs from unknowns. A borrower-supplied figure is a belief until checked.
- For each unknown that matters, decide what would turn it into something you can measure. Which document, enquiry or analysis turns this worry into a number with a boundary around it. This is the step the profession is paid for, and the step every benchmark above skips.
- Run it. Execution. Machines already do this well.
- Interpret it: what does the result now permit you to conclude, and what does it oblige you to look at next? A result closes a question, opens two, or says the question was wrong. Then go back to step four, or stop, because you have isolated what decides the deal.
So the benchmark must grade the route. Here is the whole distinction in one exercise. You have already seen the ticketing case on this page, so this is not a test of whether you remember it: both analysts below reach the same recommendation, word for word. What differs is the path, and you are asked to judge which path earned it.
Three rules for the grade fall out of that exercise.
Score the path with partial credit, not the answer pass or fail. A right answer by an unjustified route is luck; a wrong answer by a well-reasoned route shows where it went wrong, which makes it fixable. The unit of grading is the step: was it demanded by what came before, did it open a branch that was then followed, did the route stop in the right place. This is what makes expert judgment auditable without pretending it is objective. The correction worth having is a correction to the route: “that was not the thing to look at; this was, and here is why the case called for it.” That sentence is the transferable object, and a machine’s trajectory is the first place it has ever been fully visible.
Do not fix universal weights; fix case-specific ones, then lock them. A rubric that gives every consideration the same weight on every deal pretends every deal has the same shape. What decides a deal changes deal to deal, so which considerations carry weight is itself part of what is being tested. Standards count only when they are live: a standard about valuing security should not fire on an unsecured deal. So what counts, and how much, is decided by the case and settled before any model responds, which keeps the grade situation-specific and closed to hindsight. A universal checklist fixes the route before anyone has looked at the case, which is the opposite of judgment, however good each item on it is.
Test reasoning, not recall, by turning the question around. The naive task asks for a probability: “what is the chance this fails?” A bare number is hard to defend and hard to check. Because a lender’s upside is fixed, the point where the deal stops working is a calculation: the collateral must not lose more than a stated share of its value over the loan’s life. So the task becomes: here is the line, take a side, and show the evidence, assumptions and counterfactuals behind the side you took. Give the model a number and ask it to take a side, rather than asking it to invent one. The arithmetic is separated from the view, the view is specific to this deal, and it resolves later, so the record shows whose views on which questions were reliable. A model can still guess from a prior, which is why the support matters more than the side; this is a falsifiability test that complements path grading, not a replacement for it.
Everything above is about the shape of the task and the grade. Five further capabilities have to be measured as well, and existing benchmarks largely ignore them. Measured across our evaluations:
- 13% and 20% of answers used information unavailable at the anchor date (two-family evaluation).
- One line of market chatter moved a valuation by $45 million; quadrupling the company’s biggest risk factor left the price unchanged (564-run prompt-framing study).
- Half a turn of EBITDA (a valuation shift of 0.5× yearly earnings) of contamination carried over from an unrelated deal’s chatter; explicit instructions to evaluate independently removed less than a third of the effect (anchoring study).
Diagnostic reasoning. Given a set of financial documents, can the model identify what is actually happening in the business? Not what the management says is happening. Not what the analyst consensus believes. What the underlying data supports. Our end-to-end case study includes tasks across seven categories of reasoning, and diagnostic is the foundation: every downstream judgment depends on whether the model reads the situation correctly. Reading the situation includes the legal position, before any number is produced, not after.
Temporal reasoning. Can the model reason from a point-in-time information set without importing information from the future? This sounds trivial and is not. Our two-family evaluation found that 13% and 20% of answers, respectively, used information unavailable at the anchor date. A benchmark that does not enforce temporal boundaries is testing a mixture of reasoning and recall, and it cannot distinguish between them.
Sensitivity to framing. Does the model's output change when the same financial situation is presented differently? Our research shows that it does, and the magnitude is alarming. Quadrupling a company's biggest risk factor left the AI's price unchanged; one line of market chatter moved it $45 million. † A benchmark that presents each task in a single framing cannot detect this. A benchmark that holds up must test the same judgment under multiple framings and measure the variance.
The 564-run prompt framing study quantified this precisely: across controlled variants of one synthetic deal, the recommendation never moved, but the price, leverage, and what the memo noticed moved with the packaging of the question. Scaffolding away the surface-level flip did not remove the deeper framing effect.
The same deal, seen twice. Nothing about the company changed between the two panels. Only the packaging of the question did.
A finance benchmark that does not test for this is measuring the model's performance on the specific question it was asked, not the model's judgment about the underlying situation.
Resistance to anchoring. Does information from one context contaminate the model's judgment in a separate context? In practice, models are used in sessions that span multiple deals, multiple analyses, multiple information environments. Our anchoring study † shows that market chatter from one deal's context moves the model's valuation of an unrelated deal by half a turn of EBITDA, and explicit instructions to evaluate independently remove less than a third of the effect. A benchmark that tests each task in isolation misses this entirely.
Grading integrity. The benchmark's own grading function must be tested. A grader that can be gamed produces misleading scores without any visible error. A grader with a high noise floor † produces scores that reflect the grading math rather than the model's performance. The rubric must lock before any model sees the task, or the evaluation is describing the model rather than testing it. These are not features of the model being tested. They are properties of the benchmark itself, and a benchmark that has not validated its own grading pipeline is not trustworthy.
Evidence and professional judgment
Finance evaluations need rubrics because judgment has several parts. A rubric can still reward confident form over sound reasoning, so the standard must also examine evidence use, temporal discipline, and known failure modes. Evidence use is not 'did it cite something'; it is whether the number chosen survives the bad case.
The decision must be judged from the contemporaneous record. Later outcomes provide external evidence and calibration, but they do not turn an uncertain decision into one correct historical answer. A sound decision can lose money, and a weak decision can get lucky. Those are properties of a route, and each is checkable without waiting for the loan to resolve.
The precision study † shows why this distinction matters. The model changed what it repeated when the same number used a different format, while price, verdict, and stated confidence did not move significantly. Precision in a number that changed nothing downstream is the signature of an instrument, not an analysis.
What evidence-linked environments look like
An evidence-linked finance environment places the model in a dated situation with the documents and tools available at that time. It assesses the information used, the analysis performed, the tools called, and the judgment reached. The recorded route an agent takes is judgment you can read.
Four things make that assessment trustworthy. The standard is set before any model sees the task. The grade reads the contemporaneous record and nothing later. The grader is itself tested, because a grader that can be gamed, or whose noise has never been measured, fails the benchmark before the model does. And hindsight never grades; later outcomes calibrate the standard and stop there. The principles are public. The machinery is not.
This is also where our benchmark differs most from everything else on the board. Most finance benchmarks grade instruments: can the model build the spreadsheet, extract the number, reconcile the file. We grade the analysis around the instruments: whether the decisive question was identified rather than the obvious one, whether every input traces to a dated source, whether prices were observed rather than invented, whether downside scenarios replay things that actually happened instead of imagining flattering ones, and whether the verdict holds up against the later record without hindsight doing the grading. Which of those criteria applies is decided by the case, not by the benchmark.
What the grade sees
- RewardedCites where every number came from. Each claim traces to a document that existed at the decision date. Claims tied to sources count.
- PenalisedInvents an input it could not find. No observable price available, so one is assumed. A fabricated input is not an assumption. The step that rests on it earns nothing, however clean the arithmetic downstream, and the steps that stand on their own evidence keep their credit.
- RewardedWorks the downside first. What can be lost and what protects the position, before anything about the upside. That is where professional attention goes, and where the standard sits.
- PenalisedAnswers the obvious question instead of the decisive one. A perfect analysis of the recovery curve while the real credit question sat in the documents. Correct analysis of the wrong question earns credit for the arithmetic and none for the choice, and the choice is what the benchmark measures.
- RewardedFollows the branch its own result opened. A result closes a question, opens two, or says the question was wrong. Credit for noticing which, and for going there next.
- PenalisedRuns the standard pack and stops. Every deal gets the same analyses in the same order, then a number is filed. A fixed trajectory is a checklist, and a checklist decided before anyone saw the situation is the opposite of judgment.
Later outcomes can reveal missed assumptions, support calibration, and test whether the evaluation remains useful. They are evidence about the decision, not the only reward.
How grading catches models, and what it found
The score is not one grader’s impression. It is built up from named criteria, each tied to a fact in the record, so every credit and every deduction traces to a document. Partial credit per criterion is what lets a nearly-right route score as nearly right.
The reason to build it that way is that the three ways judgment scoring fails are well known. Sparse rewards starve the learner. Noisy judges inject variance. Hackable rubrics teach exploitation instead of reasoning. A grade that reads the route, credits it in degrees, and discounts what was never grounded in the record leaves confident verbosity as an unprofitable policy. The profitable policy is that of a sound analyst: commit to a verdict, ground it in the evidence you consulted, avoid the known errors, and decline when the evidence is insufficient. Outputs can be graded. The route to them can be graded better, and that is the difference the numbers below measure.
Run against five frontier models (341 tasks × 3 rollouts each, identical materials and budgets), every mechanism fired at material rates on every model:
- 15–18% of rollouts failed at least one gate: fluent prose, missing the fact that decided the case.
- 18–30% triggered at least one penalty: the errors professionals make, caught by name.
- 13–20% reasoned from information that did not exist at the anchor date: hindsight, detected and penalised.
What the mechanisms caught
Observed shares across two model families on a 0–55% scale; these ranges are not confidence intervals.
% of rollouts · scale 0–55
Hover, tap or focus a mark to inspect its values.
The spread · five models, same tasks
Published highest and lowest means · 341 tasks × 3 rollouts each · σ ≈ 0.2 throughout
Hover, tap or focus a mark to inspect its values.
The same runs show the spread a judgment benchmark must produce: mean task reward from 0.49 for the strongest model down to 0.16 for the weakest. Scores spread across the range rather than clustering. The benchmark discriminates instead of blurring. And reward rises measurably with repeated attempts (a consistent +0.07 to +0.11 best-of-three uplift on every model), which is exactly the variance a trainer converts into learning signal.
Choosing or building a finance LLM benchmark
If you are choosing a benchmark for evaluating language models on financial tasks, or building one, the checklist is short.
Does it test judgment or extraction? A benchmark that only tests extraction is useful for extraction. It tells you nothing about how the model will perform on the decisions that actually matter.
Does it grade the route or only the destination? If the prompt names the analysis to run, the benchmark grades execution, however hard the arithmetic. A benchmark that grades judgment scores which step the model chose next, and gives partial credit for a defensible step that landed wrong.
Does it enforce temporal boundaries? A benchmark that lets the model reason from hindsight is measuring a blend of judgment and recall, and it cannot tell you which is which.
Does it test sensitivity to framing? A benchmark that presents each task once in one framing is measuring the model's performance on that specific framing, not its judgment about the underlying situation. One framing is one weather reading.
Has its grading function been adversarially tested? A benchmark whose grader can be gamed, or whose noise floor has not been measured, is producing numbers that may not mean what they appear to mean.
Does it distinguish decision quality from outcome? A benchmark should assess whether the judgment was defensible from the dated evidence. It should use later outcomes as evidence and calibration, not as the sole answer key.
Does it let the case decide what carries weight? A benchmark that scores every task on the same breakdown has decided in advance what matters, which is the decision the model was supposed to make.
Appendix
Explore a topic, then expand a worked example to follow the arithmetic. Glossary links open the relevant definition.
A. The instruments
Bonds, coupons, maturity and refinancing
A. The instruments
Bond. A tradable loan. Investors hand a company money today; the company promises interest along the way and the principal back at the end.
Coupon. The interest payment on a bond. Fixed (“6% yearly”) or floating (a reference rate plus a margin).
Principal / par. The amount to be repaid at the end, by convention “100.” A bond trading “at 40” costs $40 per $100 of promise.
Maturity. The date the principal is due.
Issuer. The company that borrowed by selling the bond. High-yield means the issuer is below investment grade, riskier, so it pays more.
Prepayment. The borrower repays early.
Call. The issuer buys the bond back before maturity, at a contractually set price. The call premium is the extra over par it pays for the privilege, compensation to the lender for the future coupons given up.
Tap. Re-opening an existing bond to sell more of it on the same terms.
Refinancing. Repaying existing debt by borrowing anew. A company that cannot refinance at maturity defaults even if its business is running, which is why “is the market open to this company, at what spread?” is a survival question, not a market-trivia question.
Default. Failure to pay or to keep a term of the contract. Restructuring is the renegotiation that follows, maturities extended, coupons changed, debt swapped for equity.
Bankruptcy / Chapter 11. Bankruptcy is a court-supervised process for dealing with a company that cannot meet its obligations. Chapter 11 is the US reorganisation process: the business can keep operating while a plan allocates value among claims according to legal priority.
Amendment / waiver. A negotiated change to the loan contract, or a one-time pass on a breached term. Bondholders vote on these.
B. Capital structure
Seniority, recovery and structural subordination
B. Capital structure
A company’s capital structure is the ordered stack of everyone who has a claim on its cash and assets. Simplified, top to bottom:
- Super-senior debt ranks ahead of other financial debt. Senior secured debt is backed by collateral (specified assets the lender may enforce against), but statutory, insolvency and other super-senior claims can still rank ahead.
- Senior unsecured, next, no specific collateral.
- Subordinated, contractually agreed to stand behind the others.
- Equity, the owners; paid whatever is left, often nothing.
Seniority is the queue position. Secured vs unsecured is whether specific collateral backs you. A guarantee is another entity in the group promising to pay if the borrower doesn’t; a guarantor is that entity. Recovery is what a claim actually gets back in a failure, and the recovery waterfall is the calculation that pours the company’s remaining value down the stack, filling each claim in order, the top may get 100 cents on the dollar while the bottom gets zero.
Structural subordination is the trap of lending to the wrong box: companies are trees of legal entities, and if you lend to the parent while the assets sit in a subsidiary that never guaranteed your loan, every creditor of that subsidiary gets paid from those assets before anything flows up to you. “Senior” on your own paper, last in line in fact. A restricted subsidiary is one bound by the loan’s covenants; an unrestricted one is outside the fence, and value that migrates there, through permitted baskets and carve-outs, the contract’s negotiated exceptions, can walk away from your claim legally.
Our central claim about credit, restated with this vocabulary: the stack is one continuum of risk over the same company, and the investment decision is choosing the claim and the price, not judging the company in the abstract.
C. EBITDA versus Covenant EBITDA
Accounting earnings and the contract’s definitions
C. EBITDA versus Covenant EBITDA
EBITDA in ordinary accounting language: earnings before interest, taxes, depreciation and amortisation, a rough proxy for the yearly operating cash earnings of the business.
Covenant EBITDA is a contract-defined number that starts from something like accounting EBITDA and then applies negotiated add-backs (items the borrower may add back to make earnings look bigger, e.g. “one-time” costs, projected synergies), usually with caps (limits on how much may be added back). The definitions section of the loan agreement is the law. The two numbers can differ materially, and every covenant test runs on the contract’s number, not the accountant’s. Use the wrong one and the arithmetic can be perfect while the compliance answer is fiction.
D. Liquidity versus leverage
Cash, covenant tests and headroom
D. Liquidity versus leverage
Liquidity is cash available now. Leverage is indebtedness relative to earning power, usually debt ÷ EBITDA. A covenant is a promise in the loan contract; a leverage covenant says “debt ÷ EBITDA stays below X”; a minimum-liquidity covenant says “cash never drops below Y.” Headroom is the distance between where the company stands and the covenant limit. A breach is crossing it, which typically hands lenders rights (demand repayment, renegotiate, enforce on collateral).
They can give different answers because they measure different failure modes. Leverage is a ratio: it needs earnings to mean anything, and in a crisis where earnings go to zero or negative, the ratio stops working exactly when you need it. Cash is a level: it stays meaningful in every state of the world. The documents define each test; the facts and scenarios determine which one binds first. When the shock came, it was minimum cash that bound. Which test binds is a property of the situation, and it is the situation, not the template, that has to be read.
E. The arithmetic step by step
Five worked examples with the full reasoning
E. The arithmetic step by step
01Leverage
Leverage
Debt $30, yearly earnings $10, so $30 ÷ $10 = 3.0×. Covenant at 4.5×, so headroom exists. Now earnings fall to $0: $30 ÷ $0 is undefined, division by zero is not a large number, it is no number, and against negative earnings the ratio is nonsense (−3× is not “better” than 3×). A test that returns nonsense in the state of the world you bought protection for is not protection.
02Minimum liquidity
Minimum liquidity
Cash $25, refund outflow $20, so $25 − $20 = $5. Requirement: at least $10. $5 < $10, so breach, although the earlier, pre-shock leverage test had looked comfortable. Two tests, same company, opposite answers, the analysis is knowing which one binds first.
03A reserve is not an expected loss
A reserve is not an expected loss
A company recorded or estimated $75 of possible refund exposure; it ultimately funded less than $3. Using $3 as an upper bound, $75 ÷ $3 = 25×; because the actual payout was below $3, the true multiple was higher. The $75 was possible contingent exposure, not automatically the probability-weighted expected loss. Treating it as expected loss would confuse exposure with likelihood. The reverse error matters too: the outcome does not retroactively prove the $75 estimate was right or wrong. Outcomes grade analyses only with care; poker players call the opposite mistake “resulting,” judging a decision by how it turned out rather than by what was known when it was made.
04Return per unit of downside
Return per unit of downside
Claim 1: upside $10, downside $50, so 10 ÷ 50 = 0.2. Claim 2 on the same company: upside $10, downside $10, so 10 ÷ 10 = 1.0, five times better paid for the same corporate story. Limits of this teaching ratio: it ignores probabilities (a 0.2 with a tiny chance of the bad case can beat a 1.0 with a big one), timing, liquidity (can you exit?), correlation with everything else you hold, and sizing. Institutions use richer machinery; the ordering intuition, compare claims by what they pay per unit of what they can lose, survives into all of it.
One habit worth keeping whenever shocks stack: percentage falls compound multiplicatively rather than adding. Lose 30% of revenue, then 20% of what remains, and you are down 44%, not 50%.
05Price versus promise
Price versus promise
A bond promises $100 at maturity but trades at $40. Buying at 40: if the company ultimately pays in full, the price gain is $100 − $40 = $60 on $40 invested, plus any coupons actually paid along the way. But that $60 is the best case, not the expected result, the market price of 40 is itself the evidence that full repayment is doubtful. The honest decomposition separates: the price you pay (40), the contractual promise (100 plus coupons), the time until resolution, the probability of each outcome, what you’d recover in the failure case, and whether you could sell meanwhile (liquidity). An analyst who assumes the price is 100 because none was easy to find hasn’t made an assumption; they’ve fabricated the single input every downstream number depends on, and may reach conclusions like “short the shares” without noticing that the real debt price already shows distress is priced into the capital structure.
F. Remaining terms briefly
Underwriting, spreads, valuation and risk
F. Remaining terms briefly
Underwriting. The analysis done before committing money: what can go wrong, what protects us, are we paid enough.
Credit. Practitioner shorthand for a loan or bond position, the thing being analysed. A “calm credit” is one where nothing eventful happens.
Term sheet. The structured summary of an instrument’s terms, parties, amounts, dates, rates, protections.
Spread. The extra yield over a reference rate that the market charges a borrower for its risk; wide spreads mean expensive or closed markets.
Claim. Any right to payment from the company, a bond, a loan, a lawsuit judgment, a refund obligation. Tranche: one slice of debt issued on its own terms.
Contingent claim. An obligation that only crystallises under conditions, litigation, chargebacks (card payments reversed by the buyer’s bank), refund waves. Often invisible in bond data; it may rank ahead of a lender, or rank behind it yet still drain the cash and force a restructuring.
Covenant compliance certificate. The borrower’s periodic formal calculation demonstrating covenant compliance, using the contract’s defined terms.
Run-rate. The last twelve months annualised, with no assumed growth.
Comparable (“comp”). Another company or transaction used as a measuring stick. The judgment is in choosing which ones, and in including dead companies so the sample isn’t survivor-biased. Survivorship bias: measuring only the companies that survived, which makes history look milder than it was.
Valuation multiple. Price expressed as a ratio, e.g. enterprise value (the worth of the whole business, debt plus equity) ÷ EBITDA; lets you compare across sizes. A trough multiple / trough value is the lowest such ratio or asset price actually observed in past downturns. A haircut/discount is a reduction from a reference value; the size of that reduction needs evidence rather than invention.
Yield / return. What you earn relative to what you paid; for bonds, driven by coupons plus the pull toward par (or away from it).
Downside. What is lost in the bad case, after protections.
Short (a position). A bet that a price will fall, you profit if the security drops. “Short the shares” means betting against the equity.
Position size / sizing. How much money is committed to one investment. “In what size” is the how-much decision that follows the what and the at-what-price.
Debt capacity. How much borrowing a company’s cash flows could support, a cheap cross-check on any valuation (a business “worth” far more than anyone would lend against it deserves suspicion).
Risk-adjusted return. Return measured against the risk taken to earn it; the return-per-unit-of-downside comparison above is its simplest form, and how to measure the risk part well is an open question across the industry.
Calibration band. A committed range (“recovery lands between X and Y”); calibration is whether, across many commitments, reality lands inside the bands about as often as promised. Sharp and honest is the target; infinitely wide bands are honest and useless.
If you are evaluating models on finance tasks, or building benchmarks for financial reasoning, we will scope it with you directly.
Talk to us