Skip to benchmark
Dissei Benchmark / Financial reasoning

Financial reasoning,
measured.

How well do models turn financial evidence into sound analysis? Explore real evaluation results across seven kinds of reasoning.

About the benchmark
Observed results

Model comparison

Historical, pooled results from case-foundry evaluations.

5 models
Model results as of 2026-08-31. Coverage varies by model. Scores and gate pass rates are higher-is-better; costs are estimates.
GPT-5.6 Sol
45.53
80.6%$0.285
Claude Opus 4.8
44.26
81.4%$0.823
Muse Spark 1.2
42.41
76.4%$0.284
Kimi K3
39.76
78.6%$0.196
DeepSeek v4-flash
25.49
67.1%$0.027

Case coverage differs between models. Costs cover the policy model only.

Sorted by Score, descending.

Reasoning profile

Where performance changes

Compare two models

Seven reasoning categories · select a heading to explore.

Each cell shows score / 100.
Score out of 100 by model and reasoning category.
Model
Sol38.944.859.337.745.752.653.9
Opus 4.847.338.938.136.943.451.061.5
Muse 1.253.635.729.934.648.750.150.9
Kimi K342.636.936.731.841.843.650.3
v4-flash24.222.625.726.524.524.934.3
Quantitative

Calculate, reconcile, and interpret financial quantities.

Illustrative question

How much headroom remains under a changed earnings assumption?

Efficiency

Score is one part of the picture

Cost, tokens, and tool use from the same source snapshot.

Cost and score

Observed scores, with differing case coverage.

Pareto chart: minimise cost, maximise observed score. Logarithmic cost axis.Observed score (/ 100)01020304050550.010.020.050.10.20.51Cost (USD / attempt) · log scaleGPT-5.6 Sol: 0.285 USD / attempt, 45.53 / 100; descriptive frontierSolClaude Opus 4.8: 0.823 USD / attempt, 44.26 / 100Opus 4.8Muse Spark 1.2: 0.284 USD / attempt, 42.41 / 100; descriptive frontierMuse 1.2Kimi K3: 0.196 USD / attempt, 39.76 / 100; descriptive frontierKimi K3DeepSeek v4-flash: 0.027 USD / attempt, 25.49 / 100; descriptive frontierv4-flash
Observed frontierOther models
GPT-5.6 Sol
Cost
0.285 USD / attempt
Observed score
45.53 / 100
How to read the frontier

Models have different case coverage. This frontier describes the displayed aggregates; it is not a matched benchmark ranking.

Filled points have no displayed alternative with a lower or equal cost and a higher or equal score, with at least one strict improvement. The pale line connects those observations in cost order; it does not estimate results between them.

Tokens per attempt

Mean token usage and tool calls.

  • Sol45,272Tool calls: 13.8
  • Opus 4.8147,362Tool calls: 12.2
  • Muse 1.2212,235Tool calls: 16.1
  • Kimi K357,415Tool calls: 8.0
  • v4-flash400,399Tool calls: 29.0
Side by side

Compare two models

Explore the differences across seven kinds of reasoning.

Summary for GPT-5.6 Sol and Claude Opus 4.8
MetricSolOpus 4.8
Score / 10045.5344.26
Gate pass80.6%81.4%
Policy cost ≈$0.285$0.823
Mean turns4.346.94

Coverage differs between models. These are descriptive differences across each model’s available cases.

  • QuantitativeGPT-5.6 Sol: 38.9Claude Opus 4.8: 47.3
  • DiagnosticGPT-5.6 Sol: 44.8Claude Opus 4.8: 38.9
  • ComparativeGPT-5.6 Sol: 59.3Claude Opus 4.8: 38.1
  • StrategicGPT-5.6 Sol: 37.7Claude Opus 4.8: 36.9
  • CounterfactualGPT-5.6 Sol: 45.7Claude Opus 4.8: 43.4
  • ExplanatoryGPT-5.6 Sol: 52.6Claude Opus 4.8: 51.0
  • PredictiveGPT-5.6 Sol: 53.9Claude Opus 4.8: 61.5

Sol Opus 4.8 · Scores use the same 0–100 scale. A missing score leaves a gap.

Reasoning in practice

What the grader sees

Two worked examples. Follow the decisive credit question, then inspect a recorded model assessment.

Company and lender names removed; distinctive public figures can still identify a case. The ticketing example presents case-level findings. The first-lien pricing example presents one recorded attempt. Their results should be read independently.

Example 01 / Financial judgment · Contingent liabilities

When the cash is not the company’s

Decision window March to May 2020

The question

Whose cash is on the balance sheet, and how much of the refund obligation falls on the platform rather than on event creators?

Situation

A listed ticketing platform faces widespread event cancellations. A private-credit lender is considering senior secured financing. Some creators have already received and spent the ticket proceeds.

The evidence

Ticket proceeds, obligations to creators, refund and chargeback exposure, and dated financing terms. Each task must use only the information available at its decision date.

The issue

Claude Fable 5Model

If creators cannot fund refunds, chargebacks can fall on the platform. A lender needs to establish how much cash is available to meet that obligation before judging whether the loan can be repaid.

How the models answeredWhat they got right and what they missed

Dissei’s findings across this case. These describe patterns in the answers, rather than a scored assessment of one response.

Correct calculations

The answers computed the financial ratios correctly.

The key question

The answers focused on when live events would return. They missed whose cash the platform held and how much of the refund obligation it would have to fund.

Accurate covenant citations

The answers cited the covenant package accurately and produced competent-sounding committee memos.

A clear conclusion

The grader notes describe generic committee language without a verdict. The answers did not turn the ratios and covenant terms into a conclusion about repayment.

From the grader’s notes

“generic bank-committee language, never committed to a verdict”
“answers the obvious question instead of the real one.”

These comments come from the grader. They are not quotations from a model’s answer.

The financingStaged funding and equity participation

The May 2020 documents show these terms. A task set in March cannot rely on these later terms.

Total facility
Up to $225M
Initial loan
$125M, expected to be drawn in May, according to the filing.
Delayed draw
Up to $100M, available December 31, 2020 through September 30, 2021, subject to conditions.
Equity participation
2,599,174 Class A common shares at $0.01 per share.
Our interpretationStaged funding limited initial exposure while the refund liability became more observable. Equity participation offered compensation for risk that spread alone might not capture. The documents establish the terms; this interpretation is Dissei’s.

Refunds can consume cash needed for debt service. That does not mean every refund claim has legal priority over secured debt.

The later evidenceThe reserve, the platform’s payments and creators’ refunds
Increase in reserves
$76.5M

For potential chargebacks and refunds.
First quarter of 2020.

Paid by the platform
<$3M

Refunds and chargebacks.
Start of March through the May 11 report.

Refunded by creators
>$150M

Payments funded by event creators.
Same disclosed period.

The reserve increase is an estimate of potential chargebacks and refunds. The payment figures show cash paid during a stated period. They do not establish the final loss or prove that the reserve was excessive.

What this showsA reviewer can check whether an answer separates the potential exposure from the amounts paid by the platform and by creators. Later disclosures help test the earlier reasoning. They must not be treated as facts the model could have known at the decision date.
What to checkDecision date, ownership of cash and the meaning of loss

Use these checks when reading an answer. A pass or fail requires the recorded response; none is assigned here.

Use only what was known then
Keep May financing terms and later payments out of a task anchored in March.
Establish whose cash it is
Separate the platform’s own funds from ticket proceeds owed to creators and refunds owed to buyers.
Distinguish exposure from loss
A reserve, cash paid during a period and a final loss measure different things. The disclosed payment figures do not establish the final loss.
The resultsResults across two model families and three attempts per task

Two unnamed model families used identical environments and sealed rubrics, at three rollouts per task. The rates below describe those two families across this case. They are not an individual score for Claude Fable 5.

12.8percentage points

Difference in pass@1

Pass@1 measures success with one attempt. The families had similar rates of passing critical grading requirements, but their graded outcomes differed. Close rates alone do not establish statistical equivalence.

Reported rates for this case. Family labels preserve the order of the evaluation summary.
MeasureFamily AFamily B
Critical-gate pass rate82.3%83.0%
Low-graded outcomes24.8%11.3%

Three specific tasks scored 0.000 across every rollout and model in the reported run set. One concerned the all-in financing cost and repeatedly failed a critical criterion. This is a task-level result, not a score for the entire case.

The weaker run also produced shorter answers. This observation does not establish a cause. The summary does not include raw counts, a definition of low-graded outcomes or uncertainty intervals, so it cannot establish a general model ranking or statistical significance.

Example 02 / Explanatory · Secured credit

Why pari passu paper trades at different prices

Evidence as of December 2022

Restated question

Why do the company's first-lien instruments trade at different prices? Compare legal rights, cash flows, effective maturities, liquidity and downside recoveries, check any springing terms, and identify the evidence needed to assess relative value.

Situation

A private credit fund weighing a first lien term loan quoted around 66 in a specialty pharmaceutical company months out of Chapter 11, whose equally secured notes trade far higher.

Evidence the model could retrieve

The memorandum’s capital-structure appendix with every tranche’s quote, its key-terms sheet, the maturity and coupon table, the liquidation analysis, the scenario cases, and comparable-company tables.

Claude Opus 546.4/ 100Recorded score · original task · one attempt

Credit and legal interpretation

The price gap alone does not establish a difference in legal priority or prove that the lower-priced instrument is cheaper. Equal ranking does not mean identical cash flows, contractual rights or credit exposure over time.

Gate
Passed
Rubric
6 of 10
Pitfalls
1 of 4 triggered
Effort
12 turns · 26 tool calls
When equal recovery is a valid assumptionDefine the legal rights and the recovery scenario

These are assumptions for a simplified downside comparison, not verified facts about this case. Check them against the agreements and applicable law.

Equivalent rights to payment and security
The same obligors, collateral and guarantees, with valid, perfected and enforceable equal-ranking liens, equal payment priority and ratable sharing without instrument-specific preferences.
A common recovery event
All instruments remain outstanding when default occurs before the earliest effective maturity, with no intervening preferential repayment or change in priority.
The same recovery measure
The same recovery value, timing and form per unit of allowed claim. Accrued interest, amortisation and other allowed claim components can make that different from equal dollars per unit of original principal.
Separate repayment scenarios
Earlier-maturing debt may be repaid or refinanced before a later default. Earlier maturity creates an earlier payment entitlement, not a higher lien priority. Coupons and repayments before default can also produce different total returns.

Shared collateral does not establish one bankruptcy estate or automatically place every instrument in the same Chapter 11 class. Classification and treatment are separate legal questions. See 11 USC §1122 and §1123.

Why equally ranked instruments can trade differentlyCash flows, credit exposure, contractual rights and market conditions
Coupons and interest rates
Coupon size, fixed or floating rates, benchmark floors and reset dates change expected cash flows and sensitivity to rates. Equal ranking does not require equal yields or prices.
Maturity and refinancing
Different repayment dates expose holders to different periods of default risk and different refinancing outcomes. Check effective maturity, including any springing provision, rather than relying on the stated date alone.
Amortisation and optionality
Scheduled paydowns, cash sweeps, call protection, make-whole provisions and prepayment rights change the amount and timing of cash received.
Covenants and control rights
Financial tests, debt and lien restrictions, amendment thresholds, enforcement rights and cross-default or cross-acceleration provisions can differ despite equal lien priority.
Liquidity and investor demand
Trading depth, bid–ask spreads, transfer restrictions, investor mandates and forced selling can affect prices. The case must supply evidence before any specific technical explanation is treated as established.
Comparable price quotations
Use the same observation date and currency basis, and reconcile accrued interest, settlement conventions and the principal amount outstanding. A quoted cash-price gap is not itself a relative-value calculation.

Compare yields and spreads on a consistent basis, but also compare scenario cash flows and recoveries. Yield to maturity assumes the promised payments occur; it does not settle a distressed comparison. FINRA explains bond cash flows, interest-rate risk, default risk and liquidity risk.

Springing covenant or springing maturity?Identify the trigger and its effect on each instrument
Springing financial covenant
A financial test applies when a specified condition is met, such as a defined level of revolving-facility usage. Establish the trigger, measurement date, facilities protected, cure and waiver rights, and consequences of breach. Do not assume a revolver covenant directly binds every term loan or note.
Springing maturity
The repayment date moves forward if a specified condition is met, for example if other debt remains outstanding near its maturity. The nominally later-dated instrument may then become due earlier. Establish the exact dates, thresholds and exceptions.
Effect across facilities
Trace any cross-default or cross-acceleration provisions and applicable thresholds. A breach, acceleration and the loss of borrowing availability are different outcomes; the agreements determine which follows.

The published assessment does not supply a springing clause or its trigger. Neither type of provision is asserted as a fact of this case. A separate SEC-filed example of a springing maturity illustrates the mechanism; it is not evidence for this December 2022 case.

Inspect the recorded assessmentWhat the model got right, what it missed, and pitfall checks

Original task

At the anchor date the company has several separate instruments outstanding within its first lien class. What should an investor conclude from where each of those instruments trades in the market?

Recorded assessment summary

The model proved the four instruments rank equally and named the holder-base technical, but never put the tranches on a yield basis, skipped the recovery grid, and treated the discount as free money.

The recorded grading includes stronger claims about identical credit risk, one estate and how much of the price gap a yield comparison explains. Equal ranking alone does not establish those conclusions. They require the agreements, a stated recovery scenario and the supporting calculations.

Selected criteria from the recorded assessment. The reported score is not recalculated from this selection.

What the model got right

  • The class does not clear at one price
    Met

    Read the four first lien quotes off the capital-structure appendix and named the term loans as the low end of the class.

    Model excerptFour first-lien instruments are, on the documents, a single pari passu class secured by the same collateral and the same guarantors — yet they trade across a 15-point band (66 / 67 / 75 / 81).
    Gate · weight 25
  • Commits to what the gap represents
    Met

    Settled on one account: the gap is timing and exit optionality plus technicals, not differential credit risk.

    Model excerptThe dispersion within the class says the market is not pricing lien position at all — it is pricing timing and exit optionality.
    Gate · weight 20
  • Parity grounded in the security language
    Met

    Cited the annual report’s same-assets, same-guarantors language and the key-terms sheet’s equal-and-ratable lien, rather than asserting parity from the ‘first lien’ label.

    Model excerpt1 Lien (equal and ratable with liens securing the 1 Lien Bonds) on substantially all assets of the Issuers and the Guarantors.
    Important · weight 12
  • A driver that is not credit
    Met

    Named the loan and bond markets’ different holder bases as a technical driver, while stating that the case material does not establish the company’s holder composition.

    Important · weight 10
  • Which claim to own
    Met

    Concluded that the term loans at 66 to 68 are the cheap way to own the identical claim.

    Optional · weight 5

What it missed

  • Fixed coupons against a floating loan
    Missed

    Identified the fixed-versus-floating split but argued it points the wrong way and dropped it. Never put the four instruments on a yield basis, which is where most of the 15-point price gap disappears for an identically ranked claim.

    Model excerptIn a rising-rate environment a floater should trade better, not 9–15 points worse.
    Important · weight 10
  • Distance to the 2025 refinancing wall
    Partial

    Named the 2025 maturity wall from the debt-service table, but treated the 2028 note trading above the 2027 loans as unexplained rather than pricing each instrument’s time to a par event.

    Important · weight 10
  • One estate in the downside cases
    Missed

    Quoted the 47% to 64% liquidation recovery but never used ratable recovery from one estate to rule out differential credit risk. Instead read part of the spread as intra-class subordination risk, which the record does not support.

    Model excerptPari passu today is not pari passu tomorrow.
    Important · weight 10
  • Valuation uncertainty is shared
    Missed

    Never touched the 69% to 132% implied-recovery grid or the half-multiple it is struck on. Used a single 2.69x market-implied multiple as if it settled first lien coverage.

    Important · weight 8
  • What a real ranking gap costs
    Missed

    Did not set the intra-class spread against the far deeper discount on the junior claims to show what a genuine ranking difference does to price.

    Optional · weight 4
  • Calling it a free arbitrage
    Pitfall triggered

    Presented the loan at 66 as a 15-point discount for the same claim without netting the $200mm instalment and the 2025 wall that the whole class must fund first.

    Model excerptYou have overpaid by 15 points for the same claim.
    Pitfall · penalty -2

Pitfall checks

  • Hindsight after the anchorNot triggered

    Every dated claim was cited to case sources; nothing after December 2022 asserted as an outside fact.

  • Reading the price gap as a ranking gapNot triggered

    Affirmed identical collateral and guarantees; the gap attributed to timing, not to a weaker lien.

  • Blaming a covenant the loan does not haveNot triggered

    Covenants cited only as evidence the instruments are equivalent.

  • Calling it a free arbitrageTriggered

    Discount framed as overpayment on the notes, with the class-wide cash calls left out of the comparison.

Read the model’s answer excerpt

Recorded model output. Read alongside the assessment; this excerpt can contain errors.

Buying the two term loans at 66–67 gets you the same lien, the same collateral, the same guarantee package and the same covenant suite as the April 2025 notes at 81 — a ~15-point discount for a two-year maturity difference within a class that, if the restructuring thesis is right, gets treated as one pool anyway. If you believe the company restructures, the April 2025 note’s near maturity buys you nothing and you have overpaid by 15 points for the same claim. If you believe it refinances, the loans re-rate hardest.
Methodology

What the benchmark tests

Financial reasoning grounded in evidence, measured with case-foundry.

From evidence to a decision.

Financial work combines calculation, investigation, and judgment. Dissei Benchmark brings those capabilities into one view: how a model interprets evidence, tests assumptions, explains a mechanism, and supports a decision.

Powered by case-foundry, the benchmark evaluates how models work with financial evidence across different kinds of reasoning. This view brings together historical evaluation results to explore their strengths, trade-offs, and resource use.

Financial evidenceModel analysisEvaluation

Reasoning in practice

Quantitative
Calculate, reconcile, and interpret financial quantities.
Diagnostic
Identify the causes behind a financial outcome.
Comparative
Compare alternatives on a consistent financial basis.
Strategic
Turn evidence into a defensible course of action.
Counterfactual
Trace the consequences of changing an assumption.
Explanatory
Explain financial mechanisms using the evidence.
Predictive
Assess plausible outcomes from information available at the time.

Score

The recorded evaluation reward, shown on a 0–100 scale, describes the quality of the model’s analysis.

Gate pass

Whether the response clears the evaluation’s required checks, reported separately from the overall score.

Cost & resources

Historical policy cost, tokens, and tool use show what it takes to produce the analysis.

Build on the evidence

Bring financial reasoning into your evaluations.

Explore a case, understand the assessment, and run it in your own workflow.

Talk to us