Where equal-looking models come apart: two frontier families, 282 rollouts, one deal

Claude Sonnet 5 and GLM-5, evaluated on the same institutional credit case against fixed criteria, with deterministic checks where available: 47 tasks, three rollouts each. Mean reward 0.54 against 0.47, Sonnet 5 ahead on 34 of 47 tasks, and one answer in eight to one in five reasoned from information unavailable at the anchor date, and was penalised for it.

A model evaluated in this corpus is placed exactly in the position of an analyst. It receives the question, the date and the case materials as they existed at that time and nothing else; the rubric and pass conditions are held in a separate grader-only document, so the standard against which an answer is judged is neither visible to the model nor recoverable from the question itself. Because each criterion is fixed before any model runs and is anchored to source evidence, the grading is auditable: any contribution or deduction to the reward traces to a named criterion, and thus a document.1

Two frontier-model families were evaluated this way on the same documented transaction: Claude Sonnet 5 and GLM-5, 47 tasks, three rollouts each, 282 rollouts in all.2 Models of one family share blind spots, so a cross-family comparison is useful evidence about whether the grading measures anything the models do not have in common, though it does not prove that recognition or contamination played no part. Under this grading Claude Sonnet 5 records a mean task reward of 0.54 against 0.47 and leads on 34 of the 47 tasks; tasks that are hard for one tend to be hard for the other, and the variation between them is measured under the same fixed criteria for both. The result remains limited to this case and these runs.

Judgment here is graded as a route, not a verdict. Each task keeps the steps the model took: what it gathered and where it stopped. A verdict that is right for the wrong reason earns less than one that is right because the route was sound. Two assumptions trip up readers who know evaluation and not credit. The first is that evaluating finance means predicting outcomes. It does not: each answer is graded against the record as it stood on the day, not against how the deal ended. The second is that grading reasoning is too subjective to standardise. Here it is not: the criteria are fixed in advance and anchored to source evidence, and what this page can show is which mechanism fired, how often, and for which model, never the rubric itself. The part that transfers beyond credit is the grading, not the deal: any agent whose worth lies in its route can be scored this way.

One task, in shape. Standing at the anchor date with the lender presentation and the two latest quarterly reports in hand, state whether the borrower can meet its next interest payment from operating cash alone, and cite the line items that decide it. The case specifics and the criteria stay sealed; what the criteria reward is stated in the last section.

Higher is better
Claude Sonnet 5leads
0.54
GLM-5leads
0.47

On the headline number the two sit close. A single-number leaderboard would stop here; switch the lens to see where they part.

Equal on paper, different underneath. Switch what you measure. The averages nearly match, but on discipline, whether an answer leaned on hindsight, tripped the traps a seasoned analyst avoids, or wandered to reach the same place, the gap opens. It shows here because every answer is graded against the contemporaneous record and the route it took. Each bar runs from zero to a fixed axis per lens (0.70 reward, 40% of rollouts, 14 turns); the printed value is the measured result.

The grading mechanisms

A score in this corpus is a decomposed function rather than a holistic impression. Critical criteria act as a gate: an answer that misses one fails, whatever else it gets right. Pitfalls encode errors that professional analysts may very well make; triggering one penalises the result. Where a task grants the use of retrieval tools, a separate mechanism counts only what the answer grounds in the evidence it actually consulted. Temporal boundaries are enforced in two layers: the supplied materials stop at the anchor date, and an answer that reasons from hindsight is treated as a critical failure.

Every mechanism catches both models at material rates, and the two differ visibly in discipline. Roughly one answer in eight from Claude Sonnet 5, and one in five from GLM-5, reasoned from post-anchor information and was penalised for it. The chart below shows where the gap between them is widest. Nearly all rollouts, 88%, additionally forfeit some reward to imperfect evidence grounding.

Where the discipline gap is widest.

Bar chart · Claude Sonnet 5, GLM-5 (%)Critical criterion failure: Claude Sonnet 5 18%; GLM-5 15%. Pitfall triggered: Claude Sonnet 5 18%; GLM-5 30%. Post-anchor reasoning: Claude Sonnet 5 13%; GLM-5 20%. Equal spacing represents observation order. Missing observations remain gaps.%Critical criterion failure: Claude Sonnet 5 18%; GLM-5 15%Pitfall triggered: Claude Sonnet 5 18%; GLM-5 30%Post-anchor reasoning: Claude Sonnet 5 13%; GLM-5 20%Critical criterion failureCritical criterion failurePitfall triggeredPitfall triggeredPost-anchor reasoningPost-anchor reasoning
Claude Sonnet 5GLM-5

Hover, tap or focus a mark to inspect its values.

Investigative behaviour

The reasoning categories demand measurably different volumes and mixtures of evidence-gathering (timeline and exhibit retrieval, source reading, search), and the two models settle the same tasks at different episode cost. The chart below counts the turns and tool calls each spent per episode.

Same tasks, different effort.

Bar chart · Claude Sonnet 5, GLM-5 (per rollout)Turns: Claude Sonnet 5 7.7 per rollout; GLM-5 11.2 per rollout. Tool calls: Claude Sonnet 5 15.7 per rollout; GLM-5 19.1 per rollout. Equal spacing represents observation order. Missing observations remain gaps.per rolloutTurns: Claude Sonnet 5 7.7 per rollout; GLM-5 11.2 per rolloutTool calls: Claude Sonnet 5 15.7 per rollout; GLM-5 19.1 per rolloutTurnsTurnsTool callsTool calls
Claude Sonnet 5GLM-5

Hover, tap or focus a mark to inspect its values.

Sampling

Reward rises measurably with additional sampled attempts. Taking the best-of-three instead of a single attempt lifts both models by roughly a tenth of the scale, and that lift is the kind of variance a trainer converts into learning signal. Each case yields this signal across seven categories of reasoning at several anchor dates, so a single case produces many distinct tasks, not one question asked several ways.

Single attempt versus best of three.

Mean reward across 47 tasks from one transaction.

Single attemptBest of three

Mean reward · scale 0.450.65 · Δ change, rounded to 2 decimals

Claude Sonnet 5Δ +0.09
Claude Sonnet 5, Single attempt: 0.545 Mean reward0.545Claude Sonnet 5, Best of three: 0.638 Mean reward0.638
GLM-5Δ +0.10
GLM-5, Single attempt: 0.472 Mean reward0.472GLM-5, Best of three: 0.573 Mean reward0.5730.450.550.65

Hover, tap or focus a mark to inspect its values.

What the grading rewards

The design determines what the grading rewards. An answer that correctly declines to conclude when the evidence cannot support a conclusion earns credit for that restraint; an answer that grades well is one whose reasoning chain is sound and complete, not merely one whose verdict is correct. A critical failure decides the grade whatever else the answer gets right, and confident verbosity is in no way profitable. The profitable policy is that of a sound analyst: commit to a verdict, ground it in the evidence consulted, avoid the known errors, and decline when the evidence is insufficient. The mechanics remain sealed; the posture they enforce is the point.

We maintain no public leaderboard. The full evaluation report, the task-level results, and the environments behind them are available under evaluation agreement: tech@dissei.ai.

Notes

  1. Source evidence, date controls, and decomposed assessment are described at the public level in Built backwards from outcomes. The grading criteria remain sealed.
  2. Runs: 47 tasks × 3 rollouts per model (n = 141 each; 282 total), one documented transaction, grading fixed before any model ran, deterministic wherever a check is mechanical.
All research

© Dissei. All rights reserved. No reproduction, adaptation, or derivative use of this content or methodology without prior written permission.