A model evaluated in this corpus is placed exactly in the position of an analyst. It receives the question, the date and the case materials as they existed at that time and nothing else; the rubric and pass conditions are held in a separate grader-only document, so the standard against which an answer is judged is neither visible to the model nor recoverable from the question itself. Because each criterion is fixed before any model runs and is anchored to source evidence, the grading is auditable: any contribution or deduction to the reward traces to a named criterion, and thus a document.1
Two frontier-model families were evaluated this way on the same documented transaction: Claude Sonnet 5 and GLM-5, 47 tasks, three rollouts each, 282 rollouts in all.2 Models of one family share blind spots, so a cross-family comparison is useful evidence about whether the grading measures anything the models do not have in common, though it does not prove that recognition or contamination played no part. Under this grading Claude Sonnet 5 records a mean task reward of 0.54 against 0.47 and leads on 34 of the 47 tasks; tasks that are hard for one tend to be hard for the other, and the variation between them is measured under the same fixed criteria for both. The result remains limited to this case and these runs.
Judgment here is graded as a route, not a verdict. Each task keeps the steps the model took: what it gathered and where it stopped. A verdict that is right for the wrong reason earns less than one that is right because the route was sound. Two assumptions trip up readers who know evaluation and not credit. The first is that evaluating finance means predicting outcomes. It does not: each answer is graded against the record as it stood on the day, not against how the deal ended. The second is that grading reasoning is too subjective to standardise. Here it is not: the criteria are fixed in advance and anchored to source evidence, and what this page can show is which mechanism fired, how often, and for which model, never the rubric itself. The part that transfers beyond credit is the grading, not the deal: any agent whose worth lies in its route can be scored this way.
One task, in shape. Standing at the anchor date with the lender presentation and the two latest quarterly reports in hand, state whether the borrower can meet its next interest payment from operating cash alone, and cite the line items that decide it. The case specifics and the criteria stay sealed; what the criteria reward is stated in the last section.
On the headline number the two sit close. A single-number leaderboard would stop here; switch the lens to see where they part.
—The grading mechanisms
A score in this corpus is a decomposed function rather than a holistic impression. Critical criteria act as a gate: an answer that misses one fails, whatever else it gets right. Pitfalls encode errors that professional analysts may very well make; triggering one penalises the result. Where a task grants the use of retrieval tools, a separate mechanism counts only what the answer grounds in the evidence it actually consulted. Temporal boundaries are enforced in two layers: the supplied materials stop at the anchor date, and an answer that reasons from hindsight is treated as a critical failure.
Every mechanism catches both models at material rates, and the two differ visibly in discipline. Roughly one answer in eight from Claude Sonnet 5, and one in five from GLM-5, reasoned from post-anchor information and was penalised for it. The chart below shows where the gap between them is widest. Nearly all rollouts, 88%, additionally forfeit some reward to imperfect evidence grounding.
Where the discipline gap is widest.
Hover, tap or focus a mark to inspect its values.
—Investigative behaviour
The reasoning categories demand measurably different volumes and mixtures of evidence-gathering (timeline and exhibit retrieval, source reading, search), and the two models settle the same tasks at different episode cost. The chart below counts the turns and tool calls each spent per episode.
Same tasks, different effort.
Hover, tap or focus a mark to inspect its values.
—Sampling
Reward rises measurably with additional sampled attempts. Taking the best-of-three instead of a single attempt lifts both models by roughly a tenth of the scale, and that lift is the kind of variance a trainer converts into learning signal. Each case yields this signal across seven categories of reasoning at several anchor dates, so a single case produces many distinct tasks, not one question asked several ways.
Single attempt versus best of three.
Mean reward across 47 tasks from one transaction.
Mean reward · scale 0.45–0.65 · Δ change, rounded to 2 decimals
Hover, tap or focus a mark to inspect its values.
—What the grading rewards
The design determines what the grading rewards. An answer that correctly declines to conclude when the evidence cannot support a conclusion earns credit for that restraint; an answer that grades well is one whose reasoning chain is sound and complete, not merely one whose verdict is correct. A critical failure decides the grade whatever else the answer gets right, and confident verbosity is in no way profitable. The profitable policy is that of a sound analyst: commit to a verdict, ground it in the evidence consulted, avoid the known errors, and decline when the evidence is insufficient. The mechanics remain sealed; the posture they enforce is the point.
We maintain no public leaderboard. The full evaluation report, the task-level results, and the environments behind them are available under evaluation agreement: tech@dissei.ai.
Notes
- Source evidence, date controls, and decomposed assessment are described at the public level in Built backwards from outcomes. The grading criteria remain sealed. ↩
- Runs: 47 tasks × 3 rollouts per model (n = 141 each; 282 total), one documented transaction, grading fixed before any model ran, deterministic wherever a check is mechanical. ↩