For the team evaluating us

Questions a model team tends to ask first.

Short answers. Anything deeper is covered on a call, under mutual NDA.

What is an environment, concretely?

A financial environment is a long-horizon case built from primary-source records. An agent may need to search documents, use spreadsheets, write code, calculate, retrieve evidence and revise its view over multiple steps. It is evaluated on the trajectory, not only the final answer.

Markets are noisy. How can an outcome be a reward?

The task is framed at the moment of decision and graded on the reasoning that was defensible then, not only on the dollars that followed. A sound call that lost money still scores. A lucky one does not.

How do you keep the answer from leaking into the eval?

Point-in-time by construction. A model sees only what was knowable at the anchor, the outcome is held back, and the framing is checked for the tells that quietly hand over an answer.

Can the grader be gamed?

We attack every grader before it ships, testing for evidence dependence, numerical leakage and structural leakage. Detailed test operations remain sealed.

What happens when models saturate the tasks?

The set grows with Dissei’s lending activity, which is the primary source of new decision data. Secondary licensed origination can add further cases under its own provenance and rights. Grading standards stay fixed while the task set grows.

Why does finance need a different kind of verifier?

Financial agents already use code, spreadsheets, search, document analysis and other tools. The distinction is not the toolset, but what the system must verify. Code and mathematics often produce outputs that can be checked directly. In finance, a model can calculate correctly and still miss a covenant, use hindsight, or reach an indefensible decision from the evidence available at the time.

Existing benchmarks largely test answers or isolated components; few evaluate complete decision trajectories against point-in-time evidence and later outcomes. Dissei evaluates the full long-horizon process: the information used, the analysis performed, the tools called, and the judgement reached, against the contemporaneous record.

Later outcomes can provide external evidence and calibration, but they do not alone determine reward. A sound decision can lose money, and a weak decision can get lucky.

Why not build this in-house?

Adapting a public case framework gets you the principles, not the operating reality. Finance environments worth training against have to be written by people who have actually lent — and sourcing that judgement through a marketplace of annotators hands you recruiting, vetting, and quality control as extra jobs, on top of noisy data.

For a bank or a fund, evaluation engineering is not the business. Building that capability means hiring frontier-grade technical talent and running it as a research function — a cost centre competing with labs for the same people.

This is the only thing we do. One domain — finance, across its verticals and adjacent fields — with our own lending practice underneath it.

Who actually builds these?

People who have carried the risk lead the financial standard and author the work. Software and AI teams build the systems that turn it into data, environments, and evaluations.

Where does the data come from — and do you have the right to train on it?

Documented deals, not scraped text and not commissioned opinion. Dissei lending is the primary source, with secondary licensed origination added under its own rights. Every delivered training asset carries its right-to-train lineage and a per-asset contamination statement. Each product agreement states its access and use rights.

What do we actually receive, and where does it run?

For customer-perimeter products, we deliver the environment, its graders, and the evidence pack behind every claim into your own training stack. Sealed evaluations are hosted by us, end to end. Your training never touches our infrastructure.