Real data in, reliable environments out

The thesis behind Dissei Data: judgment is graded against what actually happened, or it is not graded at all.

Ask a frontier model to define a leveraged buyout. It will recite the mechanics, name every covenant in a credit agreement, and pass the exam. The model has read every textbook, and on anything with a settled answer, it is already fluent.

Now hand it a live decision. A borrower trips a covenant in a downturn. The facts are still incomplete, real money is at stake, and no textbook settles what comes next. Knowledge stopped being the constraint. Judgment is the constraint.

Training judgment runs into one problem: verifying it. Most of finance has no answer key. Two sensible people read the same deal and disagree. The only arbiter that ever settles it is what actually happened afterwards.

Define a leveraged buyout, and name the covenants in a standard credit agreement. A borrower trips a covenant in a downturn, the facts still incomplete and real money at stake. Extend, restructure, or call the loan?
Answered cleanly. The mechanics, the definitions, the exam questions: the model has read every textbook and recites them cold. The polish drops away. Two sensible people read the same file and disagree, and no textbook settles which one was right.
Answer key: exists, and it is unambiguous Answer key: only what actually happened next
Knowledge is not the constraint. On anything with a settled answer, the model is already fluent. This is judgment, and it can only be graded honestly against the outcome on the record. That record is the scarce asset, and it is what we build environments from.
Same model, two kinds of question. The first has an answer key in the book, and the model has read every book. The second has none, so the only arbiter is what happened next. That is the judgment we grade, and the data it takes to grade it.

Same model, two kinds of question. The first has an answer key in the book, and the model has read every book. The second has none, so the only arbiter is what happened next. That is the judgment we grade, and the data it takes to grade it.

Outcomes as the answer key

Grade against a real outcome and the score changes character. The key exists before the model runs. It lives outside any training corpus. It does not move when opinions move.

A task graded this way keeps its meaning: the same question, scored the same way a year later, still measures the same thing. That is what lets an eval double as an audit trail.

Verified experience is the scarce asset

Our thesis: the scarce asset in the next phase of AI is verified experience. Complete records of real decisions, with everything that was known at the time, and outcomes that are a matter of record. Hold that, and you can build environments where a model works the decision and gets graded against reality.

That is what Dissei Data builds. We take real financial decisions, reconstruct them as they stood on the day, and turn them into RL environments, tasks, and verifiers. The pipeline is industrial: one case or a thousand, same discipline, same controls. Around it sits a full evaluation suite: where a model holds, where it breaks, which failure modes recur, and whether a change actually moved anything.

Real data in, reliable environments out.

What ships here

The research on this site is the working record. How tasks get built. Where models come apart. What survives scrutiny and what does not.

If a claim ships here, it survived our own attempts to kill it.

This post is the thesis. Everything after it is evidence.

If you are training on finance data, or you have deals to put to work, we will scope it with you directly. [Connect with us](mailto:tech@dissei.ai)

If you are training on finance data, or you have deals to put to work, we will scope it with you directly.

Connect with us