Our data foundry

From case evidence
to checked evaluations.

Our data foundry is built for frontier-model evaluation. It combines model-assisted task construction with practitioner judgment to turn dated financial records into cases, bounded questions, grading rubrics and evaluation artifacts—not just worked examples.

Anonymization & access

Preserve the reasoning.
Change the identifying detail.

Anonymized versions replace party names and identifying language throughout a case. Counterfactual versions also change figures and regenerate the expected answers, keeping tasks tied to the transformed evidence.

These checks target deal recognition; they do not guarantee zero re-identification or contamination risk. Access to source evidence and rights to use it for training are governed by agreement.

Construction & review
  1. Prepared evidence

    Point-in-time source records and anonymized versions, with identifying details transformed across the case.

  2. Bounded tasks

    Questions built from source-linked claims, with model cross-review and practitioner supervision.

  3. Grading criteria

    Requirements for a defensible answer, tied to evidence and frozen before candidate evaluation.

  4. Checked artifacts

    Tasks and rubric versions, with probe records. Failing tasks are rewritten or retired before release.

Evidence links retained throughout

Checks target specific failure modes. They do not prove error-free grading or agreement with a domain expert.

Repeatability comes from the method: source-linked claims, decision-date boundaries, frozen rubric versions and recorded checks.

Inspect the public questions

Seven question previews show task coverage, not a runnable evaluation. Full evaluation materials require agreement.