Teaching models
financial judgment.
Carefully built training data, evaluations and environments, grounded in real financial work.
Learning, under examination.
Studies in what models learn from finance, and where their judgment breaks.
Dissei BenchmarkFinancial judgment,
Financial judgment,
measured.
Explore the results across models and financial reasoning categories.
View benchmark resultsRecorded model scores
Historical snapshot · 2026-08-31
Claude Opus 4.8: Observed score / 100 44.26. DeepSeek v4-flash: Observed score / 100 25.49. GPT-5.6 Sol: Observed score / 100 45.53. Kimi K3: Observed score / 100 39.76. Muse Spark 1.2: Observed score / 100 42.41.
Observed score / 100