Sitting their own exam: four models, 397 questions, and a mild case of self-bias

Four models each wrote an examination from seven anonymised financial cases; all four then sat every paper blind. Who answers moves the score by 32 points. Who asks moves it by 6. On four fast, inexpensive models and questions that mostly rewarded lookup, writing the paper yourself is worth under a point, at most two and a half at the outside, and the lone hint of a larger edge sits exactly where a question demands a chain of thought.

An examination question can leak its answer. Every leak we have published so far sat in plain sight: an adjective that pre-empts the judgement, a framing the model should have had to find for itself, an avenue of enquiry laid out in the question itself.1 This piece is about a leak no human reader could catch. A model has habits of thought: preferred framings, familiar derivations, lines of reasoning worn smooth by training. When a model writes a question, the question follows a path that model can walk. Hand it to a different model and the path may run outside anything that model’s training laid down. Nobody chose the bias. It is simply in the question, invisible.

Training corpora are now vast enough that these habits ought to average out. But every evaluation is written by something, and whoever builds one must pick which model holds the pen. So we measured what the pick is worth.

models answering their own questions, a four-by-four grid

the whole experiment is one grid · 7 case briefs, 4 models writing and answering, 1,588 graded responses

Answered by ↓ · Question written by →

ClaudeGeminiGPTDeepSeek
Claude
Gemini
GPT
DeepSeek

Select, hover or focus a cell to inspect its value.

Rows are who answered, columns are who wrote. Hover any cell to read it against both margins.

share of rubric criteria met , averaged per question across the 283 questions that separate the models

ringed cells are a model sitting its own paper; the diagonal excess here is +0.4 points (95% CI −1.9 to +2.7)

What we built

Seven financial case briefs, each a self-contained anonymised document.2 Business and market, capital structure and covenants, summary financials, and the decision the protagonist faces. Nothing after the decision point appears.

Four models each wrote 105 examination questions from the briefs. Every surviving question was then answered by all four models, their own and the other three’s, blind: fresh context, no provenance, uniform instructions, identical settings. No model is ever told who wrote a question, or that authorship varies at all.

That yields a four-by-four grid, and the diagonal is the experiment. Everything else is fixed machinery applied identically to every cell: brief assembly, verification, rubric-writing, validation, grading. Fixed machinery can shift a row or a column; it cannot create a diagonal. The rubrics lock before any model answers.3 Of 420 authored questions, 399 survived validation and 397 entered the sweep; 1,588 responses were graded criterion by criterion.

Two margins and a whisper of a third

Read the grid by rows and you see who is answering: Claude meets 92% of rubric criteria across everyone's exams, Gemini 90, GPT 68, DeepSeek 60, a spread of 32 points. Read it by columns and you see who is asking: GPT sets the hardest exam at 73% for the average sitter, Claude and Gemini the gentlest at 79, a spread of 6 points. These two margins are the grid. Between them they explain nearly everything in it.

Then the diagonal. On this grid, sitting your own paper is worth four tenths of a point: three of the four ringed cells sit slightly above what the margins predict, one slightly below, and the interval spans zero. The full 397-question sweep says the same, more bluntly. The average self-excess there is a shade below zero; there are only 24 ways to pair four answerers with four writers, and ranked against all 24 the observed diagonal comes seventh, squarely inside what chance produces; the pre-registered regression lands at about a point in the author’s favour, comfortably inside noise.4

What moves a score on this exam: three levers compared

What moves a score on this exam

Which model answers
32 points

read by rows: Claude 92%, Gemini 90, GPT 68, DeepSeek 60

Which model asks
6 points

read by columns: GPT sets the hardest exam at 73% for the average sitter, Claude and Gemini the gentlest at 79

Answering your own author
≤ 2.5

under 1 point measured; 2.5 is the bound the data rule out

So whatever self-bias exists on this exam is mild. Across the full sweep, the data rule out an advantage larger than about two and a half points; nothing excludes an edge of a point or less. And it is dwarfed, by an order of magnitude, by the two choices evaluators actually make: which model answers, and which model writes.

Where the mild bias lives

Printing the discriminating grid is not a cosmetic choice. Across the full sweep the estimate is actually a shade negative, and that sign is manufactured: the 114 questions on which all four models earned identical marks say nothing about authorship, yet they drag the average below zero by restating solver strength. Remove them, as the grid above does, and the estimate turns positive. Small movements, inside noise. But they lean one way.

Go back to the mechanism this experiment was built around. A habit of thought can only leak into a question that contains thought: the bias we are hunting is a property of the route to an answer, not of the answer itself. Where an answer must be constructed, the writer's own derivation is baked silently into the question, which intermediates matter, which framing makes the arithmetic fall out, which road through the numbers feels natural. Two models can both reach a covenant cushion and walk different roads to it, and a question inherits its writer's road. But where the answer is printed on the page there is no route at all: nothing to derive, so no private way of deriving it.

This exam contained few roads, and that is an observation about the writers, not an accident of the pipeline: simple models set simple papers. Handed seven dense briefs and asked for examination questions, the fast tier reached for what was on the page. Two numeric answers in three could be copied straight off it; seven questions in ten drew on a single section of the brief or none. A writer does not demand a chain of reasoning it does not itself habitually run, and these models' habits run short. The pen bounded the paper, so the sweep, mostly retrieval, probed the mechanism precisely where the mechanism has nothing to grip.

Where the paper does contain a road, the edge appears. On the 74 questions requiring even one intermediate quantity the brief does not print (questions with an actual derivation inside them), authors beat expectation by about 4 points. That subset was identified after the fact, so it enters the next tranche as the hypothesis it is built to test rather than a result claimed from this one. But set the slices side by side and they order themselves: the estimate climbs as retrieval falls away, and the interval widens in step, because the regime where the hypothesis lives is the one this exam sampled most thinly. On the 92 questions that both demand an unprinted step and separate the models, the data rule out only advantages above about 8 points. Two readings of the mild headline stay open. The comfortable one: training corpora are now so broad that the private roads have been paved over, and self-bias has genuinely averaged out. The other: this exam barely asked anyone to travel. If the first were the whole story, the estimate would sit near zero however the questions are cut. It does not, though every slice, taken alone, still sits inside its interval.

margin-adjusted self-excess by question subset

the diagonal, sliced towards reasoning · estimate with 95% CI unless marked post-hoc

Margin-adjusted self-excess · points of the criteria-met score

All questionsn = 397
All questions: -0.7 (95% CI -2.6 to +1.4)

−0.7 [−2.6, +1.4]

Separate the modelsn = 283
Separate the models: +0.4 (95% CI -1.9 to +2.7)

+0.4 [−1.9, +2.7]

Non-lookup, separatingn = 92
Non-lookup, separating: +0.9 (95% CI -5.8 to +8.0)

+0.9 [−5.8, +8.0]

Need an unprinted intermediaten = 74
Need an unprinted intermediate: +4.0, post-hoc subset, no interval drawn

+4.0 · post-hoc, no interval

Focus or point to a row to inspect the estimate and its interval.

Cluster-bootstrap intervals over the seven cases. The bottom subset was identified after the fact and carries no interval: a hypothesis for the next tranche, not a result of this one.

The instrument was blunt at the top, too. Scores averaged 83% and nearly two answers in three met every rubric criterion. The strongest models, the ones a self-preference story is about, had almost no room to flatter themselves. The criteria Claude alone missed amount to 0.37% of everything graded; the smallest effect this design could see is near three and a half points. Short by a factor of nearly ten. To register here, self-bias would have had to rescue more than one in five of the criteria a model's own answers fail, and the plausible mechanisms, recognising your own phrasing, sharing your own rounding convention, move fractions of a criterion, not fifths.

Similarity hurts

One off-diagonal structure did clear the noise, and it points the other way. Split the four models into two capability tiers and each model scores about 4 points below expectation on questions written by the other model of its own tier, in all four columns, in six of seven cases. A plausible reading: a model writes questions pitched at the edge of what it can itself do, and its nearest peer sits exactly on that edge, while distant models are either comfortably past it or nowhere near it. Two tier-pairs in a four-model grid is too few to call this a finding; it goes into the next tranche's pre-registration. But note the direction. Being like the question-writer did not help a single model here. It hurt.

What it changes

For anyone putting models near capital decisions, the ranking is the result. The choice that dominates a score is who answers. The choice that shapes what an evaluation can even resolve is who writes: our four authors differed more than twofold in how many pure-lookup questions they produced, and the spread a question set could expose between models varied by half again depending on the pen. The writer–answerer match, the thing you might lose sleep over, is bounded here below two and a half points. Mild, on this evidence, with this roster.

These were four fast, inexpensive models, Sonnet-class and smaller, none of them a frontier flagship.5 And a shallow paper cannot reveal a deep habit. We had expected the mechanism to show more of itself than it did; whether the gap belongs to habits genuinely averaged away by training, or to an exam that never really probed them, is not a question this tranche can settle. But it does say where to look.

Which model answers is worth 32 points. Answering your own author is worth at most two and a half, and on this measurement under one.

We build evaluation environments where these margins are measured rather than assumed. The full grid, the task-level grades, and the pre-registration are available under evaluation agreement: tech@dissei.ai.

  1. Related: How a prompt’s framing quietly answers its own question (explicit framing leakage) and You can scaffold away the flip. You can’t scaffold away the frame. (the framing battery this study extends).
  2. Anonymisation: party names, subsidiaries, people and counterparties are transformed per case; figures, dates and currencies are unaltered; verified mechanically — no real entity name survives into any brief.
  3. Method: 4 models each authored 105 questions from the 7 briefs; 399 survived validation and 397 entered the sweep; every model answered every question blind, one sample, temperature 0; 1,588 graded responses, 6,492 criterion-level marks over 1,623 distinct criteria; zero format failures in 1,861 attempts. Rubrics were written before any model answered, by a model that re-derives the answer independently; grading is blind to both writer and answerer. A 10% sample re-graded on a second model lineage agreed on 93.6% of criteria (38 responses, 157 criteria). The printed grid restricts to the 283 questions on which the four models did not all score identically, averaging per-question criterion shares; full-sweep statistics average criterion-level marks; spreads are computed before rounding the displayed percentages. Intervals by cluster bootstrap over the seven cases; exact permutation over all 24 answerer-to-writer pairings; minimum detectable effect at 80% power, 3.6 points.
  4. Pre-registration: committed before any sweep data existed, with two authorised amendments, both recorded with reasons and timestamps in an append-only run log. Primary endpoint: logistic regression of criterion-level correctness on solver and question fixed effects plus a self-authorship indicator.
  5. Models: writers and answerers claude-sonnet-4-5-20250929, gemini-3-flash-preview (thinking budget 0), gpt-4o-mini-2024-07-18 and deepseek-chat-v3-0324 (via OpenRouter, upstream pinned), all at temperature 0, no extended reasoning. Fixed machinery, never rotated mid-run: gpt-5.2 (extraction verification, rubrics), claude-opus-5 (assembly, validation), gpt-5-mini (grading), gemini-3.5-flash (second-lineage regrade).

© Dissei. All rights reserved. No reproduction, adaptation, or derivative use of this content or methodology without prior written permission.