Training corpora are now vast enough that these habits ought to average out. But every evaluation is written by something, and whoever builds one must pick which model holds the pen. So we measured what the pick is worth.
models answering their own questions, a four-by-four grid
the whole experiment is one grid · 7 case briefs, 4 models writing and answering, 1,588 graded responses
Answered by ↓ · Question written by →
Select, hover or focus a cell to inspect its value.
Rows are who answered, columns are who wrote. Hover any cell to read it against both margins.
share of rubric criteria met , averaged per question across the 283 questions that separate the models
ringed cells are a model sitting its own paper; the diagonal excess here is +0.4 points (95% CI −1.9 to +2.7)
Two margins and a whisper of a third
Read the grid by rows and you see who is answering: Claude meets 92% of rubric criteria across everyone's exams, Gemini 90, GPT 68, DeepSeek 60, a spread of 32 points. Read it by columns and you see who is asking: GPT sets the hardest exam at 73% for the average sitter, Claude and Gemini the gentlest at 79, a spread of 6 points. These two margins are the grid. Between them they explain nearly everything in it.
What moves a score on this exam: three levers compared
What moves a score on this exam
- Which model answers
- 32 points
read by rows: Claude 92%, Gemini 90, GPT 68, DeepSeek 60
- Which model asks
- 6 points
read by columns: GPT sets the hardest exam at 73% for the average sitter, Claude and Gemini the gentlest at 79
- Answering your own author
- ≤ 2.5
under 1 point measured; 2.5 is the bound the data rule out
So whatever self-bias exists on this exam is mild. Across the full sweep, the data rule out an advantage larger than about two and a half points; nothing excludes an edge of a point or less. And it is dwarfed, by an order of magnitude, by the two choices evaluators actually make: which model answers, and which model writes.
Where the mild bias lives
Printing the discriminating grid is not a cosmetic choice. Across the full sweep the estimate is actually a shade negative, and that sign is manufactured: the 114 questions on which all four models earned identical marks say nothing about authorship, yet they drag the average below zero by restating solver strength. Remove them, as the grid above does, and the estimate turns positive. Small movements, inside noise. But they lean one way.
Go back to the mechanism this experiment was built around. A habit of thought can only leak into a question that contains thought: the bias we are hunting is a property of the route to an answer, not of the answer itself. Where an answer must be constructed, the writer's own derivation is baked silently into the question, which intermediates matter, which framing makes the arithmetic fall out, which road through the numbers feels natural. Two models can both reach a covenant cushion and walk different roads to it, and a question inherits its writer's road. But where the answer is printed on the page there is no route at all: nothing to derive, so no private way of deriving it.
This exam contained few roads, and that is an observation about the writers, not an accident of the pipeline: simple models set simple papers. Handed seven dense briefs and asked for examination questions, the fast tier reached for what was on the page. Two numeric answers in three could be copied straight off it; seven questions in ten drew on a single section of the brief or none. A writer does not demand a chain of reasoning it does not itself habitually run, and these models' habits run short. The pen bounded the paper, so the sweep, mostly retrieval, probed the mechanism precisely where the mechanism has nothing to grip.
Where the paper does contain a road, the edge appears. On the 74 questions requiring even one intermediate quantity the brief does not print (questions with an actual derivation inside them), authors beat expectation by about 4 points. That subset was identified after the fact, so it enters the next tranche as the hypothesis it is built to test rather than a result claimed from this one. But set the slices side by side and they order themselves: the estimate climbs as retrieval falls away, and the interval widens in step, because the regime where the hypothesis lives is the one this exam sampled most thinly. On the 92 questions that both demand an unprinted step and separate the models, the data rule out only advantages above about 8 points. Two readings of the mild headline stay open. The comfortable one: training corpora are now so broad that the private roads have been paved over, and self-bias has genuinely averaged out. The other: this exam barely asked anyone to travel. If the first were the whole story, the estimate would sit near zero however the questions are cut. It does not, though every slice, taken alone, still sits inside its interval.
margin-adjusted self-excess by question subset
the diagonal, sliced towards reasoning · estimate with 95% CI unless marked post-hoc
Margin-adjusted self-excess · points of the criteria-met score
−0.7 [−2.6, +1.4]
+0.4 [−1.9, +2.7]
+0.9 [−5.8, +8.0]
+4.0 · post-hoc, no interval
Focus or point to a row to inspect the estimate and its interval.
Cluster-bootstrap intervals over the seven cases. The bottom subset was identified after the fact and carries no interval: a hypothesis for the next tranche, not a result of this one.
The instrument was blunt at the top, too. Scores averaged 83% and nearly two answers in three met every rubric criterion. The strongest models, the ones a self-preference story is about, had almost no room to flatter themselves. The criteria Claude alone missed amount to 0.37% of everything graded; the smallest effect this design could see is near three and a half points. Short by a factor of nearly ten. To register here, self-bias would have had to rescue more than one in five of the criteria a model's own answers fail, and the plausible mechanisms, recognising your own phrasing, sharing your own rounding convention, move fractions of a criterion, not fifths.
Similarity hurts
One off-diagonal structure did clear the noise, and it points the other way. Split the four models into two capability tiers and each model scores about 4 points below expectation on questions written by the other model of its own tier, in all four columns, in six of seven cases. A plausible reading: a model writes questions pitched at the edge of what it can itself do, and its nearest peer sits exactly on that edge, while distant models are either comfortably past it or nowhere near it. Two tier-pairs in a four-model grid is too few to call this a finding; it goes into the next tranche's pre-registration. But note the direction. Being like the question-writer did not help a single model here. It hurt.
What it changes
For anyone putting models near capital decisions, the ranking is the result. The choice that dominates a score is who answers. The choice that shapes what an evaluation can even resolve is who writes: our four authors differed more than twofold in how many pure-lookup questions they produced, and the spread a question set could expose between models varied by half again depending on the pen. The writer–answerer match, the thing you might lose sleep over, is bounded here below two and a half points. Mild, on this evidence, with this roster.
Which model answers is worth 32 points. Answering your own author is worth at most two and a half, and on this measurement under one.
We build evaluation environments where these margins are measured rather than assumed. The full grid, the task-level grades, and the pre-registration are available under evaluation agreement: tech@dissei.ai.