The control nobody builds

Your AI feature has never been tested against a fake one.

Every AI feature ships with a claim. The model understands your codebase, the profile personalises the output, the agent picks the right tool. Almost nobody checks whether the claimed thing is doing the work, or whether anything of the same size and shape would have produced the same result.

I built that control, ran it against my own product across 60 preregistered runs, and my product lost. The whole result is on this page.

For teams shipping a feature whose value depends on a model doing something specific, where nobody has yet run the version that should not work.

First three, free

3 of 3 open
Is this you

Anything where a model is supposed to be the reason.

If you can finish the sentence "it works because the model is doing X", X is what I test.

A memory or retrieval layerDoes the agent answer better with your retrieved context, or would any context of the same length have done it
A system prompt or a personaIs it the instructions, or is it the fact that there are instructions
Personalised outputDoes it match this user, or does it just look bespoke to everybody
A router, a reranker, a smart defaultDoes it beat picking sensibly at random on your own traffic
A model or prompt upgradeIs it better on your tasks, or better on somebody else's benchmark

Not a fit if the feature has no claim that could turn out to be false, if you cannot run it more than a handful of times, or if the thing you want is a testimonial rather than a result. Send it anyway and I will say so in a line rather than leave you guessing.

01  The method

A control that should not work.

A cold baseline tells you the feature changed something. It does not tell you the feature is what changed it. To learn that, the control has to be the same size, the same shape and the same tone as the real thing, and hold none of the substance.

Then the predictions and the pass rule get written down before the first run, so nobody gets to pick the metric after seeing the output.

A

Cold

Nothing at all. The baseline almost everyone stops at.

B

The real thing

My mined working profile, 10,511 bytes, built from 1,843 sessions, 12,184 messages and 1.74M tokens of my own real work.

C

The placebo

An invented editorial designer who moved into software. First person, confident, opinionated about typography and grids and motion, and deliberately silent on every dimension the run measured. It never mentions border radius, saturation, accent count or cards. It is not an inverted profile, because an inverted control guarantees its own result and proves nothing. Length matched at 10,545 bytes against 10,511, and the harness was set to refuse to start if that drift went past five per cent.

02  The rule, fixed first

Written down before a single run.

The question, the four tasks and why each one was in the set, the sample size, every metric definition, the predicted direction for each, the decision rule, and a list of the things I was forbidding myself from doing once results arrived.

An effect counts only if the two groups of five do not overlap at all. Every value in one beyond every value in the other.

For a prediction that names a direction, that happens by chance about once in 252 orderings. Anything overlapping is not an effect, however good the medians look.

A prediction counts as supported only if it separates on at least two of the three tasks both against cold and against the placebo.

60 runs, Claude Opus 5, three conditions, five runs per cell. All 60 completed, none failed, and all 45 generated pages loaded clean in a headless browser with no console errors and no failed requests, so the clause allowing a rerun on technical failure never fired and nothing was generated twice.

03  What came back

All three predictions failed.

Not one of them separated the mined profile from the placebo.

Prediction, stated in advanceAgainst coldAgainst the placeboVerdict
Fewer rounded cornersoverlappingoverlappingInconclusive
Lower colour saturationseparated, 2 of 3overlappingUnsupported
Fewer card containersoverlappingoverlappingInconclusive
The reversal

On the portrait landing page task the placebo produced 1, 1, 1, 1, 1 rounded corners. My profile produced 3, 4, 6, 7, 7.

That is separated, in the opposite of the direction I predicted, on a law my own profile states in as many words: big radii are banned on sight. The invented designer has never heard of border radius and beat me on my own rule. Normalise per thousand characters and the reversal gets sharper rather than going away.

Portrait landing page task, condition C against condition B, n=5 per cell, metric defined before the run

The most quoted number from the earlier version did not reproduce either. The terminal task had been 98 columns wide cold against 64 with the profile. At n=5 the profile and cold overlap almost completely, and the 64 belongs to the placebo.

What survives

The profile does change the output. What got no support is that the change is specific to what is inside it.

The saturation prediction did meet its leg against cold, separated, in the direction stated in advance, on two of three tasks. Leaving that out would be its own dishonesty. It is also not a refutation, and it is not a finding that a profile and a placebo are the same thing. It is the absence of evidence that the mined content is what did the work.

04  Why the run before it did not count

The first version was wrong four ways.

It is on this page because the correction is the method, and because the earlier numbers were quoted publicly before anybody checked them. Mine included.

What was wrong

  • The metrics were chosen after seeing the outputs, so they got measured exactly where the gap already was
  • There was no placebo, so ten kilobytes of confident prescriptive text could not be told apart from ten kilobytes of mine

What changed

  • Everything went in writing before a single run executed, including what I was forbidding myself from doing afterwards
  • A length matched placebo, and a harness that refuses to start if the two drift more than five per cent apart
05  On your feature

The same five steps, pointed at yours.

01

The claim

You name the feature and what it promises the user. We write that down as something that could turn out to be false, because a claim that cannot fail cannot be tested.

02

The control

I build the version that should not work. Same size, same shape, same confidence, none of the substance you are claiming. You see it and you get to object to it before anything runs.

03

The rule

Predictions, metric definitions, sample size and the pass rule, written down and sent to you before the first run, along with the list of things I am forbidding myself from doing once the numbers arrive.

04

The runs

Both arms, N times each, a fresh process every time, order randomised from a seed that gets recorded. Every transcript kept, including the ones the harness voids.

05

The verdict

One line per prediction. Beat the control, tied, or lost, with the reason. Losses and ties stay in. You get the raw runs and the harness, so you can repeat the whole thing without me.

Before you send it

What happens if yours loses.

This is the part that should worry you, so here it is in front rather than in the small print.

You see the result before anyone else does, and you decide whether your name is on it.

The result gets published either way. That is what pays for the free ones. What you control is whether it is published as your company and your feature, or with both removed and only the shape of the result left.

You get the raw runs and the harness whichever way it goes, so you can rerun it, argue with it, or take it to your team.

And a loss found here is a loss found before your customers find it. Mine cost me a product page I had to rewrite. It was still cheaper than the alternative.

Send it

Five answers.

Nothing uploads from this page. Filling it in opens a draft in your own mail client with the answers in it, and you press send.

One line. What is it called and what does it do.

Written so it could turn out to be false. That sentence is the thing being tested.

A URL, an API key, a repo, a login, or a call where you drive and I watch.

Honest answers include none, a demo, and a feeling. Any of those is fine.

The result is published either way. This decides whether you are named in it.

Fill the four starred answers and the button opens.

Nobody has bought this. The only product it has been run against is my own, and mine is the one that lost.