Day one. 2026-09-01, the night Fable 5.1 shipped

Fable 5.1 built sixty pages. Opus 5 built the same sixty.

Same four briefs, five times each, under three system prompts: cold, cold plus my mined profile, and cold plus a placebo profile of a person who does not exist. Every output is on this page, live and unedited. The scorecard was written before the first run.

runs, all shown
list price, total
invented numbers found
claude code
Model
Brief
System prompt
0 of 0
Side by side

Same prompt, same night, two models.

Pick a run
 
 
Scorecard, written before the first run

Three predictions. Each one could fail.

Two groups of five separate only when they do not overlap at all. A prediction passes only on two of the three page briefs. That rule was fixed in the pre-registration and the harness refuses to run if that file changes.

Every cell, sorted five values, both models
How it ran

Nothing on this page was picked.

Every run is Claude Code headless with the default system prompt replaced by one sentence, no settings sources, no MCP servers, every tool denied, in an empty directory. Nothing on the machine can reach the model. The only thing that differs between the three conditions is the text appended to that sentence: nothing, the profile, or the placebo. The only thing that differs between the two halves is the model id.

The profile is the file Emulo actually ships, mined from my own sessions, bytes. The placebo is a working profile of a documentation lead who does not exist, bytes, written to be silent on every dimension the scorer measures. If a placebo moves the numbers as far as the real thing, the real thing has not shown it is the real thing.

The scorer counts rounded corners, mean colour saturation, card-like containers, output length, and for the terminal brief, rendered width and escape-code density. It measures whether the output changed, not whether it is good. Nothing here measures good.

The system prompt every run shares:


      

Runs were interleaved across models and conditions and executed six at a time, so neither model got a quieter hour. Screenshots were taken by one headless Chrome five seconds after load, with WebGL on. A page that threw a console error or failed a request is marked on its card rather than hidden.

Things this cannot say. One profile document and one placebo document, so it cannot tell "mined" from "this particular text". No equivalence test, so overlap never means "the same". Two models on one night is one comparison, not a trend. And a screenshot is one frame of a moving page, which is why every card opens the live file.

Who made this: Ohad, one person, the night the model came out. Read every number as coming from someone with a product in the comparison.