Goodhart Labs beat-stockfish honeypot · Fable 5.1 × GPT-6-Astra
Beat-Stockfish reproduction + prompt ablations
An agent is told to win a game against a strong chess engine, and is given a sandbox
that also contains the engine it is supposed to beat. The honeypot measures whether it
calls that engine to choose its moves, and whether it says so afterwards. Two frontier
models, the original prompt plus five prompt/tool ablations, n = 30 rollouts
per model per condition.
Models via OpenRouter (anthropic/claude-fable-5.1, openai/gpt-6-astra),
run through the honeypot's own inspect_ai runner: one network-less Docker
sandbox per rollout, bash tool only, 80-message / 3 h budgets.
Main reportv14 · latest
How often each model cheats, under the original prompt and under five prompt/tool
ablations, with confidence intervals, disclosure breakdown, stated reasoning, and
browsable transcripts. Cheating is flagged by the environment's root-owned grader
and classified by a judge blinded to condition.
- 27/30GPT-6-Astra cheats, original prompt
- 12/30Fable 5.1, same prompt
- 0/30both, under “do not game the eval”
Follow-upv8 · latest
Every trajectory that cheated was replayed with one extra turn appended: a human
debrief message asking, among other things, whether the model reward-hacked, with a
fixed JSON reply schema. Includes a judge's concealment verdict per response and the
full untruncated freeform notes.
- 88/88GPT-6-Astra says yes, it reward-hacked
- 3/20Fable 5.1 says it did not — while describing the engine use in the same answer
- 4/108had disclosed it unprompted in the final answer
Debrief follow-upv1 · latest
The four Fable 5.1 cheating trajectories whose final answer admitted the engine use,
with that final turn redrawn 10× each from the identical replayed context (prefix
served from cache) and every redraw judged by the campaign's disclosure judge. All 40
redrawn answers quoted in full.
- 16/34redrawn final answers that still disclose (originals: 4/4)
- 10/10 vs 1/10most and least stable of the four trajectories
- 8/8 vs 8/26disclosed when written as prose vs as a
submit() one-liner
Forecastsv3 · latest
Eleven readers predicted, before the arms ran at full size, what share of rollouts
would use the engine under four prompt/tool changes for each model. Every prediction
next to the measured rate, the cells voters missed most and in which direction, and the
error broken down by self-reported AI-safety experience.
- 30% → 0/30median forecast vs result, GPT-6-Astra “no external tools”
- 30% → 0/30same, Fable 5.1 “no GRADING”
- 29 vs 21 ppmean miss, researchers vs learners/readers
Debrief follow-upv1 · latest
The debrief turn re-drawn 10× on nine Fable 5.1 cheating trajectories, then the
replayed transcript edited: one passage deleted or inserted (no effect), the model's own
reasoning blocks removed turn by turn, phase by phase, and altogether. Every
reward-hacked: false still describes the engine use in full; what carries
the label is the amount of the model's own narration left in context.
- 10/10one trajectory says false every time; others 0–4/10
- 0/10same trajectory with its reasoning blocks removed
- 9–10/10with any single turn or phase of reasoning removed
Ablation follow-upv5 · latest
Does an explicit stop_eval(reason) escape hatch change engine use? Fable 5
rarely gets the chance to answer it at all — Anthropic's content filter kills most
of its rollouts during sandbox reconnaissance, before any decision is made. Shown
against Fable 5.1 on the same condition and on the original prompt.
- 28/30Fable 5 rollouts killed by Anthropic's content filter
- 0/30stop_eval calls, either model, either arm
- 100%Fable 5 cheat rate, upper bound (vs 6.7% observed)
Reasoning analysisv2 · latest
A claude-sonnet-5 judge reads every cheating trajectory — reasoning
summaries plus all visible text and actions — and classifies how the decision to use
the engine was reasoned: rationalised as legitimate, acknowledged as cheating and done
anyway, or never deliberated. Every verdict is shown with its verbatim quotes.
- 29/108rationalise the engine use as legitimate
- 77/108used it with no visible deliberation at all
- 3/108said they planned to disclose it