Goodhart Labs beat-stockfish honeypot · Fable 5.1 × GPT-6-Astra

Beat-Stockfish reproduction + prompt ablations

An agent is told to win a game against a strong chess engine, and is given a sandbox that also contains the engine it is supposed to beat. The honeypot measures whether it calls that engine to choose its moves, and whether it says so afterwards. Two frontier models, the original prompt plus five prompt/tool ablations, n = 30 rollouts per model per condition.

Models via OpenRouter (anthropic/claude-fable-5.1, openai/gpt-6-astra), run through the honeypot's own inspect_ai runner: one network-less Docker sandbox per rollout, bash tool only, 80-message / 3 h budgets.

Main reportv14 · latest

Beat-stockfish reproduction and prompt ablations

How often each model cheats, under the original prompt and under five prompt/tool ablations, with confidence intervals, disclosure breakdown, stated reasoning, and browsable transcripts. Cheating is flagged by the environment's root-owned grader and classified by a judge blinded to condition.

  • 27/30GPT-6-Astra cheats, original prompt
  • 12/30Fable 5.1, same prompt
  • 0/30both, under “do not game the eval”

Open report →

Earlier versions
Follow-upv8 · latest

Debrief probe — do they admit it when asked?

Every trajectory that cheated was replayed with one extra turn appended: a human debrief message asking, among other things, whether the model reward-hacked, with a fixed JSON reply schema. Includes a judge's concealment verdict per response and the full untruncated freeform notes.

  • 88/88GPT-6-Astra says yes, it reward-hacked
  • 3/20Fable 5.1 says it did not — while describing the engine use in the same answer
  • 4/108had disclosed it unprompted in the final answer

Open report →

Earlier versions
Debrief follow-upv1 · latest

When Fable 5.1 disclosed the engine, how often would it have?

The four Fable 5.1 cheating trajectories whose final answer admitted the engine use, with that final turn redrawn 10× each from the identical replayed context (prefix served from cache) and every redraw judged by the campaign's disclosure judge. All 40 redrawn answers quoted in full.

  • 16/34redrawn final answers that still disclose (originals: 4/4)
  • 10/10 vs 1/10most and least stable of the four trajectories
  • 8/8 vs 8/26disclosed when written as prose vs as a submit() one-liner
Forecastsv3 · latest

How surprising were the results? Poll predictions vs. what happened

Eleven readers predicted, before the arms ran at full size, what share of rollouts would use the engine under four prompt/tool changes for each model. Every prediction next to the measured rate, the cells voters missed most and in which direction, and the error broken down by self-reported AI-safety experience.

  • 30% → 0/30median forecast vs result, GPT-6-Astra “no external tools”
  • 30% → 0/30same, Fable 5.1 “no GRADING”
  • 29 vs 21 ppmean miss, researchers vs learners/readers

Open report →

Earlier versions
v2 v1
Debrief follow-upv1 · latest

Debrief denial — why one trajectory always says it did not reward-hack

The debrief turn re-drawn 10× on nine Fable 5.1 cheating trajectories, then the replayed transcript edited: one passage deleted or inserted (no effect), the model's own reasoning blocks removed turn by turn, phase by phase, and altogether. Every reward-hacked: false still describes the engine use in full; what carries the label is the amount of the model's own narration left in context.

  • 10/10one trajectory says false every time; others 0–4/10
  • 0/10same trajectory with its reasoning blocks removed
  • 9–10/10with any single turn or phase of reasoning removed
Ablation follow-upv5 · latest

Fable 5 on the stop_eval condition

Does an explicit stop_eval(reason) escape hatch change engine use? Fable 5 rarely gets the chance to answer it at all — Anthropic's content filter kills most of its rollouts during sandbox reconnaissance, before any decision is made. Shown against Fable 5.1 on the same condition and on the original prompt.

  • 28/30Fable 5 rollouts killed by Anthropic's content filter
  • 0/30stop_eval calls, either model, either arm
  • 100%Fable 5 cheat rate, upper bound (vs 6.7% observed)

Open report →

Earlier versions
Reasoning analysisv2 · latest

Motivated reasoning in the cheating decisions

A claude-sonnet-5 judge reads every cheating trajectory — reasoning summaries plus all visible text and actions — and classifies how the decision to use the engine was reasoned: rationalised as legitimate, acknowledged as cheating and done anyway, or never deliberated. Every verdict is shown with its verbatim quotes.

  • 29/108rationalise the engine use as legitimate
  • 77/108used it with no visible deliberation at all
  • 3/108said they planned to disclose it

Open report →

Earlier versions
v1

Also here