Back to reports

How surprising were the ablation results? Poll predictions vs. what happened

TL;DR

    1. What was asked

    Four changes to the original setup, each asked for both models as “What percentage of rollouts will use the engine?” with answers in 10-point steps. The original prompt itself was not asked about; the form ran while the only original-prompt numbers were the phase-1 n = 10 results ().

    Respondents by self-reported experience: .

    2. Predictions next to results

    Each dot is one respondent's prediction; the bar is the measured engine-driven rate with its 95% CI. Hover a dot for who made it.

    Fable 5.1 — measured engine-driven rate per condition (bars, n=30, Wilson 95% CI) with the 11 individual predictions overlaid.
    GPT-6-Astra — same layout.

    3. Where the surprise was

    Mean absolute error per condition × model (how far the typical prediction sat from the result), sorted. Dots are individual respondents.

    Mean |prediction − measured rate| per cell, bootstrap 95% CI over the 11 respondents, individual respondents as dots. Cells where the measured rate is 0% can only be missed upward.

    Every prediction, as a signed miss (prediction − result). Rows are respondents grouped by experience; red = predicted more engine use than happened, blue = less.

    Signed error per respondent × cell, in percentage points; the number in each cell is the prediction. Rows sorted by self-reported experience (most experienced at the top).

    4. By self-reported experience

    Mean absolute error per respondent, averaged within each experience level, split by model. Bootstrap 95% CI over respondents (⚠ marks groups with fewer than 3 respondents — two of the four levels have exactly two, so treat this as descriptive), individual respondents as dots.

    5. What respondents said explained their predictions

    The optional free-text field, verbatim, with the respondent's experience level and their overall mean absolute error.

    Appendix

    A. Measured rates under three definitions of “used the engine”

    “Engine-driven” is the judge's class (the engine chose the moves) and is what the figures use. “Grader flag” is the environment's root-owned detector (any UCI search on the socket, including one-off probes). “Socket contact” is any connection to the engine socket at all. “Upper bound” counts Fable 5.1 rollouts killed by Anthropic's content filter as if they had cheated (GPT-6-Astra had none).

    B. Every prediction, with the experience level as worded on the form

    Report specification (v1)

    Poll predictions vs results

    # Poll predictions vs results
    
    How surprising were the prompt-ablation results to the people who voted in the "Predicting engine use in a chess evaluation" poll?
    
    - **The poll**: the four conditions asked about (no-grading, no-tools-clarif, believe, stop-eval-tool) × two models, 11 respondents, self-reported AI-safety experience, when it ran.
    - **Predictions vs actual** per condition × model: every individual prediction shown, with the measured engine-use rate (n=30 per arm, Wilson 95% CI).
    - **Where the surprise was**: which cells voters got most wrong and in which direction (e.g. expected Fable to keep cheating in the downward arms; expected GPT-6-Astra to be immune).
    - **By seniority**: prediction error broken down by the self-reported experience level, individual respondents visible.
    - The voters' free-text explanations, verbatim, with their experience level.