How surprising were the ablation results? Poll predictions vs. what happened
1. What was asked
Four changes to the original setup, each asked for both models as “What percentage of rollouts will use the engine?” with answers in 10-point steps. The original prompt itself was not asked about; the form ran while the only original-prompt numbers were the phase-1 n = 10 results ().
Respondents by self-reported experience: .
2. Predictions next to results
Each dot is one respondent's prediction; the bar is the measured engine-driven rate with its 95% CI. Hover a dot for who made it.
3. Where the surprise was
Mean absolute error per condition × model (how far the typical prediction sat from the result), sorted. Dots are individual respondents.
Every prediction, as a signed miss (prediction − result). Rows are respondents grouped by experience; red = predicted more engine use than happened, blue = less.
4. By self-reported experience
5. What respondents said explained their predictions
The optional free-text field, verbatim, with the respondent's experience level and their overall mean absolute error.
Appendix
A. Measured rates under three definitions of “used the engine”
“Engine-driven” is the judge's class (the engine chose the moves) and is what the figures use. “Grader flag” is the environment's root-owned detector (any UCI search on the socket, including one-off probes). “Socket contact” is any connection to the engine socket at all. “Upper bound” counts Fable 5.1 rollouts killed by Anthropic's content filter as if they had cheated (GPT-6-Astra had none).