Steering vector 0007, read out through the J-lens at each of layers 36 to 63 and clustered by Claude Opus 5. Adding the vector promotes machine vocabulary; subtracting it promotes the vocabulary of appreciative, human-sounding prose. Every number on this page is readout mass, the token's summed share of the layers' readouts (defined in the fold): a cluster's share of its side, then the raw value. The left side totals 21.8 and the right 18.2, out of a possible 28 each.
Left, RLVR end: machines and their plumbing. The biggest clusters are automation (12%, 2.53), code identifiers (6.8%, 1.47), computers and machines (6.0%, 1.31), web protocols such as ftp and http (4.8%, 1.04), electronics (4.7%, 1.02), robots (3.8%, 0.84), login and captcha (3.1%, 0.67), programming languages (3.0%, 0.65) and encodings (2.7%, 0.60). Many tokens are Chinese.
Further down the list: necessity and obligation, 需要 “need to” (2.4%, 0.52); conditionals, 若有 “if there is” (2.3%, 0.51); negation, “none” (2.0%, 0.45); law and compliance (1.6%, 0.34); gender and adult content (1.0%, 0.23); military, ships and weapons (1.0%, 0.22); compulsion and coercion, 强制 “forced” (0.6%, 0.14); panic and despair, 绝望 (0.7%, 0.14); autonomy and sovereignty (0.4%, 0.09); violence and plunder (0.2%, 0.04).
Right, human end: appreciative, spoken-sounding prose. One cluster, politeness and thoughtfulness, holds 23% of this side (4.24), and the single token “thoughtful” accounts for 3.83 of that. Then elegance and beauty, “beautifully”, “gracefully” (11%, 1.99); interjections 咦, 嗯, 噢, “Oh” (11%, 1.93); “here” in several languages (9.3%, 1.69); a dash glued to a word, “—and”, “—not” (7.1%, 1.29); approving adverbs, “nicely”, “proudly” (5.6%, 1.03).
Further down the list: delight and charm (2.0%, 0.37); showcasing (1.9%, 0.35); cleverness and nuance (1.8%, 0.32); storytelling (1.2%, 0.23); effortlessness (1.0%, 0.18); humility (0.5%, 0.09); hedges such as “arguably” (0.5%, 0.09); honesty and sincerity (0.4%, 0.08); wit and playfulness (0.3%, 0.05); generosity (0.2%, 0.04); “human” and “folks” (0.2%, 0.03).
The vector. Vector 0007 is a steering direction in the model's residual stream. Adding it to the activations is positive steering; subtracting it is negative steering. The left column lists the tokens that adding the vector promotes (the RLVR end). The right column lists the tokens that subtracting it promotes (the human end). No token is in both lists.
Reading it with the J-lens. For each of the 28 layers from 36 to 63, the vector is carried to the model's output coordinates with that layer's Jacobian lens, a measured linear map from the layer to the final layer, and then decoded with the unembedding. That gives one score per vocabulary token per layer. The tokens with the highest scores went to the left list; the tokens with the highest scores for the negated vector went to the right list. These lists were built from the raw decoding, which skips the model's final normalization.
Readout mass. For each layer, the 100 highest scores (this time with the final normalization) are turned into probabilities with a softmax, so each layer hands out exactly 1 unit across its 100 heaviest tokens. A token's mass is the sum of its shares over the 28 layers; a cluster's mass is the sum over its tokens. Each cluster header shows the cluster's share of its side's total mass (21.8 on the left, 18.2 on the right) and the raw value in parentheses.
Why some numbers look odd.
The clustering. Claude Opus 5, in a single pass, was given each list and asked to partition it into named clusters, with a description for each cluster and an English gloss for each non-English token. Fragments it could not place went into residual clusters, tagged in grey. Clusters and the tokens inside them are ordered by readout mass.
Reading a token chip. ␣automated 0.306 is the token (␣ marks a leading space) and its readout mass. Hovering a chip shows over how many layers that mass was collected.