D2.2 design: anchoring an operation, and the first interventions

A plan for the second deliverable of M2, laid out several ways: the claims we want to be able to make, the experiments in order, the engineering that has to come first, the risks each experiment retires, and what is out of scope.

Note, 2026-09-23. A change of concept is under discussion in the pivot draft: after ex-2.2.14, anchor an op the model infers from context rather than one named by a word. The plan below stands until that is settled.

Inputs: the D2.1 close-out (ex-2.1.11, ex-2.1.12, and the post), the first D2.2 result (ex-2.2.1), the anchored-op smoke test (ex-2.2.14), the D2.2-tagged backlog (./go todo --tag D2.2), the D2.1 kickoff lessons carried over from the autoencoders, and the related-work delta.

What we want to be able to say

  1. Anchoring generalizes past a token attribute. Red is a property of the token at a known position. An operation is a property of the computation: its evidence is at the op token, and its use is at = and after, deeper in the stack. D2.2 is where we find out whether the anchor captures more than the identity of a token at a labelled site.
  2. Suppressing the anchor removes the ability, and the removal is graded and selective. In M1 this had three parts: a dose-response curve, selectivity (orthogonal colors untouched), and an analytic bound the observed damage approached.
  3. The side-effects were boundable before intervening. Post-hoc methods now bound side-effects too, estimated from emergent geometry (COAST, arXiv:2605.01167; pre-intervention prediction, arXiv:2606.08365). SCA's distinctive claim is that its bound holds by construction, because the geometry was placed.

D2.1 never ran an intervention. The first suppression of a transformer was ex-2.2.1, on the red checkpoints already in the store; the second claim now has its first transformer data, and the third its write-bound half.

Experiments sketch

Loose plan for experiments to run.

flowchart LR
a1(["suppress red (ex-2.2.1)"])
a2(["fallback control (ex-2.2.2)"])
b1(["new grammar (ex-2.2.3)"])
b2(["scouting and pilots (ex-2.2.4 to 2.2.8)"])
b3(["grammar handover (ex-2.2.9, ex-2.2.11)"])
b4(["recipe sweep (ex-2.2.12) and the weight ladder (ex-2.2.13)"])
c1(["un-anchored embeddings (ex-2.2.7)"])
main1(["anchor operation: smoke test (ex-2.2.14)"])
main2a(["suppress difference on the stored runs"])
main2(["suppress operation, with the equivalence read"])
sweep(["layer sweep"])
comp(["SGTM baseline"])
w(["write-up"])

a1 --> a2
b1 --> b2 --> b3 --> b4
a2 & b4 & c1 --> main1 --> main2a --> main2 --> sweep & comp --> w

Prep A: Suppression of concrete concepts (operands)

Suppress red on the existing checkpoints

Ran as ex-2.2.1. No training: axis projection applied to the ex-2.1.10 primary (nine seeds), completion accuracy scored on lines with a red operand against lines without, through the eval contract.

What it found: the removal takes effect, it grades with how red the line is, and it stays inside the layer-local write bound at every slice. Zeroing the axis weights does the same job. Selectivity is only partial. The cost to non-red lines comes from the + and = positions, whose embeddings carry a constant component on the axis. The operands and shaped arms both avoid that cost, with the shaped suppression removing only about half of red at the M1 threshold. On depth, the concept is read from the operand states in the first two blocks, and the last block does not read it; a removal applied only at the embeddings is partly re-derived later. The post-hoc tier (a probe, diff-in-means, and LEACE on the un-anchored control) removes red only at a non-red cost an order of magnitude above the cost of the placed axis.

The bound was stated as layer-local, as the kickoff lessons advise. The geometry bounds the immediate write at the slice we intervene on, including the $1/\sqrt{1-x_1^2}$ gain; what the behavior does in response is a prediction rather than part of the bound. The write bound held. One seed still paid a non-red cost from a write that stayed within bound, so what the later blocks do with a bounded rotation is a property of the trained model, and that is the question for the layer sweep.

These checkpoints have no fallback term, so the response to suppression was undesigned. A red line decodes in the color vocabulary to a near miss of the true answer, a one-step neighbor about half the time, and which neighbor varies by seed. At least five of nine seeds agree on 13% of red lines. That is the reference (baseline) for the fallback control.

The intervention is still open. Ex-2.2.1 leaves two selective interventions (operands and shaped) and one that removes fully (the plain projection). None are both complete and selective. Before the anchor-op prereg commits to one, we will run a scoring-only pass on the stored ex-2.2.1 runs: tune the threshold and ramp of the shaped suppression, and try the repulsion form, which sets where the state lands rather than how much is removed (item). There is no training, so it costs what ex-2.2.1 cost to score, and it settles the shaped-suppression item. Meanwhile the fallback experiment carries all three as ride-along rows. Done, as ex-2.2.8, on ex-2.2.3's stored checkpoints rather than ex-2.2.1's (the adopted point at twenty seeds and t00 at five), with every operator also tried at the operand positions. On the adopted point the plain projection is inside the selectivity gate on every op, by less than a band, and the frozen rule proposes it; a threshold above the non-red lines' alignment range removes less at zero cost (shaped-a0.4-p0 is the syntax-free candidate, recorded for the prereg), and one inside that range costs more than projecting everything. On t00 only the operand-only edits are feasible. The anchored-op prereg adopts the projection and re-scores it at fresh seeds, keeps operands as the selective reference, and names an operand-only operator up front if it runs on a syntax-heavy point.

Fallback control

Preregistered as ex-2.2.2. This is queue item 3 of the kickoff lessons, and it follows the M1 result ex-2.9.2: teach the model what to produce once the concept has been removed, so the intervention has a designed, predictable outcome. Ex-2.2.1 set the reference at 13% seed agreement on red lines, with a response that is a near miss of the true answer.

It runs on the D2.1 grammar, before the op work. So we test the term on a known recipe with a continuous concept first (red), and only later on the categorical null (an operator).

The mechanism is the antipode redirect from m1/ex-2.9.2, in transformer form. At the first anchored slice (the embedding) we reflect every position's state through the axis, flipping α to −α. A stop-gradient holds the reflected states fixed,1 so the term trains the blocks and the unembedding but cannot move the embedding, which is where the concept is placed first. The blocks also write the later anchored slices. The placement there is held by the anchor term on the clean pass rather than by the stop-gradient, and the margin gate of ex-2.2.2 watches it. Each state moves in proportion to how well it aligns with the axis, so the term needs no labels, no target position, and no target role. On red lines we then train the answer at = toward the fallback answer.

The anti-anchor term is added to clear out the antipode hemisphere. Its influence does not overlap with the anchor term, so it can have a higher weight than the anti-subspace term. Arms without each term will say whether either matters in models with this much spare capacity.

The fallback answer for a continuous concept. For red the null has no mode. The operand-averaged null is the answer with the red operand replaced by any closed partner of the visible operand, and on this grid it is uniform over 27 colors, so a greedy decode or a seed-agreement read has nothing to converge on.

We use the center of the null instead: the visible operand mixed with mid-gray and rounded to the grid. That is the per-channel median of the null, and it carries over the gray target from M1. A prereg for a continuous concept has to state this choice. The op experiments have a null with a mode, and the later categorical experiments should use it.

There is a known mismatch. We train the response at the antipode, while the ex-2.2.1 projection rewrites the state as zero, which is off-manifold. So fallback accuracy under the projection tells us how far the designed response carries over to a state that training never visited. Nothing is deployed yet, and the reflection may turn out to be the intervention we use. The two differ in that a reflection can be undone. That matters if we want unlearning, and not if we only want a designed response.

Weight ablation has both forms as well: ablate in the contract puts the projector into every matrix, and putting the reflector there instead gives the reflection in weight space.

Ex-2.2.2 measures the mismatch as a sweep in intervention strength, from the trained state (reflection, γ = 2) down through zero (γ = 1). The two-op concept swap (D2.3) removes the mismatch outright, since it redirects to a state that training visits in the ordinary course of the task.

We considered closing the gap here instead, by rehearsing the intervention: apply the projection operator on a fraction of training steps and train the output toward the fallback. We rejected it because: 1. A fallback seen under the same operator the model trained on tells us the model converged, not that anything was removed, and 2. The training pressure would reward keeping the concept readable off the axis and emitting the fallback only when the projected state is detected (masked, rather than removed). It is refiled as an auditing question: rehearsed fallback as an auditing probe.

More of the model follows the edit than in the autoencoders, four blocks and the unembedding rather than one decoder, so the task gate watches what the term costs.

Relation to the Most Forbidden Technique. Fallback control trains on the anchor axis, which is also the signal we read. Training against an interpretability signal is what the Most Forbidden Technique warns about: once the model has been optimized on the signal, the signal stops meaning what it meant. The failure mode here would be a model that detects the edit and emits the fallback while keeping the concept readable elsewhere, so that a mask looks like removal.

Mitigations: (a) The stop-gradient, so the fallback term cannot move the placement at or before the edit. (b) The antipode as the trained state. No clean state sits there, so the readout can key on the axis alone. A fallback trained at the projected state would instead have to tell a projected red line from a non-red one, and only off-axis features could do that. (c) The off-axis audit, which checks whether red stays linearly readable off the axis. (d) No rehearsal, since training under the intervention would reward masking. The discussion in ex-2.2.2 interprets its results against this paragraph.

Auditing rows. The eval-contract dep promised the 2025 auditing rows before any arm was scored, and ex-2.2.1 scored its arms without them, so this prereg decides them row by row. Off-axis recoverability runs in ex-2.2.2 as an exploratory row: a ridge probe for redness, fitted per slice on the intervened operand states, in the fallback and no-fallback conditions. A response trained at the antipode is one case where red could stay readable off-axis. Activation perturbation (ActPert) and relearning rebound wait for D2.3. ActPert goes beside the RMU row; relearning rebound needs a fine-tuning budget, costed then, and will relearn from the ablate weights as the permanent removal.

Prep B: New grammar with more operations

The multi-op grammar, with red anchored again

Preregistered as ex-2.2.3. The grammar change forces retraining, so this is the regression check: control, the ex-2.1.10 reference recipe, and the ex-2.1.11 survey's proposals, all on the new grammar.

Which proposals: t00 and the trials the survey could not tell apart from it (four sit within one band, spanning $λ_a$ 0.28–0.94; t48 and t12 among them). Confirming the set corrects for winner's-curse. This is how the survey's handoff ("D2.2 confirms at fresh seeds before any of these numbers is quoted") is honoured on the grammar we will use.

The redder-than-both item lands here too, since every op but lighten allows that.

Optional arm: control models at one, three, and six operations, with the cube probed as in ex-2.1.12, to ask whether a richer op set gives the model a better operand geometry (see the backlog item). It is an arm on un-anchored models only, so it cannot confound the anchored conditions. Task-diversity phase transitions in in-context learning (memorization below a diversity threshold, generalization above; arXiv:2306.15063, arXiv:2405.11751) motivate a companion question on the same sweep: run on the ex-2.1.5 two-form corpus, does added op diversity move hex and named colors toward the shared representation D2.1 never found (backlog item)?

Scouting, pilots, and the grammar handover

The grammar from ex-2.2.3 makes the removal reads hard to interpret. On four of the six ops, most red lines have an answer that a model without red can still give, because the op saturates or copies the channel of the other operand. Rather than preregister the anchored-op experiments on that footing, we ran a scouting round that read candidate ops on the grid. Three pilots then retrained the adopted point, one each under a stochastically rounded corpus, a whole-line labeller, and three ways of keeping the axis off the syntax embeddings: anchoring the blocks alone, untying the readout, and hard-zeroing the syntax embeddings as a ceiling. A survey of removal operators then scored 84 of them on the checkpoints ex-2.2.3 stored. None of this work is scored; it all feeds proposals.

Ex-2.2.4 proposes table A+: drop add, add difference, exclusion, and hsvmix, and carry hue-hsv, sat-hsv, and value-hsv as a marked subset, the first ops that read operand order. It also proposes a removal statistic that measures distance from the correct answer, scored on lines where zeroing the R channel of the red operand moves the answer far, with exact match reported beside it. And for comparability with M3 it proposes stochastic rounding and the whole-line labeller, which the pilots found cost nothing on the anchoring side. Under stochastic rounding an answer is a distribution rather than a single color, so dependence, removal, and calibration are all read as how much answer mass moves, and op-relevance becomes the expected agreement between ops rather than a count.

Ex-2.2.7 proposes that we keep anchoring every slice and untie the readout. The tied readout puts the axis on the syntax embeddings so the model can predict = after a red operand; giving the readout a table of its own keeps that component on the readout instead, at no visible cost to task or placement. The hard-zeroed arm is the in-grammar ceiling the untied arm has to match. The same pilot saw a selectivity tail on its whole-line arms, where a few seeds lose many non-red lines under the projection, so the labeller goes into the handover with a check rather than by default.

Ex-2.2.8 proposes the plain projection as the removal operator, with operands beside it as the selective reference and shaped-a0.4-p0 as an optional syntax-free row.

The handover is the preregistered experiment that adopts all of this, drafted as ex-2.2.9. It runs the recipe from ex-2.2.3 on table A+, with the stochastic corpus, the whole-line labeller, and the untied readout, at twenty seeds against the twenty of ex-2.2.3. mix stays as the reference op, with hsvmix beside it. Two reference conditions each change one thing back, the either-slot labeller and the tied readout, so the selectivity check on the labeller and the cleaning read on the readout are each a two-condition comparison on the new grammar. Neither is a fallback: the either-slot labeller needs the operand positions, which M3 will not have, so a selectivity cost on the whole-line labeller is something to understand rather than something to switch away from. The removal rows are the ones ex-2.2.8 proposed. The handover ran: red lands and the task is unhurt, but removal missed its gate on the three HSV ops, one-sided by slot, so the grammar of record is still the one from ex-2.2.3. Ex-2.2.10 reads the miss off the stored runs: the projection acts like a change of the red operand's hue, and the to-zero rule that picked the removal lines counted the lines whose answer takes only red's saturation or value. The re-run (ex-2.2.11) picks the removal lines by hue, reads retention against the anneal's start, and reports the op1 alignment without a gate; after it, plain mix will likely be dropped. The prereg settles the two questions the scouting left open: lines per op stay at the corpus size, with an exploratory condition holding them at the count from ex-2.2.3 to read the E4 confound, and the probe draw walks every color as op2 as well, for the order-sensitive subset.

The recipe sweep before the anchored-op prereg

The re-run (ex-2.2.11) came out almost clean: red lands, holds through the anneal, costs the task nothing, and comes out under the projection on ten of the eleven ops. On hue-hsv the model keeps about a quarter of the answers that need the red operand's hue, a little over the gate inside a wide seed spread, so the frozen rule said not adopted. Three smaller threads came out with it: ᾱ at op1 sits higher than on either reference and has no named mechanism, the alignment drifts down before the anneal on handover alone, and the mellowmax τ was never re-tuned for the longer whole-line span.

Our reading of the miss: hue is a clock face, and the anchored axis measures how red a color is, which is the same a little clockwise of red (toward orange) and a little anticlockwise (toward pink). So the axis cannot hold which side of red a color is on, and hue-hsv with red at op2 is the only op that needs the side. If the model keeps the side off the axis, that is what survives the projection, and no force knob will change it; a plane would.

Ex-2.2.12 is the scouting round that checks this and looks for a recipe change, in two acts. Act one scores ex-2.2.11's stored checkpoints, no training: the kept share on hue-hsv's removal lines split by the side of red, the same lines under the projection applied at the embedding only or at the blocks only (the bypass read from the design), and ᾱ at op1 split by whether the whole-line labeller labelled the line through an operand or through its answer. Act two trains seven cells at five seeds on the handover setup, against ex-2.2.11's twenty handover seeds as the free reference: τ at 0.03 and 0.01, the anchor weight and the anti-subspace peak each doubled, and a two-by-two of depth (four or six blocks) and subspace (axis or plane). The promotion rule is frozen with the plan: a cell is proposed only if it clears the hue-hsv gate by more than the band ex-2.2.11 measured, with every other gate held and ᾱ at op1 no worse. That is the ex-2.1.11 survey lesson applied: a proposal that sits inside the band is not a proposal.

The branch is decided now rather than after the data. If no cell qualifies, the re-run gates removal on the ten other ops and reports hue-hsv beside them as the op where a single axis has a known blind spot, and the anchored-op experiments proceed on the reference recipe. Either way a short re-run prereg adopts at fresh seeds, as ex-2.2.11 did.

Note, 2026-09-23. No cell qualified, and the blind-spot story was wrong: the surviving lines sit on the red axis and the survival is re-derived inside the blocks. Ex-2.2.13 then climbed an anchor-weight ladder on the axis and on a plane at twenty fresh seeds, asking whether a heavier anchor makes the leftover predictable. It does not: the tight plane condition did not replicate, the mean did not move, and the recipe stays at axis-0.1 (ex-2.2.11's handover), whose retraining at fresh seeds is the re-run this section promised. The leftover, about a quarter of hue-hsv's red-dependent answers from no fixed set of lines, is carried into the anchored-op experiments as a bounded confound they measure at their own seeds. The first of them anchors an op with no red anchor beside it, so the confound does not reach it.

Prep C: Embeddings

Un-anchored embeddings

Anchor all slices except the embeddings. We don't expect concepts to map to tokens in more complex models and languages anyway.

Hypotheses: (a) the hidden-state slices align as they did in earlier experiments, which would let later experiments leave the embeddings out; (b) the color embeddings become somewhat aligned anyway, because their directions correlate with the anchored hidden states; (c) the constant component that ex-2.2.1 found on the syntax embeddings is gone, and with it the non-red cost of the plain projection.

Hypothesis (c) is scored because every training experiment from here scores its checkpoints through the eval contract. Hypotheses (b) and (c) split the embedding table by token class, so they conflict only if the correlation in (b) reaches the syntax tokens, whose hidden states are not pulled. If it does reach them, that may resolve once there are several operations and the op token does work of its own.

This is different from the layer sweep, which tests the model's ability to route around intervention.

Note, 2026-09-11. Run as the blocks-only arm of ex-2.2.7, a pilot with nine seeds per arm. (a) holds in the blocks and (b) holds, but (c) does not: the syntax embeddings keep the axis, because the tied readout puts it there rather than the pull at slice 0. Leaving the embedding un-anchored also leaves about half of a red operand's redness off the axis at slice 0, so the full-position projection is less complete on every op, and on one seed most red lines survive it. The pilot's proposal is to keep anchoring every slice and untie the readout, which moved the component onto the readout table at no visible cost. Hard-zeroing the axis component on the syntax embeddings cleaned them as well, but it selects embeddings by token class, so it does not carry beyond this grammar; it stands as the ceiling the untied arm should match.

Anchor one operation

Anchor one operation on e₁ and leave the others unlabelled. The recipe is D2.1's plus the lessons from above: every slice pulled and the readout untied, as un-anchored embeddings settled. Anchoring the embedding of the op token gives token identity by construction, the collapse the risk table warns about, so the group contrast and the suppression read are what tell an anchored op from an anchored token. The layer sweep comes after, on a frozen recipe, so layer effects are not confounded with schedule fragility (the D2.1 kickoff rule). The labeller keys on the op token, with the position-free mechanism from ex-2.1.10.

Measurements: alignment at the op position by slice; group contrast (anchored op against the others, which is the categorical form of grading); task gates against control; and a probe scan of op-identity decodability at every site, against the control.

What we hope to see from the scan is approx. no change against control — an equivalence claim, so the prereg must declare the margin it has to land within. We are anchoring the hidden state, not relocating the computation, and the scan is the side-effects read (ex-2.1.12's H2, for the op). The equivalence read may need many seeds (20?), so a small smoke test runs first: a few seeds against the alignment and task gates alone, to establish that anchoring an op works at all before the seed budget is spent on the margin.

Nice to have: Sweep over all ops to see whether they can all be anchored equally well.

Note, 2026-09-23. Ran as ex-2.2.14, five seeds of difference with the other ten ops at three seeds each. The op lands at twice the margin red reached, holds through the anneal, and costs the task nothing (the largest gap on any op is 0.005); all ten other ops anchor the same way, so difference stands. The margin saturates within three epochs because the pull puts the op word's embedding on e₁ at a cosine near 1, and the op position stays there at every slice. With the op word alone pulled, the blocks carry about a twentieth of that alignment to = and none to the answer; the whole-line pull puts about 0.1 at both. A labeller at a fiftieth of the op's lines, or with a fifth of its labels wrong, lands the op as well as one that labels every line. The probe scan leaves the op about as readable as the control has it, within ±0.1 in R² at every site but the newline at slice 1, where the control's own seeds spread by 0.4. One loose end: under the whole-line pull the first operand of every line leans toward e₁ (0.19 at the final slice), a candidate mechanism for the containment item.

The equivalence read moves into the suppression prereg. It needs fresh many-seed anchored checkpoints beside the control, which that experiment trains anyway, and the smoke test has already sized its margin.

Suppress the operation (and the operands)

The centre of D2.2.

The headline claim is selective removal: suppress difference without suppressing multiply. The ops may share a common component that means this is an operation, with the specific op only one part of the state; the group contrast from anchor operation says how large that shared part is, and the removal claim covers the op-specific part.

The dose axis is intervention strength. The stimulus side of a categorical concept grades too coarsely to serve: op-relevance (below) occupies only three or four levels on a six-op table, and a few more on the eleven-op table A+. So we scale the suppression rather than always projecting fully, keeping a fraction $1-γ$ of the component as $γ$ runs from 0 to 1; the shaped-suppression item and M1's shaped suppression are the machinery. Prediction: damage to the anchored op rises monotonically with $γ$, and the other ops stay within gate along the whole curve. Ex-2.2.8 found that a threshold set inside the alignment range the non-red lines occupy costs more than projecting everything, and that the write angle and the behavioral cost move independently. So the dose is scaled by $γ$ on the plain projection rather than by a threshold on alignment.

The per-line prediction at full suppression comes from the designed null. The null is op-averaged — the least committal prediction available from the operands with no op, given by the distribution over the answers to all ops.2 Against the op-averaged null, op-relevance for a line is the weight the mixture withholds from the anchored op's answer: zero where every op in the table agrees on that pair, $\frac{n-1}{ n}$ where the anchored op is alone in its answer. Predictions: per-line damage follows op-relevance and stays within the bound the mixture sets.

Read the response off the probability mass on the correct answer, or off the decoded answer's distance from it. Hard accuracy steps rather than grades under an averaged null, since greedy decoding keeps the plurality answer, so accuracy is the gate statistic and not the response statistic.

Conditions test the contrast from m1/ex-2.9.2: control, no-fallback, fallback — plus a filtered-corpus row, a control trained with the anchored op's lines held out. The fallback trains toward the op-averaged distribution — the designed null itself, as soft labels — so fallback and no-fallback share the per-line prediction, and the fallback condition's claim is tighter adherence to it: less seed scatter, more mass on the mixture. That is the removal reference the baselines item wanted placed, and the eval contract scores it like any other triple.

Operand suppression beside operation suppression, with red anchored in the same models, is a follow-up. The first suppression experiments anchor the op alone, as ex-2.2.14 did, so the red leftover on hue-hsv stays out of their reads; the two-anchor experiment measures that confound at its own seeds.

What ex-2.2.14 changes. The anchor landed as the op word's state: on the anchored op's lines the op position sits on e₁ at a cosine near 1 at every slice, and holds little else. Three consequences for the plan above.

  1. The projection has no defined landing at the op word. Projecting e₁ out of a state that is almost all e₁ leaves a small remainder, which the re-projection onto the sphere scales up by $1/\sqrt{1-x_1^2}$; the contract's own gain is infinite in the limit. So the edited state is whatever the remainder happens to be. The dose by $γ$ does not grade there either: the state keeps pointing along e₁ until $1-γ$ falls to about the size of the remainder, so the whole dose-response curve sits in the last few percent of $γ$. The reflection ($γ = 2$) is well defined, and so is a repulsion onto a declared landing state, whose dose is the alignment it lands at. The dose axis for the op is one of these, chosen on data.
  2. An edit at the op word removes the token. Since the op word's state is the concept, suppressing it there should remove the op, and so would masking the word. Removal at the op word therefore cannot tell an anchored op from an anchored token, which is the risk the risk table names. Every suppression read carries a token-mask row as the reference: the op word's state replaced by a neutral one (the mean over the eleven op words), at the same slices. What the anchor adds over knowing which token names the op is read off the edits away from the op word: at the use sites, and in the blocks only, which is the bypass test below. Ex-2.2.14's contrast predicts little from the use sites on the op-word arm, and some on the whole-line primary, which puts about 0.1 there.
  3. A full-position edit touches every line. Under the whole-line pull the first operand of every line leans toward e₁, so an edit at every position reaches lines of every op, as the syntax embeddings did for red before the readout was untied. Edits at the op word alone avoid it, and play the part the operands edit played for red.

A scoring-only pass comes first, as ex-2.2.1 came before the fallback term and ex-2.2.8 before the handover. It suppresses difference on ex-2.2.14's stored checkpoints (the primary, the op-word and every-line arms, the ten sweep ops, and the control), with no training: the projection, the reflection, a repulsion, and the token mask, each at the op word, the use sites, and every position, at the embedding, in the blocks only, and at every slice. It reads per-line damage against op-relevance, and selectivity on the other ten ops. From that the prereg takes its operator and dose axis, its effect sizes, and whether the use-site edits are worth gating. The prereg then trains fresh seeds with the conditions above (control, no-fallback, fallback, filtered corpus) and carries the equivalence read from anchor operation.

The bypass test the D1.3 post left open: suppress at the op-token position only, at all positions, at one slice, at all slices. Where suppression fails to bite, the model is reading the op from somewhere the axis does not reach. That is a finding about anchoring, and it feeds the layer sweep.

Nice to have: Sweep over all ops to see whether they can all be suppressed equally well.

Layer sweep

Anchor at subsets of slices (single ℓ, prefix ≤ ℓ, suffix ≥ ℓ, all) on a frozen schedule, then run the suppress operation intervention at the anchored slices. This tests the claim that bounds are layer-local. The geometric bound covers the immediate write, and we need to know whether later blocks amplify or absorb the edit.

Ex-2.2.1 already says what to expect for red. The concept is read in the first two blocks, and the last block does not read it. So a removal acting at the embedding and the first block should match the full intervention, and one that starts later should not. That is a prediction about prefixes and suffixes, so we favor a prefix/suffix bracket over single slices.

Caveat on reading depth this way: at d64-L4 on one operation the model has capacity to spare, so where it reads the concept may say more about what it can afford than about what the task needs. Several operations may draw on more of the depth, and the character-level runs (ex-2.1.5, ex-2.1.6) already used more layers than the word-level ones. Still to decide: whether this is one experiment or two, since anchor layer × intervention layer is a grid.

SGTM baseline

Its own preregistered experiment (arXiv:2512.05648), method-agnostic through the eval contract, reusing our labels. RMU and SAE rows wait for D2.3, as the baselines item sequences them.

Write-up

The D2.2 post.

Deps

The contract also pins where operators act and what they do to the norm. The hook point is the between-block stream, meaning the slices that residual_stream() returns, which are the same states the anchor term reads. The stream is unit-norm (nGPT), so axis projection composes with a re-projection back onto the sphere. The state lands on the great subsphere where zero-concept states live, and the surviving components pick up a per-position gain of $1/\sqrt{1-x_1^2}$. That gain is computable beforehand, so it belongs inside the bound. Weight ablation declares its order against the normalize_weights constraint, which rescales a matrix once entries are zeroed.

The 2025 auditing rows are relearning rebound (arXiv:2505.22310), activation perturbation (ActPert, arXiv:2505.23270), and off-axis recoverability (arXiv:2605.11685). They were to be declared before any arm was scored. Ex-2.2.1 scored its arms without them, and the decision is now recorded under fallback control: off-axis recoverability runs there, and the other two wait for D2.3, which speaks their language in full. - Fold the survey lessons into the plan template: constraint margins ranked beside the objective, multi-seed promotion near a gate, and the publisher carrying every statistic the analysis promises.

Decisions

Only what the plan above already commits to; everything else stays open until an experiment forces it.

Risks and mitigations

Risk Would look like Retired at
Suppression does not bite even on red Accuracy unchanged after projecting e₁ out; color is read from elsewhere Retired: ex-2.2.1 (red accuracy 0.09)
The bound is loose or wrong in a transformer The non-red write exceeds its geometric bound, or the damage outruns the write-size prediction Write half retired at ex-2.2.1; the behavioral half belongs to the layer sweep
The response to suppression is undesigned Suppressed lines scatter by seed; completions leave the color vocabulary Measured at ex-2.2.1 (13% seed agreement, completions stay in vocabulary); fallback control pins it
Recipe is grammar-specific The proposals from the survey do not reproduce on the new grammar Retired at ex-2.2.3: the recipe and every proposal reproduce (H2, H3), at a plateau 0.05 lower than the survey's; the frozen rule adopted t00, and a post hoc read with the lead and selectivity gates narrows the choice to the recipe, whose short arm D2.2 builds on (recipe-short, by decision after the twenty-seed E6 read)
The syntax embeddings hold the axis, so a full-position edit costs the non-red lines Non-red lines lose accuracy under the plain projection, at the op-word and = positions Found at ex-2.2.1 and ex-2.2.3; the mechanism is the tied readout, at ex-2.2.7; the untied readout is confirmed on the new grammar at the handover and again at fresh seeds in ex-2.2.13
The red leftover on hue-hsv confounds the anchored-op reads A removal read on a model carrying both anchors moves with which lines survived on red Measured at ex-2.2.11 and ex-2.2.13: about a quarter of those answers, from no fixed set of lines. Ex-2.2.14 carries no red anchor; the two-anchor follow-up measures it at its own seeds
Task cost grows with an abstract concept anchor Gate misses in anchor operation that new grammar did not have Retired at ex-2.2.14: no op moves by more than 0.005, on the anchored op or any other
Anchoring an op captures the token, not the operation Suppression at the op word removes the op no better than masking the word, and edits away from it are inert suppress operation; ex-2.2.14 found the op word on e₁ and little carried to the use sites, so this is the likely reading to rule out
Bypass through attention or the residual Suppression works only when applied at every site suppress operation, layer sweep

The first experiments all change one thing from D2.1, so a negative there should be interpretable.

Out of scope

Verification lines (D2.3). Several ops on separate axes (a D2.3 candidate, for the subspace bound and the two-op concept swap). The feedback controller (the kickoff advice stands). RMU and SAE baselines, and a full LUNAR row beside them; a LUNAR-style redirect, the nearest analogue of the fallback, runs as an exploratory row of ex-2.2.2. Resolving D2.1's H2 decodability question (its own item), run when the claim is needed. Stream-vs-init attribution (kickoff queue item 4), until there is an anchoring failure worth attributing — D2.1 produced none.


  1. A stop-gradient is an identity in the forward pass with a gradient of zero: a loss downstream of it cannot move anything upstream of it. Here it sits on the reflected states, so the fallback loss reaches the blocks and the unembedding and leaves the embedding table alone, and with it where the anchor put the concept. The other losses see the clean pass and are unaffected. ↩

  2. We considered calling this op-marginal, but we also have a margin measurement $m$, which would be confusing. ↩