Why some latent structures in LMs are actually actionable while others fail
Causal influence in language models isn't a binary switch but a product of three specific constraints: capacity, responsiveness, and alignment. When any one of these is low, steering a model's output becomes nearly impossible, with low capacity and responsiveness slashing effectiveness by 84% and 95% respectively. To successfully steer a concept, all three factors must be high simultaneously.
What actually drives causal influence?
The core problem in mechanistic interpretability is that finding a direction in activation space associated with a concept doesn't guarantee that moving the model along that direction will actually change the output. The research by Or Shafran and Mor Geva identifies three distinct reasons why this happens.
First is capacity, which describes how sensitive the model's final output is to movement along a specific latent structure. If the capacity is low, the model essentially ignores the shift. Second is responsiveness, which is context-dependent. A concept might be "promotable" in one prompt but completely inert in another. Third is alignment, which measures how well the global structure of a concept matches the specific representation the model is using in a given context. If alignment is low, attempting to steer the model can actually reverse the effect, suppressing the concept instead of amplifying it.
These findings were validated across 4 LM families and 50 different concepts, proving that causality is not an intrinsic property of the latent structure itself but fluctuates based on the context.
How to improve steering with causal probes
Because causally effective directions exist within a low-dimensional subspace that shifts depending on the context, standard linear probes often fail to capture the actual "actionable" part of a representation. The authors propose a method of restricting the training of linear probes specifically to this subspace.
The result is the creation of "causal probes." By focusing only on the dimensions that actually influence the output, these probes showed a 17%-118% improvement in steering across the models tested. Crucially, this gain in control doesn't come at the cost of accuracy; there was only a 3% reduction in the ability to detect the concept.
When to apply these constraints in model steering
If you are attempting to steer a model and the output isn't changing despite moving along a known concept direction, the failure is likely due to one of these three bottlenecks.
If the model is simply unresponsive, you are dealing with a responsiveness issue—the context is shielding the concept from being promoted. If the output moves in the opposite direction of your intent, it is an alignment failure. If the output barely budges regardless of the magnitude of the shift, the capacity is too low.
To fix this, the next step is to move away from global probes and instead utilize causal probes that target the context-specific subspace. This ensures that the steering vector is aligned with the model's current internal state, making the intervention actionable.
For those wanting to check the full technical breakdown, the paper is available at https://arxiv.org/abs/2610.06897.
In my jailbreak testing, low responsiveness killed steering even when I found a clear causal vector.