Perturbation probing reveals concentrated LLM safety
🔎 Our new research introduces perturbation probing, a two-pass, low-cost method that identifies the small set of feed-forward neurons causally responsible for targeted behaviors in aligned LLMs. Applied to Qwen3-4B and Qwen3.5-2B, the method found that tens of neurons (a tiny fraction of the model) control refusal and agreement behaviors, showing alignment can be highly concentrated. The study also defines the FFN/Skip ratio as a quick diagnostic predicting fragility across models.
