Activation Addition and Representation Engineering: Mathematical Foundations, Linear Subspace Projections, and Inference-Time Steering in Large Language Models
Large language models encode vast linguistic, factual, and behavioral properties within their internal hidden representations. While traditional alignment and behavioral modification rely on Supervised Fine-Tuning (SFT) or Reinforcement Learning from Human Feedback (RLHF), these gradient-based techniques modify billions of parameter weights, require substantial compute, and frequently suffer from catastrophic forgetting or alignment tax. An alternative paradigm grounded in mechanistic interpret
1 min
