What problem does it solve? Neural network neurons are polysemantic, activating for many unrelated concepts due to superposition, which makes model internals hard to interpret. This Skill guides you through using SAELens to decompose dense activations into sparse, monosemantic features that correspond to human-interpretable concepts. ## Core Features & Use Cases - Pre-trained SAE Analysis: Load SAEs from releases like gpt2-small-res-jb, encode activations into sparse features, and identify top-activating features per token. - Custom SAE Training: Configure and train Standard, Gated, TopK, or JumpReLU SAEs with L1 warm-up, ghost grads, and W&B logging, then validate with L0, CE loss recovery, and dead feature metrics. - Feature Steering & Attribution: Compute per-feature logit contributions, steer generation by adding decoder directions to the residual stream, and ablate features to test causal importance. - Use Case: You want to find which features in GPT-2 drive the prediction of "Paris". Load the layer-8 residual SAE, compute feature contributions via decoder weights and the unembedding matrix, then steer or ablate the top feature to verify its causal role. ## Quick Start Ask the AI to load the gpt2-small-res-jb pre-trained SAE with SAELens and show the top-activating features for each token in a sample prompt.