What problem does it solve?
This Skill helps you analyze and surgically reduce refusal behavior in open-weight language models when you need a model to respond more consistently without full retraining.
Core Features & Use Cases
- Refusal analysis: Inspect where and how guardrail behavior forms using mechanistic interpretability tools such as logit lens, causal tracing, and concept geometry.
- Targeted model surgery: Apply projection-based methods, SVD variants, LEACE, steering vectors, and head or neuron ablation to modify refusal behavior with controlled risk.
- Reproducible workflows: Run scripted ablation studies, compare methods, verify refusal rates, and generate research-ready outputs from YAML templates.
- Use case: A researcher wants to benchmark which removal strategy best reduces refusals on a specific instruct model while preserving perplexity and reasoning quality.
Quick Start
Use the obliteratus skill to analyze the model, recommend the best removal method, and run a reproducible ablation study for the specified open-weight checkpoint.