What problem does it solve?
This Skill helps researchers and AI engineers analyze and modify refusal behaviors in open-weight language models without retraining, addressing the challenge of understanding and changing alignment-related model behavior.
Core Features & Use Cases
- Refusal Analysis: Examine model refusal mechanisms using mechanistic interpretability techniques such as activation analysis, direction extraction, and causal tracing.
- Model Modification Workflows: Guide abliteration runs with multiple projection methods, verification metrics, hardware checks, and model-specific recommendations.
- Use Case: Apply the workflow when researching model behavior, comparing refusal mechanisms across architectures, or evaluating controlled changes to open-weight LLMs.
Quick Start
Use the obliteratus skill to analyze and run an abliteration workflow for my selected Hugging Face model with appropriate hardware checks and verification steps.