obliteratus

Remove refusal directions from open-weight LLMs via CLI-driven abliteration.

1|1|Updated Apr 26, 2026
One-click install
npx skills add https://github.com/BermudaLocals/hermes-agent-lite --skill obliteratus-bermudalocals
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: obliteratus
Source: https://github.com/BermudaLocals/hermes-agent-lite/tree/main/skills/mlops/inference/obliteratus
Command: npx skills add https://github.com/BermudaLocals/hermes-agent-lite --skill obliteratus-bermudalocals

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Remove refusal behaviors from open-weight LLMs using mechanistic interpretability techniques.

Core Features & Use Cases

  • Mechanistic interpretability driven abliteration to surgically excise refusal directions while preserving reasoning across tasks.
  • CLI-driven workflow with 9 methods, 28 analysis modules, 116 model presets, and telemetry-driven recommendations.
  • Supports modular templates, pre-run checks, and comprehensive verification to ensure model quality post-ablation.

Quick Start

Run the CLI to obliterate a target model and verify the post-ablation results before deployment.

Frequently Asked Questions about obliteratus

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I remove guardrails from open-weight LLMs while preserving reasoning?

Abliteration removes refusal directions from open-weight LLMs using mechanistic interpretability techniques. It surgically excises refusal behaviors through a CLI-driven workflow while preserving the model's core reasoning capabilities across tasks.

What is abliteration and how does it use mechanistic interpretability?

Abliteration is a mechanistic interpretability technique that identifies and removes refusal directions within open-weight LLMs. It surgically targets the specific activation patterns causing refusal behaviors to uncensor the model without degrading its underlying reasoning.

Does abliteration support different methods for uncensoring LLMs?

Yes, the abliteration workflow supports 9 distinct methods for removing guardrails from open-weight LLMs. It also includes 28 analysis modules and 116 model presets to configure the uncensoring process according to specific research requirements.

Can I verify model quality after removing refusal directions?

Yes, comprehensive verification ensures model quality post-ablation. The process includes pre-run checks and analysis modules that validate the open-weight LLM's reasoning capabilities after the refusal directions are surgically excised.

What are the limitations of using abliteration on open-weight models?

Abliteration is limited to open-weight LLMs and applies to research and development tasks. While it preserves core reasoning during uncensoring, users must still run comprehensive verification and post-ablation checks before deploying the modified model.

Why does my LLM still refuse prompts after applying abliteration?

Incomplete abliteration occurs when refusal directions are not fully excised from the open-weight LLM. Using the 28 analysis modules and telemetry-driven recommendations helps identify remaining guardrails and refine the ablation method for complete removal.