obliteratus

Analyze and transform refusal behavior in open-weight language models.

3|1|Updated Apr 19, 2024
One-click install
npx skills add https://github.com/guccang/blogclaw --skill obliteratus-guccang
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: obliteratus
Source: https://github.com/guccang/blogclaw/tree/main/cmd/hermes-agent/vendor/hermes_runtime/skills/mlops/inference/obliteratus
Command: npx skills add https://github.com/guccang/blogclaw --skill obliteratus-guccang

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps researchers and AI engineers analyze and modify refusal behaviors in open-weight language models by providing structured workflows for studying alignment mechanisms and applying model weight transformations.

Core Features & Use Cases

  • Refusal Analysis: Inspect model alignment patterns using mechanistic interpretability methods such as activation analysis, logit lens, causal tracing, and concept geometry.
  • Abliteration Workflows: Guide model transformation experiments using methods including diff-in-means, SVD, LEACE, and surgical approaches with verification metrics.
  • Use Case: A researcher evaluating an aligned open-weight model can use this Skill to analyze refusal directions, select an appropriate transformation strategy, and measure changes in refusal rate and model quality.

Quick Start

Use the obliteratus skill to analyze the refusal mechanisms of my selected Hugging Face model and recommend an appropriate workflow.

Frequently Asked Questions about obliteratus

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze refusal directions in an aligned Hugging Face model?

To analyze refusal directions in an aligned Hugging Face model, you can use mechanistic interpretability workflows like activation analysis, logit lens, causal tracing, and concept geometry to inspect alignment patterns.

What is the best way to apply abliteration workflows to open-weight language models?

The best way to apply abliteration workflows to open-weight language models is by using transformation methods like diff-in-means, SVD, LEACE, and surgical approaches, guided by structured verification metrics.

Can I use mechanistic interpretability to modify LLM refusal behaviors?

Yes, you can use mechanistic interpretability to modify LLM refusal behaviors through controlled weight transformation experiments that discover and alter refusal directions in open-weight models.

Do I need hardware evaluation to perform model analysis for refusal mechanisms?

Yes, hardware evaluation is required to perform model analysis for refusal mechanisms, ensuring your environment can handle the computational demands of inspecting and transforming open-weight language models.

How do I measure changes in refusal rate and model quality after abliteration?

To measure changes in refusal rate and model quality after abliteration, apply verification metrics that evaluate the controlled weight modifications and assess the transformed model's performance.

When do I need mechanistic interpretability for LLM alignment research?

You need mechanistic interpretability for LLM alignment research when studying alignment mechanisms, discovering refusal directions, and planning surgical weight transformations to modify model behaviors.