mechinterp-investigator

Coordinate phased research workflows to investigate and label SAE features.

1|Updated Jul 9, 2024
One-click install
npx skills add https://github.com/cesaregarza/SplatNLP --skill mechinterp-investigator
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: mechinterp-investigator
Source: https://github.com/cesaregarza/SplatNLP/tree/main/.claude/skills/mechinterp-investigator
Command: npx skills add https://github.com/cesaregarza/SplatNLP --skill mechinterp-investigator

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill orchestrates a structured, phase-driven research program to systematically investigate and label SAE features, turning opaque activations into meaningful concepts.

Core Features & Use Cases

  • Phase-driven investigation: Triage, overview, activation-region analysis, and meta-informed weapon analysis.
  • Reproducible labeling: Generates interpretable labels and documentation for SAE features across models and runs.
  • Collaboration & tooling: Integrates with the mechinterp CLI workflows to standardize analyses across teams.

Quick Start

Install the required tooling and run the overview for a target feature, e.g. poetry run python -m splatnlp.mechinterp.cli.overview_cli --feature-id {FEATURE_ID} --model ultra --top-k 20

Frequently Asked Questions about mechinterp-investigator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I systematically label sparse autoencoder features for mechanistic interpretability?

SAE feature investigation uses a phased workflow: triage, overview, activation-region analysis, and meta-informed weapon analysis. This structured research program turns opaque activations into reproducible, meaningfully labeled concepts by systematically analyzing feature behavior across models and runs.

What is the best way to structure an SAE feature analysis workflow across multiple model runs?

The best way to structure SAE feature analysis is a phase-driven research program coordinating triage, overview, activation-region analysis, and meta-informed weapon analysis. This standardizes investigations across teams, ensuring reproducible labeling and interpretable results for features extracted from different models and runs.

How do I generate reproducible labels for SAE features using CLI workflows?

Generate reproducible SAE feature labels by running the mechinterp CLI tools managed through Poetry dependencies. Execute the overview CLI for a target feature ID to produce standardized, interpretable documentation across your selected models and runs.

Do I need Poetry to manage dependencies for SAE interpretability workflows?

Yes, you need Poetry to manage dependencies for this SAE interpretability workflow. The structured research program relies on Poetry-managed environments to execute CLI tools and feature inventories, ensuring reproducible, interpretable results across different models and analysis runs.

Can I use the mechinterp workflow to analyze simulated SAE features?

Yes, you can use the mechinterp workflow to analyze simulated SAE features. The phased investigation program applies triage, overview, activation-region analysis, and weapon analysis to both real and simulated SAE features, generating interpretable labels and documentation.

What are the limitations of relying on a phased workflow for SAE feature labeling?

A limitation of phased SAE feature labeling workflows is the strict dependency on CLI tools, Poetry-managed dependencies, and accurate feature inventories. Without complete feature inventories or proper environment setup, the structured analysis cannot produce reproducible, interpretable results.