stuart-russell

Translate AI safety prompts into uncertainty-aware guidance aligned with human preferences.

100|8|Updated Apr 22, 2026
One-click install
npx skills add https://github.com/K-Dense-AI/mimeographs --skill stuart-russell
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: stuart-russell
Source: https://github.com/K-Dense-AI/mimeographs/tree/main/mimeographs/stuart-russell
Command: npx skills add https://github.com/K-Dense-AI/mimeographs --skill stuart-russell

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill guides you to apply Stuart Russell's engineering-first approach to AI safety, helping you reason about uncertainty in objectives, safety by design, burden of proof, and realization of human preferences when evaluating AI systems and governance.

Core Features & Use Cases

  • Surface Russell's core frameworks (Assistance Games, Red Line Regulation) and his mental models (The King Midas Problem, The Off-Switch Game, Breeding Horses, The Chernobyl Analogy, The Wall-E Problem) to structure analysis of AI safety, governance, and policy.
  • Provide decision-ready prompts for evaluating frontier AI deployments, regulatory proposals, and risk governance with explicit emphasis on uncertainty and human preferences.
  • Use Case: When assessing a high-stakes AI system, apply Assistance Games to align incentives and Red Line Regulation to set pre-deployment constraints.

Quick Start

Identify the user's objective, reveal unstated constraints, and reframe the task around human preferences before acting.

Frequently Asked Questions about stuart-russell

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is value alignment in AI safety and how does it handle objective uncertainty?

Value alignment in AI safety structures objectives around explicit human preferences to manage objective uncertainty. It reframes AI actions to defer to human preferences, ensuring systems remain uncertain about true goals and operate safely by design.

How do I apply assistance games to align AI actions with human preferences?

Apply assistance games by modeling AI interactions as cooperative processes where the system observes human choices to infer underlying preferences. This aligns incentives by optimizing for human satisfaction rather than rigidly optimizing a potentially flawed objective.

Does this approach to AI safety support red line regulation for frontier models?

Yes, this approach supports red line regulation by setting pre-deployment constraints and boundary conditions. It applies burden of proof concepts to ensure frontier AI systems meet formal safety criteria before deployment in high-stakes governance contexts.

Can I use these mental models for evaluating AI governance policy?

Yes, you can use mental models like The King Midas Problem and The Off-Switch Game to evaluate AI governance policy. They provide decision-ready frameworks to analyze regulatory proposals and risk governance with explicit emphasis on uncertainty.

When do I need formal proof concepts for safe AI deployment?

Formal proof concepts for safe AI deployment are needed when assessing high-stakes systems where objective uncertainty and safety-by-design matter. They establish boundary conditions and verify that actions align with governance priorities before realization.