training-data-review

Audit AI training datasets for Russian legal compliance.

17|4|Updated May 25, 2026
One-click install
npx skills add https://github.com/AlsKozlov/ru-legal --skill training-data-review
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: training-data-review
Source: https://github.com/AlsKozlov/ru-legal/tree/main/packs/ai-governance/skills/training-data-review
Command: npx skills add https://github.com/AlsKozlov/ru-legal --skill training-data-review

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill solves the critical problem of unregulated AI/ML training data usage that violates Russian federal laws, exposing organizations to legal penalties, model liability, and operational risks for non-compliance with personal data, copyright, and sanctions rules.

Core Features & Use Cases

  • Comprehensive Compliance Framework: Provides structured checklists for 152-FZ personal data legal basis assessment, anonymization guidance, and special category data handling tailored to Russian regulations.
  • Copyright & IP Risk Mitigation: Covers Russian copyright rules (no US-style fair use doctrine), licensing requirements for public, purchased, and scraped data, and high-risk source identification.
  • End-to-End Audit Workflow: Guides users through a 5-step process from training data inventory to remediation planning and formal datasheet documentation, with built-in risk scoring for data sources.
  • Use Case: A team fine-tuning a Russian-language customer service LLM can use this Skill to audit their internal chat logs, public Wikipedia dumps, and licensed news datasets to identify 152-FZ and copyright gaps before training begins.

Quick Start

Invoke the training-data-review skill to conduct a full compliance audit of your planned AI model training dataset against Russian personal data, copyright, and sanctions regulations.

Frequently Asked Questions about training-data-review

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit training data for Russian 152-FZ personal data compliance?

To audit training data for Russian 152-FZ compliance, validate the legal basis for personal data inclusion, apply anonymization guidance, and filter special category data. This Skill provides structured checklists to ensure your AI datasets meet federal requirements.

What is the process for checking copyright compliance in machine learning datasets?

Checking copyright compliance in machine learning datasets involves verifying licensing requirements for public, purchased, and scraped data, as Russian copyright law lacks a US-style fair use doctrine. This Skill identifies high-risk sources to mitigate IP infringement risks.

Can I use this to review synthetic data and user-generated content for LLM fine-tuning?

Yes, you can review synthetic data and user-generated content for LLM fine-tuning. This Skill applies compliance checks across internal, public, licensed, user-generated, and synthetic data sources to identify regulatory gaps before training begins.

How do I document training data lineage for a regulatory audit?

Documenting training data lineage for a regulatory audit requires tracking data sources, legal basis, and remediation steps. This Skill provides a 5-step workflow ending in formal datasheet documentation to satisfy Russian regulatory audit requirements.

What are the limitations of using automated data audits for post-2022 sanctions rules?

Automated data audits for post-2022 sanctions rules provide structured checklists and risk scoring but cannot replace formal legal counsel. This Skill identifies potential sanctions compliance gaps in your training data inventory but requires human review for final validation.