data-classification-scanner

Scan GitHub, S3, and Google Drive for sensitive data patterns.

1|Updated Feb 26, 2026
One-click install
npx skills add https://github.com/webrix-ai/agent-skills --skill data-classification-scanner
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-classification-scanner
Source: https://github.com/webrix-ai/agent-skills/tree/main/skills/data-classification-scanner
Command: npx skills add https://github.com/webrix-ai/agent-skills --skill data-classification-scanner

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps teams find where sensitive data lives before audits, incidents, or regulatory reviews by scanning repositories and cloud storage for PII, PHI, financial data, and exposed credentials.

Core Features & Use Cases

  • Scan GitHub repositories for sensitive patterns in code, configs, seed data, and history.
  • Inventory S3 buckets and Google Drive content to detect sensitive files, access risks, and encryption gaps.
  • Map data flows and produce a prioritized classification report with redacted findings and remediation guidance.
  • Use it for HIPAA, GDPR, PCI-DSS, and internal security reviews when you need a defensible data inventory.

Quick Start

Ask the assistant to scan the specified repositories, S3 buckets, and Google Drive folders for sensitive data and generate a classification report with findings, data flow mapping, and remediation priorities.

Frequently Asked Questions about data-classification-scanner

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scan GitHub repositories and S3 buckets for sensitive PII and PHI data?

To scan GitHub repositories and S3 buckets for sensitive PII and PHI data, this Skill applies pattern-based detection across code, configs, and cloud storage. It generates a sensitivity-ranked remediation report with redacted findings.

What is the best way to prepare for a GDPR or HIPAA compliance audit of cloud storage?

Preparing for a GDPR or HIPAA compliance audit requires a defensible data inventory. This Skill maps data flows and detects exposed credentials across GitHub and Google Drive to produce prioritized classification reports.

Can I use pattern-based scanning to find exposed credentials in my code history?

Yes, you can use pattern-based scanning to find exposed credentials in your code history. This Skill scans GitHub repositories for sensitive patterns in seed data and historical commits to identify leaked financial data.

Does Google Drive data classification support redacted evidence capture?

Yes, Google Drive data classification supports redacted evidence capture. The Skill inventories Google Drive content to detect sensitive files and encryption gaps while capturing redacted evidence for security reviews.

What limitations exist when mapping data flows across multiple cloud storage platforms?

When mapping data flows across multiple cloud storage platforms, limitations depend on available access to S3 buckets and Google Drive folders. The Skill requires explicit specification of target repositories and drives to perform pattern-based scanning.