data-miner

Discover and rank data sources with verified trigger fields and join keys.

7|3|Updated Jan 20, 2026
One-click install
npx skills add https://github.com/SantaJordan/CTOx --skill data-miner
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-miner
Source: https://github.com/SantaJordan/CTOx/tree/main/skills/data-miner
Command: npx skills add https://github.com/SantaJordan/CTOx --skill data-miner

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Data Miner helps you discover and rank the best data sources that prove which companies in a vertical are currently in pain (a trigger like an expiration, violation, filing, hiring, recall, or gap), then verifies the needed fields actually exist and can be joined.

Core Features & Use Cases

  • Vertical intake to precision-match signals: narrows to a single company type, a specific pain/trigger, geography, and the join key you need (domain, address, contact, etc.).
  • Discovery across 7 source categories: sweeps government, licensing boards, industry associations, maps/local data, open datasets, commercial APIs, and marketplaces to generate a strong candidate set.
  • Phase-2 verification + field confirmation: opens each candidate source to confirm the data is real, the trigger field exists, and a join key is available—no assumptions.
  • 7-criteria scoring and cheap-first ranking: ranks each verified source on Quality, Scale, Price, Freshness, Accessibility, Granularity, and Pain-signal strength, prioritizing cheap high-quality options.
  • Synthesis for copy-resistant targeting: combines 2+ sources into a defensible, pain-qualified target list using patterns like temporal correlation and threshold crossing.

Quick Start

Run /data-miner and describe your vertical, the specific trigger you care about, your target geography, and the join key you want to use.

Frequently Asked Questions about data-miner

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I find cheap data sources for lead generation that show which companies are in pain?

To find cheap data sources for lead generation, you can discover and rank public and commercial datasets across seven categories, verifying that trigger fields and join keys exist to build a pain-qualified target list.

What is data enrichment for GTM research and how does it verify sourcing?

Data enrichment for GTM research involves sweeping candidate datasets to confirm the trigger field exists and a join key is available, ensuring the sourced data is real and can be linked to your existing records.

How do I build a defensible prospecting dataset from multiple data sources?

You build a defensible prospecting dataset by synthesizing two or more verified sources into a pain-qualified target list, using patterns like temporal correlation and threshold crossing to create copy-resistant targeting.

Can I use web verification to confirm trigger fields exist in commercial APIs before sourcing?

Yes, web verification opens each candidate source to confirm the data is real, verifies the specific trigger field exists, and checks that a required join key like domain or address is available.

What's the best way to score and rank data sources for a sales workflow?

The best way to score data sources for a sales workflow is ranking each verified option on seven criteria: Quality, Scale, Price, Freshness, Accessibility, Granularity, and Pain-signal strength, prioritizing cheap high-quality options first.

Pain signal sourcing not working with generic company lists, what should I use instead?

When generic company lists fail to reveal pain signals, you should discover specific datasets across government, licensing boards, and industry associations that expose triggers like violations, filings, or expirations.