data-research

Automate structured data research with search, extraction, archiving, deduplication, and tracker updates.

Updated Apr 11, 2026
One-click install
npx skills add https://github.com/sixtycat2000-ctrl/gbrain-qmd --skill data-research-sixtycat2000-ctrl
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-research
Source: https://github.com/sixtycat2000-ctrl/gbrain-qmd/tree/main/skills/data-research
Command: npx skills add https://github.com/sixtycat2000-ctrl/gbrain-qmd --skill data-research-sixtycat2000-ctrl

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the process of structured data research, enabling users to efficiently search, extract, archive, deduplicate, and track data from various sources.

Core Features & Use Cases

  • Data Search and Extraction: Search through multiple sources including email, web, APIs, and attachments to extract structured data.
  • Data Archiving: Save raw data sources for future reference and ensure data integrity.
  • Deduplication: Remove duplicate entries to maintain accurate data.
  • Tracker Updates: Update and backlink canonical tracker pages with new data.
  • Built-In Recipes: Includes pre-defined recipes for investor updates, donations, and company metrics.
  • Use Case: A financial analyst can use this Skill to track and analyze investment data from emails and web sources, automating the process of data collection and analysis.

Quick Start

Run the data-research skill with the 'research' trigger to initiate the structured data research process.

Frequently Asked Questions about data-research

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate structured data extraction from emails and web sources?

Automated structured data extraction is executed through customizable YAML recipes that search emails, web sources, APIs, and attachments to collect specific data points. This streamlines capturing information from multiple sources into a unified dataset.

What is the best way to deduplicate extracted data entries?

Data deduplication is best handled by running an automated research pipeline that removes duplicate entries during the archiving process. This ensures your canonical tracker pages maintain accurate and clean data records without manual review.

Does data research work with custom YAML recipes for specific collection tasks?

Yes, custom YAML recipes support specific structured data collection tasks. You can define recipes to track investor updates, donations, and company metrics, tailoring the search, extraction, and tracker update processes to your exact requirements.

How do I update canonical tracker pages with new extracted data?

Canonical tracker pages are updated automatically by the research pipeline after searching and deduplicating new entries. The process backlinks the tracker pages with the newly archived data to ensure your records stay current.

Do I need local data sources to run structured data research pipelines?

Yes, structured data research pipelines require access to local data sources and structured data formats to execute the search, extraction, and archiving steps. This ensures the YAML recipes can accurately query and process your target information.

Why should I archive raw data sources during extraction?

Archiving raw data sources during extraction preserves the original information for future reference and ensures data integrity. This allows you to verify the extracted structured data against the source material later if discrepancies arise.