data-research

Extract structured data from emails, web sources, and APIs into tracker pages.

45|11|Updated Mar 17, 2026
One-click install
npx skills add https://github.com/beyonai/ByClaw --skill data-research-beyonai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-research
Source: https://github.com/beyonai/ByClaw/tree/main/middleware/openclaw/skills/gbrain/references/data-research
Command: npx skills add https://github.com/beyonai/ByClaw --skill data-research-beyonai

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill solves the challenge of manually tracking structured data from fragmented sources like emails, web pages, and APIs, preventing data loss and ensuring consistent record-keeping.

Core Features & Use Cases

  • Automated Research Pipeline: Executes a 7-phase process to search, extract, archive, and deduplicate data.
  • Recipe-Driven Extraction: Uses YAML-based recipes to handle diverse tasks like investor updates, expense tracking, or company metrics.
  • Use Case: Automatically monitor your inbox for investor update emails, extract key financial metrics like ARR and MRR, and update a canonical tracker page with backlinked sources.

Quick Start

Use the data research skill to initialize a new investor-updates recipe and begin tracking metrics from your email.

Frequently Asked Questions about data-research

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate structured data extraction from emails and web sources?

Automate structured data extraction by defining YAML-based recipes that configure search queries, extraction schemas, and deduplication rules. This pipeline executes a 7-phase process to search, extract, archive, and deduplicate data from fragmented sources into canonical tracker pages.

What is recipe-driven data extraction and how does it work?

Recipe-driven data extraction uses YAML configuration files to define search queries and extraction schemas for diverse workflows. This mechanism processes fragmented sources like emails and APIs, applying deduplication rules to reliably transform unstructured inputs into structured canonical tracker pages.

Can I track investor update emails and extract financial metrics automatically?

You can track investor update emails by initializing a recipe to monitor your inbox, extract key financial metrics like ARR and MRR, and update a canonical tracker page with backlinked sources. This supports diverse research workflows including company metric monitoring.

How do I set up a research pipeline for expense tracking and company metrics?

Set up a research pipeline by configuring a YAML recipe that defines your expense tracking or company metrics extraction schema. The pipeline then executes its 7-phase process to search sources, extract structured data, archive records, and deduplicate entries into a canonical tracker.

Does this data extraction pipeline work with APIs and web pages without manual entry?

The pipeline works with APIs, web pages, and emails without manual entry by using recipe-driven configuration to automate the search and extraction process. It requires YAML recipes to define the specific extraction schemas and deduplication rules needed for reliable data management.

What is the best way to prevent data loss when tracking structured data from fragmented sources?

The best way to prevent data loss from fragmented sources is to use an automated research pipeline that archives and deduplicates extracted data into canonical tracker pages. This ensures consistent record-keeping and prevents the manual tracking errors that cause data loss.