knowledge-collection

Orchestrate automated collection, normalization, and ingestion from internet and enterprise platforms.

45|11|Updated Mar 17, 2026
One-click install
npx skills add https://github.com/beyonai/ByClaw --skill knowledge-collection
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: knowledge-collection
Source: https://github.com/beyonai/ByClaw/tree/main/middleware/openclaw/skills/knowledge-collection
Command: npx skills add https://github.com/beyonai/ByClaw --skill knowledge-collection

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill solves the fragmentation of enterprise information by providing a unified, compliant, and automated pipeline to collect, archive, and ingest data from diverse sources like the public internet, DingTalk, Lark, and WeCom.

Core Features & Use Cases

  • Multi-Source Aggregation: Seamlessly routes collection requests to specialized adapters for internet sites or enterprise platforms.
  • Standardized Ingestion: Enforces a strict collection contract to ensure all gathered data is normalized, sanitized, and ready for knowledge base indexing.
  • Use Case: A project manager needs to collect meeting minutes from DingTalk, research reports from the web, and internal documents from a shared drive; this Skill orchestrates the entire flow from discovery to final knowledge base storage.

Quick Start

Use the knowledge-collection skill to crawl the provided URL and ingest the resulting content into my default knowledge base.

Frequently Asked Questions about knowledge-collection

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate enterprise knowledge collection from multiple platforms like DingTalk and the web?

Automated enterprise knowledge collection routes requests to specialized adapters for internet sites and enterprise platforms like DingTalk, aggregating fragmented information into a unified ingestion pipeline.

What is the best way to normalize and sanitize data ingested from diverse sources into a knowledge base?

Data normalization and sanitization are enforced through a strict collection contract, ensuring all gathered information from diverse sources is consistently formatted and ready for knowledge base indexing.

Does this data ingestion pipeline support multi-tenant security and auditability?

Multi-tenant security and auditability requirements are fully supported, with strict data sanitization and provenance tracking maintained across all organizational data sources.

Can I use web crawling to archive internal documents and external research reports together?

Web crawling and internal document archiving can be orchestrated together, seamlessly collecting research reports from the internet and internal documents from shared drives simultaneously.

What formats are supported for document archiving and ingestion into the knowledge base?

Document archiving enforces a standardized ingestion contract that normalizes diverse organizational data formats into consistent documents ready for downstream knowledge base services.

Why does my knowledge base ingestion fail when routing between different source adapters?

Knowledge base ingestion failures during source adapter routing are typically prevented by enforcing strict data sanitization and provenance tracking before routing to downstream services.