content-discovery

Automate content discovery across ArXiv, GitHub, and HuggingFace with semantic filtering and publishing.

1|Updated Oct 21, 2025
One-click install
npx skills add https://github.com/longkeyy/claude-discover --skill content-discovery
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: content-discovery
Source: https://github.com/longkeyy/claude-discover/tree/main/.claude/skills/content-discovery
Command: npx skills add https://github.com/longkeyy/claude-discover --skill content-discovery

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill eliminates the tedious, manual work of staying updated across multiple data sources like academic papers, open-source projects, and AI models. It automates the entire content discovery, analysis, and publishing workflow, saving you hours of research and content creation.

Core Features & Use Cases

  • AI-Powered Discovery: Automatically scan diverse data sources (ArXiv, GitHub, HuggingFace) for content relevant to your configured topics and keywords.
  • Smart Filtering & Deduplication: Leverage AI for semantic filtering and deduplication, ensuring you only receive high-quality, unique, and highly relevant content.
  • Multi-Channel Publishing: Generate high-quality Chinese summaries and insights, then automatically publish them to your Hexo blog, Telegram channel, or Discord server.
  • Use Case: An academic researcher can automatically track the latest AI papers from ArXiv, receive AI-generated summaries, and publish them to their blog without manual effort. An AI engineer can discover new LLM models and datasets from HuggingFace, filtered by quality, and share them with their team.

Quick Start

Create your first discovery task (interactive)

/mission add

Execute the discovery for all enabled tasks

/discover

Frequently Asked Questions about content-discovery

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate content discovery across ArXiv, GitHub, and HuggingFace?

Automate content discovery by configuring data sources and filtering rules, then use the /discover command to scan ArXiv, GitHub, and HuggingFace for relevant content based on your keywords. The Skill applies semantic filtering and deduplication to surface only high-quality, unique results matching your topics.

Can I deduplicate and filter academic papers and open-source projects using AI similarity?

Yes. The Skill uses AI-driven semantic filtering with a 90% similarity threshold to deduplicate content across sources. It identifies and removes near-duplicate papers, projects, and models, ensuring your discovery results contain only distinct, relevant items.

How do I publish discovered content to Hexo, Telegram, and Discord automatically?

After discovery completes, the Skill generates summaries and structured outputs with metadata and images, then publishes directly to your configured channels—Hexo blog, Telegram, or Discord. Multi-channel publishing happens in one workflow without manual steps.

What setup do I need before running discovery tasks?

Create discovery missions using /mission add to define your topics, data sources, and publishing channels via YAML/JSON rules. Once configured, execute /discover to run all enabled tasks. No external dependencies are required.

Does this work for tracking AI models and datasets on HuggingFace?

Yes. The Skill discovers new LLM models and datasets from HuggingFace filtered by your keyword rules and quality thresholds. It extracts metadata, applies semantic deduplication, and publishes findings to your chosen channels with provenance tracking.

Can I generate Chinese summaries of discovered content?

Yes. The Skill generates AI-driven Chinese summaries alongside English metadata for all discovered content. Summaries and insights are published with full provenance metadata, supporting multilingual research workflows.