document-hunter

Scan public archives and download primary-source documents with Playwright.

408|97|Updated Jan 14, 2026
One-click install
npx skills add https://github.com/bitwize-music-studio/claude-ai-music-skills --skill document-hunter
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: document-hunter
Source: https://github.com/bitwize-music-studio/claude-ai-music-skills/tree/main/skills/document-hunter
Command: npx skills add https://github.com/bitwize-music-studio/claude-ai-music-skills --skill document-hunter

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill streamlines the tedious process of locating and downloading primary source documents from free public archives using browser automation. It reduces manual searching time and ensures reproducible results.

Core Features & Use Cases

  • Systematic search across DocumentCloud, CourtListener, Scribd, Justia, and government sites for case-related documents.
  • Automated downloads of PDFs, transcripts, indictments, and reports with a structured manifest.
  • Metadata organization via a generated manifest and cited sources, enabling quick audits or research workflows.
  • Use Case: A researcher needs all public-domain court documents for a case; run document-hunter to collect and catalog them for review.

Quick Start

Run the document-hunter with the case name or album research context. Example: "document-hunter 'Dorr v. USIA'".

Frequently Asked Questions about document-hunter

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate downloading primary source documents from public archives like CourtListener?

You can automate downloading primary source documents by running browser automation scripts with Playwright that systematically scan CourtListener, Scribd, and DocumentCloud, then generate a structured manifest organizing the complete document dataset.

What is the best way to collect court documents and case-related PDFs in bulk?

The best way to collect court documents in bulk is using automated browser scanning across DOJ, SEC, and Justia sites, which captures indictments and reports while organizing metadata into a reproducible manifest for research audits.

Can I use Playwright to scrape public archives and organize a research manifest?

Yes, Playwright drives the browser automation to navigate free public archives, while Python scripts process the retrieved primary sources and produce a structured manifest and cited sources for your research workflow.

Does automated document hunting work for researching specific case names or album documentation?

Automated document hunting works effectively for both specific case-name research and album documentation by systematically scanning configured archive sites to locate and download relevant public-domain primary source documents.

What sources does automated primary document hunting support?

Automated primary document hunting supports DocumentCloud, CourtListener, Scribd, Justia, and government sites including DOJ and SEC, targeting free public archives to ensure reproducible results without paywalls.

How do I keep track of metadata when downloading archived primary sources?

You keep track of metadata when downloading primary sources through a generated structured manifest that catalogs cited sources and downloaded files, enabling quick audits and seamless integration into research workflows.