What problem does it solve?
VC firms scatter team bios, portfolio companies, and brand assets across multiple sub-pages and PDFs. Manually collecting headshots, LinkedIn URLs, logos, and titles for investment decks and memos is tedious and error-prone. This Skill automates the entire extraction and normalization pipeline.
Core Features & Use Cases
- Four-Checkpoint Cascade: Systematically extracts VC team members, external advisors, portfolio companies, and portfolio CEOs with human-confirmation gates at each stage.
- Dual Anchor Types: Supports both firm-anchored walks (one VC → team → portfolio → CEOs) and company-anchored credibility-card walks (operating company → backers → backer teams + portfolios).
- SVG-First Brand Asset Pipeline: Retrieves logos via a seven-tier cascade (inline SVG → site paths → press kits → Brandfetch → vector repos → Google CSE → raster fallback) with automatic background stripping and validation.
- Cross-Tool Fallbacks: Chains Jina Reader, Firecrawl, Tavily, OpenGraph.io, and Brandfetch with intelligent escalation and global caching to minimize cost and handle JS-gated sites.
Quick Start
Use the crawl-fetch-ingest skill to fill in the team and portfolio metadata for Sequoia Capital from their website and the attached deck PDF.