What problem does it solve?
This skill automates extracting high-quality figures from research papers by preferentially pulling images from arXiv source packages, reducing noise from logos and iconography and organizing outputs for note-taking.
Core Features & Use Cases
- Prioritized arXiv source extraction: locate and copy genuine figures from typical image directories like pics/, figures/, figs/, images/, or root image files, then fall back to extracting figures from embedded PDFs when necessary.
- PDF-based extraction as fallback: robustly extract images from PDFs when arXiv sources are unavailable or incomplete, with filtering to remove small logos and decorative elements.
- Structured output and indexing: save images to 20_Research/Papers/[领域]/[论文标题]/images/ and generate an index file detailing image sources, sizes, and formats for easy reuse in notes.
- Automated quick-start workflow: a single command to process a paper by ID or local PDF, producing image assets and a reference index.
Quick Start
Provide an arXiv ID or a local PDF path and run the extractor to save all usable figures to the target images directory and generate an index.