ky-markdown-rebuilder

Rebuild PDFs, PPTs, and scanned images into structured, page-aligned Markdown.

120|15|Updated Jul 3, 2026
One-click install
npx skills add https://github.com/KyrieCheungYep/ky-markdown-rebuilder --skill ky-markdown-rebuilder
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ky-markdown-rebuilder
Source: https://github.com/KyrieCheungYep/ky-markdown-rebuilder/tree/main
Command: npx skills add https://github.com/KyrieCheungYep/ky-markdown-rebuilder --skill ky-markdown-rebuilder

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pdftoppm, pdfinfo, soffice, markitdown, tesseract, PIL, pypdf, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill solves the common issue where standard conversion tools produce jumbled, inaccurate, or layout-broken Markdown when processing visually dense documents like PDFs, slide decks, and scanned reports.

Core Features & Use Cases

  • Visual Reconstruction: Combines text extraction with original-resolution page rendering to ensure layout, tables, and diagrams are accurately preserved.
  • Automated Validation: Uses a built-in output contract and mechanical checkers to ensure page continuity, link integrity, and structural accuracy.
  • Use Case: Use this when you need to convert a complex 50-page business proposal or a scanned technical whitepaper into a clean, searchable, and perfectly aligned Markdown document that maintains the original visual context.

Quick Start

Use the ky-markdown-rebuilder skill to rebuild the provided proposal.pdf into a page-aligned Markdown document with an outline.

Frequently Asked Questions about ky-markdown-rebuilder

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert a scanned PDF to Markdown without losing the original layout?

To convert a scanned PDF to Markdown without losing the layout, you need visual reconstruction tools that combine text extraction with original-resolution page rendering to preserve tables and diagrams accurately.

What is the best way to extract text from complex PDFs and scanned images while maintaining page alignment?

The best way to extract text while maintaining page alignment is to use automated visual inspection alongside text extraction, ensuring the output strictly validates page continuity and structural accuracy against the original document.

Can I use markitdown and tesseract to rebuild complex PDFs into structured Markdown?

Yes, combining markitdown and tesseract with visual analysis allows you to rebuild complex PDFs into structured Markdown, applying automated validation to ensure link integrity and high-fidelity layout preservation.

Does converting PPTs to Markdown preserve visual elements like tables and diagrams?

Converting PPTs to Markdown can preserve visual elements like tables and diagrams when the process uses original-resolution page rendering and visual reconstruction to capture the slide layout accurately.

Why does standard PDF to Markdown conversion produce jumbled text and broken layouts?

Standard PDF to Markdown conversion produces jumbled text and broken layouts because it lacks visual inspection and output validation, failing to map the original document's visual context to a structured format.

How to automate the conversion of a 50-page business proposal into page-aligned Markdown?

To automate converting a 50-page business proposal into page-aligned Markdown, apply a skill that uses automated visual inspection, text extraction, and a strict quality contract to verify page continuity and structural accuracy.