wiktionary-and-wikisource

Parse Wiktionary entries and Wikisource proofreading status via MediaWiki API.

15|6|Updated Feb 17, 2026
One-click install
npx skills add https://github.com/fuzheado/Wikipedia-AI-Skills --skill wiktionary-and-wikisource
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: wiktionary-and-wikisource
Source: https://github.com/fuzheado/Wikipedia-AI-Skills/tree/main/.claude/skills/wiktionary-and-wikisource
Command: npx skills add https://github.com/fuzheado/Wikipedia-AI-Skills --skill wiktionary-and-wikisource

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires wikimedia-api-access, wikimedia-commons, pywikibot, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill streamlines the process of working with Wiktionary (dictionary entries, translation tables, etymologies, audio pronunciations, lexemes) and Wikisource (proofread page workflow, OCR text extraction, quality validation, compiled works), providing tools for efficient content management and analysis.

Core Features & Use Cases

  • Wiktionary Parsing: Parse entry structures, translation tables, and extract pronunciation files.
  • Wikisource Proofreading: Check proofreading progress, extract OCR text, and manage quality levels.
  • Use Case: Imagine you need to review the proofreading progress of a large Wikisource project. Use this Skill to quickly get statistics on the number of pages without text, problematic, proofread, and validated, along with the overall progress percentage.

Quick Start

Use the 'wiktionary-and-wikisource' skill to get the proofreading progress of the work at 'Index:Example Book.pdf'.

Frequently Asked Questions about wiktionary-and-wikisource

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I check proofreading progress and quality validation status for a Wikisource project?

Check Wikisource proofreading progress by retrieving statistics on pages without text, problematic pages, proofread pages, and validated pages, along with the overall completion percentage. This allows you to quickly assess quality levels and manage large text validation workflows.

How do I extract OCR text and dictionary entry structures from Wiktionary and Wikisource?

Extract OCR text and parse Wiktionary entry structures by automating content retrieval with the MediaWiki Action API and Pywikibot. This enables efficient text extraction, translation table parsing, and pronunciation file handling for dictionary and library projects.

Do I need Pywikibot and MediaWiki API access to parse Wiktionary translation tables and etymologies?

Yes, you need Pywikibot and MediaWiki API access to parse Wiktionary translation tables and etymologies. These dependencies provide the necessary programmatic interface for comprehensive data retrieval, manipulation, and Wikimedia Commons file handling required by the parsing process.

What is the best way to manage quality assessment and proofread page workflows on Wikimedia?

The best way to manage proofread page workflows is by automating quality assessment through API-driven text extraction and validation status checks. This approach streamlines the review of compiled works by systematically tracking proofreading progress and identifying problematic pages.

Can I use this Skill to extract audio pronunciations and lexemes from Wiktionary?

Yes, you can extract audio pronunciation files and lexemes from Wiktionary. The Skill automates the parsing of dictionary entry structures, allowing you to efficiently retrieve and manage pronunciation assets alongside translation tables and etymology data.

Why does my Wikisource text extraction return incomplete data for certain scanned PDFs?

Incomplete Wikisource text extraction often occurs when OCR processing fails on problematic pages within the proofread page workflow. Use the Skill to identify pages without text or those marked as problematic, allowing you to target manual validation efforts and resolve quality issues.