data-engineer

Automate data pipeline cleaning and format conversion to Markdown and JSON.

Updated May 5, 2025
One-click install
npx skills add https://github.com/yopitek/Obsidian --skill data-engineer-yopitek
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/yopitek/Obsidian/tree/main/HQ/05_Agents/09_Data_Engineer
Command: npx skills add https://github.com/yopitek/Obsidian --skill data-engineer-yopitek

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data engineering tasks such as data cleaning, format conversion, and maintenance of vector databases are tedious and error-prone without automation.

Core Features & Use Cases

  • Automates data cleaning pipelines and format transformation (PDF/Excel/HTML/JSON to Markdown/structured data)
  • Maintains RAG-ready vector databases and supports version control for data assets
  • Supports output routing to Obsidian vault, JSON files, and vector stores for retrieval-augmented generation scenarios

Quick Start

Initialize the data-engineer workflow on a dataset and run the pipeline to generate cleaned Markdown and ready vector embeddings.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate data cleaning and PDF to Markdown conversion for RAG workflows?

Automate data cleaning and PDF to Markdown conversion by initializing a data pipeline that transforms raw files into structured formats and generates vector embeddings for retrieval-augmented generation workflows.

What is the best way to maintain a RAG-ready vector database with version control?

Maintain a RAG-ready vector database by running automated pipelines that clean data assets, apply version control, and route outputs directly to vector stores for retrieval-augmented generation.

Can I route converted data outputs to Obsidian and JSON files automatically?

You can route converted data outputs to Obsidian vaults and JSON files automatically, alongside vector database maintenance, satisfying end-to-end workflow requirements with included logging.

Does the data pipeline support Excel and HTML format conversion to structured data?

The data pipeline supports format conversion for Excel and HTML files, transforming them into Markdown or structured JSON data while performing data cleansing and validation.

How do I build a data pipeline for backup, validation, and logging across diverse sources?

Build a data pipeline by initializing the workflow on a dataset to automate backup, conversion, cleaning, validation, and routing outputs across diverse sources with comprehensive logging.