Ivan Shamaev
Community@ivanshamaev Β· Yerevan
Senior DWH Developer / Ex-BI Team Lead π BI Solutions, DWH Tools, Data Tools
Agent Skills by Ivan Shamaev
Showing 159 vetted skills indexed across 1 GitHub repositories.
dbt-core
Design dbt Core project structures with sources, refs, materializations, and snapshots.
rag-data-pipeline
Build RAG ingestion and query pipelines with hybrid retrieval and reranking.
delta-lake
Generate Delta Lake DDL, DML, and operational commands for transactional storage.
vertica_query_optimization
Diagnose Vertica 11.x EXPLAIN plans and tune projections for query performance.
mlflow-data-pipelines
Configure MLflow tracking and Model Registry for ETL and training runs.
prefect-workflows
Orchestrate Python ETL workflows with Prefect 3 flows and deployments.
sqlmesh
Plans and applies DAT (actually SQLMesh-supported model changes with automatic backfill awareness.
dbt_trino
Configure dbt profiles and incremental models for Trino and Starburst.
sqlfluff
Filters and routes S3 events to multiple destinations with customizable workflows and triggers.
pyspark_etl
Optimize and harden production PySpark ETL pipelines with explicit schemas and explain-based diagnostics.
postgresql-data-engineering
Design PostgreSQL partitioning, indexing, bulk load, and query diagnostics for data engineering.
spark_sql
Optimize and debug production Spark SQL for Hive, lakehouse, and HDFS tables.
de-architecture-decision
Generate Data Engineering Architecture Decision Records with weighted trade-off scoring.
docker-data-environments
Create Docker build patterns and local compose environments for data engineering tools.
mage-ai-pipelines
Design Mage AI pipelines with block-level composition for ETL and streaming workflows.
apache-kafka
Configure Apache Kafka topics, producers, consumers, connectors, and monitoring.
kubernetes-data-platform
Deploy Kubernetes-based data platforms for Spark and Airflow workloads.
data_vault_2
Construct Data Vault 2.0 hubs, links, satellites, and PIT tables.
airflow_dag_factory
Generate Apache Airflow DAGs declaratively from YAML using dag-factory.
de-postmortem-writer
Generate blameless data engineering postmortems with timelines, impact, and root-cause chains.
de-cost-optimization
Analyze data engineering cloud spend across query, compute, storage, and egress drivers.
clickhouse-olap
Design ClickHouse OLAP table engines and query patterns for analytics.
redpanda
Deploy and tune Redpanda clusters with rpk, Schema Registry, and tiered storage.
pyspark-structured-streaming
Build PySpark Structured Streaming pipelines for Kafka, file, and rate sources.