Ivan Shamaev avatar

Ivan Shamaev

Community

@ivanshamaev Β· Yerevan

18Followers
|
43Public Repos
|
159Published Skills

Senior DWH Developer / Ex-BI Team Lead πŸ”– BI Solutions, DWH Tools, Data Tools

Agent Skills by Ivan Shamaev

Showing 159 vetted skills indexed across 1 GitHub repositories.

ivanshamaevivanshamaev
14

dbt-core

Design dbt Core project structures with sources, refs, materializations, and snapshots.

Community
Advanced
ivanshamaevivanshamaev
14

rag-data-pipeline

Build RAG ingestion and query pipelines with hybrid retrieval and reranking.

Community
Advanced
ivanshamaevivanshamaev
14

delta-lake

Generate Delta Lake DDL, DML, and operational commands for transactional storage.

Community
Advanced
ivanshamaevivanshamaev
14

vertica_query_optimization

Diagnose Vertica 11.x EXPLAIN plans and tune projections for query performance.

Community
Advanced
ivanshamaevivanshamaev
14

mlflow-data-pipelines

Configure MLflow tracking and Model Registry for ETL and training runs.

Community
Advanced
ivanshamaevivanshamaev
14

prefect-workflows

Orchestrate Python ETL workflows with Prefect 3 flows and deployments.

Community
Advanced
ivanshamaevivanshamaev
14

sqlmesh

Plans and applies DAT (actually SQLMesh-supported model changes with automatic backfill awareness.

Community
Advanced
ivanshamaevivanshamaev
14

dbt_trino

Configure dbt profiles and incremental models for Trino and Starburst.

Community
Advanced
ivanshamaevivanshamaev
14

sqlfluff

Filters and routes S3 events to multiple destinations with customizable workflows and triggers.

Community
Advanced
ivanshamaevivanshamaev
14

pyspark_etl

Optimize and harden production PySpark ETL pipelines with explicit schemas and explain-based diagnostics.

Community
Advanced
ivanshamaevivanshamaev
14

postgresql-data-engineering

Design PostgreSQL partitioning, indexing, bulk load, and query diagnostics for data engineering.

Community
Advanced
ivanshamaevivanshamaev
14

spark_sql

Optimize and debug production Spark SQL for Hive, lakehouse, and HDFS tables.

Community
Advanced
ivanshamaevivanshamaev
14

de-architecture-decision

Generate Data Engineering Architecture Decision Records with weighted trade-off scoring.

Community
Intermediate
ivanshamaevivanshamaev
14

docker-data-environments

Create Docker build patterns and local compose environments for data engineering tools.

Community
Intermediate
ivanshamaevivanshamaev
14

mage-ai-pipelines

Design Mage AI pipelines with block-level composition for ETL and streaming workflows.

Community
Advanced
ivanshamaevivanshamaev
14

apache-kafka

Configure Apache Kafka topics, producers, consumers, connectors, and monitoring.

Community
Advanced
ivanshamaevivanshamaev
14

kubernetes-data-platform

Deploy Kubernetes-based data platforms for Spark and Airflow workloads.

Community
Advanced
ivanshamaevivanshamaev
14

data_vault_2

Construct Data Vault 2.0 hubs, links, satellites, and PIT tables.

Community
Advanced
ivanshamaevivanshamaev
14

airflow_dag_factory

Generate Apache Airflow DAGs declaratively from YAML using dag-factory.

Community
Advanced
ivanshamaevivanshamaev
14

de-postmortem-writer

Generate blameless data engineering postmortems with timelines, impact, and root-cause chains.

Community
Intermediate
ivanshamaevivanshamaev
14

de-cost-optimization

Analyze data engineering cloud spend across query, compute, storage, and egress drivers.

Community
Advanced
ivanshamaevivanshamaev
14

clickhouse-olap

Design ClickHouse OLAP table engines and query patterns for analytics.

Community
Advanced
ivanshamaevivanshamaev
14

redpanda

Deploy and tune Redpanda clusters with rpk, Schema Registry, and tiered storage.

Community
Advanced
ivanshamaevivanshamaev
14

pyspark-structured-streaming

Build PySpark Structured Streaming pipelines for Kafka, file, and rate sources.

Community
Advanced