Data Engineer with 7+ years building production-grade ETL/ELT pipelines and cloud data warehouses at enterprise scale. Specializes in AI-driven data systems with hands-on experience in autonomous error detection, LLM-powered interfaces, and full-stack delivery from ingestion to dashboards. Proven track record of reducing pipeline runtime by 40% and eliminating 10+ recurring failure points across 50+ databases.
Multi-Agent Orchestration, Autonomous Error Detection, Self-Healing Pipelines, Agent Decision Systems, Failure Recovery, Tool Use & Function Calling
Retrieval-Augmented Generation (RAG), LLM Evaluation Frameworks, Multi-Model Routing & Fallback, Prompt Engineering, Text-to-SQL, Context Window Management
Production AI Monitoring, Pipeline Health Dashboards, Automated Quality Gates, Data Lineage Tracking, Hallucination Mitigation, Latency Optimization
Vector Databases & Semantic Search, Containerized AI Deployment, API Design for LLM Systems, Cloud Data Platforms (Azure, AWS, GCP), DAG Orchestration (Dagster, Airflow)
Python, SQL (MSSQL, Spark SQL, PL/SQL), PySpark, ETL/ELT Pipeline Design, Cloud Data Warehouses (Azure Synapse, Snowflake, Databricks), Data Modeling (Star Schema), DuckDB, Parquet
Docker & Container Registry, CI/CD Pipelines (Jenkins), Kubernetes, Git Version Control, Agile/Scrum (CSM), Infrastructure as Code, Zero-Downtime Deployments