Post Job Free
Sign in

AI Data Engineer - Azure RAG Pipelines

Location:
Albany, NY
Posted:
October 05, 2026

Contact this candidate

Resume:

Harish Kumar Yandrapragada

AI Data Engineer Azure Databricks Snowflake dbt PySpark Azure Data Factory Generative AI / RAG

Albany, NY (Open to Relocate) +1-518-***-**** *********@*****.*** linkedin.com/in/harish-yandrapragada-06671a40a

Python • SQL • PySpark • Azure Databricks • Azure Data Factory • Snowflake • dbt • Delta Lake • Apache Iceberg • Kafka • Airflow • Terraform • LangChain • RAG • FAISS/Pinecone • Docker • Kubernetes

PROFESSIONAL SUMMARY

AI Data Engineer with 5+ years of experience designing, building, and optimizing scalable cloud-native ETL/ELT pipelines and AI-enabled data solutions across banking, financial services, and insurance. Hands-on expertise with Python, SQL, PySpark, Apache Spark, Azure Databricks, Azure Data Factory, Snowflake, and Delta Lake for large-scale structured and unstructured data processing, along with modern Lakehouse architecture, data modeling, governance, and quality frameworks.

Experienced developing Generative AI and Retrieval-Augmented Generation (RAG) solutions using Azure OpenAI, GPT models, LangChain, LangGraph, LlamaIndex, and vector databases (FAISS, Pinecone) to power intelligent document search, semantic retrieval, and automated knowledge assistants — including document chunking, embedding generation, vector indexing, prompt engineering, and LLM orchestration. Skilled in batch and streaming pipelines (Kafka, Spark Structured Streaming, Delta Live Tables, Airflow) with CI/CD, monitoring, and automated deployment via Azure DevOps, Git, Docker, and Kubernetes. Recognized for improving pipeline performance, reducing cloud costs, and delivering reliable, high-availability data products in collaboration with cross-functional teams.

PROFESSIONAL EXPERIENCE

AI Data Engineer Apr 2025 – Present

MSIG — Edison, NJ

•Built a production Retrieval-Augmented Generation (RAG) assistant using Azure OpenAI, LangChain, and Pinecone, enabling underwriting and claims teams to run semantic search across 500K+ policy and claims documents and cutting document lookup time from ~20 minutes to under 15 seconds.

•Architected cloud-native ETL/ELT pipelines on Azure Data Factory and Azure Databricks, ingesting 15M+ daily insurance records across policy, claims, and underwriting sources into a Delta Lake Lakehouse with schema enforcement and 99.9% pipeline reliability.

•Engineered document ingestion workflows for chunking, embedding generation, and vector indexing over 500K+ enterprise documents, tuning chunk size and retrieval strategy to improve answer relevance for downstream RAG use cases.

•Tuned PySpark transformations and Databricks cluster configurations — optimizing partitioning, caching, and autoscaling — to cut distributed job runtime by 40% and lower compute spend by 25%.

•Established CI/CD and observability with Azure DevOps, Git, and Docker, automating multi-environment deployments and embedding data quality, validation, and monitoring checks into every pipeline run.

•Provisioned and version-controlled Azure Databricks, Data Factory, and storage infrastructure as code using Terraform, standardizing dev/test/prod environment setup and eliminating manual configuration drift.

•Implemented data lineage and automated quality checks using OpenLineage and Great Expectations across RAG ingestion pipelines, improving traceability and catching schema drift before it reached production.

•Collaborated with data scientists, ML engineers, and business stakeholders to translate ambiguous requirements into secure, well-documented, production-ready AI data products.

Data Engineer May 2024 – Apr 2025

TD Bank — NJ

•Developed scalable ETL/ELT pipelines with Azure Data Factory, Databricks, PySpark, and Snowflake, consolidating banking data from 20+ source systems into a centralized analytics platform serving reporting and data science teams.

•Implemented near-real-time streaming ingestion using Kafka, Spark Structured Streaming, and Delta Live Tables, reducing end-to-end data latency from 4+ hours to under 10 minutes and enabling fresher inputs for fraud monitoring.

•Built modular, version-controlled SQL transformation models in dbt on top of Snowflake, replacing ad hoc scripts and adding automated testing that improved transformation reliability across the analytics layer.

•Curated feature datasets and ML-ready pipelines for fraud-detection and credit-risk models, standardizing and validating inputs to improve model data coverage by 30%.

•Modeled a Medallion (Bronze/Silver/Gold) Lakehouse, structuring raw, cleansed, and curated layers to power consistent downstream analytics and regulatory reporting.

•Adopted Apache Iceberg alongside Delta Lake for select high-volume datasets and enforced governance and access policies using Unity Catalog and Azure Purview across the Lakehouse.

•Leveraged Snowflake Cortex AI functions and Snowpark (Python) to prototype in-database summarization and classification on unstructured transaction and dispute notes, reducing manual triage effort.

•Enforced data validation, lineage tracking, and Role-Based Access Control (RBAC) to meet financial reporting and compliance standards, cutting data incidents by 45%.

•Automated CI/CD for Databricks notebooks and dbt jobs using GitHub Actions, adding automated testing gates before promotion to production.

•Partnered with analysts, data scientists, and business teams to gather requirements and turn them into reliable, production-grade data products.

Data Engineer Jan 2021 – Sep 2023

Bajaj Finserv — India

•Built end-to-end data pipelines in Python, SQL, and PySpark, integrating 15+ heterogeneous financial-services sources into a centralized enterprise data warehouse for analytics and reporting.

•Migrated legacy on-premises ETL workloads to a cloud-native Azure architecture, re-engineering batch jobs to improve scalability and reduce infrastructure costs by 35%.

•Designed dimensional data models and star-schema warehouse structures that streamlined analytics and reporting for lending and risk business teams.

•Orchestrated batch and incremental workflows with Apache Airflow, adding scheduling, retries, monitoring, and alerting that improved pipeline SLA adherence.

•Set up CI/CD pipelines using Jenkins and Git to automate build, test, and deployment of ETL jobs, shortening release cycle time and reducing manual deployment errors.

•Introduced automated data quality and validation checks across pipeline stages, catching 90%+ of data issues before they reached production.

•Adopted GitHub Copilot to accelerate PySpark and SQL script development and speed up code review, improving team productivity on recurring pipeline patterns.

•Accelerated Spark and SQL workloads through query, join, and partition tuning, reducing average processing time by 30%.

TECHNICAL SKILLS

Languages: Python, SQL, PySpark, Bash

Big Data & Streaming: Apache Spark, PySpark, Spark Structured Streaming, Apache Kafka, Apache Flink, Delta Live Tables, Delta Lake

Cloud & Data Platforms: Microsoft Azure (Databricks, Data Factory, Synapse Analytics, Blob Storage), Snowflake, AWS (S3, EMR, Redshift), Google Cloud Platform (BigQuery)

Data Warehousing & Architecture: Lakehouse Architecture, Medallion (Bronze/Silver/Gold), Star Schema, Data Modeling, Apache Iceberg, Data Governance

Generative AI & LLM: Azure OpenAI, OpenAI GPT, LangChain, LangGraph, LlamaIndex, RAG pipelines, prompt engineering, LLM orchestration, Snowflake Cortex AI, Snowpark ML

Vector Databases & Embeddings: FAISS, Pinecone, ChromaDB, document chunking, embedding generation, vector indexing, semantic retrieval

Orchestration & ETL: Apache Airflow, Azure Data Factory, dbt (Data Build Tool), ETL/ELT pipeline design

Infrastructure as Code: Terraform

Data Quality & Governance: Data validation, quality frameworks, OpenLineage, Great Expectations, Azure Purview, Unity Catalog, lineage tracking, Role-Based Access Control (RBAC)

DevOps & CI/CD: Azure DevOps, GitHub Actions, Jenkins, Git, Docker, Kubernetes, GitHub Copilot

BI & Visualization: Tableau, Power BI

EDUCATION

M.S. in Computer Science University at Albany, SUNY Albany, NY Aug 2023 – May 2025

B.Tech in Computer Science SRM University India July 2018 – May 2022

CERTIFICATIONS

Microsoft Certified: Azure Data Engineer Associate

Databricks Certified Data Engineer

Snowflake SnowPro

AWS Certified Data Engineer



Contact this candidate