Post Job Free
Sign in

Senior Data Engineer - Databricks Snowflake

Location:
Los Angeles, CA
Posted:
October 07, 2026

Contact this candidate

Resume:

Tharun Reddy

Dallas, TX +1-623-***-**** **************@*****.*** LinkedIn

Senior Data Engineer Cloud Data Platforms Databricks Snowflake

PROFESSIONAL SUMMARY

Senior Data Engineer with 7+ years of experience building and modernizing cloud data platforms across insurance,

healthcare, and financial-services environments. Built and supported data pipelines, cloud data warehouses,

lakehouse platforms, streaming workflows, data-quality controls, and production DataOps processes using Databricks,

Snowflake, Azure, AWS, Python, SQL, PySpark/Spark, Kafka, ETL/ELT, CDC, dimensional modeling, CI/CD, and

infrastructure as code. Progressed from Hadoop and warehouse engineering to cloud lakehouse,streaming, and

production data-platform ownership.

TECHNICAL SKILLS

Programming & Distributed Processing: Python, SQL, PySpark, Apache Spark, Scala, Bash/Shell, Apache Hive, HDFS

Cloud & Data Platforms: Azure Data Factory (ADF), ADLS Gen2, Azure Databricks, Azure Synapse, Snowflake; AWS S3,

Lambda, Glue, Redshift, Athena, EMR, Step Functions, EventBridge

Pipelines, Orchestration & Streaming: ETL/ELT, Apache Kafka, Kafka Streams, Azure Event Hubs, CDC, Apache Airflow,

dbt, Databricks Workflows/Jobs, Spark Structured Streaming

Data Architecture, Quality & Governance: Delta Lake, Lakehouse, Medallion Architecture, Dimensional Modeling,

Star Schema, Fact/Dimension Tables, SCD, Great Expectations, Unity Catalog, Data Lineage, HIPAA,GDPR

DevOps, Security & Observability: Git, GitHub Actions, CI/CD, Terraform, Docker, CloudWatch, Azure Monitor,IAM,

RBAC, Secrets Manager, Key Vault, Purview, REST APIs, OAuth 2.0

AI Engineering: FastAPI, OpenAI/Azure OpenAI APIs, RAG, Embeddings, LangChain, LangGraph, pgvector,Pinecone,

Hybrid Search, Reranking, Agentic AI, Tool Calling, MCP, LangSmith/Langfuse

PROFESSIONAL EXPERIENCE

MetLife Insurance — New York, NY Jan 2024 – Present

Senior Data Engineer

• Built and operated an enterprise cloud data platform across Azure Data Factory, ADLS Gen2, Azure

Databricks/PySpark, Delta Lake, Snowflake, dbt, and downstream analytics, integrating policy, claims,

customer, billing, and financial-reporting data across batch and streaming workloads.

• Built Databricks lakehouse pipelines using PySpark, Auto Loader, Databricks Jobs/Workflows, and Delta Lake,

organizing curated datasets across Bronze, Silver, and Gold layers for reusable analytics and warehouse

consumption.

• Engineered real-time and near-real-time ingestion flows using Apache Kafka, Azure Event Hubs, Spark

Structured Streaming, and Change Data Capture, adding production error handling, retry, logging, and

monitoring across distributed data flows.

• Developed Snowflake ingestion and transformation workflows using Snowpipe, Streams, Tasks, Dynamic

Tables, Time Travel, and RBAC, applying SQL tuning, clustering, and warehouse optimization to improve

performance and control cloud costs.

• Orchestrated enterprise ETL/ELT workflows with ADF, Apache Airflow, and dbt, implementing incremental

loads, dependency management, automated scheduling, validation, and controlled CI/CD promotion.

• Implemented data-quality and governance controls using Great Expectations, Unity Catalog, Purview, data

lineage, and role-based access controls to strengthen trusted data delivery and HIPAA/GDPR-aligned

governance.

• Developed AWS serverless data workflows using S3, Lambda, EventBridge, Step Functions, CloudWatch, IAM,

and Secrets Manager to automate integrations, monitoring, secure processing, and cross-cloud data

workflows.

• Automated infrastructure and deployment workflows with Terraform, Git, GitHub Actions, CI/CD, and Docker;

resolved production pipeline and cloud-service issues while supporting reporting dependencies and

mentoring junior engineers and analysts.

Change Healthcare — Nashville, TN Feb 2022 – Dec 2023

Data Engineer / Big Data Engineer

• Built healthcare data modernization pipelines that moved Oracle, SQL Server, REST API, and flat-file data

through Azure Data Factory and CDC into ADLS Gen2, Azure Databricks/PySpark, Delta Lake, and Snowflake

for claims, provider, member, enrollment, and eligibility analytics.

• Developed Spark/PySpark transformations in Databricks and Delta Lake to cleanse, standardize, and curate

healthcare datasets before loading reporting-ready structures into Snowflake.

• Implemented incremental-loading and CDC workflows with ADF and Kafka, combining Python/SQL validation,

reconciliation, and Great Expectations checks to improve data completeness, accuracy, consistency, and

HIPAA-oriented quality controls.

• Built Snowflake dimensional models using Star Schema, fact/dimension structures, and SCD Type 1/2

patterns, with dbt transformations supporting claims analytics, provider performance, operational reporting,

and self-service BI.

• Orchestrated batch and streaming workflows using ADF and Apache Airflow, integrating secure OAuth 2.0

REST APIs and Kafka-based data flows while managing scheduling, dependencies, production

troubleshooting, and data lineage.

• Supported Azure and Snowflake modernization by optimizing SQL and warehouse structures, maintaining Gitbased delivery practices, and collaborating with architects, analysts, and healthcare stakeholders on regulated

data-platform requirements.

Aricent Technologies — Bangalore, India Apr 2019 – Jan 2022

Data Engineer

• Developed Python/PySpark ETL scripts and Hive/HDFS batch pipelines to ingest financial transaction data

from Oracle, flat files, and REST APIs, automating 10+ manual data-movement steps and improving daily

ingestion reliability.

• Built SQL stored procedures, views, Spark/Hive transformations, validation, reconciliation, and reporting

workflows that processed 2M+ financial transactions per day for enterprise data-warehouse reporting and

production-release UAT.

• Designed dimensional data models using Star Schema, fact/dimension tables, and SCD patterns, and

documented source-to-target mappings, transformation logic, and data dictionaries for 15+ reporting

datasets.

• Contributed to Azure Data Factory cloud-migration pipelines for 12+ financial data feeds by building

parameterized linked services and transformations into Azure storage and validating end-to-end migrated

workloads with zero data loss.

• Established Git/GitHub version control, Linux/Shell automation, modular ETL patterns, and documentation

standards across pipeline codebases, reducing analyst onboarding time from 3 weeks to 1 week.

TECHNICAL PROJECTS

Enterprise RAG & Agentic AI Data Platform GitHub Project

• Built a Python-based enterprise document and data-ingestion flow covering chunking, embeddings, vector

storage with pgvector/Pinecone, hybrid search, and reranking to support retrieval-augmented generation

(RAG).

• Developed RAG and agent workflows using OpenAI/Azure OpenAI APIs, LangChain, and LangGraph, adding

tool calling for retrieval, SQL queries, APIs, and Python-based analysis through an agentic orchestration layer.

• Exposed AI workflows through FastAPI, containerized services with Docker, and implemented GitHub Actions

CI/CD with tracing and observability for LLM and retrieval workflows.

EDUCATION

Master of Science in Computer Science — Campbellsville University, KY May 2022 – Oct 2023

Bachelor of Engineering in Computer Science — JNTUH, Hyderabad, India Aug 2015 – May 2019



Contact this candidate