Tharun Reddy
Dallas, TX +1-623-***-**** **************@*****.*** LinkedIn
Senior Data Engineer Cloud Data Platforms Databricks Snowflake
PROFESSIONAL SUMMARY
Senior Data Engineer with 7+ years of experience building and modernizing cloud data platforms across insurance,
healthcare, and financial-services environments. Built and supported data pipelines, cloud data warehouses,
lakehouse platforms, streaming workflows, data-quality controls, and production DataOps processes using Databricks,
Snowflake, Azure, AWS, Python, SQL, PySpark/Spark, Kafka, ETL/ELT, CDC, dimensional modeling, CI/CD, and
infrastructure as code. Progressed from Hadoop and warehouse engineering to cloud lakehouse,streaming, and
production data-platform ownership.
TECHNICAL SKILLS
Programming & Distributed Processing: Python, SQL, PySpark, Apache Spark, Scala, Bash/Shell, Apache Hive, HDFS
Cloud & Data Platforms: Azure Data Factory (ADF), ADLS Gen2, Azure Databricks, Azure Synapse, Snowflake; AWS S3,
Lambda, Glue, Redshift, Athena, EMR, Step Functions, EventBridge
Pipelines, Orchestration & Streaming: ETL/ELT, Apache Kafka, Kafka Streams, Azure Event Hubs, CDC, Apache Airflow,
dbt, Databricks Workflows/Jobs, Spark Structured Streaming
Data Architecture, Quality & Governance: Delta Lake, Lakehouse, Medallion Architecture, Dimensional Modeling,
Star Schema, Fact/Dimension Tables, SCD, Great Expectations, Unity Catalog, Data Lineage, HIPAA,GDPR
DevOps, Security & Observability: Git, GitHub Actions, CI/CD, Terraform, Docker, CloudWatch, Azure Monitor,IAM,
RBAC, Secrets Manager, Key Vault, Purview, REST APIs, OAuth 2.0
AI Engineering: FastAPI, OpenAI/Azure OpenAI APIs, RAG, Embeddings, LangChain, LangGraph, pgvector,Pinecone,
Hybrid Search, Reranking, Agentic AI, Tool Calling, MCP, LangSmith/Langfuse
PROFESSIONAL EXPERIENCE
MetLife Insurance — New York, NY Jan 2024 – Present
Senior Data Engineer
• Built and operated an enterprise cloud data platform across Azure Data Factory, ADLS Gen2, Azure
Databricks/PySpark, Delta Lake, Snowflake, dbt, and downstream analytics, integrating policy, claims,
customer, billing, and financial-reporting data across batch and streaming workloads.
• Built Databricks lakehouse pipelines using PySpark, Auto Loader, Databricks Jobs/Workflows, and Delta Lake,
organizing curated datasets across Bronze, Silver, and Gold layers for reusable analytics and warehouse
consumption.
• Engineered real-time and near-real-time ingestion flows using Apache Kafka, Azure Event Hubs, Spark
Structured Streaming, and Change Data Capture, adding production error handling, retry, logging, and
monitoring across distributed data flows.
• Developed Snowflake ingestion and transformation workflows using Snowpipe, Streams, Tasks, Dynamic
Tables, Time Travel, and RBAC, applying SQL tuning, clustering, and warehouse optimization to improve
performance and control cloud costs.
• Orchestrated enterprise ETL/ELT workflows with ADF, Apache Airflow, and dbt, implementing incremental
loads, dependency management, automated scheduling, validation, and controlled CI/CD promotion.
• Implemented data-quality and governance controls using Great Expectations, Unity Catalog, Purview, data
lineage, and role-based access controls to strengthen trusted data delivery and HIPAA/GDPR-aligned
governance.
• Developed AWS serverless data workflows using S3, Lambda, EventBridge, Step Functions, CloudWatch, IAM,
and Secrets Manager to automate integrations, monitoring, secure processing, and cross-cloud data
workflows.
• Automated infrastructure and deployment workflows with Terraform, Git, GitHub Actions, CI/CD, and Docker;
resolved production pipeline and cloud-service issues while supporting reporting dependencies and
mentoring junior engineers and analysts.
Change Healthcare — Nashville, TN Feb 2022 – Dec 2023
Data Engineer / Big Data Engineer
• Built healthcare data modernization pipelines that moved Oracle, SQL Server, REST API, and flat-file data
through Azure Data Factory and CDC into ADLS Gen2, Azure Databricks/PySpark, Delta Lake, and Snowflake
for claims, provider, member, enrollment, and eligibility analytics.
• Developed Spark/PySpark transformations in Databricks and Delta Lake to cleanse, standardize, and curate
healthcare datasets before loading reporting-ready structures into Snowflake.
• Implemented incremental-loading and CDC workflows with ADF and Kafka, combining Python/SQL validation,
reconciliation, and Great Expectations checks to improve data completeness, accuracy, consistency, and
HIPAA-oriented quality controls.
• Built Snowflake dimensional models using Star Schema, fact/dimension structures, and SCD Type 1/2
patterns, with dbt transformations supporting claims analytics, provider performance, operational reporting,
and self-service BI.
• Orchestrated batch and streaming workflows using ADF and Apache Airflow, integrating secure OAuth 2.0
REST APIs and Kafka-based data flows while managing scheduling, dependencies, production
troubleshooting, and data lineage.
• Supported Azure and Snowflake modernization by optimizing SQL and warehouse structures, maintaining Gitbased delivery practices, and collaborating with architects, analysts, and healthcare stakeholders on regulated
data-platform requirements.
Aricent Technologies — Bangalore, India Apr 2019 – Jan 2022
Data Engineer
• Developed Python/PySpark ETL scripts and Hive/HDFS batch pipelines to ingest financial transaction data
from Oracle, flat files, and REST APIs, automating 10+ manual data-movement steps and improving daily
ingestion reliability.
• Built SQL stored procedures, views, Spark/Hive transformations, validation, reconciliation, and reporting
workflows that processed 2M+ financial transactions per day for enterprise data-warehouse reporting and
production-release UAT.
• Designed dimensional data models using Star Schema, fact/dimension tables, and SCD patterns, and
documented source-to-target mappings, transformation logic, and data dictionaries for 15+ reporting
datasets.
• Contributed to Azure Data Factory cloud-migration pipelines for 12+ financial data feeds by building
parameterized linked services and transformations into Azure storage and validating end-to-end migrated
workloads with zero data loss.
• Established Git/GitHub version control, Linux/Shell automation, modular ETL patterns, and documentation
standards across pipeline codebases, reducing analyst onboarding time from 3 weeks to 1 week.
TECHNICAL PROJECTS
Enterprise RAG & Agentic AI Data Platform GitHub Project
• Built a Python-based enterprise document and data-ingestion flow covering chunking, embeddings, vector
storage with pgvector/Pinecone, hybrid search, and reranking to support retrieval-augmented generation
(RAG).
• Developed RAG and agent workflows using OpenAI/Azure OpenAI APIs, LangChain, and LangGraph, adding
tool calling for retrieval, SQL queries, APIs, and Python-based analysis through an agentic orchestration layer.
• Exposed AI workflows through FastAPI, containerized services with Docker, and implemented GitHub Actions
CI/CD with tracing and observability for LLM and retrieval workflows.
EDUCATION
Master of Science in Computer Science — Campbellsville University, KY May 2022 – Oct 2023
Bachelor of Engineering in Computer Science — JNTUH, Hyderabad, India Aug 2015 – May 2019