Post Job Free
Sign in

Senior Data Engineer (Healthcare & Lakehouse)

Location:
McKinney, TX
Posted:
July 22, 2026

Contact this candidate

Resume:

Mahima Shrestha

Senior Data Engineer

****************@*****.*** 469-***-**** LinkedIn

PROFESSIONAL SUMMARY

Senior Data Engineer with 5+ years of experience designing and scaling enterprise data platforms across healthcare, financial services, and insurance. Expertise in Apache Spark, PySpark, SQL, Airflow, dbt, and cloud data ecosystems across GCP, AWS, and Azure. Architected Medallion Lakehouse platforms, batch and real-time pipelines, and dimensional data models processing millions of records while improving data reliability, performance, and cost efficiency. Strong healthcare data expertise spanning claims, eligibility, provider, HL7, and FHIR, with a proven ability to transform complex source data into governed, analytics-ready datasets for BI, regulatory reporting, and machine learning.

TECHNICAL SKILLS

Cloud & Data Platforms: AWS (S3, Glue, EMR, Redshift, Step Functions, IAM, Lambda, Kinesis), Azure (Databricks, Data Factory, ADLS Gen2, Azure SQL), GCP (BigQuery, Dataflow, Cloud Storage, Pub/Sub, Cloud Composer, IAM), Snowflake

Data Processing, Streaming & Orchestration: Apache Spark (PySpark, Spark SQL, Structured Streaming), Databricks, Apache Airflow, Kafka, Google Pub/Sub, dbt, Azure Data Factory (ADF), SSIS, CDC (Change Data Capture)

Data Architecture, Modeling & Storage: Medallion Architecture, Lakehouse Architecture, Star Schema, Dimensional Modeling, Data Marts, Incremental Processing, Schema Evolution, Delta Lake, Amazon S3, Azure Data Lake Storage Gen2, Google Cloud Storage

Data Quality, Governance & Security: Great Expectations, Data Reconciliation, Schema Validation, Data Lineage, HIPAA Compliance, RBAC, Column-Level Security, Data Masking, Audit Logging, Data Governance

Programming & Scripting: Python (Pandas, NumPy, Sci-kit learn, Matplotlib), SQL, R, Linux

Analytics & BI: Looker, Tableau, Power BI, Excel

DevOps & Infrastructure: Git, Docker, Kubernetes, Azure DevOps, Terraform

APIs & Healthcare Standards: REST APIs, HL7, FHIR

Development Tools: Linux, Windows, JIRA, Confluence

File Formats: JSON, Parquet, Avro, CSV, XML, Delta Lake

Machine Learning: Feature Engineering, MLOps, Feature stores, Model tuning

PROFESSIONAL EXPERIENCE

Agilon Health (Healthcare/ Value-Based Care) Senior Data Engineer Austin, TX (Remote)

Feb 2024- Present

Architected a Medallion Lakehouse on GCP leveraging BigQuery, Snowflake, dbt, and dimensional modeling to deliver governed healthcare datasets supporting HEDIS, CMS risk adjustment, and population health analytics.

Owned end-to-end ingestion of FHIR R4 and HL7 v2 clinical data streams from 5+ payer and provider systems using Kafka, Pub/Sub, APIs, and batch integrations, enabling scalable ingestion of clinical, claims, and provider data into BigQuery and Snowflake with near real-time availability for downstream analytics.

Built large-scale PySpark transformation frameworks to process payer, provider, and clinical datasets at scale, addressing data quality problems including duplicate member records, invalid diagnosis codes, and inconsistent provider identifiers.

Developed and maintained 120+ dbt models, and a governed semantic layer in Snowflake, which reduced business intelligence query complexity by 50% and improved analyst self-service in Tableau and Looker.

Established automated data quality frameworks using Great Expectations and dbt tests including schema validation, column/object level checks and contract enforcement, preventing bad data releases and improving downstream model accuracy by ~40%.

Reinforced governance using GCP IAM, Cloud KMS, Dataplex, and BigQuery row/column-level security, enforcing least-privilege access, codifying data contracts for PHI/PII datasets, and passing HIPAA/GDPR audits with zero critical findings.

Served as technical lead for workflow orchestration, designing and managing 40+ production Apache Airflow DAGs supporting critical healthcare reporting and operational workloads.

Designed and optimized Snowflake data marts and dimensional models supporting care management and quality reporting, significantly improving query performance across clinical business teams.

Built SLA monitoring, automated alerting, and self-healing recovery workflows that sustained 99%+ pipeline reliability in production, eliminating manual intervention for downstream teams.

Tuned BigQuery slot reservations and Dataflow autoscaling using caching, pruning, and workload isolation, reducing compute costs by ~25% while maintaining performance under peak loads.

Led mentoring and code-review in PySpark, SQL, and cloud data engineering which improved delivery consistency team wide.

Wintrust Financial (Banking & Financial Services) Data Engineer Chicago, IL (Remote)

Nov 2022 – Dec 2023

Built ETL Pipelines on AWS Glue, EMR, Apache Spark, and AWS S3 for transactions, account, and activity data increasing processing speed by 30% and enabling financial and risk analytics.

Designed and implemented an Amazon Redshift data warehouse and Tableau dashboards for near real-time reporting, improving query performance by 25% and reducing reporting turnaround time for business stakeholders.

Built Kafka and Spark Structured Streaming pipelines processing 120K+ daily ACH and wire transaction events with sub-minute latency addressing the gap between batch reporting cycles and real-time fraud signals, enabling faster detection and operational monitoring across banking systems.

Implemented CDC-based incremental loading, automated reconciliation, and data quality validation frameworks that reduced duplicate records by 25% and improved reporting accuracy for SOX-compliant financial reporting and strengthening audit readiness across regulatory reporting workflows.

Published 8+ governed self-service data marts in Snowflake and Amazon Redshift using dimensional data modeling and role-based access control (RBAC), supporting fraud monitoring, treasury operations, compliance reporting, and executive analytics.

Optimized Spark and Redshift workloads through partition pruning, caching, and query tuning, reducing ETL runtimes and cutting average report delivery time from 14 hours to 4 hours improving report availability for downstream risk and finance teams.

Established enterprise data governance standards across S3, Redshift, and Snowflake by implementing IAM-based access controls, encryption using AWS KMS, audit logging, and data classification policies supporting PCI-DSS and SOX compliance requirements.

Built operational monitoring and observability frameworks using Amazon CloudWatch, Airflow SLA monitoring, and automated notification workflows, reducing issue detection time and improving production pipeline reliability.

Partnered with analytics, risk, and compliance stakeholders to translate reporting requirements into scalable data solutions, reducing turnaround time for new reporting initiatives by 30% through reusable pipeline patterns and standardized data models.

Next Insurance (Insurance) Data Engineer Palo Alto, CA (Remote)

Sep 2021 – Oct 2022

Designed and maintained Azure-based data platforms leveraging Databricks, PySpark, ADF, and ADLS Gen2, processing 500K+ daily policy, claims, and customer records supporting underwriting, claims operations, financial reporting, and executive analytics.

Developed ingestion frameworks consolidating 8+ data sources including Azure SQL databases, REST APIs, partner systems, and flat files into a centralized Azure Data Lake, replacing inconsistent point-to-point feeds with a governed, accessible enterprise data layer.

Architected a Delta Lake-based Medallion Lakehouse on Azure, implementing bronze, silver, and gold layers with ACID transactions, schema evolution, and incremental processing to standardize ingestion and transformation across policy, claims, and customer domains.

Built PySpark transformation pipelines to cleanse, standardize, and enrich insurance datasets across claims, policy, and premium domains, reducing downstream reporting discrepancies by 25% and improving trust in data used for carrier and reinsurance reporting.

Optimized Databricks workloads through partitioning, file compaction, Z-Ordering, caching, and Spark performance tuning, reducing pipeline runtimes by 35% and improving dashboard responsiveness for underwriting and finance stakeholders.

Built and maintained 20+ Azure Data Factory pipelines with automated retries, dependency management, and SLA monitoring achieving pipeline reliability across critical insurance reporting workloads.

Implemented automated data quality validation and reconciliation frameworks that proactively detected anomalies before they reached production, reducing pipeline incidents and improving reliability of claims and policy reporting workflows.

Established data governance standards across Databricks and ADLS Gen2 by implementing RBAC, column-level masking, encryption controls, and audit logging to protect sensitive customer data and support regulatory compliance requirements.

Built reusable PySpark transformation libraries and standardized ingestion templates that reduced onboarding time for new data sources a contribution that outlasted the role, becoming the team's standard approach for new integrations.

Partnered with claims, underwriting, and finance stakeholders to translate reporting requirements into data solutions, reducing turnaround time for new reporting requests by 25% through reusable pipeline patterns and self-service data models.

Cedar Gate Technologies (Healthcare) Data Engineer Kathmandu, Nepal

Jun 2020 – Aug 2021

Built and operated healthcare data pipelines processing claims, eligibility, enrollment, and provider datasets using AWS Glue, PySpark, Hadoop, Hive, and EMR, supporting analytics and reporting across healthcare operations teams.

Built batch ingestion pipelines consolidating healthcare data from 5+ sources :relational databases, flat files, and third-party systems into Amazon S3 data lakes, enabling centralized storage and downstream analytics for healthcare operations.

Developed Hadoop and Hive ETL workflows processing 500GB+ of nightly healthcare data, enabling scalable transformation and aggregation of claims and clinical datasets across multiple payer sources.

Developed PySpark, HiveQL, and SQL transformation frameworks to cleanse, standardize, and validate healthcare datasets, reducing data quality issues by approximately 20% across claims and eligibility reporting pipelines.

Designed and maintained curated datasets, dimensional modeling, and reporting tables in Amazon Redshift supporting operational reporting and BI initiatives across claims, eligibility, and provider analytics for healthcare stakeholders.

Implemented reconciliation checks and automated data validation rules across healthcare pipelines, reducing downstream reporting discrepancies by approximately 25% and improving trust in enterprise data assets.

Optimized Hadoop, Hive, and Spark workloads through partitioning, query tuning, and storage strategies, reducing nightly batch processing time by 60% and cutting query runtimes from 4 hours to 90 minutes across 500GB+ healthcare datasets.

Enforced HIPAA-compliant data handling across healthcare pipelines through IAM-based access controls, encryption at rest and in transit, and secure PHI management practices.

Monitored production ETL workflows and investigated pipeline failures in collaboration with senior engineers, improving operational stability and reducing pipeline downtime across critical healthcare reporting workloads.

EDUCATION

Master of Science, Business Analytics University of Central Oklahoma, OK

CERTIFICATIONS

Microsoft Certified : Azure Data Engineer Associate

Databricks Certified : Data Engineer Associate



Contact this candidate