Post Job Free
Sign in

Senior Data Engineer - Python, PySpark, Snowflake

Location:
Austin, TX
Salary:
120000
Posted:
October 07, 2026

Contact this candidate

Resume:

SRIVIDYA NAROJU

Senior Data Engineer Python PySpark Snowflake Databricks AWS/Azure

737-***-**** ************@*****.*** LinkedIn

PROFESSIONAL SUMMARY

Senior Data Engineer with 8+ years of experience designing and developing scalable data pipelines, data platforms, and ETL/ELT solutions across cloud and enterprise environments.

Strong hands-on expertise in Python, SQL, PySpark, Apache Spark, Snowflake, Databricks, and Azure/AWS, with experience building reliable batch and real-time data processing solutions.

Experienced in developing Lakehouse and data warehouse architectures, including Delta Lake, Medallion Architecture, data modelling, schema evolution, and performance optimisation for large-scale datasets.

Proven experience building and orchestrating data pipelines using Apache Airflow, Azure Data Factory, Kafka, dbt, and cloud-native services, supporting both batch and streaming workloads.

Strong background in data quality, validation, governance, security, and production reliability, including monitoring, automated testing, access controls, and pipeline optimisation.

Experienced working across healthcare, enterprise analytics, and transportation environments, translating business requirements into scalable and production-ready data solutions.

Effective collaborator with data architects, analysts, and business stakeholders, translating complex technical designs into clear, actionable guidance for cross-functional teams.

Progressed from foundational data engineering roles to senior-level ownership of end-to-end pipeline architecture, mentoring engineers and promoting best practices in code quality and production reliability.

TECHNICAL SKILLS

Programming & Query Languages: Python, SQL, Java, Scala, PL/SQL

Data Engineering & Big Data: Apache Spark, PySpark, Apache Kafka, Kafka Streams, Spark Streaming, Hadoop, ETL/ELT, Batch Processing, Real-Time Data Processing

Cloud & Data Platforms: Snowflake, Databricks, AWS (S3, EC2, Lambda, RDS, Redshift), Azure (Data Factory, Blob Storage, Synapse, Functions), GCP, Amazon Redshift, Google BigQuery

Data Transformation & Modelling: dbt, Data Modelling, Dimensional Modelling, Delta Lake, Medallion Architecture, Apache Iceberg, Apache Hudi, Schema Evolution, Partitioning, Clustering

Orchestration & Data Quality: Apache Airflow, Metadata-Driven Pipelines, Data Validation, Data Quality Frameworks, Data Reconciliation

Databases: PostgreSQL, MySQL, SQL Server, Oracle, MongoDB, Cassandra, Cosmos DB, Epic EHR

DevOps, Infrastructure & Monitoring: Terraform, Docker, Kubernetes, Jenkins, GitHub Actions, CI/CD, Prometheus, Grafana, Git

PROFESSIONAL EXPERIENCE

Senior Data Engineer Office of Mental Health (Healthcare) New York, NY

01/2025 – Present

Project Description: Modernising and centralising clinical data infrastructure to support healthcare analytics and reporting. The project integrates data from multiple clinical and enterprise source systems, processing high-volume record and event data across cloud-based analytics platforms while maintaining HIPAA-compliant security and production reliability.

Owned the design and development of real-time clinical data ingestion pipelines using Apache Kafka, PySpark, Azure Data Lake, and Delta Lake, processing high-volume clinical records and events through a Bronze/Silver/Gold Medallion architecture.

Extracted and ingested clinical data from Epic EHR and SQL Server-based enterprise systems, standardising source data for downstream cloud analytics processing.

Led the modernisation of legacy cross-system workloads through Azure Data Factory and Databricks, converting existing processes into scalable, metadata-driven pipelines and meaningfully improving overall processing efficiency.

Architected and enforced HIPAA-compliant data masking, tokenisation, and RBAC controls across Snowflake and Delta Lake environments to protect sensitive PHI across multiple processing layers.

Designed and delivered REST API integration frameworks to securely exchange structured datasets between multiple clinical and enterprise applications and cloud analytics platforms.

Led performance optimisation of Snowflake workloads through query tuning, warehouse right-sizing, clustering strategies, and Snowpipe loading patterns, significantly improving query response times.

Drove automation of data platform deployments using Terraform, Jenkins, GitHub Actions, and Kubernetes, supporting consistent CI/CD delivery across development, testing, and production environments.

Implemented production monitoring with Prometheus and Grafana, strengthening pipeline observability and supporting consistently high production pipeline availability.

Environment: Azure Data Factory, Azure Data Lake, Delta Lake, Databricks, Snowflake, Apache Kafka, PySpark, Python, SQL, REST APIs, Terraform, Kubernetes, Jenkins, GitHub Actions, Prometheus, Grafana

Senior Data Engineer Versant Health (Healthcare) Linthicum Heights, MD

01/2024 – 12/2024

Project Description: Developed and enhanced a healthcare data platform supporting claims and member eligibility data across multiple operational data sources and a large-scale healthcare record base. The platform combined Lakehouse processing, Snowflake warehousing, real-time event processing, and automated data quality controls.

Led the build-out of scalable healthcare data processing layers using Databricks, Delta Lake, Apache Iceberg, and AWS S3, supporting large-scale claims and member eligibility record volumes across batch workloads.

Built ingestion pipelines sourcing claims and eligibility data from Oracle and SQL Server transactional databases into the Lakehouse platform for downstream processing.

Spearheaded consolidation of data from PostgreSQL, MongoDB, and Cosmos DB into Snowflake, creating consistent datasets for enterprise reporting and analytics across multiple operational sources.

Developed real-time event processing pipelines using Kafka Streams to identify rule-based anomaly patterns and support downstream healthcare claims analytics.

Applied Apache Spark optimisation techniques including partition tuning, broadcast joins, caching, and resource allocation, meaningfully improving processing efficiency and eliminating recurring data spill issues.

Integrated automated data validation frameworks with GitHub Actions and Jenkins, applying automated quality checks across the majority of production-bound datasets before loading into downstream platforms.

Environment: AWS S3, Databricks, Delta Lake, Apache Iceberg, Snowflake, Apache Kafka, Kafka Streams, Apache Spark, PySpark, Python, SQL, PostgreSQL, MongoDB, Cosmos DB, GitHub Actions, Jenkins

Data Engineer LatentView Analytics (Enterprise Analytics) India

07/2022 – 10/2023

Co-led the transition of enterprise analytics workloads from on-premises infrastructure to Azure, developing PySpark ingestion pipelines for structured and semi-structured data.

Introduced Delta Lake storage layers with schema evolution capabilities to support changing upstream data formats without disrupting downstream pipelines.

Built Kafka and Azure Data Factory pipelines for incremental loading and Change Data Capture (CDC), supporting fresher data across analytical workloads.

Developed dbt models on Snowflake using incremental strategies, automated tests, and documentation to deliver reliable datasets for Tableau and Power BI analytics.

Developed reusable REST API integration workflows using Java and Spring-based microservices to standardise ingestion from enterprise SaaS applications.

Data Engineer Tiger Analytics (Big Data Consulting) India

01/2019 – 06/2022

Developed and optimised batch ETL workloads using Apache Spark, PySpark, and Hadoop for high-volume historical transactional data.

Consolidated datasets from AWS S3, Amazon Redshift, Azure Blob Storage, and MongoDB into centralised analytics-ready data marts.

Migrated legacy cron-based workflows to Apache Airflow DAGs with dependency management, automated validation, and failure alerting.

Evaluated and prototyped Apache Iceberg to address partition evolution challenges and improve query efficiency across large historical datasets.

Optimised Spark workloads through executor memory, shuffle partition, and caching strategies to resolve production memory and processing issues.

Junior Data Engineer Indriyn Data Analytics India

05/2017 – 12/2018

Developed SQL scripts, stored procedures, and Python-based transformation workflows to process operational data into MySQL databases.

Created automated validation checks for null values, duplicate records, and data integrity issues before downstream reporting.

Built foundational Power BI dashboards and ad-hoc data views using Google BigQuery to support operational reporting.

Monitored AWS EC2 batch jobs and supported troubleshooting of data loading issues while assisting with Docker and Jenkins CI/CD workflows.

EDUCATION

Bachelor of Technology – Electronics and Communication Engineering

Jawaharlal Nehru Technological University Hyderabad (JNTUH), India 2011



Contact this candidate