Post Job Free
Sign in

Senior Data Engineer - Big Data & Lakehouse

Location:
United States
Salary:
75
Posted:
August 11, 2026

Contact this candidate

Resume:

Varshini A

Sr. Data Engineer

Phone: +1-314-***-****

Email: ********.*********@*****.***

LinkedIn: https://www.linkedin.com/in/varshini-anagurthi-74a0b1377/ PROFESSIONAL SUMMARY:

● Senior Data Engineer with 10+ years of experience in Big Data, Data Warehousing, Lakehouse Architecture, Data Modeling (Star Schema, Snowflake Schema, Data Vault, OLTP/OLAP), Medallion Architecture, Domain Data Products, and enterprise-scale data platforms across AWS, Azure, and GCP.

● Expertise in building scalable ETL/ELT Pipelines, Batch & Real-Time Streaming Solutions using Apache Spark, PySpark, Scala, Kafka, Flink, Apache NiFi, Azure Event Hubs, Airflow, Dagster, Informatica, dbt, PL/SQL, and cloud-native orchestrators, enabling low-latency analytics and high-volume data processing.

● Strong experience in SQL Development, ANSI SQL, Query Optimization, Stored Procedures, Indexing, Dimensional Modeling, Star & Snowflake Schemas, OLAP Solutions, Data Reconciliation, and enterprise data integration across SQL Server, Oracle, DB2, MySQL, PostgreSQL, Teradata, and Sybase IQ.

● Hands-on expertise with Databricks, Delta Lake, Apache Iceberg, Parquet, Snowflake, Redshift, BigQuery, Dremio, Presto, Trino, Hadoop (HDFS, Hive), and modern data lake architectures built on AWS S3, Azure Data Lake Storage Gen2 (ADLS), and GCP Cloud Storage.

● Extensive cloud experience with Azure Data Factory (ADF), Azure Synapse Analytics, Microsoft Fabric, Azure Databricks, Azure ML, AWS Glue, EMR, Lambda, IAM, CloudFormation, SageMaker, Vertex AI, and cloud-native services for scalable analytics, governance, automation, and cost optimization.

● Skilled in Python, PySpark, Scala, RESTful APIs, SOAP Web Services, JSON, Avro, FTP, SMTP, OO Python Modules & Packages, and custom ingestion frameworks supporting enterprise, IoT, and external system integrations.

● Experienced in building AI-ready Data Models, Machine Learning Solutions, MLOps/DataOps Pipelines using MLflow, AWS SageMaker, Azure ML, Google Vertex AI, LangChain, LangGraph, LLM/RAG Solutions, and predictive analytics use cases including fraud detection, risk modeling, and patient analytics.

● Proficient in Power BI (DAX, Semantic Models, Star Schema Optimization), Tableau, Looker, and SSRS for interactive analytics and executive reporting; experienced with Git, GitHub, Azure DevOps, Jenkins, Docker, Kubernetes, CloudWatch, Monitoring, Logging, Observability, JIRA, Confluence, and Agile

(Scrum/Kanban) methodologies.

TECHNICAL SKILLS:

Programming Languages Python, SQL, PL/SQL, Shell scripts, Scala, Unix Scripting Languages: Java Script, Python, Shell Script. Web Servers: Apache Tomcat4.1/5.0 Big Data Tools Hadoop, Apache Spark, MapReduce, Flink, PySpark, Hive, YARN, Kafka, Flume, Oozie, Airflow, Zookeeper, Sqoop, HBase

Cloud Services AWS, Azure, GCP

ETL/Data warehouse

Tools

Informatica, Talend, DataStage, Power BI, and Tableau Version Control &

Containerization tools

SVN, GIT, Bitbucket, Docker, and Jenkins CVS, Code Commit, GIT hub, ApacheLog4j, TOAD, ANT, Maven, JUnit, JMock, Mockito, REST HTTP Client, JMeter, Cucumber, Aginity.

Databases Oracle, MySQL, MongoDB, and DB2

Operating Systems Ubuntu, Windows, and Mac OS

Methodologies Agile/ Scrum and Traditional Waterfall PROFESSIONAL EXPERIENCE:

Centene Corp, St. Louis, MO Sep 2024 to Present

Sr. Data Engineer

Responsibilities:

● Designed and optimized scalable ETL/ELT pipelines using DBT (macros, Jinja templating, dbt models), Apache Spark, PySpark, Scala, Apache Beam, Apache Flink, Airflow,Pub/Sub,Dagster, AWS Glue, Amazon EMR (Spark/Hadoop/YARN), Delta Lake, Apache Iceberg, Hadoop, Hive, Piggybank UDFs, Sqoop, and NiFi, and Kafka for batch and real-time ingestion, reconciliation of disparate systems, raw ingestion, curated data layers, and AI-ready data models.

● Managed large-scale migration, ingestion, and modernization from Oracle, Teradata (SQL, BTEQ, AMP usage), and legacy systems to AWS and GCP ecosystems using Cloud Data Fusion, Dataflow, BigQuery, Snowflake, Redshift, Dremio, Presto, Trino, Informatica PowerCenter, and IICS, enabling scalable data warehousing, data marts, star schema, dimensional models, and Data Vault architectures.

● Built real-time and batch streaming systems using Kafka, Flink, Spark Streaming, Kinesis, AWS Lambda, API Gateway, REST APIs, and load balancers, supporting event-driven architectures, FDNS integrations, AI/ML pipelines, and low-latency data products for healthcare analytics and monitoring systems.

● Designed enterprise Data Warehousing, OLAP/OLTP systems, data marts, and lakehouse architectures using Snowflake, BigQuery, Redshift, Hive, PostgreSQL, Oracle, and Delta Lake, implementing query optimization, indexing, stored procedures, SQL tuning, scalable schemas, and data modeling (star schema, Data Vault, dimensional models) supporting KPI dashboards in Power BI, Tableau, Looker, and SSRS.

● Developed and deployed cloud-native applications and microservices using Docker, Kubernetes, GKE, Cloud Run, AWS ECS, along with Terraform, CloudFormation, Ansible, Jenkins, Git, Stash, Maven, CI/CD pipelines, and DataOps/MLOps practices, enabling automation, cloud cost optimization, and scalable infrastructure provisioning.

● Implemented AI/ML, LLM, RAG solutions, batch AI use cases, and ML pipelines using Bedrock,SageMaker, Vertex AI, Databricks, Spark MLlib, Jupyter, Python (OO modules/packages), Conda, enabling predictive healthcare analytics, anomaly detection, and AI-enabled data products including model training, inference, and evaluation workflows.

● Established secure and governed cloud ecosystems using AWS IAM, KMS, AWS Security, CloudWatch, logging/monitoring tools, encryption standards, and compliance frameworks, along with data governance, data contracts, schema enforcement, and curated data layers to ensure HIPAA-compliant, production-grade data platforms.

● Improved end-to-end data ecosystem performance through query optimization (Trino/Presto/SQL), performance tuning, automation frameworks, reconciliation pipelines, scalability improvements, and cloud cost optimization, while collaborating in Agile (SCRUM, Kanban) environments for continuous delivery and data product development.

Environment: Apache Spark, PySpark, SQL, Python, Scala, Kafka, Hadoop, Hive, AWS Glue, Amazon EMR, Amazon S3, AWS Lambda, Amazon Redshift, BigQuery, Google Cloud Dataflow, Pub/Sub, Snowflake, Airflow, dbt, Kubernetes, Docker, Terraform, CI/CD, Data Modeling, Machine Learning (SageMaker / Vertex AI / LLM / RAG). Bath & Body Works, Columbus, OH Dec 2021 - Aug 2024 Sr Data Engineer

Responsibilities:

● Designed and implemented scalable data pipelines using Google Cloud Platform (GCP) services including BigQuery, Cloud Storage, Dataproc, Pub/Sub, and Cloud Dataflow, supporting enterprise analytics with optimized partitioning, clustering, and SQL-based performance tuning.

● Developed large-scale data processing workflows using Apache Spark, PySpark, Scala, Databricks, and Python, enabling data extraction, ingestion, transformation, normalization, quality checks, and loading (ETL) for structured and semi-structured datasets.

● Built and optimized end-to-end ETL pipelines and data engineering workflows, incorporating DBT, SQL, and data modeling techniques, ensuring reliable transformation logic and analytics-ready datasets.

● Developed streaming and batch ingestion systems using Apache Kafka, Google Pub/Sub, Cloud Dataflow, and Apache Flink with AVRO serialization, enabling real-time data processing and schema enforcement.

● Executed Azure-based data platform development and migration using Azure Data Factory (ADF), Azure Data Lake Storage (ADLS), Azure Synapse Analytics, and Azure Blob Storage, supporting scalable enterprise data movement and warehousing.

● Worked with big data and distributed systems including HDFS, Hive, Sqoop, MapReduce, YARN, Oozie, Pig, optimizing legacy and modern data processing workflows for batch analytics and historical reporting.

● Built and deployed cloud-native and full-stack data applications using FastAPI, REST APIs, React, and Python, along with Git-based CI/CD pipelines for automated deployment and version control.

● Designed enterprise-grade data warehouse and analytics solutions using Snowflake, Databricks, Azure Synapse, and SQL, supporting BI dashboards via BigQuery, Looker, and Data Studio, in Agile/Scrum environments.

Environment: Google Cloud (BigQuery, Dataproc, Pub/Sub, Dataflow, GKE), Azure (ADF, ADLS, Synapse), Snowflake, Databricks, Spark, PySpark, Kafka, Flink, Hadoop, DBT, Python, SQL, FastAPI, REST APIs, Git, CI/CD, ETL, data modeling, streaming and batch processing.

Allstate, Northbrook, IL Sep 2020 - Nov 2021

Data Engineer

Responsibilities:

● Designed and maintained distributed ETL pipelines using Apache Spark, Scala, and PySpark for large-scale batch and real-time data processing across big datasets, including data ingestion, cleansing, transformation, and schema modeling, while working in cross-functional teams following Scrum methodology.

● Implemented complex transformations using Spark SQL, custom UDFs, and performance tuning techniques to optimize processing across distributed systems.

● Built and scheduled end-to-end workflows using Apache Oozie, orchestrating Hadoop ecosystem jobs including Hive, Pig, HDFS, and MapReduce, with cluster resource management via YARN.

● Developed ingestion pipelines using Kafka, Sqoop, Pig scripts, and AWS Kinesis, enabling reliable and scalable data movement into HDFS and cloud storage.

● Migrated legacy Hadoop workloads to AWS EMR and Qubole, and implemented cloud-native pipelines using AWS Glue, AWS Lambda, AWS S3, AWS Athena, AWS IAM, and Glue Data Catalog, integrating with Amazon Redshift, Amazon RDS, and Azure Cosmos DB.

● Built real-time streaming solutions using Spark Streaming, Kafka, AWS Kinesis, and AWS Lambda, and orchestrated batch and streaming workflows using Apache Airflow, AWS Step Functions, and Apache Oozie.

● Designed and deployed scalable data platforms and infrastructure using Azure Data Factory (ADF), Snowflake, Databricks, Terraform, CloudFormation, Docker, OpenShift, and Kubernetes, with data storage and modeling across Amazon S3, Amazon Redshift, Amazon RDS, and Azure Cosmos DB, and implemented monitoring using Grafana, AWS CloudWatch, and Apache JMeter in cross-functional Agile Scrum environments. Environment: Apache Spark, Scala, PySpark, HDFS, Hive, Pig, Sqoop, Kafka, YARN, Apache Oozie, Apache JMeter, Grafana, Snowflake, Azure Data Factory (ADF), Azure Data Lake, Azure SQL, Databricks, AWS. Big Lots, Columbus, OH Jan 2018 - Aug 2020

Data Engineer

Responsibilities:

● Designed and developed a custom ETL framework using Apache Spark and Scala, enabling dynamic transformations and scalable batch data processing across large datasets.

● Built end-to-end ETL pipelines using Spark, Scala, and Python, integrating AWS S3, Amazon Redshift, Teradata, AWS Lambda, SNS, and Kinesis for cloud-native data workflows.

● Implemented real-time streaming pipelines using AWS Kinesis, Spark Streaming, and AWS Lambda, achieving sub-second event processing for streaming workloads.

● Executed complex transformations using Spark SQL, including joins, aggregations, and data validation, and developed Hive/HiveQL scripts for structured data processing and legacy warehouse support.

● Built and orchestrated workflows using Apache Oozie, HDFS, YARN, and Hadoop ecosystem tools, ensuring automated multi-stage pipeline execution and cluster resource optimization.

● Developed Python ETL scripts with AWS Lambda and SNS triggers, along with SSIS packages (Lookups, Merge, Data Conversion), Informatica workflows/Mapplets, and Alteryx/Shell scripting, and implemented infrastructure automation using Terraform for cloud resource provisioning.

● Designed data ingestion and quality frameworks with XML parsing (Python), checksum validation, schema enforcement, and metadata tracking, integrated CloudWatch monitoring, and processed IoT sensor data for real-time analytics and predictive maintenance in Cassandra, Zeppelin, and enterprise data systems. Environment: Apache Spark, Hive, Scala, Python, AWS S3, Redshift, Lambda, Kinesis, CloudWatch, SSIS, SQL Server, Informatica, Apache Oozie, YARN, XML, Alteryx, Shell Scripting, Spark SQL, SNS, Cassandra, Zeppelin. PWC, India Aug 2015 - Dec 2017

Data Engineer

Responsibilities:

● Developed and maintained end-to-end ETL pipelines using Python, automating data extraction, transformation, and loading from multiple sources into SQL Server and Oracle databases.

● Built reusable Python frameworks for data validation, reconciliation, logging, exception handling, and regulatory compliance, improving data quality and workflow monitoring.

● Designed and optimized complex T-SQL and PL/SQL objects including stored procedures, triggers, views, and packages, applying indexing and query tuning for performance and data integrity.

● Developed reusable Python modules integrating Erwin Data Modeler outputs with SSIS workflows and reporting tools such as SharePoint, improving data model integration and reporting automation.

● Designed and deployed SSIS packages using advanced transformations and engineered ETL workflows to improve data integration efficiency and reliability.

● Built and optimized big data pipelines using Apache Spark (Scala, Spark SQL, DataFrames) along with Kafka, Hive, Impala, Pig, HDFS, and MapReduce, including migration of SQL-based ETL to Spark on YARN.

● Configured Oozie workflows and leveraged Hadoop ecosystem tools including YARN, HiveQL, Sqoop, and HBase to orchestrate batch and near real-time processing, supporting partitioning, bucketing, and distributed analytics.

● Migrated legacy on-premises systems to AWS, integrating Amazon S3, AWS Lambda (functions), and AWS CloudWatch, and processed JSON and Parquet datasets in Spark pipelines for analytics, data validation, and QA workflows.

Environment: Python, SQL Server 2008/2012, T-SQL, PL/SQL, Oracle 10g, SSIS, Erwin Data Modeler, Apache Spark, Kafka, Hive, Impala, Pig, HDFS, MapReduce, YARN, Oozie, Sqoop, HBase, JSON, AWS.



Contact this candidate