Post Job Free
Sign in

Senior Big Data & Lakehouse Engineer

Location:
Cincinnati, OH
Posted:
September 21, 2026

Contact this candidate

Resume:

Anjali Kandula Senior Big Data Engineer

https://www.linkedin.com/in/anjali-kandula/ phone:276-***-**** **********@*****.***@gmail.com PROFESSIONAL SUMMARY

• Experienced Data Engineer with 10+ years of expertise designing, developing, and maintaining robust ETL/ELT pipelines, data warehouses, and data lakes for large-scale analytics and reporting. Skilled in SQL, Python, Apache Spark, Databricks, Snowflake, and cloud platforms including AWS, Azure, and GCP.

• Designed and implemented enterprise-scale Delta Lake, Apache Iceberg, and Apache Hudi-based data lakes and lakehouses on AWS, Azure, and GCP, applying Lambda and Kappa Architecture patterns to support analytics, machine learning, and enterprise reporting workloads.

• Built scalable batch and real-time data pipelines using Apache Spark, PySpark, and Apache Flink, developing reusable processing frameworks in Python, Scala, and Java to handle large-scale distributed workloads across Hadoop, HDFS, and YARN environments.

• Developed Java-based Apache Spark transformation frameworks integrated with AWS EMR and Azure Databricks, processing high-volume enterprise datasets through end-to-end ETL workflows orchestrated via Apache Airflow with consistent reliability and fault-tolerant execution.

• Designed end-to-end ETL/ELT integration workflows using Informatica, Talend, IBM DataStage, Apache NiFi, AWS Glue, and Azure Data Factory, orchestrating complex pipeline dependencies through Apache Airflow and Oozie with advanced SQL and Python transformation logic.

• Implemented CDC and streaming ingestion pipelines using Apache Kafka, Kafka Streams, Confluent Platform, Amazon Kinesis, and RabbitMQ to feed downstream Snowflake, Redshift, BigQuery, and Azure Synapse warehouses with low-latency, fault-tolerant message delivery.

• Led cloud data platform modernization using AWS S3, EMR, Glue, Redshift, Athena, and Lambda alongside Azure ADLS, Databricks, Synapse, and GCP BigQuery to migrate legacy on-premise systems to scalable cloud-native Data Lakehouse Architectures.

• Deployed Java-based microservices integrated with Amazon Kinesis and AWS Lambda for real-time event processing, connecting ingestion streams to Delta Lake storage layers on AWS S3 with end-to-end pipeline orchestration managed through Apache Airflow.

• Enforced data governance and quality standards using Great Expectations, Apache Atlas, and Collibra, integrating validation checkpoints into Apache Airflow-orchestrated ETL pipelines and applying IAM-based access controls and encryption policies across AWS, Azure, and GCP platforms.

• Optimized Apache Spark execution plans, Snowflake query profiles, Redshift distribution keys, and BigQuery partitioning strategies, applying advanced SQL tuning techniques across PostgreSQL, Oracle, SQL Server, and MySQL to reduce processing time across high-volume workloads.

• Managed MongoDB, Cassandra, and DynamoDB NoSQL solutions supporting high-throughput microservices-based data platforms, integrating these storage layers with Kafka-driven streaming pipelines and Apache Spark Streaming frameworks for real-time read/write operations at enterprise scale.

• Implemented CI/CD pipelines for data engineering workflows using Jenkins and GitHub Actions, managing infrastructure-as-code through Terraform and containerizing data workloads with Docker and Kubernetes across Git, GitHub, and GitLab version-controlled repositories.

• Built Java Spring Boot-based data ingestion services connected to Apache Kafka topics, transforming and routing event streams into Apache Hudi tables on GCP storage with pipeline health monitored through Prometheus and Grafana observability dashboards.

• Monitored data platform health using ELK Stack, Prometheus, Grafana, and Splunk, configuring real-time alerting on pipeline failures and tracking performance metrics across distributed Apache Spark, Apache Kafka, and cloud-native workloads to ensure minimal production downtime.

• Developed complex data models across PostgreSQL, Oracle, SQL Server, and MySQL alongside cloud warehouses including Snowflake and Azure Synapse, designing normalized and dimensional schemas supporting operational reporting and analytical workloads while maintaining referential integrity at scale.

• Processed streaming data using Apache Flink and PySpark connected to Confluent Platform and Amazon Kinesis, storing curated outputs in Apache Iceberg tables on AWS S3 with orchestration managed through Apache Airflow DAGs and data quality enforced via Great Expectations.

• Provisioned cloud infrastructure across AWS, Azure, and GCP using Terraform, combining Docker and Kubernetes containerization to ensure consistent replication of Apache Spark and Apache Kafka environments across development, staging, and production with full audit traceability.

• Delivered analytical datasets consumed by Tableau, Power BI, and Looker dashboards by designing semantic layers and SQL and PySpark transformation workflows that translated complex business logic into reliable, reusable outputs aligned with enterprise reporting requirements.

• Built Java-based Kafka Streams consumer applications that processed real-time event data, applied business transformation logic, and persisted enriched records into Snowflake and Redshift data warehouses on AWS with monitoring handled through CloudWatch and Splunk.

• Applied microservices-based data platform design principles using Apache Kafka and Amazon Kinesis as event-driven communication backbones, decoupling ingestion, transformation, and serving layers to enable independent scalability and reduce cross-team dependency in enterprise ETL delivery.

• Configured Apache Airflow DAGs to orchestrate multi-step AWS Glue and Azure Data Factory pipeline workflows, integrating Great Expectations quality checks, IAM access controls, and alerting through Grafana to maintain governed and reliable data delivery.

• Processed large-scale Hive and HBase datasets using Apache Spark Core, Scala, and Java across Hadoop YARN clusters, applying MapReduce optimization patterns and PySpark SQL transformations to improve throughput and reduce end-to-end batch processing latency.

• Integrated GCP BigQuery, Dataflow, and Pub/Sub into enterprise streaming architectures, routing enriched event data from Apache Kafka through Python and Java transformation layers into Apache Iceberg and Delta Lake tables for analytical consumption.

• Maintained NoSQL data access layers using MongoDB, Cassandra, and DynamoDB, coordinating schema design and read/write optimization across microservices pipelines while ensuring storage consistency with upstream Apache Kafka event streams and downstream Snowflake reporting layers.

• Mentored junior engineers and led architecture design reviews covering Apache Spark optimization, Apache Kafka consumer group strategies, Apache Airflow DAG design, and Terraform-based cloud cost management to build team capability and ensure consistent delivery quality.

• Automated environment reproducibility across AWS, Azure, and GCP using Terraform and Docker, deploying containerized PySpark, Java, and Scala workloads on Kubernetes clusters with observability maintained through Prometheus, Grafana, and ELK Stack monitoring pipelines.

TECHNICAL SKILLS

• Big Data Ecosystem Tools: Apache Spark, Oozie, Zookeeper, Hadoop, Scala, Impala, Kafka, Nifi, Elastic MapReduce (EMR), Fact and Dimension Tables, Cloudera distribution, Hortonworks Ambari, HDFS, Map Reduce, YARN, PIG, Sqoop, HBase, Hive, Flink, Flume, Cassandra.

• Relational and NOSQL Databases: RDBMS, Oracle, SQL Server, PostgreSQL, DB2, DynamoDB, MongoDB, HBase, Cassandra, Cosmos DB.

• Programming Languages: SQL, PL/SQL, Python, Spark, JavaScript, HTML, SAS, Shell Scripting, Perl.

• ETL and Data Integration: IBM DataStage 11.5/9.1, Informatica, Talend 5.3/6.2, TOAD 10.5, Teradata SQL Assistant, Spring Boot, Databricks, Anaconda, Visual Source Safe, Putty, Notepad++, GitHub, Confluence, Jenkins.

• Cloud Platforms: Amazon Web Services (AWS), GCP, Azure.

• Services: Active Directory, Azure Monitoring, Azure Search, Data Factory, Key Vault, SQL Azure, Azure DevOps, Azure Synapse Analytics (DW), Azure Data Lake, AWS Lambda, Athena, AWS Elastic Compute Cloud (EC2), AWS Glue, Redshift, Kinesis, AWS Virtual Private Cloud (VPC), AWS CloudWatch, Step Functions, BigQuery, Composer, Google Cloud Storage, Cloud Dataproc.

• Visualization and Reporting: MS Excel, Power BI, Tableau, SPSS.

• Containerization, Orchestration Tools: Kubernetes, Docker, Docker Registry, Docker Hub, Docker.

• Methodologies: Agile, Scrum, Waterfall.

• Libraries: Pandas, NumPy, SciPy, Matplotlib, Seaborn, Scikit-Learn.

• Algorithms: Decision Tree, Random Forest, Logistic Regression, Gradient Boosting, Support Vector Machine (SVM), K-Nearest Neighbor (KNN).

• IDE’s & Frameworks: Anaconda Software, Jupyter Notebook, Pycharm, Django, Flask.

• Data Storage: Data Lakes, Data Warehouses, Distributed Storage Systems

• Streaming Technologies: Apache Kafka, Apache Flink, Kinesis, Event Hubs, Pub/Sub

• ETL/ELT Tools: Apache Airflow, Apache NiFi, Talend, Informatica, AWS Glue, Azure Data Factory, Google Dataflow

• Workflow Orchestration: Apache Airflow, Luigi, Prefect, Dagster, AWS Step Functions, Azure Data Factory

• Data Processing: Batch Processing, Real-Time Processing, Stream Processing, Distributed Processing

• Infrastructure & DevOps: Infrastructure as Code, CI/CD, Docker, Kubernetes, Terraform, CloudFormation, Ansible, Jenkins, GitLab CI, GitHub Actions

• Monitoring & Logging: DataDog, Splunk, ELK Stack, CloudWatch, Azure Monitor, Stackdriver, Prometheus, Grafana

• Data Governance: Data Lineage, Data Cataloging, Apache Atlas, Collibra, Alation, AWS Glue Catalog, Azure Purview

• Version Control: Git, GitHub, GitLab, Bitbucket

• Architectures: Event-Driven Architecture, Microservices, Distributed Systems

• Data Formats: Structured Data, Unstructured Data, Parquet, Avro, ORC, JSON, CSV

• Business Intelligence: Analytics Platforms, Reporting Tools

PROFESSIONAL EXPERIENCE

Client: FedEx Memphis, TN

Role: Senior Big Data Engineer June 2024 – Current

• Designed a Data Lakehouse platform for large-scale transportation analytics using Delta Lake, Apache Iceberg, and Apache Hudi on AWS Lake Formation, unifying batch and streaming access across fleet management, route optimization, and cargo tracking domains with enforced data governance through Collibra and DataHub.

• Enforced data lineage and metadata management across all ingestion layers using Collibra, DataHub, and AWS Glue Data Catalog, aligning physical data models with business domain requirements documented through Erwin for both the Data Lakehouse and data warehouse layers in the transportation platform.

• Built enterprise-scale ETL/ELT pipelines using Apache Spark, PySpark, and Spark SQL on AWS EMR and Databricks, processing high-volume transportation event data including vehicle telemetry, GPS coordinates, and shipment records following Lambda Architecture and Kappa Architecture principles for fault-tolerant analytics workflows.

• Developed Scala and Java-based Apache Spark processing modules for high-throughput batch transformation of freight billing and carrier settlement data on AWS EMR clusters configured with YARN, using Shell Scripting and AWS Lambda functions to automate cluster bootstrapping, job submission, and post-processing cleanup tasks.

• Developed real-time streaming ingestion pipelines using Apache Kafka, Kafka Streams, and Confluent Platform, integrated with AWS Kinesis for cross-system event routing in transportation operations, delivering sub-second latency and end-to-end guarantees aligned with SLA requirements across all operational data flows.

• Implemented Apache Flink and Apache Pulsar for complex event processing of vehicle sensor data and logistics status updates, consuming streams from Apache Kafka and routing processed results to Delta Lake and AWS S3 targets for downstream transportation analytics and historical reprocessing workflows.

• Configured HBase on Hadoop clusters for low-latency random read/write access to high-frequency transportation telemetry including GPS pings and vehicle diagnostics, integrating HBase with Apache Kafka consumers and Apache Storm for windowed aggregations before persisting results to Delta Lake for historical analysis.

• Built CDC (Change Data Capture) pipelines integrating AWS RDS, PostgreSQL, and DynamoDB source systems with Apache Kafka and Delta Lake sink targets, ensuring near-real-time synchronization of transportation master data including carrier profiles, route schedules, and freight contracts with AWS IAM enforced least-privilege access.

• Engineered data warehouse solutions on Snowflake, AWS Redshift, and Teradata using Star Schema modeling and Advanced SQL optimization to support enterprise reporting for transportation KPIs including on-time delivery rates, fleet utilization, and fuel consumption integrated with dbt transformation workflows.

• Integrated dbt transformation workflows with AWS Glue Data Catalog to automate schema evolution, lineage tracking, and model documentation across Snowflake and AWS Redshift warehouse layers, supporting downstream Tableau reporting and BI dashboards for transportation operations performance analysis.

• Implemented Apache Airflow and Dagster for end-to-end pipeline orchestration across batch ingestion, transformation, and delivery workflows, coordinating jobs between AWS Glue, AWS EMR, AWS Step Functions, and Databricks to ensure reliable scheduling and dependency management across the transportation data platform.

• Maintained Apache Oozie workflows for legacy Hadoop MapReduce jobs on HDFS and YARN, ensuring backward compatibility during the cloud migration phase while transitioning transportation workloads incrementally to Apache Airflow-orchestrated pipelines running on AWS EMR and Databricks cluster environments.

• Established data quality validation frameworks using Great Expectations within Apache Airflow DAGs, applying rule-based checks on transportation datasets including shipment manifests and driver logs, with quality failures triggering automated alerts via AWS CloudWatch, Prometheus, and Grafana dashboards for rapid incident response.

• Integrated IBM DataStage, Informatica, and SSIS for legacy-to-cloud data migration in the transportation domain, extracting structured data from Oracle, SQL Server, and PostgreSQL databases into AWS S3 staging zones, using Fivetran for SaaS connector-based ingestion with built-in error handling and retry logic.

• Designed and optimized Apache Hive and Trino query layers over AWS S3-backed data lakes for ad-hoc transportation analytics workloads, applying AWS Athena for serverless query execution on partitioned Parquet and ORC datasets to reduce compute costs and improve response times for fleet and logistics queries.

• Deployed containerized data engineering workloads using Docker, Kubernetes (AWS EKS), and Helm charts, with infrastructure provisioned through Terraform and configuration managed via Ansible, supporting reliable and scalable execution of transportation data pipeline workloads across development, staging, and production environments.

• Implemented CI/CD pipelines using Jenkins and Git for automated testing with PyTest, deployment validation, and rollback capabilities across Docker and Kubernetes (AWS EKS) environments, following enterprise DevOps best practices and maintaining code quality standards for the transportation data engineering platform.

• Deployed AWS CloudFormation and Terraform templates for full infrastructure-as-code provisioning of the transportation data platform, including AWS EC2, AWS ECS, Elastic Beanstalk, AWS CloudFront, and API Gateway components automated through Jenkins-driven CI/CD pipeline drift detection and validation processes.

• Maintained security and compliance controls by enforcing AWS IAM policies, AWS CloudTrail audit logging, and AWS X-Ray distributed tracing across all transportation data pipeline executions, with AWS Lake Formation column and row-level access controls aligned to SOC 2 and FMCSA data handling requirements.

• Applied AWS Secrets Manager for credential management across Apache Kafka, PostgreSQL, AWS RDS, and DynamoDB pipeline components, enforcing AWS IAM role-based security controls and maintaining comprehensive audit trails for all transportation data platform access events reviewed during internal compliance assessments.

• Implemented AWS SageMaker and MLflow integrated data pipelines supporting transportation predictive maintenance models built with XGBoost and TensorFlow, with feature stores backed by Snowflake and AWS S3, and training data prepared via PySpark feature engineering jobs running on AWS EMR and Databricks.

• Applied LangChain and LlamaIndex to build retrieval-augmented generation pipelines for transportation operations knowledge bases, using Pinecone as the vector store and Hugging Face Transformers for embedding generation, consuming structured data from AWS S3 and DynamoDB for natural language querying of shipment history.

• Modeled complex transportation network relationships using Neo4j graph database for route dependency analysis and carrier network mapping, integrating Neo4j query results with Snowflake analytical tables via Python-based ingestion scripts surfaced through Tableau dashboards reports for strategic planning decisions.

• Monitored end-to-end pipeline health using the Prometheus, Grafana setting alerting thresholds for job failures and processing latency spikes in transportation workflows, integrating dashboards with Slack notifications and AWS CloudWatch alarms to ensure timely response for critical pipeline disruptions.

• Developed Java-based AWS Lambda functions integrated with AWS Kinesis and Apache Kafka for real-time event-driven processing of transportation shipment status updates, applying AWS IAM role-based access and AWS CloudWatch monitoring to ensure secure, traceable, and highly available serverless pipeline execution.

• Mentored junior data engineers on PySpark optimization techniques, Apache Airflow DAG design patterns, and Kubernetes-based deployment workflows, standardizing PyTest unit testing frameworks and NumPy-based numerical processing utilities within Python data preparation scripts, with all technical decisions and reviews documented in Confluence and Jira.

Environment: Python, Hadoop, Kubernetes, Databricks, Apache Flink, Apache Oozie, DynamoDB, Grafana, AWS EC2, AWS Redshift, Teradata, TensorFlow, Apache Kafka, Apache Iceberg, NumPy, Apache Hudi, Dagster, Apache Airflow, Prometheus, AWS IAM, XGBoost, PyTest, Helm, PySpark, AWS Secrets Manager, Fivetran, MLflow, AWS Step Functions, AWS CloudTrail, Git, AWS Lake Formation, HBase, Apache Pulsar, Trino, Pinecone, Informatica, DataHub, AWS RDS, LangChain, AWS Athena, Scala, AWS CloudFormation, Kinesis, Java, Snowflake, Apache Hive, SSIS, Spark SQL, AWS ECS, Confluent, Delta Lake, AWS CloudFront, Docker, PostgreSQL, Hugging Face Transformers, Elastic Beanstalk, AWS Kinesis, Hadoop MapReduce, Kafka Streams, Neo4j, Jenkins, YARN, Luigi, AWS X-Ray, AWS EKS, Ansible, AWS Glue Data Catalog, AWS SageMaker, LlamaIndex, dbt, API Gateway, IBM DataStage, Terraform, AWS S3, Great Expectations, AWS Glue, AWS CloudWatch, Tableau, Shell Scripting, AWS Lambda, AWS EMR, Confluent Platform, Apache Spark, Apache Storm.

Client : Walgreens Deerfield, IL

Role: Senior Data Engineer July 2022 – May 2024

• Designed enterprise-scale data lakehouse solutions on Azure Data Lake Storage (ADLS) using Apache Iceberg and Apache Hudi table formats, enabling ACID-compliant transactional management across clinical trial, patient records, and pharmaceutical supply chain datasets with enforced HIPAA and HITECH security controls.

• Built large-scale batch and real-time ingestion pipelines using Apache Spark, PySpark, and Azure Databricks to process HL7 and FHIR-formatted healthcare messages, ensuring pipeline reliability through structured exception handling and end-to-end data lineage tracking cataloged via Alation.

• Developed ETL/ELT workflows using Azure Data Factory, Talend, and Matillion to integrate data from SQL Server, Teradata, and MongoDB into Snowflake and Azure Synapse Analytics, applying Snowflake Schema modeling patterns alongside HIPAA-compliant data masking and encryption policies across all ingestion layers.

• Implemented IBM DataStage pipelines to extract, transform, and load pharmaceutical and clinical data from legacy source systems into Snowflake and Azure Synapse Analytics, enforcing HIPAA data governance standards and validating schema conformance using Great Expectations automated rule suites.

• Built Apache NiFi data flows to automate ingestion of HL7 and FHIR healthcare interoperability messages from hospital information systems into the Azure Data Lake Storage (ADLS) landing zone, applying inline routing, transformation, and encryption to satisfy HIPAA security rule requirements before downstream processing.

• Implemented Change Data Capture (CDC) and streaming ingestion pipelines using Confluent Platform (Apache Kafka) and Azure Event Hub to process real-time clinical event streams, integrating RabbitMQ as an asynchronous messaging layer connecting pharmacy and lab order systems reporting consumers in an event-driven architecture.

• Configured Azure Event Grid alongside Azure Event Hub and Confluent Platform (Apache Kafka) to trigger event-driven pipeline executions for real-time pharmaceutical supply chain updates, routing processed clinical event data into Azure Data Lake Storage (ADLS) for structured analytics consumption.

• Designed and managed Apache Airflow and Prefect orchestration frameworks to schedule, monitor, and recover complex multi-step DAG-based pipeline workflows across Azure Databricks, Azure Functions, and Azure Data Lake environments, supplemented by Luigi for legacy pharma batch dependency orchestration.

• Developed Scala and Python-based Apache Spark transformation frameworks on Azure Databricks, utilizing Spark SQL, Apache Hive, and Presto for federated querying across Snowflake and Azure Synapse Analytics; all code was version-controlled through GitHub with peer-reviewed pull request workflows enforced via GitHub Actions.

• Implemented dbt transformation models on Snowflake and Azure Synapse Analytics to standardize pharmaceutical sales, clinical outcomes, and patient demographic data marts, applying dbt schema contracts alongside Great Expectations validation suites to maintain data integrity across all Power BI reporting layers.

• Enforced enterprise data governance standards using Apache Atlas for metadata management integrated with Alation for business glossary maintenance and data lineage cataloging, building governance frameworks in full compliance with HIPAA patient data security requirements and HITECH breach notification standards.

• Established CI/CD pipelines using Azure DevOps, GitHub Actions to automate deployment of data pipeline code, Terraform-based infrastructure provisioning, and Docker container builds to Azure Kubernetes Service (AKS), maintaining environment consistency across development, staging, and production deployments.

• Deployed containerized data pipeline components using Docker and Azure Kubernetes Service (AKS), writing Bash shell scripts and Terraform infrastructure-as-code modules to provision Azure Blob Storage, Azure SQL Database, and Azure Data Lake resources in a repeatable and auditable manner.

• Developed Java-based microservices on Azure App Service to orchestrate FHIR-compliant API data exchanges between internal analytics platforms and external healthcare partner systems, integrating Azure API Management and Azure Service Bus for governed, transport-secured interoperability meeting HIPAA data exchange policies.

• Built Java streaming consumers using Confluent Platform (Apache Kafka) and deployed them on Azure Kubernetes Service (AKS) to process high-throughput clinical event data in real time, persisting transformed outputs into Azure Data Lake Storage (ADLS) and Snowflake for downstream analytical consumption.

• Configured Azure Key Vault for secrets management and credential rotation across all data pipeline services, integrating with Azure Active Directory for RBAC enforcement and ensuring PHI access was fully logged and auditable in compliance with HIPAA access control and audit trail mandates.

• Built monitoring and observability frameworks using ELK Stack (Elasticsearch, Logstash, Kibana), Datadog, and Azure Monitor to track pipeline SLAs, data freshness, and infrastructure health across Azure Kubernetes Service, Azure Virtual Machines, and Azure App Service deployments serving healthcare data ingestion workflows.

• Developed machine learning feature engineering pipelines using PySpark, Pandas, scikit-learn, and PyTorch on Azure Databricks, integrating MLflow for experiment tracking and model versioning while applying CatBoost and FAISS-based similarity search for pharmaceutical patient cohort identification workflows.

• Authored Advanced SQL, Spark SQL, and Python-based transformation scripts to process and normalize pharmaceutical billing, claims, and EHR data loaded into Teradata and Azure Synapse Analytics, applying ER/Studio for physical and logical data modeling to maintain consistent enterprise schema documentation.

• Conducted PyTest-driven unit and integration test suites for PySpark transformation logic, dbt models, and Python-based ingestion modules, integrating all test pipelines into GitHub Actions workflows to enforce quality gates before any code was merged or deployed to production data environments.

• Collaborated with BI and Analytics teams to design Power BI semantic layer models sourced from Snowflake and Azure Synapse Analytics data marts, producing clinical trial, drug safety, and supply chain KPI reports with end-to-end data lineage fully documented in Alation and maintained in Confluence.

• Provided technical mentoring to junior data engineers on Apache Spark optimization techniques, Apache Hudi upsert strategies on Azure Data Lake Storage (ADLS), Confluent Platform Kafka consumer group design, and HIPAA-compliant data handling practices documented in Confluence and tracked via Jira.

• Coordinated cross-functional delivery with Data Science, DevOps, and compliance stakeholders using Jira for sprint planning and Microsoft Teams for operational communication, ensuring all platform deliverables met agreed SLAs, passed Great Expectations quality thresholds, and compliant with HIPAA, HITECH, and HL7/FHIR standards.

Environment: Apache Hive, Azure Blob Storage, Azure DevOps, Azure API Management, RabbitMQ, Python, Azure Synapse, Pandas, Presto, Azure Service Bus, CatBoost, Datadog, Azure Functions, MLflow, Kubernetes, Azure Key Vault, Azure Virtual

Machines, scikit-learn, Terraform, Apache NiFi, Azure Event Hub, Snowflake, GitHub, HDFS, Azure Data Lake, Azure Data Lake Storage (ADLS), PyTorch, ER/Studio, Azure SQL Database, Azure Active Directory, Matillion, dbt, Great Expectations, IBM DataStage, Spark SQL, Azure Databricks, Apache Hudi, Azure Monitor, Alation, Apache Spark, Azure Data Factory, Azure App Service, Azure Kubernetes Service, Azure Synapse Analytics, PyTest, FAISS, Bash, PySpark, Teradata, ELK Stack, GitHub Actions, Docker, Confluent Platform, Scala, Apache Storm, Apache Iceberg, Apache Airflow, Talend, Hadoop, Luigi, Java, Prefect, Power BI, MongoDB, SQL Server, Azure Event Grid.

Client: Costco Issaquah, WA

Role: Data Engineer August 2019 – June 2022

• Architected end-to-end batch and real-time data pipelines on GCP Cloud Dataflow and GCP Cloud Dataproc using Apache Beam for unified processing of large-scale retail transaction data, enforcing PCI-DSS and GDPR compliance through field-level encryption and GCP Cloud IAM access controls.

• Constructed scalable PySpark and Spark SQL processing frameworks on GCP Cloud Dataproc to transform retail customer, inventory, and sales data using Apache Hive for schema management and query optimization, enabling structured data availability for downstream analytics and Looker-based reporting dashboards.

• Delivered ETL pipelines using Informatica and SSIS to ingest retail point-of-sale, supplier, and loyalty program data into Snowflake and Teradata; implemented Star Schema dimensional models in Snowflake to support enterprise reporting aligned with retail merchandising and supply chain business domains.

• Established Delta Lake and Apache Hudi based lakehouse architectures on GCP Cloud Storage to manage large retail datasets with ACID transaction support, schema evolution, and time-travel capabilities; configured upsert and merge operations in Apache Hudi to handle CDC patterns from Oracle operational databases.

• Engineered Apache Airflow DAGs on GCP Cloud Composer to orchestrate complex multi-step retail workflows including inventory refresh, customer segmentation, and promotional data loads; integrated Oozie for legacy Hadoop-based job scheduling, maintaining dependency graphs ensuring SLA-compliant data delivery across business units.

• Spearheaded streaming ingestion pipelines using GCP Cloud Pub/Sub and Apache Flink to process real-time retail event streams including clickstream and cart abandonment data; applied Apache Storm for low-latency event processing and routed enriched results into GCP BigQuery operational dashboards maintaining GDPR-compliant data handling.

• Optimized GCP BigQuery...



Contact this candidate