DAVID R
SENIOR DATA ENGINEER PYTHON AWS BIG DATA SQL CLOUD DATA ENGINEERING
*************@*****.*** +1-551-***-****
PROFESSIONAL SUMMARY:
Senior Data Engineer with 9+ years of experience designing, developing, deploying, and supporting enterprise-scale data engineering solutions, cloud data platforms, distributed processing frameworks, ETL/ELT pipelines, and data-intensive applications.
Strong hands-on expertise in Core Python, SQL, PySpark, Apache Spark, Shell Scripting, Big Data technologies, distributed data processing, and enterprise data integration.
Extensive experience developing scalable data pipelines using Python, PySpark, Spark SQL, AWS S3, Amazon Redshift, Amazon Aurora, Kafka, AWS SQS, and cloud-native data engineering technologies.
Strong experience designing event-driven data architectures using AWS SQS, Kafka, streaming frameworks, and asynchronous processing patterns.
Experienced in developing REST APIs, API integrations, microservices, and Python-based services supporting enterprise data ingestion, processing, and integration workflows.
Strong AWS experience including Amazon S3, SQS, Redshift, Aurora, cloud-based data processing, secure data integration, and scalable data architectures.
Experienced with Kubernetes container orchestration and Argo-based deployment and workflow automation for cloud-native data engineering applications.
Extensive Big Data experience using Apache Spark, PySpark, Hadoop, Hive, HDFS, Kafka, Spark Structured Streaming, and distributed processing frameworks.
Built and maintained 150+ production-grade ETL/ELT pipelines supporting batch, micro-batch, streaming, and near-real-time workloads.
Experienced processing 100M+ records daily while maintaining performance, reliability, data quality, and SLA requirements.
Strong experience developing Python applications for data ingestion, transformation, cleansing, validation, enrichment, automation, API integration, and distributed data processing.
Extensive SQL experience including complex queries, stored procedures, views, source-to-target mappings, data profiling, data validation, reconciliation, performance tuning, and database optimization.
Experienced building event-driven and microservices-based solutions integrating APIs, queues, databases, streaming platforms, and cloud storage.
Strong experience with data quality frameworks, CDC, incremental processing, reconciliation, exception handling, metadata-driven frameworks, and automated validation.
Experienced with CI/CD, Git, GitHub, Azure DevOps, Kubernetes, Argo, automated deployment, production support, Agile/Scrum, and cross-functional collaboration.
TECHNICAL SKILLS:
PROGRAMMING
Python, PySpark, SQL, Shell Scripting, Bash
AWS CLOUD & DATA SERVICES
Amazon S3, Amazon SQS, Amazon Redshift, Amazon Aurora, AWS Data Services, Cloud Storage, Event-Driven Architecture
BIG DATA & DISTRIBUTED PROCESSING
Apache Spark, PySpark, Spark SQL, Hadoop, Hive, HDFS, Spark Structured Streaming, DataFrames
DATA ENGINEERING
Data Engineering, ETL, ELT, Data Pipelines, Data Integration, Data Ingestion, Data Transformation, Batch Processing, Micro-Batch Processing, Near-Real-Time Processing, CDC, Incremental Loading, Data Quality, Data Validation, Data Reconciliation
EVENT-DRIVEN & STREAMING
AWS SQS, Apache Kafka, Azure Event Hub, Spark Structured Streaming, Event-Driven Architecture, Message Queues, Asynchronous Processing, Real-Time Data Processing
API & MICROSERVICES
REST APIs, API Development, API Integration, Microservices, Python Services, Service Integration, Authentication, Secure API Communication
CONTAINERIZATION & DEVOPS
Kubernetes, Argo, Argo Workflows, CI/CD, Git, GitHub, Azure DevOps, Infrastructure as Code, Deployment Automation
DATABASES
Amazon Aurora, Amazon Redshift, PostgreSQL, Oracle, SQL Server, Azure SQL Database, Snowflake, GraphDB, Stardog, Neo4j
AZURE DATA ENGINEERING
Azure Data Factory, Azure Databricks, Azure Synapse Analytics, ADLS Gen2, Azure SQL Database, Azure Event Hub, Azure Functions, Azure Key Vault
DATA ARCHITECTURE
Cloud Data Architecture, Lakehouse Architecture, Enterprise Data Architecture, Data Warehousing, Dimensional Modeling, Star Schema, Snowflake Schema, Data Vault, Fact/Dimension Modeling
DATA GOVERNANCE & SECURITY
Data Governance, Data Lineage, Metadata Management, Data Quality, Auditing, Role-Based Access Control, Encryption, Key Management, Compliance
CERTIFICATIONS
Google Cloud Professional Data Engineer
NCFM – Fundamental Analysis Module
NCFM – Commodity Markets Module
Goldman Sachs Operations Virtual Job Simulation
PROFESSIONAL EXPERIENCE:
American Express, Iselin, NJ June2024–Present
Lead Data Engineer
Project Overview: Enterprise Banking Data Lakehouse, AWS/Azure Data Platform & Knowledge Graph Modernization
Led architecture, design, development, deployment, and support of enterprise-scale data engineering solutions supporting AML, KYC, fraud detection, merchant analytics, customer analytics, regulatory reporting, and risk management.
Developed scalable cloud data pipelines using Python, PySpark, Apache Spark, Spark SQL, SQL, Amazon S3, Amazon Redshift, Amazon Aurora, Kafka, AWS SQS, Azure Data Factory, and Azure Databricks.
Designed AWS-based data ingestion architectures using Amazon S3 as scalable object storage for structured and semi-structured enterprise datasets.
Developed Python and PySpark applications for ingestion, transformation, cleansing, validation, enrichment, aggregation, and downstream analytics.
Built event-driven data processing solutions using AWS SQS and Kafka to support asynchronous and near-real-time data processing.
Designed and implemented microservices-based data integration components using Python and REST APIs.
Designed and developed scalable data engineering solutions using Databricks, PySpark, AWS, and Delta Lake for high-volume data processing and analytics.
Built and maintained bronze, silver, and gold medallion architecture in Databricks to support reliable and scalable Lakehouse data platforms.
Developed high-volume PySpark data pipelines for complex transformations, data cleansing, aggregation, and business-rule processing.
Worked extensively with Delta Lake, Unity Cata log, schema evolution, partitioning, and optimized table management to improve data reliability and performance.
Developed AWS-based data solutions using S3, EMR, Glue, Lambda, and Redshift, integrating multiple enterprise data sources into centralized data platforms.
Implemented CDC, SCD Type 2, idempotent processing, and incremental data loads to maintain accurate and historical datasets.
Developed real-time and near-real-time pipelines using Spark Structured Streaming, Kafka, Kinesis, and Databricks Auto Loader.
Built and orchestrated data workflows using Databricks Workflows, Airflow, dbt, and Delta Live Tables for automated pipeline execution and monitoring.
Implemented data quality checks, validation frameworks, schema checks, and error-handling mechanisms to improve the reliability of production data pipelines.
Worked on AI-assisted data engineering use cases, including automated data quality validation, schema analysis, anomaly detection, and AI-assisted PySpark transformation development.
Used GitHub Copilot and AI-assisted coding tools to accelerate development, debugging, code optimization, documentation, and rapid prototyping of data engineering solutions.
Built rapid POCs and MVPs for data and AI use cases by combining Databricks, PySpark, AWS services, and modern AI development tools.
Developed REST API integrations for enterprise applications, data sources, and downstream services.
Implemented secure API communication, authentication, authorization, error handling, retry mechanisms, and service integration patterns.
Developed microservices supporting data ingestion, transformation, validation, and integration workflows.
Built and maintained 150+ production-grade ETL/ELT pipelines supporting batch, micro-batch, streaming, and near-real-time workloads.
Processed and transformed 100M+ transaction records daily while maintaining performance, data quality, and SLA requirements.
Implemented Kafka-based streaming architectures integrated with Spark Structured Streaming and event-driven processing frameworks.
Developed Spark and PySpark pipelines for large-scale distributed data processing.
Designed data warehouse and analytical solutions using Amazon Redshift for high-volume analytical workloads.
Integrated transactional data sources with Amazon Aurora and developed optimized SQL-based extraction and transformation processes.
Developed SQL queries, stored procedures, transformation logic, source-to-target mappings, data validation, and reconciliation processes.
Implemented metadata-driven ingestion frameworks that reduced data onboarding effort by approximately 40%.
Developed automated data quality frameworks including schema validation, reconciliation, duplicate detection, anomaly detection, completeness validation, and exception handling.
Implemented CDC and incremental loading strategies for high-volume enterprise data pipelines.
Optimized Spark workloads, SQL queries, partitioning, caching, cluster configuration, and distributed processing, improving pipeline performance by approximately 60%.
Implemented Kubernetes-based deployment patterns for containerized data engineering services.
Utilized Argo and automated workflow/deployment processes to support cloud-native application delivery and data processing workflows.
Established CI/CD deployment processes using Git, GitHub, Azure DevOps, automated pipelines, and Infrastructure-as-Code practices.
Implemented enterprise security, role-based access, encryption, secure data integration, auditing, and compliance controls.
Led migration of 20+ TB of enterprise banking data from legacy on-premises platforms to cloud data platforms.
Collaborated with Data Architects, Data Scientists, Business Analysts, application teams, and business stakeholders to define data architecture and technical solutions.
Led and mentored Data Engineers and Analysts responsible for development, deployment, production support, and optimization.
CVS HEALTH, Remote, USA Jan2022 – May2024
Senior Azure Data Engineer
Project Overview
Healthcare Data Modernization & Cloud Data Integration
Responsibilities:
Designed and developed enterprise healthcare data integration solutions using Python, PySpark, SQL, AWS S3, Amazon Redshift, Amazon Aurora, Azure Data Factory, Azure Databricks, Azure SQL Database, and ADLS Gen2.
Developed scalable data ingestion pipelines using Python and Shell Scripting to process structured and semi-structured healthcare datasets.
Developed PySpark transformation frameworks for claims, pharmacy, provider, patient, member eligibility, and clinical datasets.
Developed data layers supporting AI/ML applications, feature engineering, vector-based data processing, embeddings, and Retrieval-Augmented Generation (RAG) use cases.
Designed reusable data pipelines and frameworks using Python and PySpark, following software engineering practices such as modular design, version control, testing, and code reviews.
Led complex data engineering initiatives involving architecture, development, optimization, production deployment, and technical coordination across multiple teams.
Optimized Spark and Databricks workloads through partitioning, caching, broadcast joins, file optimization, and efficient PySpark transformations.
Implemented automated deployment and version control practices using Git, CI/CD pipelines, and infrastructure/application deployment processes.
Worked with data contracts, schema management, and standardized ingestion patterns to maintain consistency across distributed data platforms.
Built scalable data solutions capable of processing large-volume batch and streaming workloads while maintaining performance, reliability, and data quality.
Collaborated with data scientists, analysts, application teams, and business stakeholders to convert data requirements into scalable Lakehouse and AI-enabled data solutions.
Built ETL/ELT pipelines supporting batch, incremental, micro-batch, and near-real-time data processing.
Implemented Amazon S3-based data ingestion and storage patterns for high-volume healthcare datasets.
Developed SQL-based data integration and analytical workloads supporting Amazon Redshift and relational databases.
Worked with Amazon Aurora for transactional data integration and SQL-based processing.
Implemented AWS SQS-based asynchronous processing patterns for decoupling data ingestion and downstream processing components.
Developed Kafka-based streaming solutions for event-driven healthcare data processing.
Built Python-based REST API integrations for enterprise healthcare applications and external data sources.
Developed microservices supporting healthcare data ingestion, validation, transformation, and integration processes.
Implemented Kubernetes-based deployment patterns for containerized data services and processing applications.
Used Argo-based workflow and deployment automation for cloud-native application environments.
Implemented metadata-driven ETL frameworks for standardized data ingestion and transformation.
Implemented incremental loading and CDC strategies that reduced processing windows by more than 50%.
Built automated data quality frameworks including reconciliation, completeness validation, duplicate detection, and business-rule verification.
Developed complex SQL queries, stored procedures, views, transformations, and source-to-target mappings.
Optimized SQL queries, Spark workloads, and data processing pipelines, achieving approximately 50% improvement in processing performance.
Supported secure and compliant healthcare data processing, governance, auditing, and access control.
Collaborated with Data Architects, Business Analysts, application teams, QA teams, DevOps teams, and healthcare stakeholders.
UNITEDHEALTH GROUP, Remote Dec 2020 – Dec 2021
Data Engineer
Project Overview
Enterprise Healthcare Analytics Platform
Responsibilities
Designed, developed, and maintained enterprise ETL workflows supporting healthcare claims, eligibility, provider, member, and pharmacy datasets.
Built scalable data pipelines using Python, SQL, Shell Scripting, PySpark, and distributed data processing technologies.
Developed cloud-based data ingestion workflows using Amazon S3 and relational data platforms.
Developed SQL queries, stored procedures, views, database objects, and transformation logic supporting healthcare analytics.
Implemented AWS SQS-based asynchronous processing patterns for data integration workflows.
Developed API integrations for enterprise healthcare applications and data sources.
Developed Python-based services and microservices for data ingestion, validation, and transformation workflows.
Implemented Kafka-based event-driven data processing patterns for high-volume data integration.
Built reporting data marts and analytical datasets supporting operational teams, business users, and leadership.
Performed data profiling, validation, reconciliation, cleansing, and data quality activities.
Developed automated auditing and exception-handling processes to identify data integrity issues.
Optimized SQL queries and database performance to improve ETL processing and reporting execution.
Created source-to-target mappings, data dictionaries, technical documentation, and process flow documentation.
Participated in Agile development processes including sprint planning, backlog grooming, code reviews, testing, and release deployments.
HDFC Bank, Hyderabad, India July 2017-Oct2019
Data Engineer
Project: Enterprise Banking Data Warehouse & Regulatory Reporting Platform
Responsibilities
Designed and developed enterprise ETL solutions supporting retail banking, loans, deposits, credit cards, customer analytics, and regulatory reporting.
Built scalable data pipelines using Python, SQL, Shell Scripting, and ETL technologies.
Developed high-volume data integration processes for core banking systems, transactional data, customer information, and financial datasets.
Developed complex SQL queries, stored procedures, functions, views, and database objects supporting operational and regulatory reporting.
Developed data transformation and integration processes for large-scale banking datasets.
Supported cloud migration and data integration patterns involving object storage, relational databases, and analytical data platforms.
Designed dimensional data models using fact tables and dimension tables.
Supported enterprise data integration and relationship-based data analysis across customers, accounts, products, and transactions.
Optimized database performance and ETL workflows, reducing report generation time by approximately 40%.
Implemented data quality checks, validation rules, exception handling, and audit controls.
Developed source-to-target mappings, data dictionaries, technical specifications, and ETL documentation.
Collaborated with business users, risk teams, compliance stakeholders, and technical teams to understand data requirements.
EDUCATION
Master of Science in Business Analytics Saint Peter’s University Jersey City, NJ
Bachelor of Business Administration in Financial Markets St. Joseph’s Degree & PG College Hyderabad,
India