Lokesh Reddy
Email: ****************@*****.*** Contact:636-***-****
SUMMARY
Data Engineer with 6+ years of experience designing and implementing scalable, secure, and high-performing data solutions across Azure, AWS, and Big Data ecosystems.
Hands-on expertise in Azure Cloud Services including Azure Data Factory (ADF), Azure Databricks, Azure Synapse Analytics, Azure Event Hubs, Azure Blob Storage, and Azure CosmosDB.
Strong experience in AWS cloud services like S3, Redshift, Glue, Lambda, EMR, IAM, and RDS, building cost-efficient, secure data lakes and data warehouse solutions.
I am expert in building complex ETL pipelines using Azure Data Factory, including automated triggers, parameterized pipelines, secure integration with Azure Key Vault, and real-time ingestion with Event Hubs.
Highly skilled in Big Data technologies like Apache Spark, Hive, Hadoop (HDFS, MapReduce, Yarn), Kafka, Pig, Flume, Sqoop, and HBase for large-scale batch and streaming data processing.
Proficient in PySpark, Spark SQL, and Scala for building distributed data transformations and advanced analytics pipelines within Azure Databricks and AWS EMR environments.
I specialized in building Delta Lake architecture using Databricks, implementing Delta Live Tables (DLT) for real-time, reliable data ingestion and SCD Type 3 transformations.
Deep knowledge in Snowflake development, including Multi-Cluster Warehouses, Time Travel, Zero-Copy Cloning, Secure Views, Role-Based Access Control (RBAC), and performance optimization.
Strong expertise in Snowflake cost control strategies like resource monitors, auto-suspend configurations, and compute resource optimization.
Experienced in setting up real-time data integration using Azure Event Hub, Change Data Capture (CDC) mechanisms, and Azure Stream Analytics.
Skilled in integrating data platforms with Business Intelligence tools like Power BI and Tableau to design insightful dashboards, perform DAX-based time series analysis, and enable executive reporting.
Proficient in using Business Objects (Web Intelligence, Crystal Reports, Universe Designer) for enterprise-wide reporting, ad-hoc querying, and dashboard development.
Expertise in Data Governance and Data Cataloging using Azure Purview and AWS Glue Data Catalog, enabling data discovery, lineage tracking, and compliance.
Adept at setting up Hadoop clusters using Cloudera Manager, configuring services, monitoring cluster health, and optimizing resource utilization.
Developed secure, scalable data storage solutions using Azure Data Lake Gen1/Gen2, AWS S3, and HDFS, implementing security best practices like encryption and fine-grained access control.
Experienced in building CI/CD pipelines using Azure DevOps, Git, GitHub, Jenkins, Terraform, and Docker for automated, reliable data pipeline deployments.
Strong skills in scripting with Python and SQL for data validation, data mapping, ETL process automation, and workflow orchestration.
Worked extensively with Apache Oozie to create complex workflows and coordinators for managing Hadoop data pipelines.
I am proficient in Hive optimization techniques like Partitioning, Bucketing, Map-side Joins, Skew Joins, and Indexing to enhance query performance on massive datasets.
Experience integrating Spark Streaming and Kafka for real-time data ingestion and processing to support near real-time analytics.
Expertise in Teradata, Oracle, PostgreSQL, and MySQL databases for developing and optimizing complex SQL queries, stored procedures, and analytical reports.
Familiar with data serialization formats like AVRO, Protobuf, and JSON, used in batch and real-time data pipelines
I am skilled in leveraging SnowSQL and SnowPipe for efficient bulk loading and continuous data ingestion into Snowflake environments.
Hands-on experience in setting up Infrastructure as Code (IaC) using Terraform to automate cloud resource provisioning in Azure and AWS.
Knowledge of Data Vault modeling principles along with Dimensional and Star/Snowflake schemas for scalable data architecture.
Implemented robust authentication and authorization solutions using Azure Active Directory, AWS IAM roles and policies, and secure secrets management through Azure Key Vault and AWS KMS.
Designed and developed modular data transformation pipelines using DBT, improving data quality and maintainability across Snowflake and Redshift environments.
Experienced in Agile and Scrum methodologies, actively managing stakeholder expectations, participating in sprint planning, story grooming, and ensuring on-time, quality delivery of data products.
Built and optimized data pipelines that integrate PySpark with BI tools like Tableau for efficient data visualization and reporting
TECHNICAL TOOLS
Azure Services
Azure Data Factory, Azure Data Bricks, snowflake, Logic Apps, Functional App, Azure DevOps
Aws Services
Amazon S3, Amazon EC2, Amazon RDS, Amazon Redshift, AWS Glue, AWS Lambda, Amazon EMR, Amazon Kinesis, Amazon Athena, Amazon SQS
Big Data Technologies
MapReduce, Hive, Teg, PySpark, Scala, Kafka, Spark streaming, Oozie, Sqoop, Zookeeper, Flume,
Hadoop Distribution
Cloudera, Horton Works, YARN, HDFS,
Languages
SQL, PL/SQL, Python, Java, HiveQL, Scala, C++, Shell,
Web Technologies
HTML, CSS, JavaScript, XML, JSP, Restful, SOAP
Operating Systems
Windows (XP/7/8/10), UNIX, LINUX, UBUNTU, CENTOS.
Build Automation tools
Ant, Maven
Version Control
GIT, GitHub, BitBucket
IDE & Build Tools, Design
Eclipse, Visual Studio.
Databases
MS SQL Server 2016/2014/2012, Azure SQL DB, Azure Synapse. MS Excel, MS Access, Oracle 11g/12c, Cosmos DB, Google Big Query, MongoDB, Cassandra
EDUCATION: Master’s in Information Technology from Webster University.
PROFESSIONAL EXPERIENCE
Travelers, New York, NY
Dec 2023 – Present
Sr Azure Data Engineer
Responsibilities
Designed and implemented scalable data ingestion pipelines using Azure Data Factory, Azure Event Hubs, and Azure Logic Apps for ingesting structured and unstructured data from SQL databases, CSV files, JSON files, and REST APIs in the insurance domain.
Utilized PolyBase for seamless integration of heterogeneous insurance datasets across SQL Server and Azure SQL Database, optimizing analytics and reporting.
Automated and created Azure Databricks notebooks with Spark SQL, Scala, and Python for performing policyholder analytics, premium calculations, and risk modeling.
Developed distributed data processing workflows using Azure Databricks and Apache Spark for complex ETL tasks and transformations of insurance data, including underwriting, claims processing, and fraud detection.
Built custom operator tasks in Azure Integration Services using Python for managing policy updates, payment events, and tracking the claims lifecycle.
Developed real-time streaming jobs using Apache Kafka to capture policy creation, claims status changes, and fraud alerts, and store data in enterprise data repositories.
Designed ETL pipelines processing healthcare and financial data using a variety of serialization formats, including JSON, AVRO, and Protobuf, optimizing for efficient batch and streaming ingestion into cloud data lakes.
Implemented Snowflake on Azure to build scalable, secure, and high-performance data warehousing solutions, improving data ingestion, storage, and query performance for enterprise analytics.
Orchestrated insurance data workflows using Azure Logic Apps to automate actions and integrate services triggered by events such as claim submissions and renewals.
Deployed Azure Function Apps for serverless processing tasks, including quote generation, eligibility checks, and data validation for insurance applications.
Leveraged Azure Machine Learning for building and deploying predictive models for churn prediction, claims classification, and customer segmentation in the insurance industry.
Managed OLTP systems, supporting real-time transaction processing for insurance policy renewals, claims adjustments, and payments, utilizing optimized T-SQL queries for SQL Server.
Designed and maintained comprehensive insurance reports using SQL Server Reporting Services (SSRS) to track KPIs such as loss ratio, customer retention, and claim settlement time.
Worked with RDBMS platforms such as MySQL, PostgreSQL, Oracle, and SQL Server for managing structured insurance data, ensuring cross-system data consistency.
Tuned Spark workloads on Azure HDInsight using Apache YARN to enhance batch processing for claims analytics and risk assessments.
Utilized PySpark for ETL operations to clean, transform, and standardize raw insurance data from brokers and agencies.
Designed scalable data storage solutions using Azure Data Lake Storage Gen2 for long-term archival of insurance data and enabling historical analytics.
Developed robust ETL pipelines in Azure Databricks to integrate policyholder data into the gold layer of the data lake.
Implemented and managed Delta Lake on Azure to maintain versioned insurance datasets with ACID transactions, improving data consistency and integrity.
Streamlined pipeline deployment using Azure DevOps and CI/CD, reducing deployment time and enhancing pipeline reliability.
Utilized Azure Event Hubs for real-time claims updates, payment confirmations, and business data streaming into enterprise data platforms.
Built pipelines to migrate legacy SQL Server insurance data to Azure SQL Database using Azure Data Factory.
Skilled in PL/SQL for building business logic and data transformations for migration and underwriting applications.
U-SQL scripting is applied to process large volumes of insurance data within Azure Data Lake Analytics.
Adopted Agile (SCRUM) methodologies for managing insurance data projects, driving iterative development and continuous delivery.
Proficient in MDX for querying OLAP cubes and DAX for building Power BI dashboards showcasing key insurance metrics and KPIs.
Python Lambda functions in serverless workflows for lightweight processing tasks in the insurance sector.
Built globally distributed NoSQL databases using Azure Cosmos DB to support low-latency, cross-region insurance applications.
Implemented robust data validation tests (uniqueness, not null, referential integrity) using DBT, ensuring trusted datasets for analytics and reporting.
CVS Health Dallas, TX
Aug 2021 - Nov 2023
Data Engineer
Responsibilities
Designed scalable data ingestion pipelines using AWS Glue, EC2, Lambda, and S3 to ingest structured and unstructured healthcare data (e.g., electronic health records (EHR), medical claims, patient demographics) from SQL databases, CSV/JSON files, and REST APIs.
Built distributed ETL workflows using Apache Spark, PySpark, Databricks, DBT, and Delta Lake to optimize schema evolution, transformation logic, and performance tuning for healthcare analytics and claims processing.
Implemented the Medallion Architecture using AWS Data Lake and Snowflake to support clean, curated data layers, streamlined analytical workflows, with a focus on patient care and insurance claims.
Ensured data reliability with Great Expectations, DBT, and AWS Glue for validation, anomaly detection, schema drift handling, and cleansing across sensitive healthcare and insurance datasets.
Architected a cloud-based data warehouse in AWS Redshift, integrated with AWS S3 and AWS Glue, enabling scalable analytics for patient data, insurance claims, and provider transactions.
Developed Snowflake data models and schemas using DBT and Redshift, applying best practices like clustering, materialized views, and workload management for optimized query performance in healthcare and insurance data sets.
Used Apache Iceberg on Amazon S3 to manage patient monitoring data, enabling schema evolution, ACID-compliant updates, and time travel for securely tracking historical health records in a HIPAA-compliant data lake house.
Built scalable ingestion pipelines and deployed PySpark workloads on Amazon EKS, with S3 used as central raw and processed storage.
Experience with AWS Event Bridge for event-driven workflows and alerts across distributed healthcare systems.
Built AI-powered chatbots with Azure Bot Framework and Cognitive Services to automate customer interactions, manage FAQs, and improve response time in healthcare and insurance operations.
Developed and optimized Apache Spark jobs in Databricks using Python (PySpark) for advanced healthcare analytics, predictive modeling, and bulk transformation of clinical and claims datasets.
Integrated Apache Airflow, AWS Redshift, AWS Glue, and Databricks to orchestrate distributed data workflows and enable real-time insights for healthcare decision-making and fraud detection.
Configured event-driven triggers using AWS Lambda and Databricks Jobs API to automate data ingestion, transformation, and alerts for time-sensitive healthcare events and claim updates.
Automated modular ETL pipelines using DBT to enforce consistency and reusability in staging and production environments for insurance and EHR datasets.
Designed robust ETL workflows for multi-source ingestion from APIs, AWS S3, Redshift, and RDS, supporting integrated views of healthcare providers, insurance claims, and clinical documents.
Implemented CI/CD pipelines with AWS Code Pipeline, Jenkins, Terraform, DBT, and Databricks for streamlined deployments and reliable data engineering practices.
Integrated Power BI and Tableau with Databricks and Redshift to build self-service dashboards showcasing patient outcomes, claim trends, and operational KPIs for healthcare organizations.
Collaborated with data scientists, analysts, and clinical teams to gather domain-specific requirements and deliver scalable, governed solutions for healthcare intelligence and insurance workflows.
Practiced Agile delivery using JIRA, AWS Code Commit, Scrum, and Kanban, participating in sprint ceremonies, backlog grooming, and PI Planning for cross-functional healthcare/insurance projects.
Developed responsive web interfaces using Angular and TypeScript, integrating RESTful APIs, real-time updates, and secure authentication via OAuth, JWT, and Azure AD.
Leveraged AWS Compute and Storage services (EC2, Lambda, S3, Redshift, Glue) to ensure cost-effective, high-availability data platforms for healthcare and insurance analytics.
Designed HIPAA-compliant EHR storage using AWS S3 and Redshift, processing sensitive data through AWS Glue pipelines with strict access control and encryption.
Developed an interactive dashboard using Dash to visualize patient vitals, streaming alerts, and model predictions in real-time, enabling clinical users to track high-risk cases easily.
Supported the integration of historical EHR data by extending existing Ab Initio frame works, PDL-based parameterization to transform and align datasets with modern Delta Lake pipelines.
Automated data pipeline deployments by integrating DBT with CI/CD workflows (GitHub Actions, GitLab CI), reducing manual errors and deployment times.
Built real-time EHR data pipelines using AWS Lambda and Amazon Kinesis, enabling low-latency updates and near real-time insights for clinical and operational systems.
Wells Fargo, San Francisco, Ca
Oct 2019 – Jul 2021
Data Engineer
Responsibilities
Designed scalable data ingestion pipelines using AWS Glue, EC2, Lambda, and S3 to ingest structured and unstructured banking data (e.g., customer transaction histories, loan records, credit card statements, and financial reports) from SQL databases, CSV/JSON files, and REST APIs.
Built distributed ETL workflows using Apache Spark, PySpark, Databricks, DBT, and Delta Lake to optimize schema evolution, transformation logic, and performance tuning for banking analytics and transaction processing.
Implemented Medallion Architecture (Bronze, Silver, Gold) using AWS Data Lake and Snowflake to support clean, curated data layers, ensuring compliance with financial regulations (e.g., Basel III, GDPR) and enabling efficient reporting for banking operations.
Ensured data reliability with Great Expectations, DBT, and AWS Glue for validation, anomaly detection, schema drift handling, and cleansing across sensitive banking datasets like financial transactions, loan data, and customer accounts.
Architected a cloud-based data warehouse in AWS Redshift, integrated with AWS S3 and AWS Glue, enabling scalable analytics for financial transactions, credit risk analysis, customer account management, and compliance reporting.
Designed and built scalable ELT pipelines using PySpark and DBT, extracting data from Kafka and RDBMS sources, loading into S3, and transforming within Apache Iceberg tables to support analytics and reporting.
Developed Snowflake data models and schemas using DBT and Redshift, applying best practices like clustering, materialized views, and workload management for optimized query performance in banking applications.
Built AI-powered chatbots with Azure Bot Framework and Cognitive Services to automate customer interactions, manage FAQ, and improve response time in banking operations like account inquiries, transaction management, and fraud alerts.
Implemented a data Lakehouse using Apache Iceberg on Amazon S3, enabling schema evolution, time travel, and ACID-compliant data processing for financial transaction analytics.
Developed and optimized Apache Spark jobs in Databricks using Python (PySpark) for advanced banking analytics, fraud detection, and bulk transformation of transactional datasets.
Used Dremio to run federated SQL queries directly on Iceberg and Delta Lake tables stored in S3, reducing data duplication and enabling real-time reporting through Power BI.
Integrated Apache Airflow, AWS Redshift, AWS Glue, and Databricks to orchestrate distributed data workflows and enable real-time insights for banking decision-making and fraud detection.
Configured event-driven triggers using AWS Lambda and Databricks Jobs API to automate data ingestion, transformation, and alerts for time-sensitive banking events like transaction approvals, loan status updates, and payment processing.
Automated modular ETL pipelines using DBT to enforce consistency and reusability in staging and production environments for transaction data and customer datasets in banking.
Designed robust ETL workflows for multi-source ingestion from APIs, AWS S3, Redshift, and RDS, supporting integrated views of customer banking activity, transaction histories, and loan processing data.
Implemented CI/CD pipelines with AWS Code Pipeline, Jenkins, Terraform, DBT, and Databricks for streamlined deployments and reliable data engineering practices in banking environments.
Integrated Power BI and Tableau with Databricks and Redshift to build self-service dashboards showcasing financial KPIs, transaction trends, and customer insights for banking organizations.
Collaborated with data scientists, analysts, and financial teams to gather domain-specific requirements and deliver scalable, governed solutions for banking intelligence, fraud detection, and risk management workflows.
Practiced Agile delivery using JIRA, AWS Code Commit, Scrum, and Kanban, participating in sprint ceremonies, backlog grooming, and PI Planning for cross-functional banking projects.
Developed responsive web interfaces using Angular and TypeScript, integrating RESTful APIs, real-time updates, and secure authentication via OAuth, JWT, and Azure AD for banking applications.
Leveraged AWS Compute and Storage services (EC2, Lambda, S3, Redshift, Glue) to ensure cost-effective, high-availability data platforms for banking analytics and operations.
Designed secure, compliant storage solutions for sensitive banking data using AWS S3 and Redshift, ensuring strict access control and encryption for regulatory compliance in the banking industry.
Created monitoring dashboards using KQL (Kusto Query Language) integrated with Power BI, Azure Monitor, and Application Insights to visualize financial performance and alerting metrics for banking applications.
Built real-time banking transaction pipelines using AWS Lambda and Amazon Kinesis, enabling low-latency updates and near real-time insights for fraud detection and operational efficiency in banking systems.
Enhanced data ingestion pipelines by implementing incremental loading strategies with AWS Lambda and AWS S3, improving processing time for large-scale transactional datasets.
Optimized batch processing of loan application data by implementing AWS Glue jobs and PySpark transformations for faster approval and underwriting workflows.
Developed automated data validation scripts using DBT and Great Expectations to ensure high-quality financial data for credit risk models and compliance reporting.