Post Job Free
Sign in

Senior Data Engineer, Data Lakehouse Pipelines

Location:
Overland Park, KS
Posted:
October 07, 2026

Contact this candidate

Resume:

CONTACT

434-***-****

**************@*****.***

linkedin.com/in/swaroop-batchu Swaroop Batchu

github.com/BatchuSwaroop

SENIOR DATA ENGINEER

SKILLS

Programming & Frameworks PROFILE

Python, SQL, PySpark, PyArrow, FastAPI, Flask, Senior Data Engineer with 5+ years of experience designing and delivering scalable

Django, SQLAlchemy batch, streaming, ETL/ELT, data lake, lakehouse and enterprise data warehouse solutions

Databases & Warehousing across AWS, Azure and GCP. Strong expertise in Python, PySpark, SQL, Apache Spark,

PostgreSQL, SQL Server, Oracle, Amazon Redshift,

Kafka, Hadoop, Snowflake, Redshift, Synapse Analytics and cloud-native data services,

Redshift Serverless, Azure SQL Data Warehouse

Data Engineering & Big Data with proven experience in workflow orchestration, infrastructure automation, CI/CD,

AWS Glue, Glue Data Catalog, EMR Serverless, data quality, machine-learning pipelines and RAG-based solutions using LangChain,

Databricks, Apache Spark, Azure Data Factory, Hugging Face Transformers and Amazon OpenSearch Service.

HDInsight

Streaming & Messaging EXPERIENCE

Apache Kafka (MSK), Kinesis Data Firehose,

EventBridge, SQS, Azure Event Hub, Service Bus

Senior Data Engineer April 2025 - Present

Data Lake & Analytics

JPMorgan Chase, New York, NY

Amazon S3, Azure Blob Storage, Apache Iceberg, • Developed scalable financial data engineering applications using Python, building

Athena, Redshift Spectrum

reusable processing modules, validation frameworks and optimized workflows for

Orchestration & DevOps

high-volume transaction, customer and account and risk datasets.

Apache Airflow, Amazon MWAA, Step Functions,

GitHub Actions, Jenkins, CodeBuild, Bitbucket • Engineered reliable persistence and metadata layers using Azure Database for

Cloud & Infrastructure PostgreSQL, PostgreSQL and SQLAlchemy, supporting financial transaction processing,

AWS, Azure, EKS, ECS/Fargate, Docker, Azure data lineage, auditability and efficient database operations.

Container Instance, Terraform, CloudFormation, • Built high-performance financial data APIs using FastAPI to securely expose curated

ARM

datasets, pipeline status, metadata and analytical outputs to banking applications and

AI / ML & GenAI

SageMaker Pipelines, SageMaker Jumpstart,

downstream reporting systems.

Hugging Face Transformers, LangChain, RAG, • Created serverless analytics solutions using Azure Synapse Serverless SQL Pool to query

OpenSearch, Batch Inference structured and semi-structured banking data across transaction, payment, customer,

BI, Quality & Monitoring compliance and operational datasets.

Power BI, QuickSight, Great Expectations, • Designed and optimized enterprise financial data warehouse solutions using Azure

CloudWatch

Connectivity & Networking

Synapse Analytics, supporting regulatory reporting, portfolio analytics, risk analysis and

JDBC, PyODBC, IAM, VPC, NAT Gateway, VPN business intelligence workloads.

Gateway, EFS, FSx • Developed event-driven and real-time financial data pipelines using Azure Event Grid

and Azure Event Hubs, enabling high-volume transaction processing, event distribution,

CERTIFICATIONS partitioning and resilient failure recovery.

• AWS Certified Solution Architect Associate.

• Built scalable banking ETL pipelines using Azure Data Factory, performing ingestion,

• Microsoft Certified Azure Data Engineer Associate.

cleansing, transformation and cataloging of customer, account, payment and financial

EDUCATION reference data.

MASTERS, 2026 • Processed large-scale batch and distributed financial workloads using Azure Databricks,

UNIVERSITY OF CENTRAL MISSOURI improving scalability for transaction aggregation, reconciliation, risk calculations and

Bachelor in Computer Science, 2022 historical data processing.

Bharath Institute Of Higher Education & • Implemented financial lakehouse architectures using Apache Iceberg, supporting

Research

schema evolution, partition management, transactional updates and historical analysis

of banking and financial datasets.

• Containerized and deployed data-processing services on Azure Kubernetes Service

(AKS), enabling scalable and highly available execution of financial ingestion,

transformation and analytical workloads across environments.

• Provisioned reusable banking data-platform infrastructure using Terraform, supporting

secure storage, compute, networking, access controls and standardized Azure cloud

deployments.

• Developed machine-learning workflows using Azure Machine Learning Pipelines, Azure

Machine Learning Model Catalog and Hugging Face Transformers for financial data

preparation, model training and model initialization and embedding generation.

• Designed secure financial data-lake storage using Azure Data Lake Storage Gen2 (ADLS

Gen2) and orchestrated end-to-end banking workflows through Azure Data Factory

Pipelines, managing dependencies, retries, scheduling and recovery across critical data

pipelines.

• Automated code validation, testing, packaging and deployments using GitHub Actions,

while implementing financial data-quality controls and pipeline observability with Great

Expectations.

• Designed RAG pipelines using LangChain and Azure AI Search, enabling document

ingestion, chunking, metadata enrichment, semantic retrieval and contextual search

across financial documents, policies and operational knowledge repositories.

Senior Data Engineer Sep 2024 - Mar 2025

Best Buy, Richfield, MN,

• Developed scalable retail data pipelines using Python for product, order, inventory,

customer, pricing and transaction datasets, improving reliable data availability across

downstream systems.

• Designed and optimized MySQL data models using indexing, partitioning, query tuning,

stored procedures and feature datasets supporting analytics and Scikit-learn workloads.

• Built enterprise integrations using JDBC, enabling reliable movement of retail data across

databases, distributed processing platforms and downstream business applications.

• Developed REST-based data services using Flask, exposing curated retail datasets,

analytical outputs and lightweight OpenAI API-driven classification and summarization

capabilities.

• Created business-focused reporting solutions in Tableau, combining sales, inventory,

fulfillment and customer KPIs with predictive outputs generated through Scikit-learn.

• Engineered scalable analytical workloads using Databricks SQL Warehouse and Azure

Synapse Analytics, preparing governed datasets for forecasting, segmentation,

recommendations and MLflow-tracked experimentation.

• Implemented asynchronous integrations using RabbitMQ, enabling reliable event

delivery across order, fulfillment, analytics and AI-enabled processing services.

• Built real-time pipelines using Apache Flink and ETL workflows with Talend, processing

clickstream, order, inventory and customer events with Hugging Face

Transformers-based text workflows.

• Processed large-scale historical retail datasets using Apache Hadoop and Apache Hive

for sales trend analysis, customer segmentation, forecasting and machine-learning

feature generation.

• Containerized and deployed distributed data services using Kubernetes, including

scalable Flask APIs and machine-learning inference services supporting Azure-hosted

retail platforms.

• Developed predictive pipelines using Scikit-learn and Azure Machine Learning for

forecasting, segmentation and propensity modeling, with experiment tracking and

lifecycle management through MLflow.

• Built semantic retrieval and knowledge-enrichment workflows using LangChain and

FAISS over product, catalog and customer-support datasets stored in Azure Blob Storage.

• Orchestrated end-to-end retail data and ML workflows using Prefect, automated

deployments through Jenkins and supported analytical workloads in Teradata for

downstream reporting and data products.

Data Engineer June 2022 - July

2024

Cognizant, India

• Developed scalable batch and distributed data-processing solutions using Python,

PySpark and SQL, supporting banking transaction, customer, risk and enterprise

client datasets across consulting engagements.

• Designed and optimized relational data platforms using SQL Server, integrating

applications and enterprise systems through PyODBC and JDBC to support secure

financial data processing and reporting.

• Built lightweight data services and REST APIs using Flask, exposing curated banking

datasets, metadata, pipeline status and analytical outputs to downstream

applications and client platforms.

• Developed interactive financial and operational dashboards using Power BI,

providing insights into transaction trends, risk exposure and business KPIs for

banking and consulting stakeholders.

• Designed and maintained enterprise data warehouse solutions using Azure Synapse

Analytics, implementing dimensional models, bulk-loading strategies, distribution

optimization and high-performance queries for financial reporting and client

analytics.

• Engineered event-driven workflows using Azure Service Bus and Azure Logic Apps,

enabling asynchronous transaction processing, workflow coordination, automated

retries and resilient recovery across enterprise data pipelines.

• Built streaming and ETL frameworks using Azure Event Hubs, Azure Data Factory

and Azure Purview, supporting ingestion, transformation, metadata governance and

discovery of banking and client data assets.

• Processed large-scale financial and enterprise datasets using Apache Spark and

Azure Databricks, optimizing distributed joins, aggregations and reusable

transformation frameworks for analytics and reporting workloads.

• Containerized data-processing applications using Docker and deployed scalable

workloads through Azure Container Instances, ensuring consistent execution

environments and simplified operations across banking and consulting projects.

• Provisioned repeatable cloud data infrastructure using ARM Templates, creating

standardized templates for compute, storage, networking and security resources

across multiple client and financial environments.

• Configured secure cloud environments using Azure Blob Storage, Azure Virtual

Machines, Azure RBAC, Azure Virtual Network, Azure NAT Gateway and Route

Tables, supporting protected financial data storage, controlled access and private

connectivity across enterprise workloads.

• Implemented shared and high-performance storage solutions using Azure Files and

Azure NetApp Files, supporting distributed processing, analytics and data-intensive

applications across banking and client delivery environments.

• Orchestrated enterprise data pipelines using Apache Airflow and automated build,

testing and deployment workflows through Azure Pipelines and Jenkins, ensuring

consistency across data engineering releases.

• Queried large-scale data-lake environments using Azure Synapse Serverless SQL

Pool and Azure Data Lake Storage Gen2, developed Batch Inference Pipelines for

financial and enterprise prediction workloads and monitored platform health and

pipeline performance using Azure Monitor.

Data Engineer July 2021 - May

2022

Bharat Electronics Limited (BEL), India

• Engineered scalable data-processing applications using Python and PyArrow to

handle structured and columnar datasets across banking, defense electronics and

aerospace and government technology workloads.

• Developed and optimized enterprise database solutions with Oracle, enabling

reliable application and data-platform connectivity through JDBC Connector for

financial, mission-critical and operational systems.

• Built secure backend data services and web interfaces using Django, providing

controlled access to financial, defense and government datasets, metadata and

pipeline execution information.

• Created analytical dashboards and operational reports using Amazon QuickSight,

enabling banking, defense and government stakeholders to evaluate trends,

performance metrics, system readiness and operational outcomes.

• Managed enterprise transactional and analytical workloads using Oracle Database

on Amazon EC2, supporting dependable storage, availability and processing for

banking and mission-critical applications.

• Developed event-driven and serverless processing solutions using Amazon Kinesis

and AWS Lambda, enabling real-time ingestion and automated execution for

financial transactions, telemetry and operational data.

• Implemented asynchronous messaging and system-integration workflows with

Amazon SQS and Amazon SNS, supporting queues, publish-subscribe messaging

and resilient communication across distributed banking and defense applications.

• Designed Python-integrated ETL workflows using AWS Glue to extract, transform

and transfer financial, engineering and operational data across cloud and

on-premises environments.

• Executed large-scale distributed processing and analytical workloads using Amazon

EMR, supporting batch transformations across transaction, telemetry, logistics and

mission-support datasets.

• Deployed lightweight containerized Python data services with Amazon ECS and

automated repeatable infrastructure provisioning through AWS CloudFormation

across enterprise environments.

• Established secure cloud environments using Amazon EC2, Amazon S3 and AWS

Site-to-Site VPN, supporting compute, protected data storage and hybrid

connectivity for banking, aerospace and government technology workloads.

• Coordinated end-to-end Python pipeline execution using Apache Airflow while

managing source control and CI/CD activities through Bitbucket, improving

deployment consistency and operational reliability.

• Architected and optimized enterprise analytics and data warehouse solutions using

Amazon Redshift, supporting large-scale financial reporting, defense analytics and

government operational intelligence.



Contact this candidate