CONTACT
**************@*****.***
linkedin.com/in/swaroop-batchu Swaroop Batchu
github.com/BatchuSwaroop
SENIOR DATA ENGINEER
SKILLS
Programming & Frameworks PROFILE
Python, SQL, PySpark, PyArrow, FastAPI, Flask, Senior Data Engineer with 5+ years of experience designing and delivering scalable
Django, SQLAlchemy batch, streaming, ETL/ELT, data lake, lakehouse and enterprise data warehouse solutions
Databases & Warehousing across AWS, Azure and GCP. Strong expertise in Python, PySpark, SQL, Apache Spark,
PostgreSQL, SQL Server, Oracle, Amazon Redshift,
Kafka, Hadoop, Snowflake, Redshift, Synapse Analytics and cloud-native data services,
Redshift Serverless, Azure SQL Data Warehouse
Data Engineering & Big Data with proven experience in workflow orchestration, infrastructure automation, CI/CD,
AWS Glue, Glue Data Catalog, EMR Serverless, data quality, machine-learning pipelines and RAG-based solutions using LangChain,
Databricks, Apache Spark, Azure Data Factory, Hugging Face Transformers and Amazon OpenSearch Service.
HDInsight
Streaming & Messaging EXPERIENCE
Apache Kafka (MSK), Kinesis Data Firehose,
EventBridge, SQS, Azure Event Hub, Service Bus
Senior Data Engineer April 2025 - Present
Data Lake & Analytics
JPMorgan Chase, New York, NY
Amazon S3, Azure Blob Storage, Apache Iceberg, • Developed scalable financial data engineering applications using Python, building
Athena, Redshift Spectrum
reusable processing modules, validation frameworks and optimized workflows for
Orchestration & DevOps
high-volume transaction, customer and account and risk datasets.
Apache Airflow, Amazon MWAA, Step Functions,
GitHub Actions, Jenkins, CodeBuild, Bitbucket • Engineered reliable persistence and metadata layers using Azure Database for
Cloud & Infrastructure PostgreSQL, PostgreSQL and SQLAlchemy, supporting financial transaction processing,
AWS, Azure, EKS, ECS/Fargate, Docker, Azure data lineage, auditability and efficient database operations.
Container Instance, Terraform, CloudFormation, • Built high-performance financial data APIs using FastAPI to securely expose curated
ARM
datasets, pipeline status, metadata and analytical outputs to banking applications and
AI / ML & GenAI
SageMaker Pipelines, SageMaker Jumpstart,
downstream reporting systems.
Hugging Face Transformers, LangChain, RAG, • Created serverless analytics solutions using Azure Synapse Serverless SQL Pool to query
OpenSearch, Batch Inference structured and semi-structured banking data across transaction, payment, customer,
BI, Quality & Monitoring compliance and operational datasets.
Power BI, QuickSight, Great Expectations, • Designed and optimized enterprise financial data warehouse solutions using Azure
CloudWatch
Connectivity & Networking
Synapse Analytics, supporting regulatory reporting, portfolio analytics, risk analysis and
JDBC, PyODBC, IAM, VPC, NAT Gateway, VPN business intelligence workloads.
Gateway, EFS, FSx • Developed event-driven and real-time financial data pipelines using Azure Event Grid
and Azure Event Hubs, enabling high-volume transaction processing, event distribution,
CERTIFICATIONS partitioning and resilient failure recovery.
• AWS Certified Solution Architect Associate.
• Built scalable banking ETL pipelines using Azure Data Factory, performing ingestion,
• Microsoft Certified Azure Data Engineer Associate.
cleansing, transformation and cataloging of customer, account, payment and financial
EDUCATION reference data.
MASTERS, 2026 • Processed large-scale batch and distributed financial workloads using Azure Databricks,
UNIVERSITY OF CENTRAL MISSOURI improving scalability for transaction aggregation, reconciliation, risk calculations and
Bachelor in Computer Science, 2022 historical data processing.
Bharath Institute Of Higher Education & • Implemented financial lakehouse architectures using Apache Iceberg, supporting
Research
schema evolution, partition management, transactional updates and historical analysis
of banking and financial datasets.
• Containerized and deployed data-processing services on Azure Kubernetes Service
(AKS), enabling scalable and highly available execution of financial ingestion,
transformation and analytical workloads across environments.
• Provisioned reusable banking data-platform infrastructure using Terraform, supporting
secure storage, compute, networking, access controls and standardized Azure cloud
deployments.
• Developed machine-learning workflows using Azure Machine Learning Pipelines, Azure
Machine Learning Model Catalog and Hugging Face Transformers for financial data
preparation, model training and model initialization and embedding generation.
• Designed secure financial data-lake storage using Azure Data Lake Storage Gen2 (ADLS
Gen2) and orchestrated end-to-end banking workflows through Azure Data Factory
Pipelines, managing dependencies, retries, scheduling and recovery across critical data
pipelines.
• Automated code validation, testing, packaging and deployments using GitHub Actions,
while implementing financial data-quality controls and pipeline observability with Great
Expectations.
• Designed RAG pipelines using LangChain and Azure AI Search, enabling document
ingestion, chunking, metadata enrichment, semantic retrieval and contextual search
across financial documents, policies and operational knowledge repositories.
Senior Data Engineer Sep 2024 - Mar 2025
Best Buy, Richfield, MN,
• Developed scalable retail data pipelines using Python for product, order, inventory,
customer, pricing and transaction datasets, improving reliable data availability across
downstream systems.
• Designed and optimized MySQL data models using indexing, partitioning, query tuning,
stored procedures and feature datasets supporting analytics and Scikit-learn workloads.
• Built enterprise integrations using JDBC, enabling reliable movement of retail data across
databases, distributed processing platforms and downstream business applications.
• Developed REST-based data services using Flask, exposing curated retail datasets,
analytical outputs and lightweight OpenAI API-driven classification and summarization
capabilities.
• Created business-focused reporting solutions in Tableau, combining sales, inventory,
fulfillment and customer KPIs with predictive outputs generated through Scikit-learn.
• Engineered scalable analytical workloads using Databricks SQL Warehouse and Azure
Synapse Analytics, preparing governed datasets for forecasting, segmentation,
recommendations and MLflow-tracked experimentation.
• Implemented asynchronous integrations using RabbitMQ, enabling reliable event
delivery across order, fulfillment, analytics and AI-enabled processing services.
• Built real-time pipelines using Apache Flink and ETL workflows with Talend, processing
clickstream, order, inventory and customer events with Hugging Face
Transformers-based text workflows.
• Processed large-scale historical retail datasets using Apache Hadoop and Apache Hive
for sales trend analysis, customer segmentation, forecasting and machine-learning
feature generation.
• Containerized and deployed distributed data services using Kubernetes, including
scalable Flask APIs and machine-learning inference services supporting Azure-hosted
retail platforms.
• Developed predictive pipelines using Scikit-learn and Azure Machine Learning for
forecasting, segmentation and propensity modeling, with experiment tracking and
lifecycle management through MLflow.
• Built semantic retrieval and knowledge-enrichment workflows using LangChain and
FAISS over product, catalog and customer-support datasets stored in Azure Blob Storage.
• Orchestrated end-to-end retail data and ML workflows using Prefect, automated
deployments through Jenkins and supported analytical workloads in Teradata for
downstream reporting and data products.
Data Engineer June 2022 - July
2024
Cognizant, India
• Developed scalable batch and distributed data-processing solutions using Python,
PySpark and SQL, supporting banking transaction, customer, risk and enterprise
client datasets across consulting engagements.
• Designed and optimized relational data platforms using SQL Server, integrating
applications and enterprise systems through PyODBC and JDBC to support secure
financial data processing and reporting.
• Built lightweight data services and REST APIs using Flask, exposing curated banking
datasets, metadata, pipeline status and analytical outputs to downstream
applications and client platforms.
• Developed interactive financial and operational dashboards using Power BI,
providing insights into transaction trends, risk exposure and business KPIs for
banking and consulting stakeholders.
• Designed and maintained enterprise data warehouse solutions using Azure Synapse
Analytics, implementing dimensional models, bulk-loading strategies, distribution
optimization and high-performance queries for financial reporting and client
analytics.
• Engineered event-driven workflows using Azure Service Bus and Azure Logic Apps,
enabling asynchronous transaction processing, workflow coordination, automated
retries and resilient recovery across enterprise data pipelines.
• Built streaming and ETL frameworks using Azure Event Hubs, Azure Data Factory
and Azure Purview, supporting ingestion, transformation, metadata governance and
discovery of banking and client data assets.
• Processed large-scale financial and enterprise datasets using Apache Spark and
Azure Databricks, optimizing distributed joins, aggregations and reusable
transformation frameworks for analytics and reporting workloads.
• Containerized data-processing applications using Docker and deployed scalable
workloads through Azure Container Instances, ensuring consistent execution
environments and simplified operations across banking and consulting projects.
• Provisioned repeatable cloud data infrastructure using ARM Templates, creating
standardized templates for compute, storage, networking and security resources
across multiple client and financial environments.
• Configured secure cloud environments using Azure Blob Storage, Azure Virtual
Machines, Azure RBAC, Azure Virtual Network, Azure NAT Gateway and Route
Tables, supporting protected financial data storage, controlled access and private
connectivity across enterprise workloads.
• Implemented shared and high-performance storage solutions using Azure Files and
Azure NetApp Files, supporting distributed processing, analytics and data-intensive
applications across banking and client delivery environments.
• Orchestrated enterprise data pipelines using Apache Airflow and automated build,
testing and deployment workflows through Azure Pipelines and Jenkins, ensuring
consistency across data engineering releases.
• Queried large-scale data-lake environments using Azure Synapse Serverless SQL
Pool and Azure Data Lake Storage Gen2, developed Batch Inference Pipelines for
financial and enterprise prediction workloads and monitored platform health and
pipeline performance using Azure Monitor.
Data Engineer July 2021 - May
2022
Bharat Electronics Limited (BEL), India
• Engineered scalable data-processing applications using Python and PyArrow to
handle structured and columnar datasets across banking, defense electronics and
aerospace and government technology workloads.
• Developed and optimized enterprise database solutions with Oracle, enabling
reliable application and data-platform connectivity through JDBC Connector for
financial, mission-critical and operational systems.
• Built secure backend data services and web interfaces using Django, providing
controlled access to financial, defense and government datasets, metadata and
pipeline execution information.
• Created analytical dashboards and operational reports using Amazon QuickSight,
enabling banking, defense and government stakeholders to evaluate trends,
performance metrics, system readiness and operational outcomes.
• Managed enterprise transactional and analytical workloads using Oracle Database
on Amazon EC2, supporting dependable storage, availability and processing for
banking and mission-critical applications.
• Developed event-driven and serverless processing solutions using Amazon Kinesis
and AWS Lambda, enabling real-time ingestion and automated execution for
financial transactions, telemetry and operational data.
• Implemented asynchronous messaging and system-integration workflows with
Amazon SQS and Amazon SNS, supporting queues, publish-subscribe messaging
and resilient communication across distributed banking and defense applications.
• Designed Python-integrated ETL workflows using AWS Glue to extract, transform
and transfer financial, engineering and operational data across cloud and
on-premises environments.
• Executed large-scale distributed processing and analytical workloads using Amazon
EMR, supporting batch transformations across transaction, telemetry, logistics and
mission-support datasets.
• Deployed lightweight containerized Python data services with Amazon ECS and
automated repeatable infrastructure provisioning through AWS CloudFormation
across enterprise environments.
• Established secure cloud environments using Amazon EC2, Amazon S3 and AWS
Site-to-Site VPN, supporting compute, protected data storage and hybrid
connectivity for banking, aerospace and government technology workloads.
• Coordinated end-to-end Python pipeline execution using Apache Airflow while
managing source control and CI/CD activities through Bitbucket, improving
deployment consistency and operational reliability.
• Architected and optimized enterprise analytics and data warehouse solutions using
Amazon Redshift, supporting large-scale financial reporting, defense analytics and
government operational intelligence.