SANJANA MUNAGAPATI
+1-210-***-**** ***********@*****.*** LinkedIn
PROFESSIONAL SUMMARY
• 5+ years of experience as a Data Engineer designing, developing, and optimizing scalable ETL/ELT pipelines and cloud-native data platforms across Financial Services, Healthcare, Ecommerce, and Research Analytics.
• Strong expertise in Python, SQL, PySpark, Apache Spark, Apache Airflow, AWS Glue, Amazon S3, Redshift, Snowflake, Databricks, and Azure Data Factory.
• Experienced in building enterprise-scale batch and near real-time data pipelines for processing structured and semi-structured data.
• Hands-on experience developing Spark-based data transformation pipelines and optimizing distributed data processing workloads.
• Proficient in AWS data engineering services including Glue, S3, Redshift, Lambda, Athena, EMR, IAM, and CloudWatch.
• Strong experience with data modeling, data warehousing, data quality, performance tuning, workflow orchestration, and production support.
• Experienced working in Agile environments with cross-functional teams to deliver scalable, high-quality data solutions.
TECHNICAL SKILLS :
Programming Languages
Python, SQL, PySpark, Shell Scripting
Data Engineering & Big Data
Apache Spark, PySpark, Apache Airflow, ETL/ELT, Data Modeling, Data Warehousing, Batch Processing, Data Validation, Data Quality, Workflow Orchestration
Cloud Platforms
AWS: AWS Glue, Amazon S3, Amazon Redshift, AWS Lambda, Amazon Athena, Amazon EMR, IAM, CloudWatch
Azure: Azure Data Factory, Azure Databricks, Azure SQL Database, Azure Blob Storage, Azure DevOps
Data Warehousing & Databases
Snowflake, Amazon Redshift, PostgreSQL, SQL Server, MySQL, Oracle
Data Integration & APIs
REST APIs, JSON, XML, CSV, API Integration, File-Based Data Processing
Business Intelligence & Analytics
Power BI, Tableau, Microsoft Excel, KPI Reporting, Dashboard Development
DevOps & Version Control
Git, GitHub, Jenkins, CI/CD Pipelines
Monitoring & Operations
Pipeline Monitoring, Logging & Alerting, Incident Management, Production Support, Job Scheduling, Root Cause Analysis (RCA)
Software Development & Methodologies
Agile (Scrum), SDLC, Code Reviews, Unit Testing, Documentation, Cross-functional Collaboration
EXPERIENCE
Central Bank Remote, USA Nov 2024 - Present
Data Engineer
Domain: Financial Services & Compliance Platform
• Designed and developed scalable cloud-native ETL/ELT pipelines using Python, SQL, PySpark, Apache Spark, and AWS services for high-volume financial transaction and compliance datasets.
• Developed AWS Glue ETL pipelines to ingest, transform, and load enterprise financial data into Amazon S3 and Amazon Redshift.
• Built Apache Airflow DAGs to orchestrate batch data pipelines, improving workflow automation, scheduling reliability, and operational efficiency.
• Designed Spark-based transformation pipelines for cleansing, enrichment, aggregation, and validation of large-scale financial datasets.
• Integrated data from APIs, relational databases, and external regulatory systems into enterprise data lakes and analytical platforms.
• Optimized Spark jobs, SQL queries, and Redshift workloads, significantly improving ETL performance and reporting efficiency.
• Implemented automated data quality validation, reconciliation, exception handling, monitoring, and alerting solutions.
• Worked extensively with AWS services including S3, Glue, Lambda, CloudWatch, IAM, and Redshift.
• Performed production support, root cause analysis, incident resolution, and performance tuning for mission-critical pipelines.
• Collaborated with business stakeholders, compliance teams, and data analysts using Agile methodologies.
Environment: Python, SQL, PySpark, Apache Spark, AWS Glue, Amazon S3, Redshift, Lambda, Airflow, CloudWatch, Tableau, Git, Jira, CI/CD
Elevance Health Aug 2023-oct 2024
Data Engineer
Domain: Healthcare Analytics & Data Platform
• Developed enterprise healthcare ETL pipelines using Python, SQL, Azure Data Factory, and PySpark.
• Built Spark-based transformation pipelines in Azure Databricks for large-scale healthcare analytics.
• Designed scalable ingestion pipelines for APIs, flat files, relational databases, and cloud storage.
• Implemented dimensional data models and optimized SQL workloads supporting enterprise reporting.
• Built automated data validation, reconciliation, and monitoring frameworks improving data reliability.
• Optimized Spark processing and ETL execution to improve SLA performance.
• Supported production deployments, troubleshooting, CI/CD automation, and release management.
• Collaborated with business users and reporting teams to deliver analytics solutions..
Environment: Python, SQL, PySpark, Apache Spark, Azure Data Factory, Azure Databricks, Snowflake, Power BI, Git, Azure DevOps
Peace Economy Project Remote, USA Jan 2023 - Jul 2023
Data Engineer
Domain: Research & Analytics Platform
• Developed scalable ETL/ELT pipelines using Python, SQL, Azure Data Factory, and PySpark.
• Built Spark transformation pipelines for research and operational datasets.
• Designed automated ingestion pipelines integrating APIs, relational databases, cloud storage, and flat files.
• Implemented data quality validation, reconciliation, and monitoring frameworks.
• Optimized SQL queries and Spark jobs to improve ETL processing performance.
• Developed reporting datasets and Power BI dashboards supporting analytics initiatives.
• Collaborated with cross-functional engineering and analytics teams using Agile methodologies.
Environment: Python, SQL, PySpark, Apache Spark, Azure Data Factory, Azure Databricks, Power BI, Git, REST APIs
Accenture Solutions Private Limited Jan 2021 - Jul 2022
Data Engineer
Domian: Henkel - Global Ecommerce Retail Platform
• Designed and developed enterprise ETL/ELT pipelines using Python, SQL, Azure Data Factory, and PySpark for ecommerce sales, inventory, product, and customer datasets.
• Built Spark-based distributed data transformation pipelines using Azure Databricks.
• Developed scalable data ingestion frameworks integrating REST APIs, ERP systems, transactional databases, and cloud storage.
• Integrated enterprise datasets into Snowflake and Amazon Redshift data warehouses.
• Optimized Spark transformations, SQL queries, and ETL workflows, improving reporting performance by more than 35%.
• Implemented data quality validation, monitoring, reconciliation, and production support processes.
• Participated in CI/CD deployment, code reviews, release management, and Agile ceremonies.
• Collaborated with global stakeholders to deliver enterprise analytics solutions.
Environment: Python, SQL, PySpark, Apache Spark, Azure Data Factory, Azure Databricks, AWS S3, Snowflake, Amazon Redshift, Git, Jenkins, REST APIs
CERTIFICATIONS
• AI Certification
• Coursera: Python, Data Visualization, SQL
EDUCATION
George Mason University Fairfax, Virginia Aug 2022 - May 2024
Master of Science in Computer Science
Jawaharlal Nehru Technological University Hyderabad, India Aug 2017 - Jun 2021
Bachelor of Technology in Computer Science