Swathi Gajawada
Data Engineer
Email: *******.*****@*****.***
Ph #: +1-945-***-****
PROFESSIONAL SUMMARY:
Senior Data Engineer with 8+ years of experience designing, developing, and optimizing cloud-native data platforms across Banking, Healthcare, Insurance, and Enterprise Analytics domains.
Expertise in designing and implementing cloud-native data platforms on Google Cloud Platform (GCP) using BigQuery, Dataflow, Dataproc, Pub/Sub, Cloud Storage, and Cloud Composer to support enterprise-scale analytics.
Hands-on experience developing AI-powered data engineering solutions using Vertex AI, LangChain, LangGraph, OpenAI APIs, and Agentic AI frameworks for enterprise data processing and intelligent automation.
Strong experience with SQL Server, including database design, data transformation, query optimization, and performance tuning for large-volume datasets.
Hands-on experience with SSIS and Informatica for developing and maintaining enterprise ETL/ELT pipelines and data integration workflows.
Experienced in developing complex SQL queries, stored procedures, joins, CTEs, and transformation logic to support data integration, reporting, and analytics.
Experience building Retrieval-Augmented Generation (RAG) and GraphRAG solutions using Vector Databases to enable semantic search and enterprise knowledge discovery.
Developed Python-based AI pipelines integrating Large Language Models (LLMs) with cloud-native data platforms to automate document processing, metadata generation, and contextual information retrieval. Strong hands-on experience developing Apache Spark (Scala/PySpark) and Python applications for building high-performance batch and real-time ETL/ELT data pipelines processing multi-terabyte datasets.
Experienced in integrating enterprise applications through REST APIs, JSON, XML, and microservices architecture, delivering scalable, secure, and reusable cloud-based data solutions.
Proficient in SQL, data modeling, Git, CI/CD, Docker, Kubernetes, and Agile methodologies, with a strong focus on performance optimization, automation, and production support.
Experience working with cross-functional teams to design data solutions and support end-to-end SDLC in Agile and Waterfall environments
Hands-on experience with GCP ecosystem including BigQuery, Dataflow, and Pub/Sub for building scalable data solutions
Strong experience owning end-to-end data pipelines with a focus on production stability, monitoring, and issue resolution
Proven ability to work independently, handle ambiguous requirements, and deliver high-quality, production-grade data solutions
Expertise in data validation, debugging, and performance optimization across large-scale datasets
Strong expertise in cloud-based data engineering using AWS services including S3, Glue, EMR, Lambda, Redshift, Athena, and Kinesis to build scalable and cost-efficient data platforms
Proficient in Python, SQL, and PySpark with hands-on experience processing large volumes of structured and semi-structured data using distributed computing frameworks like Apache Spark
Demonstrated experience in designing and implementing data lake and data warehouse architectures to support analytics, reporting, and business intelligence use cases
Implemented end-to-end data quality frameworks, improving data reliability and consistency across systems
Hands-on experience building batch and near real-time data pipelines using technologies such as Apache Kafka and AWS Kinesis for streaming data integration
Experience working with domain-specific datasets including financial transactions, insurance claims and policy data, and healthcare patient/provider data with compliance to regulatory and data privacy standards (including HIPAA)
Hands-on experience with workflow orchestration tools such as Apache Airflow for scheduling, monitoring, and managing complex data pipelines
Strong experience in performance tuning, query optimization, and pipeline efficiency improvements, achieving up to 30–40% reduction in processing time
Proven ability to collaborate with cross-functional teams including business stakeholders, analysts, and data scientists to define KPIs and deliver scalable data solutions aligned with business goals
Experience supporting BI and analytics platforms by delivering high-quality, analytics-ready datasets enabling data-driven decision-making
Strong problem-solving skills with the ability to troubleshoot data pipeline failures, identify data issues, and implement effective solutions in fast-paced environments
TECHNICAL SKILLS:
Programming Languages
Python, SQL, Scala, Java
Cloud Platforms
GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Google Cloud Storage (GCS), Cloud Composer, Cloud Functions, Cloud Run, Vertex AI, IAM)
AWS (S3, Glue, EMR, Lambda, Redshift, Athena, Kinesis)
Azure Data Factory, Azure Synapse Analytics
Big Data Technologies
Apache Spark, PySpark, Spark SQL, Spark Structured Streaming, Apache Kafka, Kafka Streams, Apache Flink, Hadoop, Delta Lake, Apache Parquet
Data Lakehouse & Data Warehousing
Data Lakehouse Architecture, Medallion Architecture (Bronze/Silver/Gold), BigQuery, Snowflake, Redshift, Azure Synapse, Teradata, Oracle, SQL Server, PostgreSQL, MySQL, MongoDB
ETL / ELT & Orchestration
Apache Airflow, Cloud Composer, dbt, AWS Glue, Dataflow, Informatica, SSIS, Databricks Delta Live Tables
Streaming & Real-Time Processing
Apache Kafka, Kafka Streams, Spark Structured Streaming, Apache Flink, Pub/Sub, AWS Kinesis, Event Hub
AI / Agentic AI
Vertex AI, LangChain, LangGraph, RAG, GraphRAG, Vector Databases, OpenAI APIs, Hugging Face, Bedrock, Generative AI, Agentic AI Frameworks
DevOps & Infrastructure
Terraform, Jenkins, GitHub Actions, Docker, Kubernetes, CI/CD
BI & Analytics
Looker, Power BI, Tableau, QlikView
Data Governance
Data Quality Frameworks, Data Validation, Data Catalog, Data Lineage, Data Masking, HIPAA Compliance, Collibra, Alation
PROFESSIONAL EXPERIENCE:
Fifth Third Bank, Cincinnati, OH.
Feb 2025 – Present
Sr. Data Engineer
Responsibilities:
•Designed and implemented cloud-native data platforms on GCP using BigQuery, Dataflow, Pub/Sub, Dataproc, and Google Cloud Storage (GCS), processing multi-terabyte datasets daily.
•Developed AI-powered data engineering solutions using Vertex AI, LangChain, and OpenAI APIs to automate enterprise data processing and intelligent document retrieval.
•Built Retrieval-Augmented Generation (RAG) pipelines leveraging Vector Databases and embedding models to enable semantic search across enterprise knowledge repositories.
•Implemented GraphRAG-based knowledge retrieval solutions to improve contextual search accuracy and enhance enterprise analytics capabilities.
•Developed and optimized SQL Server-based ETL processes for data extraction, transformation, validation, and loading across enterprise systems.
•Developed SSIS ETL workflows with data transformation, validation, error handling, and reconciliation to support reliable data integration.
•Performed complex SQL query development and performance tuning, optimizing data transformations and improving ETL processing efficiency.
•Developed Python-based AI ingestion frameworks integrating Large Language Models (LLMs) with Snowflake and cloud data platforms for intelligent data enrichment and metadata generation.
•Integrated AI services with GCP components including BigQuery, Cloud Storage, Dataproc, and Vertex AI to support scalable AI-driven analytics solutions. strategies and retrieval optimization techniques to improve response quality and contextual relevance in AI-powered applications.
•Designed and developed scalable Python and PySpark data pipelines on Databricks and Apache Spark to process high-volume structured and semi-structured data, improving data processing performance and reliability.
•Built cloud-native ETL/ELT solutions using AWS S3, Lambda, EMR Serverless, and Step Functions, automating data ingestion, orchestration, and transformation for enterprise analytics.
•Implemented CI/CD and GitOps practices using Git, Jenkins, and GitHub Actions, automating code integration, testing, deployment, and release management to improve delivery efficiency.
•Optimized Spark jobs, SQL queries, and AWS infrastructure through performance tuning, partitioning, and resource optimization, reducing processing time and cloud costs while ensuring high availability.
•Built vector-based semantic search solutions leveraging GraphRAG and embedding models to improve enterprise knowledge discovery.
•Built robust data ingestion frameworks to integrate data from multiple banking systems, including transactional systems, APIs, and external data providers, enabling seamless consolidation of enterprise data into AWS S3 and Redshift
•Supported enterprise data migration initiatives involving Salesforce CRM datasets, including extraction, transformation, validation, and integration workflows
•Worked with Salesforce-related data models and object structures to support migration, reconciliation, and downstream reporting requirements
•Developed PySpark and SQL-based data migration pipelines to process and transform Salesforce and enterprise datasets across cloud platforms
•Performed data cleansing, deduplication, mapping, and validation activities to ensure high-quality migration of Salesforce-related data
•Collaborated with integration architects and business teams to address Salesforce migration challenges including schema mapping, data consistency, and validation rules
•Built automated migration and reconciliation workflows using AWS Spark, EMR, and S3 to support scalable Salesforce data migration processes
•Conducted advanced SQL and PL/SQL analysis for Salesforce-integrated datasets to identify data quality issues and improve migration reliability
•Supported migration planning activities by defining data cleanup strategies, transformation logic, and optimized migration paths for CRM datasets
•Worked with AWS services including Athena, EMR, Spark, S3, IAM Roles, and policies to build secure and scalable migration frameworks for enterprise applications including Salesforce integrations
•Assisted in integrating Salesforce datasets with enterprise analytics and reporting platforms to support business intelligence and operational reporting
•efficient scheduling, monitoring, and error handling of complex data pipelines
•Implemented data pipelines to process and store large-scale datasets in Parquet format on AWS S3, improving query performance and storage efficiency
•Optimized data processing using partitioned Parquet datasets, reducing query execution time and cost
•Developed PySpark jobs to convert raw data (CSV/JSON) into Parquet format for efficient downstream analytics
•Worked with real-time and near real-time data streaming technologies such as Kafka and AWS Kinesis (exposure) to support time-sensitive data processing requirements
•Performed advanced SQL query optimization and performance tuning, reducing execution time and improving overall pipeline efficiency
•Implemented data quality checks, logging, and monitoring mechanisms, ensuring high data reliability and early detection of anomalies or failures
•Troubleshot and resolved data pipeline failures, integration issues, and performance bottlenecks, ensuring system stability and minimal downtime
•Contributed to improving data architecture and pipeline design, enhancing scalability, maintainability, and performance of the data platform
Ensured data integrity, consistency, and governance across multiple systems and platforms, adhering to enterprise data standards
Elevance Health, Plano, TX
Oct 2023 – Feb 2025
Sr. Data Engineer
Responsibilities:
•Built healthcare data pipelines on GCP utilizing BigQuery, Dataflow, Pub/Sub, and Cloud Storage to support large-scale analytics workloads.
•Developed Cloud Composer workflows for orchestration of hundreds of daily ETL processes across clinical and claims datasets.
•Implemented Data Lakehouse architecture supporting structured and semi-structured healthcare data with Delta Lake.
•Developed streaming ingestion solutions using Kafka, Spark Structured Streaming, and Pub/Sub for near real-time healthcare analytics.
•Designed and implemented robust, scalable ETL pipelines using AWS Glue and PySpark to process large volumes of healthcare claims, patient, and provider data, ensuring efficient ingestion, transformation, and loading into cloud-based data platforms
•Developed comprehensive data ingestion frameworks to integrate data from multiple sources, including EHR systems, APIs, flat files, and third-party healthcare systems, enabling seamless data consolidation into AWS S3 and Redshift
•Built and optimized data transformation workflows to cleanse, standardize, and enrich healthcare datasets, ensuring consistency and usability for downstream analytics and reporting
•Implemented advanced data quality validation and reconciliation mechanisms to ensure accuracy, completeness, and compliance with healthcare regulatory standards
•Leveraged Apache Airflow for orchestration of complex workflows, enabling automated scheduling, monitoring, and error handling for data pipelines
•Utilized Parquet-based storage for healthcare datasets to improve performance of large-scale data processing and reporting
•Integrated near real-time data ingestion using AWS Kinesis to support time-sensitive reporting and analytics requirements
•Developed and maintained data pipelines using PostgreSQL, ensuring efficient data transformation, validation, and loading for healthcare datasets.
•Performed performance tuning and query optimization in PostgreSQL, improving reporting speed and handling high-volume data workloads.
•Performed query optimization and performance tuning in Redshift, significantly improving data retrieval times for reporting and analytics use cases
•Collaborated closely with business stakeholders, analysts, and data scientists to define data requirements, KPIs, and reporting metrics, ensuring alignment with business objectives
•Ensured HIPAA-compliant data handling and security practices, maintaining strict adherence to data privacy and governance standards
•Monitored pipeline performance, identified bottlenecks, and proactively resolved data latency and processing issues
Supported ad-hoc data requests and reporting needs, providing timely and accurate data insights to business users
Maintained high standards of data integrity, consistency, and reliability across all systems and pipelines
PNB MetLife India Insurance Co. Ltd, Hyderabad, India
Aug 2021 – Feb 2023
Data Engineer
Responsibilities:
Designed and developed scalable ETL/ELT pipelines using AWS Glue, PySpark, and SQL to process high-volume insurance claims, policy, and customer data, ensuring efficient ingestion, transformation, and loading into cloud data platforms
Built robust data ingestion frameworks to integrate data from multiple sources including policy administration systems, claims platforms, APIs, and third-party data providers, consolidating data into AWS S3 and Amazon Redshift
Developed and optimized data transformation workflows to cleanse, standardize, and enrich insurance datasets (claims, underwriting, policy lifecycle), enabling high-quality data for downstream analytics and reporting
Designed and maintained data marts and dimensional models (star/snowflake schema) to support reporting on claims processing, customer behavior, policy performance, and risk analytics
Leveraged Apache Airflow for workflow orchestration, enabling automated scheduling, monitoring, and failure handling of complex data pipelines
Integrated near real-time data ingestion using AWS Kinesis, supporting time-sensitive use cases such as claims status tracking and fraud detection analytics
Performed advanced SQL query optimization and performance tuning in Amazon Redshift, significantly improving query performance and reducing data processing latency
Collaborated with business stakeholders, actuaries, and analytics teams to gather requirements, define KPIs, and deliver scalable data solutions aligned with business objectives
Monitored and troubleshot data pipeline failures, data inconsistencies, and performance bottlenecks, ensuring high availability and reliability of data systems
Participated in Agile development processes, including sprint planning, backlog grooming, and regular stakeholder communication to ensure timely delivery of data solutions
ADP, Hyderabad, India
June 2018 – Aug 2021
Data Engineer
Responsibilities:
Supported data extraction, transformation, and loading (ETL) processes using Python and SQL across development and testing environments.
Assisted in developing and maintaining reporting datasets for business and operational reporting needs.
Performed data validation and accuracy checks to ensure consistency and reliability of loaded data.
Wrote and executed SQL queries to extract, reconcile, and analyze data from relational databases.
Assisted senior engineers in requirements gathering and analysis for data-related enhancements.
Created and maintained technical documentation, including data mappings and process flows.
Supported testing activities, including unit testing and data verification for ETL processes.
Participated in issue resolution and defect tracking, helping troubleshoot data discrepancies and job failures.
Developed and optimized SQL queries for data extraction, transformation, and reporting, supporting business intelligence and operational reporting requirements
Gained hands-on experience with data warehousing concepts, ETL processes, and data modeling techniques, building a strong foundation in data engineering practices
Contributed to maintaining data integrity, consistency, and accuracy across systems, supporting reliable business reporting
Education Details:
Bachelor of Computers at Osmania University (2017) – India.