ABOUT ME
Performance-focused Data Engineer with *+ years of experience in designing, building, and optimizing scalable data pipelines and architectures. Proven expertise in big data technologies, ETL processes, cloud platforms (AWS, Azure, GCP), and data modelling. Seeking a challenging role where I can leverage my technical skills and leadership experience to drive data-driven solutions, improve data infrastructure, and support strategic business decisions.
WORK HISTORY
Azure Data Engineer
Starr Insurance, Atlanta, Georgia, USA May 2024 - Current
Starr Insurance is a global insurance and investment organization. I design and implement data pipelines using Azure Data Factory (ADF), Azure Synapse Analytics, and Databricks. Develop data models optimized for performance and scalability. Implement ETL/ELT solutions for structured and unstructured data sources.
Key Responsibilities:
Created Data stage and ETL jobs for populating the data into Data Warehouse constantly from different source systems like ODS, flat files, Parquet. Experience in managing and reviewing Hadoop log files.
Provided high availability for IaaS VMs and PaaS role instances for access from other services in the VNet with Azure Internal Load Balancer. Participated in development/implementing of Cloudera Hadoop development.
Worked on Kafka streaming on subscriber side, processing the messages and inserting them into the db and Apache Spark for real-time data processing. Architected and implemented big data solutions using Hadoop, Hive, and Cosmos DB, integrated with Azure platforms to manage and process large volumes of financial data across Starr Insurance’s infrastructure.
Worked closely with data science teams to implement and operationalize machine learning models using Azure Machine Learning, Databricks MLflow, and Azure Synapse Analytics, embedding predictive analytics into data pipelines to enhance financial decision-making and implemented and optimized NoSQL databases such as MongoDB and Cassandra to store and process unstructured financial data, enabling faster retrieval and flexible storage of large-scale transaction data for analysis and Leveraged Apache NiFi for data flow management and automation, enabling the efficient movement and transformation of financial data between multiple systems, enhancing the integration of diverse data sources
Leveraged Python and libraries like Pandas and NumPy for data cleaning, transformation, and automation tasks, processing large datasets for analysis and machine learning in financial data pipelines.
Performed data processing in Azure Databricks after data ingestion into Azure services such as Azure Data Lake, Azure Storage, Azure SQL DB, and Azure SQL DW.
Built Spark programs using Scala and APIs for performing transformations and operations on RDDs.
Created Data tables utilizing PyQt to display customer and policy information and add, delete, update customer records.
Developed Monitoring and notification tools using Python. Involved in the entire lifecycle of the projects including Design, Development, and Deployment, Testing and Implementation, and support.
Developing Spark scripts, UDFS using both Spark DSL and Spark SQL query for data aggregation, querying, and writing data back into RDBMS through Sqoop.
Leveraged Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics for data ingestion into various Azure Services, including Azure Data Lake, Azure Storage, Azure SQL and Azure DW.
Performed continuous Integration/Continuous delivery (CI/CD) on Jenkins build pipeline and fixed failure issues.
Worked on migration of data from On - prem SQL Server to Cloud databases (Azure Synapse Analytics (DW) & Azure SOL DB). Responsible for implementing monitoring solutions in Ansible, Terraform, Docker, and Jenkins.
Built robust data ingestion pipelines using Logstash, Filebeat, and Kafka Connect to stream real-time logs and events into Elasticsearch clusters. Designed and implemented Infrastructure as code using Terraform, enabling automated provisioning and scaling of cloud resources on Azure.
Extensively used Azure Athena to ingest structured data from Azure Blob Storage into various systems such as Azure Synapse Analytics or to generate reports. Transferred data from Azure Blob storage to Snowflake database.
Executed tasks for upgrading cluster on the staging platform before doing it on production cluster.
Environment: Analytics, Apache, API, APIs, Athena, Azure, Azure Data Lake, Azure Synapse Analytics, Blob, CI/CD, Cloudera, Data Factory, Docker, Elasticsearch, ETL, Factory, IaaS, Jenkins, Kafka, lake, Lake, PaaS, Python, RDBMS, Scala, Services, Snowflake, Spark, Spark SQL, SQL, Sqoop
AWS Data Engineer
BDO USA, Atlanta, Georgia, USA Nov 2023 - Apr 2024
BDO USA is a leading professional services firm providing assurance, tax, and advisory services to a wide range of clients. Performed the data cleansing, aggregation, and transformation using tools like AWS Glue, Apache Spark, or Amazon Redshift. Managed data storage solutions such as Amazon S3 or Amazon Redshift.
Key Responsibilities:
Designed and developed ETL pipelines for real-time data integration and transformation using Kubernetes and Docker.
The AWS Lambda functions were written in Spark with cross - functional dependencies that generated custom libraries for delivering the Lambda function in the cloud. Extracted the data from Teradata into HDFS using Sqoop.
Actively involved in designing and developing data ingestion, aggregation, and integration in the Hadoop environment.
Wrote Spark-Streaming applications to consume the data from Kafka topics and write the processed streams to HBase.
Spearheaded HBase setup and utilized Spark and Spark SQL to develop faster data pipelines, resulting in a 60% reduction in processing time and improved data accuracy.
lnvolved in various phases of Software Development Lifecycle (SDLC) of the application, like gathering requirements, design, development, deployment, and analysis of the application.
Worked on SQL and PL/SQL for backend data transactions and validations
Involved in the automation process through Jenkins for CI/CD pipelines. Developed Sqoop scripts to import export data from relational sources and handled incremental loading on the customer, transaction data by date.
Led end-to-end data migration from on-premise Hadoop clusters to Google Cloud Platform, using Dataproc, BigQuery, and Cloud Storage, successfully transferring over 20TB of data with zero data loss.
Deployed models as Python package, as API for backend integration and as services in a Microservices architecture with a Kubernetes orchestration layer for the Docker containers. Optimized Elasticsearch cluster performance through shard tuning, heap memory management, refresh interval adjustments, and query profiling.
Provisioned high availability of AWS EC2 instances, migrated legacy systems to AWS, and developed Terraform plugins, modules, and templates for automating AWS infrastructure. Developed ETL pipelines between data warehouses using a combination of Python and Snowflake, SnowSQL, writing SQL queries against Snowflake.
Developed methods for Create, Read, Update and Delete (CRUD) in Active Record.
A security framework was created to provide for fine-grained access control for items in AWS S3.
Environment: API, AWS, BigQuery, CI/CD, Docker, EC2, Elasticsearch, ETL, HBase, HDFS, Jenkins, Kafka, Kubernetes, lake, Lambda, PL/SQL, Python, S3, Snowflake, Spark, Spark SQL, SQL, Sqoop, Teradata
GCP Data Engineer
Piramal Pharma Limited, Bangalore, India Nov 2021 - Jul 2023
Piramal Pharma Limited (PPL) is a global pharmaceutical company offering a diverse portfolio of products and services. Developed and optimized the Extract, Transform, Load (ETL) processes to ingest, process, and store data from various sources using GCP services such as Dataflow, Dataproc, and Pub/Sub.
Key Responsibilities:
Involved in requirement gathering, business analysis, and technical design for Hadoop and Big Data projects.
Assess the infrastructure needs for each application and deploy it on Azure platform.
Created reusable views and data marts in BigQuery to power Data Studio reports with consistent metrics and definitions.
Implemented dynamic cluster provisioning in Dataproc with autoscaling policies, reducing idle cluster time and optimizing GCP costs. Led requirement gathering, business analysis, and technical design for Hadoop and Big Data projects.
Created Amazon VPC to create public-facing subnet for web servers with internet access, and backend databases & application servers in a private-facing subnet with no Internet access.
Experience in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns. Used Cloud Shell for troubleshooting production data pipelines in GCP.
Developed Spark applications for the entire batch processing by using PySpark.
Architected Python scripts for automated data extraction and loading from web server output files, reducing manual data entry, and processing time by 75%. Monitored and troubleshooted GCP services using Google Cloud SDK, enabling faster incident resolution and improved system reliability.
Developed database triggers and stored procedures using T-SQL cursors and tables.
Designing and implementing data integration solutions using Azure Data Factory to move data between various data sources, including on-premises and cloud-based systems.
Enabled real-time dashboards by streaming data into BigQuery via Dataflow and Pub/Sub.
Skilled in query optimization and performance tuning, utilizing MongoDB's indexing and aggregation framework to improve query execution times and overall database efficiency.
Working with GCP cloud using in GCP Cloud storage, Data-proc, Data Flow, Big-Query, EMR, S3, Glacier and EC2 with EMR Cluster. Knowledge on Google Cloud Dataflow and Apache Beam.
Migrated previously written cron jobs to Airflow/Cloud Composer in GCP.
Environment: Airflow, Apache, Apache Beam, Azure, BigQuery, Cluster, Data Factory, EC2, EMR, Factory, GCP, PySpark, Python, S3, SDK, Spark, SQL, VPC
Data Engineer
Equifax, Bangalore, India Mar 2020 - Oct 2021
Equifax Inc. is an American multinational consumer credit reporting agency. Implemented the data validation and transformation processes to maintain data accuracy. Applied the security measures like encryption and access controls to protect sensitive information.
Key Responsibilities:
Designed SSIS control flow tasks for orchestrating the sequence and logic of ETL processes, such as conditional branching, looping, and error handling. lnvolved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analyzed them by running Hive queries.
Used AWS Lambda to perform data validation, filtering, sorting, or other transformations for every data change in a Database table and load the transformed data to another data store AWS S3 for raw file storage.
Monitored Spark cluster using Log Analytics and Ambari Web Ul. Transitioned log storage from Cassandra to Azure SQL Datawarehouse and improved the query performance. lmplemented Apache Sqoop for efficiently transferring bulk data between Apache Hadoop and relational databases (Oracle) for product level forecast.
Used Kafka functionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds. Involved in loading data into Cassandra NoSQL Database.
Good Knowledge on architecture and components of Spark, and efficient in working with Spark Core, Spark SQL, Spark streaming and expertise in building PySpark and Spark-Scala applications for interactive analysis, batch processing and stream processing. Creating job flow using Apache Airflow in python and automating the jobs.
Involved in various phases of Software Development Lifecycle (SDLC) of the application, like gathering requirements, design, development, deployment, and analysis of the application.
Set up the CI/CD pipelines using Jenkins, Maven, GitHub, Chef, Terraform, and AWS.
Migrated Oracle E-Business Suite to Google Cloud SQL with 2hr downtime
Experience in creating Kubernetes replication controllers, Clusters and label services to deployed Microservices in Docker.
Designed and implemented Elasticsearch index schemas to support scalable, high-performance search and analytics over structured and unstructured data. Worked with AWS Terraform templates in maintaining the infrastructure as code.
Conducted Performance tuning and optimization of Snowflake data warehouse, resulting in improved query execution times and reduced operational costs. Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR
Environment: Airflow, Analytics, Apache, AWS, CI/CD, Cluster, Docker, EC2, Elasticsearch, EMR, ETL, Factory, Git, Hive, Jenkins, Kafka, Kubernetes, Lake, Lambda, Oracle, PySpark, S3, Scala, Snowflake, Spark, Spark Core, Spark SQL, SQL, Sqoop, SSIS, Tableau
EDUCATION
Masters in Information Systems from Trine university, Michigan, USA (Aug 2023 - May 2025)
Bachelors in Computer science from SVS Group of institutions, Warangal, India (Aug 2018 – May 2022)
Saiteja Atika
Data Engineer
************@*****.***
PROFESSIONAL SUMMARY
RELEVANT SKILLS
5+ years of expertise designing, developing, and executing data pipelines and data lake requirements in numerous companies using the Big Data Technology stack, Python, PL/SQL, SQL, REST APIs, and the Azure cloud platform.
Hands-on experience with Amazon EC2, Amazon S3, Amazon RDS, VPC, IAM, Amazon Elastic Load Balancing, Auto Scaling, CloudWatch, SNS, SES, SQS, Lambda, EMR and other services of the AWS family. Designed, build and managed ELT data pipelines leveraging Airflow, Python, and GCP solutions.
Experienced in configuring and administering the Hadoop Cluster using major Hadoop Distributions like Apache Hadoop and Cloudera.
Working knowledge of HDFS, Kafka, MapReduce, Spark, PIG, Hive, Sqoop, HBase, Flume, and Apache ZooKeeper as tools for designing and deploying end-to-end big data ecosystems. Experience on Azure data factory with different triggers.
Develop batch processing solutions by using Data Factory and Azure Data bricks.
Created Python code to collect data from HBase and develop the PySpark implementation of the solution. Proficient in converting Hive/SQL queries into Spark transformations using Data frames and Data sets.
Experience in Windows Azure Services like PaaS, IaaS and worked on storages like Blob (Page and Block), SQL Azure. Well experienced in deployment & configuration management and Virtualization. Experience working with various SDLC methodologies like Agile Scrum, RUP and Waterfall model.
Experience in implementing Azure data solutions, provisioning storage account, Azure Data Factory, SQL Server, SQL Databases, SQL Data warehouse, Azure Data Bricks and Azure Cosmos DB. Hands-on experience in implementing, Building, and Deployment of CI/CD pipelines, managing projects often including tracking multiple deployments across multiple pipeline stages (Dev, Test/QA staging, and production).
Good experience in automating the scheduled jobs in Control-M and Airflow.
Experience in migrating on premise to Windows Azure using Azure Site Recovery and Azure backups. Used microservices and containerization technologies such as Docker and Kubernetes to build scalable and resilient SaaS applications.
Experience working with Front end technologies like Html, CSS, JS, ReactJS.
In-depth knowledge of Data Sharing in Snowflake and experienced in Snowflake Database, Schema and Table structures.
Experience in Relational Databases like Oracle, MS SQL Server, MS Access, SQL, PL/SQL and Operating Systems like Windows, Linux, and UNIX.
Ability to work effectively and efficiently as a team member as well as individually with a desire to learn new skills and technology.
Worked on production support looking into logs, hot fixes and used Splunk for log monitoring along with AWS CloudWatch.
Cloud Technologies: AWS, GCP,
AZURE.
AWS Ecosystem: S3Bucket, Athena, Glue, EMR, Redshift, Data Lake, AWS Lambda, Kinesis.
Azure Ecosystem: Azure Data Lake, ADF, Databricks, Azure SQL
Databases: Oracle, MySQL, SQL
Server, PostgreSQL, HBase,
Snowflake, Cassandra, MongoDB.
Programming Languages: Java,
Python, Hibernate, JDBC, JSON, HTML, CSS, C# /.NET.
Script Languages: Python, Shell
Script (bash, shell).
Version controls and Tools: GIT, Maven, SBT, CBT
Hadoop Components / Big Data: HDFS, Hue, MapReduce, PIG, Hive, HCatalog, HBase, Sqoop, Impala, Zookeeper, Flume, Kafka, Yarn, Cloudera Manager, Kerberos, Pyspark Airflow, Kafka, Snowflake Spark Components
Visualization& ETL tools: Tableau, PowerBI, Informatica, Talend
Operating Systems: Windows, Unix, Linux
Methodologies: Agile (Scrum),
Waterfall, UML, Design Patterns, SDLC.
Webservers: Apache Tomcat,
WebLogic.
Collaboration Tools: JIRA, Confluence, Slack, Microsoft Teams
Healthcare Data (PII/HIPAA)