Post Job Free
Sign in

AWS Data Engineer for GenAI & MLOps

Location:
Denton, TX, 76201
Posted:
August 31, 2026

Contact this candidate

Resume:

SOUMYA NANDITHA CHADALAVADA

Data Engineer AWS GenAI, Agentic AI & MLOps Cloud Infrastructure

+1-940-***-**** ************@*****.*** linkedin.com/in/nanditha-c-24496b33b/ Denton, TX

PROFESSIONAL SUMMARY

Data Engineer with approximately 5 years of experience building GCP and AWS data pipelines and AI-driven workflows at Amazon Web Services and Accenture. Comfortable working across AWS (Bedrock, MWAA, Glue, EMR, CDK), Spark, Kafka, Snowflake, and Python, with recent focus on LLMs, RAG, BigQuery ML, and agent-based systems. Have owned pipelines end to end, from dbt-based ingestion and transformation through monitoring, deployments, encryption, and cost optimization. Recent projects include automating month-end financial workflows, building a Slack-based AI assistant, and setting up Airflow and Hadoop infrastructure using Infrastructure as Code for production pipelines

PROFESSIONAL EXPERIENCE

Amazon Web Services

Data Engineer Nov 2024 – Present

Agentic AI & RAG Systems

•Wrote internal MCP servers and tooling on AWS to connect agents with our existing services, which cut down on glue code and made integrations more reliable.

•Built a multi-agent AI workflow for Month End Closure (MEC) on AWS Bedrock that automates most of the steps the team used to do by hand, taking what was around a 20-hour monthly process down to roughly an hour.

•Built a RAG pipeline on AWS Bedrock Knowledge Base with OpenSearch for vector search, used by analysts to answer process questions and speed up anomaly investigations.

•Built a Slack bot that lets the team trigger and monitor backend ETL workflows in plain English, using Lambda, API Gateway, and SQS with WebSocket streaming for live responses.

•On-Call Companion agent built on AWS to help engineers triage incidents and suggest infrastructure fixes, with a human-in-the-loop approval step before any change is applied.

•Coordinated analytics processes and reporting queue management, aligning with data governance standards for Cogito provisioning.

•Applied data security and encryption practices modeled on Epic Cogito system-wide settings.

•Architected a self-serve data platform on GCP using BigQuery, Dataplex, and BigLake to enforce data governance and accelerate analyst access across distributed teams, replacing ad-hoc Redshift queries with governed, discoverable datasets.

•Migrated batch workloads to GCP Dataflow and Apache Beam pipelines orchestrated by Cloud Composer, reducing pipeline latency in production and enabling cloud migrations away from legacy AWS Glue jobs for select data domains.

•Deployed GKE-hosted Vertex AI inference services and secured cross-project data access with VPC Service Controls and Pulumi-managed infrastructure, replacing manual CDK scripts and enforcing data masking on sensitive fields at the platform layer.

Data Engineering & Workflow Automation

•Set up Airflow (MWAA) from scratch and wrote shared Python utilities and DAG patterns that the team now uses across 50+ production pipelines.

•Wrote a master Airflow DAG that handles historical backfills with a single trigger, replacing what used to be a multi-day manual process.

•Automated several recurring operational tasks the team was doing by hand, freeing up engineering bandwidth for new data work.

•Built SCD Type 2 dimensional models in Redshift to track historical changes accurately and allow clean rollbacks when needed.

•Set up CI/CD pipelines using CDK so UAT and production deployments follow the same path, which made releases more predictable.

•Instrumented a data warehousing layer in BigQuery with BigQuery Omni and Looker dashboards to deliver business intelligence reporting across multi-cloud datasets, replacing fragmented Redshift SCD models with a unified, self-serve analytics surface.

•Engineered Pub/Sub messaging pipelines and Dataproc Spark jobs following agile sprint cycles, implementing Data Readiness Placement (DRP) checks to validate dataset completeness before downstream DAG dependencies triggered in Airflow.

•Automated data architecture governance by integrating Dataplex data quality scans into CI/CD workflows, storing validated datasets in Google Cloud Storage and enforcing schema contracts across Redshift and BigQuery environments in production.

Cloud Infrastructure & Cost Optimization

•Wrote and deployed AWS infrastructure using CDK in TypeScript, organized across multiple stacks for different environments.

•Reworked the data refresh strategy to cut roughly $30K a month in infrastructure costs.

•Upgraded our AWS Glue jobs from v3 to v4 (Spark 3.3), which reduced job failures and brought down DPU consumption.

•Handled Airflow version upgrades without production downtime by working through backward compatibility issues ahead of time.

•Tuned Spark SQL jobs and query plans to bring down pipeline runtimes and costs and built CloudWatch dashboards for MWAA health and KPI tracking.

Accenture (Client: Chevron)

Data Engineer Aug 2021 – Aug 2022

•contributed to EMR clusters running Hadoop, Hive, and Spark, and pulled data into the data lake from Oracle, MySQL, and SQL Server using Sqoop and JDBC connectors.

•Created ETL pipelines with AWS Glue, Lambda, Step Functions, and PySpark to move ORC, Parquet, and text data from S3 into Redshift for the analytics team's KPI dashboards.

•Set up Kafka for real-time streaming and used Airflow to orchestrate batch and streaming jobs, landing the data into S3 and HDFS and processing it with Spark Structured Streaming.

•Wrote Python and PySpark scripts in AWS Glue for cleansing, enrichment, and aggregation work on the source data.

•Connected Redshift to Tableau and built dashboards with calculated fields for the BI team, working with business stakeholders to sort out data issues and refine reporting requirements.

•Added data quality checks, schema drift detection, and SLA monitoring to the ETL workflows using CloudWatch metrics and SNS alerts.

•Wrote Lambda functions with scoped IAM roles and CloudWatch triggers, used alongside SQS, EventBridge, and SNS for pipeline and infra tasks.

•Replaced Kafka-based messaging systems with a Pub/Sub prototype and rewrote select ETL transforms in Rust for throughput-critical ingestion paths, improving pipeline reliability on high-volume EMR-to-cloud migration workloads.

•Designed a data architecture migration plan for moving Hive and HDFS workloads to GCP Dataproc and Google Cloud Storage, applying data masking rules on PII fields and validating readiness using DRP checklists before cutover.

•Built Looker-connected BigQuery datasets from Redshift exports to extend business intelligence coverage during cloud migrations, replacing Tableau calculated fields with governed Looker LookML models for Chevron operational KPI reporting.

Accenture (Client: Kaiser Permanente)

Data Engineer Oct 2020 – Jul 2021

•Wrote PySpark scripts to move data from S3 into Redshift and contributed to a set of serverless Lambda-based ETL pipelines registered in the AWS Glue Data Catalog.

•Developed end-to-end ETL pipelines in Informatica Power Center to load data from on-prem and cloud platforms into a central warehouse and reworked several slower processes to improve scalability.

•Used AWS Athena for ad-hoc queries on S3 data and Kinesis Data Analytics for real-time stream transformations.

•Ran Spark and Hadoop workloads on EMR clusters set up with auto-scaling to balance cost and performance for healthcare data processing.

•Added data validation, profiling, and reconciliation checks in PySpark and SQL during source-to-target migrations and documented lineage to support HIPAA-aligned data governance.

•Tuned HiveQL joins (map-side and broadcast) to cut down on shuffles and used partitioning and bucketing to keep transformations efficient.

TECHNICAL SKILLS

Cloud and Infra: AWS (S3, EC2, Lambda, Glue, EMR, Redshift, Bedrock, SageMaker, Athena, Kinesis, OpenSearch, Lake Formation, SNS, SQS, Step Functions, CDK, MWAA), GCP (BigQuery), IAM, Eventbridge, CloudWatch

Data Engineering: Spark, Airflow, dbt, Apache Iceberg, Delta Lake, Snowpark, Hadoop, Hive, Trino/Presto, Apache Flink, SQL, Oracle, Tableau, Sqoop

Languages: Python, Scala, Java, TypeScript, Go, Rust

Databases: Snowflake, Databricks, PostgreSQL, DynamoDB, MongoDB, Cassandra, MySQL

AI / ML: AWS Bedrock, LLMs (Claude, GPT, Llama), RAG, Vector Embeddings, Pinecone, MCP Servers, LangChain, LangGraph, Hugging Face, Transformers, Scikit-learn, TensorFlow

DevOps and Tools: Terraform, Docker, Kubernetes, Kafka, Git, CI/CD, Jenkins, JIRA, Datadog, Grafana, Splunk, Dynatrace, Pubsub, Postman

Google Analytics (GA4) - Event and Conversion Tracking

Google Tag Manager (GTM / sGTM) - Client and Server-Side Implementations

AWS EventBridge - Event-driven pipeline triggers and orchestration

Postgres SQL - Complex query optimization and reporting

AWS Glue Data Catalog - Metadata management, Schema Evolution, And Governance

CI/CD with GitHub Actions - Automated deployments for ETL workflows, Encryption

Frameworks & Libraries: WebSocket, JDBC

EDUCATION

University of North Texas Denton, TX

Master of Science in Data Science GPA: 3.9/4.0 Aug 2022 – May 2024



Contact this candidate