MADHURI YENUGULA
Data Engineer
+1-726-***-**** ****************@*****.*** linkedin.com/in/madhuri19 github.com/19madhu Denton, TX
PROFESSIONAL SUMMARY
Data Engineer with 3+ years of experience building and optimizing batch and streaming data pipelines across AWS and GCP using Kafka, Hadoop, and BigQuery. Strong in PySpark, Spark SQL, and Databricks Delta Lake for large-scale ETL, with hands-on delivery of event-driven ingestion using AWS Glue, Lambda, Step Functions, and Kinesis into S3 lakehouse architectures with Apache Beam. Experienced in dimensional modeling and modular SQL transformations with dbt and Snowflake, workflow orchestration with Apache Airflow, data-quality validation with Great Expectations, and Data Governance practices. Owns pipelines end to end, from ingestion and modeling through Vertex AI integration, CI/CD deployment, monitoring, and incident response
PROFESSIONAL EXPERIENCE
Data Engineer, Spark Technologies Texas, USA May 2025 – Present
•Build and maintain distributed PySpark and Spark SQL pipelines on AWS EMR and GCP Dataproc for a client enterprise analytics platform, processing batch and streaming workloads that feed downstream analytics and Machine Learning teams.
•Re-engineered legacy MapReduce jobs into optimized PySpark and Spark SQL as part of a data modernization initiative, cutting average job runtime by 40%.
•Model curated Snowflake marts using dbt for modular SQL transformations and standardized star-schema definitions, improving data-lineage visibility and reducing transformation defects across reporting teams.
•Implement Great Expectations data-quality checks and CloudWatch, Grafana, and ELK monitoring for faster incident detection, and deploy jobs through Terraform, Docker, and GitLab CI/CD.
•Developed event-driven data engineering ingestion with AWS Glue, Lambda, Step Functions, and Kinesis to onboard 30+ structured and semi-structured sources into a centralized S3 lakehouse on Delta Lake, exposing curated tables for analysis through Athena.
•Migrated legacy Oozie schedules to parameterized Airflow DAGs on Cloud Composer, adding retries, SLAs, and proactive alerting to harden pipeline reliability and reduce on-call incidents.
•Migrated on-premises data warehouse workloads to GCP by provisioning GKE clusters and Google Cloud Storage buckets with Pulumi infrastructure-as-code modules, enabling a self-serve data platform for analytics teams across distributed systems.
•Secured BigQuery datasets by enforcing VPC Service Controls, column-level data masking, and envelope encryption on sensitive fields, satisfying data security and data readiness placement requirements in production.
•Accelerated analytical reporting by building Dataflow streaming pipelines that feed BigLake external tables and Looker dashboards, integrating BigQuery ML models for in-warehouse predictions and BigQuery Omni queries across multi-cloud sources.
Data Engineer, ConnectGen, Inc. Hyderabad, India Apr 2022 – Apr 2024
•Built and maintained batch ETL workflows on AWS serverless services including Glue Jobs, Lambda, and API Gateway, transforming raw source data into structured JSON payloads consumed by Next.js and MongoDB.
•Designed Amazon DynamoDB and S3 data models with translation and cache validation logic, tuning table write patterns to keep event-driven ingestion stable under production load.
•Automated deployment of data pipelines and cloud infrastructure with AWS CDK and CloudFormation, standardizing releases and reducing manual configuration across environments.
•Troubleshot production data issues by tracing system logs and reverse-engineering existing codebases, resolving processing anomalies.
•Orchestrated containerized ETL services on Kubernetes via GKE, replacing ad-hoc EC2 deployments and cutting cloud migration downtime for batch ingestion jobs processing millions of daily records.
•Engineered a Pub/Sub-based messaging system in Rust to fan out real-time events from DynamoDB streams into Dataflow jobs, reducing end-to-end pipeline latency across distributed ingestion services.
•Governed the data lake by configuring Dataplex zones over Google Cloud Storage assets with automated data-quality scans and IAM-bound VPC Service Controls, strengthening data warehousing security posture in line with GCP Professional Data Engineer best practices.
PROJECTS
Big Data Engineering Pipeline using Databricks & AWS Nov 2024 - Dec 2024
Technologies: Databricks, Delta Lake, PySpark, Amazon S3, EMR, Hive, GitHub Actions, Parquet
•Built modular PySpark ETL pipelines in Databricks on AWS to ingest and transform 114,000+ raw files into curated Delta tables partitioned for downstream query performance.
•Enforced schema validation on every load with Delta Lake and automated multi-environment deployment through GitHub Actions CI/CD workflows.
Consolidated Analytics Data Layer and Dashboard Jan 2025 - Mar 2025
Technologies: Tableau, Power BI, SQL, D3.js, Plotly, JavaScript, Netlify
•Consolidated multiple source tables into a conformed data modeling layer feeding 10,000+ monthly records to Tableau and Power BI dashboards for business intelligence reporting.
•Delivered responsive web components with D3.js and Plotly, deployed on Netlify to support executive reporting and self-serve trend analysis.
PUBLICATION
Harnessing Climate Data for Accurate Crop Yield Predictions Jan 2025
Published in IJRASET, Vol. 13, Issue 1, January 2025 DOI: doi.org/10.22214/ijraset.2025.66566
•Engineered automated data validation and feature preparation routines over multi-source climate and crop datasets for Random Forest yield models.
TECHNICAL SKILLS
•Languages and Querying: Python (PySpark, Pandas, NumPy), HiveQL, Java, JavaScript, Shell Scripting, Go, Rust
•Big Data and Processing: SQL, Spark SQL, Spark Structured Streaming, Databricks, Delta Lake, Hadoop, Hive, Snowflake, dbt, PostgreSQL, MySQL, MongoDB, Cassandra, Airflow, Tableau, Power BI, Machine Learning
•Cloud Platforms: Amazon Kinesis, AWS (S3, EMR, Glue, Lambda, Step Functions, Athena, DynamoDB, RDS, API Gateway, IAM, CloudWatch, SNS, SQS, CDK, CloudFormation), GCP (Dataproc, Cloud Composer), Netlify
•Warehousing and Modeling: Dimensional Modeling (Star and Snowflake Schema), Data Mapping, Schema Validation, Partitioning, Parquet
•Orchestration and DevOps: Oozie, Terraform, Docker, CI/CD, JIRA, GitLab-CI, Kubernetes, Agile (SAFe)
•Monitoring and BI: Great Expectations, Data Validation, Data Lineage, ELK Stack, CloudWatch Dashboards, Data Quality, Next.js, Encryption
•Tools & Platforms: Kafka, Git, Grafana, ELK, Pubsub
CERTIFICATIONS
•AWS Certified Cloud Practitioner
•AWS Certified Solutions Architect
•IBM Data Analyst Professional Certificate
EDUCATION
University of North Texas Denton, TX
Master of Science Computer Science Aug 2024 - May 2026
Kakatiya Institute of Technology and Science Nizamabad, India
Bachelor of Technology Computer Science June 2019 - June 2023