Post Job Free
Sign in

Data Engineer for AI & Analytics

Location:
Marietta, GA
Posted:
July 22, 2026

Contact this candidate

Resume:

Vaishnavi Rudroju

Data Engineer - AI & Analytics

*********.*********@*****.*** 413-***-**** LinkedIn

PROFESSIONAL SUMMARY

Data Engineer with 5+ years of experience designing and building scalable batch and streaming data pipelines across financial services and healthcare. Hands-on expertise in Python, SQL, PySpark, Apache Kafka, Snowflake, Databricks, dbt, AWS, and ETL/ELT, delivering reliable data platforms for analytics, reporting, and large-scale transaction processing. Experienced in AI-driven data engineering, integrating LLM APIs and prompt engineering to build intelligent data workflows. Strong background in data quality, performance optimization, and collaborating with engineering, AI/ML, and business teams to deliver secure, scalable, and production-ready solutions.

TECHNICAL SKILLS

•Languages: Python, SQL, PySpark, Shell Scripting

•Big Data & Processing: Apache Spark (PySpark), Databricks, Apache Kafka, Batch & Streaming Pipelines, ETL/ELT, Data Modeling

•Cloud (AWS): S3, EC2, Lambda, CloudWatch

•Databases & Warehouses: Snowflake, Oracle, PostgreSQL

•Data Transformation & Orchestration: dbt, Control-M, Workflow Orchestration, Stored Procedures

•AI / ML: LLM APIs, Prompt Engineering, AI-Ready Dataset Preparation, Claude AI, GitHub Copilot

•Data Quality & Governance: Data Validation, Schema Validation, Reconciliation, Data Governance, Access Controls

•APIs & Services: REST APIs, RESTful Python Services

•Monitoring & Observability: Splunk, AWS CloudWatch, Logging & Alerting

•DevOps & Delivery: Git, CI/CD, JIRA, Agile/Scrum

•Domain: Financial Services, Treasury & Settlement, Post-Trade Processing, Reconciliation, Healthcare Data

EDUCATION

Master’s in Computer Science, Texas A&M University, Kingsville - Dec 2025

WORK EXPERIENCE

Charles Schwab USA Nov 2025 - Present

Software Developer – Data Engineering & AI Solutions

•Built cloud-based data engineering solutions to support AI-driven financial analytics, investment reporting, and large-scale transaction processing across enterprise platforms.

•Designed and maintained scalable batch and streaming data pipelines using Python, SQL, and Apache Kafka, improving data availability and processing efficiency by 35%.

•Integrated LLM APIs and applied prompt engineering techniques to create AI-assisted data workflows that enhanced intelligent search, operational insights, and analytics.

•Used Claude AI and GitHub Copilot to speed up development, simplify debugging, and improve overall engineering productivity.

•Created reusable PySpark, Databricks, and dbt transformation pipelines to process large financial datasets, reducing end-to-end pipeline execution time by 40%.

•Developed RESTful Python APIs for data enrichment, anomaly detection, and AI-powered validation, enabling reliable data for downstream analytics and reporting.

•Built and optimized Snowflake ELT pipelines and SQL transformation models to support customer analytics, investment reporting, operational KPIs, and AI-ready datasets, improving reporting performance by 30%.

•Improved platform reliability by implementing monitoring, logging, and alerting with Splunk and AWS CloudWatch, reducing production support issues by 25%, while leveraging AWS (S3, Lambda, and EC2) for cloud-native data processing.

•Partnered with AI/ML engineers, data engineers, and business teams to modernize reporting solutions, strengthen data governance, and deliver enterprise applications through Git-based CI/CD and Agile/Scrum practices.

UBS India Aug 2021 - Jul 2024

Software Developer- Data Engineering

•Designed and built scalable batch and near real-time data pipelines using Python, SQL, PySpark, Apache Kafka, and ETL/ELT workflows to process, validate, and reconcile millions of daily financial transactions across treasury, settlement, and post-trade systems.

•Developed complex SQL transformations, stored procedures, and reconciliation rules to support transaction matching, exception detection, break analysis, and operational and regulatory reporting.

•Implemented automated data quality and data validation frameworks using control totals, schema validation, file integrity checks, duplicate detection, and reference-data validation to improve financial data accuracy and reliability.

•Optimized high-volume SQL queries, indexing strategies, partitioned datasets, and ETL workflows, improving reconciliation accuracy by 80% while reducing manual investigation efforts by 60%.

•Engineered batch and streaming data pipelines using Apache Kafka and PySpark to process multi-million-record datasets, improving data throughput, validation efficiency, and reconciliation performance.

•Migrated 10+ legacy reconciliation processes to Gresham Clareti Transaction Control and leveraged Oracle and Snowflake for data ingestion, transformation, reconciliation reporting, and KPI analysis.

•Implemented workflow orchestration with Control-M and enhanced monitoring, logging, and alerting using Splunk, AWS CloudWatch, and SQL-based health checks to improve pipeline observability and production support.

•Developed SQL-based KPI reports and operational dashboards to monitor match rates, exception volumes, reconciliation trends, SLA performance, and pipeline health.

•Performed root-cause analysis of transaction mismatches, failed loads, and reconciliation breaks, improving platform reliability and settlement processing accuracy.

•Collaborated with Engineering, Operations, Finance, and Compliance teams to deliver scalable data engineering solutions, supporting DEV, SIT, UAT, and Production releases in an Agile/Scrum environment using Git and JIRA.

GE Healthcare India Jul 2020 - Aug 2021

Jr. Data Engineer

•Built and maintained ETL/ELT pipelines using Python and SQL to ingest, transform, clean, and validate healthcare datasets for patient monitoring, device analytics, and operational reporting, processing over 500K+ records daily.

•Developed SQL queries, stored procedures, and data transformation workflows that improved report generation performance by 35% while reducing manual data validation efforts.

•Processed large-scale healthcare datasets using PySpark and batch-processing frameworks, performing data cleansing, schema validation, and quality checks to ensure reporting accuracy and reliability.

•Automated file validation, exception handling, and routine data quality checks with Python, while optimizing SQL queries and ETL workflows to improve daily pipeline performance.

•Utilized AWS (S3, EC2, CloudWatch) for data storage, monitoring, and pipeline support, and assisted with ETL job scheduling, production support, and troubleshooting in Linux environments.

•Collaborated with analysts and cross-functional teams to integrate data from multiple healthcare applications into centralized reporting platforms, following Agile/Scrum practices using Git and JIRA.

PROJECTS

Real-Time Financial Data Streaming Pipeline using Kafka & Spark

•Built an end-to-end real-time data pipeline to process streaming transaction data using Apache Kafka and PySpark Structured Streaming.

•Implemented data ingestion, transformation, and validation for high-volume event streams.

•Stored processed data in Snowflake / AWS S3 for analytics and reporting.

•Designed monitoring and alerting for data quality issues using Python and logging frameworks.

•Enabled near real-time dashboards for transaction monitoring and anomaly detection.

AI-Powered Financial Document Intelligence Pipeline

•Built an end-to-end pipeline that ingests unstructured financial documents, extracts key entities using LLM APIs and prompt engineering, and loads structured, AI-ready output into Snowflake via dbt for analytics and semantic search.

•Implemented validation and confidence-scoring logic to flag low-quality LLM extractions for human review, improving output reliability.



Contact this candidate