JAIDEEP REDDY
Data Engineer GCP AWS PySpark Python GenAI / RAG
+1-469-***-**** *****************@*****.*** linkedin.com/in/jaideep-reddy-aluru-138564151
PROFESSIONAL SUMMARY
Results-driven Data Engineer with 5+ years of experience designing and delivering scalable ETL/ELT pipelines, enterprise data platform modernizations, and cloud migrations on Google Cloud Platform (GCP) and AWS. Experienced across Healthcare and Financial Services domains, with a strong track record of translating complex business requirements into reliable, high-quality data products. Deep hands-on expertise in Python, SQL, BigQuery, Apache Spark/PySpark, Apache Airflow, Snowflake, Databricks, Kafka, and Hadoop. Currently building Generative AI solutions using Vertex AI, Google ADK, and Retrieval-Augmented Generation (RAG) to enable natural-language data discovery. Master of Science in Data Science (University of North Texas). Effective cross-functional collaborator thriving in SAFe Agile/Scrum environments.
CORE COMPETENCIES
• ETL/ELT Pipeline Development
• Google Cloud Platform (GCP)
• Apache Spark / PySpark / Spark SQL
• Data Warehousing & Data Modeling
• Amazon Web Services (AWS)
• Apache Airflow / Cloud Composer
• Data Quality & Validation Frameworks
• BigQuery / Snowflake / Redshift
• Python / SQL / PL/SQL / Scala
• Data Migration & Integration
• Databricks / Hadoop / Kafka / HDFS
• Vertex AI / Generative AI / RAG
• Agile / SAFe Scrum Delivery
• CI/CD (Git, Jenkins)
• Looker / Power BI / Tableau
• Stakeholder Collaboration
• Data Governance & Compliance
• MongoDB / Oracle / PostgreSQL
TECHNICAL SKILLS
Programming Languages: Python, SQL, Scala, Shell Scripting, PL/SQL
Cloud Platforms: Google Cloud Platform (GCP), Amazon Web Services (AWS)
GCP Services: BigQuery, Cloud Composer, Cloud Storage, Dataproc, Dataflow, Cloud Functions, Bigtable, Cloud SQL, Vertex AI
AWS Services: Amazon S3, AWS Glue, Amazon EMR, Amazon Redshift, Amazon Athena, AWS Lambda
Big Data & Distributed Proc.: Apache Spark, PySpark, Spark SQL, Spark Streaming, Hadoop, HDFS, Hive, Databricks, Kafka, Sqoop, Cloudera
Data Engineering & ETL: ETL/ELT, Data Pipeline Development, Data Ingestion, Data Transformation, Data Migration, Data Warehousing, Data Quality, Data Validation, Data Reconciliation
Data Orchestration & DevOps: Apache Airflow, Cloud Composer, Git, Jenkins, CI/CD, Linux/Unix
Databases & Data Warehouses: BigQuery, Snowflake, Teradata, Oracle, PostgreSQL, SQL Server, MySQL, Cloud SQL, NoSQL
Generative AI & ML: Vertex AI, Google ADK, Vertex AI Agent Builder, RAG, Google RAG Corpus, LLMs, AI Agents, Natural Language Data Retrieval
BI & Visualization: Looker, Power BI, Tableau, Matplotlib, Google Analytics
PROFESSIONAL EXPERIENCE
Data Engineer CVS Health (via Consulting) — Healthcare Apr 2025 – Present
Business & Stakeholder Collaboration (Daily Activities) — 30%:
Collaborate daily with Product Owners, Business Analysts, Technical Leads, and QA teams to translate complex healthcare business requirements into scalable data engineering solutions within SAFe Agile/Scrum environments.
Participate in sprint planning, backlog refinement, daily stand-ups, UAT validation, pre-production validation, production deployments, and post-release verification, maintaining alignment between technical deliverables and business objectives.
Work closely with business users to validate migrated datasets, ensure reporting accuracy, and support business-critical healthcare analytics and reporting platforms.
Prepare and maintain technical documentation — data flow diagrams, source-to-target mappings, migration plans, ETL specifications, operational runbooks, and troubleshooting procedures — ensuring knowledge transfer and audit compliance.
Provide production support and timely issue resolution for critical healthcare applications, communicating pipeline health metrics and data quality findings to stakeholders.
Technical Responsibilities — 70%:
Design, develop, and maintain scalable ETL/ELT data pipelines using Python, SQL, BigQuery, and GCP services (Cloud Storage, Cloud SQL, Dataproc, Bigtable, Cloud Composer) to ingest, transform, validate, and deliver enterprise healthcare data at scale.
Lead Teradata-to-BigQuery migration initiatives including table conversion, SQL query remediation, data validation, reconciliation, and production deployment, enabling enterprise-scale cloud modernization of healthcare data workloads.
Develop complex SQL transformations, views, stored procedures, and optimized BigQuery workloads across large-scale structured and semi-structured healthcare datasets, improving processing efficiency and downstream reporting.
Build automated data quality, reconciliation, and validation frameworks using Python and SQL to ensure source-to-target accuracy, completeness, and integrity across millions of healthcare records.
Engineer and orchestrate cloud-based data workflows using Apache Airflow on Cloud Composer, managing DAGs, scheduling, dependencies, and end-to-end monitoring of production healthcare data pipelines.
Develop metadata-driven solutions, audit tracking mechanisms, and enterprise data governance controls ensuring healthcare data compliance, lineage, and quality standards (HIPAA-aligned).
Troubleshoot and resolve complex production failures involving BigQuery, Dataproc, Airflow, Spark, DB2 connectivity, network dependencies, and cloud infrastructure, minimizing downtime for critical healthcare systems.
Automate data ingestion, transformation, validation, and operational workflows using Python and Unix shell scripting, reducing manual processing and improving pipeline reliability and operational efficiency.
Implement CI/CD processes using Git and Jenkins, supporting automated testing, version control, deployment, and release management for data engineering solutions.
Build a Generative AI data assistant using Google Agent Development Kit (ADK), Vertex AI, and Retrieval-Augmented Generation (RAG) with Google RAG Corpus to enable business users to query enterprise healthcare data via natural language.
Design and implement RAG-based workflows to retrieve structured enterprise data context from BigQuery metadata, supporting intelligent natural-language data discovery and analytics capabilities.
Key Achievements:
Contributed to enterprise-scale Teradata-to-BigQuery migration, converting and validating large-scale healthcare data workloads and accelerating cloud modernization.
Enhanced data quality and compliance through automated reconciliation and validation frameworks processing millions of healthcare records with source-to-target accuracy.
Developed AI-powered enterprise data assistant using Google ADK, Vertex AI, and RAG, enabling natural-language querying of enterprise healthcare datasets.
Delivered cloud-native healthcare data solutions on GCP supporting business-critical analytics, reporting, and governance workloads.
Environment: GCP (BigQuery, Cloud Composer, Cloud Storage, Cloud SQL, Dataproc, Bigtable), Python, SQL, Apache Airflow, Teradata, DB2, Vertex AI, Google ADK, RAG, Git, Jenkins, Unix/Linux, ETL/ELT, Data Quality, Data Governance, Generative AI, SAFe Agile
Data Engineer Fannie Mae (via Consulting) — Financial Services Jan 2023 – Apr 2025
Business & Stakeholder Collaboration:
Collaborated with cross-functional teams in Agile/Scrum to gather requirements, translate financial business use cases into data solutions, and deliver analytics-ready datasets for mortgage and financial reporting.
Developed Python-based statistical analysis and exploratory data analysis (EDA) solutions for financial datasets, surfacing trends and actionable patterns for business stakeholders and BI teams.
Supported analytical and BI workloads using Amazon Redshift and Amazon Athena, optimizing SQL queries for large-scale financial data analysis and stakeholder reporting.
Participated in code reviews, testing cycles, production deployments, and technical documentation to maintain SDLC compliance and quality standards.
Technical Responsibilities:
Designed and developed scalable ETL/ELT pipelines using AWS, PySpark, Python, Apache Airflow, and SQL to process large-scale financial mortgage datasets.
Developed batch and near-real-time data processing solutions using Apache Spark, PySpark, Spark SQL, and Hadoop for efficient transformation and analytics of large financial datasets.
Built and orchestrated end-to-end data workflows using Apache Airflow, integrating Amazon S3, AWS Glue, Amazon EMR, AWS Lambda, and Amazon Redshift.
Developed and optimized PySpark/Spark SQL jobs on distributed environments for data cleansing, transformation, aggregation, and large-scale analytics.
Designed data ingestion pipelines extracting data from MongoDB and relational databases; used Python/PyMongo to handle semi-structured JSON/BSON documents and process via PySpark DataFrames.
Implemented source-to-target data validation and reconciliation between MongoDB, source systems, and downstream AWS datasets, ensuring data completeness, accuracy, and consistency.
Developed AWS-based data integration workflows ingesting on-premises and cloud data into Amazon S3 data lake, with transformation and loading into analytical stores.
Built AWS Glue ETL processes to catalog, transform, and prepare financial datasets for downstream analytics and reporting.
Environment: AWS (S3, Glue, EMR, Lambda, Redshift, Athena), Apache Airflow, PySpark, Spark SQL, Python, MongoDB, PyMongo, Hadoop, HDFS, SQL, Git, Linux/Unix, ETL/ELT
Jr. Data Engineer Cognizant Technology Solutions Jan 2021 – Dec 2021
Business & Stakeholder Collaboration:
Collaborated with development and business teams in Agile/Scrum, gathering requirements, performing data profiling and source-system assessments to support pipeline development.
Created technical documentation including data flow diagrams, technical specifications, workflow documentation, and standard operating procedures for knowledge transfer and maintainability.
Technical Responsibilities:
Developed scalable ETL pipelines using Python and PySpark to process structured and unstructured datasets across distributed Hadoop/Spark environments.
Designed data ingestion workflows from Oracle databases into Hadoop/Spark environments, supporting large-scale data transformation and analytics.
Developed complex SQL and PL/SQL queries, stored procedures, and data transformations to extract, cleanse, aggregate, and load enterprise data.
Built PySpark DataFrame transformations for data cleansing, enrichment, aggregation, and business-rule implementation.
Developed and deployed PySpark ETL pipelines on Databricks, leveraging distributed computing clusters for efficient large-dataset processing.
Automated recurring batch ETL workflows using Apache Spark and Oracle Scheduler, reducing manual intervention and improving operational reliability.
Tuned Oracle SQL queries and indexing strategies, improving data retrieval and ETL performance.
Participated in migration of Oracle-based ETL workloads to cloud and Spark-based distributed processing environments.
Environment: Python, PySpark, Apache Spark, Databricks, Hadoop, Oracle, SQL, PL/SQL, HDFS, ETL, Data Quality, Linux/Unix
EDUCATION
Master of Science (M.S.) — Data Science
University of North Texas Denton, TX 2023
Bachelor of Technology (B.Tech) — Electronics and Communication
Jawaharlal Nehru Technological University Hyderabad India 2021