Post Job Free
Sign in

Data Engineer

Location:
Dallas, TX
Salary:
$68/hr on c2c
Posted:
August 31, 2026

Contact this candidate

Resume:

SAI UDAY VAJRAPU

Senior Data Engineer

Allen, TX +1-646-***-**** ***.**.***@*****.*** LinkedIn

PROFESSIONAL SUMMARY

Data Engineer with 8+ years of experience designing and scaling ETL/ELT pipelines, cloud data platforms, and streaming architectures for retail, pharmaceutical, healthcare, and telecom organizations. Proven track record cutting Spark runtimes by up to 95%, reducing annual cloud spend by $500K+, and delivering governed, production-grade data for BI, machine learning, and Generative AI/RAG use cases. Skilled across the full pipeline lifecycle batch and streaming ingestion, dimensional modeling, embedding and vector pipelines, and DataOps practices (CI/CD, automated testing, observability) on Azure, AWS, Snowflake, and Databricks, with delivery held to HIPAA and GDPR standards.

CORE COMPETENCIES

ETL/ELT Pipeline Development • Data Lakehouse Architecture • Real-Time & Streaming Data (Kafka, Flink) • Distributed Computing (Apache Spark) • Cloud Data Platforms (Azure, AWS, GCP) • Snowflake & Databricks • Dimensional Data Modeling • Change Data Capture (CDC) • Data Warehousing • Workflow Orchestration (Airflow, dbt, ADF) • Data Governance & Security • Data Quality & Observability • CI/CD &

DataOps • Performance Tuning & Cost Optimization • MLOps & ML Data Pipelines • Vector Databases & RAG • Feature Engineering • Big Data (Hadoop, Hive)

TECHNICAL SKILLS

Programming & Scripting: Python, SQL, PySpark, Scala, Java, R, Shell Scripting

Big Data & Streaming: Apache Spark (RDDs, DataFrames, Spark SQL, DAGs), Hadoop (HDFS, YARN), Hive, Sqoop, Apache Kafka, Apache Flink, Delta Lake, Apache Iceberg

Cloud & Warehousing: Azure (Databricks, Data Factory, Blob Storage, Event Hubs, Synapse, RBAC), AWS (S3, Glue, EMR, Lambda, Step Functions, Redshift, Athena, SageMaker, EC2, IAM), Snowflake (Snowpark, Snowpipe, SnowSQL), Databricks (Delta Live Tables, Unity Catalog, Workflows)

Orchestration & ETL: Apache Airflow, dbt, Azure Data Factory, Talend, Informatica, Alteryx, Fivetran/Airbyte, SSIS

AI/ML & Data Science: Feature Engineering, Multi-modal Data Prep, Vector Databases (pgvector, Pinecone), Feature Stores (Feast), MLflow, RAG Data Pipelines, scikit-learn, TensorFlow, Pandas, NumPy

Governance, Quality & DevOps: Immuta, Monte Carlo, Great Expectations, Unity Catalog, Liquibase, Docker, Kubernetes (Rancher), Terraform, GitLab, GitHub Actions, Dynatrace, GDPR, HIPAA

Databases & BI: MS SQL Server, PostgreSQL, MongoDB, Oracle, Teradata, MySQL, Cosmos DB, DB2; Power BI (DAX, RLS/OLS), Tableau, Matplotlib, Seaborn

PROFESSIONAL EXPERIENCE

Data Engineer Kroger Ohio, USA Sept 2024 – Present

•Engineered an incremental ingestion backbone orchestrating Azure Data Factory runs and Unity Catalog-governed Databricks notebooks with Change Data Feed and control-table watermarks to land text, tabular, and image payloads into Delta Shares with zero duplication.

•Streamed sub-minute pricing updates built PySpark/Spark SQL jobs against Kafka and Event Hubs to reduce bulk domain feeds to the exact attributes Java pricing services required.

•Cut Spark batch runtime ~95% by salting/repartitioning to eliminate data skew, broadcasting dimension tables, enabling adaptive query execution, and right-sizing shuffle partitions, lowering Databricks DBU consumption significantly.

•Built the data foundation for Generative AI/RAG initiatives curating paired text-and-image datasets, generating embeddings, and loading pgvector and Pinecone indexes that power recommendation and RAG features.

•Delivered $500K in annual cloud savings and 9x faster response times identifying idle clusters via Dynatrace and Rancher, and right-sizing pod CPU/memory by ~60% to bring a 46-second response time under 5 seconds, while shipping governed Power BI/Tableau layers validated by Great Expectations for fraud and BSA/AML monitoring.

Data Engineer / Technology Consultant Roche California, USA Feb 2023 – Sept 2024

•Migrated Palantir Foundry workflows to Snowflake rebuilding logic as dbt models with shared macros and assertion tests, preserving lineage and reconciling outputs precisely to historical figures.

•Orchestrated dbt and Python pipelines on Airflow (AKS) with retry/backoff logic, SLA-breach alerting, and documented runbooks, giving on-call engineers a clear path to recover failed loads.

•Modeled a conformed Redshift warehouse connecting S3 to Redshift via Glue and Athena and designing star/snowflake schemas with slowly changing dimensions to support analytics and regulatory reporting.

•Strengthened data trust and compliance introducing Monte Carlo observability (freshness, volume, distribution, schema) on business-critical Snowflake tables and applying Immuta attribute-based access policies so AI/ML teams could experiment within HIPAA and GDPR bounds.

•Led Snowflake/dbt/Monte Carlo adoption across Roche/Genentech gating every pull request through dbt tests and Monte Carlo checks in a GitHub Actions/Liquibase CI/CD flow, while running enablement and coordinating access, encryption, and DLP across five time zones.

Data Engineer – R&D Anritsu Hyderabad, India May 2021 – July 2021

•Built Spark/PySpark pipelines over Hive and Spark SQL consolidating scattered marketing data sources into Delta tables and integrating legacy systems through AWS.

•Scaled Snowflake ingestion loading Parquet and CSV data via Snowpipe and external tables, provisioning warehouses/schemas, and verifying source-to-target parity throughout migration.

•Containerized and automated the orchestration stack running dbt and Airflow on AKS via Docker with GitHub Actions/GitLab CI/CD for fully reproducible deployments.

•Modernized analytics workloads transitioning from Hive/Cloudera and Foundry-style platforms to Snowflake across AWS and Azure, refining incremental dbt models and auto-publishing lineage via dbt docs.

Data Analyst HCA Healthcare Hyderabad, India Jan 2018 – April 2021

•Automated cross-system data movement building Airflow workflows with Spark-driven ETL into HDFS and Hive from varied upstream sources.

•Modernized legacy analytics migrating Teradata and Hive queries to Azure and rewriting Spark/Scala dataflows to scale with growing volumes.

•Prepared data for early ML initiatives collecting, cleansing, and engineering features from Azure and Databricks sources to feed model training and evaluation.

•Standardized warehouse documentation contributing to ETL strategy and dimensional models, and maintaining data dictionaries, source-to-target mappings, and DDL/DML.

•Improved reporting and data quality developing Tableau and Power BI dashboards and partnering with stakeholders to isolate root causes across the ETL layer and source systems.

CERTIFICATIONS

SnowPro Core Certification — Certificate ID: S0029066

EDUCATION

Master of Science – Data Science Aug 2021 – Dec 2022

Lewis University, Chicago, IL, USA

Bachelor of Technology – Electrical & Electronics Engineering Aug 2014 – May 2018

Vasavi College of Engineering, Hyderabad, India



Contact this candidate