SAI ANVESH ADEPU
GCP / BIG DATA ENGINEER
Ohio, USA Open to Relocate +1-216-***-**** ***********@*****.*** LinkedIn PROFESSIONAL SUMMARY
GCP / Big Data Engineer with 3+ years of experience building large-scale data pipelines, cloud data warehouses, and distributed streaming systems on GCP and AWS. Hands-on with BigQuery, Cloud Storage, Dataflow (Apache Beam), Pub/Sub, Apache Spark, PySpark, and Kafka, delivering pipelines that process 2TB+ daily and 100B+ records while cutting job runtime 50% and saving $30K+ annually in compute costs. Also built Spark Structured Streaming and Snowflake lakehouse pipelines processing 5TB+ daily with 3x faster queries and 99.5% data quality. Skilled in CDC, Delta Lake, dbt, and cloud data warehousing at scale. TECHNICAL SKILLS
Programming: Python, SQL, PySpark, Scala, Java, Bash GCP / Cloud Platforms: GCP (BigQuery, Dataflow, Pub/Sub, Cloud Composer, GCS); AWS (S3, EMR, Redshift, Glue, Lambda, Athena, Step Functions, Lake Formation, MSK); Azure (Data Factory, Synapse Analytics, ADLS) Big Data / Distributed Processing: Apache Spark, PySpark, Apache Beam, Hadoop, Hive, Delta Lake Streaming & CDC: Apache Kafka, Kafka Streams, Spark Structured Streaming, Pub/Sub, CDC, Debezium, Apache Flink, AWS Kinesis
Data Warehousing / Lakehouse: Snowflake, Google BigQuery, Amazon Redshift, Azure Synapse, Databricks SQL, Delta Lake, Apache Iceberg, Apache Hudi, AWS Lake Formation Data Transformation & Orchestration: dbt (Core & Cloud), Spark SQL, Apache Airflow, AWS Step Functions, Prefect, Azure Data Factory
Data Quality / Observability: Great Expectations, Monte Carlo, DataDog, Grafana, CloudWatch, PagerDuty Infrastructure / DevOps: Docker, Kubernetes, Terraform, Jenkins, GitHub Actions, Git, CI/CD Data Modeling: Star Schema, Snowflake Schema, Data Vault 2.0, Dimensional Modeling, SCD Databases: PostgreSQL, MySQL, MongoDB, Cassandra, DynamoDB, Redis Data Engineering Concepts: ETL/ELT, Data Mesh, Medallion Architecture, DataOps, MLOps, CDC, Data Governance
PROFESSIONAL EXPERIENCE
PALANTIR TECHNOLOGIES June 2025 – Present
Data Engineer Ohio, USA
● Architected real-time and batch pipelines processing 5TB+ daily using Apache Kafka, Spark Structured Streaming, and AWS S3; implemented Medallion Architecture (Bronze/Silver/Gold) on Delta Lake, cutting end-to-end latency from hours to under 5 minutes.
● Led migration of the cloud data warehouse from on-prem systems to Snowflake on AWS using Data Vault 2.0 modeling, micro-partitioning, and clustering, improving query performance 3x and reducing infrastructure costs by 40%.
● Built 50+ dbt Core transformation models with automated tests, schema documentation, and full lineage tracking, enabling self-service analytics for 200+ stakeholders across product, finance, and operations.
● Designed scalable ETL/ELT workflows with Python, PySpark, Apache Airflow, and AWS Glue, integrating REST APIs, CDC streams (Debezium), and third-party SaaS data into a centralized data lakehouse.
● Implemented an automated data quality framework using Great Expectations and Monte Carlo, raising the data accuracy and completeness SLA from 85% to 99.5% and reducing data incidents by 70%.
● Built CI/CD pipelines and infrastructure-as-code with GitHub Actions, Terraform, Docker, and Kubernetes, reducing deployment failures and cutting release cycle time by 60%.
● Established pipeline monitoring and alerting with DataDog and PagerDuty across 100+ production jobs, improving operational stability and reducing pipeline downtime by 45%.
● Partnered with ML engineers to build feature store pipelines on AWS SageMaker Feature Store and Databricks, reducing model training data preparation time by 30%. DXC TECHNOLOGY June 2021 – July 2023
Data Engineer India
● Implemented a cloud-native data lake on GCP using Cloud Storage (Parquet/ORC), BigQuery as the analytical store, and Dataflow (Apache Beam) for streaming ingestion, supporting analytics on 100B+ records.
● Designed and deployed batch ETL pipelines with Apache Airflow, PySpark, and SQL, processing 2TB+ daily from APIs, databases, and Kafka event streams into GCP BigQuery and Cloud Storage.
● Adopted Delta Lake and Apache Hudi for ACID-compliant lakehouse storage, enabling CDC patterns and upsert operations across large historical datasets.
● Optimized PySpark jobs and BigQuery SQL through partitioning, Z-ordering, broadcast joins, and predicate pushdown, cutting average job runtime by 50% and saving $30K+ annually in compute costs.
● Built dimensional data models (star schema, SCD Type 2) and dbt transformation layers to support KPI dashboards, revenue analytics, and customer behavior segmentation for 15+ business units.
● Developed real-time Pub/Sub and Dataflow streaming pipelines for event-driven ingestion, enabling near-real- time fraud detection dashboards with latency under 2 seconds.
● Partnered with data scientists to design feature stores and ML training datasets using Vertex AI Feature Store and BigQuery ML, improving model training efficiency by 35%.
● Authored Terraform modules for GCP infrastructure provisioning and established GitLab CI/CD pipelines to automate deployments across dev, staging, and production environments. EDUCATION
CLEVELAND STATE UNIVERSITY Aug 2023 – May 2025
Master of Science in Information Sciences Ohio, USA