Post Job Free
Sign in

GCP Big Data Engineer - Streaming & Warehousing

Location:
East Side, OH, 44114
Salary:
65000
Posted:
October 07, 2026

Contact this candidate

Resume:

SAI ANVESH ADEPU

GCP / BIG DATA ENGINEER

Ohio, USA Open to Relocate +1-216-***-**** ***********@*****.*** LinkedIn PROFESSIONAL SUMMARY

GCP / Big Data Engineer with 3+ years of experience building large-scale data pipelines, cloud data warehouses, and distributed streaming systems on GCP and AWS. Hands-on with BigQuery, Cloud Storage, Dataflow (Apache Beam), Pub/Sub, Apache Spark, PySpark, and Kafka, delivering pipelines that process 2TB+ daily and 100B+ records while cutting job runtime 50% and saving $30K+ annually in compute costs. Also built Spark Structured Streaming and Snowflake lakehouse pipelines processing 5TB+ daily with 3x faster queries and 99.5% data quality. Skilled in CDC, Delta Lake, dbt, and cloud data warehousing at scale. TECHNICAL SKILLS

Programming: Python, SQL, PySpark, Scala, Java, Bash GCP / Cloud Platforms: GCP (BigQuery, Dataflow, Pub/Sub, Cloud Composer, GCS); AWS (S3, EMR, Redshift, Glue, Lambda, Athena, Step Functions, Lake Formation, MSK); Azure (Data Factory, Synapse Analytics, ADLS) Big Data / Distributed Processing: Apache Spark, PySpark, Apache Beam, Hadoop, Hive, Delta Lake Streaming & CDC: Apache Kafka, Kafka Streams, Spark Structured Streaming, Pub/Sub, CDC, Debezium, Apache Flink, AWS Kinesis

Data Warehousing / Lakehouse: Snowflake, Google BigQuery, Amazon Redshift, Azure Synapse, Databricks SQL, Delta Lake, Apache Iceberg, Apache Hudi, AWS Lake Formation Data Transformation & Orchestration: dbt (Core & Cloud), Spark SQL, Apache Airflow, AWS Step Functions, Prefect, Azure Data Factory

Data Quality / Observability: Great Expectations, Monte Carlo, DataDog, Grafana, CloudWatch, PagerDuty Infrastructure / DevOps: Docker, Kubernetes, Terraform, Jenkins, GitHub Actions, Git, CI/CD Data Modeling: Star Schema, Snowflake Schema, Data Vault 2.0, Dimensional Modeling, SCD Databases: PostgreSQL, MySQL, MongoDB, Cassandra, DynamoDB, Redis Data Engineering Concepts: ETL/ELT, Data Mesh, Medallion Architecture, DataOps, MLOps, CDC, Data Governance

PROFESSIONAL EXPERIENCE

PALANTIR TECHNOLOGIES June 2025 – Present

Data Engineer Ohio, USA

● Architected real-time and batch pipelines processing 5TB+ daily using Apache Kafka, Spark Structured Streaming, and AWS S3; implemented Medallion Architecture (Bronze/Silver/Gold) on Delta Lake, cutting end-to-end latency from hours to under 5 minutes.

● Led migration of the cloud data warehouse from on-prem systems to Snowflake on AWS using Data Vault 2.0 modeling, micro-partitioning, and clustering, improving query performance 3x and reducing infrastructure costs by 40%.

● Built 50+ dbt Core transformation models with automated tests, schema documentation, and full lineage tracking, enabling self-service analytics for 200+ stakeholders across product, finance, and operations.

● Designed scalable ETL/ELT workflows with Python, PySpark, Apache Airflow, and AWS Glue, integrating REST APIs, CDC streams (Debezium), and third-party SaaS data into a centralized data lakehouse.

● Implemented an automated data quality framework using Great Expectations and Monte Carlo, raising the data accuracy and completeness SLA from 85% to 99.5% and reducing data incidents by 70%.

● Built CI/CD pipelines and infrastructure-as-code with GitHub Actions, Terraform, Docker, and Kubernetes, reducing deployment failures and cutting release cycle time by 60%.

● Established pipeline monitoring and alerting with DataDog and PagerDuty across 100+ production jobs, improving operational stability and reducing pipeline downtime by 45%.

● Partnered with ML engineers to build feature store pipelines on AWS SageMaker Feature Store and Databricks, reducing model training data preparation time by 30%. DXC TECHNOLOGY June 2021 – July 2023

Data Engineer India

● Implemented a cloud-native data lake on GCP using Cloud Storage (Parquet/ORC), BigQuery as the analytical store, and Dataflow (Apache Beam) for streaming ingestion, supporting analytics on 100B+ records.

● Designed and deployed batch ETL pipelines with Apache Airflow, PySpark, and SQL, processing 2TB+ daily from APIs, databases, and Kafka event streams into GCP BigQuery and Cloud Storage.

● Adopted Delta Lake and Apache Hudi for ACID-compliant lakehouse storage, enabling CDC patterns and upsert operations across large historical datasets.

● Optimized PySpark jobs and BigQuery SQL through partitioning, Z-ordering, broadcast joins, and predicate pushdown, cutting average job runtime by 50% and saving $30K+ annually in compute costs.

● Built dimensional data models (star schema, SCD Type 2) and dbt transformation layers to support KPI dashboards, revenue analytics, and customer behavior segmentation for 15+ business units.

● Developed real-time Pub/Sub and Dataflow streaming pipelines for event-driven ingestion, enabling near-real- time fraud detection dashboards with latency under 2 seconds.

● Partnered with data scientists to design feature stores and ML training datasets using Vertex AI Feature Store and BigQuery ML, improving model training efficiency by 35%.

● Authored Terraform modules for GCP infrastructure provisioning and established GitLab CI/CD pipelines to automate deployments across dev, staging, and production environments. EDUCATION

CLEVELAND STATE UNIVERSITY Aug 2023 – May 2025

Master of Science in Information Sciences Ohio, USA



Contact this candidate