Post Job Free
Sign in

Data Engineer Building Spark-Kafka Pipelines

Location:
Crown Point, IN, 46307
Posted:
August 31, 2026

Contact this candidate

Resume:

SAI KARTHIK PATRI

Data Engineer

+1-816-***-**** ***************@*****.*** LinkedIn GitHub

PROFESSIONAL SUMMARY

Data Engineer with 5+ years building, deploying, and operating batch and streaming pipelines on Apache Spark, Kafka, and cloud infrastructure across AWS and GCP with BigQuery. Owns the full delivery path from local development and unit testing through code review, CI/CD promotion across dev, test, and production using Pulumi, scheduled operation in Airflow and Control-M, and the monitoring and incident response that follows. Designs dimensional and lakehouse models using dbt, builds validation and reconciliation into ingestion rather than after it and tunes Snowflake for query cost and runtime. Has migrated a processing framework from AWS EMR to GCP Dataproc with Hadoop and GKE, moved on-premise relational data into Snowflake through Informatica IICS and managed regulatory and risk data at a global bank where a late or incorrect load becomes a reportable event

PROFESSIONAL EXPERIENCE

Analyte IT Services Raleigh, NC

Software Engineer Sept 2025 - Present

•Migrated the file processing framework from AWS EMR to GCP Dataproc, rebuilding Airflow DAG monitoring, validation checks, and downstream Snowflake loads so the cutover stayed invisible to the teams consuming the data.

•Architect and deliver 10+ production ETL pipelines in Python and PySpark, landing curated datasets in Snowflake that carry daily reporting for 2+ business domains.

•Develop pipelines locally against sampled data, cover transformations with unit tests, and put every change through pull request review before it moves out of the development branch.

•Engineer reusable ingestion frameworks on AWS using S3, ECS, and Lambda, cutting data refresh latency by roughly 35% while keeping the design able to scale horizontally without a matching rise in cost.

•Design reusable Informatica IICS mappings and taskflows moving data from on premise relational databases into Snowflake, with validation checks, reconciliation queries, and exception handling built into the flow rather than bolted on afterward.

•Agree column names, types, null handling, and refresh expectations with the analysts and business owners consuming each dataset before build starts, which costs less than renegotiating a schema after the first incorrect report reaches a stakeholder.

•Implemented CI/CD in GitHub Actions covering automated testing, deployment, and schema versioning.

•Automate scheduling for 40+ batch jobs in Control-M and surface SLA adherence and failure metrics in Power BI, cutting manual monitoring effort by roughly 60% and making a missed load visible before the business asks about it.

•Instrument pipelines so a failure names the step, the source, and the record count that broke, then own the rerun, the backfill, and the root cause write up rather than handing the alert to someone else.

•Architected a self-serve data platform on GCP using BigQuery, Dataplex, and Cloud Composer DAGs, enabling analysts to access governed datasets without engineering intervention across 40+ pipelines.

•Hardened data security by enforcing VPC Service Controls and data masking policies on Google Cloud Storage buckets and BigQuery tables, restricting PII exposure across all ingestion workflows.

•Accelerated cloud migration of legacy Hadoop workloads to Dataflow and Apache Beam, rewriting batch transforms in Python and deploying infrastructure as code via Terraform to cut provisioning time.

University of Missouri Kansas City Remote

Research Assistant Jul 2024 - Sept 2025

•Developed PySpark ETL pipelines on GCP ingesting and transforming high-frequency time series research data at more than 50,000 records per hour.

•Profiled and validated data across multiple datasets, reducing inconsistencies by roughly 25% and making published results reproducible from the raw inputs.

•Designed Snowflake analytical models and tuned them through clustering, caching, and micro partition pruning, cutting query runtimes by roughly 45% and making interactive analysis practical for faculty.

•Deployed and coordinated analytics jobs on AWS EMR and ECS, managing 20+ recurring pipelines through Airflow DAGs and GitHub Actions with Kubernetes-based job isolation.

•Built 3+ interactive Power BI dashboards covering research trends, system performance, and sensor metrics that fed directly into faculty decisions.

•Documented pipelines and models thoroughly enough that other researchers could rerun and extend the work without a walkthrough from the original author.

•Built BigQuery ML and Vertex AI models on top of structured research datasets, integrating Looker dashboards to surface predictions and trends to non-technical stakeholders in near real-time.

•Modeled analytical schemas in BigQuery using BigLake external tables over Google Cloud Storage, applying data warehousing best practices and dimensional data modeling to support reproducible research queries.

•Implemented a Pub/Sub messaging system to decouple ingestion from transformation stages, routing high-volume event streams into Dataflow jobs and reducing end-to-end pipeline latency across distributed datasets.

Morgan Stanley Bengaluru, India

Software Engineer (FTC) Jan 2020 - Aug 2023

•Designed event-driven pipelines on Apache Kafka, Python, Spark, and Databricks processing high-volume real-time financial data for regulatory and risk analytics.

•Led root cause analysis on recurring data issues and shipped versioned CI/CD rollbacks and fixes that brought incidents down by roughly 40%.

•Held delivery and ordering guarantees on that stream, because credit risk, transaction level, and compliance reporting all read from it and a dropped or reordered event becomes a regulatory question.

•Automated ingestion from Sybase, Oracle, and REST APIs, harmonizing schemas across 15+ heterogeneous sources into one consistent data model using dbt transformations.

•Improved pipeline observability by wiring Prometheus and Grafana into Control-M telemetry, cutting mean time to detect by roughly 30%.

•Operated inside a regulated change process where every release carried review, approval, and a documented rollback path, the discipline that keeps a bank out of trouble when a load fails at month end.

•Delivered Tableau and Power BI dashboards covering portfolio performance and risk indicators for business stakeholders.

•Extended Apache Kafka pipelines with BigQuery Omni federation and Data Readiness Placement (DRP) checks, ensuring cross-region data availability and compliance before downstream consumers could query production tables.

•Engineered data engineering utilities in Rust for high-throughput binary serialization of Sybase and Oracle payloads, cutting message-processing overhead on latency-sensitive Kafka topics in production.

•Strengthened data security posture by designing row-level data masking rules and access-scoped messaging systems policies, aligning pipeline outputs with internal governance standards across 15+ regulated data feeds.

SP Technologies Hyderabad, India

Software Developer Trainee Apr 2018 - Apr 2019

•Developed Android applications in Kotlin and Java, integrating REST APIs alongside Firebase authentication, realtime database, and cloud messaging.

•Optimized application performance by roughly 30% through profiling and memory tuning.

•Practiced Agile and TDD, holding test coverage above 90% and shortening release cycles by roughly 25%.

•Learned production support discipline early, taking defect triage and release verification alongside feature work.

•Integrated Firebase Cloud Messaging as a real-time messaging system within Android applications, modeling notification payloads and data schemas to support reliable event delivery across user devices.

•Automated build and release workflows using GitHub Actions CI/CD, applying infrastructure-as-code principles to Android project configuration and reducing manual deployment steps for production releases.

•Diagnosed and resolved recurring REST API data ingestion defects in Kotlin services, refactoring request-handling logic to enforce data validation rules and improve payload integrity before storage.

KEY PROJECTS

Web Server Log Analytics Pipeline Kafka, Faust, Snowflake, AWS, Python, Power BI, Grafana

•Built a real time ETL and analytics pipeline turning multi source web logs into OLAP insight, using Kafka and Faust for ingestion and stream processing and pushing aggregates into Snowflake.

•Modeled fact and session tables so operations teams can run ad hoc analysis without waiting on an engineer to write the query for them.

•Sized the stream processing and warehouse layers separately so ingestion spikes do not force the analytical side to scale alongside them.

•Deployed on AWS ECS and Lambda and surfaced response time trends, user patterns, and system health in Power BI and Grafana.

•Handled late arriving and malformed log events explicitly rather than dropping them, so the aggregates stay reconcilable against raw ingestion counts.

Quantum Enhanced Transformer for Financial Forecasting PySpark, Pandas, Python, Machine Learning models, Power BI, Matplotlib

•Research-A-Thon 2025.

•Implemented a feature drift detection module using distributed systems principles that triggers retraining when incoming Kafka-style incoming streams move outside expected variance, reporting gains in forecast accuracy and latency against LSTM baselines.

•Surfaced model performance and trading indicators through business intelligence dashboards in Power BI and Matplotlib analytics so results could be explained rather than only quoted.

•Versioned datasets and model runs with data governance practices so a reported number could be traced back to the exact inputs that produced it.

TECHNICAL SKILLS

•Languages: Python, SQL including CTEs and window functions, Scala, R, Java, Kotlin, Go, Rust, Shell Scripting

•Processing and Streaming: Spark SQL, Delta Lake, Snowpark, Faust, Structured Streaming, Event Driven Pipelines

•Orchestration and ELT: Control-M, Informatica IDMC and IICS, DAG design, Dependency Management, Retry and Alerting Patterns, Incremental Loads, SLA monitoring, Backfills

•Warehousing and Modeling: star and Snowflake schemas, Facts and Dimensions, Data Marts, Lakehouse and Medallion Layering, Clustering, Caching, Micro Partition Pruning, Query Tuning

•Cloud Platforms: Redshift, AWS including EMR, ECS, S3, GCP including DataProc, AWS, GCP, Emr, AWS Lambda, and Secrets Manager, GCS, GCE, IAM, and Composer

•Data Quality and Governance: Spark, Databricks, Pandas, Airflow, dbt, Snowflake, Reconciliation Queries, Exception Handling, Data Profiling, Metadata Management, Data Lineage, Access Controls, Auditing, encryption at REST and in transit, Power BI, Tableau, Oracle, Firebase, Hadoop, Validation frameworks

•DevOps and Observability: CI/CD, Jenkins, Docker, Kubernetes, Terraform, Schema Versioning, Linux administration, Apache, ELK Stack

•Analytics and BI: Matplotlib, Statistical Analysis, OLAP modeling, Ad Hoc Reporting, Agile, REST

•Tools & Platforms: Kafka, Prometheus, Grafana, TDD, Git, Pubsub

EDUCATION

University of Missouri Kansas City Kansas City, MO

Master of Science in Computer Science GPA: 3.97/4.0 Jan 2024 - May 2025

Osmania University Hyderabad, India

Bachelor of Engineering in Computer Science GPA: 7.66/10

CERTIFICATIONS AND ACHIEVEMENTS

•SRE Fundamentals with Google Certification, mentored by Salim Virji.

•Certified Production Support Analyst, Wiley Edge.

•Ultimate milestone, Google Cloud Ready Facilitator program.

•Second place, Research-A-Thon 2025, for a quantum enhanced transformer system for real time stock forecasting.

•GATE 2022 qualified candidate.



Contact this candidate