Post Job Free
Sign in

Lead ETL and Streaming Data Engineer

Location:
Sunnyvale, CA, 94089
Posted:
July 25, 2026

Contact this candidate

Resume:

Sridevi Pannerselvam

+1-669-***-**** ***********@*****.*** linkedin.com/in/sridevi1012

PROFESSIONAL SUMMARY

Lead Data Engineer with 7+ years building production data platforms for banking and payments. Deep hands-on experience with Kafka and Spark Structured Streaming (1M+ transactions/day, sub-second SLAs) and with standing up analytical platforms strictly derived from transactional systems of record — including a ~5TB, 200+ table Oracle-to-Delta Lake migration on a Medallion architecture. Strong track record in data quality, reconciliation, and schema enforcement for regulated environments, and in owning a technical domain end-to-end from loose requirements through production monitoring and operational handoff. Has delivered production systems on both Azure and GCP and led the platform evaluation between them, so tooling choices are treated as engineering tradeoffs rather than fixed assumptions.

TECHNICAL SKILLS

● Languages: Python, PySpark, SQL / Spark SQL, Scala, Shell Scripting

● Streaming & Messaging: Apache Kafka, Spark Structured Streaming, event-driven pipelines, exactly-once semantics, replay & idempotent writes

● Cloud & Lakehouse: Databricks, Delta Lake, Medallion Architecture (Bronze/Silver/Gold), Azure (ADLS, Data Factory, Synapse), GCP (BigQuery, Dataproc, Dataflow)

● Governance & Reliability: Data quality and validation frameworks, reconciliation, schema enforcement, source-to-target lineage documentation, access provisioning and approvals, CI/CD quality gating

● Storage & Databases: Oracle 11g/12c, Hive, HDFS, Elasticsearch (ELK), Redis, MongoDB

● Platforms & Tooling: Cloudera (CDH 5.8.x), Hortonworks (HDP), Git, Jenkins, SonarQube, Airflow, Autosys, Maven, Kibana, Agile/Sprint Planning

● Transferable to AWS Stack: Kafka AWS MSK (same protocol, consumer semantics, and offset/replay model); Oracle CDC and incremental extraction PostgreSQL logical replication; Delta Lake / Medallion warehouse modeling Snowflake layered architecture; ADLS S3 object-store patterns EXPERIENCE

Tata Consultancy Services Client: Wells Fargo Sunnyvale, CA, USA Lead Data Engineer

Apr 2023 - Sep 2025

Project: Credit Card Acquisition & Campaign Analytics

● Led migration of ~5TB / 200+ tables from Oracle (transactional source of truth) into a strictly derived analytical layer on Azure ADLS + Databricks Delta Lake; defined replication boundaries, incremental load patterns, partitioning, and cluster sizing, ensuring the analytical platform never fed back into transactional processing.

● Designed and implemented ETL/ELT pipelines in Databricks (Python, PySpark, Spark SQL) on a Medallion architecture, orchestrated end-to-end with Airflow DAGs — curating raw Bronze ingestion into refined Silver tables and optimized Gold models structured for downstream BI engines and analysts.

● Established governed ingestion patterns across data modeling, source system, and platform teams — securing data access approvals, documenting source-to-target lineage, and converting one-off pulls into reusable, auditable pipelines with defined ownership.

● Ran platform and cost evaluations across Delta Lake vs. BigQuery and Azure vs. GCP Databricks using Python benchmarking scripts; findings supported a ~30–40% compute cost reduction and guided cloud modernization decisions.

● Built an analytics platform measuring product adoption and campaign effectiveness across channels (email, direct mail, follow-ups); surfaced drop-offs and non-converting segments, improving campaign ROI by

~$300K.

● Integrated external credit bureau feeds (CAP, Fiserv) into Spark rule engines for eligibility and credit scoring automation.

● Led a team of data engineers across US and India — running sprint planning, coordinating cross-time-zone handoffs, and managing priorities — while remaining hands-on building complex Databricks transformations. Tata Consultancy Services Client: Bank of America Chennai, India Senior Data Engineer

Jun 2021 - Sep 2022

Project: Real Time payments

● Designed and operated the 24x7 event-driven pipeline behind US Real-Time Payments on Kafka + Spark Structured Streaming — ~1M+ transactions/day at ~100 TPS peak, meeting sub-second SLAs and cutting end-to-end latency ~90% versus the prior batch flow.

● Owned production monitoring, alerting, failure recovery, replay, and idempotent writes to maintain exactly-once behavior at the sink and support safe backfills and retries.

● Built the transactional data platform giving bank operations a 360 view of live payment status (initiated, approved, cancelled, rejected) during customer inquiries and investigations, with Airflow scheduling batch reconciliation and backfill jobs.

● Integrated Elasticsearch REST APIs via Python clients to enable fast transaction search, traceability, and downstream reporting.

● Mentored junior engineers on Kafka, Spark, and platform migration practices — running code reviews and pairing on production streaming incidents.

Tata Consultancy Services Client: Bank of America Chennai, India Data Engineer II

Jun 2019 - May 2021

Project Circuit Breaker (Fraud Detection)

● Defined data quality, validation, and reconciliation standards across streaming and batch pipelines — schema enforcement, automated reconciliation checks, and JUnit unit/integration coverage — catching upstream defects before they reached downstream consumers and saving ~8 hours of manual effort per load.

● Built near real-time fraud detection pipelines on Cloudera CDH using Spark Streaming with Redis-based enrichment; applied sliding time windows to detect velocity-based fraudulent behavior.

● Led the Cloudera platform upgrade (CDH 5.x 6.x): authored the phased rollout and dependency remediation plan, drove Spark/Hive compatibility testing, and landed the migration with zero downtime to the 24x7 payment flow.

● Owned post-migration regression validation end-to-end — schema parity, job-level output correctness, and performance baselines — confirming no silent data or latency drift before sign-off.

● Tuned Spark workloads and refactored reusable components, reducing infrastructure spend by ~$141K annually.

Tata Consultancy Services Client: Bank of America Chennai, India Data Engineer I

Jun 2018 - May 2019

Project: Trade Finance

● Built Spark/Scala batch pipelines on Cloudera CDH processing billions of records from Oracle, Hive (ORC), Parquet, CSV, JSON, and flat files, handling both read-heavy reporting and write-heavy ingestion workloads.

● Implemented a file archiving and quarantine process to isolate corrupted inputs for audit and safe reprocessing.

● Reduced HDFS small-file issues through weekly compaction and improved Hive query performance via partition pruning, bucketing, and vectorized execution. IBM Chennai, India

Software Engineer Intern

Jun 2017 - Jul 2018

● Developed an IoT-based LPG gas detection system (MQ6 sensor, Arduino Uno, ESP8266) with real-time alerting via buzzer and ThingSpeak cloud monitoring. EDUCATION

SRM Easwari Engineering College – Chennai, India

Bachelor of Engineering, Computer Science, GPA: 8.3/10 Jun 2013 - May 2017

CERTIFICATIONS

● Microsoft Azure Fundamentals (AZ-900)

● Google Cloud Fundamentals

● Confluent Kafka Fundamentals

● Databricks Lakehouse Fundamentals



Contact this candidate