Sri Charan Borra
Data Engineer
Frisco, TX 703-***-**** ***************@*****.*** LinkedIn
PROFESSIONAL SUMMARY
Senior Data Engineer with 6+ years of experience architecting and optimizing petabyte-scale data platforms, cloud lakehouses, and real-time streaming architectures across top-tier technology and financial institutions (Capital One, eBay, Amazon). Expert in designing low-latency streaming pipelines (Kafka, Kinesis, Flink) processing over 1TB of daily data and migrating complex legacy ETL infrastructures to high-performance, cost-effective AWS and GCP cloud solutions. Demonstrated track record of scaling production machine learning pipelines, implementing robust DataOps frameworks (Airflow, dbt, Terraform), and tuning Apache Spark/Delta Lake ecosystems to slash infrastructure costs by up to 35% and boost performance. TECHNICAL SKILLS
Programming Languages: Python, SQL, Bash/Shell Scripting Data Processing & Distributed Computing: Apache Spark (PySpark, Spark SQL, Structured Streaming), Databricks, Delta Lake, Apache Iceberg, Hadoop (HDFS, Hive)
Streaming & Real-Time Analytics: Apache Kafka, AWS Kinesis, Apache Flink, Spark Streaming, Event-Driven Architectures Cloud Platforms: AWS (S3, Glue, EMR, Redshift, Lambda, Step Functions, Lake Formation, MWAA, EKS, DynamoDB, Fargate), GCP
(BigQuery, Cloud Storage, Pub/Sub, Dataproc, Cloud Composer), Azure (Data Factory, Databricks, Synapse) Machine Learning & MLOps: AWS SageMaker (Feature Store), Feature Engineering Pipelines, Automated Model Deployment Data Warehousing & Orchestration: Snowflake, BigQuery, Amazon Redshift, dbt, Apache Airflow, Data Lakehouse Architecture, Dimensional Modeling, ETL/ELT Pipelines
DevOps, CI/CD & Infrastructure: Terraform, Docker, Kubernetes (EKS), GitHub Actions Certification: AWS Certified Solutions Architect – Associate PROFESSIONAL EXPERIENCE
Capital One Senior Data Engineer Plano, Texas January 2024 – Present
• Migrated 30+ legacy ETL pipelines to AWS serverless (Glue, Lambda, Step Functions) with Snowflake as analytics warehouse, reducing latency by 29% while maintaining 99.5% SLA for fraud analytics.
• Engineered real-time streaming pipelines using Amazon MSK (Kafka), Kinesis Data Streams, and Spark Structured Streaming
(PySpark), processing 1TB+ daily data with sub-5-second fraud signal latency supporting 10M+ active customers.
• Built end-to-end ML pipelines using SageMaker Pipelines, SageMaker Feature Store, and AWS Lambda, automating model training, validation, and deployment for fraud detection use cases while reducing release cycles from 12 days to 4 days.
• Optimized data lakehouse performance using Delta Lake, Apache Spark tuning, Z-Ordering, and EC2 Spot Instances, reducing query latency by 35% and saving $90K+ annually in cloud infrastructure costs.
• Modernized enterprise DataOps and CI/CD framework using dbt, Apache Airflow (MWAA), GitHub Actions, Terraform, Docker, and Amazon EKS, increasing deployment frequency and reducing MTTR by 40%. eBay Data Engineer San Jose, California April 2022 – December 2023
• Architected petabyte-scale data lake and cloud warehouse solutions on GCP using BigQuery and Google Cloud Storage (GCS), improving analytical query performance by 20% while reducing compute costs by 15%.
• Designed scalable streaming pipelines using Apache Spark and Kafka processing 200GB+ daily data, with Snowflake as central warehouse for reporting and ML feature stores.
• Orchestrated 10+ production workflows using Cloud Composer (Apache Airflow) with SLA monitoring, dependency management, and automated retries, achieving 99.2% pipeline reliability.
• Deployed scalable Dataproc Spark clusters and Google Dataflow pipelines handling 10,000+ TPS for real-time recommendation systems and anomaly detection workloads.
• Collaborated with machine learning and analytics teams to build feature engineering pipelines supporting recommendation and pricing models, improving data freshness from hourly to near real-time. Amazon Ads Data Engineer Seattle, Washington September 2020 – March 2022
• Designed and optimized Apache Spark (PySpark) ETL jobs on AWS EMR to process 50M+ daily ad impression and click events, reducing job execution time by 45% and enabling faster attribution reporting for enterprise advertisers.
• Engineered a low-latency streaming pipeline using AWS Kinesis and Apache Flink to ingest and enrich clickstream data, achieving sub-2-minute processing lag for campaign performance dashboards.
• Refactored legacy JSON-based event storage to Parquet with intelligent partitioning and compression, cutting S3 scan costs by 35% and improving Athena/Redshift Spectrum query performance by 60%.
• Built and maintained 30+ Airflow DAGs to manage cross-team data dependencies, backfills, and SLA-sensitive reporting, reducing pipeline failure rate from 8% to 1.5% over six months.
• Created dimensional data marts in Amazon Redshift powering Looker dashboards for product and sales teams, democratizing access to campaign metrics and reducing ad-hoc SQL requests by 50%. HDFC Bank Data Engineer India January 2018 – December 2018
• Designed and maintained enterprise ETL pipelines using Informatica PowerCenter, Oracle PL/SQL, and Teradata, processing 2TB+ daily banking transaction data for RBI and NPCI regulatory reporting.
• Optimized large-scale Hadoop ecosystem workflows using HDFS, Hive, and Impala for Customer 360 and risk analytics, improving query performance by 40%.
• Automated data validation and reconciliation frameworks using Python, Pandas, PySpark, and Shell scripting, reducing manual reconciliation efforts by 60%.
• Built real-time fraud detection pipelines using Apache Kafka and Spark Streaming, enabling anomaly alerting with under 10- second latency across digital payment systems.
EDUCATION
Master’s in Data Analytics Engineering George Mason University, George Mason, VA, USA. Bachelor’s in Computer Science and Engineering Jawaharlal Nehru Technological University, Hyderabad, TG, India.