Post Job Free
Sign in

Senior Big Data Engineering Leader

Location:
St. Louis, MO
Posted:
September 23, 2026

Contact this candidate

Resume:

SOWMITH NAGABHIRU

SENIOR BIG DATA ENGINEER

Wentzville, Missouri, United States **********@*****.*** 314-***-**** linkedin.com/in/sowmith-nagabhiru-7a9a6a261 PROFESSIONAL SUMMARY

Data Engineering Solutions Lead and technical lead with 10+ years of experience architecting, implementing, and operating large-scale cloud data platforms across AWS, Azure, and GCP in financial services, healthcare, retail, energy, and federal environments. Deep hands-on expertise in Databricks Lakehouse, Delta Lake, Bronze/Silver/Gold Medallion Architecture, Apache Spark, PySpark, Scala, SQL, Hadoop, Kafka, Apache Airflow, NiFi, Snowflake, AWS S3/EMR/EC2/Lambda/Glue/RDS, and Azure Data Lake. Proven delivery of scalable ETL/ELT, batch and real-time streaming pipelines, data quality controls, lineage, governance, security, performance tuning, and analytics-ready data models. Experienced leading technical design reviews, defining engineering standards, mentoring engineers, collaborating with data scientists, and communicating data solutions with technical and business stakeholders.

CORE SKILLS

• Lakehouse & Architecture: Databricks, Delta Lake, Lakehouse Architecture, Medallion Architecture (Bronze/Silver/Gold), Data Lakes, Data Warehousing, Data Vault, Kimball, Star/Snowflake Schema, 3NF, OLTP/OLAP

• Big Data & Streaming: Apache Spark, PySpark, Spark SQL, Spark Structured Streaming, Scala, Hadoop, HDFS, Hive, Kafka, MapReduce, YARN, HBase, Cassandra, Oozie, NiFi

• Cloud: AWS: S3, EMR, EC2, Lambda, Glue, RDS/Aurora, Redshift, CloudWatch; Azure: Azure Databricks, ADLS Gen2, Azure Data Factory, Synapse, HDInsight, Blob Storage, Microsoft Fabric; GCP

• ETL/ELT & Orchestration: Apache Airflow, Azure Data Factory, NiFi, Informatica, Matillion, Talend, Oozie, dbt; batch and streaming workflow automation

• Programming & Databases: Python, PySpark, Scala, Java, SQL, PL/SQL, R, Bash; PostgreSQL/Aurora, Oracle, SQL Server, Teradata, DB2, MySQL, Snowflake, Redshift, MongoDB, Cassandra

• Governance, Security & Quality: Data quality, validation, reconciliation, lineage, metadata, RBAC, IAM, encryption at rest/in transit, DLP, Azure Purview, HIPAA, SOC2, GDPR, CJIS

• Leadership & Delivery: Technical leadership, architecture/design reviews, engineering standards, mentoring, stakeholder alignment, executive-facing reporting, Agile/Jira, CI/CD, Git, GitLab, GitHub Actions, Azure DevOps, Jenkins, Terraform

• Analytics & ML: Power BI, Tableau, Spark MLlib, scikit-learn, MLflow/Kubeflow exposure, ML model integration and production collaboration PROFESSIONAL EXPERIENCE

Pilvi Systems Inc Big Data Developer (Client: Mastercard) Jul 2026 - Present

• Develop Spark applications in Scala and PySpark for large-scale financial transaction processing, including CSV-to-Parquet conversion, metadata/index validation, reconciliation, and high-volume SQL transformations across datasets exceeding 4 trillion records annually.

• Build and support Kafka -> NiFi -> Spark ingestion pipelines for production billing events, applying Spark Structured Streaming concepts, checkpointing, retry/failure handling, and operational monitoring for resilient real-time processing.

• Design NiFi pipelines to ingest ZIP/tar archives from external transfer systems, validate file integrity, route data by feeder and processing date, and trigger Spark jobs through InvokeHTTP/Livy for automated ETL execution.

• Support migration from Oracle Exadata to distributed object storage using Apache Ozone and S3-compatible storage on Cloudera Data Platform; integrate Splunk logging for success, failure, retry, and exception observability.

• Perform pipeline testing, code reviews, schema troubleshooting, file-count reconciliation, and production root-cause analysis to improve data quality, reliability, and throughput.

Compass Health Network Data Reporting & Solutions Engineer Aug 2025 - Jun 2026

• Partnered with department and organizational leaders to design governed data solutions for operational, analytical, and executive reporting needs in a healthcare environment.

• Developed, maintained, and optimized complex SQL queries, stored procedures, functions, indexes, triggers, Python/R scripts, DAX, and Power Query across multiple databases and data warehouses.

• Defined reporting data models and structures, implemented quality controls, validated data accuracy and integrity, and followed enterprise data governance policies and procedures.

• Built Power BI dashboards, performance metrics, and recurring management reporting; analyzed complex datasets to identify trends and improve decision support.

Pilvi Systems Inc / Hicks Professional Group Data Engineering Lead Developer (Client: U.S. Department of Veterans Affairs) Sep 2023 - Jul 2025

• Led evaluation, architecture prototyping, Technical Design Document authoring, and cross-team design reviews for scalable data platform initiatives; mentored engineers and promoted cloud data engineering and MDM best practices.

• Architected cloud-native data solutions using AWS S3, EMR, Lambda, Glue, EC2, Snowflake, Spark, Hadoop/HDFS, Kafka, Cassandra, and Airflow to support high-throughput batch and real-time analytics.

• Developed Spark/Scala, PySpark, Spark SQL, and Spark Streaming pipelines; consumed Kafka events, processed structured/semi-structured data, automated workflows with Airflow, and tuned distributed processing for performance and reliability.

• Designed scalable data lake and warehouse patterns using AWS S3, Snowflake, Redshift, star/3NF schemas, partitioning, and clustering; implemented SQL/PLSQL migration, reconciliation, validation, and performance optimization.

• Implemented governance and security controls including IAM, RBAC, encryption at rest/in transit, Azure Purview, DLP, audit readiness, and compliance practices supporting VA, HIPAA, SOC2, and GDPR requirements.

• Established CI/CD using GitHub Actions and Azure DevOps for Spark, Python, and containerized services; implemented Grafana/Prometheus observability and alerting standards for distributed data pipelines.

• Collaborated with data scientists to prototype and integrate ML workflows using Spark MLlib and MLflow/Kubeflow concepts; built PySpark feature pipelines and graph-based identity-resolution structures using Neo4j and Snowflake.

• Led Talend and NiFi pipeline development integrating multiple VA systems into centralized AWS data lake patterns and developed governed analytics workflows using Palantir Foundry and TIBCO EBX/DV. Shrive Solutions LLC Data Engineer Jan 2023 - Aug 2023

• Implemented Bronze/Silver/Gold Medallion Architecture on Databricks Delta Lake and built an end-to-end Lakehouse supporting ACID-compliant batch and streaming workloads.

• Designed scalable ETL/ELT pipelines integrating diverse sources into Snowflake, AWS S3, Redshift, and Aurora PostgreSQL; developed PySpark/SQL transformations for cleansing, enrichment, aggregation, and analytics-ready datasets.

• Automated orchestration with Apache Airflow and Azure Data Factory; implemented CI/CD with Azure DevOps and GitHub Actions for Databricks notebooks and PySpark pipelines across dev/test/prod.

• Optimized Databricks clusters using autoscaling, spot instances, and Photon, reducing processing costs by 30% while supporting streaming workloads ingesting 10M+ daily events.

• Established data classification, role-based access, audit logging, and quality controls; partnered with business stakeholders to define requirements and deliver governed datasets and Power BI reporting.

Wipro Technologies Senior Big Data Engineer (Client: Credit Suisse) Apr 2019 - Dec 2020

• Designed and optimized batch and streaming pipelines using Spark, Kafka, Python, SQL, Airflow, Snowflake, Aurora PostgreSQL, and cloud data warehouses; implemented ETL/ELT quality checks, lineage, monitoring, and performance tuning.

• Developed scalable Python/Spark data processing and high-performance ELT workflows; tuned SQL and Spark jobs, integrated Kafka event streams with ClickHouse/Snowflake, and standardized transformations using dbt.

• Contributed to Cloud Center of Excellence governance, DLP, privacy, retention, Terraform/IaC, CI/CD, code reviews, mentoring, and cross-functional stakeholder alignment in a regulated financial-services environment. Infosys Technologies Big Data Engineer (Client: Macy's) Dec 2017 - Jan 2019

• Migrated enterprise pipelines to Azure Databricks using Spark SQL and Scala; designed Azure Data Factory pipelines and Databricks/Delta Lake ELT workflows for large-scale retail data.

• Designed Lakehouse and data-lake patterns using Azure Data Lake Storage Gen2, Snowflake, Delta Lake, Spark Structured Streaming, Kafka, and Parquet; delivered curated data for downstream Power BI analytics.

• Implemented data quality validation, schema mapping, lineage tracing, encryption, Azure DevOps/GitLab CI/CD, Grafana/Prometheus/OpenTelemetry monitoring, and performance optimization.

Tata Consultancy Services Big Data Engineer (Client: British Petroleum) Oct 2016 - Nov 2017

• Built Oozie/Sqoop/Hive pipelines and real-time Kafka ingestion frameworks on Hadoop; used AWS EMR on EC2, S3, RDS, and AWS Data Pipeline for scalable cloud processing.

• Optimized Hive with partitioning/bucketing, managed Kafka/Zookeeper and Hadoop clusters, and developed Python-based data processing and automation. Maveric Systems Data Engineer Apr 2014 - Aug 2016

• Built AWS EMR/Hadoop data warehouse and ingestion workflows using Hive, HDFS, Pig, Sqoop, Oozie, and Tableau; implemented data quality validation and supported analytics teams and data scientists.

Nakshatra IT Solutions Data Analyst Sep 2013 - Mar 2014

• Gathered business requirements; defined ETL standards and naming conventions; documented Stage/ODS/Mart flows; developed Python/SQL Server and Informatica ETL processes and performed unit, integration, and system testing. EDUCATION

University of Illinois at Springfield Master of Science (MS), Computer Science Jan 2021 - Dec 2022 GPA: 3.4 CERTIFICATIONS

• Google Cloud Professional Data Engineer - Google Cloud, 2024

• AWS Certified Data Analytics - Specialty - Amazon Web Services, 2023

• Microsoft Certified: Azure Data Engineer Associate - Microsoft, 2023

• IBM Certified Data Engineer - Big Data - IBM, 2022



Contact this candidate