Post Job Free
Sign in

Senior Cloud Data Engineer

Location:
San Francisco, CA, 94114
Posted:
October 10, 2026

Contact this candidate

Resume:

Jaswanth Sirigiri

Overland Park, Kansas +1-361-***-**** *****.*****@*****.*** LinkedIn GitHub Website PROFESSIONAL SUMMARY

Data Engineer with 5+ years of experience building cloud-native ETL/ELT pipelines and analytical data platforms across financial services and healthcare. Skilled in Python, Azure Data Factory, Databricks, Snowflake, Git, and Azure DevOps, with expertise in metadata-driven pipelines, Delta Lake, CI/CD automation, data quality, and governance. Reduced financial data refresh latency by ~40% and pipeline execution time and cloud infrastructure costs by ~30%. TECHNICAL SKILLS

Cloud Platforms: Microsoft Azure (Data Factory, Databricks, Synapse Analytics, Purview, DevOps, Monitor, Logic Apps, Data Lake Storage Gen2, Blob Storage), AWS (S3, Glue, Redshift, Lambda), GCP BigQuery, Microsoft Fabric

Data Engineering & Streaming: Apache Spark, PySpark, Databricks, Delta Lake, medallion architecture, Apache Flink, Spark performance tuning (partition pruning, caching), pipeline orchestration Data Integration & ETL/ELT: Azure Data Factory, metadata-driven pipelines, change-data capture

(CDC), incremental loading, event-driven pipelines, Talend, SSIS migration, REST API Data Warehousing & Modeling: Snowflake, Azure Synapse Analytics, dimensional modeling, star schema, partitioning, clustering, query optimization Databases & Formats: SQL Server, PostgreSQL, MongoDB, Cosmos DB, Cassandra, ScyllaDB, Parquet, Avro, JSON

Data Quality, Observability & Governance: Data validation frameworks, data lineage, Azure Purview, anomaly detection, data observability, HIPAA and SOX compliance, data masking, RBAC, audit logging, SLA/SLO monitoring

Programming: Python, SQL, PySpark, Shell scripting DevOps: Git, Azure DevOps, Docker, CI/CD automation, automated testing Visualization: Power BI, Tableau

AI/LLM Integration: Anthropic Claude API, LLM-assisted root-cause analysis, AI-assisted data quality triage

Also familiar with: Apache Airflow, Apache Kafka, Informatica, Kubernetes Industry Knowledge: Financial Services Data

PROFESSIONAL EXPERIENCE

Charles Schwab Data Engineer Jan 2025 - Present

• Built scalable ETL/ELT pipelines across Azure Data Factory, Databricks (PySpark), and Microsoft Fabric for trading, portfolio, and client account data; implemented Delta Lake incremental loading with change-data capture (CDC), reducing data refresh latency by ~40% for client account analytics.

• Designed metadata-driven, parameterized ADF pipeline templates and reusable PySpark modules adopted by multiple downstream teams, cutting new-pipeline development time and standardizing error handling, retry logic, and logging.

• Optimized Snowflake loading and transformation workflows from heterogeneous source systems using partitioning, clustering, and query tuning, improving query performance and reducing compute costs on large-scale financial datasets.

• Developed Python-based data validation frameworks covering completeness, accuracy, and consistency checks for financial and regulatory reporting; implemented automated data lineage in Azure Purview for audit and troubleshooting.

• Partnered with risk and compliance stakeholders to translate SOX control requirements into automated pipeline checks, reducing manual audit preparation and tightening turnaround on quarterly reporting cycles.

• Established CI/CD pipelines in Azure DevOps across dev, staging, and production; maintained Git-based version control for pipeline definitions and code, and configured Azure Monitor and Log Analytics for data observability, proactive alerting, and SLA/SLO tracking on critical financial reporting pipelines.

• Authored technical runbooks and onboarding documentation for the data platform and mentored two incoming engineers on Delta Lake and Azure Data Factory best practices. UnitedHealth Group Data Engineer Jul 2020 - Dec 2023

• Designed end-to-end ETL pipelines in Azure Data Factory integrating claims systems (Epic EMR, Facets), member enrollment databases, and provider networks; implemented watermark- and CDC-based incremental loading, reducing pipeline execution time and cloud infrastructure costs by

~30%.

• Developed scalable PySpark ETL jobs on Databricks feeding Snowflake and GCP BigQuery for analytics and batch reporting; improved throughput through Spark parameter tuning, caching, and partition pruning.

• Built real-time streaming pipelines with Apache Flink, ScyllaDB, and Cassandra for low-latency access to claims and member data, reducing reporting lag from hours to minutes.

• Built Python data quality frameworks for claims adjudication workflows; enforced HIPAA compliance through data masking, role-based access controls, and PHI retention and access policies defined with governance and security teams, achieving zero compliance violations across all audit cycles.

• Migrated legacy SSIS packages to Azure Data Factory and Talend, reducing maintenance overhead; added automated failure alerts via Azure Logic Apps and Python unit and integration tests across all pipeline deployments.

• Designed star-schema dimensional models for claims, member, and provider datasets, enabling self-service Power BI and Tableau reporting and cutting ad-hoc reporting turnaround.

• Led root-cause investigations for production pipeline failures, working with upstream source-system owners (Epic EMR, Facets) to resolve schema drift and data quality issues before they affected downstream reporting SLAs.

EDUCATION

Texas A&M University – Corpus Christi Jan 2024 - Dec 2025 Master of Science, Computer Science (GPA: 3.5 / 4.0) Corpus Christi, TX PROJECTS

FinOps & Data Observability Lakehouse for Healthcare Claims

• Architected a medallion lakehouse data ingestion pipeline with Bronze, Silver, and Gold layers, processing 25,600 synthetic healthcare claims records with schema drift reconciliation and CDC-style deduplication via Delta MERGE INTO; reduced 600 duplicates to a clean 25,000-row dataset with zero data loss through quarantine-based error handling.

• Built a PySpark data quality framework with eight automated checks for null-rate thresholds, referential uniqueness, and categorical validation, along with rolling z-score anomaly detection for daily claim volume; validated the framework by catching a simulated pipeline failure at z = -9.5 against a 2.5 threshold.

• Implemented job-level cost and runtime monitoring across pipeline stages and developed a live Power BI dashboard, connected to Databricks through DirectQuery, to track data quality pass rates, per-stage compute cost, and anomaly timelines.

• Integrated the Anthropic Claude API to generate plain-English root-cause explanations for data quality failures, turning raw check failures into actionable triage steps delivered through automated Slack alerts.

CERTIFICATIONS

• AWS Certified Cloud Practitioner: Amazon Web

Services

• Data Analytics Essentials: Cisco Networking

Academy



Contact this candidate