SHIVAKUMAR THUPPALA
SENIOR DATA ENGINEER DATA SERVICES & ANALYTICS
Charleston, IL 216-***-**** **********.********.****@*****.*** linkedin.com/in/shivakumarthuppala
PROFESSIONAL SUMMARY
Senior Data Engineer and data-services professional with 5+ years of experience developing, troubleshooting, and optimizing enterprise data pipelines, database solutions, integrations, and reporting platforms. Advanced in SQL, T-SQL, Python, PowerShell, data modeling, query optimization, ETL/ELT, data quality, and production support across Azure, AWS, Snowflake, Databricks, and Power BI. Experienced tracing end-to-end data flows, resolving production issues, supporting migrations and releases, reviewing technical deliverables, and collaborating across engineering, analytics, finance, healthcare, and client-facing teams. Delivered 63% faster ETL, 48% lower query runtime, 57% faster Power BI refreshes, and 75% lower pipeline MTTR.
TECHNICAL SKILLS
Database & Integration: SQL, T-SQL, query optimization, stored procedures, relational databases, ETL/ELT, data mapping, source-to-target mapping, API integration, migrations, incremental loads
Analytics & Processing: Power BI, Python, PowerShell, Bash, PySpark, Apache Spark, Spark SQL, dimensional modeling, star/snowflake schemas, data marts
Operations & Governance: Production support, root-cause analysis, validation, monitoring, alerting, SLA management, regression testing, incident resolution, RBAC, masking, audit logging, metadata management
Cloud & DevOps: Azure Databricks, ADF, Synapse, ADLS Gen2, AWS S3, Glue, Lambda, Snowflake, Delta Lake, Airflow, Git, CI/CD, Azure DevOps, Terraform
PROFESSIONAL EXPERIENCE
Senior Data Engineer Stripe
San Francisco, CA
Jul 2024 - Present
• Designed and optimized a 10+ TB FP&A migration from Teradata to Azure using ADF, Databricks, PySpark, SQL, Delta Lake, and Bronze/Silver/Gold processing layers.
• Re-engineered complex SQL logic and stored procedures into optimized PySpark pipelines, improving transformation performance by 65%.
• Investigated production data issues by tracing source-to-target transformations, pipeline dependencies, validation failures, and downstream reporting impacts.
• Reviewed SQL transformations, pipeline configurations, data-quality rules, and deployment changes for accuracy, performance, governance compliance, and production readiness.
• Supported release validation, regression testing, monitoring, and post-deployment troubleshooting; SLA-based alerting reduced pipeline MTTR by 75%.
• Reduced Databricks compute costs by 35% through workload profiling, partition optimization, cluster right-sizing, autoscaling, and incremental processing.
Data Engineer Innovaccer
Hyderabad, India
Mar 2022 - May 2023
• Built and operated 15+ AWS Glue, Lambda, and S3 pipelines processing 300M+ healthcare records daily across clinical, claims, and operational domains.
• Partnered with business, engineering, analytics, and client-facing teams to gather requirements, document mappings, prioritize defects, and deliver reliable data solutions.
• Migrated Oracle workflows to AWS S3 and Snowflake, reducing end-to-end ingestion time from 8 hours to 90 minutes.
• Optimized Snowflake models and SQL workloads through query tuning and partition-aware processing, reducing runtime by 48% and saving approximately $12K per month.
• Implemented automated data-quality checks, monitoring, RBAC, masking, and audit logging, sustaining 99.9% pipeline availability in HIPAA and SOC 2-aligned operations.
Data Engineer EonSpace Labs
Hyderabad, India
Mar 2020 - Feb 2022
• Engineered 20+ ADF, Databricks, and Synapse pipelines processing 500M+ supply-chain records daily from APIs, JSON, CSV, and relational sources.
• Reduced ETL runtime by 63% through Spark tuning, partitioning, caching, incremental loads, and optimized SQL transformations.
• Developed a reusable ingestion framework for source-to-target mapping, schema validation, metadata capture, error handling, and standardized Delta Lake processing.
• Developed and optimized SQL and T-SQL transformations, dimensional models, and reporting datasets supporting Power BI dashboards and operational analysis.
• Improved Power BI refresh performance by 57% and reduced Azure spend by 25% through workload optimization and improved serving-layer design.
EDUCATION
Master of Science in Computer Systems Technology Eastern Illinois University
Charleston, IL GPA: 3.5
May 2025