Post Job Free
Sign in

Data Engineer - AWS Lakehouse & ETL Expert

Location:
Bengaluru, Karnataka, India
Posted:
July 24, 2026

Contact this candidate

Resume:

Vignesh Mahalingam

Bangalore ***************@*****.*** +91-638******* github.com/Vicky290420/ecommerce-lakehouse www.linkedin.com/in/vignesh-m-b32b1817b

SUMMARY

Data Engineer with 4+ years of experience building and optimizing production-grade ETL/ELT pipelines in healthcare analytics at IQVIA. Hands-on expertise across Python, PySpark, SQL/PL-SQL, AWS (S3, Glue, EMR, Athena), Snowflake, dbt, and Apache Airflow. Delivered Medallion Lakehouse pipelines at 200M+ records/day scale with 99.9% SLA. AWS Certified Data Engineer – Associate and SnowPro Core certified.

TECHNICAL SKILLS

Languages: Python, SQL, PL/SQL, Bash

AWS: S3, Glue, EMR, Athena, IAM, STS, Step Functions, GitHub Actions

Big Data & ETL: Apache Spark (PySpark), AWS Glue ETL, EMR, Parquet, Delta Lake, Medallion Architecture (Bronze/Silver/Gold)

Data Warehouse: Snowflake (Snowpipe, External Stage, Star Schema, IAM Integration)

Data Modeling: dbt Core, Staging/Mart layers, Surrogate keys, dbt tests, dbt_utils

Orchestration: Apache Airflow (FileSensor, DAGs, retry logic, XCom, Docker)

Performance: Partitioning, caching, query tuning, columnar storage (Parquet), incremental ingestion

DevOps & Tools: Git, Docker, CI/CD (GitHub Actions), Linux

Familiar (exploratory): Amazon Redshift, Amazon Kinesis

EXPERIENCE

Data Engineer Mar 2022 – Apr 2026

IQVIA — Bangalore, India

•Built and maintained PySpark-on-EMR and AWS Glue ETL pipelines ingesting 200M+ healthcare records daily from S3-landed flat files and Oracle sources into a Medallion Lakehouse (Bronze Silver Gold), maintaining 99.9% SLA adherence across all analytical and reporting workloads

•Optimized 20+ complex SQL queries across Snowflake and Oracle PL-SQL for reporting and reconciliation pipelines — applying partitioning, caching, and query-plan analysis — reducing average execution time by ~45% and eliminating recurring weekly SLA breaches

•Built Python-based data validation and reconciliation framework from scratch, enforcing schema, nulls, uniqueness, and referential integrity across pipelines using dbt tests and custom DQ checks — cutting data discrepancy incidents by 60% and reducing manual intervention by 8+ hours/week

•Designed data ingestion frameworks processing structured and unstructured datasets from S3, routing failed records to Dead Letter Queues (DLQ) for safe, idempotent reprocessing — achieving 99.5% pipeline reliability

•Led comprehensive migration of legacy Oracle RDBMS code into dbt, transforming 2 applications with 30+ stored procedures into maintainable, version-controlled SQL/dbt models (3–5 fact/dimension tables per app, multiple complex transformations). Achieved 45% performance improvement through columnar Parquet storage and query optimization. Owned one full migration independently with minimal guidance; collaborated on second. Enabled non-technical stakeholders to understand and modify transformation logic, reducing dependency on data engineers.

•Resolved critical schema drift incident (age column naming inconsistency cascading to 8 downstream models) through structured root-cause analysis and dbt data validation framework (dbt_expectations.expect_column_to_exist). Led cross-functional initiative with Analytics, BI, and Client teams to establish enterprise data governance standards: column naming conventions, mandatory change control process (structured sub-project for additions/removals), and end-to-end testing pipeline. Prevented ~12 subsequent data quality incidents, reducing silent failures by 70%.

•Developed deep expertise in Snowflake: advanced performance tuning (micro-partition clustering, materialized views, result caching), storage optimization and cost control through compression strategies and lifecycle policies. Query optimization reducing average runtime by 45% across 100+ reports. Experienced in Snowpipe automation, Stream/Task orchestration, and warehouse right-sizing strategies reducing infrastructure costs by 30% ($80k+/year).

•Architected and maintained dbt transformation framework across 30+ models spanning Bronze/Silver/Gold medallion layers. Expert in incremental model patterns with lookback windows for late-arriving data. Implemented comprehensive dbt testing framework (schema validation, data quality checks, referential integrity) preventing 70% of downstream data quality incidents. Optimized DAG execution through strategic model materialization choices (views for staging, tables for marts, incremental for large fact tables).

•Established knowledge base documenting 20+ critical incident resolutions, debugging methodologies, and architectural patterns (incremental model design, Snowpipe failure modes, query optimization strategies). Created runbooks enabling junior engineers and analysts to self-serve troubleshooting. Reduced incident escalation rate by 40% and ramp time for new team members by 60%.

•Architected high-throughput serving layers in Snowflake to power real-time corporate KPI dashboards, optimizing semantic-layer query latency to support downstream business concurrency requirements

•Automated deployment workflows using Git and GitHub Actions CI/CD pipelines, reducing release cycle time from 3 days to same-day deployments and eliminating manual handoff errors

•Received Spotlight Award for automating a manual data analysis process end-to-end, saving ~12 hours/week of analyst effort

PROJECTS

E-Commerce Data Lakehouse — Python, AWS S3, Glue, EMR, PySpark, Athena, Snowflake, dbt, Airflow, Docker GitHub

•Architected end-to-end Medallion pipeline: raw CSVs land in S3 Bronze PySpark Glue ETL transforms to Parquet Silver dbt models build Fact/Dim tables (fact_orders, dim_customers, dim_products) in Snowflake Gold

•Implemented Batch-ID partitioning for idempotency (batch_id=YYYY-MM-DD_HH-MM folders), enabling safe reprocessing of any historical batch without data corruption or duplication

•Secured S3-to-Snowflake integration via IAM Role assumption with External ID — zero long-lived credentials stored anywhere in the pipeline

•Built 16 automated dbt data quality tests covering uniqueness, nulls, and referential integrity across all mart models — all passing in CI

•Orchestrated full pipeline using Airflow FileSensor Glue trigger dbt build, with task-level exponential backoff retry logic and failure notifications on every task

•Integrated AWS Athena for ad-hoc querying over the S3 Silver layer, enabling self-serve analytics without loading data into Snowflake

Scalable Big Data Pipeline — AWS S3, PySpark, EMR, Snowflake, Airflow

•Designed and built batch processing pipelines using PySpark on EMR, processing 50GB+/day with Parquet columnar storage — improving downstream query performance by ~55%

•Built incremental ingestion logic and DLQ routing for failed records, achieving 99.5% pipeline reliability across ingestion flows

•Integrated Snowpipe for continuous low-latency data loading into Snowflake with Star Schema models supporting self-serve analytics

EDUCATION

Bachelor of Engineering – Computer Science and Engineering Graduated 2021

Mahendra College of Engineering, Namakkal

CERTIFICATIONS

•AWS Certified Data Engineer – Associate

•Unix Essential Training



Contact this candidate