VARUN R. BALAKRISHNAN
Senior Data Engineer Databricks, Snowflake & dbt AWS/GCP PySpark & Python
913-***-**** My LinkedIn Bothell, WA 98011
Summary Of Qualifications:
oSnowPro Core-certified Senior Data Engineer with strong experience designing and modernizing scalable cloud data platforms, Databricks Lakehouses, Snowflake warehouses, ELT/ETL pipelines, streaming solutions, and analytics-ready data products.
oStrong hands-on expertise in Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Snowflake, dbt, Python, advanced SQL, Apache Airflow, Kafka, AWS, GCP, and BigQuery.
oProven multi-cloud data-engineering experience across AWS and GCP, including Amazon S3, AWS Glue, Lambda, CloudWatch, IAM, Google Cloud Storage, Dataproc, Pub/Sub, Cloud Composer, Cloud Monitoring, and Cloud Logging.
oComprehensive Lakehouse experience encompassing Medallion Architecture, Bronze/Silver/Gold layers, Auto Loader, Structured Streaming, Databricks Workflows, Unity Catalog, Delta MERGE, schema evolution, ACID transactions, incremental processing, and performance optimization.
oAdvanced Snowflake and dbt expertise covering Snowpipe, external stages, Streams and Tasks, Dynamic Tables, virtual-warehouse optimization, secure views, RBAC, governed metrics, and approximately 100 modular dbt models with automated testing, documentation, and lineage.
oExtensive experience developing batch, event-driven, and near-real-time pipelines using Apache Airflow, Databricks Workflows, Kafka, Kafka Connect, AWS Glue, Google Cloud Composer, Informatica, and Pub/Sub—including approximately 30 production DAGs.
oStrong data-modeling, quality, observability, and governance background involving Kimball dimensional modeling, SCD patterns, conformed dimensions, data reconciliation, schema-drift detection, freshness monitoring, masking, hashing, lineage, Unity Catalog, and least-privilege security.
oPractical experience supporting AI-enabled analytics through chatbot-interaction pipelines, model-evaluation datasets, Snowflake Cortex AI classification, semantic reporting, and human-governed data-enrichment workflows.
oDemonstrated ability to deliver measurable technical and business outcomes, including reducing quarterly Snowflake costs by approximately 35–45%, doubling restaurant-onboarding capacity, enabling near-real-time processing of up to one million daily events, and migrating approximately 60–70 enterprise tables.
oEffective collaborator with product, analytics, AI, finance, operations, and business stakeholders, translating complex requirements into secure, scalable, well-tested, documented, and maintainable data solutions.
Certification & Education:
oSnowflake SnowPro Core Certification – verification link
oMaster of Science in Business Analytics (2024 Graduate) Arizona State University, Tempe, AZ
oBachelor of Technology in Information Technology (2018 Graduate) Thiagarajar College of Engineering, India
Professional Experience:
Client: Alervio – Seattle, WA April 2025 – Present
Role: Senior Data Engineer
oArchitected and supported a multi-source Snowflake, Databricks, and dbt data platform processing approximately 500,000 records daily through 30 Apache Airflow DAGs and 100 dbt models.
oIntegrated Toast POS, Square POS, menu, ingredient, allergen, supplier, transaction, user-event, and partner data into governed raw, standardized, dimensional, semantic, and analytics-ready layers.
oDesigned scalable ingestion frameworks using Fivetran, Snowpipe, external stages, REST APIs, and reusable Python utilities to process structured and semi-structured data from PostgreSQL, Google Cloud Storage, partner applications, and restaurant systems.
oLeveraged Databricks, Apache Spark, and PySpark to process complex restaurant, menu, ingredient, and user-event data requiring distributed transformations, large-scale backfills, and semi-structured JSON normalization.
oImplemented a Delta Lake Medallion Architecture with Bronze, Silver, and Gold layers to preserve raw source data, standardize business entities, apply quality controls, and publish curated datasets for downstream Snowflake and analytics workloads.
oDeveloped incremental ingestion patterns using Databricks Auto Loader and schema-evolution controls to process newly arriving JSON and Parquet files from Google Cloud Storage while identifying malformed records and unexpected structural changes.
oBuilt PySpark and Spark SQL transformations for data cleansing, deduplication, reference-data enrichment, ingredient normalization, allergen mapping, and preparation of analytics-ready restaurant datasets.
oApplied Delta Lake capabilities—including ACID transactions, schema enforcement, MERGE operations, versioned data, and controlled reprocessing—to improve the reliability and recoverability of incremental pipelines.
oOrchestrated Databricks processing from Apache Airflow, managing job dependencies, parameters, retries, execution status, controlled backfills, and the movement of curated data into Snowflake for dimensional modeling and reporting.
oDeveloped approximately 30 reusable Airflow DAGs to coordinate ingestion, Databricks processing, validation, dbt transformations, downstream publishing, execution logging, failure recovery, and operational notifications.
oBuilt approximately 100 modular dbt models across staging, intermediate, dimensional, semantic, and reporting layers using incremental materializations, snapshots, seeds, Jinja, macros, source-freshness checks, automated tests, and generated documentation.
oImplemented the dbt Semantic Layer to establish governed metric definitions for restaurant engagement, menu coverage, transaction activity, allergen completeness, onboarding progress, and approval status.
oApplied Kimball dimensional-modeling principles to design conformed dimensions, fact tables, bridge tables, aggregate marts, and canonical business entities supporting restaurant, menu, allergen, transaction, engagement, and approval workflows.
oImplemented Snowflake Streams and Tasks, Dynamic Tables, Snowpipe, and dbt incremental strategies to support change data capture and efficient downstream refreshes without rebuilding complete datasets.
oEstablished comprehensive data-quality and observability controls using dbt tests, Databricks validation rules, freshness checks, reconciliation queries, duplicate and malformed-JSON detection, schema-drift monitoring, Google Cloud Monitoring, and automated email and Slack alerts.
oApplied data-governance practices through Snowflake RBAC, secure views, least-privilege access, Databricks Unity Catalog, controlled schema changes, data lineage, and auditable processing across Lakehouse and warehouse layers.
oManaged Databricks notebooks, dbt models, SQL, and Python code through GIT feature branches, pull requests, peer reviews, and GitHub Actions workflows supporting automated validation and controlled deployments across development, testing, and production environments.
oOptimized Databricks and Delta Lake workloads through appropriate partitioning, optimized file sizes, selective processing, Spark execution-plan analysis, and cluster-resource tuning while improving Snowflake performance through Query Profile analysis, warehouse right-sizing, workload isolation, and micro-partition pruning.
oReduced the quarterly Snowflake bill by approximately 35–45% through workload isolation, query refactoring, incremental processing, warehouse optimization, and appropriate distribution of transformation workloads.
oLeveraged Snowflake Cortex AI Functions and Databricks-prepared datasets to classify restaurant feedback, summarize recurring concerns, and support human-governed enrichment of ingredient metadata and candidate allergen classifications.
oPartnered with product, operations, analytics, and onboarding teams to translate requirements into reusable pipelines, governed data models, standardized metrics, and validation controls—doubling restaurant-onboarding capacity from five to ten locations per month.
Environment: Snowflake, Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Medallion Architecture, Databricks Auto Loader, Databricks Workflows, Unity Catalog, Databricks Notebooks, dbt, dbt Semantic Layer, SQL, Python, Apache Airflow, Fivetran, PostgreSQL, Google Cloud Storage, Google Cloud Monitoring, Snowpipe, External Stages, Streams and Tasks, Dynamic Tables, Snowflake Cortex AI Functions, REST APIs, JSON, Parquet, GIT, GitHub Actions, Slack, RBAC, Secure Views, Kimball Dimensional Modeling
Client: Arizona State University – Tempe, AZ July 2024 – April 2025
Role: Data Engineer
oDesigned multi-cloud analytics pipelines for an AI chatbot application, processing more than 100,000 monthly interactions from application logs, conversation events, user sessions, and model-response data.
oArchitected batch and event-driven ingestion workflows using Amazon S3, AWS Lambda, Google Cloud Storage, Cloud Functions, Pub/Sub, REST APIs, Python, and SQL to capture and process structured and semi-structured chatbot data.
oImplemented secure AWS-to-Snowflake ingestion using Amazon S3, IAM-backed storage integrations, external stages, and Snowpipe for chatbot-interaction and model-response data.
oDeveloped a Databricks Lakehouse processing layer on AWS using PySpark, Spark SQL, and Delta Lake to transform high-volume chatbot events, conversation histories, user sessions, and model-evaluation data.
oImplemented Bronze, Silver, and Gold Medallion layers to preserve raw interaction data, standardize and validate chatbot events, and publish curated datasets for Snowflake, BigQuery, and analytical reporting.
oUsed Databricks Auto Loader and Structured Streaming patterns to incrementally ingest evolving JSON files from Amazon S3, incorporating checkpointing, schema inference, schema evolution, and malformed-record handling.
oDeveloped PySpark and Spark SQL transformations for JSON flattening, timestamp standardization, duplicate-event removal, session reconstruction, intent normalization, response classification, and model-version comparisons.
oApplied Delta Lake capabilities—including ACID transactions, schema enforcement, MERGE operations, time travel, and controlled reprocessing—to improve pipeline reliability, traceability, and recovery.
oCreated Databricks notebooks and reusable Python components for API integration, JSON parsing, schema validation, exception management, execution logging, reconciliation, and historical backfills.
oApplied PySpark DataFrame transformations across Databricks, Google Cloud Dataproc, and AWS Glue for distributed processing, bulk-event reprocessing, and large-scale transformation of semi-structured interaction data.
oConfigured AWS Glue Crawlers and the AWS Glue Data Catalog to discover changing JSON schemas, maintain metadata for S3 datasets, and support interoperability across AWS ingestion and processing services.
oDeveloped Databricks Workflows and integrated them with Apache Airflow through Google Cloud Composer to coordinate ingestion, distributed transformation, dbt processing, validation, backfills, and downstream publishing.
oBuilt modular dbt models across staging, intermediate, dimensional, semantic, and reporting layers for sessions, conversation intents, fallback behavior, response effectiveness, containment, model versions, engagement, and chatbot performance.
oOptimized Databricks workloads through partition-aware processing, file compaction, appropriate cluster sizing, Spark execution-plan analysis, caching, and selective processing of incremental data.
oImproved BigQuery performance through partitioning, clustering, and selective filtering; optimized Snowflake through incremental merges and pruning-aware queries; and improved AWS Glue processing through job bookmarks and partition-aware reads.
oDesigned dimensional data marts and governed semantic models for sessions, conversations, intents, responses, feedback, model iterations, and engagement events, establishing standardized metrics for adoption, effectiveness, latency, containment, and model performance.
oCreated reproducible evaluation datasets enabling AI and product teams to compare chatbot behavior across model versions, intents, response categories, fallback scenarios, user segments, and defined evaluation periods.
oImplemented data-quality and observability controls using dbt tests, Delta Lake validation rules, custom SQL checks, freshness thresholds, malformed-event detection, duplicate-session checks, timestamp validation, and source-to-target reconciliation.
oMonitored pipelines through Databricks job metrics, Google Cloud Logging, Amazon CloudWatch, and Airflow notifications, supporting investigation of failed jobs, delayed events, schema changes, and data-quality exceptions.
oProtected sensitive chatbot information through restricted raw-data layers, masked or hashed analytical identifiers, Snowflake RBAC, AWS IAM roles, least-privilege S3 policies, and Databricks Unity Catalog permissions.
oManaged Databricks notebooks, dbt models, and Python components through GIT, pull requests, code reviews, and GitHub Actions, supporting automated validation and controlled deployments across development and production environments.
oMaintained data lineage, data dictionaries, source-to-target mappings, metric definitions, validation rules, and operational runbooks to support governance, troubleshooting, and knowledge transfer.
oPartnered with AI, analytics, product, and university stakeholders to translate chatbot-performance requirements into governed datasets, reusable metrics, model-evaluation data, and Tableau-ready reporting.
Environment: Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Medallion Architecture, Databricks Auto Loader, Structured Streaming, Databricks Workflows, Databricks Notebooks, Unity Catalog, Snowflake, BigQuery, dbt, SQL, Python, Apache Airflow, Google Cloud Composer, Google Cloud Storage, Google Cloud Functions, Pub/Sub, Google Cloud Dataproc, Google Cloud Logging, Amazon S3, AWS Lambda, AWS Glue, AWS Glue Crawlers, AWS Glue Data Catalog, Amazon CloudWatch, AWS IAM, Snowflake Storage Integrations, External Stages, Snowpipe, REST APIs, JSON, Parquet, Tableau, GIT, GitHub Actions, RBAC, Data Masking and Hashing, Dimensional Modeling, Data Quality and Observability
Client: Bulls and Bears Advisory – Bangalore, India Jan 2023 – Jul 2023
Role: Data Engineer
oDesigned and implemented a streaming analytics platform integrating market signals, transactions, portfolio positions, currency movements, user events, and product-usage data into Databricks, Delta Lake, and Snowflake reporting layers.
oDeveloped event-driven and batch-ingestion pipelines processing approximately one million JSON events daily and supporting near-real-time reporting with approximately one-minute latency.
oBuilt Kafka producers, topics, partitions, consumers, and consumer groups, implementing offset management, replay-safe processing, event-level deduplication, payload validation, quarantine, and controlled reprocessing.
oDeveloped Databricks Structured Streaming pipelines using PySpark to consume Kafka events, process high-volume market and transaction data, and persist validated records within Delta Lake.
oImplemented Bronze, Silver, and Gold Medallion layers to preserve raw financial events, standardize market and portfolio data, apply business validations, and prepare curated datasets for Snowflake and Tableau reporting.
oApplied streaming checkpoints, event-time processing, watermarking, late-arriving-data handling, and idempotent processing patterns to improve reliability and prevent duplicate financial transactions during pipeline recovery.
oUsed Delta Lake schema enforcement, schema evolution, ACID transactions, and MERGE operations to support changing event structures and incremental updates to portfolio and transaction datasets.
oMaintained Kafka Connect and the Snowflake Connector for Kafka for low-latency operational feeds, while using Databricks for complex transformations, historical replay, enrichment, and large-scale analytical processing.
oDeveloped PySpark, Spark SQL, Python, and Pandas components for JSON normalization, reference-data enrichment, currency validation, portfolio calculations, exception handling, reconciliation, and controlled reprocessing of rejected events.
oDeveloped modular dbt models across staging, intermediate, dimensional, and reporting layers to standardize market feeds, transactions, portfolio positions, currency activity, trading signals, and user behavior.
oImplemented incremental dbt models and Snowflake Streams and Tasks to capture new and changed records, apply controlled merge logic, and refresh downstream analytical tables efficiently.
oDesigned conformed dimensions, transaction and activity fact tables, portfolio snapshots, aggregate marts, and reusable business entities supporting trading activity, portfolio exposure, currency movement, market-signal performance, and product engagement.
oEstablished automated data-quality controls covering missing transaction identifiers, duplicate events, invalid timestamps, unexpected currency codes, incomplete portfolio records, schema changes, referential-integrity failures, and source-to-target reconciliation.
oMonitored Kafka consumer lag, message throughput, streaming-batch duration, ingestion freshness, checkpoint status, malformed events, and processing failures through Kafka metrics, Databricks job monitoring, and Google Cloud Logging.
oOptimized Databricks and Delta Lake processing through partition-aware transformations, file compaction, selective reads, Spark execution-plan analysis, and cluster-resource tuning.
oConfigured separate Snowflake virtual warehouses for ingestion, transformation, and Tableau workloads, improving performance through right-sizing, auto-suspend and auto-resume, workload isolation, Query Profile analysis, and concurrency management.
oApplied governed access controls using Databricks Unity Catalog, Snowflake RBAC, secure views, least-privilege permissions, data lineage, and controlled access to sensitive portfolio, transaction, and user information.
oManaged Databricks notebooks, PySpark, SQL, and dbt development through GIT, code reviews, and Jenkins pipelines, supporting automated validation and controlled deployments.
oDelivered Tableau-ready data marts that replaced manual daily extracts with near-real-time reporting while maintaining data lineage, Kafka documentation, quality rules, and operational recovery procedures.
Environment: Databricks, Apache Spark, PySpark, Spark SQL, Structured Streaming, Delta Lake, Medallion Architecture, Databricks Workflows, Databricks Notebooks, Unity Catalog, Snowflake, dbt, Apache Kafka, Kafka Connect, Snowflake Connector for Kafka, Kafka Producers and Consumers, Kafka Topics, Partitions and Consumer Groups, Snowflake Streams and Tasks, Snowflake Virtual Warehouses, SQL, Python, Pandas, JSON, Parquet, Google Cloud Storage, Google Cloud Logging, GIT, Jenkins, Tableau, RBAC, Secure Views, Event-Driven Architecture, Dimensional Modeling, Data Quality and Monitoring
Client: Mega Paints – Tirunelveli, India Aug 2019 – Nov 2022
Role: Data Engineer
oContributed to an enterprise data-modernization program that migrated approximately 60–70 tables containing ERP, CRM, POS, inventory, pricing, customer, supplier, sales, and finance data into Databricks and Snowflake.
oDesigned a scalable migration architecture using Informatica for legacy-source extraction, Databricks for distributed transformation and historical processing, and Snowflake for governed warehousing and reporting.
oIntegrated data from PostgreSQL, spreadsheets, flat files, and legacy applications into controlled landing, standardized, dimensional, semantic, and reporting layers.
oProfiled source systems and created detailed source-to-target mappings covering data structures, dependencies, historical requirements, business rules, data-type conversions, reference-data standardization, surrogate keys, and identified quality issues.
oDeveloped PySpark and Spark SQL transformations in Databricks to cleanse, standardize, deduplicate, enrich, and reconcile large ERP, sales, inventory, customer, supplier, and financial datasets.
oImplemented Delta Lake Bronze and Silver layers to retain raw migration data, apply standardization and quality controls, and prepare trusted datasets for loading into Snowflake dimensional models.
oUsed Delta Lake schema enforcement, schema evolution, ACID transactions, and MERGE operations to manage changing source structures and support reliable incremental processing.
oDeveloped reusable, configuration-driven ingestion utilities using Python, Pandas, SQL, and Bash/Shell to validate and load CSV, JSON, and Parquet data while supporting audit logging, rejected records, and controlled reprocessing.
oImplemented full-load and incremental-processing patterns using extraction timestamps, high-watermark logic, business keys, audit columns, Delta MERGE operations, and Snowflake upserts to process newly created and modified records.
oBuilt modular dbt models across staging, intermediate, dimensional, semantic, and reporting layers using macros, incremental materializations, seeds, snapshots, automated tests, documentation, and source-freshness checks.
oDesigned Kimball star-schema data marts containing conformed dimensions, fact tables, aggregate tables, and SCD Type 1 and Type 2 patterns for sales, inventory, pricing, demand, customer, supplier, finance, and regional reporting.
oExecuted iterative migration cycles covering source profiling, Databricks transformation, trial loads, historical backfills, defect remediation, parallel validation, cutover preparation, and post-migration verification.
oUsed Databricks notebooks and Jobs to execute parameterized migration routines, large-scale historical backfills, exception reprocessing, and reconciliation workflows across development and testing environments.
oEstablished comprehensive migration controls using row-count comparisons, control totals, source-to-target reconciliation, duplicate and referential-integrity checks, business-rule validation, exception reporting, and financial-balance verification.
oOrchestrated daily Informatica workflows with parameterized execution, dependency checks, centralized logging, retries, failure notifications, checkpoints, restartable processing, and quarantine procedures for invalid records.
oOptimized Databricks workloads through partition-aware reads, selective transformations, file compaction, caching, Spark execution-plan analysis, and appropriate cluster-resource configuration.
oOptimized Snowflake processing through SQL tuning, Query Profile analysis, warehouse sizing, pruning-aware query design, incremental processing, caching awareness, and workload scheduling.
oProtected sensitive customer and financial information through Snowflake RBAC, least-privilege permissions, secure views, controlled data access, and auditable processing and publishing procedures.
oManaged Databricks notebooks, PySpark, Python, SQL, Bash/Shell, and dbt code through GIT, peer reviews, automated validation, and GitHub Actions for controlled deployments across development, testing, and production environments.
oPartnered with finance, sales, inventory, procurement, operations, and reporting stakeholders to replace fragmented spreadsheet-based consolidation with governed, traceable, and maintainable data pipelines and analytical datasets.
Environment: Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Databricks Notebooks, Databricks Jobs, Snowflake, dbt, Informatica, PostgreSQL, SQL, Python, Pandas, Bash/Shell, CSV, JSON, Parquet, GIT, GitHub Actions, Snowflake RBAC, Secure Views, ETL/ELT, Full and Incremental Loads, High-Watermark Processing, Delta MERGE, Data Migration, Data Quality, Source-to-Target Mapping, Data Reconciliation, Kimball Dimensional Modeling, SCD Type 1 and Type 2
Client: Temenos – Chennai, India May 2018 – Jul 2019
Role: Data Analyst
oDeveloped SQL- and Python-based data-processing and analytics workflows supporting customer acquisition, product adoption, engagement, retention, churn, and campaign-performance reporting.
oExtracted and consolidated data from customer, account, campaign, product-usage, and application-event sources into standardized analytical and operational reporting datasets.
oWrote complex Oracle SQL and PL/SQL queries using joins, common table expressions, subqueries, window functions, aggregations, and reusable business rules to transform and analyze enterprise data.
oSupported the development and maintenance of PL/SQL stored procedures, functions, views, and scheduled transformation routines used to extract, standardize, validate, and publish recurring reporting data.
oDeveloped Python and Pandas utilities to automate data extraction, cleansing, transformation, validation, reconciliation, and report preparation, reducing repetitive manual effort.
oStandardized KPI definitions and reusable calculations for customer acquisition, active users, product adoption, engagement, retention, churn, campaign response, and business growth.
oPerformed customer segmentation, cohort analysis, trend analysis, and regression-based analysis to identify behavioral patterns and factors influencing customer engagement and retention.
oBuilt and enhanced analytical dashboards and recurring reports containing KPI summaries, trend views, segment comparisons, and drill-down analysis for marketing, product, and management stakeholders.
oImplemented data-profiling, quality, and reconciliation checks to identify missing identifiers, duplicate records, invalid dates, inconsistent status values, incomplete activity data, and source-to-report discrepancies.
oSupported downstream validation of Kafka-fed application events by identifying missing, delayed, malformed, and duplicate records and coordinating discrepancies with application and data teams.
oImproved Oracle SQL and PL/SQL performance through execution-plan analysis, query simplification, selective filtering, reduced data scans, and appropriate indexing recommendations.
oAssisted with user-acceptance testing by preparing test scenarios, validating dashboard calculations, documenting discrepancies, and coordinating corrections with technical and business stakeholders.
oMaintained data dictionaries, KPI definitions, source-to-report mappings, report specifications, validation procedures, and query documentation to improve reporting consistency and knowledge transfer.
oCollaborated with marketing, product, operations, database, and reporting teams to translate analytical requirements into validated datasets, reusable calculations, and actionable business reporting.
Environment: Oracle Database, Oracle SQL, PL/SQL, Python, Pandas, SQL, Kafka-Fed Application Events, Data Extraction and Transformation, Data Profiling, Data Cleansing, Data Validation, Data Reconciliation, Source-to-Report Mapping, Regression Analysis, Cohort Analysis, Customer Segmentation, KPI Reporting, UAT and BI Dashboards
Technical Skills:
Databricks & Lakehouse: Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Medallion Architecture, Bronze/Silver/Gold Layers, Auto Loader, Structured Streaming, Databricks Workflows, Databricks Jobs, Databricks Notebooks, Unity Catalog, Delta MERGE, ACID Transactions, Schema Enforcement and Evolution, Checkpointing, Watermarking, Time Travel, Partitioning, File Compaction and Cluster Tuning
Cloud Platforms: AWS – Amazon S3, AWS Glue, Glue Crawlers, Glue Data Catalog, Glue Job Bookmarks, AWS Lambda, Amazon CloudWatch and AWS IAM; GCP – Google Cloud Storage, BigQuery, Cloud Functions, Pub/Sub, Dataproc, Cloud Composer, Cloud Monitoring and Cloud Logging
Snowflake & Warehousing: Snowflake, Snowpipe, External Stages, AWS IAM-Backed Storage Integrations, Streams and Tasks, Dynamic Tables, Virtual Warehouses, Query Profile, Warehouse Right-Sizing, Workload Isolation, Micro-Partition Pruning, Clustering Evaluation, Secure Views, RBAC, PostgreSQL and Oracle Database
dbt & Analytics Engineering: dbt, dbt Semantic Layer, Incremental Models, Snapshots, Seeds, Jinja, Macros, Packages, Schema and Custom Tests, Source Freshness, dbt Documentation, Model Lineage, Semantic Models, Governed Metrics and Analytics-Ready Data Products
Programming & Data Processing: Python, PySpark, Spark SQL, Pandas, Advanced SQL, Oracle SQL, PL/SQL, Bash/Shell, REST APIs, JSON, Parquet, Semi-Structured Data Processing, Data Cleansing, Deduplication, Enrichment and Reconciliation
Data Integration & Migration: ELT/ETL, Multi-Cloud Integration, Batch and Event-Driven Pipelines, Change Data Capture, Full and Incremental Loads, High-Watermark Processing, Historical Backfills, Data Migration, Fivetran, POS Data Integration, Configuration-Driven Ingestion, Source-to-Target Mapping and Failure Recovery
Orchestration & Streaming: Apache Airflow, Google Cloud Composer, Databricks Workflows, AWS Glue, Informatica, Apache Kafka, Kafka Connect, Snowflake Connector for Kafka, Kafka Producers and Consumers, Topics, Partitions, Consumer Groups, Offset Management, Replay-Safe Processing, Google Pub/Sub and Structured Streaming
Modeling, Quality & Governance: Kimball Dimensional Modeling, Star Schemas, Conformed Dimensions, Fact Tables, Bridge Tables, Aggregate Marts, SCD Type 1 and Type 2, Data Profiling, Data Quality, Schema-Drift Detection, Freshness Monitoring, Data Lineage, Unity Catalog, Data Masking and Hashing, Snowflake RBAC, AWS IAM and Least-Privilege Security
DevOps, Monitoring & BI: GIT, GitHub Actions, Jenkins, Pull Requests, Peer Reviews, Automated Validation, CI/CD, Environment-Based Deployments, Databricks Job Monitoring, Amazon CloudWatch, Google Cloud Monitoring, Cloud Logging