Post Job Free
Sign in

Lead Data Engineer & Cloud Lakehouse Architect

Location:
Everett, WA
Salary:
125000
Posted:
August 24, 2026

Contact this candidate

Resume:

Alee Mohsin

Lead Data Engineer Cloud & Lakehouse Architecture

Real-Time Analytics & Data Platform Strategy

Dallas, TX 75201 432-***-**** **********@*****.*** linkedin.com/in/alee-syed-b87a04423 Professional Summary

Lead Data Engineer and cloud lakehouse architect with 10+ years designing and scaling enterprise-grade data platforms across healthcare and Fortune-caliber enterprise environments. Deep expertise across Databricks, Snowflake, Apache Spark, Kafka, and Delta Lake, with hands-on multi-cloud architecture spanning Azure, AWS, and GCP. Track record of translating fragmented, high-volume data (15TB+ daily) into governed, analytics-ready lakehouse ecosystems that power AI/ML, BI, and predictive analytics at enterprise scale, while maintaining HIPAA, HITRUST, and SOC 2 compliance in highly regulated healthcare settings.

Known for architecting real-time streaming and Medallion frameworks that cut data delivery time by up to 60% and improve query performance by 40%, while leading distributed engineering teams through CI/CD modernization, Infrastructure as Code, and data governance transformation.Combines deep technical fluency in HL7/FHIR healthcare interoperability with strategic platform ownership partnering directly with data science, clinical, and executive stakeholders to convert data infrastructure into a measurable competitive advantage. Four major cloud and data-engineering certifications (Databricks, Azure, GCP, AWS) underscore a platform-agnostic, architecture-first approach to solving enterprise data problems.

Core Competencies & Technical Skills

Programming & Integration

Python, SQL, Scala, Java, Bash/Shell Scripting, REST APIs, GraphQL, Advanced SQL Optimization Big Data & Streaming

Apache Spark (PySpark, Spark SQL, Structured Streaming), Apache Kafka, Apache Flink, Delta Lake, Delta Live Tables, Databricks, Apache Beam, Kafka Connect, Debezium CDC, Azure Event Hubs, AWS Kinesis, Google Pub/Sub Multi-Cloud Platforms

Azure (Databricks, Synapse, Data Factory, ADLS Gen2, Cosmos DB, Event Hubs, Key Vault), AWS (Glue, EMR, Redshift, Lambda, S3, DynamoDB, Kinesis, Step Functions), GCP (BigQuery, Dataflow, Composer, GCS, Vertex AI) Orchestration & Automation

Apache Airflow, dbt, Azure Data Factory, AWS Glue, Databricks Workflows, Informatica, Talend, Apache NiFi, CI/CD Pipelines

Databases & Storage

Snowflake, Delta Lake, Apache Iceberg, Apache Hudi, Hadoop, Hive, MongoDB, Cassandra, PostgreSQL, MySQL, SQL Server, Oracle, Redis, Parquet, ORC, Avro

DevOps & Infrastructure

Terraform, Kubernetes (AKS/EKS/GKE), Docker, Jenkins, GitHub Actions, Azure DevOps, GitOps, ArgoCD, Infrastructure as Code, MLflow

Governance, Security & Observability

Unity Catalog, Collibra, Azure Purview, Great Expectations, Monte Carlo, OpenLineage, Data Lineage, RBAC/IAM, HIPAA, HITRUST, GDPR, SOC 2, Grafana, Prometheus, ELK Stack Architecture & Strategy

Cloud-Native Lakehouse Architecture, Data Warehousing, Data Mesh, Event-Driven Architecture, Dimensional & Data Vault Modeling, Medallion Architecture, Batch & Real-Time Processing, Enterprise Data Integration Strategy BI & Analytics

Power BI, Tableau, Looker, Qlik, Semantic Layer Modeling, Executive Dashboards, KPI Reporting, Predictive Analytics Healthcare Data Engineering

HL7, FHIR, CCDA, Epic & Cerner EHR/EMR Integration, Claims & Clinical Data Pipelines, Population Health Analytics, Healthcare Interoperability

Professional Experience

Lead Data Engineer Nov 2022 – Present

Tecton Remote

• Architected a HIPAA-compliant, cloud-native healthcare lakehouse on Azure and Databricks, unifying EHR, clinical, claims, and IoT device data into a single governed platform for enterprise analytics and AI-driven decision support.

• Designed low-latency streaming pipelines with Apache Kafka, Apache Flink, Spark Structured Streaming, and Delta Live Tables to process HL7/FHIR feeds and medical device telemetry in near real time.

• Engineered a Medallion Architecture on Delta Lake processing 15TB+ of healthcare data daily, lifting analytics query performance 40% and cutting end-to-end data delivery time 60%.

• Automated enterprise ETL/ELT workflows using Airflow, dbt, and Databricks Workflows, materially improving pipeline reliability, scalability, and freshness of clinical and claims data.

• Built enterprise data governance and observability controls with Unity Catalog, Collibra, Great Expectations, and OpenLineage, sustaining continuous HIPAA, HITRUST, and enterprise compliance readiness.

• Directed cloud cost and performance optimization across Databricks and Azure, lowering infrastructure spend while strengthening SLA adherence and processing throughput.

• Partnered with data science, analytics, and clinical stakeholders to operationalize AI/ML use cases in patient risk stratification, readmission prediction, and care-pathway optimization.

• Established CI/CD and Infrastructure as Code practices with Terraform, GitHub Actions, and Azure DevOps, reducing deployment time 70% and improving release reliability across 50+ production pipelines.

• Led and mentored a 10-person data engineering team, owning architecture decisions, sprint planning, and code review standards for enterprise-grade solution delivery.

Senior Data Engineer Jun 2019 – Oct 2022

Scion Staffing

• Designed and maintained Spark-based ETL/ELT pipelines processing multi-terabyte enterprise datasets, improving transformation performance 35% to support real-time analytics initiatives.

• Built modern lakehouse and data warehouse architectures on Snowflake, Delta Lake, Amazon Redshift, and Databricks to power high-performance analytical workloads.

• Automated orchestration across Apache Airflow, dbt, and Databricks, strengthening data reliability, scheduling efficiency, and end-to-end pipeline monitoring.

• Implemented data governance controls spanning schema validation, lineage tracking, and role-based access, raising data quality and regulatory compliance posture.

• Tuned Spark, SQL, and distributed processing workloads, improving query performance and resource utilization across large-scale data processing environments.

• Built incremental ingestion frameworks and SCD Type 2 data models enabling accurate historical tracking for enterprise reporting.

• Integrated FHIR, claims, and EHR data sources to extend healthcare interoperability and unify analytics across clinical and enterprise domains.

• Mentored junior engineers on cloud-native data engineering and CI/CD best practices, improving team delivery efficiency 30% and cutting onboarding time 40%.

Big Data Engineer Aug 2016 – May 2019

ProCogia

• Built scalable batch and real-time pipelines with Apache Spark, Kafka, AWS Glue, and EMR to process large-scale enterprise and healthcare datasets with high reliability.

• Architected Delta Lake-based data platforms enabling self-service analytics, ML workflows, and enterprise reporting capabilities.

• Modernized legacy ETL systems by migrating workloads to AWS-native architectures, improving scalability, fault tolerance, and operational efficiency.

• Developed automated data quality, reconciliation, and schema-evolution frameworks, reducing pipeline failures 45% and strengthening data accuracy across distributed platforms.

• Delivered analytics-ready data marts and curated datasets powering executive dashboards, compliance reporting, and BI initiatives.

• Implemented proactive monitoring and alerting with CloudWatch and Grafana, improving observability, platform stability, and incident response time.

• Standardized reusable Spark and Glue job templates and collaborated cross-functionally with analytics and compliance teams, shortening new-pipeline delivery cycles and aligning designs to audit requirements. Data Engineer Jul 2014 – Jul 2016

NeoReach

• Built and optimized ETL pipelines integrating EHR, laboratory, billing, and operational healthcare data, processing 20M+ records monthly at 99.9% reliability.

• Designed dimensional data models and analytics-ready marts supporting 200+ business users and executive clinical/operational dashboards.

• Automated workflow scheduling, validation, and monitoring, cutting manual operational effort 65% and improving SLA adherence.

• Enhanced SQL transformation and query optimization strategies, improving reporting performance and data accuracy across large datasets.

• Built custom HL7 and FHIR integration solutions to streamline healthcare interoperability across clinical and third-party systems.

• Implemented secure, HIPAA-aligned data processing frameworks meeting healthcare governance and auditability standards. Projects

Cloud-Native Data Governance & Observability Framework Databricks, Unity Catalog, Collibra, Great Expectations, OpenLineage, Azure DevOps, Terraform

• Architected a cloud-native healthcare lakehouse consolidating EHR, claims, pharmacy, and operational data into a unified, governed analytics ecosystem.

• Built scalable ETL/ELT and Delta Lake frameworks with Spark, Airflow, and dbt to support enterprise reporting, AI/ML workloads, and self-service analytics.

Real-Time Fraud Detection & Streaming Analytics Platform Apache Kafka, Spark Structured Streaming, AWS Kinesis, Delta Lake, Snowflake, Python

• Designed high-throughput, low-latency streaming pipelines processing enterprise and healthcare event data for real-time fraud detection.

• Built monitoring and Infrastructure as Code solutions with Grafana and Terraform, improving platform reliability, scalability, and operational efficiency.

Enterprise Self-Service BI & Data Mart Platform

Snowflake, dbt, Power BI, Tableau, Airflow, Azure Synapse, SQL

• Led enterprise data modernization through governed lakehouse and warehouse solutions delivering secure, analytics-ready data assets.

• Established governance, lineage, and quality-validation frameworks with Collibra, Purview, and Great Expectations, improving compliance and cross-domain data trust.

Certifications

• Databricks Certified Data Engineer

• Microsoft Certified: Azure Data Engineer

• Google Cloud Professional Data Engineer

• AWS Certified Data Engineer

Education

Bachelor of Science in Computer Science



Contact this candidate