Post Job Free
Sign in

Cloud Data Platform Engineering Leader

Location:
Norfolk, VA
Salary:
130000
Posted:
October 06, 2026

Contact this candidate

Resume:

Aatif Ahsan

Data Platform Engineering Leader Cloud Data Architecture Lakehouse, Streaming & AI-Ready Platforms Professional Summary

Accomplished Lead Data Engineer with 10+ years of experience architecting and modernizing cloud data platforms, data modernization solutions, and analytics ecosystems across healthcare, FinTech, and complex business environments. Experienced in designing scalable lakehouse architectures, ETL/ELT pipelines, real-time data processing solutions, and cloud data platforms using Python, SQL, Apache Spark, Databricks, Snowflake, Azure, and AWS. Specialized in building secure healthcare data ecosystems with EHR/EMR integrations, FHIR/HL7 workflows, HIPAA-compliant processing, and PHI/PII protection. Proven expertise in delivering governed data products using Palantir Foundry, implementing data quality and lineage frameworks, optimizing cloud performance and cost efficiency, and leading engineering teams to deliver reliable platforms supporting analytics and AI/ML initiatives. Technical Skills

Programming & Data Engineering:

Python (ETL automation, PySpark, APIs, data validation), SQL (advanced queries, optimization, CTEs, window functions), Scala, Java, Bash/Shell, ETL Automation, Data Transformation, API Integration, Data Processing Frameworks Data Engineering & Orchestration:

ETL/ELT Pipelines, Data Ingestion, Batch Processing, Near Real-Time Processing, Apache Airflow, dbt, Azure Data Factory, AWS Glue, Data Pipeline Automation, Workflow Orchestration, Pipeline Monitoring Big Data & Distributed Processing:

Apache Spark, PySpark, Spark SQL, Structured Streaming, Kafka, Kafka Streams, Kafka Connect, Schema Registry, Debezium CDC, Distributed Data Processing, Large-Scale Data Processing, Event-Driven Architectures Lakehouse & Data Platforms:

Databricks, Delta Lake, Unity Catalog, Databricks Workflows, Delta Live Tables, Bronze/Silver/Gold Architecture, Lakehouse Architecture, Snowflake, Modern Data Platforms, Data Products, Palantir Foundry, Palantir Ontology, Governed Data Workflows Cloud Platforms:

AWS (S3, Glue, EMR, Lambda, Redshift, Athena, Kinesis), Azure (Azure Data Factory, Azure Databricks, Synapse Analytics, ADLS Gen2, Event Hubs, Key Vault), Cloud Data Modernization, Cloud Migration, Scalable Cloud Architectures Data Warehousing & Data Modeling:

Data Warehousing, Dimensional Modeling, Star Schema, Snowflake Schema, Data Vault, Kimball Methodology, Inmon Modeling, Data Marts, Semantic Layer, Analytical Data Models, Warehouse Optimization Databases & Storage:

PostgreSQL, SQL Server, Oracle, MySQL, MongoDB, Cassandra, DynamoDB, Parquet, Avro, ORC, Data Storage Optimization, Query Performance Tuning

Healthcare Data Engineering & Governance:

EHR/EMR Integration, FHIR, HL7, Clinical Data, Claims Data, Laboratory Data, Healthcare Analytics, HIPAA Compliance, PHI/PII Protection, Healthcare Data Pipelines, Secure Data Processing Data Governance & Reliability:

Data Governance Frameworks, Data Lineage, Metadata Management, Data Cataloging, Data Quality Frameworks, Data Validation, Data Observability, Monitoring, Alerting, Access Controls, Data Reliability Engineering AI/ML Data Platforms & DevOps:

AI/ML Data Preparation, Feature Engineering, MLflow, AI-Ready Data Platforms, Docker, Kubernetes, Terraform, Infrastructure as Code, CI/CD, GitHub Actions, GitLab CI/CD

Leadership, Delivery & Collaboration:

Agile & Scrum, Jira, Confluence, Stakeholder Management, Solution Architecture, Technical Mentorship, DataOps, Engineering Documentation, Architecture Reviews, Cross-Functional Collaboration ************@*****.*** 951-***-**** Norfolk, VA, 23523, US Professional Experience

Lead Data Engineer 11/2022 – Present

Coralogix

• Architected scalable cloud data platforms by integrating 25+ healthcare, clinical, operational, and analytical data sources into governed ecosystems supporting analytics, reporting, and AI/ML initiatives.

• Spearheaded the design of modern lakehouse architectures using Databricks, Delta Lake, and medallion frameworks, enabling scalable processing of multi-terabyte workloads and accelerating analytics delivery by 35%.

• Engineered secure ETL/ELT pipelines with Python, SQL, Apache Spark, Snowflake, and Azure services to process millions of records while improving data reliability and operational efficiency.

• Developed Palantir Foundry data products, ontology-driven workflows, and governed analytical solutions to standardize complex healthcare datasets and enhance data accessibility across teams.

• Modernized healthcare data ecosystems by integrating EHR/EMR, claims, laboratory, and clinical data while implementing HIPAA- aligned security controls, PHI/PII protection, and governed access frameworks.

• Built batch and near-real-time data processing solutions using Spark, Kafka, and cloud-native services, reducing critical data delivery latency from hours to minutes.

• Established data governance practices including lineage, quality validation, observability, and monitoring frameworks, reducing production data issues by 30%.

• Optimized Spark workloads, SQL transformations, Delta Lake storage strategies, and cloud resource utilization, improving pipeline execution efficiency by 25–30% and enhancing compute efficiency.

• Led technical delivery through architecture reviews, engineering standards, CI/CD practices, mentorship, and cross-functional collaboration to deliver scalable, reliable, and AI-ready data platforms. Tech Stack: Databricks, Delta Lake, Snowflake, Apache Spark, PySpark, Python, SQL, Azure Data Factory, Azure Databricks, AWS, Kafka, Palantir Foundry, Data Governance, Data Lineage, HIPAA, FHIR, HL7 Senior Data Platform Engineer 03/2020 – 10/2022

LangChain

• Engineered scalable cloud data platforms supporting analytics, reporting, and machine learning workloads across healthcare, financial, and operational domains.

• Designed and automated enterprise ETL/ELT workflows using Python, SQL, Apache Airflow, dbt, and cloud integration services, managing 100+ production pipelines with improved reliability.

• Developed optimized Spark and Databricks processing frameworks for large-scale datasets, improving transformation efficiency and accelerating downstream analytics delivery.

• Built scalable data models, analytical datasets, and warehouse solutions using Snowflake, Delta Lake, and dimensional modeling techniques to support business intelligence initiatives.

• Implemented Kafka-based streaming architectures and event-driven pipelines, reducing analytics latency by 50% and enabling near-real-time data consumption.

• Delivered reliable data platforms by implementing automated validation, quality checks, monitoring, and recovery processes across production environments.

• Collaborated with engineering, analytics, and business teams to translate complex requirements into scalable platform solutions while contributing to architecture improvements.

Tech Stack: Databricks, Snowflake, Apache Spark, PySpark, Python, SQL, Apache Airflow, dbt, Kafka, Delta Lake, AWS, Azure, Data Modeling, Data Quality Frameworks

Cloud Data Engineer 07/2017 – 02/2020

SentiLink

• Engineered cloud-based data platforms by developing scalable ingestion and transformation frameworks that integrated 20+ structured and semi-structured data sources into centralized analytical environments.

• Designed and optimized ETL/ELT pipelines using Python, SQL, Apache Spark, and PySpark to process terabyte-scale datasets, improving data availability and supporting enterprise analytics workloads.

• Built and managed cloud data lake solutions using AWS S3, Azure Data Lake, and optimized storage formats such as Parquet, improving storage efficiency and enabling faster analytical processing.

• Developed batch and near-real-time processing workflows using Spark and distributed computing frameworks, reducing data processing cycles and improving delivery speed for downstream consumers.

• Integrated databases, APIs, and external systems into cloud-based data platforms, creating reliable data flows that supported reporting, analytics, and business intelligence initiatives.

• Optimized Spark workloads through partitioning, caching, and query tuning techniques to improve processing efficiency across large-scale transformation workloads.

• Strengthened production data reliability by implementing monitoring, validation, and troubleshooting practices, reducing pipeline failures and improving operational stability across cloud environments. Tech Stack: AWS S3, Azure Data Lake, Apache Spark, PySpark, Python, SQL, Parquet, ETL/ELT, Cloud Data Platforms, Data Integration, Batch Processing, Data Validation, Data Monitoring Data Engineer 06/2015 – 06/2017

ReversingLabs

• Developed and maintained ETL workflows using Python, SQL, and relational database technologies to integrate 15+ enterprise data sources into reliable analytical environments.

• Built data ingestion processes connecting APIs, databases, and flat-file systems while standardizing transformation logic and improving data consistency across reporting workflows.

• Created SQL-based data models, stored procedures, and curated datasets that supported business intelligence reporting and operational analytics requirements.

• Optimized SQL queries and transformation processes, improving data processing efficiency by 25% and accelerating delivery of reporting datasets.

• Implemented data validation, monitoring, and documentation practices to improve dataset accuracy, reliability, and maintainability across production environments.

• Collaborated with analysts and business stakeholders to translate reporting requirements into structured datasets and scalable data solutions.

• Supported deployment activities, version control practices, and production troubleshooting processes to improve workflow stability and engineering collaboration.

Tech Stack: Python, SQL, ETL/ELT, PostgreSQL, SQL Server, Oracle, APIs, Data Integration, Data Modeling, Data Validation, Reporting Solutions

Projects

AI-Powered Healthcare Knowledge Graph Platform

Tech Stack: Palantir Foundry, Knowledge Graphs, Databricks, Unity Catalog, Data Semantics, MLflow, Python, SQL, Azure OpenAI, FHIR

• Designed an intelligent healthcare knowledge platform that connected fragmented clinical and operational information into a unified semantic layer for advanced analytics and AI-driven insights.

• Developed reusable data intelligence workflows combining structured healthcare data, metadata enrichment, and AI-ready datasets, improving information discovery and reducing analysis turnaround time by 40%.

• Established secure foundations for AI adoption by implementing governed data access, quality standards, and trusted datasets supporting AI/ML applications.

Enterprise Data Marketplace & Self-Service Analytics Ecosystem Tech Stack: Snowflake, Databricks, Data Products, dbt, Apache Airflow, Data Catalogs, AWS, Azure, SQL, Python

• Created a self-service data ecosystem enabling business teams to discover, understand, and utilize trusted analytical assets without dependency on engineering teams.

• Engineered reusable data assets and standardized analytical models that improved reporting consistency and reduced duplicate dataset creation by 35%.

• Implemented automated publishing, documentation, and ownership workflows to increase data product adoption across enterprise teams.

Cloud-Native Data Observability & Reliability Framework Tech Stack: AWS, Azure, CloudWatch, Azure Monitor, Apache Kafka, Apache Spark, Python, SQL, Data Quality Frameworks

• Developed a cloud data reliability framework to proactively identify pipeline issues, data anomalies, and operational risks across distributed data environments.

• Implemented automated monitoring, quality checks, and alerting mechanisms that improved production visibility and reduced troubleshooting time by 45%.

• Designed scalable reliability practices for cloud data workloads, improving operational stability and supporting high-availability analytics platforms.

Certifications

• Databricks Certified Data Engineer Professional • SnowPro Core Certification

• Microsoft Certified: Azure Data Engineer Associate (DP-203) • AWS Certified Data Engineer – Associate Education

Bachelor of Science in Computer Science



Contact this candidate