Post Job Free
Sign in

Senior Data Engineer - Cloud Lakehouse Architecture

Location:
Dallas, TX
Posted:
July 22, 2026

Contact this candidate

Resume:

Asfar Rizvi

Senior Data Engineer Data Architecture Cloud & Lakehouse Platforms

Phone: 254-***-**** Email: ******.************@*****.*** Location: Dallas, TX, 75201 GitHub: github.com/Asfar-Rizvi Website: asfarrizvi.netlify.app PROFESSIONAL SUMMARY

Senior Data Engineer with 9+ years of experience designing scalable cloud-native data platforms across AWS, Azure and GCP. Skilled in building high-performance ETL/ELT pipelines, real-time streaming solutions and modern Lakehouse architectures using Spark, Kafka, Snowflake, Databricks and Airflow. Experienced in DataOps, CI/CD, cloud data architecture and data governance, delivering secure, reliable and scalable solutions that support enterprise analytics, AI and business intelligence.

SKILLS

• Programming & Scripting: Python, SQL, Scala, Java, Shell Scripting, REST APIs

• Distributed Processing & Big Data: Apache Spark, Hadoop Ecosystem (Hive, Pig, HDFS, MapReduce), Kafka, Flink, Presto, Delta Lake, Druid, Storm

• Analytics & Data Warehousing: Columnar Databases (Redshift, BigQuery, Vertica, Greenplum), Snowflake, Databricks, Tableau, Power BI, Looker, Olik, Superset, Mode Analytics, Excel (Advanced), Data Visualization Best Practices

• Databases & Storage: Relational Databases (PostgreSQL, MySQL, Oracle, SQL Server), NoSQL (MongoDB, Cassandra, DynamoDB, Redis, Couchbase), Graph Databases (Neo4j, Amazon Neptune)

• MLOps & Machine Learning: Feature Engineering, ML Data Preparation, Model Deployment Support (MLflow, SageMaker, Vertex AI), A/B Testing Data Pipelines, Real- time Analytics, Data for AI/LLM Workloads

• Batch & Real-Time Processing: ETL/ELT Development, Data Pipeline Orchestration, Data Modeling (OLTP & OLAP), Data Warehousing, Data Integration, Real-time & Batch Processing, Workflow Scheduling (Airflow, Luigi, Oozie, Azkaban)

• Cloud Platforms & Data Services: AWS (S3, Redshift, Glue, EMR, Lambda, Athena, Kinesis), Azure (Data Lake, Synapse Analytics, Data Factory, Databricks), GCP (BigQuery, Dataflow, Dataproc, Pub/Sub), Snowflake, Databricks, Terraform Cloud

• Real-Time & Messaging Systems: Apache Kafka, Flink, Storm, Kinesis, Pub/Sub

• DevOps & Infrastructure as Code (IaC): CI/CD Pipelines (Jenkins, GitLab CI, GitHub Actions, Azure DevOps), Docker, Kubernetes, Terraform, Ansible, Infrastructure as Code

(IaC), Serverless Architectures, Monitoring & Logging (Datadog, Prometheus, Grafana, Splunk, ELK Stack)

• Governance, Security & Compliance: Data Quality Frameworks, Data Cataloging, Metadata Management, Master Data Management (MDM), Data Lineage, GDPR, HIPAA, SOC 2, Role- Based Access Control (RBAC), Data Masking & Encryption PROFESSIONAL EXPERIENCE

Lead Data Engineer 12/2022 - Present

NICE

• Led a team of Data Engineers to build scalable cloud-native data platforms supporting enterprise analytics and AI initiatives.

• Designed and optimized batch and real-time data pipelines using Spark, Kafka, Python, SQL, Airflow, Snowflake and Databricks.

• Architected a modern Lakehouse platform to streamline data ingestion, transformation and analytics across business domains.

• Collaborated with product, engineering and business teams to deliver reliable, scalable and business-driven data solutions.

• Implemented enterprise data quality, lineage and observability frameworks, significantly improving platform reliability.

• Optimized Spark workloads, Snowflake performance and SQL processing to reduce execution time and cloud costs.

• Mentored engineers through code reviews, architecture guidance and Agile best practices while driving technical excellence.

• Automated deployments using CI/CD, Terraform, Docker and Kubernetes, improving release efficiency and platform stability.

• Established secure data governance with RBAC, encryption and compliance aligned with HIPAA, GDPR and SOC 2.

• Built reusable data services enabling real-time analytics, executive reporting and machine learning workloads.

Senior Data Engineer 08/2019 - 11/2022

Slickdeals

• Developed scalable Spark and Python data pipelines to process large-scale transactional and user behavior data for analytics and reporting.

• Built reliable ETL/ELT workflows using Airflow, SQL and Snowflake, improving data availability across business teams.

• Enhanced Snowflake performance through query optimization, clustering and warehouse tuning, reducing compute costs.

• Implemented data lineage, metadata tracking and dependency mapping to improve pipeline transparency and impact analysis.

• Designed automated data quality validation frameworks to ensure accurate and trusted datasets for BI and reporting.

• Collaborated with engineering, analytics and product teams to deliver scalable data solutions supporting business growth.

• Optimized batch processing and distributed workloads using Apache Spark, improving pipeline efficiency and reliability.

• Mentored junior engineers through code reviews, technical guidance, and Agile development best practices.

• Automated deployment and workflow monitoring using CI/CD and Airflow, improving release stability and operational efficiency.

ETL & Data Engineer 06/2017 - 07/2019

Foley

• Developed and optimized ETL pipelines using Hadoop, SQL and Python to process compliance, risk and operational data from multiple sources.

• Designed and maintained data warehouse models, schemas and SQL objects to support enterprise reporting and business analytics.

• Improved batch processing performance and workflow scheduling, increasing data delivery efficiency and meeting critical SLAs.

• Built reporting views, stored procedures and aggregated datasets for Tableau and Power BI dashboards used by business teams.

• Collaborated with analysts and stakeholders to translate business requirements into scalable ETL and data integration solutions.

• Performed data validation, cleansing and transformation to improve data quality, consistency, and reporting accuracy.

• Supported production ETL workflows by monitoring jobs, troubleshooting issues and implementing performance improvements.

PROJECTS

Real-Time Fraud Detection System

• Engineered a real-time fraud detection pipeline to ingest transaction streams, process events and identify anomalous activity using streaming analytics.

• Integrated Apache Kafka with Spark Structured Streaming to enable low-latency processing of transactional data.

• Built a scalable, containerized architecture to support real-time event processing and model deployment workflows.

Tech Stack: Apache Kafka, Spark Structured Streaming, Python, Machine Learning, Docker

Customer Churn Analytics Platform

• Developed a customer churn analytics pipeline to ingest, clean and analyze customer behavior data for retention insights and reporting.

• Built automated ETL workflows to transform raw datasets into analytics-ready datasets supporting churn analysis.

• Performed feature engineering and exploratory analysis to identify key drivers of customer attrition.

• Created scalable data processing workflows to support business intelligence and reporting initiatives.

Tech Stack: Python, SQL, Pandas, ETL, Data Analytics

Healthcare Data Analytics Platform

• Built a healthcare-focused analytics pipeline to process, standardize and transform clinical and claims datasets for downstream reporting and analysis.

• Implemented data modeling and transformation logic to improve data quality and analytical usability.

• Automated ingestion and transformation workflows supporting scalable healthcare data processing.

Tech Stack: Python, SQL, PostgreSQL, ETL, Data Modeling EDUCATION

• Bachelor of Science in Computer Science



Contact this candidate