Post Job Free
Sign in

Senior Big Data Engineer

Location:
Phoenix, AZ
Posted:
October 07, 2026

Contact this candidate

Resume:

Deepak G

Senior Big Data Engineer

LinkedIn phone:510-***-**** ************@*****.***

PROFESSIONAL SUMMARY

Senior Data Engineer with 10+ years of expertise designing, developing, and maintaining robust ETL/ELT pipelines, data warehouses, and data lakes for large-scale analytics and reporting. Skilled in SQL, Python, Apache Spark, Databricks, Snowflake, and cloud platforms including AWS, Azure, and GCP.

Designed end-to-end ETL/ELT pipeline development workflows using Apache Airflow, Dagster, and Prefect for data orchestration, integrating Python-based pipeline coding to automate data ingestion into BigQuery while applying data cleaning and data structuring routines aligned with requirements gathering outcomes.

Built scalable data transformation pipelines using Apache Spark and Apache Flink on Dataproc, consuming real-time events from Pub/Sub through Dataflow and persisting enriched records into BigQuery tables structured according to data modeling conventions established during business requirements gathering sessions.

Developed data warehousing strategies on BigQuery backed by Google Cloud Storage (GCS), executing SQL-based transformations and maintaining detailed source-to-target mapping documentation using Talend that guided the full ETL/ELT pipeline development lifecycle, ensuring complete traceability across all pipeline stages.

Configured automation using Jenkins, GitHub Actions, and GitLab CI/CD to deploy pipeline coding artifacts stored across Git, GitHub, and GitLab, containerizing data transformation workloads with Docker and orchestrating services through Kubernetes while provisioning infrastructure via Terraform and Puppet.

Delivered BI and dashboard reporting solutions using Power BI and Tableau connected to BigQuery data marts, building interactive dashboards that visualized metrics derived from data orchestration pipelines orchestrated through Apache Airflow and Dagster, directly supporting business intelligence and business requirements gathering outcomes.

Constructed data modeling artifacts using ERwin, ER/Studio, and Power Designer supporting data warehousing architectures, while maintaining metadata governance through Collibra, Alation, and Apache Atlas to ensure catalog accuracy and business intelligence alignment, tracking all tasks across Jira, Trello, and Confluence.

Monitored pipeline health across Prefect and Apache Flink deployments using the ELK Stack for centralized log aggregation, processing high-volume workloads through Apache Spark on Dataproc while managing Pub/Sub-driven data ingestion and data cleaning flows through Dataflow within the broader data orchestration framework.

Applied Generative AI and AI-assisted development techniques to data engineering workflows, using LLM-based tools to accelerate Python/SQL development, pipeline troubleshooting, data quality rule creation, source-to-target analysis, and technical documentation while maintaining production validation and governance standards.

TECHNICAL SKILLS

●Big Data Ecosystem Tools: Apache Spark, Oozie, Zookeeper, Hadoop, Scala, Impala, Kafka, Nifi, Elastic MapReduce (EMR), Fact and Dimension Tables, Cloudera distribution, Hortonworks Ambari, HDFS, Map Reduce, YARN, PIG, Sqoop, HBase, Hive, Flink, Flume, Cassandra.

●Relational and NOSQL: RDBMS, Oracle, SQL Server, PostgreSQL, DB2, DynamoDB, MongoDB, HBase, Cosmos DB.

●Programming Languages: SQL, PL/SQL, Python, Spark, JavaScript, HTML, SAS, Shell Scripting, Perl.

●ETL and Data Integration: IBM DataStage 11.5/9.1, Informatica, Talend 5.3/6.2, TOAD 10.5, Teradata SQL Assistant, Spring Boot, Databricks, Anaconda, Visual Source Safe, Putty, Notepad++, GitHub, Confluence, Jenkins.

●Cloud Platforms: Amazon Web Services (AWS), GCP, Azure.

●Services: Azure Search, Data Factory, Key Vault, SQL Azure, Azure DevOps, Azure Synapse Analytics (DW), Azure Data Lake, AWS Lambda, Athena, AWS Elastic Compute Cloud (EC2), AWS Glue, Redshift, Kinesis, AWS Virtual Private Cloud (VPC), AWS CloudWatch, Step Functions, BigQuery, Composer, Google Cloud Storage, Cloud Dataproc.

●Visualization and Reporting: MS Excel, Power BI, Tableau, SPSS.

●Containerization, Orchestration Tools: Kubernetes, Docker, Docker Registry, Docker Hub, Docker.

●Methodologies: Agile, Scrum, Waterfall.

●Libraries: Pandas, NumPy, SciPy, Matplotlib, Seaborn, Scikit-Learn.

●Algorithms: Decision Tree, Random Forest, Logistic Regression, Gradient Boosting, Support Vector Machine (SVM), K-Nearest Neighbor (KNN).

●IDE’s & Frameworks: Anaconda Software, Jupyter Notebook, Pycharm, Django, Flask.

●Data Storage: Data Lakes, Data Warehouses, Distributed Storage Systems

●Streaming Technologies: Apache Kafka, Apache Flink, Kinesis, Event Hubs, Pub/Sub

●ETL/ELT Tools: Apache Airflow, Apache NiFi, Talend, Informatica, AWS Glue, Azure Data Factory, Google Dataflow

●Workflow Orchestration: Apache Airflow, Luigi, Prefect, Dagster, AWS Step Functions, Azure Data Factory

●Data Processing: Batch Processing, Real-Time Processing, Stream Processing, Distributed Processing

●Infrastructure & DevOps: Infrastructure as Code, CI/CD, Docker, Kubernetes, Terraform, CloudFormation, Ansible, Jenkins, GitLab CI, GitHub Actions

●Monitoring & Logging: DataDog, Splunk, ELK Stack, CloudWatch, Azure Monitor, Stackdriver, Prometheus, Grafana

●Data Governance: Data Lineage, Data Cataloging, Apache Atlas, Collibra, Alation, AWS Glue Catalog, Azure Purview

●Version Control: Git, GitHub, GitLab, Bitbucket

●Architectures: Event-Driven Architecture, Microservices, Distributed Systems

●Data Formats: Structured Data, Unstructured Data, Parquet, Avro, ORC, JSON, CSV

●Business Intelligence: Analytics Platforms, Reporting Tools

PROFESSIONAL EXPERIENCE

Client: Cardinal Health Remote

Role: Senior Data Engineer February 2025 – Current

Designed and maintained data pipelines within Google Cloud Platform, building end-to-end data engineering workflows that moved raw insurance data through structured ingestion, transformation, and loading stages aligned with enterprise data warehousing standards and business reporting requirements.

Developed end-to-end data pipelines within the insurance domain, applying data transformation logic across structured datasets ingested into BigQuery, ensuring audit-ready change history aligned with organizational regulatory compliance obligations and maintaining schema integrity across all downstream dashboard reporting layers.

Collaborated with BI and dashboard reporting teams to validate that Tableau data sources consuming directly from BigQuery reflected accurate transformation-complete datasets, reviewing requirements against the data warehousing schema and making iterative adjustments to data structuring to resolve reporting discrepancies raised by business users.

Managed Dataproc cluster configurations for Apache Spark batch processing jobs within the insurance data pipeline, tuning executor memory, partition sizing, and shuffle behavior in Python scripts to maintain stable processing performance across high-volume data transformation workloads.

Oversaw job scheduling and dependency management entirely through Apache Airflow data orchestration workflows, ensuring Dataproc-hosted Apache Spark batch processing jobs executed in correct sequence with validated data transformation outputs flowing reliably into BigQuery for downstream Power BI consumption.

Refined executor memory allocations and shuffle behavior configurations in Python-based Apache Spark scripts on Dataproc, reducing processing failures during peak insurance batch runs while maintaining Apache Airflow-driven job scheduling stability and data pipeline dependency management accuracy.

Tracked all data engineering delivery milestones and data pipeline defects within Jira, maintaining a clear project audit trail that connected technical activities in Apache Airflow and BigQuery with insurance business priorities while supporting regulatory compliance reporting obligations.

Managed requirements gathering outcomes and data transformation rule changes within Jira, ensuring all data engineering decisions related to BigQuery schemas and Apache Airflow workflows remained traceable against insurance business objectives and regulatory compliance deliverables throughout the project lifecycle.

Supported dashboard reporting accuracy by validating Tableau data sources against BigQuery-stored data warehousing schemas, coordinating with BI stakeholders during requirements gathering sessions to align data structuring decisions with insurance reporting needs and regulatory compliance documentation standards.

Applied data orchestration best practices through Apache Airflow to coordinate Dataproc-based Apache Spark batch processing jobs, managing dependency management across multi-stage insurance data pipeline workflows and ensuring data transformation outputs were delivered reliably into Big Query for Tableau consumption.

Incorporated GenAI-assisted data engineering into day-to-day pipeline development, using LLM-based tooling to accelerate Python and SQL development, analyse Apache Airflow failures, generate data quality validation logic, and document transformation rules and lineage across BigQuery data pipelines, with all AI-generated outputs reviewed and validated before production deployment

Maintained Jira as the central tracking system for data engineering pipeline defects and delivery milestones, logging data transformation rule changes and data structuring decisions to create an auditable connection between Apache Airflow orchestration activities, BigQuery schema updates, and insurance regulatory compliance requirements.

Environment: Apache Airflow, Python, Git, Jenkins, Power BI Apache Spark, Jira, BigQuery, data orchestration, data transformation, data warehousing, BI and dashboard reporting, requirements gathering, pipeline coding, data ingestion, data structuring, Business Intelligence, Erwin, Collibra, Tablue, Informatica, Google Cloud Storage (GCS), Dataproc, Dataflow, Pub/Sub

Client: Equifax Atlanta, GA

Role: Senior Big Data Engineer September 2022 – January 202

Designed and maintained data pipelines within Google Cloud Platform, building end-to-end data engineering workflows that moved raw insurance data through structured ingestion, transformation, and loading stages aligned with enterprise data warehousing standards and business reporting requirements.

Constructed modular ETL processes using Python scripting to extract transactional insurance records from operational sources, applying field-level cleansing and normalization rules before loading structured datasets into BigQuery for downstream BI consumption and analytical processing.

Developed Apache Airflow DAGs to orchestrate multi-step data pipeline workflows across Google Cloud Platform, defining task dependencies, retry policies, and scheduling intervals that ensured reliable execution of data engineering batch jobs without manual intervention from operations teams.

Built and refined data transformation logic within BigQuery using SQL, translating raw insurance policy and claims records into reporting-ready datasets that satisfied dashboard schema requirements validated collaboratively with BI and Tableau reporting stakeholders during iterative review cycles.

Maintained Dataproc cluster configurations for Apache Spark batch processing jobs, tuning executor memory allocation, partition sizing, and shuffle behavior across Python Spark scripts to sustain stable data pipeline throughput under variable insurance data volume conditions across scheduled processing windows.

Supported data governance practices across the insurance data warehousing environment by documenting data lineage, data quality validation checkpoints, and data transformation rule changes within shared repositories, ensuring that BigQuery datasets met organizational accuracy and completeness standards before being surfaced to reporting layers.

Tracked all data engineering delivery milestones, pipeline defects, and data transformation rule updates within Jira, maintaining a structured project audit trail that connected Apache Airflow orchestration activities and BigQuery schema changes with insurance business priorities and regulatory compliance reporting obligations.

Resolved reporting discrepancies raised by business users by tracing issues back through data transformation logic in BigQuery, adjusting SQL queries and Python processing scripts to correct field mapping errors and ensure Tableau dashboards accurately reflected transformation-complete insurance datasets consumed by BI teams.

Environment: Apache Airflow, Python, Git, Jenkins, Apache Spark, Jira, BigQuery, data orchestration, data transformation, data warehousing, BI and dashboard reporting, requirements gathering, pipeline coding, data ingestion, data structuring, Business Intelligence, Erwin, Collibra, Tablue, Informatica, Google Cloud Storage (GCS), Dataproc, Dataflow, Pub/Sub

Client: Target Minneapolis, MN

Role: Data Engineer November 2019 – August 2022

Built a Snowflake-centered retail analytics platform integrating point-of-sale, e-commerce, inventory, product, customer, and campaign data so merchandising, supply-chain, and marketing teams could work from consistent analytical datasets.

Developed ingestion pipelines using AWS S3/Glue, APIs, database extracts, and file feeds, standardizing structured and semi-structured data with Python/PySpark before publishing governed Snowflake staging and curated layers.

Created dbt staging, intermediate, and mart models for sales, product, inventory, customer, and campaign analytics, applying reusable SQL, incremental processing, tests, documentation, and conformed business dimensions.

Designed dimensional models and analytical marts in Snowflake using fact/dimension tables, SCD patterns, date/product/store/customer dimensions, and reusable measures for sales, fulfillment, inventory, and customer reporting.

Implemented batch and near-real-time ingestion using Kafka, Airflow, and incremental load patterns, supporting timely order, inventory, campaign, and digital-interaction data for downstream analytics.

Optimized Snowflake and SQL workloads through warehouse sizing, clustering/pruning analysis, query refactoring, aggregate tables, and efficient data-model design for high-volume retail reporting.

Built Python/SQL data-quality controls for schema validation, duplicates, nulls, referential integrity, control totals, and source-to-target reconciliation before data was exposed to business users.

Published curated Snowflake datasets to Tableau, Power BI, and Looker, partnering with analysts to validate KPIs, drill-down logic, filters, and self-service reporting requirements.

Implemented Git/Jenkins CI/CD, automated tests, environment-specific configuration, and production monitoring for dbt models, ingestion jobs, and warehouse changes.

Implemented Snowflake role and secure-view patterns for curated customer, sales, and inventory datasets, limiting exposure of sensitive fields while preserving self-service access for approved analytics teams.

Environment: Snowflake, dbt, SQL, Python, PySpark, AWS S3, AWS Glue, Apache Airflow, Apache Kafka, Git, Jenkins, Terraform, Tableau, Power BI, Looker.

Client: Honeywell Charlotte, NC

Role: Data Engineer December 2017 – October 2019

Built an end-to-end manufacturing data ingestion pipeline using Apache Kafka and Amazon Kinesis to capture real-time sensor data from production lines, routing processed streams into Amazon S3 for downstream analytics, reducing data latency by 40%.

Designed a factory floor data lakehouse architecture on Databricks and Delta Lake, consolidating machine telemetry and quality inspection records from multiple plants, enabling unified querying through Spark SQL for production reporting teams.

Orchestrated nightly equipment maintenance ETL workflows using Apache Airflow and AWS Step Functions, pulling structured data from DynamoDB into Amazon Redshift, ensuring supply chain and inventory teams had fresh data every morning.

Developed scalable data transformation logic in Apache Spark and Scala to process high-volume manufacturing defect logs, applying data quality checks with Great Expectations and Deequ before loading results into Snowflake.

Automated data ingestion from third-party ERP and supplier systems using Fivetran and AWS Glue, transforming raw procurement records through dbt models, and surfacing finished datasets in Amazon Redshift for procurement analytics dashboards.

Delivered interactive production efficiency and OEE dashboards in Tableau and Qlik, connecting directly to Snowflake and Elasticsearch data sources, while Mode supported self-service reporting for manufacturing engineers and plant managers.

Streamlined historical manufacturing event processing by building ingestion workflows in Informatica loading cleansed data into Apache Iceberg and Delta Lake tables on Amazon S3, improving traceability for compliance audits.

Client: Tally Solutions Bengaluru, India

Role: ETL Developer April 2016 – October 2017

Developed SQL Server and SSIS/Informatica ETL workflows for ERP accounting, sales, purchase, inventory, customer, vendor, invoice, payment, and ledger data into centralized reporting structures.

Built stored procedures, staging tables, joins, and SQL/Python validation routines for cleansing, duplicate detection, referential checks, reconciliation, and exception reporting.

Supported Hadoop processing and SFTP/file ingestion while maintaining source mappings, control tables, logs, dependencies, and restartable batch execution.

Worked with reporting and application teams to understand requirements, validate source mappings and business rules, support UAT, and resolve ETL/reporting discrepancies.

Performed SQL performance tuning, failed-job recovery, data corrections, testing, release validation, and production troubleshooting for business-critical batch workflows.

Maintained technical documentation, source-to-target mappings, job dependencies, reconciliation procedures, and operational runbooks for finance and inventory reporting

Education:

University: Amritha School of Engineering June 2012 – March 2016

Degree: Bachelor’s in computer science Location – Bangalore, India



Contact this candidate