Post Job Free
Sign in

Principal Data Architect, Cloud Data Ecosystems

Location:
Leander, TX, 78641
Salary:
150000
Posted:
August 31, 2026

Contact this candidate

Resume:

Fernando Mosqueda

Leander, TX **************@*****.*** 伐+1-254-***-****

SUMMARY

A hands-on data architect with over 18 years in data engineering, warehousing, and business intelligence, focused on designing large-scale data ecosystems and driving cloud migrations. Skilled in building self-serve data platforms from the ground up with rigorous governance and security. Experienced leading cross-functional teams in Agile/DevOps environments, shaping long-term platform roadmaps, and advising executive stakeholders on data strategy, ROI, and technical risk. Known for excellent writing and documentation skills and collaborative problem-solving. EDUCATION

Sam Houston State University Sep 2004 - Apr 2010

Bachelor of Science, Computer Science GPA: 3.8

LICENSES & CERTIFICATIONS

• Databricks - Data Ingestion with Lakeflow Connect

• Databricks - Generative AI Fundamentals

• Data Analysis with Python

• Professional Fundamentals of Digital Marketing

• Event-Driven Microservices

• Project Management Professional (PMP)®

WORK EXPERIENCE

McKesson

Principal Data Architect Sep 2022 - Mar 2026

• Led the design and implementation of a large-scale data ecosystem on GCP, migrating PB-scale healthcare data from AWS to BigQuery, BigLake, and Google Cloud Storage with zero data loss.

• Architected a self-serve data platform using Dataplex for automated governance, IAM, VPC Service Controls, and data masking/encryption, enabling decentralized teams to access governed datasets.

• Built version-controlled infrastructure with Terraform and Pulumi, deploying repeatable, scalable environments for data pipelines and analytics across multiple GCP projects.

• Designed end-to-end streaming and batch pipelines using Dataflow (Apache Beam), Dataproc (Spark/Hadoop), and Cloud Composer (Airflow), processing real-time clinical and operational data.

• Established a long-term data platform roadmap aligned with global business goals, presenting North Star data strategies to the VP of Data and securing executive buy-in for key initiatives.

• Managed technical risk and ROI of data initiatives, evaluating trade-offs between cost, performance, and security for large-scale cloud migrations and new platform features.

• Led cross-functional engineering teams in an Agile/DevOps environment, mentoring engineers on GCP best practices, data modeling, and infrastructure as code.

• Integrated Confluent/Kafka with Pub/Sub for event-driven architecture, enabling low-latency ingestion of streaming data from healthcare devices and applications.

• Implemented BigQuery Omni for cross-cloud analytics, allowing querying of data across AWS and GCP without moving data, reducing egress costs and complexity.

• Developed data quality frameworks with automated validation and anomaly detection using Dataplex and custom Python scripts, ensuring high trust in business intelligence reports.

• Partnered with data scientists to deploy machine learning models using Vertex AI and BigQuery ML, supporting predictive analytics for patient outcomes and operational efficiency.

• Created executive dashboards in Looker for real-time monitoring of data platform health, migration progress, and business KPIs, facilitating data-driven decision-making.

• Defined data security frameworks including encryption at rest and in transit, RBAC, and audit logging, ensuring compliance with HIPAA and internal privacy protocols.

• Orchestrated end-to-end data pipelines with Cloud Composer, integrating dbt for transformation and Airflow for scheduling, enabling reliable and maintainable data workflows.

• Managed containerized data processing jobs on GKE/Kubernetes, optimizing resource utilization and scaling for peak workloads during migration and production operations.

• Provided strategic architecture guidance for warehouse modernization, moving from legacy on-premise systems to a modern data stack on GCP with self-serve analytics.

• Conducted architectural reviews and technical risk assessments for new data initiatives, ensuring alignment with enterprise standards and long-term scalability.

• Fostered a culture of collaboration and continuous improvement, facilitating cross-team workshops on data governance, security best practices, and Agile delivery.

• Documented data architectures, lineage diagrams, and operational runbooks, enhancing team knowledge sharing and reducing onboarding time for new engineers.

• Championed the adoption of Infrastructure as Code across the organization, standardizing on Terraform for all GCP resource provisioning and environment management.

• Collaborated with product managers and business stakeholders to translate strategic goals into technical roadmaps, balancing innovation with operational stability.

• Implemented automated testing for data pipelines and infrastructure changes, integrating CI/CD pipelines with GitHub Actions and Jenkins for reliable deployments.

• Researched and evaluated emerging GCP services (e.g., BigLake, Omni) to enhance platform capabilities, presenting findings to leadership with cost-benefit analysis.

• Spearheaded the migration of legacy data warehouses to BigQuery, optimizing query performance through partitioning, clustering, and materialized views.

• Established data cataloging and lineage tracking using Dataplex, improving data discoverability and governance for hundreds of datasets across the organization.

Xvand

Data Architect Jun 2015 - Jul 2022

• Designed and scaled a large-scale data ecosystem on GCP for e-commerce, processing petabytes of customer and transaction data using BigQuery, Dataflow, and Pub/Sub.

• Built self-serve data platforms enabling marketing and product teams to access governed datasets through Looker dashboards and custom data models.

• Implemented end-to-end cloud migration from on-premise Hadoop clusters to GCP, using Dataproc for Spark jobs and Cloud Storage for data lake storage.

• Architected streaming pipelines with Apache Kafka and Pub/Sub for real-time ingestion of order, inventory, and customer behavior data, supporting personalization engines.

• Developed data transformation workflows using Apache Beam and dbt, ensuring data quality and consistency for business intelligence and advanced analytics.

• Applied data governance practices using IAM, VPC Service Controls, and encryption, protecting sensitive customer information and meeting privacy regulations.

• Automated infrastructure provisioning with Terraform and Pulumi, enabling repeatable, version-controlled environments for development, staging, and production.

• Optimized BigQuery performance through partitioning, clustering, and query tuning, reducing latency for complex analytical queries used by cross-functional teams.

• Integrated Looker with BigQuery for real-time visualizations of sales, marketing, and operational metrics, empowering business users to make data-driven decisions.

• Collaborated with data scientists to deliver feature datasets for machine learning models using Vertex AI and BigQuery ML, improving product recommendations and forecasting.

• Managed containerized data processing workloads on GKE/Kubernetes, improving resource utilization and scalability for batch and streaming jobs.

• Developed data quality monitoring frameworks using Dataplex and custom alerts, detecting anomalies and ensuring data integrity for critical reporting.

• Documented data architectures, pipeline designs, and governance policies, supporting knowledge transfer and operational continuity across teams.

• Implemented CI/CD pipelines with GitLab CI and Cloud Build, automating deployment of data pipelines and infrastructure changes.

• Evaluated new GCP services and open-source tools, providing recommendations for modernizing the data stack and reducing operational costs.

• Mentored junior engineers on GCP best practices, data modeling, and DevOps principles, fostering a culture of learning and collaboration.

• Resolved complex production issues through systematic root cause analysis, maintaining data availability and performance within defined SLAs.

Total IT

Data Engineer Jun 2010 - Apr 2015

• Designed and built ETL pipelines using SQL Server Integration Services and Python for processing business data, enabling reporting and analytics for operational teams.

• Integrated data from multiple sources including CRM, ERP, and log files into a centralized data warehouse on Azure SQL Database and Blob Storage.

• Developed data models for star schema and dimensional modeling, supporting business intelligence dashboards in Power BI and Tableau.

• Collaborated with cross-functional teams to understand data requirements and deliver accurate, timely datasets for decision-making.

• Automated data ingestion processes using batch scripts and scheduled jobs, reducing manual effort and improving data freshness.

• Maintained data quality by implementing validation rules and error handling in ETL workflows, ensuring reliable reporting.

• Supported the migration of legacy on-premise databases to cloud-based Azure SQL and Snowflake, gaining early cloud experience.

• Contributed to a startup environment by building data pipelines from scratch, adapting to evolving business needs and tight deadlines.

• Created documentation for data sources, transformations, and reporting structures, enabling team knowledge sharing.

• Participated in Agile sprints, delivering data solutions that aligned with product roadmaps and business priorities. SKILLS

GCP Data Stack: BigQuery (BigLake, Omni), Google Cloud Storage, Dataflow (Apache Beam), Dataproc

(Spark/Hadoop), Cloud Composer (Airflow), Pub/Sub, Confluent/Kafka, Looker, Vertex AI, BigQuery ML, dbt, GKE/Kubernetes, IAM, VPC Service Controls, Data Masking/Encryption, Dataplex, Terraform, Pulumi, Infrastructure as Code

Cloud Platforms: AWS (Redshift, S3, Kinesis), Azure (Data Lake, Synapse, SQL) Databases: PostgreSQL, MySQL, MongoDB, DynamoDB, Snowflake, SQL Server Data Engineering: ETL/ELT pipelines, data warehousing, business intelligence, data governance, data security, data architecture, cloud migration, modern data stack, self-serve data platform Programming & Scripting: Python (Pandas, PySpark, Beam), SQL, R, Bash Big Data & Streaming: Kafka, Spark, Hadoop, Beam, Kinesis DevOps & CI/CD: Terraform, Pulumi, GitHub Actions, Jenkins, Cloud Build, Docker, Kubernetes, Git BI & Visualization: Looker, Power BI, Tableau

Security & Compliance: HIPAA, GDPR, RBAC, Encryption, Audit Logging Soft Skills: Executive stakeholder management, technical risk management, strategic architecture, collaboration, problem-solving, mentoring, Agile/DevOps leadership



Contact this candidate