Post Job Free
Sign in

Principal Data Architect (GCP, Data Governance)

Location:
Chicago, IL, 60628
Salary:
150000
Posted:
August 31, 2026

Contact this candidate

Resume:

Terron Sims

+1-312-***-**** ***********@*****.*** Chicago, IL LinkedIn

SUMMARY

Principal Data Architect with deep experience designing and leading large-scale data platforms on GCP, including enterprise cloud migrations, self-serve data architectures, and data governance frameworks. I have led multiple PB-scale migrations from on-premise and AWS to GCP, built cross-functional engineering teams in Agile and DevOps environments, and driven strategic architecture for data ecosystems. I focus on data security, infrastructure as code, and enabling analytics and AI through governed, self-serve data platforms. Excellent writing and documentation skills. Collaborative and strategic leader comfortable with executive stakeholder management and complex technical roadmaps.

EDUCATION

Bachelor’s Degree of Computer Science 05/2016

University of North Carolina

EXPERIENCE

Principal Data Architect 01/2022 - 05/2026

Membersy Remote

• Led the architecture and migration of a PB-scale healthcare data ecosystem from on-premise to GCP, using BigQuery, BigLake, and Google Cloud Storage as the core data platform.

• Designed self-serve data architectures with Dataplex for data mesh governance, enabling product teams to discover and access trusted datasets autonomously.

• Architected streaming pipelines with Pub/Sub and Confluent Kafka to ingest real-time clinical and claims events, integrated with Dataflow for low-latency processing and analytics.

• Implemented data governance and security frameworks including IAM, VPC Service Controls, data masking, and encryption for HIPAA-compliant environments.

• Built infrastructure as code using Terraform and Pulumi to provision GCP resources including GKE clusters, Cloud Composer environments, and BigQuery datasets.

• Led cross-functional engineering teams in an Agile and DevOps environment, establishing CI/CD pipelines with Cloud Build and GitLab for repeatable deployments.

• Developed scalable batch and stream processing jobs with Apache Beam on Dataflow and Dataproc with Spark and Hadoop for large-scale data transformations.

• Created semantic layers and curated datasets in BigQuery using dbt, enabling self-serve analytics and ML feature engineering for data scientists and analysts.

• Designed and deployed Looker dashboards for operational KPIs and executive reporting, translating business requirements into governed data products.

• Mentored senior engineers on GCP best practices, data modeling strategies, and data governance principles, fostering a culture of ownership and quality.

• Partnered with executive stakeholders to define the strategic roadmap for cloud migration, data platform modernization, and self-serve analytics capabilities.

• Automated data quality monitoring using BigQuery and custom validation scripts, ensuring freshness and accuracy across all datasets.

• Migrated key analytics workloads from Snowflake and Redshift to BigQuery, optimizing cost and performance through clustering, partitioning, and slot management.

• Integrated Vertex AI and BigQuery ML for predictive modeling, enabling real-time inference on healthcare outcomes directly from governed data.

• Established data contracts and lineage documentation for all critical datasets, improving cross-team collaboration and reducing data ambiguity.

• Drove adoption of infrastructure as code for all data platform components, reducing deployment times and increasing reliability.

• Collaborated with data science teams to build feature stores on BigLake, supporting both batch and real-time ML model training.

• Managed cloud cost optimization across GCP services, implementing budget alerts and resource right-sizing for sustainable operations.

• Conducted technical due diligence for potential acquisitions, evaluating data architectures and migration feasibility for integration into the GCP platform. Data Architect 02/2017 - 12/2021

Devbridge Chicago, IL

• Architected and led the migration of a PB-scale e-commerce data warehouse from on-premise Hadoop to GCP BigQuery, using Dataflow and Dataproc for transformation pipelines.

• Designed self-serve data platforms on BigQuery with dbt for transformation and Looker for visualization, enabling marketing and product teams to self-serve analytics.

• Implemented streaming data ingestion with Pub/Sub and Confluent Kafka for real-time order and payment events, processed through Apache Beam on Dataflow.

• Established data governance frameworks including IAM policies, VPC Service Controls, and data masking for customer PII, ensuring compliance with privacy regulations.

• Automated infrastructure provisioning using Terraform for BigQuery, Google Cloud Storage, and Cloud Composer environments in a DevOps pipeline.

• Orchestrated complex ETL workflows with Cloud Composer and Airflow, managing dependencies and monitoring pipeline health with automated alerts.

• Led cross-functional engineering teams in an Agile environment, coordinating sprints and ensuring delivery of data platform milestones.

• Developed batch processing pipelines with Apache Spark on Dataproc for historical data loads and large-scale aggregations.

• Built data quality and reconciliation frameworks using Great Expectations and custom SQL validations, capturing lineage and freshness metadata.

• Created Looker dashboards for revenue, customer segmentation, and supply chain metrics, translating business requirements into actionable insights.

• Integrated third-party SaaS data via REST and GraphQL APIs into BigQuery, unifying CRM, billing, and support datasets for executive reporting.

• Implemented cost optimization strategies for BigQuery including clustering, partitioning, and materialized views, reducing query spend by a significant margin.

• Collaborated with data science teams to prepare feature datasets for churn prediction and recommendation models using Vertex AI.

• Documented data architectures, lineage diagrams, and operational runbooks for team onboarding and knowledge transfer.

• Managed cloud migration from AWS to GCP, including re-architecting storage and compute layers for improved performance and cost.

• Developed infrastructure as code for GKE clusters and Kubernetes deployments, supporting containerized data applications and microservices.

• Conducted technical interviews and vendor evaluations for data tools and cloud services, influencing platform decisions.

• Troubleshot production pipeline failures and data quality issues, performing root cause analysis and implementing corrective measures.

• Supported executive stakeholder presentations with data-driven insights and migration progress updates. Data Engineer 05/2016 - 12/2016

RedHat Chicago, IL

• Built ETL pipelines using Apache Spark and Hadoop for processing business and operational data, supporting real-time analytics dashboards.

• Integrated multiple data sources including SQL Server, logs, and SaaS systems into a centralized data warehouse for decision support.

• Developed and maintained data storage solutions using Google Cloud Storage and BigQuery, enabling scalable analytics workloads.

• Collaborated with cross-functional teams to ensure high-quality data for reporting, improving accuracy of insights for business units.

• Automated data ingestion from external systems into the data platform, reducing manual efforts and improving data freshness.

• Contributed to data-driven projects in a startup environment, working with product and engineering teams to enhance user experience through analytics.

• Created real-time dashboards using Tableau and Looker to monitor system performance and business KPIs.

• Supported analytics initiatives by providing expertise in data modeling and SQL query optimization.

• Documented data schemas and pipeline workflows, enabling team knowledge sharing and smoother onboarding.

• Participated in Agile sprints, contributing to code reviews and continuous improvement of data engineering practices.

SKILLS

GCP Data & Analytics: BigQuery, BigLake, BigQuery Omni, Google Cloud Storage, Dataflow, Apache Beam, Dataproc, Spark, Hadoop, Cloud Composer, Airflow, Pub/Sub, Confluent, Kafka, Looker, Vertex AI, BigQuery ML, dbt, GKE, Kubernetes, Dataplex

Infrastructure & DevOps: Terraform, Pulumi, Cloud Build, GitLab CI, Jenkins, Docker, IAM, VPC Service Controls, Data Masking, Encryption

Programming & Scripting: Python, PySpark, SQL, Java, Scala, Bash, pandas Data Governance & Security: data masking, encryption, IAM, VPC Service Controls, Dataplex, data lineage, data contracts, RBAC, HIPAA, PII governance

Cloud Migration & Architecture: enterprise-scale migrations, PB-scale, AWS to GCP migration, self-serve data platforms, infrastructure as code

Databases & Storage: PostgreSQL, MySQL, Snowflake, Redshift, SQL Server, Google Cloud Storage, HDFS Visualization & BI: Looker, Tableau, Power BI, Metabase



Contact this candidate