Terron Sims
+1-312-***-**** ***********@*****.*** Chicago, IL LinkedIn
SUMMARY
Principal Data Architect with deep experience designing and leading large-scale data platforms on GCP, including enterprise cloud migrations, self-serve data architectures, and data governance frameworks. I have led multiple PB-scale migrations from on-premise and AWS to GCP, built cross-functional engineering teams in Agile and DevOps environments, and driven strategic architecture for data ecosystems. I focus on data security, infrastructure as code, and enabling analytics and AI through governed, self-serve data platforms. Excellent writing and documentation skills. Collaborative and strategic leader comfortable with executive stakeholder management and complex technical roadmaps.
EDUCATION
Bachelor’s Degree of Computer Science 05/2016
University of North Carolina
EXPERIENCE
Principal Data Architect 01/2022 - 05/2026
Membersy Remote
• Led the architecture and migration of a PB-scale healthcare data ecosystem from on-premise to GCP, using BigQuery, BigLake, and Google Cloud Storage as the core data platform.
• Designed self-serve data architectures with Dataplex for data mesh governance, enabling product teams to discover and access trusted datasets autonomously.
• Architected streaming pipelines with Pub/Sub and Confluent Kafka to ingest real-time clinical and claims events, integrated with Dataflow for low-latency processing and analytics.
• Implemented data governance and security frameworks including IAM, VPC Service Controls, data masking, and encryption for HIPAA-compliant environments.
• Built infrastructure as code using Terraform and Pulumi to provision GCP resources including GKE clusters, Cloud Composer environments, and BigQuery datasets.
• Led cross-functional engineering teams in an Agile and DevOps environment, establishing CI/CD pipelines with Cloud Build and GitLab for repeatable deployments.
• Developed scalable batch and stream processing jobs with Apache Beam on Dataflow and Dataproc with Spark and Hadoop for large-scale data transformations.
• Created semantic layers and curated datasets in BigQuery using dbt, enabling self-serve analytics and ML feature engineering for data scientists and analysts.
• Designed and deployed Looker dashboards for operational KPIs and executive reporting, translating business requirements into governed data products.
• Mentored senior engineers on GCP best practices, data modeling strategies, and data governance principles, fostering a culture of ownership and quality.
• Partnered with executive stakeholders to define the strategic roadmap for cloud migration, data platform modernization, and self-serve analytics capabilities.
• Automated data quality monitoring using BigQuery and custom validation scripts, ensuring freshness and accuracy across all datasets.
• Migrated key analytics workloads from Snowflake and Redshift to BigQuery, optimizing cost and performance through clustering, partitioning, and slot management.
• Integrated Vertex AI and BigQuery ML for predictive modeling, enabling real-time inference on healthcare outcomes directly from governed data.
• Established data contracts and lineage documentation for all critical datasets, improving cross-team collaboration and reducing data ambiguity.
• Drove adoption of infrastructure as code for all data platform components, reducing deployment times and increasing reliability.
• Collaborated with data science teams to build feature stores on BigLake, supporting both batch and real-time ML model training.
• Managed cloud cost optimization across GCP services, implementing budget alerts and resource right-sizing for sustainable operations.
• Conducted technical due diligence for potential acquisitions, evaluating data architectures and migration feasibility for integration into the GCP platform. Data Architect 02/2017 - 12/2021
Devbridge Chicago, IL
• Architected and led the migration of a PB-scale e-commerce data warehouse from on-premise Hadoop to GCP BigQuery, using Dataflow and Dataproc for transformation pipelines.
• Designed self-serve data platforms on BigQuery with dbt for transformation and Looker for visualization, enabling marketing and product teams to self-serve analytics.
• Implemented streaming data ingestion with Pub/Sub and Confluent Kafka for real-time order and payment events, processed through Apache Beam on Dataflow.
• Established data governance frameworks including IAM policies, VPC Service Controls, and data masking for customer PII, ensuring compliance with privacy regulations.
• Automated infrastructure provisioning using Terraform for BigQuery, Google Cloud Storage, and Cloud Composer environments in a DevOps pipeline.
• Orchestrated complex ETL workflows with Cloud Composer and Airflow, managing dependencies and monitoring pipeline health with automated alerts.
• Led cross-functional engineering teams in an Agile environment, coordinating sprints and ensuring delivery of data platform milestones.
• Developed batch processing pipelines with Apache Spark on Dataproc for historical data loads and large-scale aggregations.
• Built data quality and reconciliation frameworks using Great Expectations and custom SQL validations, capturing lineage and freshness metadata.
• Created Looker dashboards for revenue, customer segmentation, and supply chain metrics, translating business requirements into actionable insights.
• Integrated third-party SaaS data via REST and GraphQL APIs into BigQuery, unifying CRM, billing, and support datasets for executive reporting.
• Implemented cost optimization strategies for BigQuery including clustering, partitioning, and materialized views, reducing query spend by a significant margin.
• Collaborated with data science teams to prepare feature datasets for churn prediction and recommendation models using Vertex AI.
• Documented data architectures, lineage diagrams, and operational runbooks for team onboarding and knowledge transfer.
• Managed cloud migration from AWS to GCP, including re-architecting storage and compute layers for improved performance and cost.
• Developed infrastructure as code for GKE clusters and Kubernetes deployments, supporting containerized data applications and microservices.
• Conducted technical interviews and vendor evaluations for data tools and cloud services, influencing platform decisions.
• Troubleshot production pipeline failures and data quality issues, performing root cause analysis and implementing corrective measures.
• Supported executive stakeholder presentations with data-driven insights and migration progress updates. Data Engineer 05/2016 - 12/2016
RedHat Chicago, IL
• Built ETL pipelines using Apache Spark and Hadoop for processing business and operational data, supporting real-time analytics dashboards.
• Integrated multiple data sources including SQL Server, logs, and SaaS systems into a centralized data warehouse for decision support.
• Developed and maintained data storage solutions using Google Cloud Storage and BigQuery, enabling scalable analytics workloads.
• Collaborated with cross-functional teams to ensure high-quality data for reporting, improving accuracy of insights for business units.
• Automated data ingestion from external systems into the data platform, reducing manual efforts and improving data freshness.
• Contributed to data-driven projects in a startup environment, working with product and engineering teams to enhance user experience through analytics.
• Created real-time dashboards using Tableau and Looker to monitor system performance and business KPIs.
• Supported analytics initiatives by providing expertise in data modeling and SQL query optimization.
• Documented data schemas and pipeline workflows, enabling team knowledge sharing and smoother onboarding.
• Participated in Agile sprints, contributing to code reviews and continuous improvement of data engineering practices.
SKILLS
GCP Data & Analytics: BigQuery, BigLake, BigQuery Omni, Google Cloud Storage, Dataflow, Apache Beam, Dataproc, Spark, Hadoop, Cloud Composer, Airflow, Pub/Sub, Confluent, Kafka, Looker, Vertex AI, BigQuery ML, dbt, GKE, Kubernetes, Dataplex
Infrastructure & DevOps: Terraform, Pulumi, Cloud Build, GitLab CI, Jenkins, Docker, IAM, VPC Service Controls, Data Masking, Encryption
Programming & Scripting: Python, PySpark, SQL, Java, Scala, Bash, pandas Data Governance & Security: data masking, encryption, IAM, VPC Service Controls, Dataplex, data lineage, data contracts, RBAC, HIPAA, PII governance
Cloud Migration & Architecture: enterprise-scale migrations, PB-scale, AWS to GCP migration, self-serve data platforms, infrastructure as code
Databases & Storage: PostgreSQL, MySQL, Snowflake, Redshift, SQL Server, Google Cloud Storage, HDFS Visualization & BI: Looker, Tableau, Power BI, Metabase