Sandeep Sreekumar
+1-331-***-**** ******************@*****.*** Chicago, IL, USA
SUMMARY
Principal Data Architect with over 10 years of experience designing and building scalable data ecosystems, cloud data platforms, and enterprise analytics solutions. Experienced in leading large-scale cloud migrations, self-serve data platform design, and data governance and security frameworks. Strong background in GCP data stack including BigQuery, Dataflow, Dataproc, Cloud Composer, and Pub/Sub with proven ability to lead cross-functional engineering teams in Agile and DevOps environments. Recognized for excellent writing and documentation skills, strategic thinking, and collaboration with business and technical stakeholders to deliver data-driven solutions.
EDUCATION
Master’sDegree,Computer Science GPA: 3.9 Sep 2011 - Jun 2013 Illinois Institute of Technology
EXPERIENCE
Principal Data Architect Jan 2023 - Jun 2026
Oracle Health
• Architected and led the migration of enterprise-scale PB-scale data workloads from on-premise and AWS environments to GCP, using BigQuery, Cloud Storage, and Dataproc for scalable analytics.
• Designed and implemented a self-serve data platform on GCP using BigQuery, BigLake, and dbt, enabling business teams to access curated datasets with governance and security controls.
• Built real-time and batch data pipelines using Dataflow, Apache Beam, Cloud Composer, and Pub/Sub to process large volumes of healthcare, clinical, and operational data.
• Orchestrated complex data workflows using Cloud Composer and Airflow, automating pipeline dependencies, monitoring, and SLA management for production-critical systems.
• Implemented data governance frameworks using Dataplex, IAM, VPC Service Controls, data masking, and encryption to ensure compliance with HIPAA and other regulatory requirements.
• Developed containerized data processing solutions using GKE and Kubernetes, improving deployment consistency and resource utilization across data platform components.
• Led cross-functional engineering teams in an Agile/DevOps environment, establishing sprint planning, code reviews, CI/CD pipelines, and operational runbooks for data platform releases.
• Designed and deployed Vertex AI and BigQuery ML solutions for predictive analytics, including feature engineering pipelines and ML model training datasets for clinical and operational use cases.
• Optimized BigQuery performance through partitioning, clustering, materialized views, slot management, and query tuning, reducing costs and improving query response times substantially.
• Implemented streaming data ingestion from Confluent Kafka and Pub/Sub into BigQuery and BigLake, enabling near-real-time analytics on patient monitoring and operational events.
• Built data quality and observability solutions using automated validation checks, anomaly detection, and lineage tracking to ensure data reliability and trustworthiness.
• Established data cataloging and metadata management practices using Dataplex, enabling data discovery, business glossary, and lineage for enterprise data assets.
• Designed and implemented security controls including data masking, encryption at rest and in transit, and VPC Service Controls to protect sensitive healthcare data.
• Collaborated with data scientists, analysts, and business stakeholders to define data platform roadmaps, prioritize features, and align technical architecture with business strategy.
• Provided technical leadership and mentorship to data engineers, conducting code reviews, architectural guidance, and training on GCP data services and best practices.
• Owned the end-to-end lifecycle of data platform components, from architecture design and development through deployment, monitoring, and continuous improvement.
• Developed reusable data transformation frameworks using Python, PySpark, and SQL on Dataproc and Dataflow, standardizing ETL/ELT patterns across the organization.
• Integrated with external systems via REST APIs, event-driven architectures, and Pub/Sub topics, enabling seamless data exchange with third-party applications and partners. Senior Data Engineer Jul 2015 - Dec 2022
Accenture
• Designed and implemented enterprise-scale data warehouses and lakehouse architectures on GCP using BigQuery, Cloud Storage, Dataproc, and Dataflow for large-scale analytics across finance, retail, and e-commerce domains.
• Led migration of on-premise data warehouses to GCP, including schema design, data transfer using Cloud Storage and Dataflow, and validation of petabyte-scale datasets.
• Built scalable ETL/ELT pipelines using Python, PySpark, and SQL on Dataproc and Cloud Composer, processing structured and unstructured data from CRM, ERP, and transactional systems.
• Developed self-serve analytics platforms using Looker and dbt on BigQuery, enabling business users to create reports and dashboards with governed data access.
• Implemented data governance and security controls including IAM roles, VPC Service Controls, data masking, and encryption for compliance with GDPR and financial regulations.
• Optimized BigQuery performance through partitioning, clustering, and slot management, reducing query costs and improving report generation times for business intelligence teams.
• Integrated real-time data streams from Kafka and Pub/Sub into BigQuery using Dataflow, supporting near-real-time dashboards and operational analytics.
• Designed and deployed containerized data processing workflows using Docker and GKE, improving scalability and portability of data pipeline components.
• Collaborated with data scientists and ML engineers to prepare feature engineering datasets and deploy BigQuery ML models for predictive analytics in retail demand forecasting.
• Created architecture documentation, data flow diagrams, and operational runbooks to support knowledge transfer and platform governance.
• Automated database administration tasks using scripting and process improvements, reducing manual effort and improving data availability.
• Built data validation and reconciliation frameworks to ensure data accuracy across multiple source systems.
• Developed interactive dashboards in Tableau and Power BI for executive reporting on key business metrics.
• Implemented CI/CD pipelines using Git, Jenkins, and Terraform for automated deployment of data platform configurations.
• Provided production support and incident troubleshooting for mission-critical data pipelines, ensuring SLA adherence.
Data Engineer Sep 2013 - Jun 2015
Nielsen
• Designed and implemented data warehousing solutions using BigQuery, Cloud Storage, and SQL Server, enabling analytics and reporting for media and advertising clients.
• Integrated data from multiple sources including transactional databases, flat files, and APIs into centralized data repositories on AWS and GCP.
• Built ETL pipelines using Python and SQL to automate data ingestion and transformation, improving data availability for business intelligence teams.
• Optimized database performance through query tuning, indexing, and partitioning on BigQuery and SQL Server, reducing query execution times.
• Developed data quality checks and validation processes to ensure accuracy and consistency of reporting datasets.
• Collaborated with analysts and business users to define data requirements and deliver dashboards and reports for operational decision-making.
Data Engineer Feb 2013 - Sep 2013
Civis Analytics
• Assisted in preprocessing and cleansing text data for publishing and media analytics, preparing structured datasets for business reporting.
• Supported language analytics by extracting text features and performing SQL-based analysis on customer feedback data.
• Worked in a startup environment to maintain data processing workflows, validate incoming data, and generate reports for marketing and audience engagement.
• Performed data quality checks, documentation, and dataset validation using SQL, ensuring reliable inputs for analytics and decision-making.
SKILLS
Cloud & Data Engineering: GCP, BigQuery, BigLake, Google Cloud Storage, Dataflow, Apache Beam, Dataproc, Spark, Hadoop, Cloud Composer, Airflow, Pub/Sub, Confluent Kafka, Docker, Kubernetes, GKE, AWS, Azure, Snowflake, PostgreSQL, MySQL, SQL Server.
Data Governance & Security: Dataplex, IAM, VPC Service Controls, Data Masking, Encryption, Data Lineage, RBAC, HIPAA, GDPR, Audit Logging.
Analytics & BI: Looker, Tableau, Power BI, Sigma, dbt, SQL, Data Modeling, KPI Development, Reporting Automation. Programming & Automation: Python, PySpark, SQL, Bash, Git, Jenkins, Terraform, CI/CD, Agile, DevOps. AI/ML: Vertex AI, BigQuery ML, Feature Engineering, Predictive Analytics, ML Model Training.