SANJAY BALAJI
Houston, TX **************@*******.*** 713-***-**** https://www.linkedin.com/in/sanjay-balaji-bb9240422/ PROFESSIONAL SUMMARY
AI/ML Data Architect with 10+ years of experience defining enterprise data strategies, modernizing cloud data platforms, and delivering trusted data products across insurance, financial services, healthcare, and telecommunications. Proven record of designing scalable lakehouse, data warehouse, data mesh, data fabric, streaming, and event-driven architectures that support operational analytics, machine learning, generative AI, and intelligent automation. Experienced in establishing architecture standards for data ingestion, transformation, storage, semantic modeling, governance, lineage, observability, security, and consumption across Azure, AWS, and hybrid environments. Combines strategic architecture leadership with hands-on expertise in Python, SQL, PySpark, Databricks, Snowflake, Azure Data Factory, Airflow, dbt, Kafka, Delta Lake, Unity Catalog, Terraform, Kubernetes, and modern API integration patterns. Designs conceptual, logical, and physical data models, including dimensional, star-schema, canonical, Data Vault, master-data, and domain-oriented models supporting analytics and transactional workloads. Builds AI-ready data foundations for predictive modeling, real-time inference, RAG applications, vector search, LLM orchestration, feature engineering, and governed agentic workflows. Trusted partner to engineering, product, analytics, compliance, security, and executive stakeholders, translating business priorities into multi- year architecture roadmaps and practical implementation plans. Experienced in leading architecture reviews, mentoring engineers, defining reusable reference architectures, governing data contracts, improving platform reliability, and optimizing cloud cost. Known for influencing without authority, balancing long-term architectural quality with near-term business delivery, and creating systems that improve data trust, operational efficiency, regulatory readiness, and enterprise decision-making. CORE COMPETENCIES
Enterprise Data Architecture AI/ML Architecture Generative AI Architecture Lakehouse Architecture Data Mesh Data Fabric Medallion Architecture Data Warehousing Data Products Semantic Layers Master Data Management Data Governance Metadata Management Data Catalogs Data Lineage Data Contracts Data Quality Data Observability Cloud Modernization Streaming Architecture Event-Driven Architecture API Integration Microservices Platform Engineering Infrastructure as Code FinOps Architecture Review Boards Technology Roadmaps Executive Communication Technical Leadership Mentoring PROFESSIONAL AND TECHNICAL SKILLS
AI & Machine Learning : Python, Java, Scikit-learn, TensorFlow, PyTorch, Generative AI, RAG, LLM Orchestration, Agentic AI, Embeddings, Vector Databases, Feature Stores, Azure AI Foundry, Snowflake Cortex, Responsible AI MLOps & LLMOps : MLflow, Kubeflow, Azure Machine Learning, Vertex AI, Experiment Lineage, Model Input Validation, Reproducible Training Datasets, Production Scoring Workflows, Model Evaluation, Prompt Governance, Hallucination Monitoring Backend, Microservices & APIs : Python, Java, JavaScript, TypeScript, REST APIs, API Integration, Microservices, Event-Driven Architecture, Kafka, Confluent Kafka, Azure Event Hubs, Pydantic v2, CDC, Batch and Streaming Pipelines Cloud & Infrastructure : Microsoft Azure, Amazon Web Services, Google Cloud Platform, Hybrid Cloud, Multi-Cloud Architecture, Terraform, Docker, Kubernetes, AKS, EKS, Azure DevOps, GitHub Actions, GitLab CI, Jenkins, CI/CD, Infrastructure as Code, IAM, RBAC, ABAC, Zero Trust, Encryption, Secrets Management, Network Isolation, Disaster Recovery, High Availability Data Engineering & Storage : Azure Databricks, Snowflake, Azure Synapse Analytics, Microsoft Fabric, Azure Data Lake Storage Gen2, AWS S3, Redshift, BigQuery, PostgreSQL, AWS Aurora, SQL Server, Oracle, MongoDB, DynamoDB, Apache Spark, PySpark, Databricks SQL, Apache Flink, Apache Beam, Hadoop, HDFS, Hive, Presto, Athena, Dataproc, Dataflow, Azure Data Factory, Apache Airflow, Cloud Composer, dbt, Azure Stream Analytics, Apache NiFi, Prefect, Dagster, AWS Glue, Delta Lake, Apache Iceberg, Hudi, Data Lakes, Data Warehouses, Data Marts, Operational Data Stores, Data Vault, Dimensional Modeling, Semantic Modeling, MDM, Unity Catalog, DataHub, Collibra, Alation, Atlan Testing, Monitoring & Observability : Data Quality Rules, Schema Validation, Schema Registries, Data Contracts, Source-to-Target Reconciliation, Data Lineage, Metadata Management, Data Classification, Monte Carlo, Dynatrace, CloudWatch, CI/CD Validation, Python Unit Testing, Pipeline Validation, Audit Logging, SLA Monitoring, Cost Anomaly Monitoring PROFESSIONAL EXPERIENCE
Staff AI/ML Data Architect
MetLife Remote Apr 2023 – Present
• Own the architectural direction for an enterprise insurance data and AI platform supporting claims intelligence, operational analytics, predictive risk scoring, executive reporting, and AI-assisted decision workflows across claims, underwriting, billing, and customer-service domains.
• Defined a multi-year modernization roadmap that moved fragmented analytical workloads toward a governed Azure and Databricks lakehouse architecture, improving data availability and reducing duplicated data-processing logic across engineering and analytics teams.
• Designed medallion-style ingestion, standardized, curated, and semantic data layers using Azure Data Lake Storage, Databricks, Delta Lake, PySpark, SQL, and Unity Catalog, shortening the delivery cycle for new analytical data products by approximately 35%.
• Established canonical insurance data models covering customers, policies, claims, coverage, payments, providers, risk indicators, and case activities, giving business and technical teams a consistent vocabulary for enterprise reporting and AI use cases.
• Architected conceptual, logical, and physical models using dimensional, star-schema, Data Vault, and domain-oriented patterns, balancing historical analytics requirements with operational access, performance, reuse, and regulatory traceability.
• Led the architecture of reusable batch, CDC, API, and event-driven ingestion patterns using Azure Data Factory, Databricks, Kafka, Event Hubs, and Python, enabling near-real-time claim updates while reducing dependence on brittle point-to-point integrations. 2
• Partnered with data science and ML engineering teams to create AI-ready training and inference pipelines with governed feature sets, experiment lineage, model input validation, reproducible datasets, MLflow integration, and monitored production scoring workflows.
• Designed generative AI and RAG architecture patterns for claims-document summarization, policy lookup, adjuster assistance, and operational knowledge search using secure LLM gateways, vector databases, prompt templates, retrieval controls, human approval checkpoints, and audit logging.
• Established responsible AI standards covering model evaluation, prompt governance, hallucination monitoring, sensitive-data handling, explainability, access control, fallback behavior, and traceable human intervention for insurance decision-support applications.
• Implemented governance-by-design through Unity Catalog, automated metadata registration, data lineage, schema validation, Pydantic- based data contracts, quality gates, role-based access, data classification, retention controls, and environment isolation across development, testing, and production.
• Built a data observability framework that monitored freshness, volume, schema drift, null rates, reconciliation, lineage breaks, SLA adherence, pipeline failures, and cost anomalies, reducing recurring production data incidents by approximately 40%.
• Introduced cloud FinOps practices that attributed Databricks compute, warehouse consumption, storage, model execution, and LLM token usage to business domains and teams, helping platform owners lower avoidable processing spend by approximately 20%.
• Chaired architecture and design reviews for multiple delivery teams, documenting reference architectures, decision records, integration maps, non-functional requirements, and technical standards while mentoring senior engineers in scalable modeling, distributed processing, and governance practices.
• Presented architecture roadmaps, platform tradeoffs, implementation risks, and business-value cases to executive, product, operations, security, and compliance stakeholders, influencing investment decisions and improving alignment between enterprise strategy and technical delivery.
Lead Data and AI Solutions Architect
Heartland Remote Jan 2020 – Mar 2023
• Led the data architecture for cloud-based financial-services platforms supporting merchant onboarding, card and payment operations, compliance surveillance, risk investigations, customer servicing, transaction analytics, and regulatory reporting.
• Assessed legacy SQL Server, Oracle, file-based, and application-owned data flows and produced a modernization blueprint spanning cloud lakehouse, warehouse, event-streaming, API, master-data, and governance capabilities.
• Designed a scalable Azure data ecosystem using Azure Data Lake Storage Gen2, Databricks, Delta Lake, Azure Data Factory, Synapse Analytics, Event Hubs, Azure Functions, AKS, and Power BI to support both operational and analytical workloads.
• Introduced domain-oriented data products for merchants, accounts, transactions, settlements, disputes, fees, risk events, and communications, with defined owners, service-level expectations, quality indicators, access policies, and lifecycle responsibilities.
• Created canonical and dimensional financial models that unified definitions across payment-processing, merchant-service, fraud, compliance, and finance teams, reducing reconciliation disputes and inconsistent KPI calculations by roughly 30%.
• Architected high-throughput streaming pipelines using Kafka-compatible Event Hubs, Spark Structured Streaming, and Delta Lake, enabling near-real-time transaction monitoring, risk signal generation, operational alerts, and downstream analytics.
• Implemented batch and incremental ETL/ELT pipelines with Python, PySpark, SQL, dbt, Airflow, and Azure Data Factory, using idempotent loading, dependency controls, retry policies, backfill workflows, partition strategies, and automated reconciliation.
• Designed API and backend-for-frontend integration patterns that shielded core warehouses from application concurrency spikes, introduced caching and pagination, enforced service boundaries, and improved response stability for customer-facing and internal applications.
• Partnered with machine learning teams to support fraud detection, merchant-risk scoring, anomaly detection, account prioritization, and communication-monitoring use cases through governed feature pipelines, reproducible training datasets, and real-time inference patterns.
• Established MDM and reference-data practices for merchant, customer, account, product, and organizational hierarchies, improving entity resolution, golden-record quality, downstream reporting consistency, and auditability.
• Implemented enterprise data governance controls including cataloging, lineage, classification, encryption, tokenization, role-based access, schema contracts, change approval, quality validation, retention, and evidence collection for regulated financial data.
• Automated cloud resources, access boundaries, pipeline deployments, catalog configuration, and database changes using Terraform, Azure DevOps, Git, CI/CD, and policy-as-code, reducing manual environment configuration and deployment errors by approximately 45%.
• Defined architecture standards for resiliency, disaster recovery, replication, backup validation, high availability, observability, performance tuning, cost allocation, secrets management, network isolation, and Zero Trust data access.
• Guided three cross-functional engineering teams through solution design, backlog refinement, implementation reviews, performance analysis, and production readiness while coaching engineers on distributed systems, modeling standards, Python quality, and reusable platform patterns.
• Worked with product, compliance, operations, security, finance, and executive leaders to translate regulatory and business requirements into architecture roadmaps, investment options, delivery phases, and measurable operational outcomes. Senior Data and ML Platform Architect
EviCore by Evernorth Franklin, TN Sep 2017 – Dec 2019
• Architected healthcare data platforms supporting authorization workflows, clinical review, utilization management, service requests, provider operations, member records, reporting, and analytical decision support across regulated healthcare environments.
• Developed an enterprise data architecture that unified clinical, claims, eligibility, provider, authorization, operational, and reference data from internal applications, payer feeds, partner APIs, files, and third-party healthcare systems.
• Designed a cloud and hybrid data platform using Azure services, SQL Server, Databricks, Spark, Python, PySpark, data lake storage, and governed analytical layers to modernize legacy reporting and manual data-processing workflows.
• Created conceptual, logical, and physical healthcare data models with dimensional, normalized, and canonical patterns, supporting operational applications, analytical marts, clinical reporting, machine learning, and longitudinal patient views. 3
• Established source-of-truth and master-data strategies for members, providers, facilities, plans, procedures, diagnoses, authorizations, and organizational hierarchies, improving data consistency across clinical and administrative systems.
• Designed secure batch and near-real-time integration patterns using REST APIs, messaging, CDC, file ingestion, and orchestrated pipelines, enabling dependable movement of structured and semi-structured healthcare data.
• Partnered with application teams to define API-first and event-driven service boundaries for workflow updates, clinical decisions, document events, and reporting data while reducing direct coupling between transactional systems and analytical consumers.
• Supported AI/ML use cases involving utilization prediction, operational prioritization, document classification, anomaly detection, and decision support by designing governed feature pipelines, labeled datasets, reproducible transformations, and monitored scoring interfaces.
• Designed healthcare interoperability patterns based on HL7 and FHIR concepts, mapping external clinical and payer data into enterprise models while preserving provenance, terminology, validation, and access constraints.
• Embedded HIPAA-aligned controls into the platform through encryption, least-privilege access, audit trails, sensitive-field classification, tokenization, environment separation, retention rules, and secure data-sharing processes.
• Established data-quality frameworks covering source-to-target reconciliation, referential integrity, code-set validation, duplicate detection, schema drift, timeliness, completeness, and business-rule conformance, reducing manual data review by approximately 35%.
• Introduced automated data deployments and testing through Git, CI/CD, SQL migration scripts, Python unit tests, pipeline validation, and controlled promotion across development, QA, staging, and production environments.
• Improved query and pipeline performance through partitioning, indexing, data compaction, caching, workload management, incremental processing, and Spark tuning, lowering processing windows by approximately 30%.
• Created architecture diagrams, lineage views, source-to-target mappings, interface specifications, runbooks, technical requirements, and architecture decision records that strengthened cross-team understanding and audit readiness.
• Mentored developers, data engineers, analysts, and technical leads while facilitating design reviews and workshops with healthcare stakeholders, security teams, infrastructure groups, and delivery managers. Data Platform and Automation Architect
Comcast Philadelphia, PA Sep 2014 – May 2017
• Designed data and integration solutions for high-volume customer-service, call-center, digital-assistant, order-management, document- processing, identity-verification, and operational-reporting platforms.
• Built scalable data pipelines that captured and standardized customer interactions, intent signals, service events, account activities, document metadata, application logs, and operational outcomes for analytics and workflow automation.
• Defined relational, dimensional, and document-oriented models for customer, account, product, order, interaction, agent, device, and service domains, improving data reuse across support, operations, engineering, and reporting teams.
• Developed batch and near-real-time ingestion processes using Python, SQL, Java, APIs, messaging, and distributed processing patterns, enabling dependable movement of structured and semi-structured data from multiple source systems.
• Designed event-driven architectures for customer requests, order updates, document events, notifications, and identity-verification workflows, improving responsiveness and reducing reliance on scheduled point-to-point integrations.
• Created reusable REST API and microservice integration patterns that separated customer-facing applications from underlying data stores, strengthened security, and allowed backend services to evolve independently.
• Contributed to early machine learning and natural-language-processing initiatives for customer-intent classification, automated routing, recommendation, anomaly detection, and assisted-service experiences.
• Developed feature-engineering and training-data processes that transformed interaction history, service outcomes, account events, and operational signals into structured datasets for predictive modeling and classification.
• Implemented data-quality controls covering duplicate detection, schema validation, referential integrity, missing values, format standardization, reconciliation, and exception handling, improving the reliability of downstream reports and automation.
• Supported migration of legacy workloads toward cloud-ready, containerized, and distributed architectures using Docker, infrastructure automation, CI/CD, centralized logging, monitoring, and repeatable environment configuration.
• Improved database and pipeline performance through query-plan analysis, indexing, partitioning, caching, incremental loads, concurrency controls, and workload-aware processing, reducing selected reporting runtimes by approximately 40%.
• Established access-control, encryption, audit, retention, and sensitive-data-handling patterns for customer and identity-related information in collaboration with cybersecurity and infrastructure teams.
• Produced architecture diagrams, data-flow documentation, interface contracts, deployment guides, runbooks, and technical standards that helped application, analytics, QA, and operations teams deliver changes consistently.
• Collaborated with product managers, customer-service leaders, developers, data specialists, and infrastructure engineers to translate service challenges into practical platform designs and phased delivery plans.
• Mentored junior engineers through code reviews, modeling sessions, troubleshooting, and design discussions, helping establish stronger software engineering, testing, documentation, and production-support practices. SELECTED ARCHITECTURE IMPACT
• Designed enterprise lakehouse, warehouse, streaming, API, and AI-ready platform patterns across insurance, financial services, healthcare, and telecommunications.
• Established reusable data governance, lineage, metadata, quality, contract, observability, and access-control frameworks.
• Led cloud modernization initiatives spanning Azure, AWS, Databricks, Snowflake, Spark, Kafka, Airflow, dbt, Terraform, and Kubernetes.
• Enabled predictive analytics, real-time inference, RAG, generative AI, vector-search, and agentic AI solutions through governed data architecture.
• Guided multiple engineering teams through architecture reviews, technical standards, modernization roadmaps, and production-readiness decisions.
4
• Improved data-processing performance, platform reliability, cloud-cost visibility, delivery speed, and data trust through automation and architecture standardization.
EDUCATION
Penn State University
Master of Software Engineering, Software Engineering Aug 2019 – Aug 2021
Drexel University
Bachelor of Science in Computer Science, Game Programming and Development 2013 – 2017