SRIKAR MURTHY MOKARALA
Senior Data Engineer
Email: *********@*****.***
Phone: 469-***-****
Professional Summary:
•Senior Data Engineer with 6+ years of experience architecting and delivering enterprise-scale data pipelines, cloud data warehouses, MLOps platforms, and Generative AI/LLM solutions across AWS, GCP, and Azure for Fortune 500 organizations in healthcare, telecom, and consulting.
•Proven expertise building big end-to-end data and machine learning ecosystems from ingestion and transformation through model deployment and production monitoring using PySpark, Databricks, Delta Lake, BigQuery, Snowflake, Vertex AI, SageMaker, and Gemini/LLM integrations.
•Adept at translating complex business requirements into scalable, cost-optimized cloud architecture, driving measurable performance gains, and mentoring cross-functional engineering teams.
•Recognized for rapid ramp-up in high-velocity contract environments, strong stakeholder communication, and consistent delivery of production-grade, AI-augmented data solutions.
Key Highlights:
•6+ years architecting cloud-native data and AI/ML platforms across GCP, AWS, and Azure for Fortune 500 clients.
•Deep hands-on expertise in Databricks, Delta Lake, PySpark, and SparkSQL for large-scale distributed data processing across every project.
•Proven track record designing and shipping Generative AI / LLM solutions (Gemini, prompt engineering, unstructured-text extraction) for enterprise decision-making.
•Strong MLOps foundation — model registries, versioning, experiment tracking, and CI/CD-driven ML deployment pipelines on Vertex AI and SageMaker.
•Delivered measurable business impact through faster pipeline runtimes, leaner infrastructure spends, and more reliable, production-ready data.
•Trusted technical partner to data science, DevOps, and business stakeholder teams throughout the full SDLC; mentors junior engineers on cloud and Spark best practices.
Core Competencies
Cloud Platforms:
AWS (EMR, Glue, Redshift, S3, Lambda, Kinesis, Step Functions, SageMaker), GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Vertex AI, Cloud Composer), Azure (Synapse, ADF)
Big Data & Data Engineering:
Databricks, Delta Lake PySpark, SparkSQL Apache Beam, Kafka, Airflow (Cloud Composer), Snowflake, Data Fusion, REST APIs
Databases & Warehouses:
BigQuery, Redshift, Snowflake, Synapse, SQL Server, Oracle, PostgreSQL
AI / ML / MLOps / Gen AI:
Vertex AI, SageMaker, Databricks ML, Feature Engineering, Model Registry, Experiment Tracking, Gemini, LLM Integration, Prompt Engineering
Programming & BI:
Python, SQL, Scala, Shell, Tableau, Power BI, Looker
CI/CD & DevOps:
GitHub Actions, Azure DevOps, Jenkins, Docker, Kubernetes, Git
Contract Value Proposition
•6+ years of enterprise contract & consulting delivery at the senior data engineer level.
•Rapid onboarding into new client environments, data platforms, and tooling with minimal ramp-up time.
•End-to-end ownership: from requirements gathering through production deployment, ML/AI integration, and post-go-live support.
•Strong cross-functional collaboration with data scientists, BAs, product owners, and DevOps teams.
•Delivered cloud migration, modernization, greenfield builds, and Gen AI-enabled solutions across healthcare, telecom, and consulting verticals.
•Hands-on Databricks and Delta Lake experience woven across every engagement, alongside native AWS/GCP tooling.
•Clear, proactive stakeholder communication and measurable delivery outcomes.
Professional Experience
Role: AI Data Engineer Sept 2024 – Present
Client: CVS Health, Texas
Responsibilities:
•Partnered with business stakeholders to translate reporting and analytics requirements into scalable GCP-based data solutions supporting enterprise care management workflows.
•Designed and implemented BigQuery-centric data architectures with optimized schemas, partitioning, and SQL tuning to improve analytical performance and reduce operational costs.
•Built scalable ETL/ELT pipelines using Dataproc, Cloud Data Fusion, Dataflow, Pub/Sub, and Snowflake for batch and real-time ingestion across multiple data domains.
•Developed a Generative AI solution using Gemini Flash 2.5 to process unstructured call notes, extracting member-level features and insights to enhance care management decision-making.
•Designed prompt-engineering and evaluation frameworks for LLM-based extraction pipelines, improving accuracy and consistency of Gen AI outputs in production.
•Productionized ML models on GCP using Vertex AI, establishing a centralized model registry for versioning, lifecycle management, and scalable inference pipelines.
•Built feature engineering pipelines using BigQuery, Dataform, and Cloud Storage to deliver consistent, reusable feature sets for ML model training.
•Leveraged Databricks and Delta Lake for cross-platform data validation and exploratory analysis, ensuring consistency between GCP-native and multi-cloud data assets.
•Orchestrated complex workflows using Cloud Composer (Airflow), ensuring reliable scheduling, dependency management, and monitoring across production DAGs.
•Developed Python-based ingestion frameworks and REST API integrations supporting near real-time and batch data processing.
•Implemented CI/CD pipelines using GitHub Actions with YAML-based workflows enabling automated deployments to Google Cloud Storage on code merges.
•Established monitoring and alerting for production ML pipelines using Vertex AI Pipelines and Cloud Monitoring, reducing model drift and pipeline-failure incidents.
•Delivered curated datasets supporting Tableau and Looker dashboards and mentored junior engineers on GCP and data pipeline best practices.
Environment: GCP, BigQuery, Snowflake, Dataproc, Vertex AI, Cloud Data Fusion, Dataflow, Pub/Sub, Composer (Airflow), Dataform, Gemini, Databricks, Delta Lake, Python, SQL, Docker, Kubernetes, GitHub Actions, MLOps, Gen AI
Role: Senior Data Engineer Jun 2022 – Aug 2024
Client: Charter Communications, Denver, CO
Responsibilities:
•Designed and built end-to-end data pipelines using PySpark, SparkSQL, AWS EMR, Glue, Redshift, S3, Lambda, and Step Functions for batch and streaming workloads.
•Developed and optimized Databricks-based transformation pipelines using Delta Lake and PySpark notebooks to produce clean, business-aligned analytical datasets in Redshift, improving pipeline reliability and reducing processing time.
•Led migration of legacy on-prem Hadoop and Teradata workloads to AWS Redshift, EMR, and Databricks, improving scalability and reducing infrastructure costs.
•Enabled model versioning, experiment tracking, and artifact management using SageMaker Experiments and Model Registry, integrating Databricks notebooks for feature engineering and model training at scale.
•Optimized analytics and ML workloads using Redshift Spectrum and S3-based data lakes following Lake House architecture patterns built on Databricks and Delta Lake.
•Tuned Spark jobs on EMR and Databricks clusters for performance, memory usage, and cost optimization, meaningfully reducing average job runtime.
•Designed data quality and validation frameworks to catch schema drift and anomalies before they reached downstream reporting layers.
•Collaborated with data analysts and data scientists to deliver clean, well-modeled datasets aligned to business KPIs and ML model requirements.
•Implemented CI/CD pipelines using Jenkins and Azure DevOps to automate deployments and support production release cycles across multiple environments.
•Mentored a team of junior data engineers on PySpark, AWS Glue, and Databricks best practices, conducting code reviews and knowledge-sharing sessions.
•Partnered with cross-functional stakeholders to gather requirements and translate them into scalable, production-ready AWS data architectures.
Environment: AWS (S3, EMR, Glue, Redshift, Lambda, Step Functions, SageMaker), Databricks, Delta Lake, PySpark, SQL, Python, Airflow, Jenkins, GCP, Git, Azure DevOps
Role: Data Engineer Jun 2020 – Dec 2021
Client: Capgemini, India
Responsibilities:
•Delivered data engineering and analytics solutions for enterprise clients across distributed onshore/offshore delivery teams.
•Designed and developed Spark-based data pipelines for large-scale processing, data profiling, cleansing, and validation using SQL, Python, and Hive.
•Built and optimized data transformation workflows using Databricks notebooks and PySpark for exploratory data analysis and reusable data-cleansing routines.
•Assisted with AWS Redshift-based data warehouse development and data ingestion workflows supporting analytics consumption.
•Contributed to NLP and sentiment analysis use cases using Python and Spark ML libraries, applying feature-extraction techniques to unstructured text.
•Built dashboards and visualizations using Tableau to communicate insights and model outputs to business users.
•Partnered with QA and business analyst teams to validate data accuracy and ensure alignment with client reporting requirements.
•Documented pipeline designs and contributed to internal best-practice guidelines for Spark, Databricks, and AWS-based data engineering.
Environment: AWS (Redshift, S3), Databricks, Spark, Python, SQL, Tableau, Hadoop, Hive
Role: Data Analyst Jan 2020 – Jun 2020
Client: Igate Global Solutions, Hyderabad, India
Responsibilities:
•Extracted, transformed, and analyzed data from Hadoop and relational databases to support operational reporting.
•Developed SQL and Hive queries to create summary datasets, metrics, and trend analyses for downstream reporting teams.
•Supported data quality checks and schema design enhancements in collaboration with data engineering teams.
•Gained early hands-on exposure to distributed processing using Spark and Databricks notebooks for ad-hoc data exploration.
•Assisted senior analysts in preparing Excel-based reporting summaries for business stakeholders.
Environment: SQL, Hive, Hadoop, Python, Excel, Spark
Education:
•Master of Science (M.S.) — Western Illinois University, USA - Jan 2022 – May 2023
•Bachelor of Technology (B.Tech) — Jawaharlal Nehru Technological University, India - Aug 2016 – May 2020