Post Job Free
Sign in

Lead Data Scientist - ML, NLP, GenAI

Location:
Berkeley, CA
Salary:
120000
Posted:
October 05, 2026

Contact this candidate

Resume:

Emma Beck Lead Data Scientist

***********@*****.*** 516-***-**** Berkeley, CA 94707

PROFILE

Lead Data Scientist with 9+ years of hands-on experience architecting and deploying production-grade Machine Learning, NLP, and Generative AI solutions across healthcare, fintech, and digital product domains. Demonstrated expertise in building end-to- end ML pipelines from feature engineering and model development to MLOps deployment and real-time inference, leveraging Python, PyTorch, AWS SageMaker, Apache Spark, and Kafka at scale. Proven ability to translate complex data science into measurable business outcomes, including 5M+ in documented cost savings, fraud prevention, and revenue recovery. Skilled at leading cross-functional teams, mentoring junior scientists, and partnering with product, engineering, and compliance stakeholders to deliver AI systems aligned with HIPAA, GDPR, CCPA, and PCI DSS standards. Passionate about applying Causal Inference, Explainable AI, and advanced Generative AI techniques to solve high-impact business problems. SKILLS

•Machine Learning & Predictive Analytics: Supervised & Unsupervised Learning, Classification, Regression, Churn Modeling, Demand Forecasting, Anomaly Detection, Fraud Detection, Pricing Optimization, Recommendation Systems, Feature Engineering, Model Evaluation, XGBoost, LightGBM, Scikit-learn

•Deep Learning & Generative AI: PyTorch, TensorFlow, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Prompt Engineering, Fine-Tuning, PEFT, LoRA, BERT, GPT, Hugging Face Transformers, LangChain, FAISS, Pinecone, Computer Vision, YOLOv5, OpenCV

•NLP & Text Analytics: Natural Language Processing, Entity Recognition, Text Classification, Named Entity Recognition (NER), Sentiment Analysis, Document Understanding, Clinical NLP, Biomedical Text Mining, Transformer Architectures, Semantic Search, Vector Embeddings

•Statistical Analysis & Experimentation: A/B Testing, Bayesian Hypothesis Testing, Causal Inference, Time Series Forecasting, ARIMA, Prophet, LSTM, Quantile Regression, Multi-Armed Bandits, Power Analysis, Sequential Testing, CUPED, Mixed Integer Linear Programming

•Healthcare & Clinical Data Science: EHR/EMR Analytics, Clinical Decision Support, Patient Risk Stratification, Readmission Prediction, Healthcare Predictive Analytics, HL7, FHIR, HIPAA Compliance, Clinical Entity Extraction, Real-Time Patient Monitoring, Care Quality Improvement

•Data Engineering & Big Data: Apache Spark, Kafka, PySpark, Hadoop, Flink, Databricks, Snowflake, dbt, ETL/ELT Pipelines, Airflow, Batch & Real-Time Streaming, Lakehouse Architecture, SQL, R, Python

•MLOps & Cloud Platforms: MLflow, Docker, Kubernetes, Kubeflow, CI/CD, Model Monitoring, Drift Detection, Feature Store, Tecton, Terraform, REST APIs, AWS SageMaker, S3, Glue, Azure ML, Azure Databricks, GCP, BigQuery, Vertex AI

•Governance, Compliance & Visualization: Explainable AI, SHAP, LIME, Model Governance, AI Ethics, GDPR, CCPA, PCI DSS, Tableau, Power BI, Looker, Executive Dashboarding, Agile Methodology, Stakeholder Management PROFESSIONAL EXPERIENCE

Mavenir Systems, Lead Data Scientist 10/2022 – Present

•Architected and deployed a predictive healthcare analytics platform on AWS SageMaker, incorporating patient risk stratification models that reduced hospital readmissions by 18% and generated $5M in annual cost savings.

•Engineered a clinical NLP pipeline using Hugging Face Transformers for entity extraction across 500K+ EHRs per week, improving clinical data capture accuracy by 35% and significantly reducing downstream manual review effort.

•Automated end-to-end healthcare document extraction, normalization, and prediction workflows, cutting clinician review time by 65% across high-volume patient record processing.

•Established robust MLOps workflows using MLflow, Docker, and automated retraining pipelines, maintaining production model AUC above 0.92 with continuous drift monitoring and alerting.

•Delivered AI-driven insights through Power BI and Tableau dashboards integrated with real-time scoring APIs, improving care team response times by 25% and increasing clinical AI adoption rates.

•Led and mentored a team of 4 junior data scientists on model governance, Explainable AI practices, SHAP/LIME interpretability, and HIPAA-compliant ML development workflows.

•Ensured all ML systems met HIPAA, GDPR, and CCPA regulatory requirements through proactive collaboration with compliance, legal, and engineering teams.

HTD Health, Senior Data Scientist 06/2019 – 09/2022

•Designed and deployed a real-time fraud detection platform processing 2M+ daily transactions using PySpark and Kafka, improving detection accuracy by 22% and preventing $8M in annual fraud losses.

•Built a hybrid recommendation engine combining collaborative filtering with BERT-based semantic embeddings, increasing user engagement by 15% year-over-year across digital product surfaces.

•Automated ETL and ELT data pipelines using Airflow and dbt, reducing batch processing time by 90% and significantly improving data reliability and freshness for downstream modeling.

•Developed real-time anomaly detection dashboards in Looker and Tableau for 2M+ daily transactions, giving executives visibility into high-risk behavioral patterns and accelerating compliance response.

•Partnered with compliance and legal teams to design and embed PCI DSS, GDPR, and CCPA controls directly into production data science workflows and model deployment processes.

•Optimized large-scale transaction processing infrastructure to support low-latency model scoring and high-throughput streaming analytics at enterprise scale.

Ontra, Data Scientist 11/2016 – 05/2019

•Developed a customer churn prediction model using XGBoost with advanced feature engineering, enabling targeted retention campaigns that recovered $2.5M in annual revenue.

•Built multi-horizon demand forecasting models using ARIMA, Facebook Prophet, and LSTM neural networks, improving forecast accuracy by 20% and delivering $1.2M in annual operational savings.

•Designed and executed A/B tests and pricing optimization analyses using SQL and R, driving a 12% increase in repeat purchases and informing key growth strategy decisions.

•Applied Mixed Integer Linear Programming (MILP) to logistics and supply chain optimization, reducing shipping costs by $1.5M annually.

•Containerized ML services using Docker and deployed via REST APIs, improving model serving latency by 40% and accelerating cross-team product integration timelines.

•Partnered with product teams to define North Star Metrics, instrumentation frameworks, and measurement strategies for new AI and analytics features.

PROJECTS

Customer Churn Prediction & Retention Intelligence Platform

•Built an end-to-end churn prediction system using XGBoost and LightGBM on a 3M-record customer dataset, achieving AUC of 0.91 with SHAP-based feature importance analysis for interpretability.

•Engineered 80+ behavioral, transactional, and engagement features using Python and Pandas; applied SMOTE to address severe class imbalance (5% churn rate).

•Deployed the model via a Flask REST API integrated with a real-time Tableau dashboard, enabling marketing teams to proactively target high-risk segments and reduce churn by an estimated 18%. Multi-Horizon Demand Forecasting System with Uncertainty Quantification

•Developed an ensemble forecasting pipeline combining ARIMA, Facebook Prophet, and Temporal Fusion Transformers (TFT) to predict demand at SKU-level across 500+ products and 12 regions.

•Implemented quantile regression and prediction intervals for uncertainty estimation, enabling supply chain teams to make risk- aware inventory decisions.

•Reduced MAPE by 23% over the baseline ARIMA model; automated weekly retraining and reporting with Airflow on AWS, cutting manual forecasting effort by 70%.

Large-Scale A/B Testing & Experimentation Framework

•Designed a reusable experimentation platform supporting frequentist A/B tests, Bayesian hypothesis testing, and multi-armed bandit algorithms for adaptive traffic allocation.

•Implemented statistical power analysis, sequential testing (alpha-spending), and CUPED variance reduction to detect smaller effect sizes with 30% fewer required samples.

•Applied the framework to 20+ product experiments, driving a cumulative 14% improvement in conversion rates and informing pricing strategy through rigorous causal measurement. Real-Time Fraud Detection with Graph Neural Networks

•Built a fraud detection system combining traditional ML (XGBoost) with a Graph Neural Network (GNN) to capture hidden transactional relationships, improving detection precision by 19% over baseline.

•Streamed 1M+ daily events through Apache Kafka and scored models in sub-50ms using an optimized FastAPI inference service deployed on Kubernetes.

•Integrated SHAP explainability into fraud alerts, providing compliance teams with human-readable model reasoning for every flagged transaction, meeting GDPR Article 22 requirements. RAG-Powered Clinical Decision Support System

•Built a Retrieval-Augmented Generation (RAG) system using LangChain, FAISS, and a fine-tuned clinical LLM to surface evidence-based treatment recommendations from a corpus of 200K+ medical documents.

•Integrated HL7/FHIR APIs to ingest real-time patient data and personalize recommendations based on individual clinical context, lab values, and medication history.

•Reduced clinician time-to-decision by 30% in pilot testing; implemented HIPAA-compliant data handling with audit logging, access controls, and PII masking pipelines.

CERTIFICATES

Google Professional Machine Learning Engineer Azure Data Scientist Associate EDUCATION

Bachelor's of Computer Science



Contact this candidate