Post Job Free
Sign in

Senior Data Scientist

Location:
Richmond, TX, 77469
Posted:
August 19, 2026

Contact this candidate

Resume:

Aaron Allen

Senior Data Scientist

Causal Inference · Statistical Modeling · Clinical AI

945-***-****

Richmond, TX (Open to Remote)

*************@*****.***

www.linkedin.com/in/aaronallen094

SUMMARY

Data Scientist with 11 years of experience building models and analytical systems that drive real business and clinical decisions across healthcare, retail analytics, clinical AI, and housing finance. Career moved from financial services analytics at NFCU into consulting DS work at Metability across retail loyalty and healthcare claims, then into increasingly technical senior work at Excella covering genomic survival modeling at Tempus and causal mortgage risk modeling at Fannie Mae. The consistent thread across every senior engagement has been causal inference, not just measuring what happened but isolating what actually caused it, and that methodological discipline has shaped every major project since joining Excella in 2020. Healthcare is the deepest domain, running from population health claims analytics through clinical NLP and genomic ML, and that depth has opened natural pathways into clinical AI specifically. More recently the focus has shifted to statistical evaluation of GenAI systems, applying the same experimental rigor to measuring LLM output quality that was previously applied to traditional ML models. SKILLS

Languages: Python, R, SQL

Machine Learning: scikit-learn, XGBoost, LightGBM, PyTorch, statsmodels, Prophet, KMeans, Hierarchical Clustering Statistical Methods: Bayesian Inference, A/B Testing, Sequential Testing, Survival Analysis, Cox Proportional Hazards, Kaplan-Meier, Difference-in-Differences, Propensity Score Matching Causal Inference: DoWhy, CausalML

NLP / GenAI: spaCy, Hugging Face Transformers, LLM Evaluation Model Interpretability: SHAP

Bioinformatics: Bioconductor

MLOps: MLflow, DVC, Docker, GitHub Actions

Cloud / AWS: SageMaker, EMR, S3, Redshift, Athena

Data Engineering: PySpark

Data Analysis: Pandas, NumPy, scipy

Visualization: Tableau

EXPERIENCE

Excella

Senior Data Scientist (Consulting) · Jun 2020 – Jul 2026 Client Engagement: Tempus (Genomic & Clinical AI) — tempus.com 2022 – 2026

• Built survival analysis models using Cox proportional hazards regression on genomic and clinical datasets covering thousands of cancer patients, with Kaplan-Meier curves used for cohort visualization alongside the predictive models to help oncologists identify effective therapy options for individual patient profiles.

• Designed multimodal ML pipelines in PyTorch combining EHR data, RNA sequencing outputs, and molecular biomarker signals, with treatment response models improving AUC by around 11 percentage points over clinical-only baselines across the primary cancer types evaluated.

• Developed clinical NLP pipelines using spaCy and Hugging Face Transformers to extract structured diagnosis, treatment, and outcome information from unstructured oncology notes, turning free text clinical documentation into clean features for downstream modeling.

• Applied causal inference using DoWhy to isolate the true effect of treatment decisions on patient outcomes, controlling for confounders in observational clinical data and giving research teams more reliable evidence than retrospective analysis alone could provide.

• From 2022 onward, built statistical evaluation frameworks for Tempus's LLM based clinical document tools, designing experiments to measure response accuracy, clinical relevance, and factual grounding before any outputs fed into physician-facing workflows.

• Used R and Bioconductor for RNA sequencing data analysis and genomic feature work, managed model training on AWS SageMaker, versioned datasets with DVC and S3, and tracked all experiments in MLflow with Docker ensuring reproducible environments across research runs.

Client Engagement: Fannie Mae (Mortgage Risk) — fanniemae.com 2020 – 2022

• Developed mortgage default risk models using XGBoost and LightGBM on loan level datasets covering millions of mortgages, improving GINI by around 8 points over the previous generation model and giving risk management teams better visibility into portfolio credit quality across economic scenarios.

• Applied SHAP to all credit risk models to generate loan level explainability for model governance and regulatory examination, ensuring every risk score had a documented and auditable attribution of the features driving individual loan decisions.

• Used CausalML to measure the true causal impact of loan modification programs on default outcomes, isolating program effects from selection bias in observational data and giving the policy team reliable evidence on which interventions actually reduced foreclosure rates.

• Designed and ran experiments on credit policy changes using Bayesian inference, giving risk and policy teams a rigorous framework to evaluate how underwriting guideline updates affected eligible mortgage pool quality and downstream default performance.

• Led statistical evaluation of GenAI tools applied to mortgage document processing from 2024 onward, building evaluation datasets and metric frameworks to measure extraction accuracy, response consistency, and downstream risk model impact before production deployment.

• Managed analytical workflows on AWS with Redshift and Athena for large scale loan data querying, S3 for dataset management, SageMaker for model training, and GitHub Actions automating monthly model retraining pipelines triggered by loan data refreshes.

Metability LLC

Data Scientist (Consulting) · Nov 2016 – Apr 2020

Client Engagement: Kroger / 8451 (Retail Loyalty Analytics) — 8451.com

• Developed customer segmentation models using KMeans and hierarchical clustering on Kroger Plus loyalty card purchase histories, identifying distinct shopper personas by category preferences, purchase frequency, and basket composition that the personalization team used to improve digital coupon redemption rates by around 18% on targeted segments.

• Built store-level demand forecasting models using Prophet and statsmodels on weekly sales data across thousands of product-store combinations, reducing MAPE by around 3 percentage points over baseline and giving category managers reliable advance visibility into replenishment needs across high-velocity SKUs.

• Measured true promotional sales lift using difference-in-differences and propensity score matching, controlling for seasonal effects and pre-existing purchase trends to give CPG brand partners reliable evidence on which promotions drove genuine incremental volume rather than just pulling forward existing demand.

• Designed and analyzed A/B tests on personalized digital coupon targeting strategies using Python, scipy, and Bayesian sequential testing, with results directly informing how the team allocated personalization spend across millions of Kroger loyalty card households.

• Managed analytical workflows on AWS with S3 for data storage and Redshift for large scale SQL querying across loyalty transaction datasets, building reusable query layers that cut recurring data preparation time by around 40% and became standard practice across the broader analytics team. Client Engagement: Inovalon (Healthcare Claims Analytics) — inovalon.com

• Built patient risk stratification models using scikit-learn and XGBoost on medical and pharmacy claims data covering several million members, helping health plan clients identify high-risk individuals earlier and prioritize care management outreach before conditions escalated.

• Developed predictive models for healthcare cost drivers using LightGBM, engineering features from ICD-10 diagnosis codes, procedure codes, and pharmacy claims to forecast medical cost trends and give actuarial teams earlier visibility into plan performance.

• Ran statistical analyses in R on HEDIS quality measure performance across multiple health plan clients, identifying care gaps at the population level and modeling the expected impact of specific intervention programs on preventive care compliance rates.

• Applied SHAP across all predictive models so clinical and actuarial stakeholders could validate model reasoning before any outputs informed care program decisions, which was a non-negotiable requirement on every client engagement.

• Built feature engineering pipelines using PySpark on AWS EMR across large scale claims datasets, stored processed outputs in S3, and managed model training through SageMaker as the team standardized on AWS from mid-2017 onward.

Navy Federal Credit Union

Data Analyst · Jan 2015 – Sep 2016

• Pulled and transformed large member datasets using SQL across multiple relational databases, building the recurring reports and ad hoc analyses that business and product teams relied on weekly for lending performance, account activity, and member retention tracking.

• Built Tableau dashboards that gave leadership visibility into member behavior trends, loan portfolio health, and product adoption rates, replacing a lot of manual Excel reporting that had been slowing the team down.

• Used Python and Pandas to clean, reshape, and analyze member transaction data, doing the kind of exploratory work that surfaced patterns in spending behavior and account usage that SQL alone was not flexible enough to catch.

• Supported basic customer segmentation work using member demographics and transactional history, helping the marketing team identify which member groups were underutilizing products and where outreach was most likely to drive engagement.

EDUCATION

Bachelor's Degree in Computer Science

ITT Technical Institute, TX · 2010 – 2014



Contact this candidate