Post Job Free
Sign in

Senior GenAI & LLM Systems Engineer

Location:
Charlotte, NC
Posted:
July 10, 2026

Contact this candidate

Resume:

DANIEL HAYDEN

Senior Machine Learning Engineer · GenAI & LLM Systems · Open to Remote (US)

Stafford, VA (Remote) • 703-***-**** • ***************@*****.*** • http://www.linkedin.com/in/danielhayden1214 PROFESSIONAL SUMMARY

I'm a Senior ML Engineer with 14 years of experience building real production AI systems, not just prototypes. My background spans GenAI, LLM inference, multi-agent architectures, and RAG pipelines across healthcare, fintech, and fraud detection. I've spent most of my career at the intersection of cutting-edge ML and serious engineering constraints, so I know how to ship models that actually hold up in regulated, high-stakes environments. I own the full stack from model selection and fine-tuning through API design, MLOps, and production monitoring. I work well independently, communicate clearly across teams, and I'm looking for a strong remote-first team where I can do meaningful work. CORE TECHNICAL SKILLS

Languages Python, SQL

LLMs & GenAI OpenAI API, Hugging Face Transformers, llama.cpp, LangChain, LlamaIndex, PydanticAI, Prompt Engineering, Fine-tuning, Quantization (GGUF/AWQ)

RAG & Vectors PostgreSQL pgvector, Embedding pipelines, Semantic search, OpenSearch, Retrieval-Augmented Generation MLOps Docker, Kubernetes, Helm, Terraform, GitHub Actions, ArgoCD, Model versioning, Drift detection, A/B testing Cloud (AWS) EKS, ECS Fargate, ECR, S3, RDS, Glue, Athena, API Gateway, OpenSearch, Lambda, CloudWatch, Kinesis Frameworks FastAPI, Pydantic, REST, gRPC, Scrapy, BeautifulSoup, scikit-learn, TensorFlow, SparkML Observability Prometheus, Grafana, CloudWatch, model performance dashboards Other Model Context Protocol (MCP), GitOps, SpiceDB authorization, OAuth PROFESSIONAL EXPERIENCE

Senior Machine Learning Engineer Excella Arlington, VA (Remote) May 2018 - Jun 2026 Project: GPT4ALL (privacy-first offline LLM ecosystem for consumer hardware)

Built a local-first LLM inference platform in Python that lets users run open-source models entirely offline with no cloud dependency. Designed for both personal users and enterprise teams where data privacy is a hard requirement.

Got sub-second response times on 7B to 13B parameter models running on CPU-only hardware by combining llama.cpp with GGUF quantization and smart batching through Hugging Face Transformers.

Designed clean API layers with FastAPI and PydanticAI that handled prompt execution, model switching, embedding generation, and local agent orchestration with strict input/output validation throughout.

Built a local RAG pipeline backed by PostgreSQL pgvector with document and web ingestion via Scrapy and BeautifulSoup, so users could query private knowledge bases without touching any external vector database.

Added SpiceDB for role-based access control, which made it possible to support secure multi-user setups and enterprise self-hosted deployments with a full audit trail.

Containerized all services with Docker and handled scalable deployments through Kubernetes and Helm, with optional AWS cloud hosting via ECS Fargate, EKS, and ECR.

Set up full GitOps CI/CD with GitHub Actions, ArgoCD, and Terraform so builds, tests, and infra changes were all automated and repeatable. Wired up Prometheus and Grafana for production observability. Project: Aidoc (real-time clinical AI for radiology triage and decision support)

Led the build of always-on GenAI services that processed medical imaging findings alongside patient context and flagged critical conditions for radiologists in real time, with the goal of faster, more consistent triage.

Used OpenAI models with PydanticAI to generate structured, clinician-readable explanations and escalation recommendations. Output schemas were tightly validated to hold up against medical safety requirements.

Built a multi-agent system with separate agents for triage, prioritization, explanation, and escalation, orchestrated through LangChain and LlamaIndex. This reduced decision variability and took real cognitive load off the clinical team.

Applied Model Context Protocol (MCP) to standardize how context flowed between LLM agents, ML models, and clinical rules engines so every decision was traceable and auditable, which was a strict requirement in this regulated environment.

Built ETL pipelines on AWS Glue to normalize clinical data from multiple sources, used Amazon Athena for cohort analysis, and ran Amazon Comprehend to extract clinical entities from unstructured text for richer LLM context.

Deployed containerized inference workloads on Amazon EKS for 24/7 uptime and exposed versioned REST APIs through AWS API Gateway so PACS, RIS, and hospital systems could consume AI outputs reliably.

Managed all infra with Terraform across dev, staging, and production. Set up OpenSearch pipelines to track inference latency, agent decisions, and model behavior for clinical safety compliance and ongoing audit. Project: Zest AI (explainable ML for credit underwriting and lending compliance)

Built and shipped credit risk models using gradient boosting and ensemble methods that improved lender approval rates in production across multiple bank customers while staying within regulatory explainability and fair-lending requirements.

Integrated SHAP-based feature attribution into every model so compliance and audit teams could see exactly why a lending decision was made. This was a core product requirement and had to be bulletproof.

Owned the full MLOps lifecycle on AWS covering data ingestion, feature engineering, training, validation, deployment, and drift detection across multi-tenant environments with version-controlled rollouts through GitHub Actions.

Exposed predictions and explanations through secure REST APIs that plugged directly into underwriting decision engines, with controlled and reversible rollout processes for each model update. Machine Learning Engineer Perficient Columbia, MD Feb 2012 - Apr 2018 Project: Sift (real-time fraud detection using global ML signals)

Built real-time fraud detection models in Python using scikit-learn and TensorFlow, combining classification, anomaly detection, and ensemble methods. Model scores were served live through gRPC and REST APIs inside Sift's core decisioning engine.

Ran large-scale feature engineering on AWS EMR with Spark and Redshift, with Kafka and Kinesis handling near-real-time signal aggregation across extremely high-volume transaction streams.

Monitored model health in production and worked closely with product and DevOps to reduce false positives over time. Getting the balance right between catching fraud and not blocking legitimate users was a constant challenge worth solving.

Project: K Health (AI-powered virtual care with clinical ML for triage and diagnosis)

Built clinical ML models for symptom analysis, patient triage, and diagnostic suggestions that ran in real time inside a virtual care platform used by millions of patients around the clock.

Developed NLP pipelines to parse patient-entered symptom text into structured clinical inputs that fed both diagnostic models and automated documentation workflows.

Deployed and scaled inference services on Docker and Kubernetes with Jenkins CI/CD, and worked directly with clinicians to validate triage logic against real-world safety and workflow requirements. EDUCATION

Bachelor of Science, Information Technology 2008 - 2012 Stratford University

AT A GLANCE

14+ Years ML Experience 5 Production LLM Systems Healthcare · Fintech · Fraud AWS-Native · Full MLOps



Contact this candidate