Post Job Free
Sign in

Staff AI Engineer - Foundation Models & LLM Ops

Location:
Burlingame, CA
Salary:
180000
Posted:
September 19, 2026

Contact this candidate

Resume:

Qiang Zhang

Staff AI Engineer Foundation Models LLM Systems AI Infrastructure

Burlingame, CA 940-***-**** **********.*****@*******.*** Linkedin GitHub

Professional Summary

Staff AI Engineer and Member of Technical Staff specializing in foundation models, LLM post-training, model evaluation, and

production AI systems. Leads research and engineering across reasoning, coding, instruction following, and structured generation,

including evaluation infrastructure that processes 100K+ benchmark samples and millions of generated tokens while cutting

validation cycles by 50%. Core contributor to Structured Outputs and experienced in building reliable, high-scale systems, including

99.99%-available payment platforms at Stripe. Brings deep research and distributed-systems expertise to improve model quality,

accelerate development, and deliver dependable AI products.

Technical Skills

• Foundation Models & Generative AI: Large Language Models (LLMs), Foundation Models, Generative AI, Transformer

Architectures, Reasoning Models, Multimodal AI, Vision-Language Models (VLMs), AI Agents, Agentic Workflows, Multi-Agent

Orchestration, Retrieval-Augmented Generation (RAG), GraphRAG, Embedding Models, Vector Search, Semantic Retrieval, Prompt

Engineering, Context Engineering, Tool Calling, Function Calling, Structured Outputs, Schema-Constrained Generation, Agentic AI

Applications, Production RAG Systems, AI Copilots

• LLM Post-Training & Alignment: LLM Post-Training, Supervised Fine-Tuning (SFT), Instruction Tuning, Reinforcement Learning

from Human Feedback (RLHF), Preference Optimization, Reward Modeling, Model Alignment, Behavioral Tuning, Model Behavior

Optimization, Synthetic Data Generation, Parameter-Efficient Fine-Tuning (PEFT), LoRA, QLoRA, Knowledge Distillation, Model

Quantization

• LLM Evaluation & Research: Model Evaluation, Capability Benchmarking, Behavioral Evaluation, Reasoning Evaluation-

, Coding Evaluation, Instruction-Following Evaluation, Structured-Generation Evaluation, Tool-Use Evaluation, Regression Testing-

, Continuous Evaluation, Benchmark Engineering, Failure Analysis, Hallucination Analysis, Factuality Assessment, Robustness

Testing, Safety Evaluation, AI Evaluation And Guardrails, Human Feedback Analysis, Quality Measurement, Error Analysis-

, Experimental Design, LLM Applications

• Machine Learning & Deep Learning: Machine Learning, Deep Learning, Natural Language Processing (NLP), Computer

Vision, Statistical Learning, Predictive Modeling, Representation Learning, Self-Supervised Learning, Transfer Learning-

, Reinforcement Learning, Classification, Regression, Ranking, Recommendation Systems, Time-Series Forecasting, Optimization,

Feature Engineering, Model Validation, Statistical Modeling

• AI/ML Frameworks & Inference: PyTorch, TensorFlow, JAX, Keras, Hugging Face Transformers, Hugging Face Datasets, Hugging

Face Accelerate, Scikit-learn, OpenCV, NumPy, Pandas, SciPy, XGBoost, LightGBM, ONNX Runtime, TensorRT, vLLM, MLflow,

Weights & Biases

• AI Infrastructure, Training & Serving: Distributed Training, Distributed Inference, GPU Computing, Multi-GPU Training, Model

Parallelism, Data Parallelism, GPU Resource Management, Model Serving, Inference Optimization, Batch Inference, Dynamic

Batching, Continuous Batching, GPU Memory Optimization, KV-Cache Optimization, Quantized Inference, Low-Latency Inference,

High-Throughput Inference, Model Deployment, Enterprise GenAI Applications

• MLOps & LLMOps: MLOps, LLMOps, Experiment Tracking, Model Versioning, Model Evaluation Pipelines, Training Pipelines,

Data Pipelines, Continuous Integration, Continuous Evaluation, Continuous Deployment, Model Registry, Production Validation,

Model Monitoring, AI Observability, Experiment Automation, Model Performance Tracking

• Distributed Systems & Backend Engineering: Distributed Systems, Microservices, Service-Oriented Architecture, Event-Driven

Architecture, API Design, REST APIs, gRPC, Asynchronous Processing, Concurrent Programming, High-Throughput Services-

, Low-Latency Systems, Fault-Tolerant Architecture, Scalability Engineering, Load Balancing, Caching, Service Decomposition,

Performance Optimization

• Cloud & Platform Engineering: Amazon Web Services (AWS), Google Cloud Platform (GCP), Microsoft Azure, AWS SageMaker,

Google Vertex AI, Kubernetes, Docker, Kubeflow, Terraform, Pulumi, Infrastructure as Code, Container Orchestration, GPU Workload

Management, CI/CD, Production Deployment, Influencing Across Teams

• Data & AI Data Platforms: PostgreSQL, MySQL, Redis, Kafka, Vector Databases, Knowledge Graphs, Data Modeling-

, Data Pipelines, ETL/ELT, Streaming Systems, Database Optimization, Query Optimization, Schema Design, Indexing, Feature

Engineering, Caching, Apache Spark, Databricks, Airflow

• Programming Languages: Python, Go, Java, SQL, TypeScript, JavaScript, Rust, C#, Bash, React, Java, CUDA, C++

PROFESSIONAL EXPERIENCE

OpenAI

Member of Technical Staff – Foundation Models & AI Systems Mar 2024 - Present

LLM Post-Training API Model Research Reasoning Models Model Evaluation

• Led post-training and API model research for frontier foundation models across four core capability areas-reasoning, coding, instruction

following, and structured/tool-based generation-improving developer-facing model quality.

• Architected LLM evaluation infrastructure processing 100K+ benchmark samples and millions of generated tokens, expanding

automated coverage across 5+ evaluation dimensions and reducing model-validation cycles by approximately 50% while strengthening

AI evaluation and guardrails.

• Automated benchmark execution, regression analysis, behavioral testing, and quality reporting across model iterations, accelerating

evaluation throughput by approximately 2 and shortening research iteration cycles by 50% across development and validation

workflows.

• Served as a core contributor to Structured Outputs, developing evaluation-driven methodologies for JSON Schema, function calling,

and structured responses, improving measured schema-conformance reliability by approximately 20% across targeted workloads.

• Analyzed 5+ major LLM failure categories, including hallucination, instruction conflict, malformed structured output, reasoning

inconsistency, and tool-use errors; converted recurring failures into targeted test suites and guardrails that increased measured reliability

by approximately 15%.

• Contributed to o3-mini training and post-training, evaluating reasoning consistency, instruction adherence, and model behavior across

large-scale experiments and identifying quality gaps for subsequent training iterations.

• Contributed to GPT-4.1 as a research core contributor, analyzing 4+ capability areas spanning coding, reasoning, instruction following,

and long-context/API workflows to quantify model-quality changes across research iterations.

• Created Python research-engineering tooling for experiment orchestration, benchmark execution, result aggregation, and failure

analysis, cutting manual evaluation effort by approximately 40% and increasing researcher iteration capacity.

• Standardized reusable evaluation components and quality metrics across multiple model experiments, reducing duplicated evaluation

engineering by approximately 30% and improving consistency between release candidates.

• Connected research findings with production-oriented API workflows across 3 major capability areas-structured generation, reasoning,

and tool use-shortening research-to-validation handoffs by approximately 25%

Stripe

Staff Software Engineer – Distributed Systems & Platform Engineering Jul 2023 - Mar 2024

Distributed Systems Backend Architecture Reliability Engineering

• Led technical architecture for high-volume backend services, increasing platform scalability by approximately 35% through service

decomposition, workload balancing, and performance engineering across distributed systems.

• Optimized distributed services across 6 core technologies-Go, Java, Kafka, PostgreSQL, Redis, and Kubernetes-raising processing

capacity by approximately 30% under production workloads.

• Reworked asynchronous processing, caching, and fault-isolation paths, increasing transaction-processing efficiency by approximately

30% while maintaining consistency requirements.

• Introduced monitoring, automated recovery, and failure-isolation mechanisms across critical services, lowering recurring incidents by

approximately 30%.

• Delivered reusable platform components and engineering standards that reduced implementation effort by approximately 25% across

backend development workflows.

• Directed architecture reviews and cross-team technical decisions for multiple production services, influencing technical direction across

teams and extending scalability and reliability practices across the engineering organization.

Senior Software Engineer – Distributed Backend Systems Jul 2019 - Jun 2023

Backend Services APIs Performance Engineering

• Delivered highly available backend services supporting millions of daily payment transactions, maintaining 99.99% availability across

critical workloads.

• Increased service throughput by approximately 30% through concurrency optimization, service decomposition, workload balancing,

and backend performance tuning.

• Refined database schemas, indexing, query execution, and caching strategies, lowering critical API latency by approximately 25%.

• Expanded production observability across 3 primary telemetry areas-metrics, logging, and distributed tracing-improving

incident-diagnosis efficiency by approximately 35%.

• Strengthened automated testing, CI/CD, and deployment validation, cutting regression-related production issues by approximately

25%.

• Participated in system design, capacity planning, and reliability initiatives across multiple distributed services, enabling a 20% increase

in request capacity while maintaining zero downtime

Software Engineer – Backend Infrastructure Jul 2017 - Jun 2019

Backend Development APIs Platform Engineering

• Implemented backend services and APIs for payment workflows and internal platforms, increasing transaction-processing throughput

by approximately 15%.

• Created reusable service libraries and platform components covering multiple recurring backend capabilities, reducing implementation

effort by approximately 20%.

• Tuned SQL queries, indexing, database access, and caching paths, lowering latency across selected high-traffic services by

approximately 30%.

• Introduced automated testing and continuous deployment practices, reducing production incidents by approximately 25% and

improving release consistency.

Luxe

Data Scientist – Machine Learning Systems Apr 2015 - Jul 2017

Real-Time Dispatch ETA Prediction Applied Machine Learning

• Developed machine-learning solutions for 2 core real-time systems-dispatch and ETA prediction, improving prediction quality and

operational efficiency by approximately 20%.

• Established end-to-end ML workflows across 5 stages-data preparation, feature engineering, training, validation, and production

integration-reducing experimentation cycles by approximately 30%.

• Applied statistical modeling and optimization to routing, resource allocation, and fleet-utilization problems, contributing to

approximately 15% greater operational efficiency.

• Analyzed prediction errors and operational outcomes across multiple decision variables, increasing the effectiveness of data-driven

dispatch decisions by approximately 20%.

• Translated predictive-model outputs into real-time operational workflows with engineering and operations teams, reducing the gap

between model development and deployment by approximately 25%

EDUCATION

University of California, Los Angeles (UCLA) 2013 - 2015

Master's Degree, Statistics

Nankai University 2009 - 2013

Bachelor's degree, Math & Statistics



Contact this candidate