Qiang Zhang
Staff AI Engineer Foundation Models LLM Systems AI Infrastructure
Burlingame, CA 940-***-**** **********.*****@*******.*** Linkedin GitHub
Professional Summary
Staff AI Engineer and Member of Technical Staff specializing in foundation models, LLM post-training, model evaluation, and
production AI systems. Leads research and engineering across reasoning, coding, instruction following, and structured generation,
including evaluation infrastructure that processes 100K+ benchmark samples and millions of generated tokens while cutting
validation cycles by 50%. Core contributor to Structured Outputs and experienced in building reliable, high-scale systems, including
99.99%-available payment platforms at Stripe. Brings deep research and distributed-systems expertise to improve model quality,
accelerate development, and deliver dependable AI products.
Technical Skills
• Foundation Models & Generative AI: Large Language Models (LLMs), Foundation Models, Generative AI, Transformer
Architectures, Reasoning Models, Multimodal AI, Vision-Language Models (VLMs), AI Agents, Agentic Workflows, Multi-Agent
Orchestration, Retrieval-Augmented Generation (RAG), GraphRAG, Embedding Models, Vector Search, Semantic Retrieval, Prompt
Engineering, Context Engineering, Tool Calling, Function Calling, Structured Outputs, Schema-Constrained Generation, Agentic AI
Applications, Production RAG Systems, AI Copilots
• LLM Post-Training & Alignment: LLM Post-Training, Supervised Fine-Tuning (SFT), Instruction Tuning, Reinforcement Learning
from Human Feedback (RLHF), Preference Optimization, Reward Modeling, Model Alignment, Behavioral Tuning, Model Behavior
Optimization, Synthetic Data Generation, Parameter-Efficient Fine-Tuning (PEFT), LoRA, QLoRA, Knowledge Distillation, Model
Quantization
• LLM Evaluation & Research: Model Evaluation, Capability Benchmarking, Behavioral Evaluation, Reasoning Evaluation-
, Coding Evaluation, Instruction-Following Evaluation, Structured-Generation Evaluation, Tool-Use Evaluation, Regression Testing-
, Continuous Evaluation, Benchmark Engineering, Failure Analysis, Hallucination Analysis, Factuality Assessment, Robustness
Testing, Safety Evaluation, AI Evaluation And Guardrails, Human Feedback Analysis, Quality Measurement, Error Analysis-
, Experimental Design, LLM Applications
• Machine Learning & Deep Learning: Machine Learning, Deep Learning, Natural Language Processing (NLP), Computer
Vision, Statistical Learning, Predictive Modeling, Representation Learning, Self-Supervised Learning, Transfer Learning-
, Reinforcement Learning, Classification, Regression, Ranking, Recommendation Systems, Time-Series Forecasting, Optimization,
Feature Engineering, Model Validation, Statistical Modeling
• AI/ML Frameworks & Inference: PyTorch, TensorFlow, JAX, Keras, Hugging Face Transformers, Hugging Face Datasets, Hugging
Face Accelerate, Scikit-learn, OpenCV, NumPy, Pandas, SciPy, XGBoost, LightGBM, ONNX Runtime, TensorRT, vLLM, MLflow,
Weights & Biases
• AI Infrastructure, Training & Serving: Distributed Training, Distributed Inference, GPU Computing, Multi-GPU Training, Model
Parallelism, Data Parallelism, GPU Resource Management, Model Serving, Inference Optimization, Batch Inference, Dynamic
Batching, Continuous Batching, GPU Memory Optimization, KV-Cache Optimization, Quantized Inference, Low-Latency Inference,
High-Throughput Inference, Model Deployment, Enterprise GenAI Applications
• MLOps & LLMOps: MLOps, LLMOps, Experiment Tracking, Model Versioning, Model Evaluation Pipelines, Training Pipelines,
Data Pipelines, Continuous Integration, Continuous Evaluation, Continuous Deployment, Model Registry, Production Validation,
Model Monitoring, AI Observability, Experiment Automation, Model Performance Tracking
• Distributed Systems & Backend Engineering: Distributed Systems, Microservices, Service-Oriented Architecture, Event-Driven
Architecture, API Design, REST APIs, gRPC, Asynchronous Processing, Concurrent Programming, High-Throughput Services-
, Low-Latency Systems, Fault-Tolerant Architecture, Scalability Engineering, Load Balancing, Caching, Service Decomposition,
Performance Optimization
• Cloud & Platform Engineering: Amazon Web Services (AWS), Google Cloud Platform (GCP), Microsoft Azure, AWS SageMaker,
Google Vertex AI, Kubernetes, Docker, Kubeflow, Terraform, Pulumi, Infrastructure as Code, Container Orchestration, GPU Workload
Management, CI/CD, Production Deployment, Influencing Across Teams
• Data & AI Data Platforms: PostgreSQL, MySQL, Redis, Kafka, Vector Databases, Knowledge Graphs, Data Modeling-
, Data Pipelines, ETL/ELT, Streaming Systems, Database Optimization, Query Optimization, Schema Design, Indexing, Feature
Engineering, Caching, Apache Spark, Databricks, Airflow
• Programming Languages: Python, Go, Java, SQL, TypeScript, JavaScript, Rust, C#, Bash, React, Java, CUDA, C++
PROFESSIONAL EXPERIENCE
OpenAI
Member of Technical Staff – Foundation Models & AI Systems Mar 2024 - Present
LLM Post-Training API Model Research Reasoning Models Model Evaluation
• Led post-training and API model research for frontier foundation models across four core capability areas-reasoning, coding, instruction
following, and structured/tool-based generation-improving developer-facing model quality.
• Architected LLM evaluation infrastructure processing 100K+ benchmark samples and millions of generated tokens, expanding
automated coverage across 5+ evaluation dimensions and reducing model-validation cycles by approximately 50% while strengthening
AI evaluation and guardrails.
• Automated benchmark execution, regression analysis, behavioral testing, and quality reporting across model iterations, accelerating
evaluation throughput by approximately 2 and shortening research iteration cycles by 50% across development and validation
workflows.
• Served as a core contributor to Structured Outputs, developing evaluation-driven methodologies for JSON Schema, function calling,
and structured responses, improving measured schema-conformance reliability by approximately 20% across targeted workloads.
• Analyzed 5+ major LLM failure categories, including hallucination, instruction conflict, malformed structured output, reasoning
inconsistency, and tool-use errors; converted recurring failures into targeted test suites and guardrails that increased measured reliability
by approximately 15%.
• Contributed to o3-mini training and post-training, evaluating reasoning consistency, instruction adherence, and model behavior across
large-scale experiments and identifying quality gaps for subsequent training iterations.
• Contributed to GPT-4.1 as a research core contributor, analyzing 4+ capability areas spanning coding, reasoning, instruction following,
and long-context/API workflows to quantify model-quality changes across research iterations.
• Created Python research-engineering tooling for experiment orchestration, benchmark execution, result aggregation, and failure
analysis, cutting manual evaluation effort by approximately 40% and increasing researcher iteration capacity.
• Standardized reusable evaluation components and quality metrics across multiple model experiments, reducing duplicated evaluation
engineering by approximately 30% and improving consistency between release candidates.
• Connected research findings with production-oriented API workflows across 3 major capability areas-structured generation, reasoning,
and tool use-shortening research-to-validation handoffs by approximately 25%
Stripe
Staff Software Engineer – Distributed Systems & Platform Engineering Jul 2023 - Mar 2024
Distributed Systems Backend Architecture Reliability Engineering
• Led technical architecture for high-volume backend services, increasing platform scalability by approximately 35% through service
decomposition, workload balancing, and performance engineering across distributed systems.
• Optimized distributed services across 6 core technologies-Go, Java, Kafka, PostgreSQL, Redis, and Kubernetes-raising processing
capacity by approximately 30% under production workloads.
• Reworked asynchronous processing, caching, and fault-isolation paths, increasing transaction-processing efficiency by approximately
30% while maintaining consistency requirements.
• Introduced monitoring, automated recovery, and failure-isolation mechanisms across critical services, lowering recurring incidents by
approximately 30%.
• Delivered reusable platform components and engineering standards that reduced implementation effort by approximately 25% across
backend development workflows.
• Directed architecture reviews and cross-team technical decisions for multiple production services, influencing technical direction across
teams and extending scalability and reliability practices across the engineering organization.
Senior Software Engineer – Distributed Backend Systems Jul 2019 - Jun 2023
Backend Services APIs Performance Engineering
• Delivered highly available backend services supporting millions of daily payment transactions, maintaining 99.99% availability across
critical workloads.
• Increased service throughput by approximately 30% through concurrency optimization, service decomposition, workload balancing,
and backend performance tuning.
• Refined database schemas, indexing, query execution, and caching strategies, lowering critical API latency by approximately 25%.
• Expanded production observability across 3 primary telemetry areas-metrics, logging, and distributed tracing-improving
incident-diagnosis efficiency by approximately 35%.
• Strengthened automated testing, CI/CD, and deployment validation, cutting regression-related production issues by approximately
25%.
• Participated in system design, capacity planning, and reliability initiatives across multiple distributed services, enabling a 20% increase
in request capacity while maintaining zero downtime
Software Engineer – Backend Infrastructure Jul 2017 - Jun 2019
Backend Development APIs Platform Engineering
• Implemented backend services and APIs for payment workflows and internal platforms, increasing transaction-processing throughput
by approximately 15%.
• Created reusable service libraries and platform components covering multiple recurring backend capabilities, reducing implementation
effort by approximately 20%.
• Tuned SQL queries, indexing, database access, and caching paths, lowering latency across selected high-traffic services by
approximately 30%.
• Introduced automated testing and continuous deployment practices, reducing production incidents by approximately 25% and
improving release consistency.
Luxe
Data Scientist – Machine Learning Systems Apr 2015 - Jul 2017
Real-Time Dispatch ETA Prediction Applied Machine Learning
• Developed machine-learning solutions for 2 core real-time systems-dispatch and ETA prediction, improving prediction quality and
operational efficiency by approximately 20%.
• Established end-to-end ML workflows across 5 stages-data preparation, feature engineering, training, validation, and production
integration-reducing experimentation cycles by approximately 30%.
• Applied statistical modeling and optimization to routing, resource allocation, and fleet-utilization problems, contributing to
approximately 15% greater operational efficiency.
• Analyzed prediction errors and operational outcomes across multiple decision variables, increasing the effectiveness of data-driven
dispatch decisions by approximately 20%.
• Translated predictive-model outputs into real-time operational workflows with engineering and operations teams, reducing the gap
between model development and deployment by approximately 25%
EDUCATION
University of California, Los Angeles (UCLA) 2013 - 2015
Master's Degree, Statistics
Nankai University 2009 - 2013
Bachelor's degree, Math & Statistics