Post Job Free
Sign in

Senior Software Engineer

Location:
Santa Clara, CA
Posted:
October 02, 2026

Contact this candidate

Resume:

Zhiyu Li

*********@*****.*** +1-432-***-**** Seattle, WA linkedin.com/in/zhiyul

PROFESSIONAL SUMMARY

Senior AI Infrastructure & Systems Engineer with over 7 years of expertise architecting high-performance distributed training systems, scalable microservice backends, and containerized orchestration ecosystems. Proven track record optimizing extra-large foundation models across world-class infrastructure, including NVIDIA Blackwell GPU clusters and 10-exaflop Google Cloud TPU Pods, consistently pushing Model FLOPs Utility (MFU) past 85%. Specialized in building production-grade post-training Reinforcement Learning (RL) frameworks (PPO/GRPO), engineering decoupled rollout- trainer pipelines, and designing highly available backend APIs and distributed file systems for enterprise ML pipelines. Technical anchor with deep mastery across JAX, PyTorch, C++, Python, Java, C#, Megatron-Core, and Ray. SKILLS

Programming Languages Python, Go (Golang), Java, C#, C++, TypeScript, Lua, Bash, R AI / Machine Learning JAX, PyTorch, TorchServe, Triton Inference Server, Whisper ASR, LLaMA, CodeGen, RAG, Hugging Face, OpenNLP, MLOps, LangChain, TensorFlow, LLMs, Prompt Engineering, Semantic Search, GPU, Megartron-Core, Ray, DeepSpeed, NeMo, Keras Backend & API Development Microservices, RESTful API, gRPC, Protocol Buffers (Protobuf), OpenAPI Standard (OAS) Cloud & Infrastructure Kubernetes, Docker, Azure, Google Cloude, AWS (Lambda, EC2, S3, ECS, IAM, CloudWatch), EKS, Terraform, CI/CD

Data Engineering & Streaming Databricks, Apache Spark, Apache Kafka, Redis, Supabase, PostgreSQL, MySQL, NoSQL, Vector Databases, Data Pipelines

Frontend Frameworks React.js, Next.js, Nuxt.js, Telemetry Dashboards, UI/UX Components Testing & Methodologies A/B Testing, Unit Testing, Integration Testing, Mock Frameworks, Agile/Scrum System Performance & Safety Performance Profiling, Memory Management, Thread Pooling, High Availability, Fault Tolerance, TDIR, UEBA

WORK EXPERIENCE

Senior Software Engineer NVIDIA Jan 2025 – Jun 2026 Santa Clara, CA

Engineered scalability and performance optimizations for core post-training Reinforcement Learning (RL) frameworks, focusing on large-scale LLM alignment algorithms including PPO and GRPO.

Architected decoupled rollout-trainer pipelines to eliminate GPU data-starvation bottlenecks, seamlessly overlapping auto-regressive generation and distributed training phases across massive GPU clusters.

Optimized 4D-parallelism strategies (tensor, pipeline, data, and context parallelism) using Megatron-Core and Ray backends to maximize Model FLOPs Utility (MFU) during large-scale RL alignment loops.

Integrated low-latency inference accelerations, including speculative decoding and advanced quantization techniques, into the RL actor loop to drastically slash step times during training-time rollout generation.

Enhanced cluster communication and orchestration layers, tuning network dynamic load balancing and distributed memory management to ensure linear scaling across hundreds of Blackwell GPU nodes. Software Engineer Google Jun 2023 – Jan 2025

Greater Seattle Area, WA

Optimized extra-large model distributed training across multiple Cloud TPU Pods within a dedicated multi-pod infrastructure team.

Open-sourced a new end-to-end GPT3-175B model training pipeline in JAX, offering high performance and unlimited scalability tailored for MLPerf submission leaderboards.

Achieved top Model FLOPs Utility (MFU) surpassing 60% for bf16 and 85% for int8 by adopting cross-team contributions in algorithms, networks, distributed file systems, and compilers.

Diagnosed and enhanced system performance by 25% through deep investigation into int8 quantization, parallelism strategy optimization, compiler tuning, and network dynamic load balancing.

Conducted scaling experiments demonstrating nearly linear scalability across single and multi-slice configurations on a 10-exaflop TPU cluster utilizing tens of thousands of v5e/v5p TPUs. Senior Software Engineer SoundHound Inc. Mar 2020 – Jun 2023 Santa Clara County, CA

Served as Team Lead for Language Model Research and Project Lead of Neural Language Models for the core speech engine and ML pipeline systems.

Led innovations and deployments of GPT-like Neural Language Models across 4 languages, improving speech engine accuracy by 20% on average for automotive, smart device, and restaurant partners.

Accomplished an end-to-end Kubeflow K8s pipeline for Neural Language Models covering data processing, training, fine-tuning, benchmarking, and TensorFlow Lite deployment in a C++ environment.

Designed a configuration hierarchy using Hydra for scaling new language services and ablation experiments; extended MLOps plug-ins for knowledge distillation, hyperband/grid search hyperparameter tuning, and TFLite quantization.

Prototyped a restaurant conversation LLM in PyTorch by prompt-tuning 6B GPT-J and 20B GPT-NeoX models; achieved GPU memory efficiency via DeepSpeed ZeRO-powered model parallelism.

Implemented a K8s cluster file storage system merging local and remote S3 storage using Boto3, expanding an open- source multi-storage streaming library to handle Kubeflow data loading and checkpointing.

Developed a distributed data processing pipeline on hundreds of GB of corpus data using Hadoop and Spark; built text post-fine-tuning services for an end-to-end RNN transducer speech engine. Software Engineer CERT Division at the Software Engineering Institute May 2019 – Sep 2019 Pittsburgh, PA

Led a team of 5 engineers to create a customized tool platform for malware feature research utilized by cybersecurity experts.

Developed a pipeline class module in Python allowing customized designs of arbitrary data transformation steps, model training, prediction, analysis, and visualization with partial execution flexibility.

Designed an abstract class system for engineering tools to ensure seamless extensibility and compatibility across both Scikit-Learn and Keras libraries.

Implemented configuration and parser class modules in Python to manage command-line parameters and JSON files for customized parameters saving and loading.

Graduate Student Researcher Carnegie Mellon University May 2019 – Sep 2019 Pittsburgh, PA

Researched structural features in deep neural networks, designing tree-structured multi-channel CNN networks, implementing Tree-LSTM, and developing customized tree-matrix attention mechanisms.

Optimized and fine-tuned a pre-trained BERT model with layer-wise decreasing learning rates and bucket data loaders, ranking 288/3165 in a Kaggle competition with under 1.5 hours of training time.

Applied spell correction and forbidden words estimation preprocessing to decrease the out-of-vocabulary rate by 30%.

EDUCATION

Master’s Degree in Computer Science

Carnegie Mellon University Pittsburgh, PA

2018 - 2019

Bachelor’s Degree in Engineering Physics

Shanghai Jiaotong University Shanghai, China

2009 - 2013



Contact this candidate