Post Job Free
Sign in

Senior Software Engineer, AI Systems & Data

Location:
Santa Clara, CA
Salary:
200,000
Posted:
August 02, 2026

Contact this candidate

Resume:

Hechao Li Software Engineer

********.****@*****.*** +1-509-***-**** Milpitas, CA linkedin.com/in/hechaoli SUMMARY

Senior Software Engineer with 10+ years of experience building large-scale distributed systems, AI-driven backend platforms, and production-grade data infrastructure across OpenAI, Netflix, and Meta . Led architecture and development of high-throughput systems including agentic AI platforms, LLM-integrated search systems, and distributed blockchain infrastructure, processing millions of requests and data points at scale. Specialized in LLM systems (RAG, agent orchestration), high-cardinality data pipelines, and low- latency distributed services, with strong focus on scalability, reliability, and real-time performance in cloud- native environments.

PROFESSIONAL EXPERIENCE

OpenAI, Member of Technical Staff 06/2025 – Present San Francisco, CA Project: Deep Research API

Autonomous Agentic Intelligence Platform Enterprise-scale agentic AI platform enabling multi-step autonomous research workflows across web-scale and enterprise data sources. The system orchestrates distributed microservices for autonomous browsing, document ingestion, vector retrieval, and LLM-based synthesis, generating citation-backed reports from hundreds of sources.

•Led backend architecture for a distributed agent orchestration system using Agents SDK, designing DAG- based planning, execution lifecycle, and fault-tolerant workflows for long-running AI tasks.

•Designed scalable document ingestion pipelines handling PDFs, HTML, and structured data with OCR, chunking, and embedding generation for large-scale semantic indexing.

•Built high-throughput vector retrieval system using embedding models and vector databases, enabling low- latency semantic search across millions of document chunks.

•Implemented recursive agent reasoning loop combining LLM planning, tool invocation, and feedback- driven query refinement to improve multi-hop reasoning accuracy.

•Developed RAG pipelines integrating retrieval, summarization, and synthesis to produce citation-grounded outputs with high factual consistency.

•Designed context window optimization framework using hierarchical summarization and ranking strategies to maximize signal under token constraints.

•Built production-grade LLM orchestration layer leveraging GPT-4.5 / o3 models, supporting structured outputs, tool usage, and multi-step reasoning.

•Improved accuracy and reliability by implementing retrieval validation, citation verification, and hallucination detection pipelines.

•Scaled asynchronous execution using event-driven architecture with distributed task queues and background job orchestration.

•Instrumented end-to-end AI observability pipelines tracking token usage, retrieval quality, latency distributions, and model performance metrics.

Technologies Used: Python, TypeScript, FastAPI, gRPC, LLMs (GPT-4.5, o3), RAG, Vector DB, Agents SDK, MCP, Kafka, Redis, PostgreSQL, Docker, Kubernetes, AWS/GCP, Prometheus, Grafana Hechao Li ********.****@*****.***

Netflix, Senior Software Engineer 05/2022 – 05/2025 Los Gatos, CA Project: AI-Powered Conversational Search (Unicorn Ranking Platform) Large-scale conversational search system transforming Netflix discovery into a natural language-driven experience using LLMs, semantic retrieval, and unified ranking pipelines .

•Led backend development of LLM-integrated search orchestration services, enabling intent extraction, query rewriting, and conversational refinement.

•Designed RAG-based pipeline combining LLM query understanding with semantic retrieval over large- scale content embeddings.

•Built scalable vector search infrastructure using Faiss/OpenSearch, supporting approximate nearest- neighbor retrieval over millions of embeddings.

•Developed distributed content ingestion pipelines using Maestro, processing scripts, subtitles, and metadata into embeddings and knowledge graph features.

•Implemented Unified Contextual Ranker (UniCoRn) inference integration, combining user behavior, query context, and content signals for personalized ranking.

•Designed microservices-based query pipeline (Java/Spring Boot) with REST/gRPC APIs and service mesh routing for high availability.

•Built feature engineering pipelines integrating knowledge graph embeddings and user signals into ranking models.

•Optimized latency-critical serving path, achieving sub-second response times under high concurrent traffic.

•Integrated OpenAI LLM Gateway for intent parsing and response generation, enabling conversational UX.

•Scaled backend using cloud-native infrastructure (AWS, Kubernetes, Envoy) with auto-scaling and fault isolation.

•Collaborated with ML platform teams using Metaflow, Ray, and distributed training systems for model lifecycle management.

Technologies Used: Java, Spring Boot, LLMs (OpenAI GPT), RAG, Vector Search (Faiss/OpenSearch), Kafka, gRPC, REST APIs, Kubernetes, AWS, Envoy, Metaflow, Ray, DynamoDB, Cassandra Facebook, Software Engineer 10/2018 – 05/2022 Menlo Park, CA Project: Diem (Libra)

Distributed Blockchain Infrastructure Permissioned BFT blockchain platform enabling global payments with distributed validator nodes, deterministic execution, and cryptographically verifiable state .

•Designed distributed transaction processing pipeline spanning admission control, mempool propagation, consensus, execution, and storage layers.

•Built JSON-RPC admission services for transaction validation, signature verification, and stateless pre- check enforcement.

•Developed shared mempool system with gossip protocol, ensuring consistent transaction propagation across validator nodes.

•Implemented DiemBFT consensus protocol (HotStuff-based), achieving fault-tolerant agreement under Byzantine conditions.

•Built execution pipeline using MoveVM, ensuring deterministic state transitions and preventing double- spending.

•Designed storage layer (DiemDB) on RocksDB with Merkle Tree-based authenticated state and versioned ledger.

•Optimized consensus throughput and latency via batching, leader rotation, and pipeline parallelism.

•Architected multi-region distributed deployment across AWS/GCP/Azure, ensuring high availability and fault isolation.

Hechao Li ********.****@*****.***

Technologies Used: Rust, DiemBFT, MoveVM, RocksDB, Merkle Trees, gRPC, JSON-RPC, Distributed Systems, Consensus Algorithms, Multi-Cloud

VMware, Member of Technical Staff 02/2017 – 10/2018 San Francisco, CA Project: Wavefront

Distributed Metrics & Observability Platform High-scale time-series analytics platform processing millions of data points per second with sub-second query latency, enabling high-cardinality observability and real- time anomaly detection .

•Designed proxy-based ingestion system using Java, enabling secure TLS push model with buffering and disk spillover for lossless data collection.

•Built distributed time-series storage engine (TSDB) supporting high-cardinality metrics and real time querying at scale.

•Implemented T-Digest-based histogram storage, enabling accurate percentile computation (p95, p99) for latency analytics.

•Optimized query engine (WQL) with align functions and streaming aggregation algorithms for cross-metric correlation.

•Developed AI-driven anomaly detection (AI Genie) using dynamic baselines and time-series forecasting models.

•Built observability pipelines powering dashboards, alerts, and real-time analytics across distributed systems.

Technologies Used: Java, TSDB, T-Digest, Kafka, Kubernetes, Prometheus, Time-Series Analytics, Machine Learning, Distributed Systems

Carnegie Mellon University, Teaching Assistant 01/2016 – 12/2016 Pittsburgh, PA EDUCATION

Carnegie Mellon University, Master's Degree, Computer Science 08/2015 – 12/2016 Pittsburgh, PA Beihang University, Bachelor's Degree, Computer Science 09/2011 – 07/2015 Beijing, China SKILLS

Languages (Python, Java, TypeScript, C++, Rust, SQL) Frameworks (FastAPI, Spring Boot, Node.js, gRPC, REST APIs) AI & Data (LLMs (OpenAI GPT, o3), RAG, Agentic AI, Vector Search, Embedding Models, Time- Series Analytics, T-Digest, NLP, Information Retrieval, Knowledge Graph) Cloud & DevOps (AWS, GCP, Docker, Kubernetes, Terraform, Kafka, Envoy, Prometheus, Grafana) Distributed Systems (Microservices Architecture, Event-Driven Systems, Consensus (BFT), High-Throughput Pipelines, Fault Tolerance, Multi- Cloud Infrastructure) Machine Learning Systems (Model Serving, Feature Engineering, Ranking Systems

(UniCoRn), Anomaly Detection, Context Optimization, LLM Evaluation & Monitoring)

Hechao Li ********.****@*****.***



Contact this candidate