Thriveen Kumar
AI/ML Engineer
314-***-**** ***************@*****.*** San Francisco, CA LinkedIn Summary
AI/ML Engineer with 5+ years building scalable ML infrastructure and production-grade AI systems. Experience with cloud-native and edge deployments (AWS, Edge AI, GPUs) across large-scale environments.Expert in designing batch/streaming pipelines and feature- engineering frameworks using Spark, Kafka, Kubernetes, etc. Strong skills in MLOps (Triton, TorchServe, MLflow) and GenAI architectures (LLMs, retrieval-based systems).Deployed Retrieval- Augmented Generation (RAG) pipelines integrating semantic vector search with contextual web indexing, powering grounded summarization in an AI assistant. Implemented chain-of-thought prompting and LangChain-style agents to improve answer relevance.Integrated on-device lightweight language models for low-latency tasks, routing complex queries to cloud LLM endpoints (AWS Bedrock, Vertex AI) to maintain sub-second response.Built hybrid retrieval pipelines (dense FAISS + BM25) with re-ranking layers to boost precision and factual consistency, mitigating hallucinations. Led experiments on Responsible AI (reducing bias) and content filtering (PII detection) to enforce safe, compliant outputs. Professional Experience
Perplexity Apr 2024 – Present
AI/ML Engineer San Francisco, CA
Architected RAG pipelines combining FAISS vector database semantic search and web scraping to power an AI assistant summarization feature. Introduced LangChain vector agents for multi-step query analysis, improving answer relevance by 15%.
Deployed on-device language model inference for latency-sensitive tasks and dynamically routed complex queries to cloud LLM endpoints (AWS Bedrock, GCP Vertex) to maintain sub-second response times.
Designed scalable inference pipeline using Triton Server and Kubernetes/Docker GPU nodes, improving throughput by 25% at peak load.
Integrated on-device LLMs (TensorRT-optimized) for latency-sensitive tasks and routed complex queries to AWS Bedrock and cloud endpoints to meet sub-second SLAs.
Designed a hybrid retrieval pipeline using FAISS for dense vector search and lexical matching (BM25), with a reranking layer
(heuristic + cross-encoder) to boost precision, reduce hallucinations, and improve end-to-end response quality in RAG workflows.
Developed passage extraction, ranking, and citation normalization pipelines to enforce source attribution rules, reducing hallucination-prone summaries and increasing factual consistency in A/B evaluations.
Implemented audio-to-intent pipelines combining speech recognition, intent classification, and NLU components, carefully managing end-to-end latency budgets for conversational browsing workflows.
Added privacy-preserving filters (PII detection, token anonymization) to protect user data, aligning with Responsible AI guidelines.
Instrumented production ML evaluation frameworks (A/B testing) with monitoring of latency, accuracy, and user satisfaction, enabling data-driven R&D and performance optimization.
Built privacy-aware filtering systems incorporating PII detection and secure token management to ensure browser context safety while supporting compliant personalization capabilities.
Collaborated with cross-functional teams (frontend, UX, SRE) on telemetry, staged feature rollouts, and CI/CD pipelines; built observability dashboards (Prometheus, Grafana) to ensure scalable, reliable production systems.
Led rapid 0 to 1 feature experiments for Comet’s agentic browsing capabilities, delivering measurable engagement uplift of approximately 18% during initial launch expansion phases. Nvidia Feb 2020 – Jul 2023
Machine Learning Engineer India
Developed GPU-accelerated ML models (NLP + CV) in PyTorch/TF, cutting inference latency 35% for fraud detection and autonomous systems.
Fine-tuned transformer models (BERT, GPT, mBERT, XLM-R) for multilingual NLP tasks, boosting accuracy 20%. Worked with Hugging Face Transformers (GPT-3.5, Stable Diffusion).
Engineered real-time data pipelines (Kafka, Spark) for fraud detection, integrating anomaly detection and graph analytics into a scalable streaming system.
Architected distributed inference services (Triton, TensorRT) on Kubernetes/Docker GPU clusters, implementing microservices architecture for low-latency model deployment.
Built full-stack personalization pipelines (Kafka streams, Redis, vector database) with recommendation models and real-time API endpoints.
Established MLOps workflows: MLflow tracking, CI/CD, drift monitoring, and blue-green rollouts for reliable model updates.
Led multimodal media intelligence projects: deep learning, speech-to-text, and content classification integrated into dashboards.
Deployed cloud-native GPU inference clusters (AWS EKS) with autoscaling, multi-AZ high availability, and IAM-based access control for secure production deployment.
Built end-to-end cloud-native pipelines (S3, Redshift, DynamoDB, Lambda) linking ML inference to payment and AdTech engines.
Delivered improved models for NLP/CV/autonomous driving using PyTorch, TensorRT, ONNX, improving overall accuracy by 19%.
Developed end-to-end ML pipelines (data ingestion, training, model deployment) on cloud (S3, Redshift, Lambda), cutting execution time 40% while improving accuracy.
Created scalable video analytics with NVIDIA DeepStream SDK, reducing latency 33% for real-time retail/surveillance applications. Technical Skills
Programming & Data Languages: Python, SQL, Scala, Java, JavaScript, TypeScript, C++, Go, Linux, Data Structures, Algorithms, REST/GraphQL APIs, Git (Version Control).
Full-Stack & Frameworks: React, Node.js, Angular, Microservices, Frontend and Backend development, Cloud APIs. Machine Learning & Core AI: Machine Lerning, Deep Learning (PyTorch, TensorFlow), Reinforcement Learning, Generative AI (LLMs, RAG), Prompt Engineering, Feature Engineering, Model Training & Deployment, A/B Testing, NLP, Computer Vision, Statistical Modeling.
Generative AI & LLM Engineering: LLM (e.g. GPT/BERT), Retrieval-Augmented Generation, Prompt Engineering, LangChain agents, Embeddings, Vector Databases (FAISS, Pinecone), Hallucination mitigation, LLM evaluation metrics. Natural Language Processing: Intent Classification, Named Entity Recognition, Conversational AI Systems, Semantic Search, Context Grounding.
Deep Learning & Inference Systems: PyTorch, TensorFlow, Hugging Face Transformers, Torch Serve, NVIDIA Triton Inference Server, ONNX.
Search & Vector Infrastructure: FAISS, Pinecone, Hybrid Search, BM25, Vector Indexing, Retrieval Optimization, Redis Caching. Data Engineering & Streaming: Spark, Kafka, Airflow, Databricks, ETL and batch/real-time pipelines, Data Lake (S3, Redshift, DynamoDB), Elasticsearch, Redis.
Cloud & Platform Engineering: AWS (EC2/S3/Lambda/SageMaker/Bedrock), Google Cloud (BigQuery, Vertex AI), Azure, Kubernetes/Docker (EKS, GKE), Terraform, CI/CD pipelines, MLOps (MLflow), Observability (Prometheus, Grafana). MLOps & Production AI Systems: Model Versioning, Experiment Tracking, ML Workflow Orchestration, Data Drift Detection, Automated Retraining, CI/CD Pipelines, Blue Green Deployment, Observability (Prometheus, Grafana). Distributed Systems & Scalability: Scalable system design, low-latency inference, autoscaling, fault-tolerant architecture, high- availability, microservices architecture.
Soft Skills: Cross-functional collaboration, team communication, problem-solving, technical documentation. Education
Master of Science in Information Technology
Webster University
CERTIFICATIONS
AWS CERTIFIED MACHINE LEARNING ENGINEER – ASSOCIATE Amazon Web Services (AWS) 2026