Post Job Free
Sign in

Production RAG and LLM Systems Engineer

Location:
Columbus, OH
Posted:
July 10, 2026

Contact this candidate

Resume:

Rajesh Racha

AI / ML Engineer

Kent, USA

+1-234-***-****

*************@*****.***

linkedin.com/in/rajesh-racha/

AI/ML engineer with 5+ years building applied ML, NLP, and generative AI systems across healthcare, hospitality, and financial services. Current focus is production RAG pipelines, LLM workflows, and the engineering around them — ingestion, evaluation, guardrails, and deployment. Comfortable owning a use case from ambiguous business problem to deployed FastAPI service running on AWS or Azure, and equally comfortable on the data side with PySpark, embeddings, and vector stores. Looking for AI engineering roles where the team is shipping real products, not just demos.

Employment history

FEB 2025 – PRESENT, GENERATIVE AI DEVELOPER

UnitedHealth Group, Cleveland, OH

JUN 2022 – AUG 2024, AI / ML ENGINEER

TCS (Client: Hilton), Hyderabad, India

Built and maintain a production retrieval-augmented generation pipeline over internal healthcare documents using LangChain, FAISS, ChromaDB, and GPT-based APIs, serving document Q&A and summarization workflows for internal teams across ~120K documents. Designed the ingestion side end-to-end: parsing, semantic chunking, metadata tagging, embedding generation, and vector indexing across structured and unstructured enterprise content, with nightly refresh. Developed multi-step LLM workflows with guardrails around context selection, prompt structure, and response consistency for document Q&A and conversational assistant use cases.

Built evaluation and observability for the RAG pipeline using LangFuse for production tracing and an offline eval harness covering faithfulness, relevance, latency, and failure-case review,so prompt and retrieval changes had to clear a quantitative gate before production rollout.

Built NLP pipelines with BERT, SBERT, spaCy, and NLTK for semantic search, classification, and entity recognition on domain-specific text

— reused across multiple downstream LLM workflows. Exposed retrieval, summarization, and inference behind FastAPI services with secure API access, deployed and monitored on AWS SageMaker and containerized runtimes alongside data scientists and senior engineers. Fine-tuned LLaMA-2 7B with QLoRA (4-bit quantized base, low-rank adapters) to evaluate an open-weight alternative to the GPT-4 generation layer — benchmarking answer quality and cost against the production Azure OpenAI path before retaining GPT-4 for the shipped pipeline.

Wrapped internal FastAPI retrieval endpoints as MCP tools to evaluate agent-based access patterns for the care-ops copilot, informing the team's roadmap for tool-calling interfaces.

Cut LLM token spend on the highest-traffic workflow ~40% through systematic prompt engineering — instruction restructuring, few-shot tuning, and grounding constraints — plus retrieval tuning and response caching, with no measurable quality drop on the eval set. Built a prompt engineering workflow for GPT-4 (Azure OpenAI): chain-of-thought decomposition for multi-policy queries, few-shot grounding examples, and strict citation-format constraints, validated against a golden eval set before every deploy. Built evaluation checks for RAG outputs covering relevance, faithfulness, latency, and failure-case review, so prompt and retrieval changes had to clear a quantitative gate before production rollout. Built applied AI and NLP solutions for hospitality use cases — guest-feedback analysis, document classification, semantic search, and a virtual assistant workflow over internal SOPs.

Delivered RAG and document search pipelines using Azure OpenAI, LangChain, and vector databases with hybrid sparse + dense retrieval to provide grounded answers from internal knowledge sources. Fine-tuned Hugging Face transformer models for domain-specific classification and entity extraction, lifting category-level F1 from ~0.71

(baseline keyword rules) to ~0.89 on guest-feedback categorization. Engineered data prep and feature engineering workflows with Python, SQL, PySpark, and Databricks supporting model training, evaluation, and batch inference at multi-TB scale.

Stood up MLOps with Azure ML, MLflow, Docker, Kubernetes, Jenkins, and Azure DevOps including infrastructure-as-code (Terraform) for repeatable environments,cutting model deployment turnaround from ~2 weeks to ~3 days. Wrapped models behind FastAPI services integrated with Java backends and worked with DevOps on reliable enterprise delivery. Built Power BI and Tableau dashboards to track model performance, adoption, and operational KPIs, used in monthly business reviews with client stakeholders.

Established MLflow experiment tracking and versioning to standardize model comparison and reproducible runs. NOV 2020 – MAY 2022, DATA SCIENTIST / ML ENGINEER

Goldman Sachs, Hyderabad, India

Education

M.S., COMPUTER SCIENCE

Kent State University, Kent, OH

B.E., ELECTRICAL ENGINEERING

Mahatma Gandhi Institute of Technology, Hyderabad, India Courses

GENERATIVE AI ENGINEERING

IBM

CLOUD PRACTITIONER ESSENTIALS

AWS

Skills

RAG, prompt engineering, LLM workflows, document Q&A, summarization, embeddings, semantic search, OpenAI API, Azure OpenAI, GPT-4, LLaMA 2, LangChain, BERT, SBERT, Hugging Face Transformers, spaCy, NLTK, XGBoost, scikit-learn, PyTorch, TensorFlow, fine-tuning, model evaluation, FAISS, Pinecone, ChromaDB, PostgreSQL, SQL, Pandas, NumPy, C, PySpark, Apache Spark, Databricks, Snowflake, AWS SageMaker, Azure ML, C++, GCP Vertex AI, MLflow, typescript, Next.js, FastAPI, Docker, Kubernetes, Jenkins, Azure DevOps, CI/CD, model monitoring, SHAP, LIME, A/B testing, regression analysis, Power BI, Tableau, rust, R, Go, react Links

Github: https://github.com/racharajeshAI

Portfolio: https://racharajeshai.github.io/portfolio-site/ Built supervised ML models (XGBoost, logistic regression) for financial-data use cases including classification, anomaly detection, and fraud- risk signals on transaction data.

Lifted precision at the top-risk band by ~22% over the previous heuristic baseline on the fraud-signal model, with full SHAP-based explainability artifacts produced for compliance review. Performed exploratory analysis, data cleaning, and feature engineering using Pandas, NumPy, SQL, scikit-learn, PCA, encoding, and scaling to build reliable training datasets.

Processed large datasets — 100M+ rows — with Apache Spark and PySpark, enabling repeatable training and scoring workflows beyond single-machine limits.

Applied SHAP and LIME to explain model decisions and support transparent review of high-impact predictions for business and compliance stakeholders.

Designed validation workflows including regression analysis, A/B testing, and threshold review to compare candidate model changes before release.

Deployed and monitored ML services on AWS and Azure environments, and partnered with engineers, analysts, and business stakeholders via Git, Jira, and Confluence.



Contact this candidate