Post Job Free
Sign in

Senior AI/ML Engineer: GenAI & RAG MLOps

Location:
Chicago, IL
Posted:
October 06, 2026

Contact this candidate

Resume:

Uday L

Senior AI/ML Engineer Generative AI & RAG Agentic AI MLOps Azure & AWS

***********@*****.*** +1-916-***-****

PROFESSIONAL SUMMARY

Senior AI/ML Engineer with 9+ years of experience delivering machine learning, data engineering, and AI applications across healthcare, banking, insurance, and telecom.

Delivers ML solutions from data preparation and feature engineering through evaluation, API integration, cloud deployment, and production monitoring.

Builds enterprise retrieval-augmented generation (RAG) applications with Azure OpenAI Service, Azure AI Search, and Amazon Bedrock to ground answers in approved policy and operational content.

Develops LangGraph workflows that coordinate retrieval and tool execution with state management, response validation, and human-review controls.

Applies predictive modeling and clinical NLP to healthcare risk, insurance model validation, operational analytics, and customer behavior analysis.

Engineers distributed data and feature pipelines with PySpark, Databricks, and Delta Lake for reproducible training and reliable batch scoring.

Integrates containerized models and AI APIs with enterprise applications using FastAPI, Azure ML, AKS, and AWS services.

Uses MLflow and evaluation workflows to track experiments, version model and prompt artifacts, monitor drift, and support controlled releases.

Works with clinical, business, risk, and engineering teams to define use cases, review model limitations, and resolve production data and AI issues.

TECHNICAL SKILLS

Programming & APIs: Python, SQL, Bash, FastAPI, Flask, Pydantic, REST APIs

ML & Deep Learning: PyTorch, TensorFlow, Scikit-learn, XGBoost, LightGBM; classification, regression, forecasting, anomaly detection, recommendation systems

GenAI & NLP: Azure OpenAI Service, Azure AI Foundry, Amazon Bedrock, LangChain, LangGraph, Model Context Protocol (MCP), Hugging Face, BERT, spaCy, NLTK, retrieval-augmented generation (RAG), large language models (LLMs), embeddings, prompt engineering

Retrieval & Data: Azure AI Search, FAISS, semantic search, PySpark, Spark, Databricks, Delta Lake, Airflow, Pandas, NumPy, Hadoop, Hive, Kafka, Flink, Parquet

Cloud & Model Serving: Azure ML, AKS, Azure Container Registry, AWS, SageMaker, EC2, S3, EMR, Glue, Glue Data Catalog, Athena, Redshift, Lambda, Step Functions, KServe, Triton Inference Server, TensorFlow Serving, TorchServe

MLOps & Delivery: MLflow, DVC, Docker, Kubernetes, Helm, Terraform, GitHub Actions, Jenkins, Azure DevOps, Git, GitHub; CI/CD, evaluation, agent tracing

Monitoring & Governance: Prometheus, Grafana, Azure Monitor, CloudWatch, Evidently AI, SHAP, LIME, Tableau; Azure Key Vault

Databases & Storage: PostgreSQL, MySQL, MongoDB, Hive, Delta Lake

WORK EXPERIENCE

Truist Financial, Charlotte, NC Senior GenAI Engineer Dec 2024 - Present

Engineered Amazon Bedrock knowledge assistants for lending-policy and deposit-servicing questions, combining answers grounded in approved content with controlled tool execution.

Assembled banking documents and evaluation datasets with Databricks and PySpark, masking sensitive data and removing duplicates before assistant testing.

Configured metadata filters and prompt context for RAG retrieval, keeping banking responses grounded in relevant policy content and traceable to their sources.

Developed LangGraph workflows that coordinated retrieval and tool execution, using workflow state and conditional routing to control multi-step banking-support tasks.

Implemented response validation and human-review boundaries within agent workflows to control how banking-support outputs were returned.

Standardized FastAPI and Pydantic interfaces for banking assistants, giving enterprise applications consistent request validation, inference settings, and structured responses.

Added retries, timeouts, and error handling to model requests so failed integrations could be handled without disrupting assistant service behavior.

Maintained evaluation questions and reference answers covering lending policies and deposit servicing, providing consistent test cases for model and prompt comparisons.

Evaluated answer relevance, completeness, grounding, and citations against reference answers, combining automated checks with banking SME review before model or prompt changes.

Tested agent tool selection, argument quality, and stopping behavior, tracing execution to diagnose routing and tool-use failures before workflow changes were released.

Authored prompt-injection and sensitive-data disclosure tests to identify unsafe assistant behavior before release.

Tracked prompt versions, model identifiers, and evaluation artifacts in MLflow to reproduce experiments and compare banking assistant configurations.

Versioned evaluation data with DVC and retained reference snapshots in Amazon S3 to keep assistant comparisons tied to the same source data.

Benchmarked candidate ML components in SageMaker against reference datasets to compare performance during model review.

Packaged banking assistant services with Docker and supported Kubernetes deployment checks for configuration and service health.

Checked dependencies, configuration changes, and assistant behavior through GitHub Actions and Jenkins CI/CD workflows before deployment.

Monitored latency, token use, retrieval failures, and tool errors with Prometheus and Grafana, using traces and feedback to prioritize fixes.

Documented intended use, evaluation coverage, limitations, and release controls with banking SMEs and model-risk and security teams.

Environment: Python, Amazon Bedrock, LangGraph, Databricks, PySpark, Hugging Face, FastAPI, Pydantic, MLflow, DVC, Docker, Kubernetes, SageMaker, EC2, S3, SQL, MongoDB, Prometheus, Grafana, GitHub Actions, Jenkins, Terraform, Agile

Sentara Health, Norfolk, VA ML Engineer Jul 2023 - Nov 2024

Delivered healthcare risk-prioritization models and clinical-policy assistants, connecting data preparation, evaluation, APIs, deployment, and monitoring.

Curated claims and member features in Databricks and Delta Lake, checking schemas and data quality before training and scoring.

Extracted diagnoses, treatments, and provider details from clinical notes with Hugging Face BERT, turning unstructured text into reusable inputs for healthcare analytics.

Built PyTorch and gradient-boosting models for healthcare risk and utilization to support case prioritization and operational review.

Validated model performance and decision thresholds with clinical review before releasing healthcare risk and case-prioritization outputs.

Orchestrated training and retraining with Airflow, retaining model artifacts and evaluation records in MLflow for repeatable releases.

Linked DVC-versioned training datasets to model artifacts and evaluation records so release reviews could trace results to the underlying data.

Built clinical-policy RAG with Azure OpenAI Service and Azure AI Search, grounding healthcare answers in retrieved policy and operational documents.

Developed LangChain ingestion and chunking pipelines that preserved clinical-policy metadata, keeping retrieved passages traceable to their source documents.

Indexed care guidelines and operational documents with embeddings and Azure AI Search to enable semantic retrieval.

Implemented LangGraph workflows for retrieval, response validation, and tool execution across healthcare knowledge sources.

Evaluated RAG answers and traced LangGraph execution to separate retrieval, tool-use, and generation failures and guide fixes.

Integrated assistants and predictive models with care-management applications through FastAPI REST services that returned structured outputs.

Defined Pydantic request validation and API contracts to keep healthcare AI integrations consistent.

Deployed Docker-packaged healthcare models on AKS with versioned inference endpoints for scalable care-management integration.

Supported Helm configuration of Kubernetes model services, checking deployment settings and documenting rollback steps with platform engineers.

Reviewed feature drift with Evidently AI and investigated inference health and latency through Azure Monitor.

Collaborated with clinical and security teams on explainability, subgroup performance, access controls, auditability, and human review.

Environment: Python, PyTorch, Hugging Face, BERT, Azure OpenAI Service, Azure AI Search, LangChain, LangGraph, FastAPI, Pydantic, Docker, Kubernetes (AKS), Helm, KServe, Triton Inference Server, MLflow, DVC, Apache Airflow, Databricks, PySpark, Delta Lake, Evidently AI, Azure Monitor, Prometheus, Terraform, SQL, GitHub Actions

GEICO, Chevy Chase, MD Associate Data Scientist Jan 2022 - Jun 2023

Developed insurance prediction models with Scikit-learn and XGBoost, carrying policy and claims features through training and validation.

Applied point-in-time feature controls to exclude information unavailable at prediction time, reducing leakage risk in insurance training and validation datasets.

Transformed policy and claims data into model features with PySpark and Delta Lake, checking schemas and removing duplicates before training.

Established reusable Python and NumPy preprocessing pipelines so insurance models could be compared using the same data transformations.

Logged parameters, metrics, and dependencies in MLflow to compare candidate models and reproduce insurance experiments.

Compiled Azure ML registration records and validation evidence to preserve model lineage and review history before engineering handoff.

Assessed precision-recall, calibration, and segment performance, using SHAP and LIME to explain influential features and individual insurance predictions.

Supported Docker-packaged inference on AKS, checking health probes and autoscaling with engineers to maintain insurance prediction endpoints.

Validated rollback behavior before inference endpoint deployment and assisted with container image releases through Azure Container Registry.

Investigated inference errors, latency, and feature changes with Azure Monitor and Application Insights, and contributed to Azure DevOps CI/CD releases.

Developed RAG ingestion and evaluation components for quote-support content, checking retrieval, grounding, and content refresh.

Assisted with secure inference deployment using Azure Key Vault, managed identities, and data filtering, and documented recovery steps with engineers.

Environment: Python, Azure Machine Learning, Azure Databricks, PySpark, Delta Lake, Scikit-learn, XGBoost, MLflow, Apache Airflow, Docker, Azure Container Registry, AKS, Azure DevOps, Azure Monitor, Application Insights, Azure Key Vault, SHAP, LIME, SQL

Comcast, Philadelphia, PA Data Engineer Oct 2018 - Dec 2021

•Built AWS telemetry pipelines for broadband reporting and downstream ML, agreeing schemas and data-quality rules with network and analytics teams.

•Organized raw and curated telemetry in Amazon S3 with Parquet and date-based partitions, aligning storage with historical retention and time-based reporting queries.

•Implemented Kafka and Flink ingestion workflows to validate and process device telemetry before downstream storage and analysis.

•Maintained schemas and table metadata in Glue Data Catalog so telemetry stayed discoverable and queryable as source structures changed.

•Automated AWS Glue ETL with PySpark for analytical and model-input datasets, using incremental loads, reconciliation, and controlled reruns to validate outputs and recover processing.

•Prepared Amazon S3 training and scoring datasets for SageMaker workflows, keeping data splits and schema checks consistent.

•Designed Amazon Redshift tables and SQL queries for device-level analysis, tuning distribution and sort strategies around recent-data reporting patterns.

•Queried historical telemetry with Amazon Athena to validate pipeline outputs and help analysts investigate network-quality trends.

•Wrote Python-based AWS Lambda functions for incoming file-metadata validation and lightweight ingestion checks.

•Configured CloudWatch dashboards and alerts for job failures, latency, and data freshness to help investigate telemetry issues and recover workflows.

•Managed Git-versioned code and configuration, checking releases and documenting dependencies, rerun procedures, and recovery steps.

Environment: Python, SQL, PySpark, Apache Kafka, Apache Flink, AWS Glue, Glue Data Catalog, Amazon S3, Amazon Athena, Amazon Redshift, AWS Lambda, Amazon SageMaker, Amazon CloudWatch, Parquet, Git, Linux

Impetus Technologies, HYD, India Jr Data Engineer Apr 2016 - Aug 2018

Supported telecom data engineering and retention analytics, preparing subscriber datasets and validating distributed workflows with senior engineers.

Optimized SQL and Hive queries to produce consistent subscriber-engagement datasets for churn analysis and reporting.

Aggregated Hadoop event data with PySpark into subscriber-level reports and reusable churn features for the retention analytics team.

Cleaned telecom records with Python and Pandas, examining missing values and subscriber patterns before analysis.

Compared Scikit-learn and XGBoost churn models using subscriber behavior, reviewing elevated cancellation risk with retention analysts.

Created Tableau dashboards for acquisition, engagement, and churn KPIs, tracing discrepancies through EMR and Hive transformations to investigate upstream data issues.

Contributed to Git-based change reviews, workflow validation, and production troubleshooting with senior engineers on distributed data and ML processes.

Environment: Python, PySpark, SQL, Hive, Pandas, Scikit-learn, XGBoost, Tableau, AWS EMR, Hadoop, Git

Education

Amity University, Bachelor of Computer Science



Contact this candidate