Name: Praneeth Babu Thota
Email: *****************@*****.*** Phone: +1-814-***-****
Gen AI /ML/ Data Scientist
PROFESSIONAL SUMMARY:
●Experienced AI/MLops Data Scientist with overall 5+ years of experience, specializing in transforming business needs into analytical models and scalable AI/ML solutions.
●Expert in Generative AI applications and prompt engineering, leveraging LLMs like GPT, BERT, and T5 for sentiment analysis, fraud detection, document summarization, and conversational AI.
●Proficient in building advanced Generative AI models, multimodal Vision-Language workflows, and scalable ML systems.
●Proficient in GPU efficiency and optimization, leveraging TensorFlow, PyTorch, and cloud GPU clusters (Azure ML, AWS SageMaker, GCP Vertex AI) to accelerate model training and reduce compute costs
●Proficient in supervised, unsupervised, and reinforcement learning algorithms, including regression, classification (SVM, Random Forest, XGBoost), clustering, PCA, and GANs.
●Proficient in developing PL/SQL procedures, triggers, packages, and API logic to support model inference, compliance monitoring, and real-time transaction analysis.
●Skilled in integrating PL/SQL with cloud ecosystems (Azure, AWS, Snowflake, Redshift) to enhance the performance and scalability of AI-driven applications.
●Skilled in leveraging GitHub Copilot for accelerated API development, code reviews, and faster delivery of AI/ML pipelines.
●Hands-on experience with Agentic AI workflows, orchestrating LLM reasoning with automated actions for compliance, fraud detection, and customer support.
●Hands-on experience in AI-powered chatbots and automation, using Kore.ai for virtual assistants with sentiment detection, contextual awareness, and enterprise integration.
●Strong expertise in MLOps automation, leveraging Kubernetes, Docker, CI/CD pipelines (GitLab, Jenkins), Kubeflow, MLflow, and AutoML for scalable deployments.
●Big Data & Cloud Computing specialist, skilled in Azure ML, Databricks, AWS SageMaker, GCP Vertex AI, Snowflake, Hadoop, Spark (PySpark), and SQL databases.
●Extensive experience in NLP & Computer Vision, including OCR-based document processing, image classification, object detection, and synthetic data generation.
●Advanced programming skills, proficient in Python (TensorFlow, PyTorch, Pandas, NumPy, Flask), R, JavaScript, and SQL for AI model development and integration.
●Skilled in data engineering & visualization, working with Tableau, Power BI, Matplotlib, Seaborn, and Apache Airflow for data pipelines and real-time dashboards.
●Agile & DevOps methodologies, applying Scrum, version control (Git/GitHub), Terraform for cloud automation, and end-to-end AI solution deployment.
TECHNICAL SKILLS
●Programming & Libraries: Python (Pandas, NumPy, TensorFlow, PyTorch, OpenCV, Scikit-learn, Flask), R, JavaScript, SQL
●Data Science & AI/ML: Predictive Analytics, NLP, Computer Vision, Supervised & Unsupervised Learning, Reinforcement Learning, GANs, VAEs, Large Language Models (LLMs)
●Big Data & Cloud Computing: Azure (Databricks, Data Factory), AWS (SageMaker, Glue, Redshift), GCP (Big Query, Vertex AI), Snowflake
●DevOps & Automation: Kubernetes, Docker, CI/CD Pipelines (GitLab, Jenkins), Terraform
●Visualization & Reporting: Power BI, Tableau, Looker, Matplotlib, Seaborn, Excel
●Database & Storage: SQL (PostgreSQL, MySQL, MongoDB, CosmosDB), Hadoop, Spark (PySpark)
WORK EXPERIENCE:
Client: Blue Cross Blue Shield, United States Oct 2024 – Till Date
Role: Gen AI/Data Scientist
PROJECT DESCRIPTION: Built and deployed AI-driven Healthcare Insurance models using Azure ML, Databricks, and Snowflake.
Responsibilities:
●Architected enterprise-grade Retrieval-Augmented Generation (RAG) pipelines using Azure OpenAI and Azure Cognitive Search to enable contextual reasoning over large-scale healthcare datasets.
●Built production-grade healthcare AI APIs using FastAPI and Flask, supporting secure streaming responses, OAuth2/JWT authentication, and role-based access controls.
●Implemented Agentic AI workflows using retrieval-augmented generation (RAG) and orchestration tools to autonomously process regulatory data and trigger real-time compliance alerts.
●Architected HIPAA-compliant Retrieval-Augmented Generation (RAG) systems using LangChain and LangGraph to power clinical knowledge assistants for providers and care coordinators.
●Integrated Genesys data (InfoMart, UCS, GAAP, Speech Miner) into Azure ML pipelines for call center analytics, real-time insights, and sentiment modeling.
●Optimized Genesys WFM routing and operations data, improving contact center performance and KPIs using Azure Databricks and ML models.
●Built and deployed conversational AI bots using Kore.ai, integrating LLMs (GPT, T5) for customer service automation with sentiment-driven, context-aware responses.
●Integrated Azure OpenAI Service (GPT-4 class models) within secure VNet environments for PHI-safe inference and contextual reasoning.
●Developed FastAPI-based AI microservices exposing LLM-powered summarization, semantic search, and decision-support APIs.
●Optimized GPU utilization for large-scale TensorFlow and PyTorch deep learning models, reducing training time by fine-tuning batch sizes, model architecture, and resource allocation on Azure ML GPU clusters.
●Automated MLOps workflows using AWS Step Functions, Azure DevOps, and Prefect, handling training, validation, deployment, and retraining pipelines.
●Developed CI/CD pipelines for ML models using Jenkins, GitHub Actions, integrating Python, Bash, and PowerShell to manage deployments across dev/stage/prod.
Environment: Azure Machine Learning, Azure Databricks, Snowflake, Python (TensorFlow, PyTorch, Scikit-learn), SQL, PL/SQL, Spark (PySpark), MLflow, AutoML, XGBoost, LightGBM, Azure DevOps, Airflow, Power BI, Tableau, Pandas, Matplotlib, spaCy, Hugging Face Transformers, BERT, GPT Models.
Client: Genpact, India. Sep 2020 – Nov 2023
Role: Data Scientist
Project Description: Designed predictive analytics models for, implemented NLP-based sentiment analysis, and data processing using Databricks and Hadoop.
Responsibilities:
●Delivered end-to-end GenAI solutions from scoping to production deployment, building a unified AI platform that orchestrated diverse enterprise AI/ML use cases using CrewAI, LangGraph, AutoGen, and other agentic frameworks—aligned with Databricks-native agent orchestration patterns.
●Built NLP models for sentiment analysis on patient reviews using AI-driven techniques and synthetic data generation.
●Built PL/SQL procedures for disease prediction models using structured medical data from SQL and NoSQL sources.
●Containerized AI microservices using Docker and deployed on Kubernetes (EKS/AKS).
●Implemented PL/SQL scripts for data extraction and transformation in AWS Glue and Redshift environments.
●Implemented cross-cloud integration between Azure and AWS search services for distributed AI workloads.
●Designed scalable data workflows, utilizing AWS Glue and Python (Pandas, NumPy) for structured and unstructured medical data processing.
●Extracted and processed structured and unstructured medical data from SQL and NoSQL databases (MongoDB).
●Implemented AI-powered automation for medical diagnostics using SageMaker AutoML and distributed computing techniques.
●Collaborated with stakeholders to productionize AI-enabled automation use cases across enterprise operational workflows.
●Developed CI/CD pipelines for AI model deployments using GitLab CI/CD, Kubernetes, and Docker.
Environment & Tools: Azure, GCP, AWS, Spark, Databricks, Hadoop, SQL, PL/SQL, MongoDB, TensorFlow, Keras, Data Robot, Power BI, Tableau, Kubernetes, Docker, GitLab CI/CD, Python, Scala, Apache Airflow, Snowflake, ADF.
Education Details:
Bachelor of Technology – Jawaharlal Nehru Technological University Kakinada
Masters - University of Central Missouri