AI/ML Engineer Data Scientist
Name: Chandan Aithagoni
Mobile: +1-512-***-****
E-mail: ****************@*****.*** LinkedIn: linkedin.com/in/aithagoni-chandan
Professional Summary:
AI/ML Engineer and Data Scientist with 5 years of experience designing, developing, and deploying production-grade AI/ML solutions across banking, healthcare, and retail domains.
•Strong hands-on expertise in Python, SQL, Scikit-learn, PyTorch, TensorFlow, Hugging Face Transformers, Pandas, PySpark, and Apache Spark for machine learning, data processing, and AI application development.
•Hands-on experience building Generative AI and LLM-powered applications using RAG, Agentic AI, prompt engineering, embeddings, semantic search, and vector databases.
•Designed and developed AI agents and multi-step agentic workflows using LangChain and LangGraph, integrating LLMs with enterprise APIs, databases, retrieval systems, and external tools.
•Built production-ready Retrieval-Augmented Generation (RAG) pipelines covering document ingestion, chunking, embedding generation, vector search, contextual retrieval, prompt construction, and response generation.
•Developed NLP and document-processing solutions using BERT and Transformer-based models for text classification, semantic search, information extraction, document analysis, and contextual retrieval.
•Strong experience developing end-to-end machine learning pipelines, including data ingestion, preprocessing, feature engineering, model training, hyperparameter tuning, evaluation, deployment, and production monitoring.
•Built real-time fraud detection and investigation solutions using Python, Scikit-learn, Spark, Kafka, and AI/ML techniques to identify suspicious transactions, reduce false positives, and support analyst decision-making.
•Designed and deployed scalable REST APIs and AI/ML services using FastAPI and Flask, integrating machine learning models and AI workflows with enterprise applications.
•Built batch and real-time data pipelines using Kafka, Spark/PySpark, Databricks, Airflow, SQL, and AWS to process large-scale structured and unstructured datasets.
•Implemented MLOps and LLMOps practices including MLflow experiment tracking, model versioning, evaluation, monitoring, prompt optimization, CI/CD, logging, and production troubleshooting.
•Containerized AI/ML applications using Docker and deployed services on Kubernetes to support scalable and reliable production inference.
•Strong experience with SQL and NoSQL databases including PostgreSQL, MySQL, MongoDB, Snowflake, and vector-search solutions for analytical, ML, and GenAI applications.
•Experienced in translating business requirements into scalable, production-ready AI solutions while collaborating with Data Engineers, Software Engineers, DevOps teams, analysts, and business stakeholders.
•Strong problem-solving experience across model performance optimization, data quality, root-cause analysis, production support, scalability, and continuous improvement of AI/ML systems.
EDUCATION:
Master’s in Management Information Systems, Auburn University at Montgomery, USA
Bachelor’s in Electronics and Communication Engineering, JNTUH, India
TECHNICAL SKILLS:
Category
Skills
Programming
Python, SQL
Machine Learning
Scikit-learn, XGBoost, Classification, Regression, Clustering, Feature Engineering, Model Evaluation
Deep Learning & NLP
PyTorch, TensorFlow, BERT, Transformers, NLP
Generative AI
LLMs, RAG, LangChain, LangGraph, LlamaIndex, Prompt Engineering, Agentic AI
Vector Databases
Pinecone, ChromaDB, MongoDB Atlas Vector Search
Data Processing
Pandas, NumPy, PySpark, Spark SQL
Big Data
Apache Spark, Kafka, Hadoop, Hive
MLOps
MLflow, Model Deployment, Model Monitoring, CI/CD
Cloud
AWS (S3, Lambda, EMR)
Data Engineering
ETL, Airflow, Data Modeling, Data Quality
Databases
Snowflake, PostgreSQL, MySQL, MongoDB
APIs & Deployment
FastAPI, Flask, Docker, Kubernetes
Visualization
Power BI, Tableau
Tools
Databricks, Git, Jenkins, JIRA
PROFESSIONAL EXPERIENCE
Client: Citi Bank, Charlotte, NC. Feb’25 - Present
Role: AI/ML Engineer/ Data Scientist
Project: Real-time Fraud Detection
Worked on building a real-time fraud detection system combined with an NLP-based investigation tool to help fraud analysts review suspicious transactions faster.
Responsibilities:
Built a production-grade fraud intelligence platform combining real-time transaction scoring with retrieval-based case investigation for analyst workflows.
Developed fraud detection models using scikit-learn to assign risk scores for each transaction.
Fine-tuned BERT models using PyTorch to analyze transaction descriptions and identify suspicious patterns.
Built and deployed FastAPI inference services for fraud scoring and retrieval workflows, enabling integration with internal banking applications.
Designed and implemented vector database solutions using MongoDB Atlas Vector Search and Pinecone to support semantic retrieval of fraud cases and investigation documents.
Developed semantic search using BERT embeddings to improve document retrieval relevance across fraud records.
Integrated OpenAI API with retrieval pipelines and semantic search workflows to generate context-aware fraud investigation summaries and analyst recommendations.
Designed agentic AI workflows using LangGraph to automate fraud investigation steps and reduce analyst review time by 40%.
Developed Retrieval-Augmented Generation (RAG) pipelines using embeddings and vector search to improve contextual accuracy and reduce irrelevant responses.
Developed LLM evaluation workflows using relevance, groundedness, and response quality metrics to validate retrieval performance and improve answer accuracy across fraud investigation use cases.
Built retrieval-based investigation workflows using embeddings and semantic search over fraud case data.
Integrated semantic search and context-aware response generation for analyst decision support.
Reduced false positives and improved fraud detection precision by 20 - 30% through feature engineering enhancements and detailed error analysis.
Tracked model experiments, hyperparameter tuning results, and evaluation metrics using MLflow to improve fraud detection accuracy and model reproducibility.
Designed both batch and streaming ETL pipelines using Kafka, Spark, and AWS S3 to ensure consistent data availability.
Used Databricks notebooks with PySpark for large-scale fraud data processing, feature engineering, and exploratory analysis.
Built and optimized Spark SQL transformations in Databricks to prepare transaction data for real-time fraud detection models.
Containerized machine learning services using Docker and deployed them on Kubernetes for scalable inference.
Applied controlled retrieval and filtering strategies to improve contextual relevance and enforce response quality standards in AI-generated outputs.
Developed and deployed end-to-end NLP pipelines using LangChain and LangGraph for extracting insights from unstructured data sources.
Used AWS services like S3, Lambda, and EMR for data storage, processing, and model execution.
Utilized MLflow for experiment tracking, model versioning, and performance comparison across fraud detection model training cycles.
Improved fraud detection capability significantly by identifying suspicious transactions earlier and reducing manual review effort.
Implemented Elasticsearch/OpenSearch pipelines for semantic search and relevance ranking, integrating BM25 scoring with embedding-based retrieval to improve precision in fraud case investigations.
Designed and optimized BM25 ranking workflows alongside transformer-based embeddings to balance keyword relevance with semantic context in large-scale document search systems.
Conducted search relevance evaluation using precision/recall, NDCG, and custom ranking metrics to validate retrieval quality and improve analyst decision support across fraud detection and customer intelligence platforms.
Environments: Python, LangChain, LangGraph, RAG, Vector Search, PySpark, Databricks, Scikit-learn, PyTorch, FastAPI, Kafka, Apache Spark, MongoDB, SQL, AWS (S3, Lambda, EMR), Docker, Kubernetes, Airflow, Git, Jenkins, Hugging Face Transformers.
Client: Best Buy, Alpharetta, GA Sep’23 – Jan’25
Role: Machine Learning Engineer / Python Developer
Project: Retail Personalization & Customer Intelligence Platform
Built a data-driven personalization system for an e-commerce platform to improve customer retention and targeted marketing. Developed machine learning models for churn prediction and purchase propensity using customer behavior and transaction data. Enabled personalized recommendations and optimized business decisions through predictive analytics and forecasting.
Responsibilities:
Built end-to-end customer analytics pipelines using PySpark and SQL to process large-scale data from web activity, transactions, and app usage.
Designed and developed churn prediction and purchase propensity models using XGBoost and classification algorithms to identify high-risk and high-value customers.
Improved campaign targeting efficiency by 18% through customer segmentation and predictive analytics models.
Performed feature engineering on customer behavior data including engagement patterns, visit frequency, and purchase trends to improve model performance.
Implemented scalable batch inference workflows to generate customer propensity scores for downstream reporting and campaign systems.
Conducted cohort analysis and customer segmentation using Spark SQL in Databricks to identify retention opportunities and targeted marketing strategies.
Developed time-series forecasting models using Prophet to predict customer activity and product demand for campaign planning and inventory optimization.
Built and optimized data pipelines to ingest, transform, and prepare structured customer data for analytics and machine learning workflows.
Integrated model outputs into dashboards and reporting systems to enable business teams to drive data-driven marketing decisions.
Collaborated with cross-functional teams including product, marketing, and analytics to translate business requirements into scalable ML solutions.
Improved performance of data pipelines by optimizing Spark jobs and resolving bottlenecks, enabling faster processing of large datasets.
Delivered measurable impact by reducing customer churn by 17% and improving the effectiveness of targeted marketing campaigns.
Environment: Python, PySpark, Spark SQL, Databricks, SQL, Snowflake, XGBoost, Prophet, Tableau, AWS, Kafka, Airflow, Git, FastAPI.
Client: Common Spirit Health, India May’21– Apr’22
Role: Data Scientist/ Machine Learning Engineer
Project: Patient Recommendation
Worked on building a recommendation system and NLP-driven chatbot to improve patient engagement, personalize service recommendations, and automate support interactions.
Responsibilities:
Built a recommendation system using Python and scikit-learn to suggest healthcare services based on patient history and interaction data.
Implemented collaborative filtering and content-based approaches to personalize recommendations for patients.
Improved patient query response accuracy by 22% through NLP preprocessing and intent refinement.
Processed patient and service data using Pandas and SQL to prepare datasets for model training.
Developed supervised learning models to predict patient preferences and improve recommendation relevance.
Designed and implemented an NLP-based chatbot to handle patient queries related to services, appointments, and support requests.
Used Dialogflow to define intents, entities, and conversation flows for patient interaction scenarios.
Integrated the chatbot with backend systems such as patient records and service databases for real-time responses.
Built Python-based backend services to connect chatbot inputs with machine learning models and data sources.
Developed document processing workflows to extract key information from healthcare-related documents and records.
Conducted model evaluation and error analysis to improve recommendation accuracy and chatbot response quality.
Automated retraining workflows to keep models updated with new patient interaction data.
Built pipelines to process and structure healthcare text data for downstream ML and chatbot applications.
Worked with SQL Server and ETL processes to manage and transform structured healthcare data.
Built dashboards using Tableau to visualize patient engagement, recommendation trends, and system performance.
Collaborated with business and healthcare teams to align recommendation logic with real-world use cases.
Improved chatbot response handling by refining intent mapping and managing edge-case scenarios.
Supported deployment of models and chatbot services using CI/CD workflows.
Improved patient engagement by delivering more relevant recommendations and faster support responses.
Environment: Python, scikit-learn, TensorFlow, Pandas, SQL Server, Apache Spark, Dialogflow, Tableau, AWS, ETL, R, DB2, Teradata, Git, Agile.
Client: INFOTEX IT SOLUTIONS, India Nov’20 - April’21
Role: Data Scientist
Project: Machine Learning for Business Reporting
Worked on building data pipelines, machine learning models, and reporting systems to support business analytics and improve decision-making using structured data.
Responsibilities:
Built machine learning models using Python and scikit-learn for classification tasks on business datasets.
Worked on data preprocessing using Pandas and NumPy to clean, transform, and prepare data for modeling.
Implemented supervised learning algorithms like logistic regression, decision trees, KNN, and Naive Bayes for prediction tasks.
Used SQL and PL/SQL to create tables, write queries, and manage structured data in relational databases.
Performed exploratory data analysis using Matplotlib and Seaborn to identify trends and patterns.
Worked with Spark Data Frames and Spark SQL to process large datasets and improve data handling efficiency.
Built basic machine learning workflows using Spark MLlib for scalable model training.
Developed data models using Erwin Data Modeler to support data warehouse design.
Created QlikView dashboards to visualize business metrics and reporting insights.
Integrated data from multiple systems into a unified reporting layer for business users.
Used AWS services to set up storage and support data processing workflows.
Developed scripts in Python to automate data extraction, transformation, and loading processes.
Collaborated with business analysts to understand reporting requirements and translate them into technical solutions.
Supported migration of data from OLTP systems to data warehouse environments.
Environment: Python, scikit-learn, Pandas, NumPy, SQL, PL/SQL, Apache Spark, Hadoop, Hive, AWS, MongoDB, DB2, Informatica, QlikView, R, Excel, Erwin Data Modeler