KRISHNA CHARAN KANDLAKUNTA
United States ****************@*****.*** 972-***-**** LinkedIn Portfolio SUMMARY
Data Scientist with around 5 years of experience specializing in architecting end-to-end Machine Learning lifecycles and Generative AI solutions. Expert in developing scalable LLM-based applications (RAG architecture), Deep Learning models (CNN, RNN), and predictive analytics using Python, PyTorch, and AWS SageMaker. Proven track record of transforming complex data into strategic insights, driving measurable improvements in operational efficiency and $1.5M+ in annual cost savings. Adept at deploying production-grade MLOps pipelines while ensuring rigorous data governance. SKILLS
Languages: Python, R, SQL, SAS
Libraries: Pandas, NumPy, PyTorch, PySpark, SciPy, spaCy, Statsmodels Machine Learning: Scikit-learn, TensorFlow, Statistical Modeling, Regression, Time Series Forecasting, Classification, Recommendation Systems, Clustering, MLflow, FastAPI, Flask, Streamlit, NLP, Deep Learning, LLMs, RAG, LangChain, Vector Databases, NLTK, SHAP, LIME, GenAI, GPT-3.5, BERT
Big Data and Databases: ETL/ELT, Airflow, Hadoop, Spark, Hive, Kafka, Snowflake, PostgreSQL, MySQL, SQL Server, MongoDB Data Visualization: Tableau, Power BI, Looker, Matplotlib, Seaborn, Plotly Cloud & DevOps: AWS (S3, EC2, Lambda, Glue, SageMaker, CloudWatch), Azure (ADF, ADLS, Databricks, Synapse), GCP, CI/CD, Docker, Kubernetes, GitHub Actions, Jenkins, Grafana, Terraform Other Skills: A/B Testing, Hypothesis Testing, Agile/Scrum, Jira, Confluence, Role Based Access Control, REST APIs EXPERIENCE
CitiGroup, United States April 2025 – Present
Data Scientist
• Architected a secure, enterprise-grade RAG system using Llama 3 and LangChain to automate internal compliance auditing; reducing manual document review cycles by 65% and achieving a 92% precision in regulatory risk identification.
• Developed a real-time transaction monitoring framework using PyTorch and RNN (LSTMs) to detect anomalous patterns; processed 500K+ daily transactions, reducing false positives by 22% and driving $1.8M in projected annual savings.
• Engineered a full-cycle MLOps environment on AWS SageMaker, integrating MLflow for experiment tracking and GitHub Actions for CI/CD; slashed the research-to-production deployment timeline from 8 weeks to 5 days through automated canary testing.
• Designed distributed ETL pipelines using PySpark and Apache Airflow to aggregate high-frequency trading data in Snowflake; implemented delta-loading strategies that reduced data latency by 40% and improved downstream model training efficiency.
• Integrated SHAP and LIME frameworks into credit-scoring models to provide transparent feature-importance justifications, ensuring compliance with Fair Lending regulations while increasing model approval rates by 15% for under-served segments.
• Deployed high-concurrency model APIs using FastAPI and Docker (EKS), utilizing Kong API Gateway for load balancing; achieved sub- 100ms latency for real-time recommendation engines serving 1M+ active banking users. Wipro Technologies, India August 2021 – February 2023 Data Scientist II
• Built a distributed NLP pipeline using spaCy and Transformers (BERT) to analyze multi-lingual customer feedback across 10+ global markets; achieved a 91% F1-score in sentiment classification, driving a 15% increase in CSAT scores.
• Designed and deployed a hybrid recommendation system (Collaborative + Content-based) using TensorFlow Recommenders for an e-commerce client; boosted cross-sell conversion rates by 22% and increased average order value (AOV) by $45.
• Engineered a real-time fraud detection system using Isolation Forests and Autoencoders (Deep Learning) on streaming Kafka data; successfully flagged fraudulent transactions with 96% precision, preventing an estimated $800K in annual losses.
• Developed a global supply chain demand forecasting model using LSTMs and Prophet; improved inventory turnover by 19% and reduced stock-out incidents by 30% during peak seasonal periods.
• Containerized and orchestrated 12+ microservices-based ML models using Docker and Kubernetes (EKS); implemented Prometheus and Grafana for real-time drift monitoring, reducing model degradation issues by 40%.
• Mentored a team of 4 junior data scientists and translated complex technical roadmaps for C-suite stakeholders; led the delivery of 3 major AI products that generated an aggregate $2.5M in operational ROI. Wipro Technologies, India June 2019 – August 2021
Data Scientist
• Conducted in-depth exploratory data analysis (EDA) and built multivariate regression models to identify key drivers of customer churn; reduced churn rate by 12% through targeted retention campaigns.
• Developed a custom Python library for automated feature selection and engineering, reducing the model development lifecycle by 25% for the data science team.
• Designed executive-level Tableau dashboards that consolidated KPIs from MySQL and PostgreSQL, providing real-time visibility that accelerated strategic decision-making cycles by 20%.
• Implemented K-Means and Hierarchical Clustering to segment a database of 5M+ users, enabling personalized marketing that improved click-through rates (CTR) by 14%.
• Refactored legacy SQL scripts and implemented window functions/CTEs, resulting in a 60% reduction in report generation time for monthly business reviews.
• Actively participated in Scrum ceremonies and utilized Jira for sprint tracking, consistently delivering 100% of committed story points across 24 consecutive sprints.
EDUCATION
Master of Science in Data Science
University Of North Texas Denton, Texas
• Coursework: Machine Learning, Statistical Modeling, NLP, Big Data (Hadoop, Spark), Cloud Data Engineering, Data Visualization, Advanced Database Systems.
PROJECTS
SAMVAADH: AI Research Podcast Generator
• Built an AI-driven platform that transforms research papers into natural, conversational podcast episodes using multi-agent LLM pipelines and controlled dialogue generation.
• Designed a retrieval-augmented generation (RAG) workflow integrating OpenAI models, fine-tuned GPT4-mini variants, and semantic search to extract and summarize content from 1,000+ academic transcripts and arXiv papers.
• Implemented customizable episode generation with user-selectable length, topic depth, and narration style, supported by multi- agent orchestration and function-calling.
• Deployed the application using Streamlit, Python, and cloud inference APIs, incorporating text-to-speech (TTS) to deliver high- quality, on-demand educational podcasts.
Financial Sentiment & Market Trend Analyzer
• Built a pipeline to scrape and process 10K+ daily financial news articles from multiple sources and social feeds using BeautifulSoup and NLTK; achieved 88% accuracy in trend prediction.
• Integrated findings into a Power BI dashboard to visualize the correlation between social sentiment and stock price volatility. Self-Healing MLOps Pipeline for Financial Risk
• Designed a distributed Deep Learning model using PyTorch to predict credit default probability, outperforming baseline XGBoost models by 12% in F1-score on a dataset of 10M+ records.
• Developed a custom automated drift detection service using MLflow and Prometheus that monitors Kolmogorov-Smirnov (KS) statistics in real-time on streaming Kafka data; reduced model downtime by 50%.
• Implemented Infrastructure as Code (IaC) using Terraform to manage Kubernetes (EKS) clusters, enabling "Blue-Green" deployments that ensured zero-downtime during critical model updates and version rollbacks.