Nafeesa Samreen Kolthuru
+1-513-***-**** ******************@*****.*** linkedin.com/in/nafeesam
PROFILE SUMMARY
Data Analytics and Data Science professional with 7 years of experience leveraging data to drive business decisions across financial services, healthcare, and supply chain domains. Expertise in SQL, Python, Tableau, machine learning, statistical analysis, forecasting, risk analytics, fraud detection, KPI reporting, and data visualization. Proven track record of developing predictive models, conducting experimentation, building executive dashboards, and translating complex analytical findings into actionable business insights. Experienced in working with large-scale datasets on AWS and collaborating with cross- functional stakeholders to improve operational efficiency, reduce risk exposure, and support strategic decision-making. CORE COMPETENCIES
Data Analysis & Business Intelligence, SQL Querying & Data Exploration, Statistical Analysis, Predictive Modeling, Machine Learning, Risk & Fraud Analytics, Forecasting & Time Series Analysis, A/B Testing & Experimentation, KPI Reporting & Dashboard Development, Tableau & Power BI, Data Visualization, Stakeholder Communication, Root Cause Analysis, AWS Analytics Solutions WORK EXPERIENCE
Senior Data Analyst — Risk Analytics TIAA Financial Nov 2023 – Present
• Led end-to-end AI development lifecycle for predictive risk models from requirements gathering and data sourcing through governance review and production deployment on AWS (S3, Lambda, Glue) improving underwriting accuracy and reducing loss ratios by 12%; presented model findings and risk tradeoffs directly to senior leadership.
• Automated fraud review workflows using Python, SQL, and AI-driven analytics, reducing manual review effort by 30% and improving operational efficiency through scalable automation solutions.
• Developed AI-powered analytics solutions using AWS Bedrock and large language models to automate document analysis and anomaly detection, accelerating business process efficiency and supporting risk management initiatives.
• Designed A/B-style incrementality experiments on time-series forecasting models (ARIMA, Prophet, LSTM) to evaluate scenario outcomes and support strategic capital allocation decisions; delivered high-profile readouts to cross-functional senior stakeholders.
• Developed SQL-driven KPI reporting infrastructure and Tableau dashboards for fraud exposure and premium flow monitoring, enabling real-time portfolio visibility and executive decision-making.
• Automated catastrophe loss estimation using AWS (S3, Glue, EMR) and weather APIs, improving risk assessment efficiency by 25%; scoped and delivered initiative independently, proactively surfacing blockers and managing transparent communication with leadership throughout. Data Scientist — Supply Chain & Clinical Analytics Cardinal Health Dec 2022 – Oct 2023
• Developed and deployed ML forecasting models (ARIMA, Prophet, XGBoost) on AWS SageMaker for supply chain scenario analysis and operational planning, boosting efficiency by 20%; collaborated with product and operations stakeholders to align model outputs with strategic decisions.
• Developed analytics solutions using Python and NLP techniques to analyze large-scale clinical datasets, identifying key safety trends and delivering actionable recommendations to business stakeholders.
• Built Tableau and Power BI dashboards with anomaly detection tracking clinical trial KPIs, enabling faster cross-functional decisions and supporting regulatory submissions partnered directly with product owners to define requirements and validate findings.
• Performed advanced data analysis on EHR and clinical trial datasets using SQL and Python, developing patient stratification frameworks that supported strategic healthcare decisions. Data Analyst — Digital & Financial Analytics South Indian Bank Dec 2020 – Nov 2022
• Designed and ran end-to-end A/B tests and incrementality experiments on digital onboarding features from hypothesis definition through analysis and rollout improving customer adoption rates by 18%; owned full project lifecycle and presented findings to business stakeholders.
• Built predictive models (multivariate regression, classification, cohort analysis) on AWS to forecast loan defaults and optimize operations, reducing defaults by 15% and saving $2M annually led initiative from data ingestion to stakeholder presentation with minimal oversight.
• Championed AI literacy and self-serve analytics by building Tableau and Power BI dashboards for branch performance, churn, and digital adoption KPIs, and training non-technical teams on dashboard usage to drive data-informed decision-making.
• Data Analyst — Pharma Analytics & Engineering Hetero Drugs Ltd., Jun 2018 – Nov 2020
• Built automated ETL pipelines (Airflow, Talend, AWS Glue) processing large-scale clinical and manufacturing datasets with Python and SQL; ensured compliance-ready data delivery for FDA/EMA submissions.
• Deployed real-time pharmacovigilance data pipelines (Kafka, Spark) enabling faster safety signal detection and created insight dashboards that improved operational efficiency by 15%.
• Used SQL and Python to conduct root cause analysis on manufacturing quality failures, identifying process variables that reduced batch rejection rates by 18%.
• Partnered with QC, regulatory, and operations teams to define KPI frameworks and translate clinical data into actionable reporting, enabling faster data-driven decisions. TECHNICAL SKILLS
Visualization & BI: SQL, Tableau, Power BI, Looker Studio, Plotly, Dash, Streamlit, Excel(Pivot Tables, Power Query, XLOOKUP, Data Analysis).
Languages & Querying: Python (expert), SQL (expert), R, Spark/PySpark — Snowflake, Teradata, PostgreSQL, MySQL, MongoDB.
AWS Services: SageMaker, Bedrock, JumpStart, Textract, Lambda, S3, Athena, Glue, EMR, Redshift, EC2. Machine Learning & Statistical Methods: Machine Learning, NLP, Deep Learning, A/B Testing, Classification, Regression, Forecasting (ARIMA, Prophet, LSTM), XGBoost, LightGBM, scikit-learn, TensorFlow, PyTorch, Keras. AI & LLMs: Generative AI, LLMs (Claude, LLaMA, Titan), LangChain, Prompt Engineering, Vector Databases, Embeddings, Retrieval-Augmented Generation (RAG), AI Workflow Automation. Data Pipelines & MLOps: dbt, Airflow, Kafka, Hadoop, Docker, Kubernetes, MLflow, FastAPI, Flask, CI/CD, Git. Statistical Analysis & Experimentation: Hypothesis Testing, Statistical Significance, Confidence Intervals, A/B Testing, Cohort Analysis, Regression Analysis.
Data Quality & Governance: Data Lineage, Data Validation, Great Expectations, Data Cataloging, Compliance Reporting
(FDA/EMA), Data Dictionaries.
Collaboration & Project Tools: Jira, Confluence, Git, Agile/Scrum, Stakeholder Reporting, Cross-functional Communication.
EDUCATION
Master of Science, Computer Science — Wright State University 2024 PROJECTS
Credit Card Fraud Detection Python, scikit-learn, SMOTE, Random Forest Built an end-to-end fraud detection model on a dataset of 284,807 transactions with a 0.17% fraud rate. Handled class imbalance using SMOTE, trained and compared Logistic Regression and Random Forest models, achieving 91% precision and 85% recall on fraud detection.