Nithin Reddy Kuchukulla Data Scientist
*********************@*****.*** +1 (507)-497- 2070 https://www.linkedin.com/in/nithin-reddy-kuchukulla-5a6857284/
PROFESSIONAL SUMMARY
Data scientist with over 5+ years of experience applying machine learning, statistical analysis, and advanced research methods to solve complex business problems across financial and health domains. Skilled in working with high volume and high dimensional datasets and proficient in SQL, Python, Snowflake, and Teradata to create scalable data pipelines, automated modeling workflows, and reliable analytical processes. Experienced in building classification, regression, clustering, time series, NLP, and deep learning models that generate meaningful insights and drive measurable business impact. Adept at exploratory data analysis, feature engineering, hypothesis testing, and design of experiments to uncover patterns and validate decisions. Known for translating ambiguity into clear technical solutions, communicating insights in a business friendly way, and partnering closely with finance and operational teams to support strategic planning, improve forecasting, and strengthen decision making. Committed to developing strong understanding of business operations and the MRO market to deliver insights that keep customer operations running effectively.
TECHNICAL SKILLS
Programming Language: Python, R, SQL, Scala, Java, Julia, Bash, MATLAB, SAS, javascript, shell scripting, DSA.
Gen AI & LLM’s: GPT-4, Prompt engineering, Feature Engineering, RAG pipelines, BERT, Fine-Tuning, embeddings, LangChain, LlamaIndex, LoRA, QLoRA, vector DBs (FAISS, Milvus, Pinecone, pgvector), function/tool calling, guardrails.
Machine Learning & Deep Learning: Supervised/Unsupervised, NumPy, pandas, scikit-learn, XGBoost, LightGBM, SVM, GLMs/GAMs, CatBoost, statsmodels, PyTorch, TensorFlow, Keras, ONNX, statsmodels, CNNs, RNN, ANNs, LSTM, GRU.
NLP & Computer Vision: Hugging Face Transformers, spaCy, NLTK, Gensim, fastText, sentence-transformers, OpenCV, torchvision, Ultralytics/YOLO, Detectron2, embeddings (word2vec/GloVe/SBERT), NER, POS, topic modeling (LDA/BERTopic).
Time Series & Recommenders: ARIMA/SARIMA, Prophet, TBATS, LSTM/Transformer TS, statsmodels, Matrix factorization, implicit ALS, factorization machines, bandits.
Data Engineering: Apache Spark/PySpark, Hadoop, Kafka, Flink, Beam, Airflow, dbt, Great Expectations, ETL/ELT, .
Big Data & Warehouses: Hadoop, Spark, MapReduce, Pig, Luigi, Snowflake, BigQuery, Redshift, Databricks Lakehouse, Hive.
Cloud: AWS (GovCloud, S3, EC2, Lambda, Glue, EMR, Redshift, Athena, Kinesis, SageMaker), GCP (GCS, Dataflow, Pub/Sub, Dataproc, BigQuery, Vertex AI), Azure (Blob, Synapse, Data Factory, Databricks, Azure ML).
DevOps: Git/GitHub/GitLab, CI/CD (Jenkins, GitLab CI), Docker, Kubernetes (EKS/GKE/AKS), Helm, Terraform, Ansible.
SAP Data Stack: SAP BW, SAP CRM, HANA Data Lake, Business Data Cloud, DataSphere
Data Visulization & BI Tools: Matplotlib, seaborn, Plotly, Bokeh, Altair, ggplot2, SAP, Tableau, Power BI, Looker, Apache Superset.
Database: PostgreSQL, MySQL, SQL Server, NoSQL, SQLite, MongoDB, Cassandra, DynamoDB, Neo4j, TimescaleDB, Elasticsearch/OpenSearch.
Data Management & ETL: ETL, Data Mapping, Data Transfer, Data Migration, Data Integration, Metadata Management, Data Dictionary, Data Lineage, Informatica, Talend.
Statistics & Modeling: statistical modeling, regression (linear, logistic, Poisson/Negative Binomial), GLM/GLMM, mixed-effects models, Bayesian modeling (hierarchical), survival analysis (Cox), time series (ARIMA, Prophet), clustering (k-means, HDBSCAN), Markov chains, HMM.
Attribution & Measurement: Attribution Modeling, MMM, MTA, Markov, Shapley, Incrementality, ROAS, CAC, LTV:CAC.
Product & UX Analytics: funnel analysis, cohort analysis, retention/activation (DAU/WAU/MAU), feature adoption, discoverability, task success rate, time-to-task, error rate, completion rate, NPS, CSAT, SUS, session replay analytics.
Structured Credit & Securitization: ABS, RMBS, CLO, CDO, securitization, waterfall modeling, cash-flow engines, prepayment modeling, default modeling, loss severity, roll-rate analysis.
Causal Inference & Experimentation: A/B testing, multivariate testing, sequential testing, multi-armed bandits, CUPED, power analysis, MDE, SRM diagnostics, randomization checks, guardrail metrics, causal inference, difference-in-differences (DiD), regression discontinuity (RDD), propensity score matching/IPW, instrumental variables, uplift modeling, heterogeneous treatment effects.
MLOps: MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, Databricks, DVC, Feast (feature store).
Model Evaluation & Explainability: ROC-AUC, PR-AUC, F1, RMSE/MAE/MAPE, Lift, SHAP, LIME, PDP/ICE, calibration.
Data Management & Security: data quality, data lineage, RBAC, row-level security, HIPAA, PHI, PII.
Quality & Monitoring: unit/pytest, black/flake8, pre-commit, data validation, drift monitoring (Evidently, WhyLabs, Arize), logging/observability.
Collaboration: Jupyter/Colab, VS Code, notebooks to prod, Agile/Scrum, Jira, Confluence, storytelling & stakeholder comms.
EDUCATION
University of Wisconsin-Milwaukee, Wisconsin, USA Sep 2023 – Dec 2024
Master’s in Information Technology (AI & Data Analytics) GPA: 3.6/4.0
WORK EXPERIENCE
Data Scientist COREBRIDGE FINANCIAL Houston, Texas, USA Feb 2025 - Present
Key Responsibilities:
Improved financial insights by partnering with business teams to understand complex problems and translate them into machine learning solutions that strengthened pricing and forecasting decisions and increased measurable impact by 35%.
Conducted detailed exploratory data analysis on customer and marketplace data across financial domains and uncovered patterns, trends, and anomalies that improved opportunity identification by 30%.
Worked with high volume and high dimensional datasets from multiple sources and applied strong SQL and Python skills to clean, manipulate, and combine structured and unstructured data stored in platforms like Snowflake, Teradata, and Oracle, improving data processing speed by 50%.
Developed high performing features by applying feature engineering, feature selection, and dimension reduction which increased model accuracy by 25% and enhanced interpretability for finance partners.
Strengthened analytics reliability by applying statistical hypothesis testing, design of experiments, outlier detection, and research driven methods to validate assumptions and improve evidence based financial recommendations by 20%.
Improved decision support across pricing, demand planning, and customer analytics by combining business knowledge with technical depth in statistical analysis, machine learning, data visualization, and experimentation.
Data Scientist WALGREENS BOOTS ALLIANCE Deerfield, Illinois, USA Mar 2024 - Dec 2024
Key Responsibilities:
Built and deployed scalable automated workflows for model development, model training, model scoring, and model validation that reduced manual effort by 40% and supported repeatable enterprise analytics.
Delivered predictive and prescriptive insights using classification, regression, clustering, time series forecasting, NLP, and deep learning methods which improved forecast stability and business planning accuracy by 18%.
Designed and executed experiments to measure Health and operational impact, collected necessary data, and translated statistical results into simple recommendations that supported strategic decisions.
Produced clear, compelling presentations and visualizations that summarized analytical findings, highlighted business implications, and enhanced communication with Health and operational leaders.
Built scalable data pipelines that ensured consistent data quality and improved analytic performance by 45%, while supporting large scale modeling and reporting needs.
Applied research experience to drive thoughtful investigation into customer behavior, product usage patterns, and market trends across health domains which improved insight discovery by 28%.
Developed strong understanding of business operations, go to market strategies, the maintenance repair and operations domain, and customer purchasing behaviors to improve model alignment with real business needs.
Data Scientist/Data Analyst MANIPAL CIGNA HEALTH INSURANCE Mumbai, India Sep 2021 - Aug 2023
Key Responsibilities:
Developed and maintained automated dashboards to track KPIs such as claim processing time, customer satisfaction, and network efficiency for various departments and applied Python and R for advanced data analysis, cleaning, and statistical modelling to inform strategic decisions and improve claims processing efficiency.
Conducted customer segmentation analysis with machine learning algorithms to uncover patterns and optimize targeting strategies for health products and retention efforts and analysed financial health data, including claims costs and revenue, to optimize pricing strategies and market share using Excel and Python for detailed financial modelling.
Monitored and reported on operational performance metrics using Hadoop and Big Data technologies, ensuring data accuracy for real-time decision-making in healthcare operations and utilized AWS Redshift to manage and analyse large health datasets, ensuring high data quality and optimizing query performance for better insights.
Analysed large health-related datasets to extract actionable insights, improving operational strategies across customer health, claims, and network performance using SQL, Excel, and Tableau.
Led A/B testing for new health products, services, and marketing campaigns using Google Analytics and Excel, delivering detailed performance analysis to guide decision-making and collaborated with customer service and marketing teams to create customer satisfaction surveys, analysing feedback data using SQL and Python to drive service improvements.
Automated ETL workflows using Talend, Apache NiFi, and Apache Airflow for seamless integration of data from CRM systems, financial ledgers, and customer service portals and utilized NoSQL databases like MongoDB to manage unstructured customer feedback, loan application documents, and scanned ID proofs for verification processes.
Data Analyst MANAPPURAM FINANCE’S Mumbai, India July 2019 – Aug 2021
Key Responsibilities:
Developed predictive models using Python, R, and machine learning libraries like scikit-learn and TensorFlow to forecast loan defaults, optimize interest rate pricing, and improve customer segmentation.
Managed cloud data infrastructure using Snowflake and Amazon Redshift, ensuring secure, scalable storage and efficient access to financial and operational data.
Performed data preprocessing and statistical analysis with Pandas and NumPy, and generated visual reports using Matplotlib and Seaborn to support decision-making on loan product performance.
Developed containerized microservices for data transformation using Docker, implemented Git for version control, and leveraged Apache Hive for querying large-scale lending datasets stored in Hadoop.
Built and maintained real-time and batch data pipelines using Apache Spark, Hadoop, and Apache Airflow to process gold valuation data, KYC verification logs, and transaction records.
Created interactive dashboards and performance reports using Tableau, Power BI, and Excel to visualize financial KPIs such as gold loan disbursement trends, branch-wise performance, and customer acquisition metrics.