Post Job Free
Sign in

Data Science Research Analyst & ML Engineer

Location:
Chicago, IL
Posted:
September 23, 2026

Contact this candidate

Resume:

AB BASIT SYED

Chicago, IL 517-***-**** ******@***.*** LinkedIn GitHub

EDUCATION

Michigan State University — Master of Science, Data Science Michigan, USA May 2026

University of Kashmir — Bachelor of Engineering, Computer Engineering Kashmir, India Nov 2019

SKILLS

Programming & Data: Python, SQL, R, Pandas, NumPy, SciPy, Scikit-learn

Statistics & Experimentation: Statistical Analysis, A/B Testing, Experiment Design & Evaluation, Hypothesis Testing, Root-Cause Analysis, Regression, Time-Series Analysis

Machine Learning & AI: Predictive Modeling, Classification, Feature Engineering, Time-Series Forecasting, XGBoost, LightGBM, Random Forest, RAG, FAISS

Data & ML Systems: AWS (Lambda, Step Functions), MLflow, FastAPI, Model Deployment & Monitoring, Streamlit, Docker, Git

Visualization & Reporting: Tableau, Power BI, Excel (pivot modeling, VBA-assisted analysis)

EXPERIENCE

Research Analyst — Henry Ford Health & SLIM Lab Michigan, USA Jan 2025 – Present

•Reduced EEG/ERP analysis turnaround time 60% (10 to 4 hours per batch) by building automated Python pipelines for time-series signal processing, feature extraction, and data quality validation across a 65-participant research study.

•Identified the sources of intervention-effect variance across 4 experimental groups using hypothesis testing and pre/post comparison, presenting findings to the research team and directly informing 3 of 4 adopted methodology refinements.

Data Scientist (Supply Chain) — Meijer & Michigan State University Michigan, USA Jan 2026 – May 2026

•Reduced modeled stockout cost 67% while increasing service level 7% by evaluating 5 replenishment strategies and building a demand-forecasting and inventory-optimization solution across 14 months of distribution-center data covering 16,000+ SKUs, presenting recommendations to Meijer stakeholders and faculty advisors.

•Reduced holdout forecast MAE 18% versus a seasonal-naive baseline through structured model evaluation comparing gradient-boosted trees, Ridge regression, and baseline forecasting approaches.

Sr Data Analyst — Teleperformance Gurgaon, India Sep 2023 – Jun 2024

•Increased onboarding adoption 15% by analyzing multi-quarter user engagement data and developing propensity-based customer segments through feature engineering in Python, SQL, and Scikit-learn.

•Improved customer retention 12% by combining hypothesis-driven analysis with Logistic Regression and Random Forest models, using Excel-based exploratory analysis and summary reporting to align payments and product teams on customer risk signals.

•Reduced production ML pipeline latency 60% by building and operationalizing an AWS Lambda and Step Functions pipeline for real-time model inference and monitoring across payments and financial-services workflows.

Data Analyst — Fidelity National Information Services (FIS) Gurgaon, India Feb 2020 – Sep 2023

•Reduced customer churn 8% across 5 geographic segments by analyzing behavioral and retention patterns in Python and SQL, identifying root-cause drivers of attrition, and translating findings into retention and marketing strategy for cross-functional stakeholders.

•Influenced 4+ product roadmap decisions by designing and evaluating A/B experiments and applying XGBoost-based customer lifetime value analysis to quantify pricing, engagement, and lifecycle opportunities, presenting results directly to product and marketing leadership.

•Improved reporting accuracy from 74% to 93% while reducing manual reporting effort 70% across 10+ Tableau and Power BI dashboards by standardizing KPI definitions, datasets, and validation checks for business stakeholders.

PROJECTS

Fraud Decisioning Platform Python, LightGBM, FastAPI GitHub Streamlit

•Built and deployed a real-time fraud decisioning pipeline across 590K+ transactions and 400+ features, achieving 0.92 AUC-ROC, 89% precision@Top-1000, and 80% recall on a 3.5% fraud-rate dataset, with FastAPI model serving, drift monitoring, and capacity-aware prioritization that reduced estimated false-positive review cost 35%.

RAG Resume Screening GPT-4, FAISS, LangChain, RAGAS GitHub Streamlit

•Improved candidate-selection accuracy from 35.8% to 55.8% across 5 evaluation folds by building a LangChain RAG-Fusion pipeline with GPT-4, sentence-transformer embeddings, FAISS vector search, Reciprocal Rank Fusion, and RAGAS evaluation.



Contact this candidate