Post Job Free
Sign in

Data Engineer & Analytics Specialist

Location:
South San Francisco, CA, 94083
Posted:
August 10, 2026

Contact this candidate

Resume:

Nitin Vemuri

925-***-**** ***********@*****.*** San Ramon, CA github.com/nitinvemuri

SUMMARY

Data Engineer and Analytics professional with hands-on experience designing ETL pipelines, building ML-powered data products, and delivering scalable dashboards and reporting tools for operational and scientific stakeholders. Proficient in Python, SQL, and cloud platforms (AWS, Azure). MSBA candidate at UC Davis with a track record of converting raw, complex datasets into production-ready data assets.

TECHNICAL SKILLS

Languages: Python, SQL, Java, HTML5, JavaScript ML / Data: FAISS, scikit-learn, XGBoost, Pandas, NumPy, Matplotlib, Plotly, SHAP

Pipelines & Cloud: ETL/ELT (Python), AWS, Azure, Docker, Kubernetes, Git Visualization: Power BI, ArcGIS, Streamlit, Plotly

Databases: Relational SQL, Restaurant365 ERP, Power BI datasets Other: Bloom Filters, HNSW/ANN indexing, REST APIs, Bootstrap, Angular

EXPERIENCE

Data Engineer / Business Analyst — Practicum Aug 2025 – Present

El Dorado Irrigation District · Placerville, CA

• Designed and deployed a full ETL pipeline in Python and SQL to ingest, clean, and normalize 15 years of multi-source South Fork American River environmental monitoring data — resolving schema inconsistencies, duplicates, and missing values across time-series records.

• Built automated data validation scripts to enforce schema standards and flag anomalous readings before downstream consumption, ensuring data integrity for compliance and legal documentation purposes.

• Engineered interactive Power BI dashboards tracking key river variables (flow rate, turbidity, dissolved oxygen, temperature) with embedded ArcGIS geospatial layers, enabling non-technical biologists to self-serve trend analysis across the watershed.

• Delivered production-grade reporting outputs actively used by EID operations staff and serving as audit-ready documentation for regulatory and legal proceedings — requiring strict data quality and reproducibility standards.

Data Engineer / Developer Sep 2023 – Aug 2024

Aptiva Corp · Remote

• Built Python scripts to automate data quality checks, identify integrity issues, and reconcile discrepancies across multiple systems — reducing manual validation effort and standardizing data inputs for downstream reporting.

• Developed KPI dashboards and operational reports in response to stakeholder-gathered requirements, translating business needs into data-driven solutions used across cross-functional teams.

• Investigated production data issues by analyzing application logs and system outputs to identify root causes and recommend corrective actions, supporting system reliability and data pipeline uptime.

• Supported UAT for production releases — documenting business requirements, tracking action items, and validating data outputs against expected results prior to deployment.

• Contributed to cloud-based automation initiatives using Python, SQL, AWS, and Azure, improving system efficiency and reducing manual operational overhead.

Business Analyst Intern Jun 2023 – Aug 2023

Aptiva Corp · Remote

• Led intern cohort capstone project end-to-end — managing scope, task assignments, timelines, and deliverable quality from kickoff through shareholder presentation to senior leadership.

• Supported data gathering and analysis using Python and Excel to inform business decisions and track project milestones across workstreams.

Data & Procurement Analyst Oct 2024 – May 2025

Trio Donuts · Roseville, CA

• Operated Restaurant365 ERP platform end-to-end for procurement and inventory management — creating purchase orders, managing vendor relationships, reconciling stock levels, and generating procurement performance reports.

• Monitored inventory data daily to identify shortages and discrepancies, coordinating supplier resolution to maintain uninterrupted supply chain operations.

UI & Automation Intern Jun 2022 – Aug 2022

Vikasha Corp · Remote

• Built automated reporting and reminder systems using Python, HTML5, Bootstrap, and Angular; developed internal employee management software integrating Workday and Replicon platforms.

PROJECTS

NutriAI — FAISS-Powered Meal Planning Engine — Personal Project 2024–2025

Python · FAISS · Bloom Filters · Thompson Sampling · USDA FoodData Central API

• Designed and built a full ML data pipeline ingesting 11,000+ USDA foods, encoding each as a 13-dimensional nutritional vector, and indexing into a FAISS HNSW approximate nearest-neighbor structure returning candidates in <5ms per query.

• Engineered a multi-stage ranking pipeline (candidate generation pre-rank full rank) with Bloom filters for O(1) allergen and clinical-condition screening, and a Thompson sampling bandit for 7-day diversity optimization balancing exploration vs. exploitation.

• Debugged and resolved a critical state-management bug in the week-level planning architecture; validated across 5 clinical personas (hypertension, diabetes, vegan/IBS, Jain) achieving 80% constraint-satisfaction rate.

Non-Profit Financial Intelligence Platform — Aggie Hacks 2026 — Finalist 2026

Python · XGBoost · scikit-learn · SHAP · Streamlit · Plotly · Pandas

• Built an end-to-end analytics data pipeline processing 7 years of IRS Form 990 data across thousands of U.S. nonprofits — engineering a composite Financial Resilience Score from four weighted sub-metrics for grant-making decisions.

• Trained a distress prediction pipeline (Logistic Regression, Random Forest, XGBoost) with SHAP explainability to surface at-risk organizations; built a shock simulator modeling revenue-loss scenarios.

• Delivered a production-ready multi-page Streamlit dashboard with interactive benchmarking and radar charts, replacing hours of manual analyst review for Fairlight Advisors.

SF Housing Price Prediction Model — UC Davis MSBA — Group Project 2025

Python · scikit-learn · Random Forest · LassoCV · PCA · K-Means

• Built the core model-evaluation framework used across 6 models (Linear Regression, LassoCV, PCR, K-Means, 3 Random Forests), computing CV R and RMSE for consistent apples-to-apples comparison.

• Contributed to LassoCV feature selection reducing ~150 sparse features to 71–91; best model (tuned RF) achieved CV MAE of $597,327 — 13% improvement over baseline.

EDUCATION

M.S. Business Analytics Aug 2025 – Aug 2026

University of California, Davis · GPA in progress

B.A. Economics Dec 2022

University of California, Davis

Certifications: Java Developer (Revature) · Web Design (Berkeley Extension) · Biotechnology / Bio-Manufacturing (Ohlone College)



Contact this candidate