Post Job Free
Sign in

Product Data Scientist (Civic Data Science)

Location:
Hayward, CA
Posted:
August 18, 2026

Contact this candidate

Resume:

NIKHIL AGRAWAL

San Francisco, CA ************@*****.*** 1-341-***-**** GitHub LinkedIn Portfolio

SUMMARY

Product-focused Data Scientist with an M.S. in Statistics and experience turning ambiguous stakeholder questions into measurable analytical problems, reliable datasets, evaluation systems, and user-facing data products. Built Python and SQL analytics, evaluation workflows across approximately 1,600 civic meetings, and canonical data systems spanning 1.7M+ profiles and approximately 4M historical records. Combines statistical inference, experimentation, product analytics, data quality, and production engineering with founder experience translating customer and operational problems into business decisions.

TECHNICAL SKILLS

Analytics & Statistics: Python, SQL, R, Pandas, NumPy, SciPy, Scikit-learn, Statistical Inference, A/B Testing, Experimental Design, Hypothesis Testing, Confidence Intervals, Bootstrapping, Regression, Segmentation, Error Analysis Product Data Science: Product Analytics, KPI Definition, Funnel Analysis, Cohort Analysis, Experiment Analysis, Business Trend Monitoring, Data Visualization, Statistical Communication

Data & Scale: PySpark, Apache Spark, Hive, MongoDB, PostgreSQL, ETL, Data Modeling, Data Validation, Temporal Data, Parquet Visualization & Delivery: Tableau, Dashboards, PowerPoint, Analysis Reports, Stakeholder Presentations Engineering: FastAPI, Flask, Docker, Git, Azure, REST APIs PROFESSIONAL EXPERIENCE

Data Scientist, Open Governance Initiative (OpGov.ai)Jan 2025 – Present San Francisco Bay Area

•Civic researchers lacked a reliable way to evaluate answers across fragmented meeting records, so built and validated hybrid RAG/KG-RAG workflows over approximately 1,600 meetings and a versioned evaluation system measuring retrieval, attribution, citations, and abstention, giving product stakeholders measurable quality gates for release decisions.

•Incompatible county schemas and rolling data snapshots made longitudinal analysis unreliable, so engineered a canonical platform spanning 1.7M+ profiles and approximately 4M historical records and reconciled a 734,457-record update, identifying 20,060 new profiles while isolating one ambiguous identity rather than risking an incorrect merge.

•Product and engineering teams needed a consistent way to investigate usage and quality issues, so built Python/SQL analytics for search behavior, feature usage, and failure patterns while translating stakeholder goals into KPI definitions and acceptance criteria, creating a repeatable basis for prioritization and regression analysis. Founder & Technical Lead, R P Era Aug 2018 – Aug 2023 India

•Customers arrived with loosely defined engineering requirements, so translated discovery into technical specifications, material choices, estimates, timelines, and production plans across 100+ engineering and dental projects, creating a repeatable path from business need to accepted technical outcome.

•Quotations, inventory, production status, invoicing, and delivery information were distributed across manual workflows, so built a Python/Flask order-management system and operational reporting processes that centralized pipeline visibility and improved project, stock, and delivery decision-making.

•Customer, sales, inventory, and production data were underused in day-to-day planning, so developed analytical and reporting workflows to surface bottlenecks and higher-value opportunities, enabling more informed prioritization across capacity, inventory, vendors, and customer engagements.

CREATIVE PROJECTS

Multi-LOB Web Funnel Experimentation & Conversion Lab, Python SQL Tableau Statistical Experimentation

•Built an experimentation pipeline over synthetic multi-service marketplace web events to diagnose funnel abandonment and evaluate UX changes using SQL, power analysis, A/B testing, confidence intervals, CUPED-adjusted lift, and device, market, and product-segment analysis.

•Delivered a Tableau decision dashboard and experiment readout separating overall conversion lift from segment-level effects, enabling defensible launch recommendations while exposing guardrail regressions hidden by aggregate metrics. Experiment Reliability & Decision Control Plane, Python SQL Statistical Testing Experiment Monitoring

•Built a versioned experiment-quality system that automatically evaluated sample-ratio mismatch, instrumentation loss, invariant metrics, sequential monitoring, metric drift, and heterogeneous treatment effects before experimental results reached decision-makers.

•Created statistical release gates and root-cause reports separating true product effects from data-quality and experimental-design failures, reducing the risk of product decisions based on misleading averages or broken telemetry. Urban Mobility Access & Event Impact Lab, Python SQL Geospatial Analysis Time Series Statistical Inference

•Combined public trip, weather, geographic, transit, and event data to estimate how external shocks affected neighborhood mobility patterns using rolling forecasts, geospatial segmentation, regression, and quasi-experimental analysis.

•Produced an operational heatmap and decision report identifying when and where mobility-access gaps emerged, translating statistical results into resource-planning recommendations for transportation and event stakeholders. Scalable Product Telemetry Lakehouse & BI Reliability Monitor, PySpark Spark SQL Hive Parquet SQL Tableau

•Built a large-scale product telemetry lakehouse using PySpark, Hive, Parquet, partitioning, Spark SQL, and automated data-quality checks, transforming tens of millions of raw events into reusable session, funnel, cohort, and KPI marts.

•Benchmarked scalable query strategies and served validated metrics to Tableau dashboards for business-trend monitoring, anomaly detection, and executive reporting, demonstrating an end-to-end path from distributed processing to product decisions. EDUCATION

Master of Science in Statistics - Data Science Concentration California State University - East Bay GPA: 3.9/4.0 2023 – 2025 Hayward, CA

Relevant Coursework: Experimental Design, Statistical Inference, Regression, Resampling Methods, Machine Learning, Database Systems, Natural Language Processing

Bachelor of Engineering - Mechanical Engineering, University of Mumbai 2014 – 2018 Mumbai, IN NIKHIL AGRAWAL 1 / 1



Contact this candidate