Post Job Free
Sign in

Applied Geospatial Data Scientist

Location:
Eloy, AZ
Salary:
200000
Posted:
September 01, 2026

Contact this candidate

Resume:

OMER MEJIA

Brooklyn, NY 646-***-**** *****.****@*******.*** https://www.linkedin.com/in/omer-mejia-367375414/

PROFESSIONAL SUMMARY

Applied data scientist with 15+ years of experience delivering reliable, reproducible Python analytics across large-scale datasets, with deep strength in automation, documentation, and stakeholder-facing delivery. Proven ability to implement end-to-end workflows from data acquisition and cleaning through processing, model execution, validation, and polished outputs for decision-makers. Strong foundation in statistics, machine learning, and scalable data processing, with hands-on experience structuring repeatable pipelines, reusable notebooks, and version-controlled code using Git/GitHub. Adept at evaluating data quality, surfacing caveats and risks early, and conducting rigorous accuracy assessments and quality-control checks. Collaborative, remote-ready engineer who communicates technical concepts clearly to non-technical partners while supporting mission-driven work across land, water, infrastructure, and environmental use cases.

KEY SKILLS

Geospatial Data Science (Python): Python, raster data, vector data, geospatial analysis, remote sensing, spatial joins, reprojection, tiling/windowing, large geospatial datasets, Earth observation

GIS Tools and Open-Source Stack: ArcGIS, QGIS, GDAL, Rasterio, GeoPandas, Shapely, Fiona, pyproj, OGR

Modeling, GeoAI, and Statistics: machine learning, deep learning, spatial statistics, feature engineering, cross-validation, error analysis, uncertainty assessment, accuracy assessment, model monitoring

Data Engineering and Reproducibility: data acquisition, data cleaning, ETL/ELT, data dictionaries, metadata, reproducible project structures, documentation, notebooks, packaging reusable components

Automation and Workflow Orchestration: automation, scripting, reusable functions, batch processing, parameterized notebooks, CI checks, testable pipelines, performance optimization, standardization

Cloud and Scalable Processing: AWS, object storage patterns, distributed processing concepts, containerization concepts, scalable workflows, cost/performance tradeoffs, prototype-to-production handoffs

Collaboration and Delivery: Git/GitHub, peer review, technical memos, client deliverables, proposals support, risk assessments, milestone planning, remote collaboration, stakeholder communication

EXPERIENCE

Senior Data Scientist (Geospatial Analytics), UnitedHealth Group, Brooklyn, NY — Jan 2022– Present

- Implemented repeatable Python geospatial analytics workflows from data acquisition through processing, analysis, model execution, validation, and delivery to support location-driven program decisions.

- Acquired, organized, cleaned, and evaluated large raster-like gridded datasets and vector boundary layers from public and partner sources, improving usability through standardized schemas and checks.

- Automated error-prone processing steps with reusable scripts, functions, and parameterized notebooks, reducing manual reruns by 45% and improving consistency across projects.

- Conducted spatial statistics and rigorous accuracy assessments using established evaluation methods, documenting assumptions, limitations, and performance caveats for senior technical review.

- Developed quality-control gates to review outputs for completeness and technical integrity, catching projection mismatches, invalid geometries, and missing metadata before stakeholder delivery.

- Designed and iterated machine learning workflows that integrated spatial features, tabular predictors, and engineered distance/adjacency signals, improving model utility for regional decision contexts.

- Evaluated model outputs for performance issues and data limitations, escalating risks related to sampling bias, label leakage, and spatial autocorrelation to senior staff early in delivery cycles.

- Maintained reproducible project structures with clear data dictionaries, metadata, and technical handoff materials, enabling new contributors to onboard in days instead of weeks.

- Partnered with AI engineers and platform teams to move prototypes into scalable workflows, aligning compute needs and artifact storage patterns with AWS operating constraints.

- Established Git/GitHub branching and peer review norms for analytics repositories, increasing review participation and reducing production defects tied to untested changes.

- Produced concise technical memos and stakeholder-ready summaries translating computational findings into clear tradeoffs, limitations, and recommended actions for non-technical audiences.

- Coordinated milestones and scoped workstreams for multi-team initiatives, converting ambiguous stakeholder needs into measurable deliverables, risk assessments, and trackable acceptance criteria.

- Tested new datasets and methods, sharing practical findings and “gotchas” with the broader team to increase reuse and reduce duplicated exploratory effort.

- Collaborated effectively in a fully distributed environment, providing forthright feedback, documenting decisions, and improving team delivery predictability across concurrent engagements.

Senior Data Scientist (Applied Analytics), Capital One, Brooklyn, NY — Jan 2015– Dec 2021

- Built end-to-end Python data science workflows spanning ingestion, cleaning, feature engineering, modeling, and monitored delivery, supporting multiple stakeholder groups under tight timelines.

- Implemented repeatable analytical outputs using notebooks and reusable modules, standardizing code patterns and documentation practices that improved auditability and reproducibility.

- Applied statistics and machine learning techniques to complex, noisy datasets, performing error analysis and validation to ensure reliable conclusions and defensible recommendations.

- Established quality-control checks and data validation routines to detect drift, missingness, and outliers, reducing downstream rework and improving stakeholder trust in outputs.

- Produced thorough documentation of data sources, assumptions, limitations, and results, enabling consistent reuse of analytical assets across teams and quarters.

- Led peer review practices using Git/GitHub workflows, improving code quality through structured reviews, consistent style guidelines, and regression test expectations.

- Automated repetitive reporting and data preparation steps with robust scripting, improving delivery speed by 30% while lowering the likelihood of manual processing errors.

- Collaborated with cross-functional partners to translate client and internal needs into scoped milestones, clear definitions of done, and proactive risk tracking.

- Communicated technical findings to non-technical audiences through memos and presentations, emphasizing caveats, data quality constraints, and decision-relevant implications.

- Evaluated competing methods and datasets for fitness, balancing interpretability, performance, and operational constraints while escalating limitations to senior technical leadership.

- Supported proposal-like internal initiatives by estimating effort, identifying dependencies, and crafting practical implementation paths from prototype to scalable workflow.

- Mentored teammates on reproducible research practices, including project structures, code organization, and documentation patterns that improved long-term maintainability.

Senior Software Engineer (Data & Analytics Platforms), Etsy, Brooklyn, NY — Oct 2010– Dec 2014

- Developed Python-based data processing components to acquire, clean, and transform multi-source datasets into analytics-ready forms, improving consistency for downstream modeling.

- Implemented repeatable pipelines and reusable utilities that standardized common transformations, enabling faster iteration and reducing duplicated code across teams.

- Collaborated with analysts and engineers to define data requirements, data quality expectations, and validation checks that improved trust in reported metrics.

- Built robust logging and monitoring patterns for batch workflows, accelerating root-cause analysis and reducing time-to-recovery during pipeline failures.

- Produced documentation of processing steps, code assumptions, and data definitions, creating clearer handoffs and enabling smoother cross-team collaboration.

- Applied statistical thinking to interpret experimental results and observed shifts in key metrics, flagging likely confounders and data integrity issues to stakeholders.

- Automated recurring data preparation and reporting tasks, reducing manual effort while increasing repeatability and lowering the chance of human error.

- Used Git/GitHub-style version control practices to manage changes safely, incorporating code reviews to improve maintainability and reduce regressions.

- Optimized processing performance through profiling and algorithmic improvements, lowering runtime and compute overhead for large historical backfills.

- Partnered with product and operations teams to translate questions into measurable analyses, ensuring outputs aligned with business decisions and time constraints.

- Assisted with technical demonstrations and presentations, communicating methods, limitations, and takeaways clearly to mixed technical and non-technical audiences.

- Contributed to an inclusive, learning-oriented engineering culture through peer feedback, shared standards, and collaborative problem-solving across disciplines.

EDUCATION

Master's Degree in Computer Science, Alliance University (2013–2015)

Bachelor's Degree in Computer Science, Alliance University (2006–2010)



Contact this candidate