Somil Urmil Shah
United States 857-***-**** Email LinkedIn GitHub Google Scholar Open to remote/relocate Summary
Data Engineer with 5 years of building reliable pipelines, governed data marts, and analytics-ready platforms across financial services, product analytics, research telemetry, and cloud migrations. Hands-on with Python, SQL, Scala, Spark/PySpark, dbt, Airflow, Kafka, AWS, GCP, Snowflake, and BigQuery, spanning batch, streaming, warehouse, and distributed data systems. Strong focus on data modeling, quality frameworks, performance optimization, and BI-ready datasets for Tableau, Power BI, and AI-enabled analytics workflows. Education
Master of Science in Information Systems Northeastern University, Boston, MA Sep 2023 – Dec 2025 Bachelor of Engineering in Computer Engineering University of Mumbai, India Jul 2016 – Oct 2020 Experience
Data Engineer Dreamline AI, Community Dreams Foundation, Boston, MA (Contract) Apr 2026 – Present
• Built Python ingestion pipeline with Scrapy, Playwright, pdfplumber, and APIs, extracting 150+ federal, state, and utility incentive programs into an 18-field Pydantic schema powering ZIP-keyed lookups across 40+ Tampa Bay ZIPs.
• Engineered Claude LLM extractor with function calling to parse raw HTML and PDFs into schema-valid JSON, enforcing null-on-missing and per-record confidence scoring to suppress hallucinated amounts and cut manual corrections 40%.
• Orchestrated Celery, Redis, and PostgreSQL refresh jobs with a field-level diff engine and Pydantic/Airtable review queue, holding 95%+ spot-check accuracy and serving FastAPI ZIP lookups under 2 seconds via cached results. Data & AI Engineer, RA Northeastern University, INTERACT Animal Lab, Boston, MA Oct 2024 – Dec 2025
• Built Python/Pandas pipelines for 1.5 TB animal interaction telemetry across sensor exports, behavioral observations, video metadata, device logs, and experiment-session records, adding 6+ quality checks for missing IDs, duplicates, nulls, corrupted exports, and schema drift.
• Modeled interaction-event datasets with session, device, behavior, enrichment, and experiment dimensions, cutting cross-session preparation from 6+ hours to under 15 minutes.
• Built RAG literature and enrichment workflows for 800+ ACI papers and zoo logs using AI extraction, searchable metadata, text chunks, offline-first SQLite, sync strategies, and analytics-ready records. Data Engineer Intern Plainsight, Seattle, WA (Remote) Jan 2025 – Apr 2025
• Built BigQuery/dbt marts across HubSpot CRM, LinkedIn Ads, GA4 events, and product activity data, developing 12+ models for acquisition, conversion, campaign, and operational KPIs.
• Automated ELT workflows processing 100K+ rows/day and implemented dbt tests for freshness, uniqueness, not-null, accepted values, and relationships.
• Wrote Python/Pandas validation scripts and Looker Studio dashboards to flag duplicate leads, missing campaign IDs, schema changes, and row-count anomalies, reducing ad-hoc reporting by 70%. Senior Data Engineer Citi, Mumbai, India Sep 2022 – Aug 2023
• Developed Snowflake/dbt payment marts, models, tests, snapshots, and reusable macros for exception analytics across payment events, settlement status, return/reject codes, delayed transactions, failed-payment rate, and settlement-delay metrics.
• Designed and optimized fact/dimension models for payments, clients, accounts, settlement events, and exception codes, improving reporting query performance by 35%.
• Developed PySpark/Spark SQL jobs on AWS EMR to process multi-million-row payment batches from S3 using partition pruning, deduplication, joins, and late-arriving event handling.
• Supported Kafka-based payment-event ingestion and automated 6+ reconciliation checks using SQL, Python validation scripts, Airflow retries, and CloudWatch logs, reducing recurring data quality issues by 20%. Data Engineer Tata Consultancy Services, Mumbai, India Sep 2020 – Sep 2022
• Developed SQL, PL/SQL, and Java ETL workflows for insurance agent lifecycle data, including onboarding, licensing, renewals, commissions, status changes, and policy-agent mappings.
• Built Java batch jobs and Oracle PL/SQL procedures to extract, validate, transform, and reconcile agent master, commission, audit, and reporting datasets.
• Modernized legacy Oracle reporting workflows by staging extracts in AWS S3, loading Redshift with COPY jobs, and validating reporting tables before cutover.
• Supported Oracle-to-Redshift migration for 2B+ lifecycle records, validating row counts, checksums, null rates, primary keys, and commission/status aggregates.
• Optimized Oracle SQL queries, stored procedures, joins, indexes, batch routines, trigger-based audit trails, change tracking, and Linux cron jobs for systems serving 2M+ users. Technical Skills
Programming & Querying: Python, SQL, Scala, Java, Bash, Pandas, NumPy, Pydantic Data Engineering: PySpark, Spark SQL, Kafka, Airflow, dbt, ETL/ELT, CDC, Batch Pipelines, Streaming Pipelines, Data Lakes Cloud & DevOps: AWS (S3, Glue, Redshift, EMR, Athena, Kinesis, Lambda, CloudWatch, Bedrock), GCP (BigQuery, Dataflow, Cloud Composer), Docker, Terraform, GitHub Actions, CI/CD Databases, Warehouses & Modeling: Snowflake, BigQuery, Redshift, PostgreSQL, MySQL, MongoDB, Star/Snowflake Schema, SCD, Partitioning, Query Optimization
AI, Quality & Analytics: RAG, LangChain, Vector DBs, Embeddings, OpenAI API, Gemini API, Claude, FastAPI, Great Expectations, dbt Tests, Power BI, Tableau, Looker Studio