HARISH ARVINDH RAMALINGAM
571-***-**** ******.********@*****.***
F-1 STEM OPT, authorized to work through 2028; sponsorship required thereafter SUMMARY
Data Engineer with 4+ years designing and operating batch and streaming data platforms across banking, enterprise hardware, and financial services. Strong in Python, SQL, and PySpark, with production experience on AWS, Databricks, Snowflake, dbt, Airflow, and Kafka. Comfortable owning pipelines end to end, from ingestion and modeling through orchestration and data quality, with hands-on side projects applying LLM-based tooling to data extraction and retrieval. TECHNICAL SKILLS
Languages: Python (PySpark, pandas), SQL (CTEs, window functions, query optimization), Spark SQL, Bash Processing & Orchestration: Apache Spark, Apache Airflow, dbt (tests, docs, incremental models), Trino, Apache Kafka, ETL/ELT design Data Platforms & Modeling: Databricks, Snowflake, Amazon Redshift, Hive/Hadoop, PostgreSQL, MongoDB; OLAP, dimensional (star-schema) modeling Storage & Table Formats: AWS S3, Apache Iceberg, Delta Lake, Parquet; time travel, schema evolution Ingestion, Cloud & Quality: AWS (Glue, EMR, Lambda), Fivetran, CDC, Docker, Git, GitHub Actions CI/CD; Great Expectations, data contracts, schema validation, SLA & data-freshness monitoring
AI / LLM Tooling: LangChain, Ollama, pgvector, RAGAS, Pydantic, FastAPI; RAG, embeddings, LLM-based extraction PROFESSIONAL EXPERIENCE
Aon Hewitt Remote, USA
Data Engineer May 2025 – Present
• Built PySpark ingestion standardizing daily enrollment, contribution, and eligibility files from 40+ employer and insurer feeds into S3 and Snowflake, with an event-driven Lambda staging each file on arrival.
• Modeled participant-level star-schema marts in dbt, joining enrollment, contribution, and claims in Snowflake into the reporting source of truth for benefits analysts and client-delivery teams.
• Enforced data contracts with dbt and Great Expectations tests in GitHub Actions CI/CD, using ACID-compliant transactional loads so a failed check leaves the previously validated Snowflake mart intact instead of a partial write.
• Rebuilt the nightly refresh as Airflow DAGs with per-feed dependencies and retries, lifting on-time delivery from about 92% to 98% and holding the morning data-ready SLA despite late-arriving sources.
• Converted the heaviest marts to incremental dbt models processing only new and changed records, cutting the nightly refresh from about 90 to 60 minutes and lowering Snowflake compute cost.
• Helped build an embedding pipeline chunking plan documents into a pgvector store, refreshed nightly via Airflow, powering an internal LLM-backed search tool for client-service teams.
Xerox Germantown, MD
Data Engineer Feb 2023 – Mar 2024
• Developed telemetry ingestion for Xerox’s managed print fleet, landing meter reads, toner levels, and error codes from 50K+ devices into S3 cataloged in AWS Glue, with CRM data via Fivetran.
• Cleaned and conformed manufacturer-specific telemetry in PySpark on Databricks, deduplicating repeated scans and flagging backward meter reads into trusted Delta tables.
• Modeled billing marts in dbt joining meter reads against contracted per-page rates, establishing the billable-volume source of truth internal finance invoiced clients from.
• Wrote dbt and Great Expectations tests enforcing billing-critical rules such as monotonic meter reads and active-contract checks, resolving a recurring class of invoice disputes flagged by finance.
• Scheduled daily runs as Airflow DAGs with per-site dependencies and retries, preventing late-reporting clients from stalling or corrupting the billing run. HSBC Bangalore, India
Data Engineer Apr 2020 – Jul 2022
• Set up batch ingestion from PostgreSQL and MongoDB into Hive on Hadoop, then built PySpark jobs producing the regulatory and AML reporting datasets due each morning.
• Contributed to Kafka ingestion for near-real-time transaction monitoring and built curated star-schema tables in Redshift for downstream analyst reporting.
• Helped migrate Hive/HDFS workloads to AWS EMR and S3 (Parquet) during the on-prem cluster retirement, rewriting HiveQL into Spark SQL and cutting average batch runtime about 40%.
• Automated source-to-warehouse reconciliation in Python and SQL, catching upstream data drops before they reached regulatory reports. PROJECTS
LLM-Powered ETL for Document Extraction Python, LangChain, Ollama, pgvector, Pydantic, RAGAS, FastAPI
• Built a pipeline ingesting unstructured invoices and extracting vendor, line items, amounts, and dates via LangChain and Ollama into Postgres, with Pydantic schema validation rejecting malformed extractions before load.
• Grounded extraction with a RAG pipeline retrieving context from a pgvector store behind a FastAPI endpoint, and built an evaluation harness scoring faithfulness and field-level accuracy across prompt and chunking changes. Lakehouse Pipeline on Apache Iceberg Spark, Apache Iceberg, S3, Trino, dbt
• Built an incremental ingestion pipeline writing streaming and batch data into Apache Iceberg tables on S3 with Spark, using schema evolution and partition-based upserts to handle late-arriving and changed records.
• Queried the tables through Trino and modeled downstream marts in dbt, using Iceberg snapshots for time-travel validation and comparing compaction strategies to control small-file overhead.
EDUCATION
Rochester Institute of Technology Rochester, NY
Master of Science, Computer Engineering 2025 · GPA: 3.55/4.00 Kumaraguru College of Technology India
Bachelor of Engineering, Electronics and Communication Engineering 2021 · GPA: 8.25/10 1