Post Job Free
Sign in

Data Engineering & AI-Powered Data Platforms

Location:
Paterson, NJ
Salary:
120000
Posted:
August 30, 2026

Contact this candidate

Resume:

Ankita Singh

+1-872-***-**** ******.*****.******@*****.*** H1B (2025-2027) linkedin/ankitansingh/ GitHub ankita.github.io Summary

Data Engineer with 7+ years of experience building scalable ETL/ELT pipelines, data models, and quality frameworks using Python, SQL, AWS Glue, S3, Spark, Airflow, dbt, and Snowflake. Experienced in real-time streaming, API and data-serving workflows, schema evolution, and production reliability, with hands-on experience applying machine learning and generative AI to enterprise data platforms. Skills

• Languages: Python, SQL, Bash/Unix Shell, Scala, NodeJS

• Data Platform Engineering: ETL/ELT, Data Modeling, Canonical Datasets, Distributed Data Processing, Data Lineage, Data Quality, Schema Evolution, Query Optimization, Fault-Tolerant Pipelines, Batch Processing, Event-Driven Data Pipelines, Data Warehousing, Data Architecture, Amazon Redshift, Amazon Kinesis, Amazon Kinesis Data Firehose

• APIs & Data Stack: REST APIs, gRPC, FastAPI, Apache Spark, Apache Airflow, dbt, Kafka, Snowflake, AWS Glue, Hadoop, Hive

• AI/ML: Machine Learning, LangChain, GPT-4, AWS Bedrock, Vertex AI, SageMaker, Vector Databases, Embeddings, RAG

• IoT/Real-Time: AWS IoT Core, AWS TwinMaker (Digital Twin), Grafana, CloudWatch, Sensor Telemetry Integration

• Cloud & Infra: AWS (Glue, S3, EC2, EMR, ECR, Lambda, Lake Formation, Bedrock, DynamoDB, IoT), GCP (BigQuery, Vertex AI), Azure, Terraform, Docker, Kubernetes, CI/CD, AWS IAM, Non-Relational Databases

• Analytics & BI: pandas, NumPy, SciPy, Jupyter, matplotlib, Tableau, Power

• Software Engineering: BISDLC, Object-Oriented Programming, Agile, Git, Bitbucket, CI/CD, Microservices, Backend Development Experience

GanaIT (Coaction Specialty Insurance) Feb 2026 - Present Lead Data Engineer

• Owned the design and technical delivery of enterprise Policy ODS enhancements by translating business requirements into scalable data integration solutions through impact analysis, XML/XPath mapping, and end-to-end data lineage across Policy ODS and Core ODS, improving data consistency for downstream reporting systems.

• Collaborated with Product Owners, Business Analysts, architects, and cross-functional engineering teams to define scalable data architecture, prioritize platform enhancements, resolve integration challenges, and deliver trusted underwriting data solutions supporting enterprise applications.

• Own production reliability for AWS Glue-based ETL workflows, troubleshooting pipeline failures through CloudWatch logs, SQL debugging, stored procedures, and cross-layer dependency analysis while strengthening error handling and ensuring reliable data delivery across Policy ODS pipelines.

• Lead a 10-member distributed engineering team across onshore and offshore locations, driving Agile delivery, sprint planning, production releases, code quality, and technical execution for mission-critical enterprise data platform initiatives.

• Reduce production deployment errors by 30–35% by implementing Redgate Flyway-based database version control and automated schema migration practices, improving release reliability and deployment consistency across monthly production releases.

• Designed and maintained canonical enterprise datasets supporting underwriting operations, reporting, and downstream analytics by standardizing XML-to-relational data models, transformation logic, and source-to-target mappings across Policy ODS and Core ODS. Wissen Technologies(Morgan Stanley) Dec 2024 - Jan 2026 Data Engineer

• Developed end-to-end dbt transformation pipelines on Snowflake, transforming custodian data (State Street, BlackRock Aladdin) into IBOR, PBOR, Trades, and Performance datasets supporting investment risk and portfolio analytics.

• Cut job runtime to ~2-3 minutes and eliminated backtracking/debugging overhead by migrating legacy DB2/Informatica ETL workloads to dbt on Snowflake, orchestrated with Apache Airflow for scheduling and dependency management.

• Designed and improved modular Bronze/Silver/Gold dbt architectures using reusable macros, snapshots, incremental/upsert custom materializations, SCD Type 1 and Type 2 implementations, and flexible schema evolution patterns, reducing model redundancy by 50%.

• Integrated ML-driven anomaly detection into Snowflake pipelines using Python and dbt tests, reducing undetected data quality issues by ~35%.

• Performed data validation and reconciliation using Snowflake SQL, dbt tests, MD5 hashing, and NULL validation, ensuring accuracy across 200M+ of financial trading records.

• Improved engineering quality across an 8-member Agile team by reviewing 20+ pull requests per sprint, implementing CI/CD and software development lifecycle best practices, and mentoring developers to ensure maintainable, production-ready data pipelines.

• Collaborated with data consumers, portfolio managers, and engineering teams to evolve shared data models, resolve production issues, and enable scalable self-service analytics across investment platforms. National Renewable Energy Laboratory Jun 2022 - Jun 2024 Software Data Engineer Denver, CO, USA

• Designed and developed an enterprise Retrieval-Augmented Generation (RAG) platform by extending LangChain's RecursiveURLLoader to automate document ingestion, recursive web crawling, semantic chunking, embedding generation, vector indexing with ChromaDB, and grounded retrieval using AWS Bedrock (Llama), enabling AI-powered knowledge discovery for 2,000+ internal users.

• Designed reusable Python modules and automation utilities to support document ingestion, vector indexing, and cloud-native AI workflows, improving maintainability and reducing manual operational effort

• Engineered real-time streaming data pipelines by ingesting telemetry from 40–50 industrial sensors through Apache Kafka into an AWS S3 landing zone powering AWS TwinMaker, Grafana, and CloudWatch, enabling event-driven monitoring and visualization of hydrogen plant operations.

• Designed and implemented a Python-based AWS Glue ETL pipeline to ingest and transform 100,000+ hydrogen plant sensor records into an AWS Lake Formation data lake, reducing manual data preparation by approximately 60%

• Developed secure REST APIs and cloud services using Django, AWS Lambda, API Gateway, DynamoDB, and AWS IAM roles and permissions, enabling authenticated access to enterprise cloud resources through an SSO-enabled portal.

• Automated AWS S3 infrastructure provisioning using Terraform, standardizing infrastructure deployment and reducing environment setup time by approximately 30%.

Amdocs Jul 2018 - May 2021

Software Development Engineer

• Ensured data integrity for a go-live by migrating 1M+ customer, billing, and subscription records from Huawei legacy systems to PLDT across Oracle and MySQL with SQL-based reconciliation.

• Performed root cause analysis and resolved production data discrepancies across 30-40 relational database tables prior to go-live, improving deployment quality and reducing production risk.

• Reduced manual validation effort by ~30% by automating 10+ UNIX batch jobs with Shell scripting and SQL.

• Collaborated with Business Analysts, DBAs, and client stakeholders during Agile releases and production migration war rooms, supporting successful enterprise billing deployments Education

Illinois Institute of Technology Aug 2021 - May 2023 Master of Science, Computer Science Chicago, USA

Savitribai Phule Pune University Aug 2014 - May 2018 Bachelor of Engineering, Electronics and Telecommunications Pune, India PROJECTS

LinkedIn Data Mart Project

• MySQL, PowerBI, Pentaho,

• Developed an ETL process using Pentaho to parse LinkedIn JSON files and load structured data into MySQL tables.

• Created Power BI and Tableau visualizations with search capabilities to support data exploration and insights. Big Data Customer Segmentation Analysis source

• Python, NoSQL, Data Mining, Predictive Modeling

• Generated customer segmentation insights using Neo4j graph analytics, NoSQL data, and RFM modeling to support personalized marketing recommendations.

ChitGenie source

• GenAI, Multilingual model, NLP, Vertex AI, Google Wallet, GPT-4, Big Query, Machine learning

• Ideated and led the development of ChitGenie, an AI-powered WhatsApp-integrated receipt assistant with multimodal parsing, multilingual NLP, and digital wallet integration.

• Built scalable systems enabling smart receipt tracking, budget insights, and conversational commerce using GPT-4, Vertex AI, Firebase, BigQuery, and Google Wallet APIs.

Publications

• Data Transmission Through Visible Light (LiFi).in the International Journal of Innovations in Engineering Research and Technology, Volume 5, Issue 5, ISSN 2394-3696. LI-FI Technology



Contact this candidate