Aman Kumar Sahu
github.com/amansahu*** ****.*******@*****.*** 240-***-**** linkedin.com/in/aman205 aman-sahu.tech PROFESSIONAL SUMMARY
Data Engineer with 2+ years building real-time financial data infrastructure at Fortune 500 scale. MS in Data Science and 4x hackathon winner, specializing in agent-ready systems combining SQL, knowledge graphs, vector pipelines, and streaming infrastructure.
EDUCATION
University of Maryland, College Park
MS in Data Science. GPA: 3.73.
SRM Institute of Science and Technology
B.Tech in Mechatronics. GPA: 3.6/4.0.
EXPERIENCE
Founding Data Engineer May 2025 – Present
Connyct New York, NY
• Built a RAG pipeline scraping 10,000+ university records via Playwright into S3, generating Sentence Transformers embeddings indexed into Elasticsearch to power LLM agent retrieval
• Orchestrated ETL workflows with Dagster and AWS Step Functions, stabilizing data availability from under 50% to 95% across 50+ volatile API endpoints and storing half a million records
• Automated pipeline deployments via GitHub Actions CI/CD, enabling consistent releases with zero manual intervention Data Engineer Aug 2022 – Sep 2024
Tata Consultancy Services (Fortune 500) -- PNC Bank Hyderabad, India
• Built ETL pipelines ingesting data from SQL Server, Oracle, and MongoDB into Hadoop, transforming 10TB daily with Spark SQL to serve 30 downstream business teams
• Rewrote 150 Spark SQL transforms using partition pruning and broadcast joins, cutting nightly batch runtime from 8 hours to 3 and saving $50K annually in compute costs
• Built a Kafka and PySpark Streaming alert system processing 1M daily credit card transactions under 5-minute latency, preventing an estimated $2M in annual over-limit losses
• Automated 50 ETL workflows with CA7 scheduling, reducing manual intervention by 70% PROJECTS
AlphaQuery -- Financial RAG System
https://github.com/amansahu205/alphaquery-financial-rag Dagster, Pinecone, MongoDB, Docker, LangChain, FastAPI
• Built an ETL pipeline ingesting 9,806 SEC filings across S&P 500 companies into 540K+ Pinecone vectors, powering a 4-agent RAG system with section-aware chunking and temporal filtering Inflect -- AI Financial Research Platform https://github.com/amansahu205/inflect-ai Airflow, Snowflake, Kafka, Neo4j, FinBERT, TA-Lib
• Built an Airflow-orchestrated pipeline pulling SEC filings, Kafka market data, and news feeds into Snowflake, layering FinBERT sentiment and a Neo4j knowledge graph to deliver citation-backed stock analysis in under 2 seconds Credit Risk Scoring with Fairness Testing
https://github.com/amansahu205/credit-risk-platform LightGBM, SHAP, Fairlearn, PostgreSQL, scikit-learn
• Built a credit risk scoring pipeline on 300K+ loan applications using LightGBM, achieving 88% AUC with SHAP adverse action explanations and correcting a 12% approval rate disparity to meet ECOA compliance thresholds CERTIFICATIONS
• Microsoft Azure Fundamentals (AZ-900)
• AWS Solutions Architect (In Progress)
SKILLS
Data Engineering: Python (Pandas, NumPy, Pydantic), SQL, Spark SQL, PySpark, Kafka, Airflow, Dagster, dbt, Hive, Hadoop Cloud & Infra: AWS (S3, Lambda, SageMaker, Step Functions, ECR), GCP (BigQuery, Cloud Run), Docker, GitHub Actions ML & AI: XGBoost, LightGBM, SHAP, Fairlearn, TensorFlow, PyTorch, LangChain, LangGraph, Pinecone, FinBERT, Neo4j Databases: PostgreSQL, MongoDB, Redis, DynamoDB, Elasticsearch, Snowflake