Fatima Firdouse
# *****************@*****.*** +91-918**-***** ï LinkedIn § GitHub + Hyderabad, India
Summary
AI and Data Science enthusiast with hands-on experience in solving real-world problems using data, building end-to-end
systems involving data pipelines, machine learning, and user-focused applications. Worked with large-scale datasets (500k+
records), built bias detection and matching systems, and deployed real-time applications. Comfortable working in ambiguous
environments, breaking down problems, and focusing on impact over complexity.
Skills
Programming & Analysis Python(OOP), Pandas, NumPy,Seaborn, Matplotlib, EDA, Feature Engineering
Machine Learning Scikit-learn, Supervised & Unsupuervised Algorithms, Model Evaluation
NLP & Information Retrieval TF-IDF, Cosine Similarity, BM25, Semantic Search, LangChain, RAG Pipelines
Data & Vector Stores MySQL, ChromaDB, Hybrid Search(BM25 + Semantic), Cross-Encoder Reranking
Deployment & Cloud AWS EC2, FastAPI, Streamlit, Git & GitHub
Observability & Evaluation Langfuse (LLM Tracing), Ragas, Jupyter, HuggingFace, Groq (Llama 3.1 8B)
Experience
Data Science Job Simulation Jan 2026
Boston Consulting Group (BCG) — Forage
• Designed and automated an ETL workflow on AWS EC2 using Python and Pandas/NumPy to ingest and transform
500k+ customer records, cutting preprocessing time by 40% and ensuring reproducible data pipelines.
• Engineered domain-specific features and trained a Random Forest churn model — isolating the top 10 churn
predictors and enabling a 15% increase in targeted retention actions.
• Synthesised model outputs into a concise executive summary for senior non-technical stakeholders, demonstrating the
ability to translate complex ML findings into actionable business decisions.
Projects
HireIQ — AI-Based Ranking and Bias Analysis Platform Live § GitHub Apr 2026
• Engineered a dual-flow ETL/ML inference pipeline in Python that parses data, scores candidate-job matches (0-100),
and runs bias detection across recruiter and candidaye workflows.
• Built a FastAPI backend on AWS EC2 (5+ endpoints), added Langfuse observability, and delivered Streamlit
dashboards for client.
• Implemented context-aware bias detection using LLM prompting, generating explainable audit reports for each job
description.
DocMind — Production ML Retrieval Pipeline with Observability Live § GitHub Mar 2026
• Architected a multi-stage retrieval pipeline using LangChain and Python, orchestrating data ingestion, transformation,
and a 3-stage ranking system (BM25 keyword + semantic vector search fused via RRF, re-scored with cross-encoder
reranking) — achieving a Ragas Context Precision score of 1.0 (maximum possible) on automated evaluation.
• Deployed the complete system as a publicly accessible live app on AWS EC2, serving real document queries at
sub-2s response time across 5+ FastAPI REST endpoints.
• Integrated Langfuse tracing to capture every retrieval call, generation step, and latency spike on a live dashboard —
reducing debugging time by 40% and demonstrating production-grade observability practices.
Education
B.Tech in Artificial Intelligence and Data Science (Pursuing) Nov 2022 – Jun 2026
Dr. VRK Women’s College of Engineering & Technology, Telangana