NIKHILA BOPPANA
Jersey City, NJ +1-551-***-**** **************@*****.*** LinkedIn
PROFESSIONAL SUMMARY
• Data Engineer with 4+ years of experience architecting and delivering scalable, production-grade data pipelines and cloud data platforms from the ground up across financial services, healthcare, and enterprise domains.
• Modern data stack expertise: deep hands-on skills in Python, advanced SQL, Apache Spark, Snowflake, Databricks, dbt, and Apache Airflow on AWS, Azure, and GCP.
• End-to-end delivery: builds ETL/ELT pipelines, dimensional data models, and medallion lakehouse architectures with a strong focus on performance tuning, cost optimization, data quality, and CI/CD automation.
• Versatile contributor: combines strategic ownership of end-to-end data architecture with rapid, hands-on execution suited to fast-paced contract and enterprise environments.
• AI-ready platforms: integrates machine learning and GenAI (RAG, vector search, embeddings) workflows into production data platforms to accelerate insights and delivery. TECHNICAL SKILLS
Programming & Query Languages: Python, Advanced SQL (T-SQL, PL/SQL, window and analytical functions, query optimization), Java, Scala, Shell scripting (Linux/Unix) Cloud Platforms: AWS (S3, Glue, Redshift, Athena, IAM), Azure (Data Factory, Synapse, Microsoft Fabric - pipelines, notebooks, Lakehouse/Warehouse), Google Cloud Platform / BigQuery (working knowledge) Data Warehousing & Lakehouse: Snowflake, Databricks, Delta Lake, medallion architecture (bronze/silver/gold), data lake, data warehouse, data marts, Amazon Redshift; performance tuning and cost/resource optimization ETL/ELT & Orchestration: dbt, Apache Airflow, Azure Data Factory, AWS Glue, ETL/ELT pipelines, Change Data Capture
(CDC), batch and near real-time processing
Streaming & Integration: Apache Kafka (event-driven streaming), REST APIs, semi-structured data (JSON, Parquet), data integration
Data Modeling: Dimensional modeling (star schema, fact and dimension tables, SCD Type 2), 3NF normalization, data marts, MDM
Governance, Quality & Security: Data quality validation, data governance, Unity Catalog, metadata management, data lineage, security and access controls (IAM), HIPAA compliance AI/ML & GenAI: RAG, embeddings, vector search (Pinecone, FAISS), model-ready datasets, LLM-based data pipelines DevOps & Delivery: Git, GitHub, Azure DevOps, Jenkins, GitHub Actions, CI/CD, Docker, unit and integration testing, Agile
(Scrum), JIRA
Business Intelligence: Power BI, Tableau
PROFESSIONAL EXPERIENCE
Data Engineer, Techzert, New York City, NY May 2025 - Present Project: AI-Powered Financial Intelligence Assistant
• Improved data availability by 40% for large-scale financial datasets by architecting cloud-native ingestion and ELT pipelines with Python, Spark (Databricks), and dbt into Snowflake.
• Reduced document-retrieval latency by 30% by building Python-based RAG and machine learning workflows with vector search, embeddings, and REST APIs.
• Increased data trust and reliability by implementing automated data quality checks (null validation, uniqueness, referential integrity) and dbt source-freshness tests across analytics-ready datasets.
• Enabled scalable financial reporting by designing dbt staging and mart models (star schema) and tuning Snowflake and PostgreSQL query performance.
• Cut manual intervention and improved pipeline uptime by orchestrating and monitoring production workflows in Apache Airflow and Azure Data Factory with proactive troubleshooting.
• Accelerated delivery by shipping production changes through Git, Azure DevOps, and CI/CD with unit and integration testing in an Agile (Scrum) team.
Graduate Research Assistant, Saint Peter's University, Jersey City, NJ Feb 2025 - May 2025 Project: Financial Document Intelligence using RAG
• Enabled semantic search over large financial document collections (earnings reports, SEC filings) by designing RAG pipelines with Python, FAISS, and Pinecone.
• Reduced query latency by ~20% while maintaining retrieval precision by evaluating and optimizing vector indexing methods and retrieval configurations.
• Improved retrieval accuracy across the document corpus by building preprocessing workflows with chunking strategies and embedding generation.
Software Engineer, Data Engineering, Deloitte, Hyderabad, India Jun 2022 - Aug 2024 Project: Statewide Data Modernization and Reporting Platform, State of Connecticut
• Improved data processing efficiency by ~30% on multi-billion-row datasets by tuning Spark performance (partitioning, caching, query optimization).
• Enabled consistent cross-agency reporting by building Databricks (PySpark) pipelines in a medallion (bronze/silver/gold) architecture with Delta Lake.
• Reduced reporting query latency and warehouse compute cost by optimizing Snowflake performance (clustering, partitioning, warehouse tuning) and designing fact and dimension models with materialized views using advanced SQL.
• Ensured governed, compliant data by applying Unity Catalog access controls, metadata, and lineage across multi- agency systems.
• Delivered scalable analytics with end-to-end ownership in an Agile (Scrum) team, providing technical guidance to junior engineers.
Project: Enterprise Sales and Supply Chain Data Integration, Newell Brands
• Improved data availability for analytics by ~35% by developing and migrating ETL pipelines integrating ERP (supply chain, inventory, sales) data to scalable, cost-efficient AWS storage and processing (S3, Glue, Redshift).
• Eliminated discrepancies across inventory and sales reporting by building Kafka-based event-driven ingestion and dbt transformations.
• Reduced reporting latency by ~25% by optimizing SQL transformations feeding Power BI and Tableau dashboards, and unified cross-domain reporting by integrating CRM and operational datasets. Data Engineering Associate, Dizi Shades, Hyderabad, India Jun 2021 - May 2022 Project: Healthcare Data Integration and Analytics Platform
• Enabled reliable production data delivery by building and automating ETL pipelines (Python, SQL) from FHIR REST APIs and flat files into AWS S3 and Amazon Redshift with Apache Airflow.
• Reduced redundant processing versus full refreshes by implementing incremental loading with timestamp-based watermark logic (CDC pattern).
• Improved query efficiency for analytical workloads by designing fact and dimension models for patient encounter and diagnosis (ICD-10) data.
• Ensured healthcare compliance by performing data quality validation (null handling, schema validation, deduplication) and handling sensitive patient data per HIPAA.
EDUCATION
Saint Peter's University, NJ, USA, Master's in Data Science (GPA: 3.93) Koneru Lakshmaiah Education Foundation, India, B.Tech. in Computer Science and Engineering CERTIFICATIONS AND ACHIEVEMENTS
• Confluent Certified Data Streaming Engineer Foundations, Confluent (2026)
• Salesforce Certified AI Associate, Salesforce
• Aviatrix Certified Engineer (ACE), Multi-Cloud Networking, Aviatrix
• Outstanding Performance Award, Deloitte, for scalable data flows and APIs enabling the OEC360 platform go-live, improving backend performance and cross-system data reliability.
• Deloitte Spot Award, for data integration and backend optimization improving system reliability and data availability for downstream analytics.