Post Job Free
Sign in

Metadata-Driven Data Engineer for Pipelines

Location:
Manhattan, NY, 10007
Posted:
August 24, 2026

Contact this candidate

Resume:

Viswack Reddy Yelugoti

Denton, TX 940-***-**** *************@*****.*** LinkedIn GitHub

PROFESSIONAL SUMMARY

Data Engineer with hands-on experience building production data pipelines on Microsoft Fabric, Azure, AWS, and Databricks. Currently a Data Engineering Intern at Simmons Bank, implementing metadata-driven ingestion pipelines and SCD Type 2 transformation logic using PySpark within a regulated banking environment. Skilled in ETL/ELT development, data quality and monitoring practices, and translating business requirements into scalable technical solutions. Databricks Certified Data Engineer Associate (in progress); AWS Certified Data Engineer – Associate.

EDUCATION

University of North Texas Jan 2025 - Nov 2026

Master of Science, Computer Science - Cybersecurity Specialization Amrita School of Engineering Aug 2020 - Jun 2024

Bachelor of Technology, Electronics and Computer Engineering EXPERIENCE

Simmons Bank Jun 2026 - Aug 2026

Data Engineering Intern Dallas, TX

• Built a reusable, config-driven ingestion framework in Microsoft Fabric implementing a Bronze Silver Gold medallion architecture for banking-style source files.

• Designed a single metadata-driven ingestion notebook supporting five load strategies (SCD Type 2, Daily Snapshot, Append, Daily Append, Delta), selected dynamically via pipeline parameters with no code changes.

• Implemented SCD Type 2 transformation logic to preserve historical record changes, enabling accurate point-in-time reporting for downstream analytics.

• Built schema enforcement, primary key validation, and file-level audit tracking, logging run-level metrics (row counts, status, errors) to a LoadAudit table and file-level status to a control table with watermark tracking.

• Developed and tested a governance chatbot (DG Assistant) using Azure AI Search to improve data discoverability and self-service access to the Governance Hub catalog.

• Wrote PySpark transformation logic within Fabric notebooks to clean, validate, and reshape source data across all five load strategies before writing to Silver and Gold layers.

• Partnered with the data engineering team to translate business and compliance requirements into scalable technical solutions. University of North Texas Jun 2025 - May 2026

Student Assistant, Data Analyst Denton, TX

• Supported alumni fundraising analytics using SQL and Microsoft Excel, querying and consolidating donor and giving records for reporting.

• Cleaned and standardized alumni and donation data by resolving duplicate records, correcting formatting inconsistencies, and validating missing or invalid fields prior to reporting.

• Built Excel-based summaries and pivot reports to track fundraising trends and support decision-making for the advancement team.

• Wrote SQL queries to extract, clean, and validate alumni and donation data from university databases for accuracy and consistency. PROJECTS

Batch Processing ETL Pipeline with Apache Airflow Jul 2025 - Aug 2025

• Developed a microservices-based containerized streaming data pipeline using Docker multi-container architecture with Apache Airflow for workflow orchestration, job scheduling, and observability monitoring.

• Designed production-grade DAGs with SLA monitoring, dynamic task generation, XCom for inter-task communication, Jinja-templated retry logic, and failure handling in event-driven streaming environments.

• Leveraged PySpark for distributed, in-memory processing of real-time brewery data streams with horizontal scalability and optimized partition strategies.

• Integrated AWS S3 as a scalable data sink with automated Parquet conversion, time-based partitioning, lifecycle policies, and data versioning for lineage tracking.

Real-Time Financial Data Streaming and Analysis Using Kafka Apr 2025 - May 2025

• Engineered a cloud-native, event-driven ETL pipeline fetching company profiles from the Finnhub REST API, processing 1000+ records with CI/CD automation and exponential backoff for rate-limit management.

• Architected an AWS S3 data lake with IAM role-based access control and serverless Lambda triggers, resolving access permission challenges via Boto3 SDK integration.

• Implemented distributed data cleaning using PySpark DataFrames with Delta Lake ACID transactions to handle schema evolution, null values, and data quality validation.

• Configured a Databricks Lakehouse SQL Warehouse with Unity Catalog for governance, enabling self-service BI through query optimization and incremental loading.

Real Estate Analytics ETL Data Pipeline Aug 2025 - Sep 2025

• Built a scalable, cloud-native ELT pipeline with infrastructure as code, processing thousands of Zillow property records from REST API extraction through PySpark transformation to interactive dashboards.

• Utilized AWS Glue serverless ETL jobs with Apache Spark and the AWS Glue Data Catalog for metadata management, implementing data cleaning, window-function business logic, and deduplication.

• Integrated Snowflake with zero-copy cloning and time travel for optimized columnar storage, implementing star-schema dimensional modeling, materialized views, and query caching.

• Developed Power BI dashboards with DAX measures and Power Query M, delivering real-time market trend analysis, investment scoring, and geospatial visualization.

SKILLS

• Programming Languages: Python, SQL, PySpark

• Data Engineering: Microsoft Fabric, Apache Airflow, Apache Spark, Azure Data Factory, AWS Glue, Docker, Databricks, DBT, REST APIs

• Cloud Platforms: AWS (S3, Lambda, IAM), Microsoft Azure (Data Factory, Synapse, Blob Storage, AI Search)

• Databases: SQL, NoSQL

• Data Visualization: Power BI, MS Excel

• Version Control & DevOps: Git, GitHub, CI/CD, IT Automation

• Methodologies: Agile Methodology, Software Development Life Cycle

• Analytics & Governance: Data Analytics, Data Governance, Data Quality Monitoring CERTIFICATIONS

• Databricks Certified Data Engineer Associate: Databricks (In Progress)

• AWS Certified Data Engineer – Associate: AWS

• Microsoft Azure Data Fundamentals (AZ-900): Microsoft

• IBM Spark Fundamentals I & II: IBM Developer Skills Network

• Introduction to SQL: DataCamp

• Astronomer Apache Airflow 3 Fundamentals Certification: Astronomer

• Data Visualisation: Empowering Business with Effective Insights: TATA



Contact this candidate