Le Ba Vinh
Data Engineer
**********@*****.*** **69.839.257 github.com/aqua-277353 linkedin.com/in/vinh-le-628a69323 Summary
Results-driven Data Engineer with hands-on expertise in SQL, Python, and Microsoft Fabric. Specialized in building, maintaining, and monitoring scalable ELT pipelines and data quality frameworks. Dedicated to transforming raw data into reliable, structured assets within the Medallion architecture to empower business intelligence and reporting operations.
Education
University of Information Technology (UIT)
BSc in Information Systems
• GPA: 3.2/4.0
• Relevant coursework: Databases & Data Warehousing, Big Data, Machine Learning, Business Intelligence. Work Experience
NaviWorld
Data Engineer Intern
• Engineered an end-to-end, automated HR data pipeline on Microsoft Fabric (Medallion architecture), utilizing PySpark and SQL for scalable data ingestion and transformation.
• Integrated automated quality validation and modular transformation logic, ensuring the system is highly maintainable and easily adaptable to evolving HR business requirements.
• Impact: Designed and deployed Power BI dashboards atop the Gold layer, replacing manual HR workflows with a self-sustaining ecosystem for immediate operational insights. Skills
• Programming: Python, SQL (T-SQL, PostgreSQL), PySpark.
• Data Architecture: ELT Pipelines, Medallion Architecture, Data Modeling, Data Quality Validation.
• Cloud & Platforms: Microsoft Fabric, Azure (Databricks, ADLS Gen2).
• Tools & Frameworks: Apache Airflow, Docker, Git, Power BI. Certifications
• English: TOEIC Listening & Reading (685), TOEIC Speaking & Writing (240). Projects
Cloud-based Retail Purchase Behavior Analysis
• Architected a cloud-native analytics platform on Azure to process and analyze a large-scale retail dataset.
• Executed feature engineering and data transformation tasks using PySpark and Databricks to prepare data for downstream machine learning applications.
• Impact: Processed 541K+ records to deliver product correlation insights and optimized transformation logic, reducing data preparation time by 75%.
YouTube Metadata ELT Pipeline
• Developed a containerized ELT pipeline extracting video performance trends and metadata from the YouTube API.
• Designed the analytical schema in PostgreSQL and orchestrated the workflow using Apache Airflow and Docker.
• Impact: Established automated alerting and data quality checks, ensuring high reliability for the reporting layer.