Quoc Tran Nguyen
Senior Data Engineer **********@*****.*** 714-***-**** Santa Ana, CA Linkedin Summary
My name is Tran, data engineer with 8 years of experience specializing in high-scale data processing and the development of robust ETL pipelines using SQL and Python. Beyond core data engineering, I am passionate about the practical application of Machine Learning, Deep Learning, and AI. I thrive on tackling complex, high-impact problems using state-of-the-art technologies to build intelligent, data-driven solutions. I can communicate fluently in English during daily stand-ups, code reviews, and technical presentations with distributed global teams. Technical Skills
• Cloud & Big Data: Azure (Synapse, Databricks, ADF, ADLS, SQL DB), Apache Spark (PySpark)
• Programming Languages: Python (Pandas, Numpy, Flask, FastAPI), SQL (T-SQL, PL/SQL), Java
• Databases: MS SQL Server, PostgreSQL, Oracle Apex/Database
• Data Engineering: Dimensional Modeling (Star/Snowflake), ETL/ELT, CI/CD (Azure DevOps, Git)
• Communication: Fluent English communication, Cross-functional Collaboration, Technical Documentation. Education
Ho Chi Minh City University of Technology (HCMUT), Bachelor of Engineering in Computer Science
2015 – 2018
Experience
Senior Data Engineer – FPT Software, Vietnam 2022 – Present
• Logistics and Supply Chain Simulation (Schaeffler Germany):
– Served as a Senior Data Engineer to design and implement large-scale data solutions for global manufacturing plants, ensuring compliance with rigorous international automotive security and quality standards.
– Medallion Architecture: Architected an automated pipeline to ingest SAP Master Data into a multi-layered ADLS Gen2 Data Lake, streamlining data from raw to gold layers using PySpark and Delta Lake.
– SAP BOM Transformation Engine: Designed recursive PySpark transformation logic on Azure Databricks to flatten multi-level SAP Bill of Materials (BOM) hierarchies, converting complex parent-child relationships into analysis-ready datasets for supply chain forecasting.
– Large-scale Manufacturing ETL: Built high-performance ETL/ELT pipelines using Azure Data Factory
(ADF) and Azure Synapse to consolidate terabytes of SAP Master Data (MARA, MARC, MBEW, etc.) and complex BOM structures.
• FedEx Canada Data Migration (Oracle to Azure Synapse):
– Architected and executed a large-scale data migration strategy to transition legacy Data Marts from Oracle and Enterprise Data Warehouse (EDW) to Azure Synapse Analytics.
– Developed complex Azure Data Factory (ADF) pipelines to automate the extraction, loading, and transformation of high-volume historical and incremental data, ensuring 100% data integrity.
– Optimized T-SQL scripts and Synapse distribution strategies to handle massive data loads, significantly improving query performance and reporting efficiency for the Canadian logistics market.
– Streamlined data orchestration using ADF, reducing manual migration effort and ensuring a seamless transition from on-premises to the cloud environment. Data Engineer – Bosch Global Software Technologies, Vietnam 2018 – 2022
• Tax Data Ingestion & Banking API Service:
– Leveraged PySpark to architect a robust data ingestion engine, automating the migration and loading of complex, high-volume tax datasets into Oracle Database.
– Developed and deployed a high-availability FastAPI service to provide secure, real-time tax data access for banking enterprise applications, ensuring low-latency performance and data security.
– Optimized data retrieval logic and API endpoints to handle concurrent requests from financial institutions, replacing manual reporting with a scalable automated solution.
• Internal HR AI Chatbot:
– Architected a high-performance API using FastAPI and Transformer models to automate HR policy inquiries, reducing administrative workload by 40%.
– Engineered a robust NLP pipeline utilizing Pandas and SpaCy for advanced text cleaning, lemmatization, and Bag of Words (BoW) feature extraction.
– Designed a scalable MongoDB document store for conversational logs and semi-structured knowledge bases, ensuring rapid data retrieval and 24/7 access.
• Automotive Web Applications (Oracle Application Express):
– Developed enterprise-grade web applications utilizing the Oracle Application Express (APEX) ecosystem.
– Engineered a high-performance Data Model using Star Schema on Oracle Database to consolidate and structure complex automotive datasets for optimized reporting.
• Automotive Telematics Reporting Engine:
– Architected and implemented a high-performance reporting engine using Java Spring Boot and Flying Saucer (XHTMLRenderer) to generate complex, multi-page automotive PDF reports (e.g., vehicle health logs, telematics summaries).
– Engineered complex Oracle SQL schemas and optimized PL/SQL procedures to aggregate and transform millions of telemetry records, ensuring sub-second data retrieval for real-time document generation.