Han Xu
Sunnyvale, CA 323-***-**** *********@*****.***
github.com/hanxudata linkedin.com/in/han-xu-data SKILLS
Programming Languages:Python, SQL, R, MATLAB, HTML, JavaScript Frameworks:Apache Spark, Apache Hadoop, Apache Airflow Tools:AWS, MySQL, PostgreSQL, Cassandra, Git, Tableau, Power BI Packages:Pandas, Numpy, SciPy, Scikit, Plotly, Flask, NLTK Other Skills:Distributed Computing, Machine Learning, A/B Testing, Exploratory Data Analysis EXPERIENCE
Data Engineering Fellow,Insight, San Francisco May 2020 - Present
● Built a system to detect fake reviews and devalue low-quality products for Amazon.
● Developed a batch processing pipeline with Spark, stored the result in MySQL and automated the whole process with Airflow scheduling.
● Developed a web application for users to check the adjusted ratings of Amazon products with Flask. Data Analyst Intern,Remoov, San Francisco May 2019 - Aug 2019
● Worked with the CEO to perform data analysis for operations, sales, marketing and general performance of the company on its 3-years data via Power BI, built 15 real-time dashboards.
● Managed the PostgreSQL database on AWS.
● Improved the operation efficiency, and increased the sales by 30% by optimizing the sales channel strategy.
Associate Consultant,A.T. Kearney, Shanghai Sep 2016 - Nov 2016
● Developed investing strategies for one of the biggest phosphate corporations of China.
● Performed research on three different industries, analyzed the financial situation, developing prospect and difficulty of acquisition of more than one hundred companies.
● Attained the knowledge of consulting and the capacity of researching different industries. PROJECTS
Movie Recommendation https://bit.ly/movie-project Jul 2019 - Aug 2019
● Predicted the Netflix movie ratings in a MovieLens dataset via Python and Spark.
● Built data processing pipeline with Spark RDD and Spark SQL for big data OLAP.
● Applied Alternating Least Squares algorithm to find the optimal hyperparameters. User Churn Prediction https://bit.ly/user-churn Mar 2019 - Apr 2019
● Developed algorithms to predict user churn probability based on labeled data via Python.
● Utilized data cleaning, categorical feature transformation and standardization to preprocess dataset.
● Trained 3 supervised machine learning models, applied regularization to overcome overfitting.
● Analyzed feature significance to get top 10 factors that influenced the results. EDUCATION
New York University Sep 2018 - May 2020
M.S.in Industrial Engineering, Data Track GPA 3.7/4.0 University of Southern California Aug 2013 - May 2018 B.S.in Electrical Engineering