Ayanjyoti Thakuria
704-***-**** *.*.*.***@*****.*** ********@****.***
www.linkedin.com/in/ayanthakuria github.com/athakuria02 public.tableau.com/profile/ayanjyoti.thakuria#!/ EDUCATION
The University of North Carolina at Charlotte Charlotte, NC Master’s in Computer Science May 2020
Assam Don Bosco University Assam, India
Bachelor of Technology, Computer Science August 2012 Technical Experience
1. Statistical Analysis Techniques:
Linear Regression Generalized Linear Models (Logistic Regression, Multinomial Regression) Decision Trees Bagging Random Forest Boosting Support Vector Machine Clustering (K-Means, Hierarchical) K-Nearest Neighbor Dimensionality Reduction (Principal Component Analysis) Text Mining Hypothesis Testing A/B Testing Tensor Flow Naive Bayesian Statistics Deep Learning NLP 2. Statistical and Data Management Tools:
Python (Scikit-learn, NumPy, Pandas), MYSQL, Oracle, Hadoop, Pig, Hive, Apache Ambari, PySpark, Kafka, SaS, Cassandra,Hbase, MongoDB, AWS(Lambda, DynamoDB, IAM, S3)
3. Data Visualization Tools:
Tableau, Matplotlib, Seaborn, Plotly, Cufflinks, Ggplot2, Data Studio 4. Programming Languages:
Python, Java, SQL, Linux
Certification: Python Bootcamp, Data Analysis and Visualization Bootcamp in Udemy, Hadoop Professional Experience
Capgemini India Pvt Ltd. – Consultant Bangalore, India (Apr 2013 – Jun 2016) 1. Led a team of 6 members as an offshore team lead, implemented agile methodology, organized scrum teams. 2. Designed and Created reliable end to end data pipelines to move telecom data across multiple data warehouses and real time systems. 3. Hands-on Experience on Bigdata Hadoop.
4. Worked with ETL pipelines for supporting Data extraction, Data Migration, Data Transformation, validated using SQL queries and loading using Informatica Power Center 9.x.
5. Collaborated with business users and subject-matter experts to establish distinct levels of with regards to business usability and performance requirements and to develop an ETL architecture using Agile methodology. Harman Connected Services – Senior Product Engineer Bangalore, India (Jun 2016 – Dec 2018) 1. Managed a team of 8 members as an offshore team lead, implemented an ETL platform and perform data mining techniques such as Classification, clustering, Regression.
2. Designed and implemented ETL and Data pipelines using Informatica 9.0 using Agile Methodology. 3. Experienced in Quality Assurance including Manual and Automation ETL testing Process for Database, Informatica, and Shell scripting. 4. Worked on meta data, data management, data creation, manipulation, integration and transformation of unstructured data. 5. Developed performance tuning scripts in UNIX and SQL queries improving execution efficiency by 30%. 6. Implemented Machine Learning algorithms like K-means, Classification, regression, PCA for trend predictions on Insurance claim dataset. 7. Validated all the backend functionalities using SQL queries, joins and advanced SQL. Execute data load (Full/delta) and fix errors if any. Academic Project
Predicted Attendance at a Basketball game using Python (Sklearn) 1. Implemented Polynomial Linear regression on features to analyze the probability of attendance at a basketball game. 2. Implemented cross validation to separate train and test data. 3. Forecasted by implementing Ridge and Lasso Techniques and compared the results with linear regression results. 4. Ridge and Lasso predicted better results which improved accuracy by 10-12% Loss Ratio prediction for Auto Insurance Using Python (Sklearn) 1. Implemented Random Forest and KNN regression to calculate loss ratio for 330 portfolios containing more than 400K records. 2. Creates own feature engineering technique to reduce and create new features. 3. This method of feature engineering provided an accuracy of 95.7%. 4. Our model came 2nd in the Kaggle competition created for this project. New York Cycle Usage Data Analysis using Talend, S3 and Snowflake 1. Created ETL pipelines in Talend to populate the files in S3. 2. Built an auto ingestion data pipeline using Snowflake to read files from S3, transformed and aggregated data into tables in Snow tables. 3. Created end to end data pipelines to push the data to data warehouses in snowflake for analysis using snowSQL. Sales Revenue analysis using Python
1. Designed a dashboard to perform a high-speed country-specific sale-revenue analysis using Python. 2. Investigated the role of interactive visualization in model-driven decision making 3. Evaluated the sales revenue at each stage of the sales cycle using Matplotlib, Seaborn, Plotly and Cufflinks. Used Tableau to create an Interactive Dashboard for Searching Datasets 1. Implemented an interactive dashboard to create search datasets from around 37K Json files. 2. Cleaned and manipulated data using python to create a master dataset. 3. Added various actions on clicking to perform various operations on the dashboard. 4. The Dashboard could be used like a search engine to search and go the location using the URL to get the exact dataset. Created a Database Project for an online application using MySQL 1. Created Joins, Stored procedures, Stored functions and triggers to create a Database system using MYSQL In AWS RDS instance. 2. Implemented Agile methodology to develop the application. 3. The database system showed great performance, usability and data integrity.