MAHESH MANJU
Email-id: **************@*****.***, *******@*****.**.**
Mobile: 779-***-****, 897*******
PROFESSIONAL SUMMARY:
Having 2.5 years of extensive experience as Bigdata Engineer in Snipe IT Solutions.
Design and creating real-time data streaming solutions using Apache Spark core, spark SQL & Data Frames, Spark Streaming and Hive.
Experience in building Data pipelines using Big Data Technologies.
Hands-on experience in writing MapReduce programs and user-defined functions for Hive.
Experience in NoSQL technologies like Mongodb.
Excellent understanding knowledge on Hadoop and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, Resource Manager(YARN).
Experience in importing and exporting data using Sqoop from HDFS to Relational Database System(RDMS) and from RDMS to HDFS.
Proficient at using Spark APIs to Cleanse, Explore, Aggregate, Transform, and store machine sensor data.
Hands-on experience with systems-building languages such as Scala and java.
Has worked on ETL tools like Pentaho Kettle.
Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
Experience working on processing unstructured data using Hive.
Responsible for creating Hive tables based on business requirements and Implemented Partitioning, Dynamic Partitions and Buckets in HIVE for efficient
access.
TECHNICAL HIGHLIGHTS:
Big Data Ecosystems: Hadoop, MapReduce, HDFS, Spark, Spark SQL, Hive, Sqoop, Flume, Oozie.
Programming Languages: Java, Scala.
Databases: NoSQL(Mongodb), SQL, Postgresql.
ETL Tools: Pentaho Kettle.
Tools: Eclipse, R, Rstudio.
Platforms: Linux(Ubantu), Windows.
Application Servers: Apache Tomcat – 8.0
WORK EXPERIENCE:
Organization: Snipe IT Solutions
Role – Hadoop Developer
Duration - June14th 2015 - Till date
Project Title: DHC (Digital Health Care)
Technical Environment: Apache Spark, HDFS, Spark SQL and Scala
Databases: Postgresql, Mongodb.
ETL Tool and Visualization Tool : Pentaho, R, Rstudio
Description:
DHC application is one repository where user can store all his personal health records. Main intent of this application is to store the digitized data .User can store his health related records like, routine check-up reports, disease, drugs, and lab-report and doctor prescription. It helps in understanding user’s personal health history, which is helpful for doctors for further medication and it is more useful in emergency cases. Doctor can use this application for updating prescription for his patients, get the records of his patients and also schedule an appointment.
Responsibilities:
Designing, building, installing, configuring and supporting Hadoop.
Storing data in relational database PostgreSQL, mapping these data into MongoDB which acts as a warehouse
Scheduling Times slot in ETL and Dumping Data from ETL to MongoDB.
Collecting and aggregating large amounts of data using Apache Spark and staging data in MongoDB for further analysis
Querying the stored data using Spark shell and Spark sql
Visualization design on R-studio.
Project Title: Hotel Billing System
Technical Environment: Apache Spark, HDFS, Spark SQL, Hive, Scala and Java
Databases: Postgresql, Mongodb
ETL Tool and Visualization Tool : Pentaho, R, Rstudio
Description:
It’s a project in which the billing report of every hotel is maintained and stored in cloud. This includes food billing and room services, if admin have multiple hotels he or she can login with their particular id and can maintain the details. Admin can see the transactions with the help of dashboards which are obtained after series of analysis.
Responsibilities:
Loading Billing data from PostgreSQL to ETL.
Scheduling Times slot in ETL and Dumping Data from ETL to MongoDB.
Querying the stored data using Spark shell and Spark sql
Experience in analyzing data with Hive.
Writing Scripts using Scala for analytics.
Preparing High charts for Visualization using R.
Project Title: Sentiment Analysis the Twitter data
Technical Environment: Apache Flume, HDFS, Hive, R, Rstudio.
Description:
sentiment analysis of particular Movie across different nations using data from twitter. It deals with using the dictionary file to score the sentiment of each tweet by the number of positive words compared to number of negative words, and then assigned a positive, negative or neutral sentiment value to each tweet.
Responsibilities:
For fetching the twitter data, I am using Apache Flume.
Data would be loaded directly to HDFS .
Data loaded in HDFS is still in unstructured format and not good for Ad-hoc analysis. So I will be converting the JSON data to tabular format and store it in HIVE.
Hive-serde.jar using maven package in Hadoop cluster.
Following analysis using Hive and Tableau:-
a) Maximum tweets count per user.
b) Count of re-tweets.
c) Geographically mapping people’s sentiments towards Movie - bahubali
Finally generate the report like Word Cloud and Bar Chart using R tool.
ADDITIONAL INFORMATIONS:
EDUCATION:
Bachelor of Engineering in Information Science and Engineering,
College - PESIT South Campus - Bangalore, University of VTU, Bangalore, Karnataka, India, 2015
Diploma In computer Science and Engineering University of KEA, Bangalore, Karnataka, India, 2008.
DECLARATION:
I hereby declare that the above information is correct to the best of my knowledge and I take complete responsibility for any false information.
PLACE: BANGALORE MAHESH MANJU