Post Job Free
Sign in

Data Manager

Location:
Duluth, GA
Posted:
June 08, 2017

Contact this candidate

Resume:

Hadoop and Spark Developer

Aparna Bolla

717-***-****

*********@*******.***

Professional Summary:

Overall 8 years of professional IT experience in Software Development. This also include 4 years of experience in Big Data using Hadoop technologies and solutions in various domains such as Media, Finance and Banking.

Hands on experience in Hadoop Technologies such as HDFS, Hive, Pig, Spark, Sqoop, Impala, Flume, Solr, Avro, Chukwa, Kafka.

Hands on experience in writing MapReduce jobs in Hive, Pig and Efficient in building hive, pig and MapReduce scripts.

In depth understanding of Hadoop Architecture and its various components such as Resource Manager, Application Master, Name Node, Data Node.

Capable of processing large sets of structured, semi-structured and unstructured data and supporting systems application architecture.

Experienced in developing PigLatin and HiveQL scripts for Data Analysis and ETL purposes and extended the default functionality by writing User Defined Functions (UDFs), (UDAFs) for custom data specific processing.

Tuning Hadoop applications for high performance and throughput. Troubleshoot and debug Hadoop ecosystem run time issues.

Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice-versa.

Hands on experience in in-memory data processing with Apache Spark and Scala.

Experience with Kafka in understanding and performing thousands of megabytes of reads and writes per second on streaming data.

Experience in collecting log data from many source, aggregating it and writing it to HDFS using Flume.

Good knowledge in NoSQL database including MongoDB, Cassandra and HBase.

Hands on experience in AWS Cloud in various AWS services such as Redshift cluster, Route 53 domain configuration.

Experienced with different file formats like CSV, Text files, Sequence files, XML, JSON and Avro files.

Used Oozie operational services for batch processing and scheduling workflows dynamically with Python scripting.

Experienced in using Zookeeper to provide a centralized infrastructure and services that enable synchronization across a cluster.

Developed Apache Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying.

Detailed understanding of Software Development Life Cycle (SDLC) and sound knowledge of project implementation methodologies including Waterfall and Agile.

Strong knowledge in Dimensional modeling using Star schema and Snowflake schema.

Good Knowledge on OLAP and OLTP systems, Dimensional modeling, extraction, transformation and loading (ETL) process.

Good experience in Java/J2EE programming.

Working experience with SOA, Web Services as well as understanding of MVC and other design patterns.

Experience working with REST services.

Good Knowledge on User Interface development with AngularJS, Bootstrap, HTML5 and AJAX technology.

Good experience on Application Server like Tomcat/Weblogic.

Good experience in using Tableau, Zeppelin Tools.

Good Knowledge on Python Scripting Language.

Good Knowledge on Search engine experience like Lucene/Solr.

Excellent communication, analytical skills and attitude to learn new technologies in the IT industry.

TECHNICAL SKILLS:

Languages

C, C++, Java, HQL, Pig Latin, Python, Scala

Big Data technologies

HDFS, YARN, MapReduce, Hive, Pig, Sqoop, Flume, Impala, Oozie, Zookeeper, Storm, Kafka, Spark.

Hadoop Distributed systems

Cloudera, Hortonworks, MapR

Frameworks

Spring, Hibernate, Struts.

J2EE Technologies

JSP, Servlets, EJB, JMS, JPA, JDBC, AOP, Java Mail.

IDE/Tools

Eclipse, Maven, Ant

App/Web Servers

Apache Tomcat, JBOSS, WebSphere, WebLogic.

Databases

Oracle, SQL Server, MySQL

NoSQL Databases

HBase, MongoDB, Cassandra

Operating systems

UNIX, Windows, LINUX.

Project Details:

Project_1: Advance Publications – Los Angeles, CA Dec 2015 - present

Role : Hadoop and spark Developer

Description:

Advance Publications, Inc., is an American media company that traditionally collects and generates vast amounts of data. They collect raw data and make this data easily accessible for public.

Responsibilities

Worked on analyzing Hadoop cluster and different Big Data analytic tools including Pig, Hive, Cassandra, and Sqoop.

Developed a spark pipeline to transfer data from datalake to Cassandra in cloud to make the data available for decision engine to publish customized offers real time.

Created concurrent access for hive tables with shared and exclusive locking that can be enabled in hive with the help of Zookeeper implementation in the cluster.

Implemented Kafka Custom encoders for custom input format to load data into Kafka Partitions.

Supported code/design analysis, strategy development and project planning.

Used Spark API over Hadoop YARN to perform analytics on data in Hive.

Implemented Oozie coordinated workflow to execute Sqoop incremental job on a regular basis.

Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.

Designed and implemented Cassandra no SQL based database.

Worked on Spark code using Scala and Spark-SQL for faster testing and processing of data.

Developed Spark Streaming script which consumes topics from distributed messaging source Kafka and periodically pushes batch of data to Spark for real time processing.

Assisted with data capacity planning and node forecasting.

Used AWS services like EC2 and S3 for small data sets.

Responsible for running Hadoop streaming jobs to process terabytes of xml data.

Involved in extracting customer's big data from various Transfer Protocol using Flume into Hadoop HDFS.

Extracted files from Cassandra through Sqoop and placed in HDFS for further processing.

Collaborated with the infrastructure, network, database, application and BI teams to ensure data quality and availability.

Implementation of POC on Hadoop stack and different big data analytic tools, migration from different databases (i.e. Teradata, Oracle, MySQL) to Hadoop.

Tested raw data and executed performance scripts.

Designed and Developed Automation framework and Participated in the scrum meeting.

On time completion of tasks and the project per quality goals.

Environment: Hadoop, PIG, Hive, Apache Sqoop, Oozie, Zookeeper, Hortonworks, Flume, Kafka, Spark, Cassandra.

Project_2 : CNA Financial - Connecticut May 2014- Nov 2015

Role : Hadoop Developer

Description:

This application involves in designing and development of the Data Warehouse the main scope is to analyze the data from all policies to find out the customer and company profitable plan with services which assists in determining new strategic plans into market.

Responsibilities:

Ingested the raw data, populated staging tables and stored the refined data.

Created internal table, Externals tables in HIVE, and merged the data sets using Hive joins.

Worked with parsing XML files using Map reduce to extract customer related attributed and store it in HDFS.

Designing and developing tables in HBase and storing aggregating data from Hive.

Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.

Assisted in Cluster maintenance, Cluster Monitoring and Troubleshooting, Manage and review data backups and log files.

Created Hive queries for the market analysts to analyze the emerging data and comparing it with fresh data with EDW reference tables.

Involved in the regular Hadoop Cluster maintenance such as patching security holes and updating system packages.

Moved the data from Hive tables into MongoDB collections.

Worked on Hive scripts to extract, transform, load (ETL) and store the data.

Shared responsibility with administration of Hive and Pig.

Worked in Apache Tomcat for deploying and testing the application.

Designed and Developed Dashboards using Tableau.

Worked with different file formats like Text files, Sequence Files, Avro.

Managed and reviewed Hadoop log files using Flume.

Extracted files from MongoDB through Sqoop placed in HDFS and processed.

Extended support for application to work with Hive, Pig, HBase and Sqoop.

Enabled speedy reviews by using Oozie for automated data loading into the HDFS and PIG to process the data.

Actively participated in all the SCRUM meetings as part of Agile.

Co-ordinating with other programmers in the team to ensure that all the modules complement each other well.

Environment:

Hive, Pig, HBase, MySQL, Flume, Eclipse, Map-Reduce, MongoDB, Cassandra, NetBeans

Project_3 : KEY Bank - Indianapolis, IN July 2013– April 2014

Role : Hadoop Developer

Description:

Helped this regional bank streamline business processes by developing, installing and configuring Hadoop ecosystem components that moved data from individual servers to HDFS.

Responsibilities:

Worked on importing data from various sources and performed transformations using MapReduce, Hive to load data into HDFS.

Continuous monitoring and managed the Hadoop cluster using Cloudera Manager.

Worked on Oozie workflow to run multiple jobs.

Implemented Storm builder topologies to perform cleansing operations before moving data into HBase.

Monitoring and tuning MapReduce Programs running on the cluster.

Involved in extracting customer's big data from various data sources into Hadoop HDFS. This included data from mainframes, databases and logs data from servers.

Building scalable distributed data solutions using Hadoop.

Created HBase tables to store variable data formats coming from different portfolios.

Performed real time analytics on HBase using Java API and Rest API.

Implemented HBase Co-processors to notify Support team when inserting data into HBase Tables.

Designed and developed Map Reduce jobs to process data coming in different file formats like XML, CSV, JSON.

Configured Sqoop jobs to import data from RDBMS into HDFS using Oozie workflows.

Worked on compression mechanisms to optimize MapReduce Jobs.

Analyzed the customer behavior by performing click stream analysis and to ingest the data using Flume.

Solved small file problem using Sequence files processing in Map Reduce.

Implemented business logic by writing UDF's in Java and used various UDF's from Piggybanks and other sources.

Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.

Gained experience with NoSQL database.

Developed Hive Scripts, Pig scripts, Unix Shell scripts, programming for all ETL loading processes and converting the files into parquet in the Hadoop File System.

Environment:

Hadoop - Pig, Hive, Oozie, HBase, HDFS, Java, Cloudera manager, LINUX.

Project_4 : IDBI Bank- India May 2011 – June 2013

Role : Java/J2EE Developer

Description:

This is a financial service company in India, formerly known as Industrial Development Bank of India. This project involves collecting payment information, online order processing and payment processing.

Responsibilities:

Implemented the core java programming i.e. object oriented programming concepts for the banking modules

Developed screens based on JQuery to dynamically generate HTML and display the data to the client side.

Led the migration of monthly statements from UNIX platform to MVC Web-based Windows application using Java, JSP, Struts technology.

Implemented Action Classes and server side validations for account activity, payment history and Transactions.

Implemented views using Struts tags, JSTL and Expression Language.

Implemented session beans to handle business logic for fund transfer, loan, credit card & fixed deposit modules.

Prepared use cases, designed and developed object models and class diagrams.

Developed SQL statements to improve back-end communications.

Used Spring Batch for reading, validating and writing the daily batch files into the database

Incorporated custom logging mechanism for tracing errors, resolving all issues and bugs before deploying the application in the WebSphere Server.

Generated the Web Services and helped the clients to understand the systems.

Used the JDBC for data retrieval from the database for various inquiries.

Designed the application using MVC framework for easy maintainability.

Involved in creating tables, stored procedures in SQL for data manipulation and retrieval using SQL Server, Oracle and DB2.

Established communication among external systems using Web Services (SOAP).

Developed JUnit test cases for regression testing and integrated with ANT build.

Perform application testing involving multiple up/downstream systems, create test cases and test plans from scratch, analyze test results and produce detailed issue reports.

Code Review & Debugging using Eclipse Debugger.

Update clients with weekly status reports.

Environment:

Junit, Java Script, Web Services (SOAP), jQuery, Ajax, JSON, SVN, Oracle SQL Developer.

Project_5 : Oriental Insurance-Hyderabad, India June 2009 -April 2011

Role : Java Developer

Description:

Oriental Insurance has always been a major contributor to the development and growth of India’s life insurance sector, came up with a requirement that supports their executives in processing and maintaining billing details of their customers.

Responsibilities:

Assisted in designing and programming for the system, which includes development of Process Flow Diagram, Entity Relationship Diagram, Data Flow Diagram and Database Design.

Developed front end screens using JSP, HTML, CSS and JavaScript.

Involved in developing Java APIs, which communicates with the Java Beans.

Extensively involved in the development of persistence layer using Hibernate and used SQL server as backend database

Implemented MVC architecture using Java, Custom and JSTL tag libraries.

Implemented the data access using Hibernate and wrote the domain classes to generate the Database Tables.

Implemented MVC architecture and DAO design pattern for maximum abstraction of the application and code reusability.

Created Stored Procedures using SQL/PL-SQL for data modification.

Used XML, XSL for Data presentation, Report generation and customer feedback documents.

Used Java Beans to automate the generation of Dynamic Reports and for customer transactions.

Implemented Logging framework using Log4J.

Involved in code review and documentation review of technical artifacts.

Environment:

Java, Spring, JSP, Hibernate, XML, HTML, JavaScript, JDBC, CSS, SOAP Web services.



Contact this candidate