Post Job Free
Sign in

Us Citizen Big Data

Location:
New Hyde Park, NY
Posted:
March 29, 2024

Contact this candidate

Resume:

ANGELINA BUSHRA

917-***-****

***********@*****.***

US Citizen, New York City, NY

PROFESSIONAL SUMMARY

Over 6 years of extensive IT experience and more than 4 years of experience as a Hadoop development, operations and management of small to medium clusters using distributions like CDH, HDP and ECS.

Hands on Experience in Installing, configuring and using Hadoop Ecosystem components like HDFS, Hadoop MapReduce, Yarn, Zookeeper, Sqoop, Impala, Flume, Hive, Pig, HBase, Spark, Pig, Oozie.

Involved in converting HQL queries into Spark transformations using Spark RDDs.

Experienced in installing and configuring Flume, Hive, Pig, Sqoop and Oozie on the Hadoop cluster.

Experience in deploying a Hadoop clusters using Cloudera 5.X integrated with Ambari for management, monitoring and alerting.

Experienced in launching and setting up of Hadoop clusters on AWS as well as physical servers, which includes configuring different Hadoop components

Developed and monitored of Puppet Configuration Manager and automated configuration of Hadoop ecosystem.

Good knowledge on implementation and design of big data pipelines and implementing of ETL/ELT processes

Hands on experience in using MapReduce programming model for Batch processing of data stored in HDFS.

Responsible for designing and building a Data Lake using Hadoop and its ecosystem components.

Very good experience in complete project life cycle (design, development, testing and implementation) of client-server and web applications

Developed Spark Applications by using Scala, Java and Implemented Apache Spark data processing project to handle data from various RDBMS and Streaming sources

Worked with the Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Spark MLlib, Data Frame, Pair RDD's, Spark YARN

Experienced in Apache Spark for implementing advanced procedures like text analytics and processing using the in-memory computing capabilities written in Scala

Experience using middleware architecture using Sun Java technologies like J2EE, Servlets, and application servers like Web Sphere and Web logic.

Used Different Spark Modules like Spark core, Spark RDD's, Spark Data frame, Spark SQL.

Converted Various Hive queries into Spark transformations and actions that are required.

Experienced in working on apache Hadoop open source distribution with technologies like HDFS, Map-reduce, Python, Pig, Hive, Hue, HBase, SQOOP, Oozie, Zookeeper, Spark, Spark-Streaming, Storm, Kafka, Cassandra, Impala, Snappy, Green Plum and MongoDB, Mesos.

Experienced in Tableau enabled the JDBC and ODBC connectivity from those to Hive tables.

Designed neat and insightful dashboards in Tableau.

Experienced in installing and configuring Kerberos for the authentication of users and Hadoop daemons.

Hands on experience in Linux Hadoop activities on RHEL &Cent OS.

Knowledge on Cloud technologies like AWS Cloud.

Experienced in benchmarking, backup and disaster recovery of Name Node Metadata.

Experienced in working with popular frame works like Spring MVC and Hibernate.

proficient in source code management tools Git.

Excellent interpersonal and communication skills, creative, research-minded with problem solving skills.

Ensure that critical customer issues are addressed quickly and effectively.

Apply troubleshooting techniques to provide solutions to our customer's individual needs.

Troubleshoot, diagnose and potentially escalate customer inquiries during their engineering and operations efforts.

PROFESSIONAL EXPERIENCE

Big Data Developer

Visa, Austin, Texas

March 2022 – present

Responsibilities:

Experience working in an Agile development environment, especially using Scrum framework. Understanding of Agile principles, iterative development, user stories, and sprint planning. Participated in unit testing and sprint reviews. Familiarity with tools like Jira for project management and collaboration.

Implemented data change tracking and incremental updates, enhancing data quality and historical analysis using Apache Hudi.

Designed real-time streaming application with Apache Flink and Kafka for processing and analyzing data with low latency, enabling immediate insights and responses.

Experienced in working with Azure data factory and Azure Synapse for seamless data integration, processing, orchestration in a cloud environment.

Utilized Databricks collaborative analytics platform to process and analyze data at scale, enabling data-driven decision-making in the Azure environment.

Have worked with both batch processing and real-time streaming data, using Hive in conjunction with technologies like Kafka or spark streaming.

Skilled in developing and deploying Azure Functions to automate data transformation and manipulation processes, improving efficiency and scalability.

Implemented data processing and analysis pipelines using PySpark, a python API for Apache Spark, to handle large-scale datasets efficiently and leverage distributed computing capabilities.

Utilized columnar storage formats, like parquet or ORC, to optimize data compression and query performance for petabyte-scale datasets.

Proficient understanding of relational and cloud-based database systems, including Snowflake, AWS RDS, PostgreSQL, Amazon Redshift, and DynamoDB.

Created technical documentation to help other team members to give KT about the project with full explanation of ETL processing, test plans, integration with Databricks and spark and prod deployment in Symphony.

HADOOP DEVELOPER

CreditSuisse, Raleigh, North Carolina

May 2021 - March 2022

Responsibilities:

Worked with the business team to gather the requirements and participated in the Agile planning meetings to finalize the scope.

Involved in creating Hive tables, loading data and writing hive queries.

Used Hive and Spark SQL for analyzing the financial data by extracting datasets for meaningful information such as SF classification, Balance Type, Super Type, Cash noncash, Days to Maturity, Months to Maturity, Equity FID, Product classification, Counterparty Classification, Counterparty levels, Entity Level hierarchy configuration.

Used Pig as ETL tool to do Transformations, even joins and some preaggregations before storing the data.

Created ControlM workflow and Coordinator jobs to kick off the jobs on time for data availability.

Involved in defining job flows. Used Hive to analyze the partitioned and bucketed data and compute.

Developed complex hive queries using joins and partitions for huge data sets as per business requirements.

Involved in importing the data from various data sources into HDFS using Sqoop and applying various transformations using Hive, Apache Spark and then loading data into Hive tables or AWS S3 buckets.

Used Bitbucket as a repository for storing the code and integrated with Jira software for integration purpose.

Worked with Spark Ecosystem using Scala, Hive queries on different data formats like Test file and parquet.

Run impala queries for testing purpose by initializing Kerberos authentication.

Worked with senior engineers for building scalable distributed data solutions using Hadoop, Spark and Scala.

Worked with senior engineers on configuring kafka for streaming data.

Created applications using Kafka, which monitors consumer lag within Apache Kafka clusters.

HADOOP DEVELOPER

Mitsubishi Financial Group (MUFG), New York City, NY October 2019 - May 2021

Responsibilities

Hadoop installation, configuration of multiple nodes using Cloudera platform.

Installed, configured and maintained Hortonworks HDP 2.2 using Ambari and manually through command line.

Worked on analyzing Hadoop clusters using different big data analytic tools including Kafka, Pig, Hive and Map Reduce.

Collected and aggregated large amounts of log data using Apache Flume and staging data in HDFS for further analysis.

Real time streaming the data using Spark with Kafka.

Involved with Continuous Integration team to setup tool GitHub for scheduling automatic deployments of new/existing code in Production.

Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala.

Worked within the Apache Hadoop framework, utilizing Opinion Lab statistics to ingest the data from a streaming application program interface (API), automated processes by creating Oozie workflows, and draw conclusions about consumer sentiment based on data patterns found using Hive for external client use.

Wrote the Storm topology with HDFS Bolt and Hive Bolts as destinations.

Expertise in writing Storm topology development, maintenance and bug fixes.

Developed Hadoop streaming Map/Reduce works using Java.

Implemented test scripts to support test driven development and continuous integration.

Worked on tuning the performance of Pig queries.

Involved in loading data from Linux file system to HDFS and imported and exported data into HDFS using Sqoop.

Good knowledge on building Apache spark applications using Scala.

Experience working on processing unstructured data using Pig.

Implemented Partitioning, Dynamic Partitions, Buckets in Hive.

Implemented Spark using Scala and SparkSQL for faster testing and processing of data.

Good knowledge with NoSQL databases like HBase, Cassandra

Installed, administered, upgraded and managed distributions of Cassandra and involved in Cassandra performance tuning.

Plan, deploy, monitor, and maintain Amazon AWS cloud infrastructure consisting of multiple EC2 nodes and VMWare VMs as required in the environment.

Supported Map Reduce Programs those are running on the cluster.

Managed and reviewed Hadoop log files.

Involved in scheduling of Oozie workflow engine to run multiple pig jobs.

Responsible for developing data pipeline using flume, Sqoop and Pig to extract the data from weblogs and store in HDFS.

Data scrubbing and processing with Oozie.

Developed Pig Latin scripts to extract data from the web server output files to load into HDFS.

Involved in developing Hive DDLs to create, alter and drop tables.

Created and maintained technical documentation for launching Hadoop clusters and for executing Hive queries and Pig Scripts.

Responsible for Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark SQL, Data Frame, Pair RDDs, Spark YARN.

Evaluated existing data platform and apply technical expertise to create a data modernization

roadmap and architect solutions to meet business and IT needs.

Ensured technical feasibility of new projects and successful deployments, orchestrating key resources and infusing key data technologies (e.g. Azure Data technologies like, Azure Data Lake, Azure Blog Storage, Azure SQL DB, Analysis Services)

Built a prototype Azure application that accesses 3rd party data services via Web Services. The solution dynamically scales, automatically adding/removing cloud-based compute, storage and network resources based upon changing workloads.

Built a consumer price comparison application on the Azure platform that requires no on-premises hardware. The application is designed to scale to meet the needs of an unpredictable fluctuating worldwide demand.

Positioned Aflac to extend their on-premises computational functions to the Azure platform.

Worked with systems engineering team to plan and deploy new Hadoop environments and expand existing Hadoop clusters.

Expertise in building Cloudera, Hortonworks Hadoop clusters on bare metal and Amazon EC2 cloud.

Experienced in installation, configuration, troubleshooting and maintenance of Kafka & Spark clusters.

Experience in setting up Kafka cluster on AWS EC2 Instances.

Good understanding on cluster configurations and resource management using YARN

Environment: Hadoop, HDFS, MAPREDUCE, HIVE, PIG, OOZIE, SQOOP, AMBARI, STORM, GFS, ZOOKEEPER, KAFKA, Microsoft Azure.

HADOOP DEVELOPER

Nationwide Insurance, Columbus, Ohio June 2016 - September 2018

Responsibilities

Creation of business development offerings featuring all aspects of OMS architecture, deployment and solutions.

Experience Microsoft Azure date storage and Azure Data Factory, Data Lake.

Developed different kind of custom filters and handled pre-defined filters on HBase data using API.

Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data.

Configured Azure Traffic Manager to build routing for user traffic

Infrastructure Migrations: Drive Operational efforts to migrate all legacy services to a fully Virtualized Infrastructure.

Implemented HA deployment models with Azure Classic and Azure Resource Manager.

Configured Azure Active Directory and managed users and groups

Worked on tuning Hive and Pig to improve performance and solve performance related issues in Hive and Pig scripts with good understanding of Joins, Group and aggregation and how it does Map Reduce jobs

Implemented concepts of Hadoop eco system such as YARN, MapReduce, HDFS, HBase, Zookeeper, Pig and Hive.

In charge of installing, administering, and supporting Windows and Linux operating systems in an enterprise environment.

Involved in Installing and configuring ranger for the authentication of users and Hadoop daemons.

Experience in methodologies such as Agile, Scrum, and Test-driven development.

Worked with cloud services like Amazon Web Services (AWS) and involved in ETL, Data Integration, Datawarehouse, and Migration, and installation on Kafka.

Used Flume extensively in gathering and moving log data files from Application Servers to a central location in Hadoop Distributed File System (HDFS)Used Python and Django creating graphics, XML processing, data exchange and business logic

Created Oozie workflows to run multiple MR, Hive and pig jobs.

Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.

Develop Spark code using Scala and Spark-SQL for faster testing and data processing

Involved in the development of Spark Streaming application for one of the data sources using Scala, Spark by applying

Environment Hadoop, HDFS, Pig, Sqoop, Shell Scripting, Ubuntu, Linux Red Hat, Spark, Scala, Hortonworks, Cloudera Manager, Apache Yarn.



Contact this candidate