PARTHIBAN VENKATESAN
Email ID: ***********@*****.***
Mobile No: +1-469-***-****
LinkedIn:
https://www.linkedin.com/in/parthiban-venkatesan-a830b983
PROFESSIONAL SUMMARY
** **** ***** ** ** experience in software development, data analysis, Data Engineering, ETL, and requirements gathering, I have worked across diverse industries, including banking, healthcare, retail, hospitality, and gaming. My expertise spans data warehousing, on-prem to cloud migration, and building data pipelines for both on-prem and cloud environments.
Experience working with business users to convert technical data into business insights. Have led integration and migration projects that involve a large team of people and a wide variety of data. Moving legacy systems from on-prem databases (e.g., Teradata,SQL Server, Oracle) to build scalable and reliable cloud platform based in Azure and databricks, while exploring the other cloud options in Snowflake, GCP etc..
Recent project involve leading a team and design, develop a comprehensive data purging framework to manage large-scale data archival and purging across transactional and historical databases built in the multiple sql servers.
Write stored procedures, triggers, and cursors in PL-SQL to automate purging operations based on configurable retention policies and business rules.
Design and Architect an efficient plan to perform the purging activity that is capable of handling billions of records to process utilization of SSIS packages.
Experienced in architecting and developing ETL and ELT pipelines in Azure Data Factory (ADF), SSIS, and Informatica PowerCenter, as well as performance tuning in Databricks with PySpark and Spark SQL.
Expertise in database management with SQL and SSMS, and have designed scalable and secure data solutions, ensuring efficient data flow, orchestration, and performance optimization. Additionally, I have implemented automated workflows using Airflow and Google Composer, performed POCs using AWS and GCP, and provided solutions utilizing Snowflake's advanced features like zero-copy cloning, time travel, and snowpipes.
Furthermore, I bring leadership experience, leading onshore and offshore teams, and ensuring smooth project execution through project management practices using tools like Jira and Confluence. My industry experience in Banking, Travel and Healthcare also adds value in managing large-scale data migration and compliance projects.
Consistently received excellent feedback and commendations from business users for delivering high-quality, reliable data solutions that significantly improved decision-making processes and operational efficiency (Can be shared as required)
TECHNICAL SKILLS
Databases & Data Warehousing:
Teradata Oracle SQL Server DB2 Snowflake BigQuery Databricks
Analytics & BI Tools
Power BI SAS
Big Data & Processing
PySpark SQL Python
Cloud Platforms & Services
Azure Databricks ADLS Azure Data Factory GCP
ETL & Data Integration Tools
Informatica PowerCenter (9.5.1, 10.x) Informatica IICS Datastage 11.5 Dat Fusion (GCP) Workato SSIS Data Integration Hub
Operating Systems & Platforms
Mainframes Unix
Data Governance & Quality
Data Governance Data Cleansing
Scheduling & Automation Tools
CA7 AutoSys Control-M WCC ENDEVOR
Development & Collaboration Tools
Jira GitHub Confluence
Data Modeling & Design
Erwin
Database & Client Utilities
SQL Assistant Teradata Studio SSMS Visio
Methodologies
Waterfall Agile
+
EDUCATION
Bachelor’s in engineering, Electrical, Electronics And Communications Engineering, SA Engineering College, Chennai – Mar 2007
CERTIFICATIONS
Texas McCombs School of Business Texas McCombs School of Business Post Graduate Program in Artificial Intelligence and Machine Learning
Business Applications, Artificial Intelligence and Machine Learning
Post Graduate Program in Artificial Intelligence and Machine Learning: Business Applications, Artificial Intelligence and Machine Learning Sep 2021 - Aug 2022Sep 2021 - Aug 2022
oActivities and societies: https://eportfolio.mygreatlearning.com/parthiban-venkatesan
WORK EXPERIENCE
Randstad/ Bank of America - Plano, TX Jul 2024 – Current
Technical Architect
Responsibilities:
Develop framework to perform the purge activity for the legacy data and DBs for collateral and mortgage systems
Understand the data, provide design solutions, architecting the framework with the necessary tools and utilities involved.
Build Stored procedures, triggers in SQL server for performing the purge activity
Utilize SSIS to build packages that execute the stored procs
Write Teradata components using utilities BTEQ, TPT, MLOAD, FLoad for the migration activity based on design and file formats
Develop a comprehensive data purging framework to manage large-scale data archival and purging across transactional and historical databases build in the multiple sql server for the collateral and mortgage data of the bank
Create high-performance stored procedures, triggers, and cursors to automate purging operations based on configurable retention policies and business rules using the PL-SQL Queries
Design and Architect an efficient plan to perform the purging activity that is capable of handling billions of records to process
Purging process involve utilization of SSIS packages, Reports that helps users to choose the application and the databases to be in the purge.
Build pipelines in the SSIS packages that connects to sources like databases, files etc.. create connection managers and layers
Utilize datastage quality stage in address matching process and test different match specifications using multiple level of passes by exact to fuzzy matches.
Good experience with SSMS for database management, data analysis, and SQL query performance tuning.
Create the Indexing, layout architectures and data
Create Datastage sequences and the jobs for bringing data from multiple files and sources like BMH and tables in sql server.
Create audit, exception, and reconciliation layers to validate pre- and post-purge data states, with built-in alerting, logging, and recovery mechanisms.
Conducted performance tuning of T-SQL and write queries to handle billions of records
Propose the data archival steps with cloud as alternatives
Good understanding on ADLS, Azure Data Factory, Azure Databricks, Synapse equivalent to Amazon RDS.
Work on POC projects to utilize the cloud services in various cloud providers.
Good understanding and utilization of file formats like json, parquet, avro etc in the cloud
Knowledge on the performance tuning techniques in pyspark like Coahlesce, repartitioning,bucketing.
Understanding on Joins like Broadcast join, Hash join and different performance techniques like salting, wide transformation, filtering.
Create tables and datasets in Bigquery using webconsole and CLI python
Implement Snowflake as a datawarehouse on top of the Azure storage, write Snowflake queries.
Create autosys jobs and design the scheduling of jobs
Environment: Teradata, Azure Data Factory, Azure Databricks, Synapse, BigQuery,Oracle, SSIS, SQL server, Confluence, Github,Python, Autosys Job Scheduler, Datastage,ADLS, SAS and Power BI reports
MGM Resorts - Plano, TX Nov 2020 – Jul 2024
Senior Teradata and ETL Data Engineer
Responsibilities:
Understand the various business data streams and their corresponding databases, responsible as a SME to assist the business users to build dashboards in power bi or cognos feeding data from Teradata/Azure
Create new mapping sheets for any additional stream of business data related to various operations like guest services, call centers, Gaming, Labor, ticket reservation
Write Teradata components using utilities BTEQ, TPT, MLOAD, FLoad for the migration activity based on design and file formats
Convert the SAS reports to Teradata and PowerBI reporting that includes end to end data flow and reporting structure and dataframes in SAS and rewrite the code equivalent in sql and BI reports.
Performed POCs using GCP google cloud platform and Azure to migrate on-prem datasets of hotel and slots data utilizing google bigquery
Create tables and datasets in Bigquery using webconsole and CLI python
Implement Snowflake as a datawarehouse on top of the Azure storage, write Snowflake queries.
Good understanding of snowflake concepts zero copy cloning, time travel, data sharing.
Create ETL pipelines in GCP using Data Fusion, perform computing, customization, performance tuning
Migrate the existing code and logic in On-prem Teradata to Azure databricks, convert Teradata SQLs to Spark Sqls, and pyspark from on-prem to cloud
Developed data pipelines in ADF to ingest and process large datasets.
Orchestration of jobs in ADF for new marts and migrating from on-prem
Performance tuning of pyspark jobs in azure databricks, using narrow transformations, repartitioning and coahlesce techniques
Create the medallion architecture based bronze-silver-gold data flow
Developed data pipelines to ingest and process large datasets.
Orchestration of jobs in ADF for new marts and migrating from on-prem
Work with admins for resource and storage accounts setup
Perform modelling activities on migration of the marts
Designed, built, and maintained high-performance databases for reporting and analysis purposes.
Write SQL queries in Teradata to create views that is sourced to the dashboards built on Power BI and Cognos
Write new and modify existing python scripts that pull data from multiple sources and databases and convert to csv and populate to Teradata using the utilities
Perform UT on the Azure Databricks using the python spark sql and test SCD1 and SCD2
Debug, support the existing on-prem jobs and ETL built on Datastage
Run Datastage jobs, debug code in datastage and convert transformation to databricks and Teradata
Create new tables and views in Teradata for various business streams adhering to Data model
Responsible to analyze the existing data and the data flow in EDW that holds the clients enterprise data built on Teradata 16
Maintain the jobs schedule in Atomic, monitor daily jobs and provide support and maintain during the schedule maintenance
Perform GAP analysis between the middleware application and the database sources built on DB2 and sourced from Oracle and Teradata
Analyze data present outside of EDW built on databases like Teradata, Oracle PL SQL, SQL Server etc
Utilize Confluence for uploading documentation to be shared across the team
Coding version control is uploaded to Github, create branches and peers in github
Write and modify the ETL scripts in databricks, perform scheduling and monitoring in ADLS
Modify the existing and create new python scripts for file handling
Create and track projects in Jira, write user stories and attend scrum meetings.
Attend Business meetings to gather requirements, create a design strategy, convert functional requirements to technical design and write design documents
Write complex Teradata and SQL queries that include various join strategies to arrive at the design, create Unix shell scripts to handle files created on .csv, xml etc
Perform Data analysis before creating the design and perform data cleansing of files
Create source to target mappings before coding by understanding the Physical and Logical data model working along with Data modeler
Worked as a data modeler to bring in venue details using Erwin data model
Designed and implemented data models for various Curam modules, including Participant Management, Case Management, and Eligibility Determination for the feeds on voluntary and social activities.
Developed custom data structures and relationships to meet client-specific requirements and improve data integrity.
Lead migration activity analysis for Teradata to Cloud
Senior ETL Analyst Aug 2019 – Nov 2020
Compunnel Software Group Inc - Plano, TX
Responsibilities:
Responsible to analyze the existing data and the data flow in EDW that holds the clients enterprise data built on Teradata 16
Perform GAP analysis between the middleware application and the database sources built on DB2 and sourced from Oracle and Teradata
Analyze data present outside of EDW built on databases like Teradata, Oracle PL SQL, SQL Server etc that uses code in Mainframes, UNIX etc
And provide ETL solutions using Oracle, Teradata, Datastage and Informatica Powercenter 10.1 to bring in to EDW
Utilize Confluence for uploading documentation to be shared across the team
Maintain spaces for knowledge and space documentation
Coding version control is uploaded to Github, create branches and peers in github
Modify the existing and create new python scripts for file handling
Utilize business intelligence tools Tableau for retails data reporting
Attend Business meetings to gather requirements, create a design strategy, convert functional requirements to technical design and write design documents
Lead team on design and development and support until warranty
Write Stored procedures and Triggers using T-SQL in SQL server, PL-SQL in Oracle, create packages to perform ETL
Create error handling workflows and mappings to generate email alerts
Write complex Teradata and SQL queries that include various join strategies to arrive at the design, create Unix shell scripts to handle files created on .csv, xml etc
Write Teradata components using utilities BTEQ, TPT, MLOAD, FLoad for the migration activity based on design
Utilize Teradata TASM for maintaining workload management, performance monitoring and increase efficiency of performance
Perform data and object migration between systems using Data mover in TeradataWrite Teradata Stored procedures and Triggers for performing CDC on each tables
Create Informatica mapping, workflows, using various transformations for performing ETL aspects of the data migration
Utilize Erwin data model to make changes and include new tables into the model
Performance tuning on Informatica mappings and workflows in powercenter, utilize Push down optimization
Implement performance tuning and write advance SQL and PL-SQL queries in finding out the equivalent Physical columns within EDW for the corresponding business termsCreate workflow and jobs in Workato, to access data in web using rest API and populate them in databases
Use SQL Server Management Studio to create tables in SQL server and populate data from APIs and create ETL jobs using SSIS
Environment : Teradata, Oracle, SSIS, Informatica, Unix, SQL server, Workato, Confluence, Github,Python
Senior ETL Programmer Analyst Teradata, Oct 2017 – Aug 2019
Systems Technology Group, Inc, Ford Motor Credit Company - Dearborn, MI
Responsibilities:
Worked in Data Integration and migration projects within different databases of Teradata by utilizing the ETL Capabilities
Validate the data compatibility between systems across various databases
Create designs in Teradata and ETL to migrate new data into the system
Responsible to work in Enterprise data warehouse (EDW) and Analyze the jobs created in Mainframes, Teradata and other (Extract transform Load) ETL utilities like Datastage used in the client environment that access the EDW to bring new data and create reports on existing data
Perform migration activity on the components using (Job Control Language) JCLS, Parmcards, BTEQ, Fastload, MLoad
Understand the Logical and Physical data model, relationship of tables and Create mapping documents based on the analysis on existing jobs to create new components
Perform Proof of concepts on converting SQLs into Datastage stages for enabling lineage
Attend Business meetings, create a design strategy and convert functional requirements to design
Interpret the various business terms and identify/mine their equivalent technical terms inside the data warehouse built on Teradata
Write various PL-SQL queries to perform updating the existing data based on the business requirements
Implement performance tuning and write advance SQL queries in finding out the equivalent physical columns for the corresponding business terms
Work with DBAs to create DDLs, ensure proper standards are setup, Skewness is avoided in queriesTweak underperforming queries and work with DBA to setup spaces
Create Datastage components utilizing various stages of IBM Datastage
Perform Proof of Concept in converting complex SQL codes into Datastage mappings
Environment : Teradata, Mainframes, Datastage, Autosys
Cognizant Technology Solutions, PEPSI CO - Plano, TX Jan 2016 - Sep 2017
ETL Developer
Responsibilities:
Working as a Data Analyst and Developer for a retail domain, responsible for gathering requirements, develop SQL code, Unit testing
Understand the data model, write source to Target mappings for the ETL process
Perform data quality checks based on the test data
Create Mappings, workflows in Informatica power center 9.5.1, 10.1 and Cloud perform Publication, Subscriptions in cloud and usage of different transformations or create packages in SSIS and create ETL jobs
Write Shell scripts in handling csv, xml, txt files in Unix, use SED,AWK commands in file handling and writing trigger scripts
Involve in performance tuning of T-SQL Queries, creating mapping source to target mapping sheets, handling stats on the tables
Generate reports by writing ad-hoc T-SQLs
Responsible for creating and doing unit test, ensure performance tuning using Informatica Power Center 9.5.1 and 10.1
Develop new components in Informatica Data Integration Hub (DIH)
Teradata Utilities Fast load, Mload, Bteq, Tpump and responsible for Performance tuning
Worked in Oracle 11g, using load utilities like SQLPLUS, SQLLDR
Environment: Oracle, Teradata, SQL Server, SSIS, Informatica Power center 9.5.1 and 10.1, Data Integration Hub (DIH), SQL Assistant, UNIX, Oracle, Db visualizer.
Wipro Technologies Ltd, CVS Care Mark - Richardson, Dallas, TX Feb 2015 – Jan 2016
Programmer Analyst, ETL Analyst
Responsibilities:
As an ETL Lead responsible to gather requirements, Design and assist develop and review the code on the health care projects in Teradata 13, SQL Server with Informatica Power center 9.5.1 and SSIS as ETL on UNIX operating system
Responsible for creating design documents, source to target mappings as per functional requirements and presentation to client side Leads and SMEs for approvals
Perform release management process in Serena
Good knowledge on Health care subject areas CLIENT, MEMBER, ACCOUNT, GROUP, DRUG, PRESCRIBER
Create T- SQLs, Tables is SQL server and Teradata components BTEQ, MLOAD, FLOAD and Informatica workflows, sessions and mapping based on data model and requirements
Environment: Teradata, SQL Server, SSIS, Informatica Power center 9.5.1, SQL Assistant, UNIX, Oracle, Design and documentation
Wipro Technologies Ltd, Lloyds Banking Group - Manchester, UK Jan 2011 – Jan 2015
Subject Matter Expert
Responsibilities:
For Bank's data warehouse and Teradata lead, Responsible to successfully handle and deliver the data migration and integration activity of two major banks and also handle SDLC projects on Retail portfolio
In depth knowledge in various databases that holds each type of data in the warehouse
Working on Teradata 12, Mainframes, ETL and FSLDM
The core database of the ware house was built on Teradata FSLDM, Expertise in FSLDM on subject areas- Agreement, Party, Feature, Event, Product
Knowledge on using different type of Indexes like Join Index, Multi Table JI, Single table JI
Environment: Teradata, TD Utilities, ETL, Mapping, SQL Assistant, UNIX, Mainframes zOS, JCLs, Connect Direct.
Infosys Technologies LTD, Bank of America, Chennai, India Jul 2007 – Dec 2010
Senior Systems Analyst – Teradata Developer
Responsibilities:
Responsible to build Teradata components like BTEQ, MLOAD, FLOAD in Mainframes
Build Mainframe components JCL, PROCS, parmcards
Environment: Teradata, TD Utilities, Oracle –PL SQLs, ETL, Mapping, SQL Assistant, UNIX, Mainframes zOS, Change man, File Aid, JCLs, CA7, Data Migration and System migration.