Post Job Free
Sign in

Senior Data Engineer and ETL Architect (Azure/Databricks)

Location:
Frisco, TX
Posted:
January 07, 2026

Contact this candidate

Resume:

PARTHIBAN VENKATESAN

Email ID: ***********@*****.***

Mobile No: +1-469-***-****

LinkedIn:

https://www.linkedin.com/in/parthiban-venkatesan-a830b983

PROFESSIONAL SUMMARY

** **** ***** ** ** experience in software development, data analysis, Data Engineering, ETL, and requirements gathering, I have worked across diverse industries, including banking, healthcare, retail, hospitality, and gaming. My expertise spans data warehousing, on-prem to cloud migration, and building data pipelines for both on-prem and cloud environments.

Experience working with business users to convert technical data into business insights. Have led integration and migration projects that involve a large team of people and a wide variety of data. Moving legacy systems from on-prem databases (e.g., Teradata,SQL Server, Oracle) to build scalable and reliable cloud platform based in Azure and databricks, while exploring the other cloud options in Snowflake, GCP etc..

Recent project involve leading a team and design, develop a comprehensive data purging framework to manage large-scale data archival and purging across transactional and historical databases built in the multiple sql servers.

Write stored procedures, triggers, and cursors in PL-SQL to automate purging operations based on configurable retention policies and business rules.

Design and Architect an efficient plan to perform the purging activity that is capable of handling billions of records to process utilization of SSIS packages.

Experienced in architecting and developing ETL and ELT pipelines in Azure Data Factory (ADF), SSIS, and Informatica PowerCenter, as well as performance tuning in Databricks with PySpark and Spark SQL.

Expertise in database management with SQL and SSMS, and have designed scalable and secure data solutions, ensuring efficient data flow, orchestration, and performance optimization. Additionally, I have implemented automated workflows using Airflow and Google Composer, performed POCs using AWS and GCP, and provided solutions utilizing Snowflake's advanced features like zero-copy cloning, time travel, and snowpipes.

Furthermore, I bring leadership experience, leading onshore and offshore teams, and ensuring smooth project execution through project management practices using tools like Jira and Confluence. My industry experience in Banking, Travel and Healthcare also adds value in managing large-scale data migration and compliance projects.

Consistently received excellent feedback and commendations from business users for delivering high-quality, reliable data solutions that significantly improved decision-making processes and operational efficiency (Can be shared as required)

TECHNICAL SKILLS

Databases & Data Warehousing:

Teradata Oracle SQL Server DB2 Snowflake BigQuery Databricks

Analytics & BI Tools

Power BI SAS

Big Data & Processing

PySpark SQL Python

Cloud Platforms & Services

Azure Databricks ADLS Azure Data Factory GCP

ETL & Data Integration Tools

Informatica PowerCenter (9.5.1, 10.x) Informatica IICS Datastage 11.5 Dat Fusion (GCP) Workato SSIS Data Integration Hub

Operating Systems & Platforms

Mainframes Unix

Data Governance & Quality

Data Governance Data Cleansing

Scheduling & Automation Tools

CA7 AutoSys Control-M WCC ENDEVOR

Development & Collaboration Tools

Jira GitHub Confluence

Data Modeling & Design

Erwin

Database & Client Utilities

SQL Assistant Teradata Studio SSMS Visio

Methodologies

Waterfall Agile

+

EDUCATION

Bachelor’s in engineering, Electrical, Electronics And Communications Engineering, SA Engineering College, Chennai – Mar 2007

CERTIFICATIONS

Texas McCombs School of Business Texas McCombs School of Business Post Graduate Program in Artificial Intelligence and Machine Learning

Business Applications, Artificial Intelligence and Machine Learning

Post Graduate Program in Artificial Intelligence and Machine Learning: Business Applications, Artificial Intelligence and Machine Learning Sep 2021 - Aug 2022Sep 2021 - Aug 2022

oActivities and societies: https://eportfolio.mygreatlearning.com/parthiban-venkatesan

WORK EXPERIENCE

Randstad/ Bank of America - Plano, TX Jul 2024 – Current

Technical Architect

Responsibilities:

Develop framework to perform the purge activity for the legacy data and DBs for collateral and mortgage systems

Understand the data, provide design solutions, architecting the framework with the necessary tools and utilities involved.

Build Stored procedures, triggers in SQL server for performing the purge activity

Utilize SSIS to build packages that execute the stored procs

Write Teradata components using utilities BTEQ, TPT, MLOAD, FLoad for the migration activity based on design and file formats

Develop a comprehensive data purging framework to manage large-scale data archival and purging across transactional and historical databases build in the multiple sql server for the collateral and mortgage data of the bank

Create high-performance stored procedures, triggers, and cursors to automate purging operations based on configurable retention policies and business rules using the PL-SQL Queries

Design and Architect an efficient plan to perform the purging activity that is capable of handling billions of records to process

Purging process involve utilization of SSIS packages, Reports that helps users to choose the application and the databases to be in the purge.

Build pipelines in the SSIS packages that connects to sources like databases, files etc.. create connection managers and layers

Utilize datastage quality stage in address matching process and test different match specifications using multiple level of passes by exact to fuzzy matches.

Good experience with SSMS for database management, data analysis, and SQL query performance tuning.

Create the Indexing, layout architectures and data

Create Datastage sequences and the jobs for bringing data from multiple files and sources like BMH and tables in sql server.

Create audit, exception, and reconciliation layers to validate pre- and post-purge data states, with built-in alerting, logging, and recovery mechanisms.

Conducted performance tuning of T-SQL and write queries to handle billions of records

Propose the data archival steps with cloud as alternatives

Good understanding on ADLS, Azure Data Factory, Azure Databricks, Synapse equivalent to Amazon RDS.

Work on POC projects to utilize the cloud services in various cloud providers.

Good understanding and utilization of file formats like json, parquet, avro etc in the cloud

Knowledge on the performance tuning techniques in pyspark like Coahlesce, repartitioning,bucketing.

Understanding on Joins like Broadcast join, Hash join and different performance techniques like salting, wide transformation, filtering.

Create tables and datasets in Bigquery using webconsole and CLI python

Implement Snowflake as a datawarehouse on top of the Azure storage, write Snowflake queries.

Create autosys jobs and design the scheduling of jobs

Environment: Teradata, Azure Data Factory, Azure Databricks, Synapse, BigQuery,Oracle, SSIS, SQL server, Confluence, Github,Python, Autosys Job Scheduler, Datastage,ADLS, SAS and Power BI reports

MGM Resorts - Plano, TX Nov 2020 – Jul 2024

Senior Teradata and ETL Data Engineer

Responsibilities:

Understand the various business data streams and their corresponding databases, responsible as a SME to assist the business users to build dashboards in power bi or cognos feeding data from Teradata/Azure

Create new mapping sheets for any additional stream of business data related to various operations like guest services, call centers, Gaming, Labor, ticket reservation

Write Teradata components using utilities BTEQ, TPT, MLOAD, FLoad for the migration activity based on design and file formats

Convert the SAS reports to Teradata and PowerBI reporting that includes end to end data flow and reporting structure and dataframes in SAS and rewrite the code equivalent in sql and BI reports.

Performed POCs using GCP google cloud platform and Azure to migrate on-prem datasets of hotel and slots data utilizing google bigquery

Create tables and datasets in Bigquery using webconsole and CLI python

Implement Snowflake as a datawarehouse on top of the Azure storage, write Snowflake queries.

Good understanding of snowflake concepts zero copy cloning, time travel, data sharing.

Create ETL pipelines in GCP using Data Fusion, perform computing, customization, performance tuning

Migrate the existing code and logic in On-prem Teradata to Azure databricks, convert Teradata SQLs to Spark Sqls, and pyspark from on-prem to cloud

Developed data pipelines in ADF to ingest and process large datasets.

Orchestration of jobs in ADF for new marts and migrating from on-prem

Performance tuning of pyspark jobs in azure databricks, using narrow transformations, repartitioning and coahlesce techniques

Create the medallion architecture based bronze-silver-gold data flow

Developed data pipelines to ingest and process large datasets.

Orchestration of jobs in ADF for new marts and migrating from on-prem

Work with admins for resource and storage accounts setup

Perform modelling activities on migration of the marts

Designed, built, and maintained high-performance databases for reporting and analysis purposes.

Write SQL queries in Teradata to create views that is sourced to the dashboards built on Power BI and Cognos

Write new and modify existing python scripts that pull data from multiple sources and databases and convert to csv and populate to Teradata using the utilities

Perform UT on the Azure Databricks using the python spark sql and test SCD1 and SCD2

Debug, support the existing on-prem jobs and ETL built on Datastage

Run Datastage jobs, debug code in datastage and convert transformation to databricks and Teradata

Create new tables and views in Teradata for various business streams adhering to Data model

Responsible to analyze the existing data and the data flow in EDW that holds the clients enterprise data built on Teradata 16

Maintain the jobs schedule in Atomic, monitor daily jobs and provide support and maintain during the schedule maintenance

Perform GAP analysis between the middleware application and the database sources built on DB2 and sourced from Oracle and Teradata

Analyze data present outside of EDW built on databases like Teradata, Oracle PL SQL, SQL Server etc

Utilize Confluence for uploading documentation to be shared across the team

Coding version control is uploaded to Github, create branches and peers in github

Write and modify the ETL scripts in databricks, perform scheduling and monitoring in ADLS

Modify the existing and create new python scripts for file handling

Create and track projects in Jira, write user stories and attend scrum meetings.

Attend Business meetings to gather requirements, create a design strategy, convert functional requirements to technical design and write design documents

Write complex Teradata and SQL queries that include various join strategies to arrive at the design, create Unix shell scripts to handle files created on .csv, xml etc

Perform Data analysis before creating the design and perform data cleansing of files

Create source to target mappings before coding by understanding the Physical and Logical data model working along with Data modeler

Worked as a data modeler to bring in venue details using Erwin data model

Designed and implemented data models for various Curam modules, including Participant Management, Case Management, and Eligibility Determination for the feeds on voluntary and social activities.

Developed custom data structures and relationships to meet client-specific requirements and improve data integrity.

Lead migration activity analysis for Teradata to Cloud

Senior ETL Analyst Aug 2019 – Nov 2020

Compunnel Software Group Inc - Plano, TX

Responsibilities:

Responsible to analyze the existing data and the data flow in EDW that holds the clients enterprise data built on Teradata 16

Perform GAP analysis between the middleware application and the database sources built on DB2 and sourced from Oracle and Teradata

Analyze data present outside of EDW built on databases like Teradata, Oracle PL SQL, SQL Server etc that uses code in Mainframes, UNIX etc

And provide ETL solutions using Oracle, Teradata, Datastage and Informatica Powercenter 10.1 to bring in to EDW

Utilize Confluence for uploading documentation to be shared across the team

Maintain spaces for knowledge and space documentation

Coding version control is uploaded to Github, create branches and peers in github

Modify the existing and create new python scripts for file handling

Utilize business intelligence tools Tableau for retails data reporting

Attend Business meetings to gather requirements, create a design strategy, convert functional requirements to technical design and write design documents

Lead team on design and development and support until warranty

Write Stored procedures and Triggers using T-SQL in SQL server, PL-SQL in Oracle, create packages to perform ETL

Create error handling workflows and mappings to generate email alerts

Write complex Teradata and SQL queries that include various join strategies to arrive at the design, create Unix shell scripts to handle files created on .csv, xml etc

Write Teradata components using utilities BTEQ, TPT, MLOAD, FLoad for the migration activity based on design

Utilize Teradata TASM for maintaining workload management, performance monitoring and increase efficiency of performance

Perform data and object migration between systems using Data mover in TeradataWrite Teradata Stored procedures and Triggers for performing CDC on each tables

Create Informatica mapping, workflows, using various transformations for performing ETL aspects of the data migration

Utilize Erwin data model to make changes and include new tables into the model

Performance tuning on Informatica mappings and workflows in powercenter, utilize Push down optimization

Implement performance tuning and write advance SQL and PL-SQL queries in finding out the equivalent Physical columns within EDW for the corresponding business termsCreate workflow and jobs in Workato, to access data in web using rest API and populate them in databases

Use SQL Server Management Studio to create tables in SQL server and populate data from APIs and create ETL jobs using SSIS

Environment : Teradata, Oracle, SSIS, Informatica, Unix, SQL server, Workato, Confluence, Github,Python

Senior ETL Programmer Analyst Teradata, Oct 2017 – Aug 2019

Systems Technology Group, Inc, Ford Motor Credit Company - Dearborn, MI

Responsibilities:

Worked in Data Integration and migration projects within different databases of Teradata by utilizing the ETL Capabilities

Validate the data compatibility between systems across various databases

Create designs in Teradata and ETL to migrate new data into the system

Responsible to work in Enterprise data warehouse (EDW) and Analyze the jobs created in Mainframes, Teradata and other (Extract transform Load) ETL utilities like Datastage used in the client environment that access the EDW to bring new data and create reports on existing data

Perform migration activity on the components using (Job Control Language) JCLS, Parmcards, BTEQ, Fastload, MLoad

Understand the Logical and Physical data model, relationship of tables and Create mapping documents based on the analysis on existing jobs to create new components

Perform Proof of concepts on converting SQLs into Datastage stages for enabling lineage

Attend Business meetings, create a design strategy and convert functional requirements to design

Interpret the various business terms and identify/mine their equivalent technical terms inside the data warehouse built on Teradata

Write various PL-SQL queries to perform updating the existing data based on the business requirements

Implement performance tuning and write advance SQL queries in finding out the equivalent physical columns for the corresponding business terms

Work with DBAs to create DDLs, ensure proper standards are setup, Skewness is avoided in queriesTweak underperforming queries and work with DBA to setup spaces

Create Datastage components utilizing various stages of IBM Datastage

Perform Proof of Concept in converting complex SQL codes into Datastage mappings

Environment : Teradata, Mainframes, Datastage, Autosys

Cognizant Technology Solutions, PEPSI CO - Plano, TX Jan 2016 - Sep 2017

ETL Developer

Responsibilities:

Working as a Data Analyst and Developer for a retail domain, responsible for gathering requirements, develop SQL code, Unit testing

Understand the data model, write source to Target mappings for the ETL process

Perform data quality checks based on the test data

Create Mappings, workflows in Informatica power center 9.5.1, 10.1 and Cloud perform Publication, Subscriptions in cloud and usage of different transformations or create packages in SSIS and create ETL jobs

Write Shell scripts in handling csv, xml, txt files in Unix, use SED,AWK commands in file handling and writing trigger scripts

Involve in performance tuning of T-SQL Queries, creating mapping source to target mapping sheets, handling stats on the tables

Generate reports by writing ad-hoc T-SQLs

Responsible for creating and doing unit test, ensure performance tuning using Informatica Power Center 9.5.1 and 10.1

Develop new components in Informatica Data Integration Hub (DIH)

Teradata Utilities Fast load, Mload, Bteq, Tpump and responsible for Performance tuning

Worked in Oracle 11g, using load utilities like SQLPLUS, SQLLDR

Environment: Oracle, Teradata, SQL Server, SSIS, Informatica Power center 9.5.1 and 10.1, Data Integration Hub (DIH), SQL Assistant, UNIX, Oracle, Db visualizer.

Wipro Technologies Ltd, CVS Care Mark - Richardson, Dallas, TX Feb 2015 – Jan 2016

Programmer Analyst, ETL Analyst

Responsibilities:

As an ETL Lead responsible to gather requirements, Design and assist develop and review the code on the health care projects in Teradata 13, SQL Server with Informatica Power center 9.5.1 and SSIS as ETL on UNIX operating system

Responsible for creating design documents, source to target mappings as per functional requirements and presentation to client side Leads and SMEs for approvals

Perform release management process in Serena

Good knowledge on Health care subject areas CLIENT, MEMBER, ACCOUNT, GROUP, DRUG, PRESCRIBER

Create T- SQLs, Tables is SQL server and Teradata components BTEQ, MLOAD, FLOAD and Informatica workflows, sessions and mapping based on data model and requirements

Environment: Teradata, SQL Server, SSIS, Informatica Power center 9.5.1, SQL Assistant, UNIX, Oracle, Design and documentation

Wipro Technologies Ltd, Lloyds Banking Group - Manchester, UK Jan 2011 – Jan 2015

Subject Matter Expert

Responsibilities:

For Bank's data warehouse and Teradata lead, Responsible to successfully handle and deliver the data migration and integration activity of two major banks and also handle SDLC projects on Retail portfolio

In depth knowledge in various databases that holds each type of data in the warehouse

Working on Teradata 12, Mainframes, ETL and FSLDM

The core database of the ware house was built on Teradata FSLDM, Expertise in FSLDM on subject areas- Agreement, Party, Feature, Event, Product

Knowledge on using different type of Indexes like Join Index, Multi Table JI, Single table JI

Environment: Teradata, TD Utilities, ETL, Mapping, SQL Assistant, UNIX, Mainframes zOS, JCLs, Connect Direct.

Infosys Technologies LTD, Bank of America, Chennai, India Jul 2007 – Dec 2010

Senior Systems Analyst – Teradata Developer

Responsibilities:

Responsible to build Teradata components like BTEQ, MLOAD, FLOAD in Mainframes

Build Mainframe components JCL, PROCS, parmcards

Environment: Teradata, TD Utilities, Oracle –PL SQLs, ETL, Mapping, SQL Assistant, UNIX, Mainframes zOS, Change man, File Aid, JCLs, CA7, Data Migration and System migration.



Contact this candidate