Post Job Free
Sign in

Senior Cloud Data Engineer

Location:
Cincinnati, OH
Posted:
July 21, 2026

Contact this candidate

Resume:

Data Engineer

Name: Vamshi P

Phone: 513-***-****

Email: ***************@*****.***

PROFESSIONAL SUMMARY:

●Data Engineer with 6+ years of experience designing scalable ETL/ELT pipelines using PySpark, SQL, Python, AWS, Databricks, and cloud-native data engineering technologies.

●Experienced in processing large-scale structured and unstructured datasets using Spark-based distributed computing frameworks.

●Skilled in building cloud-native data platforms, optimizing data workflows, implementing data quality checks, and supporting enterprise analytics solutions in healthcare and financial domains.

●Performed query optimization and debugging of SQL/PLSQL code to enhance database performance and ensure data integrity

●Skilled in building automated data workflows using AWS services such as Glue, Redshift, S3, Lambda, and Athena, ensuring efficient and reliable data processing.

●Adept at performing data quality checks, validation, and reconciliation to maintain accuracy and compliance across enterprise data systems.

●Experience working with PL/SQL packages for modular and reusable database development

●Experienced in creating interactive dashboards and visual analytics using Power BI, Tableau, and QuickSight to support data-driven decision-making.

●Collaborates effectively with business and technical teams to translate analytical requirements into actionable insights.

●Focused on optimizing data processes, improving performance, and ensuring seamless data flow for analytical and operational needs.

●Strong communicator with a proven track record of delivering high-quality solutions within Agile project frameworks.

TECHNICAL SKILLS:

Languages

Python, SQL, PL/SQL

Cloud & Tools

AWS (Redshift, S3, RDS, Glue, Athena, Lambda)

Data Analytics & ETL

Data Cleaning, Data Validation, Data Reconciliation, ETL Pipelines, Pandas, PySpark

Databases

Oracle, Redshift, PostgreSQL, MySQL

Big Data Technologies:

Apache Spark, Hadoop Ecosystem, Hive, Spark SQL

Transformation Tools

dbt (Data Build Tool), PySpark, Pandas

Visualization & BI

Power BI, Tableau, Excel (Pivot Tables, VLOOKUP, Macros)

Orchestration Tools:

Apache Airflow, Workflow Automation, Scheduling Pipelines

WORK EXPERIENCE:

Blue Cross Blue Shield Data Engineer

Remote Jul 2025 - Present

Responsibilities:

●Designed scalable ETL/ELT pipelines using PySpark and Databricks on AWS, processing structured and semi-structured datasets exceeding 5TB daily with optimized Spark transformations and partitioning strategies.

●Optimized SQL and PySpark-based transformation workflows, improving pipeline performance and reducing processing time through partitioning and query tuning.e.

●Ensured regulatory compliance and data governance by implementing robust data validation, auditing, and lineage tracking across enterprise data systems.

●Worked on distributed data processing using Spark ecosystem components and optimized ETL workflows for scalable cloud-based analytics platforms.

●Developed and optimized SQL and PL/SQL scripts, including stored procedures and data transformation logic, improving performance and data reliability.

●Implemented data quality checks and reconciliation workflows, ensuring 99% accuracy and consistency across financial and operational datasets.

●Developed Databricks notebooks using PySpark for large-scale ETL processing with Delta Lake tables enabling ACID compliant data pipelines.

●Automated data loading and processing tasks using Airflow and AWS Lambda, reducing manual intervention by 40%.

●Created Power BI and QuickSight dashboards to visualize KPIs and operational metrics, improving transparency for business stakeholders.

●Collaborated with BI developers and data scientists to deliver clean, structured datasets for analysis and model development.

●Documented ETL processes, data dictionaries, and lineage for improved traceability and compliance.

●Monitored and optimized pipeline performance using AWS CloudWatch and Glue job metrics, ensuring timely data availability.

●Integrated healthcare data from REST APIs and cloud storage sources into Databricks ETL pipelines.

●Collaborated within Agile Scrum teams using Git-based version control and CI/CD deployment processes.

Environment: AWS (S3, Glue, Redshift, Athena, Lambda, CloudWatch), PySpark, Databricks, SQL, Python, Apache Airflow, Power BI, QuickSight, ETL/ELT Pipelines

NVIDIA – Data Analyst

Hyderabad, India Apr 2021 – Aug 2023

Responsibilities:

●Designed and maintained AWS-based ETL pipelines using Glue, Redshift, and S3 for structured and semi-structured datasets.

●Conducted large-scale data remediation and validation, achieving over 99% accuracy in analytical datasets.

●Built and optimized SQL, Python, and PL/SQL scripts, including stored procedures and data transformation logic, for reporting and data processing

●Created BI-ready datasets in Redshift to support dashboards, predictive analytics, and performance monitoring.

●Partnered with business stakeholders to define data quality and validation rules, reducing reporting inconsistencies by 35%.

●Developed and published Power BI and Tableau dashboards, delivering actionable insights to leadership.

●Automated recurring data ingestion and transformation workflows using Airflow and AWS Lambda, increasing efficiency by 30%.

●Monitored and troubleshot ETL jobs using AWS CloudWatch, proactively resolving data delivery issues.

●Supported migration from on-premise SQL systems to AWS cloud, improving scalability and reducing maintenance costs.

●Collaborated with QA and analytics teams to ensure high-quality datasets were available for machine learning experiments.

Environment: AWS (Redshift, S3, Glue, Athena, Lambda, CloudWatch), SQL, Python (Pandas), Power BI, Tableau, Airflow, Excel

Edarlabs Learning Pvt. Ltd. Data Analyst

India 2018 – 2021

Responsibilities:

●Automated ETL workflows using SQL and PL/SQL scripts, including functions and procedures, reducing manual data prep time by 60%

●Performed data audits, cleaning, and validation across structured datasets, improving reliability and compliance.

●Designed Excel and Power BI reports for tracking content performance, student engagement, and operational KPIs.

●Built and maintained data extraction scripts from APIs and internal databases for analytics dashboards.

●Collaborated with QA and product teams to identify data gaps and standardize reporting formats.

●Optimized SQL queries and indexing strategies to improve report generation speed by 25%.

●Developed data quality metrics dashboards for internal monitoring and decision-making.

●Documented workflows, metadata, and data dictionaries to support onboarding and team knowledge sharing.

●Partnered with content and technology teams to ensure accurate and consistent reporting across all EdTech platforms.

Environment: AWS (S3, Redshift), Python, SQL, Power BI, Excel, Jira

Education:

●MSIS in Trine University, GPA:3.8

●Bachelor of Technology in Electronics and Communication Engineering, GPA: 3.5/4.0

JNTU, Hyderabad, India



Contact this candidate