Data Engineer
Name: Vamshi P
Phone: 513-***-****
Email: ***************@*****.***
PROFESSIONAL SUMMARY:
●Data Engineer with 6+ years of experience designing scalable ETL/ELT pipelines using PySpark, SQL, Python, AWS, Databricks, and cloud-native data engineering technologies.
●Experienced in processing large-scale structured and unstructured datasets using Spark-based distributed computing frameworks.
●Skilled in building cloud-native data platforms, optimizing data workflows, implementing data quality checks, and supporting enterprise analytics solutions in healthcare and financial domains.
●Performed query optimization and debugging of SQL/PLSQL code to enhance database performance and ensure data integrity
●Skilled in building automated data workflows using AWS services such as Glue, Redshift, S3, Lambda, and Athena, ensuring efficient and reliable data processing.
●Adept at performing data quality checks, validation, and reconciliation to maintain accuracy and compliance across enterprise data systems.
●Experience working with PL/SQL packages for modular and reusable database development
●Experienced in creating interactive dashboards and visual analytics using Power BI, Tableau, and QuickSight to support data-driven decision-making.
●Collaborates effectively with business and technical teams to translate analytical requirements into actionable insights.
●Focused on optimizing data processes, improving performance, and ensuring seamless data flow for analytical and operational needs.
●Strong communicator with a proven track record of delivering high-quality solutions within Agile project frameworks.
TECHNICAL SKILLS:
Languages
Python, SQL, PL/SQL
Cloud & Tools
AWS (Redshift, S3, RDS, Glue, Athena, Lambda)
Data Analytics & ETL
Data Cleaning, Data Validation, Data Reconciliation, ETL Pipelines, Pandas, PySpark
Databases
Oracle, Redshift, PostgreSQL, MySQL
Big Data Technologies:
Apache Spark, Hadoop Ecosystem, Hive, Spark SQL
Transformation Tools
dbt (Data Build Tool), PySpark, Pandas
Visualization & BI
Power BI, Tableau, Excel (Pivot Tables, VLOOKUP, Macros)
Orchestration Tools:
Apache Airflow, Workflow Automation, Scheduling Pipelines
WORK EXPERIENCE:
Blue Cross Blue Shield Data Engineer
Remote Jul 2025 - Present
Responsibilities:
●Designed scalable ETL/ELT pipelines using PySpark and Databricks on AWS, processing structured and semi-structured datasets exceeding 5TB daily with optimized Spark transformations and partitioning strategies.
●Optimized SQL and PySpark-based transformation workflows, improving pipeline performance and reducing processing time through partitioning and query tuning.e.
●Ensured regulatory compliance and data governance by implementing robust data validation, auditing, and lineage tracking across enterprise data systems.
●Worked on distributed data processing using Spark ecosystem components and optimized ETL workflows for scalable cloud-based analytics platforms.
●Developed and optimized SQL and PL/SQL scripts, including stored procedures and data transformation logic, improving performance and data reliability.
●Implemented data quality checks and reconciliation workflows, ensuring 99% accuracy and consistency across financial and operational datasets.
●Developed Databricks notebooks using PySpark for large-scale ETL processing with Delta Lake tables enabling ACID compliant data pipelines.
●Automated data loading and processing tasks using Airflow and AWS Lambda, reducing manual intervention by 40%.
●Created Power BI and QuickSight dashboards to visualize KPIs and operational metrics, improving transparency for business stakeholders.
●Collaborated with BI developers and data scientists to deliver clean, structured datasets for analysis and model development.
●Documented ETL processes, data dictionaries, and lineage for improved traceability and compliance.
●Monitored and optimized pipeline performance using AWS CloudWatch and Glue job metrics, ensuring timely data availability.
●Integrated healthcare data from REST APIs and cloud storage sources into Databricks ETL pipelines.
●Collaborated within Agile Scrum teams using Git-based version control and CI/CD deployment processes.
Environment: AWS (S3, Glue, Redshift, Athena, Lambda, CloudWatch), PySpark, Databricks, SQL, Python, Apache Airflow, Power BI, QuickSight, ETL/ELT Pipelines
NVIDIA – Data Analyst
Hyderabad, India Apr 2021 – Aug 2023
Responsibilities:
●Designed and maintained AWS-based ETL pipelines using Glue, Redshift, and S3 for structured and semi-structured datasets.
●Conducted large-scale data remediation and validation, achieving over 99% accuracy in analytical datasets.
●Built and optimized SQL, Python, and PL/SQL scripts, including stored procedures and data transformation logic, for reporting and data processing
●Created BI-ready datasets in Redshift to support dashboards, predictive analytics, and performance monitoring.
●Partnered with business stakeholders to define data quality and validation rules, reducing reporting inconsistencies by 35%.
●Developed and published Power BI and Tableau dashboards, delivering actionable insights to leadership.
●Automated recurring data ingestion and transformation workflows using Airflow and AWS Lambda, increasing efficiency by 30%.
●Monitored and troubleshot ETL jobs using AWS CloudWatch, proactively resolving data delivery issues.
●Supported migration from on-premise SQL systems to AWS cloud, improving scalability and reducing maintenance costs.
●Collaborated with QA and analytics teams to ensure high-quality datasets were available for machine learning experiments.
Environment: AWS (Redshift, S3, Glue, Athena, Lambda, CloudWatch), SQL, Python (Pandas), Power BI, Tableau, Airflow, Excel
Edarlabs Learning Pvt. Ltd. Data Analyst
India 2018 – 2021
Responsibilities:
●Automated ETL workflows using SQL and PL/SQL scripts, including functions and procedures, reducing manual data prep time by 60%
●Performed data audits, cleaning, and validation across structured datasets, improving reliability and compliance.
●Designed Excel and Power BI reports for tracking content performance, student engagement, and operational KPIs.
●Built and maintained data extraction scripts from APIs and internal databases for analytics dashboards.
●Collaborated with QA and product teams to identify data gaps and standardize reporting formats.
●Optimized SQL queries and indexing strategies to improve report generation speed by 25%.
●Developed data quality metrics dashboards for internal monitoring and decision-making.
●Documented workflows, metadata, and data dictionaries to support onboarding and team knowledge sharing.
●Partnered with content and technology teams to ensure accurate and consistent reporting across all EdTech platforms.
Environment: AWS (S3, Redshift), Python, SQL, Power BI, Excel, Jira
Education:
●MSIS in Trine University, GPA:3.8
●Bachelor of Technology in Electronics and Communication Engineering, GPA: 3.5/4.0
JNTU, Hyderabad, India