Post Job Free
Sign in

Data Engineer - PySpark, AWS, Kafka

Location:
Grand Prairie, TX
Posted:
September 19, 2026

Contact this candidate

Resume:

SUMMARY

Lavanya Vemireddy

********@***********.*** +1-940-***-**** TX, USA LinkedIn Data Engineer around 3 years of experience building cloud-based pipelines and analytics platforms across SaaS and telecommunications environments. Applies PySpark, AWS Glue, Airflow, Kafka, and Redshift to improve lifecycle-data availability and reporting reliability, while using Spark, Snowflake, dbt, and PostgreSQL to modernize high-volume usage and billing workflows. Experienced in validating transformations, managing schema changes, automating deployments, and strengthening data quality across distributed development and production environments. TECHNICAL SKILLS

Programming & Querying: Python, SQL, Scala, Bash, PySpark, Spark SQL Data Engineering & Processing: ETL, ELT, Batch Processing, Stream Processing, Apache Spark, Data Transformation, Data Pipelines, Change Data Capture

Databases & Data Warehousing: Snowflake (Snowpipe, Streams, Tasks), PostgreSQL, Amazon Redshift, Data Warehousing, Dimensional Modeling, Star Schema, Data Lakehouse

Cloud & Big Data Platforms: AWS (S3, Glue, EMR, Lambda), Databricks (Delta Lake, Unity Catalog), Apache Hadoop, Apache Hive Orchestration & Data Integration: Apache Airflow, Apache Kafka, dbt, REST APIs, Workflow Automation, Schema Evolution, API Integration

DevOps, Quality & Governance: Docker, Kubernetes, Terraform, Git, CI/CD, Data Quality, Data Validation, Data Lineage, Data Cataloging, Monitoring, Logging

PROFESSIONAL EXPERIENCE

Data Engineer, HighLevel Jan 2026 - Present Remote, USA

• Modeled unified customer lifecycle datasets from CRM contacts, conversations, opportunities, appointments, and subscription events, defining dimensional structures for consistent product, customer success, and revenue analytics.

• Developed PySpark pipelines using AWS Glue and AWS S3, reducing daily processing time by 24% while standardizing tenant-level transformations across expanding multichannel SaaS customer datasets.

• Configured Apache Airflow workflows for ingestion, dependency validation, retries, and Amazon Redshift loading, supporting dependable daily refreshes for product, finance, customer success, and operations teams.

• Integrated Apache Kafka event streams with Python consumers, improving lead and conversation data availability by 19% for near-real-time funnel monitoring across customer-facing CRM platform modules.

• Automated Terraform deployments and CI/CD validations for data infrastructure, transformation configurations, and environment changes, maintaining consistent releases during frequent updates to customer-facing SaaS platform capabilities.

• Monitored pipeline freshness, schema evolution, data quality, and processing failures through centralized logging, reducing issue- resolution time by 16% before downstream metrics and reports were affected. Data Engineer, Tech Mahindra Nov 2022 - Jul2024 Remote, India

• Evaluated telecom source tables, Apache Hive datasets, and interface specifications, documenting transformation rules, field mappings, dependencies, and reconciliation requirements for a subscriber usage data modernization engagement.

• Built Scala and Spark SQL processing jobs, improving batch throughput by 21% while reducing recurring transformation failures by 13% across high-volume subscriber usage and billing datasets.

• Configured AWS S3 landing zones and Snowpipe ingestion into Snowflake, enabling incremental processing through Streams and Tasks for continuously arriving network usage records and reference files.

• Applied dbt models and reusable SQL tests across subscriber, usage, and billing marts, strengthening transformation consistency before client reconciliation reviews and scheduled downstream reporting cycles.

• Validated migrated Snowflake datasets against PostgreSQL control totals and source extracts, shortening acceptance turnaround by 12% through automated completeness, duplication, referential integrity, and transformation checks.

• Containerized Python validation utilities with Docker and coordinated Git-based releases, reducing deployment preparation effort by 17% across development and testing environments used by distributed delivery teams. Data Analyst (Intern), TCS May 2022 - Oct2022 Remote, India

• Queried customer, transaction, and operational datasets using SQL to identify missing records, duplicate entries, and inconsistent field values affecting scheduled business reports.

• Developed Python scripts for data cleansing, format standardization, and validation, reducing repetitive manual preparation effort by 15% across recurring analytical datasets.

• Supported ETL workflows by reviewing source-to-target mappings, validating transformed PostgreSQL tables, and reconciling record counts before datasets were released for downstream reporting.

• Documented data-quality issues, transformation rules, and validation results while collaborating with senior analysts to resolve discrepancies across development and testing environments. CERTIFICATIONS

• AWS Academy Graduate - Cloud Foundations - Training Badge By AWS (Link) EDUCATION

Master of Science in Computer Science, University of North Texas, Denton Aug 2024 - May 2026 TX, USA Bachelor of Technology, Lakireddy Bali Reddy College of Engineering Jan 2021 - May 2024 AP, India



Contact this candidate