MANASA PELLURU
United States • 210-***-**** • *******@************.*** • LinkedIn • Github
SUMMARY Data Analytics Engineer with over 4 years of experience building automated data pipelines, scalable cloud warehouses, and enterprise business intelligence models. Expert at using dbt, Python, and advanced SQL to transform raw engineering data into clean, production-ready analytics. Proven track record of optimizing big data infrastructure via PySpark and Airflow while delivering high-impact dashboards in Tableau and Power BI. Skilled at establishing robust data quality frameworks that eliminate operational errors, unify corporate metrics, and drive measurable business growth. SKILLS Data Transformation, Modeling & Scripting: dbt, Advanced SQL (Complex CTEs, Window Functions, Query Performance Tuning), Python (Pandas, NumPy, Scipy), R (Statistical Computing). Cloud Data Warehousing & Architecture: Snowflake & Databricks, Google BigQuery AWS Redshift, Dimensional Modeling (Star and Snowflake Schemas), Data Lakehouse Design (Delta Lake). Data Ingestion, Pipelines & Orchestration: Pipeline Automation, Apache Airflow, Automated Ingestion, REST APIs, Python Scripting, Change Data Capture (CDC).
Distributed Processing & Big Data Infrastructur: Apache Spark, PySpark, Distributed Computing, Data Streaming Frameworks (via Apache Kafka), Memory Management Optimization. Business Intelligence & Data Storytelling: Interactive Dashboard, Tableau, Power BI, Microsoft Fabric, Looker, Semantic Layer Design, Unified Corporate Metrics Configuration. Data Governance, Quality & DevOps: Version Control, CI/CD Pipelines, Git, GitHub Actions, Data Quality Validation, Schema Evolution Management, Data Lineage Tracking. EXPERIENCE Avnet, United States Data Analytics Engineer October 2025 – Present
• Architected automated end-to-end data pipelines using Apache Airflow and Fivetran to ingest raw data into Snowflake, eliminating 120+ hours of monthly manual data operations and accelerating reporting delivery by 72 hours.
• Engineered modular data transformation models using dbt (Data Build Tool) and advanced SQL (CTEs, Window Functions), optimizing star schemas to slash cloud database compute costs by $115,000 annually and improve query times by 35%.
• Constructed interactive dashboards in Tableau and Power BI for 3 key business units, designing a unified semantic layer that standardized corporate metrics and reduced executive report generation time by 60%.
• Optimized high-throughput data processing workflows using PySpark running on Databricks clusters, successfully scaling the infrastructure to process over 500M+ rows of daily log data with zero system latency.
• Implemented automated data validation frameworks using Great Expectations and dbt-tests within a CI/CD pipeline via GitHub Actions, reducing downstream reporting anomalies and dashboard bugs by 42%. Adons Softech, India Data Analyst June 2021 - December 2023
• Designed and maintained cloud storage architectures across Google BigQuery and AWS Redshift to support global operations, migrating 15 terabytes of legacy data and improving enterprise data accessibility across departments by 50%.
• Developed custom automated extraction workflows via Python scripting and REST APIs to fetch high-frequency financial data, increasing data ingestion reliability by 28% across critical business lines.
• Designed a unified corporate metrics matrix within Microsoft Fabric and Looker, centralizing business logic for over 150 corporate users and cutting down cross-departmental reporting discrepancies by 80%.
• Partnered directly with product and finance operations teams to parse user retention trends and revenue leaks using Pandas and NumPy, which successfully uncovered system drop-off points and recovered $45,000 in operational cost leakages.
• Instituted schema evolution management and end-to-end data lineage tracking using Git, ensuring complete data governance transparency and achieving a 95% compliance rating during internal system data audits. EDUCATION Master of Science in Computer Science January 2024 – December 2025 Northern Arizona University Flagstaff, AZ
PROJECTS Cloud Data Platform Migration & Clickstream Warehouse Architecture
• Engineered a multi source data ingestion pipeline using Apache Airflow, Apache Spark, and Snowflake to aggregate over 100 million daily raw unstructured clickstream assets into structured lakehouse environments.
• Formulated optimized data transformation models and advanced dimensional schemas within Snowflake, leveraging complex SQL optimization and staging workflows to increase analytical querying speed by 37% across enterprise reporting layers. Real Time Financial Ingestion & Multi Channel Analytics Framework
• Constructed an event driven data processing architecture leveraging Apache Kafka, PySpark, PostgreSQL, and Databricks, enabling continuous streaming workflows that accelerated actionable data asset availability by 43% for reporting teams.
• Developed a comprehensive data aggregation layer within Databricks incorporating schema evolution enforcement and robust statistical validation methodologies to increase reporting forecast precision by 29% across critical business lines.