Sreeja K
Data Engineer
+1-913-***-**** *************@*****.*** www.linkedin.com/in/sreeja-k-099096222
PROFESSIONAL SUMMARY
Data Engineer with 6 years of experience designing, building, deploying, and supporting scalable cloud data platforms and production data solutions across GCP, BigQuery, Dataflow, Dataproc, Cloud Storage, Pub/Sub, Cloud Composer, Python, SQL, and Apache Spark. Experienced in translating ambiguous business and technical requirements into production-ready architectures, data pipelines, analytical systems, and reusable engineering frameworks, collaborating closely with architects, application teams, analysts, DevOps engineers, and business stakeholders. Hands-on experience with AI/LLM solutions, Generative AI, RAG-oriented data preparation, prompt engineering, and AI-assisted software development using tools including GitHub Copilot, Claude, ChatGPT, Claude Code, and Cursor. Strong background in end-to-end solution delivery, system design, production deployment, CI/CD, Docker, Kubernetes, cloud security, monitoring, troubleshooting, data governance, and reliability engineering. Experienced in building reusable Python utilities, automation frameworks, pipeline components, and infrastructure patterns that improve delivery consistency and operational efficiency.
TECHNICAL SKILLS
•Cloud Platform: Google Cloud Platform (GCP)
•GCP Data & AI Services: BigQuery, Cloud Dataflow, Dataproc, Google Cloud Storage (GCS), Pub/Sub, Cloud Composer, Cloud IAM, Cloud Monitoring, Cloud Logging
•Programming: Python, SQL, Scala, Bash, PySpark, Spark SQL
•AI / Generative AI: LLMs, Generative AI, AI Agents, RAG, Prompt Engineering, AI-assisted Engineering, GitHub Copilot, Claude, ChatGPT, Claude Code, Cursor
•Data Engineering: ETL/ELT, Data Pipelines, Data Integration, Batch Processing, Real-Time Processing, Incremental Loading, CDC, Data Transformation, Data Validation, Data Quality, Data Reconciliation
•Architecture & Delivery: Cloud Architecture, System Design, Technical Requirements, Solution Design, Production Deployment, Technical Documentation, Reusable Frameworks, Automation, Production Support
•Data Warehousing & Modeling: BigQuery, Data Warehousing, Dimensional Modeling, Data Modeling, Analytical Data Models, Fact and Dimension Tables, Data Marts
•Orchestration: Cloud Composer, Apache Airflow, Dataproc Workflows, GCP Workflows
•Data Integration: REST APIs, JDBC, SFTP, JSON, CSV, Parquet, Avro
•DevOps & Cloud-Native: Git, GitHub, GitHub Actions, Jenkins, Docker, Kubernetes, CI/CD, Infrastructure Automation
•Security & Governance: GCP IAM, Service Accounts, Cloud Security, Data Governance, Data Lineage, Access Controls, Encryption
•Monitoring & Operations: Cloud Monitoring, Cloud Logging, Error Handling, Alerting, Troubleshooting, Production Support, SLA Monitoring
PROFESSIONAL EXPERIENCE
UHG, Minneapolis MN
Sep 2024 - Present
Data Engineer
•Owned end-to-end delivery of GCP data and AI solutions, including discovery, technical scoping, architecture, development, deployment, production stabilization, and optimization.
•Partnered with customer engineering, Product, Analytics, Security, GRC, and business stakeholders to translate ambiguous requirements into scalable production solutions and drive adoption.
•Designed GCP data-platform architectures using BigQuery, Dataflow, Dataproc, GCS, Pub/Sub, Cloud Composer, Python, SQL, and Spark across data, analytics, and AI workflows.
•Built and supported scalable Snowflake/SQL ETL pipelines for high-volume transactional and operational datasets, incorporating data reconciliation, validation, stored procedures, regulatory-style reporting controls, access governance, and Tableau reporting to deliver accurate, audit-ready data for downstream business and compliance analytics.
•Developed Python backend services and APIs integrating data platforms, AI services, enterprise applications, and customer workflows with secure authentication, validation, monitoring, and error handling.
•Implemented RAG and evaluation pipelines measuring LLM accuracy, relevance, latency, reliability, and workflow impact to refine prompts, retrieval, and application behavior.
•Built scalable ETL/ELT and real-time pipelines integrating databases, REST APIs, applications, files, and event streams into BigQuery and AI/analytics platforms.
•Developed Cloud Composer/Airflow workflows supporting retries, idempotency, backfills, schema evolution, dependency management, failure recovery, and monitoring.
•Implemented Pub/Sub/Dataflow event-driven architectures handling duplicate events, late-arriving data, partial failures, and downstream synchronization.
•Optimized BigQuery/Dataflow workloads using partitioning, clustering, SQL optimization, resource tuning, and cost-management practices.
•Built reusable Python frameworks, technical playbooks, deployment patterns, and infrastructure components to standardize solution delivery.
•Implemented production CI/CD with GitHub Actions, Jenkins, Docker, Kubernetes, and infrastructure automation across development, testing, and production.
•Applied GCP IAM, service accounts, encryption, access controls, governance, lineage, and security controls for enterprise production environments.
•Communicated technical trade-offs, risks, and delivery decisions across Product, Research/AI, Engineering, Security, GRC, Architecture, and customer stakeholders, incorporating field feedback into solution and roadmap improvements.
•Collaborated with data architects, analysts, application developers, DevOps engineers, and business stakeholders to deliver scalable GCP data solutions.
PWC, Hyderabad India.
Jan 2022-Nov 2023
Data Engineer
•Designed and developed scalable GCP-based data pipelines for ingesting, transforming, and processing structured and semi-structured enterprise data.
•Developed Python/SQL ETL and reporting workflows integrating relational database sources including DB2, Oracle, SQL Server, and MySQL, supporting trade/transaction data onboarding, migration, reconciliation, stored procedures, reporting datasets, and Business Objects/Tableau-style enterprise reporting within Agile/DevOps and CI/CD environments.
•Developed BigQuery ETL/ELT workflows using Python and SQL for large-scale analytical processing.
•Built data ingestion pipelines using Cloud Storage, Cloud Dataflow, Dataproc, Pub/Sub, and Cloud Composer.
•Developed distributed data processing applications using Apache Spark, PySpark, and Dataproc for high-volume batch processing.
•Developed complex SQL queries and Python programs for data transformation, cleansing, standardization, enrichment, aggregation, and validation.
•Integrated data from SQL Server, Oracle, MySQL, APIs, flat files, JSON, CSV, and SFTP into GCP data platforms.
•Designed and implemented BigQuery data models supporting reporting, analytics, and downstream business intelligence applications.
•Implemented batch, incremental, and real-time processing patterns using Dataflow, Dataproc, Pub/Sub, and Airflow.
•Developed and maintained Cloud Composer/Airflow DAGs for scheduling, orchestration, dependency management, retries, monitoring, and recovery.
•Optimized BigQuery queries and data pipelines using partitioning, clustering, efficient joins, predicate filtering, and query optimization techniques.
•Implemented data quality processes covering completeness, accuracy, consistency, duplicates, null values, referential integrity, and source-to-target reconciliation.
•Implemented Cloud Monitoring and Cloud Logging for pipeline monitoring, error detection, alerting, and production troubleshooting.
•Applied GCP security practices using IAM, service accounts, access controls, and secure data-access patterns.
•Developed automated CI/CD workflows using Git, GitHub, Docker, and GitHub Actions.
•Participated in Agile/Scrum development, code reviews, unit testing, integration testing, deployment, and production support.
•Collaborated with data architects, analysts, application teams, and DevOps engineers to translate business requirements into scalable GCP data solutions.
•Created technical documentation covering data models, source-to-target mappings, pipeline dependencies, transformation logic, and operational procedures.
Hudda Infotech Private Limited Hyd India
Apr 2020-Dec 2021
Data Engineer
•Developed and maintained cloud-based data pipelines for structured and semi-structured data from databases, APIs, files, and enterprise applications.
•Built ETL workflows using Python, SQL, Apache Spark, and PySpark for extraction, transformation, validation, and loading.
•Developed GCP ingestion and processing workflows using BigQuery, Cloud Storage, Dataflow, Dataproc, and Pub/Sub.
•Created BigQuery tables, views, datasets, and analytical data models for reporting and downstream analytics.
•Developed Python and SQL transformations for filtering, joins, aggregations, deduplication, data cleansing, and business-rule validation.
•Supported batch and real-time data processing using Dataflow, Dataproc, Spark, and Pub/Sub.
•Developed Airflow/Cloud Composer workflows for pipeline scheduling, orchestration, dependency management, and operational monitoring.
•Integrated data from SQL Server, Oracle, MySQL, REST APIs, CSV, JSON, and SFTP sources.
•Implemented incremental processing using timestamps, watermarks, and change-data processing patterns.
•Performed data validation and reconciliation between source and target systems.
•Optimized Spark and SQL workloads through partitioning, efficient joins, filtering, and distributed processing techniques.
•Supported monitoring, troubleshooting, error handling, and production support for scheduled data pipelines.
•Used Git for source-code management and collaborative development.
•Participated in testing, deployment, code reviews, and Agile development activities.
•Supported cloud data migration and modernization initiatives involving relational and non-relational data sources.
KEY ACHIEVEMENTS
•Architected 15TB+ enterprise data lakes supporting 500K+ daily transactions
•Built ETL/ELT pipelines processing 8M+ records daily with 99.9% reliability
•Reduced data processing time by 75% (6 hours to 90 minutes)
•Improved SLA compliance to 99.5% through real-time streaming architectures
•Deployed 100+ automated DAGs and 50+ reusable infrastructure templates
•Achieved 87% accuracy in predictive ML models for customer churn
EDUCATION
Master’s degree: M.S. in Computer Science from University of Central Missouri.
Bachelor’s Degree: B.Tech in Information Technology from Bhoj Reddy Engineering College for Women.