DIVYA VALLAPU
Location: CA, USA Phone: 510-***-**** Email: *************@*****.*** LinkedIn
SUMMARY
• Site Reliability Engineer with 5+ years of experience building and managing scalable cloud-native infrastructure across healthcare and enterprise environments using Kubernetes, Amazon EKS, Docker, Terraform, and GitOps practices.
• Expertise in CI/CD automation and deployment reliability using Argo CD, Jenkins, GitHub Actions, and Helm, enabling efficient and scalable software delivery.
• Strong background in observability, monitoring, and incident response using Prometheus, Grafana, Datadog, OpenTelemetry, ELK Stack, and Loki to improve system reliability and performance.
• Proficient in infrastructure automation with Python, Bash, and Ansible, supporting highly available distributed systems built on Kafka, RabbitMQ, Redis, PostgreSQL, and microservices architectures.
• Experienced in cloud security and reliability engineering, including IAM, OAuth2, secrets management, vulnerability scanning, SLI/SLO implementation, root cause analysis (RCA), postmortems, and production stability improvements.
WORK EXPERIENCE
CVS Health Site Reliability Engineer Jul 2024 – Present
• Spearheaded the Healthcare Cloud Reliability Modernization initiative by designing resilient Kubernetes and Amazon EKS environments to support scalable pharmacy and patient-care applications across multi-cloud infrastructure.
• Engineered GitOps-driven CI/CD workflows using Argo CD, Terraform, GitHub Actions, and Helm, reducing deployment turnaround time by 38% across production services.
• Optimized observability, AIOps-driven monitoring, and incident management processes by implementing Prometheus, Grafana, Open Telemetry, Loki, and Datadog dashboards with anomaly detection and intelligent alerting for proactive monitoring and distributed tracing.
• Automated infrastructure operations, self-healing operational workflows, and configuration management using Python, Ansible, and Bash scripting, improving operational efficiency and lowering recurring manual support tasks by 34%.
• Strengthened platform security and API reliability by integrating IAM policies, OAuth2 authentication, secrets management, vulnerability scanning, and load-balancing strategies across cloud-native services.
• Orchestrated highly available distributed systems using Kafka, RabbitMQ, Redis, PostgreSQL, and FastAPI-based microservices while supporting AIOps-based root cause analysis (RCA) and incident remediation, contributing to a 29% improvement in production system stability during peak healthcare transaction workloads. VMware Site Reliability Engineer Oct 2019 – Dec 2022
• Contributed to the Cloud Infrastructure Stability Automation initiative by supporting Kubernetes-based application environments, monitoring Linux servers, and assisting with deployment reliability across internal enterprise platforms.
• Automated routine operational tasks using Python, Bash, and Jenkins pipelines, reducing manual deployment efforts by 32% across development and staging environments.
• Monitored distributed workloads using Prometheus, Grafana, and ELK Stack to identify performance bottlenecks and improve system visibility for engineering teams.
• Assisted in managing Docker containers, Helm configurations, and Kubernetes resource scaling to improve service availability during high-traffic production windows.
• Implemented Infrastructure as Code practices using Terraform and Git-based workflows, accelerating environment provisioning time by 28% for internal testing environments.
• Supported incident response activities by analyzing application logs, participating in RCA discussions, and documenting operational runbooks for recurring infrastructure issues.
• Integrated GitHub Actions and Argo CD pipelines for automated deployment validation, helping improve release consistency and lowering deployment rollback incidents by 21%.
• Strengthened platform reliability by configuring Redis caching, monitoring API health metrics, and supporting secure access management processes, contributing to a 26% improvement in application response stability. PROJECTS
CI/CD Pipeline Automation Framework
• Built automated deployment pipelines using Jenkins, GitHub Actions, Docker, Kubernetes, and GitOps workflows for application build, automated testing, continuous delivery, and production release management.
• Implemented Git-based branching, infrastructure deployment strategies, version control practices, and rollback management to support scalable continuous integration and cloud-native DevOps operations.
• Streamlined artifact validation, deployment verification, environment health checks, and release monitoring processes across distributed development, QA, and staging environments. API Reliability and Performance Monitoring
• Developed REST API health monitoring workflows using FastAPI, Python, Datadog integrations, and cloud-native observability tools for proactive service monitoring and incident detection.
• Implemented latency tracking, uptime monitoring, request validation, and performance analysis for backend service reliability, scalability, and high-availability distributed application environments.
• Assisted in debugging API failures, analyzing production incidents, and improving service observability across distributed systems using centralized logging and monitoring dashboards. SKILLS
• Programming & Scripting: Python, Bash, SQL, Go (Basics), Automation Scripting
• Cloud & Infrastructure: AWS, Azure, Linux, IAM, DNS, Networking Fundamentals, Load Balancing
• Containers & Orchestration: Docker, Kubernetes, Helm, Amazon EKS
• Infrastructure as Code & Configuration Management: Terraform, Ansible, GitOps
• CI/CD & Deployment Automation: Jenkins, GitHub Actions, Argo CD, Git
• AIOps & Intelligent Operations: AIOps, Event Correlation, Predictive Monitoring, Automated Incident Remediation, Log Analytics, Anomaly Detection, Root Cause Analysis (RCA), Alert Noise Reduction, Self-Healing Systems
• Monitoring & Observability: Prometheus, Grafana, OpenTelemetry, ELK Stack, Loki, Datadog
• Reliability & Distributed Systems: SLI/SLO/SLA, Incident Response, RCA, Postmortems, Kafka, RabbitMQ, Redis, High Availability, Scalability, Fault Tolerance
• Databases, APIs & Security: PostgreSQL, MySQL, MongoDB, REST APIs, FastAPI, OAuth2, JWT, Secrets Management, Vulnerability Scanning
EDUCATION
Master of Science in Computer Science Jan 2023 – May 2024 Florida Institute of Technology, Melbourne, FL