Post Job Free
Sign in

Kubernetes SRE with CI/CD and Observability

Location:
Sunnyvale, CA
Salary:
95000
Posted:
July 27, 2026

Contact this candidate

Resume:

DIVYA VALLAPU

Location: CA, USA Phone: 510-***-**** Email: *************@*****.*** LinkedIn

SUMMARY

• Site Reliability Engineer with 5+ years of experience building and managing scalable cloud-native infrastructure across healthcare and enterprise environments using Kubernetes, Amazon EKS, Docker, Terraform, and GitOps practices.

• Expertise in CI/CD automation and deployment reliability using Argo CD, Jenkins, GitHub Actions, and Helm, enabling efficient and scalable software delivery.

• Strong background in observability, monitoring, and incident response using Prometheus, Grafana, Datadog, OpenTelemetry, ELK Stack, and Loki to improve system reliability and performance.

• Proficient in infrastructure automation with Python, Bash, and Ansible, supporting highly available distributed systems built on Kafka, RabbitMQ, Redis, PostgreSQL, and microservices architectures.

• Experienced in cloud security and reliability engineering, including IAM, OAuth2, secrets management, vulnerability scanning, SLI/SLO implementation, root cause analysis (RCA), postmortems, and production stability improvements.

WORK EXPERIENCE

CVS Health Site Reliability Engineer Jul 2024 – Present

• Spearheaded the Healthcare Cloud Reliability Modernization initiative by designing resilient Kubernetes and Amazon EKS environments to support scalable pharmacy and patient-care applications across multi-cloud infrastructure.

• Engineered GitOps-driven CI/CD workflows using Argo CD, Terraform, GitHub Actions, and Helm, reducing deployment turnaround time by 38% across production services.

• Optimized observability, AIOps-driven monitoring, and incident management processes by implementing Prometheus, Grafana, Open Telemetry, Loki, and Datadog dashboards with anomaly detection and intelligent alerting for proactive monitoring and distributed tracing.

• Automated infrastructure operations, self-healing operational workflows, and configuration management using Python, Ansible, and Bash scripting, improving operational efficiency and lowering recurring manual support tasks by 34%.

• Strengthened platform security and API reliability by integrating IAM policies, OAuth2 authentication, secrets management, vulnerability scanning, and load-balancing strategies across cloud-native services.

• Orchestrated highly available distributed systems using Kafka, RabbitMQ, Redis, PostgreSQL, and FastAPI-based microservices while supporting AIOps-based root cause analysis (RCA) and incident remediation, contributing to a 29% improvement in production system stability during peak healthcare transaction workloads. VMware Site Reliability Engineer Oct 2019 – Dec 2022

• Contributed to the Cloud Infrastructure Stability Automation initiative by supporting Kubernetes-based application environments, monitoring Linux servers, and assisting with deployment reliability across internal enterprise platforms.

• Automated routine operational tasks using Python, Bash, and Jenkins pipelines, reducing manual deployment efforts by 32% across development and staging environments.

• Monitored distributed workloads using Prometheus, Grafana, and ELK Stack to identify performance bottlenecks and improve system visibility for engineering teams.

• Assisted in managing Docker containers, Helm configurations, and Kubernetes resource scaling to improve service availability during high-traffic production windows.

• Implemented Infrastructure as Code practices using Terraform and Git-based workflows, accelerating environment provisioning time by 28% for internal testing environments.

• Supported incident response activities by analyzing application logs, participating in RCA discussions, and documenting operational runbooks for recurring infrastructure issues.

• Integrated GitHub Actions and Argo CD pipelines for automated deployment validation, helping improve release consistency and lowering deployment rollback incidents by 21%.

• Strengthened platform reliability by configuring Redis caching, monitoring API health metrics, and supporting secure access management processes, contributing to a 26% improvement in application response stability. PROJECTS

CI/CD Pipeline Automation Framework

• Built automated deployment pipelines using Jenkins, GitHub Actions, Docker, Kubernetes, and GitOps workflows for application build, automated testing, continuous delivery, and production release management.

• Implemented Git-based branching, infrastructure deployment strategies, version control practices, and rollback management to support scalable continuous integration and cloud-native DevOps operations.

• Streamlined artifact validation, deployment verification, environment health checks, and release monitoring processes across distributed development, QA, and staging environments. API Reliability and Performance Monitoring

• Developed REST API health monitoring workflows using FastAPI, Python, Datadog integrations, and cloud-native observability tools for proactive service monitoring and incident detection.

• Implemented latency tracking, uptime monitoring, request validation, and performance analysis for backend service reliability, scalability, and high-availability distributed application environments.

• Assisted in debugging API failures, analyzing production incidents, and improving service observability across distributed systems using centralized logging and monitoring dashboards. SKILLS

• Programming & Scripting: Python, Bash, SQL, Go (Basics), Automation Scripting

• Cloud & Infrastructure: AWS, Azure, Linux, IAM, DNS, Networking Fundamentals, Load Balancing

• Containers & Orchestration: Docker, Kubernetes, Helm, Amazon EKS

• Infrastructure as Code & Configuration Management: Terraform, Ansible, GitOps

• CI/CD & Deployment Automation: Jenkins, GitHub Actions, Argo CD, Git

• AIOps & Intelligent Operations: AIOps, Event Correlation, Predictive Monitoring, Automated Incident Remediation, Log Analytics, Anomaly Detection, Root Cause Analysis (RCA), Alert Noise Reduction, Self-Healing Systems

• Monitoring & Observability: Prometheus, Grafana, OpenTelemetry, ELK Stack, Loki, Datadog

• Reliability & Distributed Systems: SLI/SLO/SLA, Incident Response, RCA, Postmortems, Kafka, RabbitMQ, Redis, High Availability, Scalability, Fault Tolerance

• Databases, APIs & Security: PostgreSQL, MySQL, MongoDB, REST APIs, FastAPI, OAuth2, JWT, Secrets Management, Vulnerability Scanning

EDUCATION

Master of Science in Computer Science Jan 2023 – May 2024 Florida Institute of Technology, Melbourne, FL



Contact this candidate