Vyshnavi Site Reliability Engineer
Chesterfield, MO 63005 ************@*****.*** +1-314-***-**** LinkedIn PROFESSIONAL SUMMARY
• Site Reliability Engineer with 5 years of experience leveraging Python, PowerShell, and Bash to automate operational tasks, reduce manual intervention and enhance deployment pipelines across hybrid cloud environments, leveraging multi- cloud expertise with AWS (EC2, S3, CloudWatch, IAM) and GCP (Compute Engine, GKE, Cloud Logging).
• Applied ITSM / ITIL practices such as change management, incident tracking, and reduced ticket resolution delays by integrating alerts with Slack and Microsoft Teams.
• Specialized in reducing toil through Python- and Terraform-based automation, implementing SLO-driven reliability frameworks, and optimizing CI/CD pipelines for zero-downtime deployments.
• Proficient in incident response automation, observability stack design, and reliability analytics (SLOs, SLIs, error budgets) to proactively improve service health.
• Experienced in chaos engineering, failover readiness testing, and incident postmortem analysis to strengthen production resilience and recovery times (MTTR, MTTD).
• Skilled in managing Kubernetes clusters (AKS) with Helm, RBAC policies, and advanced troubleshooting; containerized many apps using Docker and maintained rolling updates with minimal disruption.
• Managed and optimized Kubernetes clusters across AWS, Azure, and GCP, ensuring high availability, scalability, and seamless deployment of production workloads in multicloud environments.
• Established centralized logging and observability stacks using Prometheus, Grafana, and Azure Monitor, tuning alerts and dashboards for proactive incident detection and RCA for various critical incidents.
• Contributed to continuous improvement initiatives by analyzing system metrics, implementing automation, and applying SRE best practices to enhance overall system performance.
• Administered Linux (Ubuntu, RHEL) and Windows Server environments, performing patching, log rotation, and system tuning; built automation for image hardening and baseline configurations. TECHNICAL SKILLS
Reliability Engineering
SLOs, SLIs, SLAs, Error Budgets, Toil Reduction, Reliability Reviews, Postmortem Analysis, Incident Automation
Observability & Monitoring
Prometheus, Grafana, ELK, Splunk, Kibana, Datadog, Azure Monitor, CloudWatch, Stackdriver
(GCP), Application Insights, OpenTelemetry
Chaos & Resilience Testing Gremlin, LitmusChaos, Load Testing (k6, Locust) Infrastructure as Code
(IaC)
Terraform, Ansible, ARM Templates
Cloud Platforms AWS (EC2, S3, Lambda), Azure (AKS, App Services, Functions), GCP (GKE, Compute Engine) CI/CD & DevOps
Jenkins, GitHub Actions, Azure DevOps, Git, Deployment Strategies (Blue-Green, Canary, Rolling), Quality Gates
Languages & Scripting Python, PowerShell, Bash, YAML, JSON Containerization &
Orchestration
Kubernetes (AKS/EKS/GKE), Helm, RBAC, Calico, Docker Security & Compliance
Vulnerability Scanning (Aqua, Twistlock), Secrets Management (Vault, AWS Secrets Manager), IAM, TLS/SSL
OS Administration Linux (Ubuntu, RHEL), Windows Server, System Hardening, Patch Automation PROFESSIONAL EXPERIENCE
Site Reliability Engineer - DevOps
Accenture New York Jan 2025 – Current
• Developed and optimized Python scripts to automate log analysis and routine health checks across Linux (Ubuntu, RHEL) and Windows Server environments, reducing manual intervention and improving system reliability during on-call rotations.
• Administered Kubernetes (AKS, Helm, RBAC) clusters and Docker containers to support high-availability applications, implementing blue-green and rolling update strategies that improved release success rates and reduced deployment rollbacks by 15%.
• Designed and maintained multi-cloud infrastructure across Azure, AWS (EC2, S3, CloudWatch), and GCP (Compute Engine, GKE, Cloud Logging), ensuring cost-effective resource utilization and achieving consistent uptime above 99% for business-critical workloads.
• Defined and monitored SLOs, SLIs, and error budgets for key production services, enabling data-driven reliability improvements and release risk management.
• Implemented Prometheus and Grafana dashboards with alert rules for latency, request failure rates, and resource utilization across clusters.
• Conducted chaos engineering experiments (using LitmusChaos) to validate system failover and resiliency under simulated outages.
• Introduced postmortem documentation process for every Sev-1 incident, tracking action items and preventing recurrence.
• Collaborated with developers to integrate OpenTelemetry traces into microservices for full request lifecycle visibility.
• Optimized infrastructure cost efficiency by automating unused resource cleanup and scaling policies, saving 20% in monthly cloud spend.
• Deployed CI/CD pipelines through Azure DevOps and GitHub Actions to standardize software delivery, integrating testing and compliance checks that cut build failures and supported cross-functional agile teams in faster release cycles.
• Oversaw VMware virtualization infrastructure for scaling workloads, optimizing VM provisioning, and integrating with Azure Automation to streamline patching cycles, which reduced downtime windows during maintenance.
• Led incident response and RCA sessions for high-priority service disruptions, documenting workflow automation improvements aligned with ITIL practices that enhanced SLA adherence and reduced repetitive incidents over a 6-month period.
• Conducted change management reviews and capacity planning for hybrid infrastructure, collaborating with security and network teams to ensure smooth rollouts of critical updates across Kubernetes and VMware environments. Site Reliability Engineer - DevOps
Mphasis India June 2019 - May 2023
• Developed auto-remediation workflows and self-healing infrastructure using Azure Functions and Python scripts to spot configuration drift and automatically restore system health in both staging and production.
• Built and maintained Docker-based microservices deployed via Helm to AKS, EKS, and GKE, resolving production issues related to pod autoscaling, service discovery, and inter-cluster routing.
• Created and enforced CI/CD pipelines using GitHub Actions and Azure DevOps with deployment strategies including Blue-Green, Canary, and Rolling Updates, integrating automated rollback and infrastructure validation across 4 environments.
• Architected, deployed, and managed highly available Kubernetes clusters across AWS, Azure, and GCP, automating scaling, monitoring, and maintenance workflows using Python scripts, which improved deployment speed by 30% and ensured consistent application uptime across multi-cloud environments.
• Managed Kubernetes workloads across 6 AKS clusters in dev, test, and prod, handling upgrades, ingress controller tuning, persistent volumes, and custom network policies using Calico.
• Maintained Azure-based test environments and configured Azure App Services, Azure Functions, and SQL Databases to mimic production for performance testing and release validation.
• Designed and implemented Python-based automation frameworks for Kubernetes resource management, CI/CD pipeline orchestration, and routine SRE tasks, improving operational efficiency by 35% and reducing manual intervention for repetitive processes.
• Configured end-to-end observability using Prometheus, Grafana, OpenTelemetry, and Jaeger to capture metrics, logs, and distributed traces, improving root cause analysis speed and reducing mean time to resolution (MTTR) by 40%.
• Automated VMware snapshot management, backup, and disaster recovery processes, improving system reliability and ensuring compliance with organizational uptime requirements while reducing manual administrative effort.
• Helped standardize root cause documentation and action item tracking in Confluence after each RCA, ensuring no high- priority incident was closed without follow-through.
• Took lead role in RCA and postmortem analysis for various P1 incidents using structured SRE templates, SLO/SLI data, and impact assessments reviewed by cross-functional engineering leads in an Agile/Scrum environment.
• Wrote Bash scripts to automate image builds, container pushes, and deployment triggers for Docker-based apps in Azure DevTest Labs used by various cross-functional teams.
• Collaborated with SRE and DevOps teams to integrate Jenkins pipelines for daily CI builds and Git hooks, catching all the possible broken builds before hitting QA cycles.
• Mentored interns and junior engineers, improving team productivity and knowledge retention. EDUCATION
• Master of Science, Computer Science (June 2023 - December 2024) University of Central Missouri, Warrensburg, MO
• Bachelor of Computer Applications, Computer Science (July 2017 - March 2020) Osmania University, Hyderabad, Telangana