Chenna Keshava N
Email: ***********@*****.***
Mobile: 936-***-****
LinkedIn: www.linkedin.com/in/chenna-keshava-n-a45480415 Site Reliability Engineer
PROFESSIONAL SUMMARY
DevOps Engineer with 4+ years of experience architecting cloud and DevOps practices across AWS, Azure, Kubernetes, Terraform, improving release reliability and governance enterprise platforms.
Delivered infrastructure as code, monitoring, and incident response workflows with Jenkins, GitHub Actions, Prometheus, and Grafana, increasing platform resilience and consistent delivery outcomes enterprise-wide.
Standardized container orchestration, security controls, and release management across Docker, OpenShift, Helm, and GitOps, strengthening compliance, observability, and service availability outcomes across cloud platforms.
Optimized automation, scripting, and cloud engineering practices with Python, Bash, PowerShell, Ansible, and Terraform, reducing manual effort and accelerating environment provisioning for delivery teams.
Facilitated strong communication skills during team meetings, resulting in improved collaboration and clarity among team members, which led to increased project efficiency and a more cohesive work environment.
Employed a customer-centric mindset while addressing client concerns, which significantly enhanced customer satisfaction and loyalty, ultimately driving repeat business and fostering long-term relationships with key clients. TECHNICAL SKILLS
Cloud Platforms - AWS (EKS, EC2, Lambda, S3, RDS, CloudFormation, CloudWatch, VPC), Azure (AKS, Azure DevOps, Functions, App Services, ARM Templates, Azure Monitor)
SRE & Observability - Datadog, Prometheus, Grafana, ELK Stack, PagerDuty, CloudWatch, Azure Monitor, New Relic, SLO/SLI Management, Incident Response, MTTR Optimization, monitoring gaps, alerting strategies
Databases - PostgreSQL, MySQL, MongoDB, DynamoDB, Redis, Azure SQL Database
Containerization & Orchestration - Kubernetes, Docker, Helm, ArgoCD, Flux, GitOps, Container Registry, Service Mesh Infrastructure as Code: Terraform, Ansible, CloudFormation, ARM Templates, Bicep
Operating Systems - Linux (Ubuntu, CentOS, RHEL), Windows Server, Amazon Linux, Linux kernel internals
CI/CD & Automation - Azure DevOps, Jenkins, GitLab CI, GitHub Actions, Azure Pipelines, AWS CodePipeline, automating routine processes
Security & Compliance - HashiCorp Vault, Prisma Cloud, SonarQube, Azure Security Center, IAM (AWS/Azure), DevSecOps, PCI-DSS, SOC 2
Programming & Scripting - Bash, PowerShell, Go, YAML, JSON
Collaboration Tools - GitHub, GitLab, Jira, Confluence, Slack, Agile/Scrum
System Performance & Debugging - system performance metrics, root cause analysis, analyzing system logs, resilient code, architecture, design patterns, scaling systems, system-level debugging, kernel panic analysis, kdump
Networking - TCP/IP network programming
Storage Solutions - distributed storage systems, object storage, block storage, file storage
Schedulers - memory allocation, driver subsystem PROFESSIONAL EXPERIENCE
American Express January 2025 – Present
Sr. DevOps Engineer New York, NY, USA
Spearheaded the integration of kdump with CI/CD practices in AWX, enhancing system observability and reliability, resulting in a 30% reduction in incident response time across production environments.
Modernized driver subsystem architecture for object storage, achieving a 50% increase in read/write speeds and significantly enhancing user experience across cloud storage solutions.
Engineered resilient code using established design patterns, which minimized system vulnerabilities and enhanced the security posture of applications serving millions of users daily.
Enhanced monitoring capabilities through deep understanding of Linux kernel internals, which improved incident response times by 50% and bolstered system observability across critical infrastructure components.
Orchestrated bare-metal cloud infrastructure enhancements utilizing TCP/IP network programming and OVN/OVS-based networking stack, achieving a 30% increase in throughput and significantly reducing latency across distributed services.
Designed secure CI/CD pipelines with Jenkins, GitHub Actions, and GitLab CI, improving application releases, traceability, rollback readiness, and audit-aligned deployment governance across enterprise platforms.
Engineered Terraform and CloudFormation modules across AWS VPC, IAM, EC2, and EKS, strengthening reusable infrastructure patterns and improving environment consistency across standardized landing zones.
Automated container deployments with Kubernetes, Docker, Helm, and Argo CD, improving release repeatability, service uptime, and production support efficiency across scalable containerized production environments.
Enhanced observability with Prometheus, Grafana, CloudWatch, Splunk, and Datadog, improving incident visibility, alert quality, and operational response across payment platforms and reliability engineering practices. CVS Health July 2023 – December 2024
DevOps Engineer Woonsocket, RI, USA
Consolidated block storage and file storage systems into a unified architecture, which reduced operational complexity and cut costs by 20% while maintaining high availability and performance.
Streamlined scaling systems to address monitoring gaps, resulting in a 60% improvement in resource utilization and ensuring optimal performance during peak operational periods.
Spearheaded system-level debugging and kernel panic analysis initiatives, reducing downtime by 99% and ensuring high availability of services critical to business operations and customer satisfaction.
Executed root cause analysis on system performance metrics and Hardware GPU troubleshooting, leading to a 25% reduction in system failures and enhancing overall system reliability and user satisfaction.
Integrated Azure DevOps pipelines with AKS, Azure Key Vault, and Terraform, improving secure application delivery, configuration control, and healthcare platform reliability across delivery teams.
Configured Docker, Kubernetes, and OpenShift deployment workflows with Helm charts, improving environment parity, application scalability, and production release stability across standardized healthcare runtime environments.
Validated DevSecOps controls with Snyk, SonarQube, Trivy, and vulnerability management processes, strengthening remediation workflows and compliance readiness across regulated delivery pipelines for healthcare platforms.
Streamlined incident management with ServiceNow, Jira, runbooks, SLIs, and SLOs, improving triage consistency, escalation accuracy, and operational service restoration across support teams and applications. Accenture October 2021 – November 2022
Cloud Engineer Hyderabad, India
Architected a new scheduler framework that improved resource allocation efficiency, resulting in a 30% increase in throughput and enabling the platform to handle higher workloads seamlessly.
Optimized alerting strategies and memory allocation processes, which improved system performance metrics by 45% and reduced operational overhead associated with alert fatigue.
Pioneered improvements in distributed storage systems by conducting post-mortems on failures, which led to a 35% increase in data integrity and reliability across storage solutions.
Automated routine processes by analyzing system logs, resulting in a 40% decrease in manual intervention and enabling teams to focus on strategic initiatives that drove business growth.
Established multi-cloud provisioning standards across AWS, Azure, GCP, Terraform, and Ansible, improving reusable environments, governance alignment, and delivery velocity across client portfolios and programs.
Secured cloud identity and access with IAM, RBAC, MFA, SSO, and Key Vault, improving policy enforcement and reducing configuration risk across enterprise cloud environments.
Modernized logging and monitoring practices with ELK, OpenTelemetry, New Relic, and Dynatrace, improving root- cause analysis and distributed system reliability across distributed cloud service landscapes.
Coordinated migration readiness with Kubernetes, serverless, networking, DNS, VPN, and load balancing patterns, improving platform scalability and cloud adoption outcomes across enterprise transformation programs. EDUCATION
Master's in Computer Science and Engineering - University of Bridgeport
Bachelor's in Computer Science and Engineering - JNTUH