Narsa Reddy Katpally
Email: ********************@*****.***
Mobile: +1-313-***-****
LinkedIn: www.linkedin.com/in/narsa-reddy-k/
Site Reliability Engineer
PROFESSIONAL SUMMARY
DevOps Engineer with 4+ years of experience automating CI/CD pipelines, Terraform infrastructure, Kubernetes platforms, and cloud operations across AWS, Azure, GCP for reliable delivery.
Skilled in infrastructure automation, observability, incident response, and DevSecOps practices, aligning secure deployments with enterprise reliability, compliance, and operational standards across distributed environments efficiently.
Experienced with Docker, Kubernetes, Helm, GitHub Actions, Jenkins, Ansible, and Terraform, improving release consistency, configuration control, scalability, and platform resilience for production workloads continuously.
Focused on scripting, monitoring, networking, and cloud governance, supporting automated deployments, root cause analysis, and dependable operations for business-critical applications across hybrid platforms securely.
Utilized strong communication skills to facilitate team meetings, ensuring all members contributed ideas and feedback, which led to enhanced collaboration and improved project outcomes across various departments.
Employed a customer-centric mindset to resolve client issues promptly, resulting in increased customer satisfaction scores and fostering long-term relationships that significantly boosted repeat business and company reputation. TECHNICAL SKILLS
Cloud Platforms - AWS (EKS, EC2, Lambda, S3, RDS, CloudFormation, CloudWatch, VPC), Azure (AKS, Azure DevOps, Functions, App Services, ARM Templates, Azure Monitor)
CI/CD & Automation - Azure DevOps, Jenkins, GitLab CI, GitHub Actions, Azure Pipelines, AWS CodePipeline
Security & Compliance - HashiCorp Vault, Prisma Cloud, SonarQube, Azure Security Center, IAM (AWS/Azure), DevSecOps, PCI-DSS, SOC 2
Databases - PostgreSQL, MySQL, MongoDB, DynamoDB, Redis, Azure SQL Database
Operating Systems - Linux (Ubuntu, CentOS, RHEL), Windows Server, Amazon Linux
Containerization & Orchestration - Kubernetes, Docker, Helm, ArgoCD, Flux, GitOps, Container Registry
Service Mesh Infrastructure as Code - Terraform, Ansible, CloudFormation, ARM Templates, Bicep
SRE & Observability - Datadog, Prometheus, Grafana, ELK Stack, PagerDuty, CloudWatch, Azure Monitor, New Relic, SLO/SLI Management, Incident Response, MTTR Optimization, monitoring capabilities, system performance metrics, monitoring gaps, alerting strategies, SLIs
Programming & Scripting - Python (Boto3), Bash, PowerShell, Go, YAML, JSON
Collaboration Tools - Git, GitHub, GitLab, Jira, Confluence, Slack, Agile/Scrum
System Architecture - design patterns, scaling systems
System Administration & Infrastructure - Linux kernel internals, scheduler, memory allocation, driver subsystem, system-level debugging, kdump, kernel panic analysis, bare-metal cloud infrastructure, TCP/IP network programming, distributed storage systems, object storage, block storage, file storage PROFESSIONAL EXPERIENCE
Goldman Sachs June 2025 – Present
Senior DevOps Engineer New York, NY, USA
Spearheaded the integration of CI/CD practices using AWX, enabling automated deployments that reduced release cycles by 50% and increased system reliability to 99.99% across multiple microservices in production.
Architected a robust solution for block storage and file storage integration, improving data redundancy and availability, which contributed to a 99.99% service level agreement compliance across distributed systems.
Spearheaded efforts in scaling systems to address monitoring gaps, achieving a 60% increase in system capacity while maintaining performance integrity and reducing operational costs by 20%.
Orchestrated system-level debugging initiatives and kernel panic analysis that improved overall system reliability, leading to a 99.98% uptime across critical services and enhancing customer satisfaction ratings.
Engineered a bare-metal cloud infrastructure leveraging TCP/IP network programming and an OVN/OVS-based networking stack, enhancing system throughput by 35% while reducing latency across distributed applications.
Designed secure CI/CD pipelines with Jenkins, GitHub Actions, and GitLab, improving controlled deployments, artifact traceability, and release reliability across enterprise platforms and delivery teams.
Engineered Terraform and CloudFormation modules for AWS infrastructure, standardizing VPC, IAM, EC2, and RDS provisioning while strengthening scalable environment consistency across repeatable cloud environments.
Automated Kubernetes deployments with Docker, Helm, and ArgoCD, improving container orchestration, rollout stability, and service availability for production workloads and platform resilience across clusters.
Integrated Prometheus, Grafana, CloudWatch, and Splunk observability, enabling faster incident response, root cause analysis, and dependable operational visibility across distributed services and platforms securely. Cardinal Health January 2024 – May 2025
DevOps Engineer Dublin, OH, USA
Implemented SLIs and refined architecture to enhance system observability, achieving a 35% improvement in incident response times and significantly increasing platform reliability for mission-critical applications.
Revolutionized alerting strategies and memory allocation processes, which led to a 50% decrease in false positives and enhanced system responsiveness during peak operational loads.
Consolidated distributed storage systems with comprehensive post-mortems that identified bottlenecks, resulting in a 30% improvement in data retrieval times and increased operational efficiency across the platform.
Pioneered the implementation of system performance metrics that automated routine processes and improved Hardware GPU troubleshooting, resulting in a 50% reduction in incident resolution time and enhanced user experience.
Configured Ansible automation for Linux, Windows, and application environments, reducing manual configuration drift while improving repeatable deployments and system maintainability across enterprise operations teams.
Validated DevSecOps controls with IAM, RBAC, Vault, and SonarQube, improving secure access governance, secrets management, and compliant release readiness across audited deployment workflows consistently.
Standardized Docker container builds through Nexus, Artifactory, and Git-based workflows, improving dependency control, image consistency, and deployment readiness across application delivery environments and releases.
Optimized incident response workflows with Datadog, PagerDuty, ELK, and RCA practices, improving service restoration, monitoring accuracy, and operational accountability across production support operations daily. Accenture May 2021 – June 2023
Cloud Engineer Hyderabad, India
Engineered an advanced scheduler and kdump integration that streamlined resource allocation, resulting in a 45% reduction in job processing times and enhanced throughput across the distributed computing environment.
Optimized the driver subsystem for object storage, enhancing throughput and reliability, which resulted in a 40% improvement in data access times and significantly increased customer satisfaction metrics.
Modernized monitoring capabilities by integrating advanced design patterns, which enhanced observability and reduced alert fatigue, resulting in a 25% increase in proactive incident management effectiveness.
Streamlined system logs analysis to optimize Linux kernel internals, achieving a 40% increase in operational efficiency and significantly reducing the time required for root cause identification during system outages.
Streamlined Azure DevOps pipelines with Terraform, Bicep, and PowerShell, improving cloud provisioning, repeatable releases, and environment governance across Azure platforms for scalable delivery operations.
Established Kubernetes platform automation on AKS and GKE with Helm, improving workload scalability, configuration alignment, and cloud-native deployment reliability across containerized service environments consistently.
Developed Python and Bash scripts for infrastructure automation, backup validation, and troubleshooting, improving operational efficiency and reducing recurring manual tasks across hybrid cloud operations.
Implemented CloudWatch, OpenTelemetry, SLOs, and SLA tracking for cloud services, improving observability practices, reliability reporting, and platform performance insights across distributed platform operations effectively. EDUCATION
Master's in Information Studies - Trine University