Post Job Free
Sign in

Senior Cloud Observability & Operations Engineer

Location:
Amravati, Maharashtra, India
Salary:
1145000
Posted:
September 16, 2026

Contact this candidate

Resume:

Kunal Bharatbhai Davada

SENIOR CLOUD OPERATIONS & OBSERVABILITY ENGINEER AWS Azure

Multi-Cloud Infrastructure L2 NOC Operations Network Operations Cloud Monitoring Observability Incident Management Production Support Infrastructure Operations DevOps Collaboration ITSM ***************@*****.*** 766-***-****

prasad Nagar colony, near saturna area, plot 41 lane 3 1995-06-27 Indian Immediate joiner linkedin.com/in/kunal-davada-6003793a6 Profile

Senior Cloud & Observability Engineer with 8+ years of experience in Cloud Operations, Infrastructure Operations, Network Operations, Monitoring, Observability, Incident Management and Production Support across AWS, Azure and VMware environments. Started career in Cloud Support & Operations, progressed into AWS Administration, and later moved into Multi-Cloud, Network and Infrastructure Operations, with recent experience focused on Observability, Production Operations, Major Incident Management and close collaboration with DevOps teams.

Experienced in monitoring and troubleshooting production infrastructure and applications using Dynatrace, New Relic, Grafana, Splunk, Amazon CloudWatch, SolarWinds, Zabbix and HostMonitor. Strong knowledge of metrics, logs, traces, synthetic monitoring, service alerts, anomaly alerts, infrastructure health alerts, low-disk alerts and application availability monitoring.

Experienced in complete Incident, Problem and Change Management, including P1/P2 incidents, alert triage, escalation, troubleshooting, RCA support, SLA/SLI/SLO adherence, customer communication and ticket closure.

Strong exposure to AWS, Azure, VMware, Citrix, Okta, AWS IAM, Terraform, Spacelift, ServiceNow, Jira, PagerDuty, Salesforce CMDB and Power BI, along with production deployment, migration, server patching, runbook execution, health checks, smoke testing and post-deployment validation. Additional experience with Amazon Connect, AWS Lex, Genesys, Zendesk and Webhooks, supporting cloud/contact-center and application-related operational workflows. Experience

Multi Cloud & Network Operations Engineer / L2 NOC Engineer, Foundever crm india private limited

Worked as a Multi-Cloud & Network Operations / L2 NOC Engineer in a 24x7 Command Centre environment, responsible for production infrastructure monitoring, observability, incident management, network and cloud operations, application troubleshooting, deployment support, server patching and coordination with DevOps and L2/L3 technical teams. Supported business-critical environments across multiple domains including Healthcare, FinTech, Automobile, Insurance and Logistics. Worked across AWS, Azure, VMware and enterprise infrastructure environments, with strong focus on monitoring, observability, service availability, incident response and production stability.

21/08/2023 –

30/06/2026

Chennai City,

Tamilnadu state

Cloud & Infrastructure Operations

•Provided L2 production support for cloud, infrastructure, application and network environments in a 24x7 operational environment.

•Supported AWS, Microsoft Azure and VMware vSphere environments and hybrid infrastructure.

•Performed operational troubleshooting for cloud infrastructure, servers, applications, services and connectivity-related issues.

•Supported AWS services including EC2, VPC, IAM, CloudWatch, EBS, AMI, Security Groups, NACLs, Internet Gateway, NAT Gateway and Elastic Load Balancing.

•Supported Azure services including Azure Virtual Machines, VNet, Subnets, Network Security Groups, Azure Monitor and Log Analytics.

•Used Azure Jump Server for controlled access to Azure-hosted environments where required.

•Supported VMware vSphere environments and virtual infrastructure health.

•Provided operational support for Windows and Linux servers, application servers and web servers.

•Supported infrastructure availability, capacity, health and performance monitoring.

•Assisted in identifying infrastructure-level issues before they affected application availability.

Monitoring & Observability

•Performed 24x7 infrastructure and application monitoring using multiple enterprise observability and monitoring platforms.

•Worked with Dynatrace, New Relic, Grafana, Splunk, Amazon CloudWatch, SolarWinds, Zabbix and HostMonitor.

•Monitored host, application, service, network and cloud infrastructure health.

•Investigated alerts related to:

•CPU and memory utilization

•Low disk space

•Service availability

•Application failures

•Network connectivity

•Synthetic monitoring

•IsAlive/service health checks

•Anomaly detection

•Application performance

•HTTP errors

•Infrastructure health

•Worked with synthetic monitoring and synthetic availability checks to validate whether services were accessible and responding as expected.

•Used Grafana dashboards to analyze infrastructure and application metrics and monitor service health.

•Used Dynatrace for application performance monitoring, service health, transaction analysis and tracing.

•Used Splunk for centralized log analysis, error investigation and troubleshooting.

•Used New Relic for monitoring dashboards, alert policies, application/service monitoring and performance analysis.

•Used Amazon CloudWatch for AWS metrics, logs, alarms, dashboards and infrastructure monitoring.

•Used SolarWinds for network and infrastructure monitoring.

•Used Zabbix and HostMonitor for infrastructure and host-level monitoring.

•Performed alert validation and determined whether an alert represented a genuine production issue or monitoring noise.

•Correlated alerts with logs, metrics and application/service behaviour before escalation.

•Supported alert deduplication and appropriate incident creation to reduce unnecessary escalation.

Alert & Error Investigation

•Investigated monitoring alerts from initial detection through technical troubleshooting and resolution.

•Analyzed metrics, logs, traces, service health and application responses to identify the probable source of incidents.

•Troubleshot application and infrastructure errors including:

•HTTP 404

•HTTP 403

•HTTP 500

•HTTP 503

•DNS resolution failures

•Unknown host errors

•Connectivity failures

•Service unavailable conditions

•Application configuration issues

•Performance degradation

•Low disk conditions

•Service/process failures

•Performed basic-to-advanced operational troubleshooting before escalating to specialized L2/L3 teams.

•Coordinated with application, database, network, cloud and DevOps teams when deeper investigation was required.

Incident Management

•Managed the complete end-to-end incident lifecycle from alert detection through final closure.

•Created and updated incidents based on monitoring alerts and customer- reported issues.

•Performed incident classification, categorization, prioritization and impact assessment.

•Handled P1, P2, P3 and P4 incidents according to business impact and operational procedures.

•Participated in Major Incident / P1-P2 bridge calls during critical production outages.

•Coordinated with L2/L3, DevOps, Cloud, Network, Application and DBA teams during major incidents.

•Provided real-time incident updates including:

•Business impact

•Technical findings

•Actions performed

•Teams engaged

•Current status

•Expected resolution/ETA

•Followed defined SLA requirements for incident acknowledgement, response and resolution.

•Applied SLI/SLO/SLA concepts when evaluating service health and operational performance.

•Focused on restoring service quickly and reducing MTTR.

•Supported post-incident analysis and recurring incident identification. PagerDuty & On-Call Operations

•Used PagerDuty for real-time alerting, on-call management and incident escalation.

•Supported alert-to-incident integration workflows.

•Managed severity-based escalation and routing.

•Supported escalation across L1 L2 L3 teams.

•Validated alerts before escalation.

•Supported alert deduplication and noise reduction.

•Monitored active incidents and ensured appropriate ownership and escalation.

Problem Management & RCA

•Participated in Problem Management activities for recurring production issues.

•Identified recurring alert and incident patterns.

•Supported Root Cause Analysis (RCA) by correlating logs, metrics, traces and incident history.

•Coordinated with technical teams to identify permanent corrective actions.

•Documented troubleshooting findings and recurring issue patterns.

•Supported improvements to monitoring, alerting and operational procedures.

DevOps Collaboration

•Worked closely with DevOps and infrastructure teams during production deployments and infrastructure changes.

•Supported infrastructure implementation and controlled production changes.

•Worked with Terraform and Spacelift as part of infrastructure-as-code and deployment workflows.

•Supported deployment validation and production readiness activities.

•Coordinated with DevOps teams during:

•Infrastructure deployment

•Application deployment

•Migration activities

•Production implementation

•Change execution

•Post-deployment validation

•Performed operational health checks following deployments.

•Supported smoke testing and service validation after implementation.

•Verified monitoring and alerts after deployment to ensure services were functioning normally.

Server Patching & Maintenance

•Supported Windows and Linux server patching activities in both UAT and Production environments.

•Followed approved runbooks for patch implementation.

•Performed pre-patching health checks and service validation.

•Supported patch implementation during approved maintenance windows.

•Performed post-patching:

•Service checks

•Application validation

•Server health checks

•Monitoring verification

•Connectivity validation

•Supported rollback validation and recovery activities where required.

•Maintained change and implementation documentation.

•Coordinated with application, infrastructure and DevOps teams during patching activities.

Change, Deployment & Migration Support

•Supported the complete change management lifecycle including planning, approval, implementation, validation and closure.

•Assisted in production and non-production deployment activities.

•Supported end-to-end migration activities.

•Performed pre-implementation checks and post-implementation validation.

•Conducted smoke testing and health checks after deployments.

•Verified service availability and application functionality after changes.

•Supported rollback validation when implementation did not produce expected results.

•Ensured production changes were executed according to approved runbooks and change procedures.

ITSM & Ticket Management

•Worked extensively with ServiceNow for:

•Incident Management

•Request Management

•Change Management

•Problem Management

•Ticket lifecycle management

•Used Jira / Jira Service Management for incident and operational workflow management.

•Created, updated and closed tickets based on defined processes.

•Maintained accurate troubleshooting notes, actions and resolution details.

•Ensured tickets were resolved within applicable SLA.

•Maintained proper incident traceability and operational documentation.

•Used Confluence for operational documentation, SOPs and knowledge articles.

Customer Communication

•Communicated with customers and stakeholders during active incidents.

•Provided service status updates and resolution information.

•Used approved service/no-reply email communication workflows to confirm whether the reported issue had been resolved.

•Followed up with customers for confirmation before closing applicable incidents.

•Maintained professional communication during production-impacting incidents.

•Ensured customer-facing updates reflected the latest technical status. Identity & Access Management

•Supported Okta for authentication and access-related operational activities.

•Supported MFA and authentication troubleshooting.

•Assisted with user access provisioning and de-provisioning.

•Supported AWS IAM users, roles and permission-related requests.

•Investigated access failures and authentication-related issues.

•Coordinated with appropriate teams for access approval and resolution. CMDB & Reporting

•Worked with Salesforce CMDB for configuration and customer/service records.

•Maintained accurate infrastructure and service-related records.

•Used Power BI for operational reporting and analysis.

•Supported reporting related to customer records, service information and operational performance.

Contact Center & Application Technologies

•Worked with/support exposure to Amazon Connect, AWS Lex, Genesys and Zendesk.

•Supported operational workflows associated with cloud/contact-center platforms.

•Exposure to Amazon Connect IVR, AWS Lex chatbot/conversational workflows and Webhook integrations.

•Supported troubleshooting and operational monitoring of application/contact-center services.

•Worked with REST/HTTP-based application endpoints and service availability checks.

Key Technologies

AWS Azure VMware vSphere Citrix Dynatrace New Relic Grafana Splunk SolarWinds Zabbix HostMonitor Amazon CloudWatch PagerDuty ServiceNow Jira Confluence Terraform Spacelift Okta AWS IAM Salesforce CMDB Power BI Amazon Connect AWS Lex Genesys Zendesk Webhooks Windows Linux Networking Associate Engineer, Cognizant technology solutions Worked as an AWS Administrator, providing cloud administration, identity and access management, monitoring and operational support across enterprise environments in the Healthcare, Insurance and Automobile domains.

AWS Cloud Administration

14/12/2022 –

08/08/2023

Pune City,

Maharashtra state

•Provided day-to-day operational administration and support for AWS cloud environments.

•Supported production and non-production AWS environments.

•Worked with AWS infrastructure and cloud administration activities according to approved operational procedures.

•Supported cloud service availability and operational health.

•Assisted in troubleshooting AWS infrastructure and access-related issues.

•Followed defined SOPs, security policies and AWS operational best practices.

AWS IAM & Access Management

•Managed AWS IAM users, groups, roles and policies.

•Supported privileged-access management activities.

•Performed user onboarding based on approved access requests.

•Processed access modification requests after appropriate approvals.

•Supported user deactivation and access removal.

•Ensured access was provisioned according to defined role and security requirements.

•Troubleshot authentication and authorization-related issues.

•Coordinated with application, security and infrastructure teams for access- related incidents.

Cloud Monitoring & Operations

•Supported AWS infrastructure monitoring and responded to operational alerts.

•Investigated cloud service alerts and coordinated resolution.

•Supported availability and health monitoring of AWS resources.

•Assisted in identifying infrastructure-related issues before they affected production services.

•Escalated complex technical issues to appropriate L2/L3 teams. Incident & Service Management

•Managed AWS administration incidents, service requests and access- related tickets.

•Performed ticket analysis, categorization, troubleshooting and resolution.

•Maintained SLA adherence for assigned requests and incidents.

•Documented troubleshooting steps and resolution details.

•Coordinated with internal teams when incidents required additional technical support.

•Ensured tickets were updated with accurate technical information before closure.

Security & Compliance

•Followed AWS security best practices for IAM and access management.

•Ensured user access was provided only after required approvals.

•Supported security and compliance procedures related to cloud administration.

•Maintained proper access and operational documentation.

•Supported privileged-access controls in production environments. Cross-Functional Collaboration

•Worked with Cloud, Infrastructure, Application, Security and Operations teams to resolve production issues.

•Coordinated access-related changes and cloud administration activities.

•Supported operational changes according to approved procedures.

•Communicated technical status and requirements to relevant stakeholders. Key Technologies

AWS IAM EC2 CloudWatch VPC Security Groups Cloud Administration Identity & Access Management Privileged Access Incident Management Service Requests SLA Security & Compliance Production Support

L1 Cloud Operations and Infrastructure Support, Brain vision technology Started professional career in L1 Cloud Operations and Infrastructure Support, providing 24x7 operational support for AWS and Azure environments. Developed foundational expertise in cloud monitoring, incident management, network troubleshooting, access management, customer support and production operations.

AWS & Azure Cloud Operations

08/2018 – 12/2022

Pune City

Maharashtra state

•Provided L1 operational support for AWS and Microsoft Azure cloud environments.

•Monitored infrastructure and services using Amazon CloudWatch and Azure Monitor.

•Responded to infrastructure and application alerts.

•Performed initial investigation and troubleshooting of cloud-related incidents.

•Escalated complex issues to L2/L3 teams according to defined escalation procedures.

•Assisted with cloud resource provisioning activities.

•Supported basic cloud infrastructure administration.

•Assisted with access-management requests.

•Supported routine cloud operational activities and service health checks. Monitoring & Alert Management

•Monitored infrastructure using CloudWatch and Azure Monitor.

•Investigated alerts related to server health, resource utilization and service availability.

•Performed initial alert validation and troubleshooting.

•Checked service status and basic infrastructure health before escalation.

•Monitored production environments and ensured alerts were appropriately handled.

•Maintained incident records for monitoring-related issues. Incident Management

•Managed incidents, service requests and changes using Jira.

•Performed initial incident analysis and troubleshooting.

•Classified and prioritized tickets based on impact and urgency.

•Escalated complex incidents to appropriate L2/L3 technical teams.

•Maintained SLA adherence for assigned incidents and service requests.

•Updated tickets with investigation details and resolution information.

•Followed defined incident and escalation procedures. Network & Infrastructure Support

•Assisted with basic network setup and infrastructure support.

•Performed initial troubleshooting of connectivity-related issues.

•Supported troubleshooting involving:

•DNS

•Network connectivity

•VPN

•IP/subnet-related issues

•Basic routing/connectivity

•Coordinated with network and infrastructure teams for complex issues.

•Supported application availability issues caused by infrastructure or connectivity problems.

Server & System Operations

•Supported routine Windows and Linux server operations.

•Performed routine server health checks.

•Supported patching activities.

•Conducted backup checks.

•Performed log reviews for initial troubleshooting.

•Assisted with service validation and operational checks.

•Escalated system-level issues requiring specialized administration. Access Management

•Assisted with user access and cloud access-management requests.

•Supported basic authentication and authorization troubleshooting.

•Coordinated access requests with appropriate teams.

•Followed approval and security procedures for access-related activities. Customer Support

•Used Salesforce for customer case tracking and service updates.

•Created and maintained customer/service records.

•Provided customer-facing service updates based on incident status.

•Updated cases with troubleshooting details and resolution information.

•Coordinated with technical teams to provide accurate customer updates.

•Supported customer communication until issue resolution. Documentation & Runbooks

•Created and maintained runbooks and troubleshooting documentation.

•Documented common operational procedures and recurring issues.

•Followed documented SOPs for routine cloud and infrastructure activities.

•Maintained accurate operational records.

•Supported knowledge sharing for recurring troubleshooting procedures. Security & Operational Compliance

•Followed defined security, privacy and operational policies.

•Ensured cloud and infrastructure activities were performed according to approved procedures.

•Maintained accurate incident and customer records.

•Followed escalation and change-management procedures. Key Technologies

AWS Azure Amazon CloudWatch Azure Monitor Cloud Operations Infrastructure Support Network Support Jira Salesforce Incident Management Service Requests Change Management Access Management Windows Linux DNS VPN Cloud Monitoring Server Operations Patching Backup Checks Log Analysis Runbooks Education

Bachelor of commerce (computer application),

Sant gadge Baba amravati University

•Generations of computer

•Information technology

01/07/2013 –

28/11/2016

Amravati City,

Maharashtra state

•Basics knowledge about.net

Hsc, Maharashtra board 02/07/2012 –

10/04/2013

Skills

CORE KEY SKILLS

Cloud & Infrastructure: AWS, Microsoft Azure,

VMware vSphere, Hybrid Cloud, Cloud

Infrastructure Operations

AWS: EC2, VPC, IAM, CloudWatch, S3, RDS,

EBS, AMI, ELB/ALB/NLB, Lambda, API Gateway

Azure: Azure VM, VNet, NSG, Azure Monitor,

Log Analytics, Entra ID/Azure AD, Azure Virtual

Desktop, Jump Server

Monitoring & Observability: Dynatrace, New

Relic, Grafana, Splunk, Amazon CloudWatch,

SolarWinds, Zabbix, HostMonitor, Synthetic

Monitoring, Application Monitoring,

Infrastructure Monitoring, Metrics, Logs, Traces,

Alert Management

Production & Operations: Incident

Management, Major Incident Management,

Problem Management, Change Management,

P1/P2/P3/P4 Incident Handling, RCA, Alert

Triage, Escalation Management, SLA, SLI, SLO,

MTTR

DevOps & Infrastructure Operations: Terraform,

Spacelift, Infrastructure as Code, Deployment

Support, Production Implementation, Migration

Support, Server Patching, Runbook Execution,

Health Checks, Smoke Testing, Post-Deployment

Validation, Rollback Validation

ITSM & Incident Tools: ServiceNow, Jira, Jira

Service Management, PagerDuty, Confluence,

Salesforce CMDB, Power BI

Identity & Access: AWS IAM, Okta, MFA,

Authentication, Authorization, User Provisioning,

De-provisioning, Access Management

Virtualization & End User Computing: VMware

vSphere, Citrix

Network & Server Operations: DNS, TCP/IP,

VPN, LAN/WAN, VPC Networking, Subnetting,

Connectivity Troubleshooting, Windows Server,

Linux Server

Operational Tools: PowerShell Tool, CSDR Tool,

Service Restart, Service Status Validation,

Operational Remediation, Runbook-Based

Operations

Application & Contact Center: Amazon

Connect, AWS Lex, Genesys, Zendesk,

Webhooks, REST/HTTP APIs, IVR, Chatbot

Programming & Scripting: Basic Exposure – not

a primary area of expertise; primarily

experienced in operational tools, runbook-based

activities and production troubleshooting.

Industry Domain Experience:

Healthcare Banking FinTech Automobile Logistics Insurance



Contact this candidate