SWATI RASTOGI
Email: *****.*******@*****.***
Contact: +91-991*******
Place: Noida
Profile
Lead Reliability & Platform Engineer 15+ Yrs Experience Enterprise AI Platforms SRE Cloud Infrastructure DevOps Observability Production Operations LLM Application Reliability
Technical Skills
CI/CD & Automation: GitLab CICD, GitHub Actions, Azure DevOps (ADO), Jenkins, PowerShell, Python
Cloud Platforms: AWS (EC2, RDS, IAM, Lambda, S3, Load Balancers, ECS, Landing zones, AWS Networking, CDN),Azure Cloud (app services, VNet, App gateway, Functions, Redis Cache, Postgres etc), GCP.
Infrastructure as Code (IaC): Terraform
Containerization & Orchestration: Docker, Kubernetes, OpenShift
Database & Technology Expertise: Redis Cache, Postgres DB flexible server, MS SQL, Snowflake, .NET 3.5
Security & Compliance: AWS Security Hub, CyberArk, AWS Backup, Wiz, Armor Code
Observability & Reliability : Grafana, PagerDuty, Application Insights, CloudWatch, Open Telemetry
Key Skills: Site Reliability Engineering, Cloud Infrastructure (AWS, Azure, GCP), Infrastructure-as-Code (Terraform), CI/CD Automation, Observability (Grafana, Application Insights, CloudWatch, OpenTelemetry), Incident Response, Security (Identity Secrets, Policy-as-Code), Containerization (Docker, Kubernetes), Cost Optimization, Mentoring
Certifications
Red Hat Certified Specialist in Containers and Kubernetes (Dec 2021-Dec 2024)
HashiCorp Certified: Terraform Associate (003) (May 2023-May 2025)
CAREER HIGHLIGHTS
S&P Global – Noida, India
Lead Reliability & Platform Engineer (Nov. 2023 onwards)
Job Profile and Responsibilities
AI
Currently working in Spark Assist project as SRE & Platform Engineer. Spark Assist is an AI assistant for S&P Global employees.
Ensured high availability and production stability of AI/LLM-based services across enterprise environments.
Supported model and service deployments, including release validation, rollback planning, and production readiness checks.
Created SLO/SLA/SLI documents for model onboarding.
Drove SLI/SLO adoption and practices across engineering teams, improving service reliability and performance metrics.
Managed scalable cloud infrastructure to support AI workloads, user traffic growth, and high-volume platform usage.
Implemented observability and monitoring for latency, error rates, availability, throughput, API health, and infrastructure performance.
Led incident response and root-cause analysis, creating runbooks to reduce downtime and improve recovery processes.
Optimized platform performance by identifying bottlenecks, improving response times, and supporting capacity planning.
Automated deployment and operational workflows to reduce manual steps and improve release consistency.
Collaborated with AI engineering, security, infrastructure, and product teams to deliver reliable and secure AI platform capabilities.
Established operational governance for reliability, supportability, and compliance readiness for the AI platform.
Azure
Spark Assist is hosted on Azure App Service.
Managed Azure App Service operations, configuration management, and production stability.
Configured deployment slots for safe blue/green/canary releases, reducing downtime with quick rollback during production deployments.
Maintained App Service Plan capacity and scaling, optimizing compute resources for performance, availability, and cost efficiency.
Implemented autoscaling and capacity management for fluctuating AI workloads.
Monitored application health using Azure Monitor and Application Insights.
Used Log Analytics to troubleshoot application issues, identify failure patterns, and support root-cause analysis during incidents.
Managed secure integrations using Managed Identity and Azure Key Vault.
Supported connectivity with Azure services such as Azure OpenAI, Azure Storage, Azure Cache for Redis, and other backend dependencies used by Spark Assist.
Configured network security controls, including VNet integration, private endpoints, access restrictions, and secure inbound/outbound connectivity.
Supported production readiness practices: TLS/certificates, diagnostic settings, alerts, and environment runbooks.
Maintained IaC using Terraform, to enable consistent prod & non-prod deployments.
Created Secured Internet facing app service using app gateway + akamai WAF.
Partner with Wipro to enable operations support, incident response, ServiceNow (SNOW) request handling, and release management KT for Spark Assist production deployments.
Supported Zero Trust and least-privilege adoption across multiple teams, configuring secure network segmentation and access controls.
Optimized cloud spend and implemented cost management strategies for Azure and AWS resources.
AWS
Added certificate manager & monitoring services in EKS cluster. Used Flux to monitor & manage deployments.
Designed and optimized CI/CD pipelines for AWS deployments, automating processes through Auto
Scaling Groups (ASG), which increased reliability.
Upgraded Terraform infrastructure from v0.14.0 to v1.5.4 and maintained an AWS Security Hub Score of 95%+
Implemented AWS Back Up & Compliance locks.
Created an automated solution for customized CloudTrail logs to process & send data via mail on regular basis.
Lead the migration of an internal product (fortdocs) across AWS accounts, utilizing ADO, Gitlab/Github, Terraform.
Extensive hands-on experience with containerization (Docker) and orchestration (Kubernetes, OpenShift, EKS), including upgrades, deployments, and operational management.
Worked on Gitlab->GitHub migration.
IBM – Noida, India
Senior DevOps Consultant (Aug. 2021 - Oct 2023)
Job Profile and Responsibilities
To advise the client on application development methodology and tools.
Coordinating with product operations team to resolve issues, developing, and running scripts, and troubleshooting services in a hosted environment.
Worked on AWS Services like EKS, EC2, S3, Data kinesis, VPC, Lambda functions, CFT,Cloudwatch,RDS,CDN,AFT,Landing Zones, mTLS using AWS App Mesh etc.
Worked on Creating Monitoring & alerting channel for Cloud functions & compute instance in GCP using Terraform.
Worked on upgrading Kubernetes version for the existing k8s objects. Deployments using ArgoCD on EKS cluster.
Precisely (Erstwhile Pitney Bowes Software) – Noida, India
Senior Software Engineer ll (Sep. 2014 – Aug 2021)
Job Profile and Responsibilities
Manage and administer Git,Nexus, Artifactory, Gitlab, Jira,Confluence, CICD processes in AWS Cloud.
Implement DevOps best practices in Product teams.
Deployment using Kubernetes and Ansible.
Hosted & maintained Jenkins Build servers on Openshift platform.
Accenture Services Pvt Ltd. – Gurgaon, India
Software Engineer (July. 2012 – Sep. 2014)
Job Profile and Responsibilities
Created Stored Procedures in T-SQL to transfer Data, providing billing calculations as required functionally.
Created SSRS Reports,SSIS Packages to display account information
Wipro Technologies – Hyderabad, India
Project Engineer (Aug. 2010 – July. 2012)
Job Profile and Responsibilities
Involved in Requirement Analysis and preparing Technical Specifications.
Created Stored Procedures, SSIS Package in T-SQL to transfer Data.
EDUCATIONAL CREDENTIALS
Integrated M.Tech (Biotechnology) from Amity University Noida in 2010.