Post Job Free
Sign in

Senior DevOps/SRE Cloud AI Infrastructure Engineer

Location:
50 Penn West, OK, 73112
Posted:
September 16, 2026

Contact this candidate

Resume:

Kotikalapudi Sai Teja

Senior DevOps Engineer SRE Cloud Engineer AI/ML Infrastructure

Hyderabad, India +91-906******* *****************@*****.*** https://www.linkedin.com/in/sai-teja-kotikalapudicontacttejajain/ github.com/saiteja-kotikalapudi PROFESSIONAL SUMMARY

Senior DevOps/SRE/Cloud Engineer with 7years of experience building, automating, and operating production-grade cloud-native infrastructure on AWS at enterprise scale. Currently at DAZN (world's leading sports streaming platform) managing 15 EKS clusters and 500+ microservices serving millions of concurrent users across 200+ countries. Hands-on experience provisioning and managing NVIDIA GPU nodes on Kubernetes for AI/ML workloads, deploying LLM (Llama 3, Mistral) and VLM (LLaVA, InternVL) models in production environments, and implementing agentic AI workflow pipelines using LangChain and CrewAI. Expert in Terraform, Kubernetes

(EKS/OpenShift), CI/CD (Jenkins, GitLab, ArgoCD), observability (Prometheus, Grafana, OpenTelemetry), and SRE practices (SLI/SLO, error budgets, chaos engineering). Delivered 81% faster deployments, 35% cost reduction ($180K/year), 99.95% uptime, and 65% MTTR reduction.

PROFESSIONAL EXPERIENCE

Senior DevOps / SRE Engineer

DAZN — ) November 2023 – Present Global Sports Streaming (200+ Countries, Millions of Users Hyderabad, India Platform, Infrastructure & AI/ML:

■ Architected multi-region AWS infrastructure using Terraform (50+ modules), Terragrunt, and Helm — managing 200+ EC2, 50+ RDS

(Multi-AZ), and 15 EKS clusters with 99.95% SLA.

■ Provisioned and managed NVIDIA GPU nodes on EKS clusters for AI/ML workloads — deployed LLM (Llama 3, Mistral) and VLM (LLaVA, InternVL) models in production Kubernetes environments with optimized CUDA scheduling and GPU resource quotas.

■ Implemented agentic AI workflow deployments using LangChain and CrewAI on GPU-backed Kubernetes pods — enabling autonomous multi-step AI agent pipelines for real-time video analytics and content intelligence.

■ Operated EKS/OpenShift running 500+ microservices with HPA, VPA, Karpenter, Istio service mesh, and zero-downtime deployments — resource utilization improved 45%.

■ Configured ALB with Auto Scaling (target tracking, predictive scaling) — 40% off-peak savings while maintaining SLOs during 10x traffic spikes.

■ Managed multi-VPC networking with Transit Gateway, VPC peering, NAT gateways, and Route53 failover for cross-region high availability. CI/CD & Automation:

■ Built CI/CD using Jenkins, GitLab CI/CD, and ArgoCD GitOps with security scanning (SonarQube, Trivy, Snyk), Docker builds, and blue/green/canary delivery — releases reduced from 2 weeks to 2 days, 100+ deployments/week.

■ Built automated CI/CD pipelines for LLM/VLM model versioning, containerization, and deployment — reducing AI model release cycles from days to hours with automated rollback on inference degradation.

■ Developed 50+ Python/Boto3 scripts and 30+ Ansible playbooks for resource management, cost optimization, AMI building (Packer), and toil reduction — saving 20+ hours/week.

■ Implemented Helm/Kustomize overlays for standardized GitOps deployments across all environments.

■ Integrated CodeBuild/CodeDeploy/CodePipeline for serverless and EC2 automation with rollback triggers across 10+ services. SRE & Observability:

■ Built full observability stack using Prometheus/Thanos, Grafana (100+ dashboards), OpenTelemetry, Loki, PagerDuty — coverage across 500+ services, 200+ SLI metrics, MTTR reduced 65%.

■ Built GPU utilization dashboards (Prometheus DCGM exporter + Grafana) for real-time monitoring of AI/ML inference workloads across all GPU nodes.

■ Defined SLIs, SLOs, and error budgets for Tier-1/Tier-2 services — customer incidents reduced 40%.

■ Led 24/7 on-call with incident response (SEV1-4), blameless postmortems, action tracking — recurring incidents down 55%.

■ Conducted chaos engineering (AWS FIS, Litmus Chaos) identifying 20+ failure modes pre-production; built automated runbooks for top 25 incidents — manual intervention reduced 70%.

Security:

■ Established IAM least-privilege, WAF (10K+ blocks/day), OPA/Gatekeeper, CIS benchmark scanning.

■ Managed secrets via HashiCorp Vault and Secrets Manager with automated rotation.

■ Integrated DevSecOps: container scanning (Trivy, Snyk), SAST (SonarQube) in CI/CD — 95% vulnerabilities caught pre-production. DevOps Engineer

Apsis Technologies — Weyyak September 2021 – October 2023 (OTT Video Streaming, 5M+ Users) Bengaluru, India

■ Designed infrastructure using Terraform (30+ modules) with Terragrunt — 100% IaC coverage with drift detection and DR capabilities.

■ Managed EKS/OpenShift for 5M+ concurrent users with blue-green/canary deployments and zero-downtime rollouts.

■ Optimized CloudFront CDN with Lambda@Edge — latency reduced 60%, bandwidth costs down 45% ($95K savings).

■ Configured Jenkins/GitLab CI/CD with Maven builds, security scanning, Ansible deployments — 50+ deployments/week.

■ Managed JFrog Artifactory with RBAC, retention policies, vulnerability scanning across 8 teams.

■ Implemented ELK Stack (100GB+ daily logs) with OpenTelemetry and PagerDuty — detection time reduced 70%.

■ Built Kafka pipelines for real-time event processing with Spark analytics; defined SLIs/SLOs for streaming quality (buffering, start time, error rate).

■ Automated DR using AWS Backup, Lambda, cross-region replication — RPO 1hr, RTO 4hr. DevOps Engineer

Apsis Technologies — March 2019 – August 2021 Zurich Farmers Insurance (15M+ Policies) Bengaluru, India

■ Deployed Java/Tomcat on EC2 with Auto Scaling, health checks, rollback — 99.9% success rate.

■ Managed AWS VPC networking (subnets, security groups, NACLs, NAT, VPN) for hybrid connectivity.Automated

■ Implemented Jenkins distributed architecture (10+ agents, GitFlow) — build throughput improved 300%.

■ Implemented Prometheus/Grafana tracking 500+ metrics with baseline SLIs for critical services. Associate Software Engineer

Cogent E-Services July 2018 – February 2019 Hyderabad, India

■ Developed and optimized complex SQL queries, stored procedures, and views — query performance improved 40% through indexing and execution plan analysis.

■ Built Power BI dashboards with DAX calculations and data modeling — real-time insights for 50+ stakeholders across 5 departments.

■ Automated data processing using SQL/Python — manual reporting reduced 60%, delivery improved from weekly to daily.

■ Supported application deployments including database migrations, environment setup, and release coordination. CERTIFICATIONS & PROFESSIONAL DEVELOPMENT

■ AWS Solutions Architect Associate (SAA-C03) — Exam Scheduled Q2 2026

■ 200+ hours hands-on labs: Kubernetes, Terraform, AWS Networking, Chaos Engineering, SRE (A Cloud Guru, KodeKloud, Udemy)

■ AWS re:Invent 2023/2024 CNCF KubeCon Workshops DevOps Institute SRE Foundation ArgoCD & Ansible Courses (KodeKloud) EDUCATION

Bachelor of Commerce (B.Com) Chaitanya Degree College, Kakinada, India 2015 – 2018 DECLARATION

I hereby declare that the information furnished above is true to the best of my knowledge and belief. Place: Hyderabad

K Sai Teja



Contact this candidate