Kotikalapudi Sai Teja
Senior DevOps Engineer SRE Cloud Engineer AI/ML Infrastructure
Hyderabad, India +91-906******* *****************@*****.*** https://www.linkedin.com/in/sai-teja-kotikalapudicontacttejajain/ github.com/saiteja-kotikalapudi PROFESSIONAL SUMMARY
Senior DevOps/SRE/Cloud Engineer with 7years of experience building, automating, and operating production-grade cloud-native infrastructure on AWS at enterprise scale. Currently at DAZN (world's leading sports streaming platform) managing 15 EKS clusters and 500+ microservices serving millions of concurrent users across 200+ countries. Hands-on experience provisioning and managing NVIDIA GPU nodes on Kubernetes for AI/ML workloads, deploying LLM (Llama 3, Mistral) and VLM (LLaVA, InternVL) models in production environments, and implementing agentic AI workflow pipelines using LangChain and CrewAI. Expert in Terraform, Kubernetes
(EKS/OpenShift), CI/CD (Jenkins, GitLab, ArgoCD), observability (Prometheus, Grafana, OpenTelemetry), and SRE practices (SLI/SLO, error budgets, chaos engineering). Delivered 81% faster deployments, 35% cost reduction ($180K/year), 99.95% uptime, and 65% MTTR reduction.
PROFESSIONAL EXPERIENCE
Senior DevOps / SRE Engineer
DAZN — ) November 2023 – Present Global Sports Streaming (200+ Countries, Millions of Users Hyderabad, India Platform, Infrastructure & AI/ML:
■ Architected multi-region AWS infrastructure using Terraform (50+ modules), Terragrunt, and Helm — managing 200+ EC2, 50+ RDS
(Multi-AZ), and 15 EKS clusters with 99.95% SLA.
■ Provisioned and managed NVIDIA GPU nodes on EKS clusters for AI/ML workloads — deployed LLM (Llama 3, Mistral) and VLM (LLaVA, InternVL) models in production Kubernetes environments with optimized CUDA scheduling and GPU resource quotas.
■ Implemented agentic AI workflow deployments using LangChain and CrewAI on GPU-backed Kubernetes pods — enabling autonomous multi-step AI agent pipelines for real-time video analytics and content intelligence.
■ Operated EKS/OpenShift running 500+ microservices with HPA, VPA, Karpenter, Istio service mesh, and zero-downtime deployments — resource utilization improved 45%.
■ Configured ALB with Auto Scaling (target tracking, predictive scaling) — 40% off-peak savings while maintaining SLOs during 10x traffic spikes.
■ Managed multi-VPC networking with Transit Gateway, VPC peering, NAT gateways, and Route53 failover for cross-region high availability. CI/CD & Automation:
■ Built CI/CD using Jenkins, GitLab CI/CD, and ArgoCD GitOps with security scanning (SonarQube, Trivy, Snyk), Docker builds, and blue/green/canary delivery — releases reduced from 2 weeks to 2 days, 100+ deployments/week.
■ Built automated CI/CD pipelines for LLM/VLM model versioning, containerization, and deployment — reducing AI model release cycles from days to hours with automated rollback on inference degradation.
■ Developed 50+ Python/Boto3 scripts and 30+ Ansible playbooks for resource management, cost optimization, AMI building (Packer), and toil reduction — saving 20+ hours/week.
■ Implemented Helm/Kustomize overlays for standardized GitOps deployments across all environments.
■ Integrated CodeBuild/CodeDeploy/CodePipeline for serverless and EC2 automation with rollback triggers across 10+ services. SRE & Observability:
■ Built full observability stack using Prometheus/Thanos, Grafana (100+ dashboards), OpenTelemetry, Loki, PagerDuty — coverage across 500+ services, 200+ SLI metrics, MTTR reduced 65%.
■ Built GPU utilization dashboards (Prometheus DCGM exporter + Grafana) for real-time monitoring of AI/ML inference workloads across all GPU nodes.
■ Defined SLIs, SLOs, and error budgets for Tier-1/Tier-2 services — customer incidents reduced 40%.
■ Led 24/7 on-call with incident response (SEV1-4), blameless postmortems, action tracking — recurring incidents down 55%.
■ Conducted chaos engineering (AWS FIS, Litmus Chaos) identifying 20+ failure modes pre-production; built automated runbooks for top 25 incidents — manual intervention reduced 70%.
Security:
■ Established IAM least-privilege, WAF (10K+ blocks/day), OPA/Gatekeeper, CIS benchmark scanning.
■ Managed secrets via HashiCorp Vault and Secrets Manager with automated rotation.
■ Integrated DevSecOps: container scanning (Trivy, Snyk), SAST (SonarQube) in CI/CD — 95% vulnerabilities caught pre-production. DevOps Engineer
Apsis Technologies — Weyyak September 2021 – October 2023 (OTT Video Streaming, 5M+ Users) Bengaluru, India
■ Designed infrastructure using Terraform (30+ modules) with Terragrunt — 100% IaC coverage with drift detection and DR capabilities.
■ Managed EKS/OpenShift for 5M+ concurrent users with blue-green/canary deployments and zero-downtime rollouts.
■ Optimized CloudFront CDN with Lambda@Edge — latency reduced 60%, bandwidth costs down 45% ($95K savings).
■ Configured Jenkins/GitLab CI/CD with Maven builds, security scanning, Ansible deployments — 50+ deployments/week.
■ Managed JFrog Artifactory with RBAC, retention policies, vulnerability scanning across 8 teams.
■ Implemented ELK Stack (100GB+ daily logs) with OpenTelemetry and PagerDuty — detection time reduced 70%.
■ Built Kafka pipelines for real-time event processing with Spark analytics; defined SLIs/SLOs for streaming quality (buffering, start time, error rate).
■ Automated DR using AWS Backup, Lambda, cross-region replication — RPO 1hr, RTO 4hr. DevOps Engineer
Apsis Technologies — March 2019 – August 2021 Zurich Farmers Insurance (15M+ Policies) Bengaluru, India
■ Deployed Java/Tomcat on EC2 with Auto Scaling, health checks, rollback — 99.9% success rate.
■ Managed AWS VPC networking (subnets, security groups, NACLs, NAT, VPN) for hybrid connectivity.Automated
■ Implemented Jenkins distributed architecture (10+ agents, GitFlow) — build throughput improved 300%.
■ Implemented Prometheus/Grafana tracking 500+ metrics with baseline SLIs for critical services. Associate Software Engineer
Cogent E-Services July 2018 – February 2019 Hyderabad, India
■ Developed and optimized complex SQL queries, stored procedures, and views — query performance improved 40% through indexing and execution plan analysis.
■ Built Power BI dashboards with DAX calculations and data modeling — real-time insights for 50+ stakeholders across 5 departments.
■ Automated data processing using SQL/Python — manual reporting reduced 60%, delivery improved from weekly to daily.
■ Supported application deployments including database migrations, environment setup, and release coordination. CERTIFICATIONS & PROFESSIONAL DEVELOPMENT
■ AWS Solutions Architect Associate (SAA-C03) — Exam Scheduled Q2 2026
■ 200+ hours hands-on labs: Kubernetes, Terraform, AWS Networking, Chaos Engineering, SRE (A Cloud Guru, KodeKloud, Udemy)
■ AWS re:Invent 2023/2024 CNCF KubeCon Workshops DevOps Institute SRE Foundation ArgoCD & Ansible Courses (KodeKloud) EDUCATION
Bachelor of Commerce (B.Com) Chaitanya Degree College, Kakinada, India 2015 – 2018 DECLARATION
I hereby declare that the information furnished above is true to the best of my knowledge and belief. Place: Hyderabad
K Sai Teja