- Proven experience deploying and managing Kubernetes clusters for AI/ML workloads. Experience of at scale deployments with Azure Kubernetes. Experience level
- 5 Years or more Positions
- 2 Proven experience deploying and managing Kubernetes clusters for AI/ML workloads.
- Experience of at scale deployments with Azure Kubernetes Service, RedHat OpenShift, Microk8s and Helm Charts.
- Expertise with infrastructure and resource management and virtualization tools such as VMWare/EXSi, KVM, Ansible, Redfish.
- Strong understanding of Run:AI platform, including job scheduling, quota management, and GPU virtualization.
- Knowledge of NVIDIA AI Enterprise components including, NIM, NeMO, TAO, Triton and Nucleus Servers
- Familiarity with DGX systems, Jetson, and NVIDIA’s AI Factory components.
- Proficiency in Python, C++, and optionally .NET/C# for enterprise integration.