Source description
About the role
Team leadership: Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of AI model training and inference systems. Observability: Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into model serving pipelines and infra. Automation & Tooling: Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem CPU+GPU environments. Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements. Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments. Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows. Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with Site Reliability Engineering, DevOps, or Infrastructure Engineering Leadership roles AND 8+ years experience with Kubernetes, Docker, and container orchestration, AND 6+ years experience with programming/scripting skills not limited to Python, Go, or Bash Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience AND 10+ years experience with Kubernetes, Docker, and container orchestration, AND 10+ years' experience with public cloud platforms like Azure/AWS/GCP and infrastructure-as-code OR equivalent experience 6+ years people management experience. 8+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.). Knowledge of CI/CD pipelines for Inference and ML model deployment. Solid knowledge of distributed systems, networking, and storage. Experience running large-scale GPU clusters for ML/AI workloads (preferred). Familiarity with ML training/inference pipelines. Experience with high-performance computing (HPC) and workload schedulers ( Kubernetes operators). Background in capacity planning & cost optimization for GPU-heavy environments
More at Microsoft
Related open roles
Cloud Solution Architecture
São Paulo · Onsite
Cloud Solution Architect - Cloud & AI Platforms (CAIP) Factory
United States · Onsite
Director Architecture, Azure Management Solutions
San Francisco Bay Area · Seattle · Austin · Onsite
Data Center Critical Environment Technician Manager
Toronto · Onsite
Senior Data Center Technician - Night Shift
Phoenix · Onsite
Critical Environment Operations Specialist
United States · Onsite