Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

HPC Operations Engineering Manager

United States · OnsitePosted 5 months ago
InfrastructureStaff+Full TimeH-1B track record
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Team leadership: Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of AI model training and inference systems. Observability: Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into model serving pipelines and infra. Automation & Tooling: Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem CPU+GPU environments. Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements. Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments. Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows. Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with Site Reliability Engineering, DevOps, or Infrastructure Engineering Leadership roles AND 8+ years experience with Kubernetes, Docker, and container orchestration, AND 6+ years experience with programming/scripting skills not limited to Python, Go, or Bash Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience AND 10+ years experience with Kubernetes, Docker, and container orchestration, AND 10+ years' experience with public cloud platforms like Azure/AWS/GCP and infrastructure-as-code OR equivalent experience 6+ years people management experience. 8+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.). Knowledge of CI/CD pipelines for Inference and ML model deployment. Solid knowledge of distributed systems, networking, and storage. Experience running large-scale GPU clusters for ML/AI workloads (preferred). Familiarity with ML training/inference pipelines. Experience with high-performance computing (HPC) and workload schedulers ( Kubernetes operators). Background in capacity planning & cost optimization for GPU-heavy environments

More at Microsoft

Related open roles

View all roles