Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Pre-Training Infrastructure

United States · OnsitePosted 8 months ago
InfrastructureSeniorFull TimeH-1B track record
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Design, implement, test, and optimize distributed training infrastructure in Python and C++ for large-scale GPU clusters. Profile, benchmark, and debug performance bottlenecks across compute, memory, networking, and storage subsystems. Optimize collective communication libraries (e.g., NCCL) for emerging NVLink and InfiniBand topologies. Collaborate with hardware teams to optimize for next-generation accelerators (NVIDIA, AMD, and beyond). Gather data and insights to develop the pretraining compute roadmap. Care deeply about conversational AI and its deployment. Actively contribute to the development of AI models powering our innovative products. Find solutions to overcome roadblocks and deliver your work to users quickly and iteratively. Enjoy working in a fast-paced, design-driven product development cycle. Embody our Culture and Values. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience in distributed computing and large-scale systems. Experience with GPU programming (CUDA, NCCL) and frameworks such as PyTorch. Proven ability to profile, benchmark, and optimize performance-critical systems. Experience in leading technical projects and supporting architectural decisions with data. Experience building infrastructure for large-scale machine learning or generative AI workloads. Experience in networking (InfiniBand, NVLink), storage systems, or distributed training parallelisms. Track record of contributing to high-performance computing or large-scale AI infrastructure projects.

More at Microsoft

Related open roles

View all roles
Pre-Training Infrastructure at Microsoft · Padmi