Source description
About the role
Lead end-to-end software architecture and performance optimization for AI accelerator platforms. Prototype and validate software capabilities across kernels, compiler and runtime layers, distributed training and inference frameworks, and serving infrastructure. Analyze workload behavior at scale to identify performance bottlenecks, numerical correctness issues, and system-level efficiency opportunities. Use workload insights to guide hardware-software co-design decisions across architecture, silicon, systems software, networking, and product teams. Define requirements, evaluate tradeoffs, and deliver practical solutions for production-scale AI systems. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. These requirements include, but are not limited to, the following specialized security screenings: PhD in Computer Science,Computer Architecture, Electrical Engineering, Machine Learning, High-Performance Computing, or a related field OR Master's Degree in Computer Science, Electrical Engineering, Computer Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Computer Science, Electrical Engineering, Computer Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience. Experience designing, building, or optimizing systems software for AI, machine learning, high-performance computing, or distributed systems. Experience analyzing large-scale AI training runs, including loss curves, convergence behavior, gradient flow, activation statistics, and numerical stability. Experience debugging training correctness issues such as gradient divergence, NaNs/Infs, optimizer behavior, mixed-precision instability, distributed synchronization bugs, or hardware/software numerical differences.
More at Microsoft