Source description
About the role
Hands on debug in data center (onsite and virtual) Develop and implement a robust supplier quality management strategy to ensure the data center hardware is manufactured at the highest level of quality standards. Leadership to work across data centers, development, and supplier to resolve critical & high severity issues. Conduct hands on debug in global data centers (onsite and virtual) including GPU sub-system failure analysis. Drive the continuous improvement process based on Root Cause Analysis (RCA) and identified opportunities. Manage multiple NPI builds and quality phase-gate deliverables for the manufacturing team throughout the engineering development lifecycle, from concept through production readiness. Establish Critical-to-Quality performance metrics to measure and improve product quality. Act as the voice of quality in the hardware change management process, ensuring quality requirements are considered and met. Doctorate in Electrical Engineering, Computer Engineering, or related field AND 3+ years technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, or related field AND 6+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience. 12+ years relevant technical engineering experience OR Master's Degree in Electrical Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience. 12+ years of work experience in managing product quality in the electronic industry. 5+ years of direct engineering experience in hardware system issue resolution for GPU systems. Versed in filtering through applicable debug data, like telemetry and logs to identify and investigate HW failure signatures. Patent or track record of engineering excellency. Experience with Liquid Cooling Systems in Data Centers 12+ years of experience in working with the modern server architectures - includes understanding of GPU, GPU system hardware, Memory or CPU and methods for failure analysis, debugging or validation. 12+ years of proven success of leading resolution of critical quality issues across data centers 8+ years of system level server debugging with an understanding of power, system and network environments 3+ years of direct GPU related engineering experience in issue debug/test log review. Leadership skills and ability to collaborate with diverse teams and drive a call to action.
More at Microsoft
Related open roles
Fabric Interconnect Design Verification Engineer
San Francisco Bay Area · Seattle · Austin · Onsite
Senior Firmware Engineer
San Francisco Bay Area · Seattle · Portland · Onsite
Principal Signal Integrity Simulation Engineer
Seattle · Onsite
Firmware Engineer II
San Francisco Bay Area · Seattle · Austin · Onsite
Senior Silicon Test Engineer
San Francisco Bay Area · Seattle · Austin · Onsite
Critical Environment Electrical Engineer
Phoenix · Onsite