Source description
About the role
Ability to drive both Design for Reliability (DfR) and Reliability, Availability & Maintainability (RAM) processes with consistency and rigor for both current and next generation Cloud & AI hardware and infrastructure solutions. Work across functionally across different disciplines (Hardware, Firmware, Architecture, Safety, Serviceability etc.) to influence business and operational decisions. Works with key stakeholders to set appropriate reliability and availability targets and allocates the targets to lower-level sub-systems and components. Develops appropriate tests at system, sub-system and/or/component as needed to demonstrate the budgeted reliability targets. Carries out effective Design Failure Modes & Effects Analysis (DFMEA) to identify and mitigate critical risks via design, operational and diagnostic improvements. Utilizes different modeling methods such as Reliability Block Diagram (RBD), Markov, Discreet Event Simulations (DES) and Fault Tree Analysis (FTA) to quantify risks and/or support business decisions. Develops Prognostics & Health Management (PHM) models for Remaining Useful Life (RUL) prediction based on telemetry data. Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 2+ years technical engineering experience OR Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 4+ years technical engineering experience OR Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience OR 12+ years relevant technical engineering experience. M.S. in Electrical or Electronic Engineering, Reliability Engineering, or equivalent discipline 5+ years of experience performing Reliability Engineering on Complex Systems Background in Applied Reliability Engineering statistics for repairable and non-repairable systems. Experienced in using software tools such as Reliasoft, JMP and python scripts for performing different reliability analyses. Passionate individual who is able to work collaboratively in a team environment and across internal divisions, industry (OEM, ODM), and with customers Prior experience driving reliability efforts for Cloud & AI hardware including infrastructure (Networking, Power and/or Cooling). Experienced in Prognostics & Health Management (PHM) techniques Certified Reliability Engineer (CRE)
More at Microsoft
Related open roles
Fabric Interconnect Design Verification Engineer
San Francisco Bay Area · Seattle · Austin · Onsite
Senior Firmware Engineer
San Francisco Bay Area · Seattle · Portland · Onsite
Principal Signal Integrity Simulation Engineer
Seattle · Onsite
Firmware Engineer II
San Francisco Bay Area · Seattle · Austin · Onsite
Senior Silicon Test Engineer
San Francisco Bay Area · Seattle · Austin · Onsite
Critical Environment Electrical Engineer
Phoenix · Onsite