Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Senior System Reliability Engineer

Seattle · OnsitePosted 6 days ago
HardwareSeniorFull TimeH-1B track record
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Ability to drive both Design for Reliability (DfR) and Reliability, Availability & Maintainability (RAM) processes with consistency and rigor for both current and next generation Cloud & AI hardware and infrastructure solutions. Work across functionally across different disciplines (Hardware, Firmware, Architecture, Safety, Serviceability etc.) to influence business and operational decisions. Works with key stakeholders to set appropriate reliability and availability targets and allocates the targets to lower-level sub-systems and components. Develops appropriate tests at system, sub-system and/or/component as needed to demonstrate the budgeted reliability targets. Carries out effective Design Failure Modes & Effects Analysis (DFMEA) to identify and mitigate critical risks via design, operational and diagnostic improvements. Utilizes different modeling methods such as Reliability Block Diagram (RBD), Markov, Discreet Event Simulations (DES) and Fault Tree Analysis (FTA) to quantify risks and/or support business decisions. Develops Prognostics & Health Management (PHM) models for Remaining Useful Life (RUL) prediction based on telemetry data. Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 2+ years technical engineering experience OR Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 4+ years technical engineering experience OR Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience OR 12+ years relevant technical engineering experience. M.S. in Electrical or Electronic Engineering, Reliability Engineering, or equivalent discipline 5+ years of experience performing Reliability Engineering on Complex Systems Background in Applied Reliability Engineering statistics for repairable and non-repairable systems. Experienced in using software tools such as Reliasoft, JMP and python scripts for performing different reliability analyses. Passionate individual who is able to work collaboratively in a team environment and across internal divisions, industry (OEM, ODM), and with customers Prior experience driving reliability efforts for Cloud & AI hardware including infrastructure (Networking, Power and/or Cooling). Experienced in Prognostics & Health Management (PHM) techniques Certified Reliability Engineer (CRE)

More at Microsoft

Related open roles

View all roles