Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Site Reliability Engineer II - CTJ - Top Secret

Seattle · Washington DC · OnsitePosted 2 months ago
InfrastructureMid-levelFull TimeH-1B track record
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Live Site Operations: Serve as a Designated Responsible Individual (DRI) in a 24x7 on-call rotation, monitoring service health and responding to incidents within SLA timelines. Automation & Deployment: Contribute to automation efforts and validate code functionality in non-production environments to ensure smooth deployments. Compliance & Security: Support compliance processes by verifying security, privacy, and accessibility standards during onboarding of new technologies. Continuous Learning: Stay current with industry trends and internal tools to improve reliability, performance, and observability at scale. Engineering Best Practices: Apply proven development and scaling practices to meet performance and customer requirements. Cross-Team Collaboration: Communicate effectively with engineering partners to align on goals and deliver user-centric solutions. Incident Response & Postmortems: Address complex live site issues, implement mitigations, and document learnings through postmortems. Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration These requirements include, but are not limited to the following specialized security screenings: Candidates must have an active TS and be willing and eligible to upgrade to TS/SCI (with polygraph) or have an active TS/SCI and be willing and eligible to upgrade to TS/SCI (with polygraph). This role will require candidates to maintain the TS/SCI (with polygraph) clearance. Failure to maintain or obtain the appropriate clearance and/or customer screening requirements may result in employment action up to and including termination. Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment. Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 2+ years technical experience working with large-scale cloud or distributed systems. Demonstrated experience applying software engineering principles to production systems, including designing, building, or improving services and platforms. Proficiency in one or more programming languages such as C#, Go, Java, or Python, with the ability to develop and maintain production-quality code. Experience with automation that results in measurable improvements (e.g., reduced toil, fewer manual steps, improved system reliability). Experience with debugging and troubleshooting complex distributed systems in production environments. Ability to independently identify problems and implement solutions that improve system reliability and operational efficiency. Hands-on experience with CI/CD pipelines, testing, deployment, and reliability tooling.

More at Microsoft

Related open roles

View all roles