Source description
About the role
Live Site Operations: Serve as a Designated Responsible Individual (DRI) in a 24x7 on-call rotation, monitoring service health and responding to incidents within SLA timelines. Automation & Deployment: Contribute to automation efforts and validate code functionality in non-production environments to ensure smooth deployments. Compliance & Security: Support compliance processes by verifying security, privacy, and accessibility standards during onboarding of new technologies. Continuous Learning: Stay current with industry trends and internal tools to improve reliability, performance, and observability at scale. Engineering Best Practices: Apply proven development and scaling practices to meet performance and customer requirements. Cross-Team Collaboration: Communicate effectively with engineering partners to align on goals and deliver user-centric solutions. Incident Response & Postmortems: Address complex live site issues, implement mitigations, and document learnings through postmortems. Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration These requirements include, but are not limited to the following specialized security screenings: Candidates must have an active TS and be willing and eligible to upgrade to TS/SCI (with polygraph) or have an active TS/SCI and be willing and eligible to upgrade to TS/SCI (with polygraph). This role will require candidates to maintain the TS/SCI (with polygraph) clearance. Failure to maintain or obtain the appropriate clearance and/or customer screening requirements may result in employment action up to and including termination. Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment. Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 2+ years technical experience working with large-scale cloud or distributed systems. Demonstrated experience applying software engineering principles to production systems, including designing, building, or improving services and platforms. Proficiency in one or more programming languages such as C#, Go, Java, or Python, with the ability to develop and maintain production-quality code. Experience with automation that results in measurable improvements (e.g., reduced toil, fewer manual steps, improved system reliability). Experience with debugging and troubleshooting complex distributed systems in production environments. Ability to independently identify problems and implement solutions that improve system reliability and operational efficiency. Hands-on experience with CI/CD pipelines, testing, deployment, and reliability tooling.
More at Microsoft
Related open roles
Cloud Solution Architecture
São Paulo · Onsite
Cloud Solution Architect - Cloud & AI Platforms (CAIP) Factory
United States · Onsite
Director Architecture, Azure Management Solutions
San Francisco Bay Area · Seattle · Austin · Onsite
Data Center Critical Environment Technician Manager
Toronto · Onsite
Senior Data Center Technician - Night Shift
Phoenix · Onsite
Critical Environment Operations Specialist
United States · Onsite