Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Cloud Solution Architecture - Infrastructure

Singapore · OnsitePosted 28 days ago
InfrastructureSeniorFull Time
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Trusted Advisor & Customer Advocacy Act as a trusted technical advisor, helping customers improve the reliability, resiliency, security, performance, and operational maturity of mission-critical workloads running on Azure. Advise customers and stakeholders on architecture, operations, and best practices aligned with the Azure Well-Architected Framework. Communicate complex technical concepts and recommendations in clear, actionable terms to both technical and executive audiences. Perform proactive health assessments, risk reviews, and operational analysis to identify opportunities for improvement and escalation prevention. Develop and maintain deep technical understanding of assigned customer environments, architectures, dependencies, and mission-critical workloads. Deliver onboarding assessments and help define service delivery and improvement plans aligned with customer objectives. Operate effectively within a global follow-the-sun support model, collaborating with teams across multiple regions and time zones to ensure continuity of service for mission-critical workloads. Collaborate effectively across teams, cultures, and organizational boundaries to drive customer success and operational improvements. Lead complex troubleshooting efforts across infrastructure, platform, and application layers, including critical and high-severity incidents. Operate effectively in high-stakes, customer-impacting incidents, combining platform expertise and customer business context to accelerate mitigation, recovery, and restoration of service. Facilitate Root Cause Analysis (RCA) activities for critical incidents, helping customers identify corrective and preventative actions that reduce future risk. Analyze support cases, operational telemetry, incident trends, and platform events to identify recurring risks and recommend proactive remediation measures. Drive reduction of reactive operational demand through reliability-focused recommendations, operational maturity improvements, resiliency best practices, and service optimization initiatives. Promote operational excellence across reliability, availability, security, performance, recoverability, and capacity management. Improvements in workload reliability, resiliency, security, and operational maturity. Adoption of recommended architecture, operational practices, and remediation plans. Reduction in customer-impacting incidents, repeat escalations, and operational risk. Faster mitigation and recovery of critical incidents. Effective coordination across global teams, ensuring seamless customer support and operational continuity across regions and time zones. Increased customer satisfaction and trusted advisor influence. Positive business outcomes through improvements in reliability, security, performance, capacity management, and service resilience. Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related field, AND 7+ years of relevant experience supporting mission-critical production environments; OR equivalent practical experience Experience supporting mission-critical production environments Experience leading or coordinating Sev A / P1 incidents Experience providing recommendations to enterprise customers Experience improving reliability, resiliency, performance, security, or operational maturity Experience working across multiple time zones and globally distributed teams Experience coordinating multiple technical teams to resolve customer issues Experience with telemetry, monitoring, logging, and root-cause analysis Experience with DR, HA, BCP, and recovery planning

More at Microsoft

Related open roles

View all roles