Source description
About the role
Operate and improve production infrastructure with a focus on reliability, security, performance, and cost efficiency.
Define, measure, and improve reliability using SLIs, SLOs, SLAs, error budgets, and DORA metrics .
Build and improve monitoring, alerting, dashboards, logging, and incident response processes.
Participate in incident management, root cause analysis, postmortems, and follow-up remediation.
Automate infrastructure and operational workflows using modern IaC and scripting tools.
Work closely with engineering teams to improve service reliability, deployment quality, and operational readiness.
Turn ambiguous infrastructure, reliability, and operational problems into clear, scalable, and measurable solutions.
Engage with backend codebases through code reviews, pull requests, and occasional feature or tooling work to build shared context with product engineering teams.
More at VRChat
