Source description
About the role
Team Introduction: OSE builds software that helps TikTok engineering teams run services with autonomy and confidence. We are moving from reactive operational support toward self-service reliability platforms, trusted service context, and workflows that scale across the organization. Our work spans the reliability lifecycle: helping teams define quality alerts, route and investigate operational signals, manage incidents, and turn learning into durable improvements. We build platforms that make reliable operations easier for both engineers and on-call teams.
OSE sits at the intersection of software engineering and site reliability. We build platform products for monitoring, incident operations, and service metadata. We partner deeply with high-impact engineering teams and design automation that remains accountable to the people who operate production systems.
Engineers also embed with partner teams through a scheduled 12-hour on-call rotation. This builds direct understanding of live operational workloads and informs better platform design. The rotation is no more frequent than once every four weeks. We are building toward AI reliability workflows. In this role, you will help make engineering and operational work more agent-ready through clear product contracts, safe automation, and evidence-driven iteration. AI can assist with triage, drafting, and bounded actions. Engineers remain responsible for validation, production decisions, and outcomes.
- Build and evolve reliability platforms that engineers use daily: APIs, workflows, configuration surfaces, and operator experiences.
- Improve the alarm lifecycle from authoring through triage, handling, remediation, and feedback into better rules and metadata.
- Design for safe automation by applying appropriate dry-run paths, approval gates, rollback options, and traceability to agent-assisted or scripted actions.
- Keep humans accountable for policy decisions, high-risk changes, and production outcomes. Agents can propose and assist. People own merges, rollouts, and incident resolution.
- Validate before you ship: tests, static checks, staged rollout, and self-verification evidence where the platform supports it.
- Instrument workflows so adoption, quality, and failure modes are observable without confusing activity with causality.
- Partner with internal teams to prove patterns, gather feedback, and hand off repeatable workflows so partners can self-serve.
- Embed with partner teams through a 12-hour on-call rotation to understand live operational workloads and translate recurring friction into platform improvements. Rotations occur no more frequently than once every four weeks.
- Participate in reliability culture: incident learning, post-incident follow-up, on-call hygiene, and documentation that downstream teams and agents can trust.
- Contribute to AI platform capabilities where appropriate: agent-usable skills, product contracts, read-only query paths, and governed write paths with clear permissions.
More at TikTok
Related open roles
Software Engineer, Machine Learning Infrastructure
San Francisco Bay Area · Onsite
Site Reliability Engineer Graduate (Video & Edge Global Engineering) - 2027 Start (BS/MS)
Sydney · Onsite
Site Reliability Engineer - Global E-Commerce
Singapore · Onsite
Site Reliability Engineer - Cloud Infrastructure
Dublin · Onsite
Site Reliability Engineer- Video and Edge Global Engineering (English & Mandarin Speaking)
Sydney · Onsite
Site Reliability Engineer Intern (Compute Platform) - 2026 Summer (BS/MS)
San Francisco Bay Area · Onsite