Padmi
GRAIL logo
GRAIL

multi-cancer early detection (MCED) · next-generation sequencing (NGS)

Staff Site Reliability Engineer (SRE) | Dev Ops Engineer #4770

United States · OnsitePosted 3 months ago
InfrastructureStaff+Full TimeH-1B track record
Apply at GRAIL

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Design, build, and operate highly available, fault-tolerant cloud infrastructure across AWS, GCP, and/or Azure

Architect and maintain scalable CI/CD pipelines and deployment frameworks for enterprise-grade software delivery

Lead infrastructure-as-code adoption and maturity using tools such as Terraform, CloudFormation, and Ansible

Own Kubernetes reliability across multi-cluster environments, including upgrades, scaling, and workload lifecycle management

Establish and evolve observability platforms (metrics, logs, traces) and define SLO/SLI frameworks across teams

Lead incident response for critical outages, drive root cause analysis, and implement preventative improvements

Optimize infrastructure for cost, performance, and scalability, partnering closely with engineering and finance stakeholders

Define and enforce DevOps, reliability, and security best practices across the organization

Partner cross-functionally with engineering, data, QA, security, and IT teams to design resilient systems

Mentor engineers and contribute to technical leadership through design reviews, standards, and knowledge sharing

These responsibilities summarize the role’s primary responsibilities and are not an exhaustive list. They may change at the company’s discretion.

What Success Looks Like in Your First Year

Conduct a comprehensive assessment of the current infrastructure, drive infrastructure-as-code adoption to 95%+ across critical systems, and establish clear health and reliability baselines for the Kubernetes platform

Standardize observability using modern tooling and implement an SLO/SLI framework adopted across multiple product teams, including defined SLAs for critical data systems

Strengthen security and compliance posture across cloud environments by implementing consistent baselines, launching a compliance-as-code framework, and reducing mean time to resolution (MTTR) for production incidents

Define, document, and drive adoption of engineering standards, best practices, and operational guidelines across platform and product teams

Develop and align stakeholders on a forward-looking platform reliability and infrastructure roadmap

Demonstrate measurable mentorship and technical leadership impact across the engineering organization

Evaluate and provide recommendations on emerging infrastructure needs, including support for AI/ML and advanced data workloads

More at GRAIL

Related open roles

View all roles