Source description
About the role
Design and build self-service platform, such as Message Queue service, Observability service, etc, that would enable developers to focus on coding, testing, and managing their cloud-native applications.
Automate and simplify lifecycle operations, such as provisioning and scaling using Infrastructure-as-Code and GitOps workflow.
Implement observability and alerting systems for the platform services, using tools such as Prometheus, Grafana, or Elastic Observability to meet service-level objectives (SLOs).
Collaborate with security teams to integrate and enforce security controls and compliance requirements of the platform services.
Work with application teams to improve platform usability, streamline onboarding, and reduce operational toil.
Respond to incidents and perform post-incident reviews , driving continuous improvement and operational excellence.
Contribute to the reliability engineering culture , fostering shared responsibility for system availability and performance.
Collaborate with cross functional teams to define platform service requirements.
More at Centre for Strategic Infocomm Technologies
