Source description
About the role
Execution & Collaboration
Respond to production incidents, perform triage and troubleshooting, and contribute to post-incident analysis.
Identify and automate manual processes to improve efficiency and reduce risk.
Enhance and evolve monitoring tools and platforms to improve observability.
Promote and apply best practices for reliability, scalability, and performance across engineering.
Implement and support cloud automation using Terraform, Ansible, or CloudFormation.
Work within change management protocols to provide maximum uptime for production systems.
Participate in on-call rotation, providing 24x7 support for incidents and contributing to root cause analysis.
Partner with developers, architects, vendors, and IT teams to ensure reliable system operations.
Research and remediate vulnerabilities in coordination with security teams.
Maintain documentation of infrastructure, monitoring, runbooks, and incident response procedures.
Standards & Process
Apply company policies and procedures when handling operational tasks and incidents.
Suggest and implement improvements to operational processes and monitoring practices.
Contribute to technical diagrams, documentation, and runbooks for system reliability.
Learning & Growth
Expand expertise in cloud services (Azure, AWS, or GCP) and container platforms (EKS, ECS, AKS).
Build proficiency with observability and monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios).
Develop scripting and automation skills using Python, Bash, PowerShell, or similar.
Participate in planning discussions by contributing technical input on system stability and reliability.
