Site Reliability Engineer

Details of the offer

Design and implement resilient system architectures that support high availability and scalability. Develop automation tools and scripts to enhance operational efficiency and reduce manual effort. Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs. Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation. Collaborate with development and operations teams to establish best practices in system reliability and incident management. Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines). Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery. Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements. Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance. Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency. Demonstrated experience in system architecture and design, prioritizing reliability, and scalability. Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems. Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management. Strong expertise in Linux system administration. Proven experience in troubleshooting application support issues with a focus on performance and connectivity. Familiarity with networking concepts and effective troubleshooting techniques. Excellent problem-solving abilities and a proactive approach to operational challenges. Ability to work independently while effectively collaborating within a team environment.

Nominal Salary: To be agreed

Source: Grabsjobs_Co

Job Function:

Engineering

Requirements

Similar offers

See more similar offers

Quantity Surveyor

Prepare detailed cost plans for the project, including materials, labor, and overheads.Prepare and issue tender documents, including bills of quantities (BQ)...

Srk Builders Sdn Bhd - Kuala Lumpur

Published a month ago

Quantity Surveyor

The Quantity Surveyor will be responsible for managing the costs relating to projects. This includes new builds, renovations, and maintenance work. The role ...

Chs Interior Decoration Sdn Bhd - Kuala Lumpur

Published a month ago

Tender Engineer

Job Responsibilities Prepare and manage tender documentation, including technical proposals and method statements Utilize AutoCAD and other documentation sof...

Chis International Technical Resources Sdn Bhd - Kuala Lumpur

Published a month ago

Project Engineer (Tender)

Job Responsibilities Prepare and manage tender documentation, including technical proposals and method statements Utilize AutoCAD and other documentation sof...

Chis International Technical Resources Sdn Bhd - Kuala Lumpur

Published a month ago

Built at: 2024-11-21T22:07:51.069Z