Site Reliability Engineer

Details of the offer

Engineering - Software (Information & Communication Technology)
Position: Site Reliability Engineer (SRE)
Location: Malaysia
Duration: 2 years (direct contract & convertible to permanent)
Experience: 6 months - 15 years (Multiple headcounts)As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability.
You will,
Design and implement resilient system architectures that support high availability and scalability.
Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
Define, track, and analyse SLOs and SLIs to ensure reliability and performance meet business needs.
Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation.
Collaborate with development and operations teams to establish best practices in system reliability and incident management.
Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery.
Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.
Requirements:
Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency.
Demonstrated experience in system architecture and design, prioritizing reliability, and scalability.
Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems.
Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management.
Strong expertise in Linux system administration.
Proven experience in troubleshooting application support issues with a focus on performance and connectivity.
Familiarity with networking concepts and effective troubleshooting techniques.
Excellent problem-solving abilities and a proactive approach to operational challenges.
Preferred Skills:
Familiarity with monitoring tools and performance optimization techniques.
Experience in scripting or automation for system administration tasks.
Knowledge of networking concepts and troubleshooting methodologies.
Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services.
Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.#J-18808-Ljbffr

Nominal Salary: To be agreed

Source: Whatjobs_Ppc

Job Function:

Engineering

Requirements

Similar offers

See more similar offers

Lead Engineer, Mechanical Design

Remote Position: No Region: Asia Country: Malaysia State/Province: Kedah City: Kulim General Overview Functional Area: Engineering Career Stream: ...

Celestica - Malasia

Published a month ago

Quantity Surveyor (Qs) Diperlukan

Kami syarikat kontraktor perkhidmatan dan pembinaan memerlukan Quantity Surveyor (QS) Projek di Sungai Petani, Kedah. Berkelulusan Diploma / Ijazah Ukur Ba...

Sri Budinar Solution Sdn Bhd - Malasia

Published a month ago

Civil Technician

Responsibilities: Work hand in hand with other working colleagues to solve maintenance issuePerform minor fixes, install appliances and equipment related to ...

Tsg Expert M Sdn Bhd - Malasia

Published a month ago

Senior Engineer Software Test Mes

In your new role you will: Become part of the Manufacturing Execution Systems (MES) team Maintain and continuously extend the set of automated software test ...

Infineon Technologies - Malasia

Published a month ago

Built at: 2024-12-12T07:37:50.849Z