Site Reliability Engineer (Mandarin Speaker)

Details of the offer

Overview:
As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability.
Key Responsibilities:
Design and implement resilient system architectures that support high availability and scalability.
Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation.
Collaborate with development and operations teams to establish best practices in system reliability and incident management.
Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery.
Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.
Requirements:
Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems.
Familiarity with monitoring tools and performance optimization techniques.
Proficiency in programming languages such as Python / Golang / Java, or similar, focusing on operational efficiency.
Demonstrated experience in system architecture and design, prioritizing reliability, and scalability.
Strong expertise in Linux system administration.
Experience in scripting or automation for system administration tasks.
Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services.
Proven experience in troubleshooting application support issues with a focus on performance and connectivity.
Familiarity with networking concepts and effective troubleshooting techniques.
Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.
For more information, kindly contact Sunny Khoo via WhatsApp at 012-5164406 or via email ****** . Thank you.#J-18808-Ljbffr

Nominal Salary: To be agreed

Job Function:

Engineering

Requirements

Similar offers

See more similar offers

Technical Service Technician

Job Description Dispenser Technician for, Technical Services To do new installation, maintenance services and Calibration for our new customers and existing ...

Sika - Selangor

Published 23 days ago

Field Service Engineer (Automation)

Job Summary: Primarily responsible for managing and maintaining automated equipment at the client site. Here is the job requirements and daily responsibiliti...

Career Wise - Selangor

Published 23 days ago

Sales Engineer (Prefer Man)

The Sales Engineer is responsible for driving sales growth by providing technical support and solutions to customers in the construction and machinery sector...

Litec Machinery Sdn Bhd - Selangor

Published 23 days ago

Installation Technician

Company Description Jobs Xpert Brilliant Sdn Bhd is a Recruitment Firm. We assist our clients with all the talent search processes. Our client is a provider...

Jobs Xpert Brilliant Sdn Bhd - Selangor

Published 23 days ago

Built at: 2024-12-25T15:16:11.215Z