注册一亩三分地论坛,查看更多干货!
您需要 登录 才可以下载或查看附件。没有帐号?注册账号 
x
Site Reliability Engineer (SRE)
Location: San Francisco, CA
Employment Type: Full-time. check 1point3acres for more.
About the Role
. 1point3acresWe are looking for a Site Reliability Engineer (SRE) to support the deployment, operation, maintenance, monitoring and troubleshooting of large-scale Kubernetes-based AI training clusters.
You will be responsible for improving cluster reliability, automation, resource utilization and operational efficiency across AI infrastructure environments.
Responsibilities. Χ
- Deploy, operate and maintain Kubernetes-based AI training clusters.
- Design and develop monitoring and automation capabilities for cluster management platforms.
- Continuously improve cluster management and control capabilities.
- Troubleshoot issues related to containers, Linux operating systems, networks and storage.
- Analyze cluster performance and resource utilization.
- Manage business-level resource quotas and support capacity planning.
- Participate in on-call operations and promptly respond to infrastructure incidents and user issues.
- Perform root-cause analysis and implement solutions to improve system reliability.
- Collaborate with engineering and infrastructure teams to improve cluster performance and operational efficiency.
Requirements
- Bachelor's degree or above in Computer Science, Computer Engineering or a related field.
- 3+ years of industry experience in SRE, DevOps, Infrastructure Engineering, Platform Engineering or a related area.
- Strong Linux operation, maintenance, troubleshooting and performance analysis skills.
- Proficiency in Python, Go, Shell or similar programming/scripting languages.
- Strong understanding of Kubernetes architecture and hands-on experience with Kubernetes.
- Practical experience with Kubernetes CNI, CSI, Load Balancing, networking and storage.
- Strong troubleshooting and problem-solving skills.
Preferred Qualifications. check 1point3acres for more.
- Experience building or optimizing large-scale AI/ML training clusters.
- Experience with GPU infrastructure, AI accelerators or HPC.
- Experience with Kubernetes scheduling and resource management.
- Experience with monitoring and observability systems.
- Experience with large-scale distributed systems.
- Strong communication and cross-functional collaboration skills.
Keywords: SRE · Kubernetes · Linux · Python · Go · CNI · CSI · GPU · AI Infrastructure · HPC · Cloud · Distributed Systems · DevOps · Platform Engineering
How to Apply
Please send your resume to jacky.guo@renruihr.com with the subject line “Site Reliability Engineer – San Francisco”. |