查看: 20| 回复: 0
跳转到指定楼层
上一主题 下一主题
收起左侧

[找工就业] Site Reliability Engineer (SRE) | AI Infrastructure | Kubernetes | San Francisco

全局:

2026(7-9月)-CS硕士+3-5年 | 猎头|BayArea湾区 码农类General全职@RenRui

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x
Site Reliability Engineer (SRE)
Location: San Francisco, CA
Employment Type: Full-time. check 1point3acres for more.
About the Role
. 1point3acresWe are looking for a Site Reliability Engineer (SRE) to support the deployment, operation, maintenance, monitoring and troubleshooting of large-scale Kubernetes-based AI training clusters.
You will be responsible for improving cluster reliability, automation, resource utilization and operational efficiency across AI infrastructure environments.
Responsibilities. Χ
  • Deploy, operate and maintain Kubernetes-based AI training clusters.
  • Design and develop monitoring and automation capabilities for cluster management platforms.
  • Continuously improve cluster management and control capabilities.
  • Troubleshoot issues related to containers, Linux operating systems, networks and storage.
  • Analyze cluster performance and resource utilization.
  • Manage business-level resource quotas and support capacity planning.
  • Participate in on-call operations and promptly respond to infrastructure incidents and user issues.
  • Perform root-cause analysis and implement solutions to improve system reliability.
  • Collaborate with engineering and infrastructure teams to improve cluster performance and operational efficiency.
Requirements
  • Bachelor's degree or above in Computer Science, Computer Engineering or a related field.
  • 3+ years of industry experience in SRE, DevOps, Infrastructure Engineering, Platform Engineering or a related area.
  • Strong Linux operation, maintenance, troubleshooting and performance analysis skills.
  • Proficiency in Python, Go, Shell or similar programming/scripting languages.
  • Strong understanding of Kubernetes architecture and hands-on experience with Kubernetes.
  • Practical experience with Kubernetes CNI, CSI, Load Balancing, networking and storage.
  • Strong troubleshooting and problem-solving skills.
Preferred Qualifications. check 1point3acres for more.
  • Experience building or optimizing large-scale AI/ML training clusters.
  • Experience with GPU infrastructure, AI accelerators or HPC.
  • Experience with Kubernetes scheduling and resource management.
  • Experience with monitoring and observability systems.
  • Experience with large-scale distributed systems.
  • Strong communication and cross-functional collaboration skills.
Keywords: SRE · Kubernetes · Linux · Python · Go · CNI · CSI · GPU · AI Infrastructure · HPC · Cloud · Distributed Systems · DevOps · Platform Engineering

How to Apply
Please send your resume to jacky.guo@renruihr.com with the subject line “Site Reliability Engineer – San Francisco”.

上一篇:Anthropic Fireworks RS面经交流
下一篇:付费求高通在职微波工程师帮忙 Mock Interview
您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表