查看: 272| 回复: 1
跳转到指定楼层
上一主题 下一主题
收起左侧

[我司要招人] [Hiring] Inference Optimization Engineer | LLM / CUDA / vLLM | 北美 / 国内均可

全局:

2026(10-12月)-CS硕士+3-5年 | 内推| 码农类General全职@网宿科技

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x
工作地点:北美或中国一线城市,灵活可选

我们正在寻找 Inference Optimization Engineer / 大模型推理优化工程师,主要负责大语言模型及多模态生成模型在 GPU 集群上的高性能推理与系统优化。

主要工作

  • 构建和优化 LLM / 多模态模型推理引擎,实现低延迟、高吞吐的生产级部署
  • CUDA / Triton Kernel 开发与性能优化
  • vLLM / SGLang 等推理框架开发与优化
  • Multi-GPU / Multi-node 分布式推理、计算与通信优化
  • Quantization / Sparsification 等模型压缩与推理加速
  • 基于 C++ / Python 构建高性能推理系统
  • 分析真实生产环境中的 GPU、显存、通信、调度等性能瓶颈,并进行软硬件协同优化
我们希望你

  • 3+ 年 LLM inference、GPU optimization、AI Infra、HPC 或相关经验
  • 熟练使用 Python,熟悉 C++
  • 在以下至少一个方向有较深入的实践经验:
    • CUDA / Triton / TileLang 等 GPU 编程与 Kernel 优化
    • vLLM / SGLang 等推理框架
    • Quantization / Sparsification / Distillation
    • Multi-GPU / Multi-node 分布式推理与通信优化
  • 有大规模生产级推理系统、GPU 集群优化、开源项目或相关论文经验加分
关于我们
我们是一家全球边缘云与 CDN 基础设施平台,运营全球最大规模边缘网络之一:2,800+ 节点、覆盖 87+ 国家和地区、200+ Tbps 容量、200,000+ 台服务器,为全球 99% 可居住地区提供毫秒级分发。
母公司为亚洲上市科技集团,长期投入研发。我们聚焦分布式系统、网络协议、边缘计算、云原生、安全及 AI Infra,团队和客户遍布全球,在亚洲性能领先。
如果你的背景偏 LLM Inference / GPU Systems / CUDA / Distributed AI Infra,欢迎联系或私信交流。
WeChat: valleyview55
Email:
您好!
本帖隐藏的内容需要积分高于 10 才可浏览
您当前积分为 0。
使用VIP即刻解锁阅读权限或查看其他获取积分的方式
游客,您好!
本帖隐藏的内容需要积分高于 10 才可浏览
您当前积分为 0。
VIP即刻解锁阅读权限查看其他获取积分的方式
Unlock interview details and practice with AI
Curated Interview Questions from Top Companies


Inference Optimization Engineer
Location: Flexible — North America or major technology hubs in China. 1point 3acres
Position Overview. 1point3acres
We are looking for an Inference Optimization Engineer to build and optimize high-performance, low-latency, and high-throughput inference systems for large language models and multimodal generative models running on GPU clusters.
You will work at the intersection of LLM inference, GPU acceleration, distributed systems, and high-performance computing, helping bring large-scale AI workloads into production with industry-grade performance and efficiency.
Responsibilities

  • Design, build, and optimize inference engines for large language models and multimodal generative models, delivering low-latency, high-throughput, and highly reliable production deployment on GPU clusters.
  • Drive end-to-end inference performance optimization, including CUDA / Triton kernel development, vLLM / SGLang framework optimization, distributed inference strategies, and acceleration techniques such as quantization and sparsification.
  • Develop and optimize the GPU inference acceleration stack, improving compute and communication efficiency across multi-GPU and multi-node environments, including PCIe / GPU communication, memory management, scheduling, and high-concurrency inference architectures.
  • Explore and implement next-generation high-performance inference solutions, building scalable inference systems and core components in C++ and Python.
  • Work closely with customers and internal engineering teams to identify performance bottlenecks across training and inference workloads. Apply software-hardware co-design principles to improve efficiency, support large-model deployment, and contribute to the broader AI toolchain and technology ecosystem.
Requirements

  • 3+ years of relevant experience in LLM inference, GPU acceleration, high-performance computing, AI infrastructure, or related areas. Master's degree or above preferred in Computer Science, Artificial Intelligence, Electronics, Information Engineering, Communications, Automation, Software Engineering, or a related field.
  • Strong proficiency in Python and solid knowledge of C++ and modern C++ features, with hands-on experience developing and optimizing high-performance software.
  • Deep practical experience in at least one of the following areas:
    • GPU programming and kernel optimization using CUDA, Triton, AscendC, TileLang, or similar technologies;
    • Model compression and inference acceleration, including quantization, sparsification, or distillation;
    • Development or optimization of inference frameworks such as vLLM, SGLang, or similar systems;
    • Parallel computing, distributed inference, and compute-communication optimization across multi-GPU or multi-node environments.
  • Strong performance-analysis and debugging skills, with the ability to identify bottlenecks across GPU compute, memory, communication, scheduling, and system architecture.
Preferred Qualifications
Experience in one or more of the following areas is highly valued:

  • Building or operating large-scale production LLM or multimodal inference systems;
  • Deep development experience or open-source contributions to vLLM, SGLang, or comparable inference frameworks;
  • CUDA, Triton, TileLang, or other GPU kernel development and optimization;
  • Strong understanding of Transformer architectures, attention mechanisms, KV cache management, and inference performance characteristics;
  • Hands-on implementation of FP8, INT8, INT4, or other quantization and sparsification techniques;
  • Large-scale inference optimization across multi-GPU and multi-node GPU clusters;
  • Relevant publications, open-source projects, or meaningful technical community contributions.
About Us
We are a global edge cloud and CDN infrastructure platform operating one of the world’s largest edge networks: 2,800+ PoPs across 87+ countries, 200+ Tbps capacity, and 200,000+ servers, delivering millisecond latency to 99% of the inhabited world.
Backed by a publicly listed Asian technology group with long-term R&D investment, we tackle hard problems in distributed systems, networking, edge computing, cloud-native, and security. Our teams and customers are global, with leading performance in Asia.

Contact Info:
您好!
本帖隐藏的内容需要积分高于 10 才可浏览
您当前积分为 0。
使用VIP即刻解锁阅读权限或查看其他获取积分的方式
游客,您好!
本帖隐藏的内容需要积分高于 10 才可浏览
您当前积分为 0。
VIP即刻解锁阅读权限查看其他获取积分的方式
Unlock interview details and practice with AI
Curated Interview Questions from Top Companies

上一篇:求affirm内推-capacity modeling职位
下一篇:Fireworks AI 招data engineer/data scientist
全局:
请教LZ:

”母公司为亚洲上市科技集团“,这公司名称都不告知吗?
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
⚠️: 通过大量刷回复等方式使帖子靠前的内推帖,将会被自动关闭下沉

⚠️ 每个人只可以有一个active的内推帖子,其他内推帖会被删除

本版积分规则

>
快速回复 返回顶部 返回列表