查看: 602| 回复: 2
跳转到指定楼层
上一主题 下一主题
收起左侧

[Hiring Manager 直招] 招聘Speech / Multimodal / Omni LLM Research Scientist

全局:

2026(7-9月)-CS博士+fresh grad 无实习或全职 | 内推|美国其他地区 码农类General全职@

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x
大家好!我组 目前正在招聘一名 Junior 到 Senior/Founding 级别的 Research Scientist。

关于公司:
  • Healthcare/ Medical Tech 头部Top5 公司深耕行业30年,持续盈利且增长迅速,业务稳定,巨量数据,场景非常扎实,目前和多个顶尖Speech AI公司积极合作训练模型。
  • 当前AI团队迅速扩张,积极招聘 Senior Engineer+ Researcher (Headcount 10+),整体package 和成长空间都很有竞争力。
  • 亿级收入企业,白人占80%的公司,非中国小公司。
关于岗位:
  • PhD/Master 学历, 方案 speech/ Omni LLM 理解和生成,具体如下
  • 至少一篇相关方案的 顶会 1篇 First Author paper 和 publication
  • 至少3+Year 相关方案的 Industry/Academic Research经验,不需要healthcare/medical相关经验
  • 熟悉Pytorch 整个大模型training的流程
  • 需要Relocate 到美国中部
  • 支持Sponsor ( H1b/transfer )
Hiring Manager直招,如果您有兴趣且符合条件,欢迎直接发Resume发到 yukileong001@gmail.com,并在帖子下方回复告知。期待与您联系!(后面岗位official link开了,会补充到这里。)

----. 1point 3 acres
Speech / Multimodal / Omni LLM Research Scientist
Title: Speech / Multimodal / Omni Research Scientist (Junior to Founding/Principle level available)
Team: Core Model Research, reporting to AI Director / Head of Research.google  и
Focus: Speech & Multimodal AI / Multimodal ML/Speech-to-Speech Translation / Omni / Any-to-Any Foundation Models

About the Role
We build real-time speech translation and interpretation technology, and we train our own models to power it. We're looking for an exceptional research scientist to push the frontier of our core models — and we're open on where your depth lies. Whether you go deep on speech translation (ST / ASR / TTS, simultaneous and low-latency systems) or on any-to-any (omni) multimodal models (unified understanding and generation across speech, text, and vision), we want to talk. These two paths converge on the same prize, and you'll help us get there.

End-to-end, voice- and prosody-preserving, low-latency speech-to-speech translation is fundamentally a unified speech/multimodal modeling problem. The speech-translation track reaches it by deepening the speech line; the omni track reaches it by unifying modalities under one model. They meet at speech tokenization and streaming low-latency generation. We hire the best researcher from either side, and help them grow into the other.-baidu 1point3acres

Key Responsibility - What you'll do
Shared across both tracks
  • Set the modeling and architecture strategy for our core models, and make pragmatic, evidence-driven choices.
  • Lead training and post-training of our models (pretraining and/or SFT, RLHF/RLVR).
  • Own evaluation across quality and latency, partnering with professional interpreters and linguists.
  • Set technical direction and collaborate with the data, MT, and real-time-systems teams.
Track A — Speech translation depth
  • Own the speech-translation architecture (cascaded ASR→MT→TTS vs. end-to-end Speech/Audio LLM vs. hybrid).
  • Crack simultaneous / streaming translation: "when to emit" policies (wait-k, AlignAtt, Local Agreement) that optimize first-chunk latency against quality from incomplete input.
  • Advance ASR robustness, expressive / voice-preserving TTS, multilingual and low-resource quality, and terminology consistency.
  • Beat speech-translation data scarcity (pretraining, multi-task, distillation, synthesis, self-supervision).
Track B — Omni multimodal depth
  • Own the unified architecture (autoregressive over a shared discrete vocabulary, masked discrete diffusion, or hybrid), unifying understanding and generation in one model.
  • Design cross-modal tokenization — for speech, choosing semantic vs. acoustic tokens with streaming low-latency generation in mind (avoiding the "complete-then-decode" bottleneck).
  • Solve modality interference via modality-specific routing / MoE, and advance cross-modal alignment and interleaved multimodal modeling.
  • Drive high-quality any-to-any generation (speech and image), not just any-to-text understanding.
 
Qualifications
  • PhD in speech/CS/NLP/ML, or equivalent research track record.
  • 3+ years of research experience with deep, demonstrated expertise in at least ONE track above — and genuine excitement about where the two meet.
  • Hands-on experience training modern LLMs / foundation models (pretraining and/or post-training, RLHF).
  • Strong engineering: Python + PyTorch; able to design and run experiments at scale.
  • Track record of first-author publications at top venues ((NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, Interspeech, ICASSP, IWSLT).
 
Preferred Qualifications
  • Experience shipping production real-time speech systems.
  • Familiarity with simultaneous policies (wait-k, AlignAtt, MMA, Local Agreement) and tools like Whisper, wav2vec/HuBERT, SeamlessM4T.
  • Experience building omni/Any-to-Any models (GPT-4o-style) or work in the lineage of Chameleon / AnyGPT / NExT-GPT / Qwen-Omni.
  • Speech tokenization and streaming speech generation (the bridge between both tracks).
  • MoE architectures / modality routers; discrete-diffusion multimodal generation.
  • Low-resource / massively multilingual MT; large-scale multimodal data construction.
  • IWSLT participation, open-source contributions, or agentic / tool-use modeling.

评分

参与人数 1大米 +1 收起 理由
127849172401 + 1 给你点个赞!

查看全部评分


上一篇:Uber 内推,长期有效,strong candidate直推HM
下一篇:Geotab AI Platform 部门招聘 - Toronto / Atlanta
全局:
感谢楼主!已发邮件,已加米,我的email是cl38**@nyu.edu
回复

使用道具 举报

🔗
mainghay 2026-7-22 05:59:29 | 只看该作者
全局:
已发 eeming*****@gmail.com 感谢!
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
⚠️: 通过大量刷回复等方式使帖子靠前的内推帖,将会被自动关闭下沉

⚠️ 每个人只可以有一个active的内推帖子,其他内推帖会被删除

本版积分规则

>
快速回复 返回顶部 返回列表