查看: 196| 回复: 2
跳转到指定楼层
上一主题 下一主题
收起左侧

[找工就业] 求问Physical AI/Robotics公司System Design怎么准备?

全局:

2026(10-12月)-CS硕士+3个月-1年 | Other| MachineLearningEng全职@

注册一亩三分地论坛,查看更多干货!

您需要 登录 才可以下载或查看附件。没有帐号?注册账号

x
楼主硕士期间做的VLA相关的方向,现在在一家robotics初创工作。公司比较草台,想提桶跑路。

上周面了一个西雅图的具身初创,VO三场面试里面有一场System Design。完全不知道要怎么准备。我只准备了一些VLA架构设计之类的,结果面试官让我design整个从数据收集到模型训练到deployment的pipeline,场景是在50个不同的工厂里部署。我直接懵圈,只能磕磕巴巴答一点。

想问下大家这是正常的题型吗?有没有好心人分享一下自己的经历🥺该怎么准备这种面试

上一篇:LLM/ML学习笔记 适用new grad找工和低yoe想转llm/ml的 欢迎contribute!
下一篇:找人一起准备Perplexity
全局:
经典没做过就不会的题,只要有经验的人
回复

使用道具 举报

地里匿名用户
🔗
匿名用户-IGMJ2  昨天 03:15
引用:Deploytovalue.com

Deploying Embodied AI Across 50 Factories

目标不是教你训练一个单独的机器人模型,而是设计一个可以在 50 个不同工厂中持续运行、持续学习、持续部署的具身智能系统。. 1point3acres
核心问题:

* 50 个工厂的数据应该怎样统一采集?
* 不同机器人、摄像头、传感器如何接入同一个平台?
* 哪些数据应该上传,哪些数据应该留在本地?
* 怎样从真实世界数据和 simulation 中构建训练集?
* 怎样训练、评估和发布新的 embodied AI model?
* 怎样避免一次坏模型影响 50 个工厂?
* 工厂断网以后机器人还能不能继续运行?. 1point 3acres
* 怎样衡量一个新模型到底有没有创造业务价值?
.--
案例展开:一家公司需要在 50 个工厂部署一套机器人抓取、搬运和检测系统。工厂使用不同机器人、不同传感器、不同生产线环境,但希望通过一个统一的 AI 平台持续改善机器人表现。

⸻

Lesson 1 — Designing the 50-Factory System
. Waral dи,
场景

公司已经在 Factory 01 成功部署了一个机器人抓取系统。
.1point3acres
现在管理层提出新的要求:

把它推广到 50 个工厂。

第一反应往往是:

train model
↓
copy model
↓
deploy everywhere

但现实很快会变得复杂。. 1point 3acres

Factory 01 使用 ABB。

Factory 12 使用 Fanuc。

Factory 24 使用 UR。

Factory 37 光照不同。

Factory 42 的摄像头安装角度不同。 ..

Factory 49 的零件材料反光。

因此系统不能被设计成:
. 1point 3 acres
One Robot
+
One Model
+
One Factory

而应该设计为:

Central AI Platform
        ↓. From 1point 3acres bbs
Factory Edge-baidu 1point3acres
        ↓-baidu 1point3acres
Robot Abstraction
        ↓
Physical Robot

Core Architecture

                 CENTRAL PLATFORM
        Data Lake
        Dataset Registry
        Simulation
        Training
        Model Registry
        Deployment Control
        Fleet Observability
                    │
      ───────── Secure Network ─────────
                    │
     Factory 01   Factory 02   ... Factory 50
          │            │               │
       Edge Node     Edge Node       Edge Node
          │            │               │
      Robots       Robots          Robots

Key Principle

Cloud 是 control plane。

Factory 是 execution plane。

机器人必须能够在 cloud 暂时不可用的情况下继续安全运行。

Lab
. 1point3acres
设计:
. 1point3acres
Central Platform
Factory Edge
Robot Runtime

三层 architecture。
.
列出每一层负责什么。

⸻

Lesson 2 — Defining a Robot Episode
. .и
具身智能真正重要的数据单位不是 image。

也不是 video。

而是:. check 1point3acres for more.

Episode

一个 Episode 表示机器人完成一次任务的完整经历。

例如:. 1point3acres

Robot sees object
↓-baidu 1point3acres
Robot selects grasp
↓
Robot moves
↓
Robot grips
↓. 1point 3acres
Robot transports
↓
Robot releases
↓
System evaluates outcome

一个 episode 应该包含:

episode_id
factory_id
robot_id
task_id
camera frames
depth
LiDAR.
joint state
force / torque
gripper state. ----
model version
software version
calibration version
actions. 1point3acres.com
outcome
failure reason
human intervention
safety event

例如:

{
  "factory_id": "F17",
  "robot_id": "R03",
  "task": "pick_part",
  "model": "pick-v2.14",
  "success": false,
  "failure_reason": "grasp_slip",
  "human_intervention": true
}
-baidu 1point3acres
Why This Matters

如果半年以后出现事故,仅有一段录像没有多少意义。

你必须知道:

哪个工厂
哪个机器人
哪个模型
哪个 calibration
哪个 task
哪个 environment

Lab

设计自己的 Episode Schema。-baidu 1point3acres

要求必须能够回答:

为什么这个机器人在 Factory 17 失败了?

⸻

Lesson 3 — Edge Data Collection

50 个工厂最大的错误之一是:. check 1point3acres for more.
. ----
Camera Stream
↓ ..
Upload Everything
↓
Cloud. Χ

真实 production 环境里,这会迅速产生巨大的:

network cost
storage cost
processing cost
labeling cost. 1point 3acres

因此每一个 Factory Edge 都应该运行:-baidu 1point3acres

Experience Collector
. 1point3acres.com
结构:

Robot.google  и
↓
Episode Recorder
↓
Event Detector
↓
Priority Scorer. 1point3acres.com
↓
Local Buffer. 1point3acres.com
↓
Uploader

Data Selection

正常成功任务:

sample 0.1%–1%
. From 1point 3acres bbs
异常任务:
. 1point3acres
failure                100%. ----
human takeover         100%
safety event           100%
low confidence         100%
OOD                     100%
unexpected behavior     100%
. 1point3acres
真正重要的不是:

收集最多数据。
. check 1point3acres for more.
而是:

收集最有学习价值的数据。

Lab

设计一个:

episode_priority_score.1point3acres

例如:
. 1point 3 acres
failure             +50-baidu 1point3acres
human intervention  +40. ----
OOD                 +30.
low confidence      +20
random novelty      +10

决定哪些 episode 上传。

⸻. check 1point3acres for more.

Lesson 4 — Building the Dataset Pipeline

上传的数据不能直接进入训练。. From 1point 3acres bbs

. 需要经过:

Raw Episodes.
↓.1point3acres
Validation
↓
Synchronization
↓
Deduplication
↓
Quality Check
↓
Normalization
↓
Episode Store
↓
Dataset Registry

最终形成:

dataset_pick_v31

Dataset Balance

如果 Factory 01 每天生成 10 倍数据,模型可能最终变成:. From 1point 3acres bbs

Factory 01 specialist.--

因此训练数据需要平衡:

factory
robot
task
object
environment
failure type

例如:

Factory 01     10%
Factory 02      8%. ----
Factory 03      6%. 1point3acres
...
Simulation     30%
. Χ
Lab

给你一个不平衡的数据集:

Factory 01   65%
Factory 02    3%
Factory 03    2%
.... 1point 3 acres

重新设计 sampling strategy。
.
⸻

Lesson 5 — Automated Labeling

工业机器人数据完全靠人工标注通常不可持续。
. .и
很多 label 可以由机器自己产生。

例如:
. From 1point 3acres bbs
gripper closed
+
weight changed
+
source object disappeared
+
destination sensor activated. From 1point 3acres bbs

可以自动判断:

pick_success = true

整个 labeling pipeline:

Episode
↓
Automatic Rules
↓
Vision Model
↓
Confidence Score
↓.--
High Confidence → Accept
Low Confidence → Human Review

Labels

可以包括:

object
pose
grasp
task success
failure type
collision
drop
slip
human intervention

Lab

设计:
. From 1point 3acres bbs
auto-label
human-review

两级 labeling system。

⸻

Lesson 6 — Simulation and Digital Twin

很多危险情况不能在真实工厂故意制造。

例如:

human enters workspace
camera failure
sensor freeze
object falls
robot slips
network delay
PLC timeout

因此需要 Simulation Farm。
. ----
Real Factory. check 1point3acres for more.
↓
Scene Reconstruction.--
↓
Digital Twin.1point3acres
↓
Simulation
↓
Synthetic Episodes. 1point3acres.com

Domain Randomization

simulation 中随机变化:

lighting
camera angle.1point3acres
friction
object weight.--
object texture
background
sensor noise. Χ
robot speed
occlusion

目标是让模型学到:

哪些东西应该保持稳定,哪些东西只是环境变化。

Lab
.1point3acres
给 Factory 17 设计 20 个 simulation variations。

⸻

Lesson 7 — Designing the Model Stack

不要把所有问题塞给一个巨大模型。
. 1point 3acres
更合理的系统:

Vision / VLA Model
↓
Task Planner
↓
Skill Policy
↓
Motion Planning. check 1point3acres for more.
↓
Low-Level Controller

例如:

Instruction:.1point3acres
Pick the red valve.
VLA:. Waral dи,
identify valve
Planner:. 1point 3acres
locate
approach.google  и
grasp
lift
move. Χ
place
Controller:-baidu 1point3acres
joint trajectory
servo control
.google  и
大型模型负责:

semantic understanding
task planning
high-level decisions

确定性 controller 负责:

motion
velocity
joint control. 1point 3acres
safety envelope

Lab

把一个:

Pick → Inspect → Place

任务拆成不同模型和 controller 层。

⸻

Lesson 8 — Global Model + Factory Adapter

50 个工厂不应该训练 50 个完全独立模型。

一种更加 scalable 的结构:

Global Model. 1point3acres.com
+
Factory Adapter
+
Robot Adapter
+
Task Adapter.--

例如:
. check 1point3acres for more.
Global VLA v7
Factory:
Factory17 Adapter v3
Robot:
Fanuc Adapter v5
Task:
Pick Policy v12

最终:

Runtime Policy =
Global.
+.
Factory. check 1point3acres for more.
+. 1point 3acres
Robot
+
Task

Why
-baidu 1point3acres
这样 Factory 37 学到的新能力,有机会进入 global model。

其他工厂也可以获得改善。

这形成:

Fleet Learning
. ----
Lab. check 1point3acres for more.
-baidu 1point3acres
设计 Factory 01、17 和 42 的 adapter strategy。

⸻

Lesson 9 — Training Pipeline

完整训练过程:.google  и
. .и
Dataset Registry. Χ
↓
Training Job
↓
Behavior Cloning
↓
Fine-tuning
↓. .и
Simulation RL
↓
Offline Evaluation
↓
Candidate Model

可能使用:

Behavior Cloning
Imitation Learning
DAgger
Offline RL
Reinforcement Learning
VLA fine-tuning. Χ
LoRA

其中极有价值的一种数据是:
. Χ
AI action
↓
Human overrides
↓
Correct action

Human takeover 本质上就是:
. check 1point3acres for more.
免费获得的 expert demonstration。

Lab

设计一个 Human Intervention Dataset。

⸻.

Lesson 10 — Evaluation Before Deployment
. .и
模型 accuracy 不足以决定机器人能否上线。. 1point3acres.com

真正需要关注:
.--
task success
failure rate
collision
near miss.--
human intervention
cycle time
latency
OOD rate
recovery success

例如:

Task Success       99.4%
Human Intervention 0.6%
Collision          0
P99 Latency        78 ms
Recovery Success   96%

但是更重要的是:

per factory
per robot
per task. 1point 3acres

Factory fleet average 99% 没意义。

因为可能:. 1point 3acres

Factory 01  99.9%
Factory 17  91.2%

Lab

设计一个 Production Readiness Scorecard。

⸻

Lesson 11 — Shadow Deployment

新模型不应该第一次上线就控制机器人。. Χ

先运行:

Production Model
↓
controls robot
.1point3acres
同时:

Candidate Model
↓
observes
↓
makes decisions
↓
does NOT control robot

比较:
. Χ
production_action
candidate_action

如果出现:

Production:. 1point 3acres
pick object A
Candidate:
pick object B. 1point3acres.com

就记录:
-baidu 1point3acres
model disagreement
. 1point 3 acres
并进入 review queue。
.
Lab
.
定义:
. ----
shadow disagreement rate

以及什么情况下允许模型进入下一阶段。

⸻

Lesson 12 — Canary Deployment Across 50 Factories. ----
. 1point3acres
不要:

Model v10
↓
50 factories

应该:

Simulation
↓
Test Cell
↓. 1point3acres
1 Factory. 1point3acres
↓
3 Factories
↓
10 Factories. .и
↓
25 Factories
↓
50 Factories.--
. 1point3acres
这叫 deployment rings。

例如:

Ring 0 Simulation
Ring 1 Test Cell ..
Ring 2 Factory 03. Χ
Ring 3
Factory 03
Factory 11. .и
Factory 27
Ring 4
10 factories
Ring 5
50 factories. Waral dи,

任何阶段失败:

halt rollout. 1point3acres
↓
rollback

Lab

设计 Model v8.4 的 rollout plan。

. 1point3acres⸻

Lesson 13 — Safety Governor
. Waral dи,
AI Policy 不能拥有最终控制权。

Architecture:

AI Policy
↓
Safety Governor
↓
Robot Controller. ----

Safety Governor 应该是 deterministic。
-baidu 1point3acres
例如:

if human_distance < SAFE_DISTANCE:. 1point 3acres
    stop()
if target_outside_workspace:. Χ
    reject()
if velocity > MAX_SPEED:
    reject().1point3acres
if confidence < threshold:
    request_human()

模型可以说:

MOVE. 1point 3 acres

Safety Governor 可以回答:

DENIED-baidu 1point3acres
. .и
Lab

设计 10 条不能被 AI model bypass 的 safety rules。

⸻
. Waral dи,
Lesson 14 — Offline Factory Operation

工厂不能因为 cloud unavailable 就停产。

因此:

Internet Down
↓. 1point3acres
Edge Runtime Continues
.--
本地需要拥有:-baidu 1point3acres

current production model. 1point3acres
fallback model
config
calibration
episode buffer
safety system
rollback package
. From 1point 3acres bbs
网络恢复后:. 1point3acres
. 1point 3acres
sync telemetry
sync episodes
sync deployment status
sync model metadata
. .и
Lab.--

设计:.--

Factory 28 断网 12 小时。
. Waral dи,
系统应该如何运行?

⸻

Lesson 15 — Fleet Control Plane

中央平台必须知道整个 fleet 当前状态。

例如:

50 FACTORIES
Robots Online          1,842
Robots Degraded           18
Robots Offline              6
Model v7.3. .и
Production Factories       42. Χ
Model v7.4
Shadow                      5.
Model v7.2
Rollback                    3

每台机器人需要记录:

hardware. check 1point3acres for more.
firmware. check 1point3acres for more.
model
adapter
calibration
runtime ..
GPU
health
deployment.google  и

Lab

设计 Factory Fleet Dashboard。

⸻

Lesson 16 — Physical Observability

普通软件 observability:
. ----
logs
metrics
traces ..

机器人系统还需要:

physical observability

完整 observability:

Software
CPU
GPU
network
latency
Model
confidence.google  и
OOD
entropy
disagreement
Robot
joint. check 1point3acres for more.
force
temperature
gripper
Physical
collision
drop
slip
Business. Waral dи,
units/hour
defects
downtime
human interventions
.--
这里开始出现一个非常重要的问题:

模型表现变好了,业务真的变好了吗?

Lab.--
. From 1point 3acres bbs
建立:

Model Metric
→ Robot Metric.google  и
→ Business Metric

之间的 mapping。

⸻

Lesson 17 — Closing the Learning Loop

整个系统最后应该形成:
. ----
50 Factories
↓
Robot Experiences
↓
Failure Mining
↓
Novelty Mining
↓
Dataset
↓.--
Simulation
↓. ----
Training
↓
Evaluation
↓
Shadow
↓. check 1point3acres for more.
Canary
↓
Fleet Deployment. .и
↓
50 Factories

Factory 37 遇到一种新的 failure:

reflective packaging
causes bad grasp

Factory 37 上传数据。. Waral dи,

模型重新训练。
.
新版本通过测试。

逐步发布。

最终:

Factory 01–50

都获得这个经验。

这就是:

Fleet Learning

单个机器人学习。

最后变成整个 robot fleet 学习。
. ----
⸻

Final Mission
. check 1point3acres for more.
你现在负责一个:

50 Factory Embodied AI Platform
. From 1point 3acres bbs
公司要求:

99%+ task success
zero uncontrolled collision. 1point 3acres
offline factory operation
centralized fleet management
continuous model improvement. .и
safe rollback

你需要设计:

1. Robot Episode Schema.--
2. Factory Edge Architecture
3. Data Collection Strategy
4. Dataset Pipeline
.1point3acres5. Simulation Pipeline
6. Model Architecture
7. Training Pipeline.--
8. Evaluation Gates
9. Shadow Deployment.--
10. Canary Deployment-baidu 1point3acres
11. Safety Governor
12. Fleet Control Plane
13. Physical Observability
14. Continuous Learning Loop ..

最终 architecture:

                   ┌───────────────────────────┐
                   │   Embodied AI Platform    │
                   │                           │
                   │ Data                      │
                   │ Simulation                │
                   │ Training                  │
                   │ Evaluation                │
                   │ Model Registry            │
                   │ Fleet Control             │
                   └─────────────┬─────────────┘
                                 │
                        Secure Control Plane. 1point 3 acres
                                 │. Waral dи,
              ┌──────────────────┼──────────────────┐
         Factory 01         Factory 17        Factory 50
             │                  │                  │. check 1point3acres for more.
          Edge AI            Edge AI            Edge AI.1point3acres
             │                  │                  │
           Robot              Robot              Robot
              └──────────────────┬──────────────────┘
                           Physical World
                                 │
                                 ▼.google  и
                              Episodes
                                 │
                                 ▼
                         Continuous Learning

理解一个核心思想:具身智能真正困难的地方,不只是让机器人“会做”。
更困难的是让数百甚至数千台机器人在真实世界中安全地运行、持续地产生数据、持续改善,并且每一次模型变化都可以被观察、验证、控制和回滚。

这才是从 Robotics Demo 走向 Embodied AI Production System。
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 注册账号
隐私提醒:
  • ☑ 禁止发布广告,拉群,贴个人联系方式:找人请去🔗同学同事飞友,拉群请去🔗拉群结伴,广告请去🔗跳蚤市场,和 🔗租房广告|找室友
  • ☑ 论坛内容在发帖 30 分钟内可以编辑,过后则不能删帖。为防止被骚扰甚至人肉,不要公开留微信等联系方式,如有需求请以论坛私信方式发送。
  • ☑ 干货版块可免费使用 🔗超级匿名:面经(美国面经、中国面经、数科面经、PM面经),抖包袱(美国、中国)和录取汇报、定位选校版
  • ☑ 查阅全站 🔗各种匿名方法

本版积分规则

>
快速回复 返回顶部 返回列表