引用:Deploytovalue.com
Deploying Embodied AI Across 50 Factories
目标不是教你训练一个单独的机器人模型,而是设计一个可以在 50 个不同工厂中持续运行、持续学习、持续部署的具身智能系统。. 1point3acres
核心问题:
* 50 个工厂的数据应该怎样统一采集?
* 不同机器人、摄像头、传感器如何接入同一个平台?
* 哪些数据应该上传,哪些数据应该留在本地?
* 怎样从真实世界数据和 simulation 中构建训练集?
* 怎样训练、评估和发布新的 embodied AI model?
* 怎样避免一次坏模型影响 50 个工厂?
* 工厂断网以后机器人还能不能继续运行?. 1point 3acres
* 怎样衡量一个新模型到底有没有创造业务价值?
.--
案例展开:一家公司需要在 50 个工厂部署一套机器人抓取、搬运和检测系统。工厂使用不同机器人、不同传感器、不同生产线环境,但希望通过一个统一的 AI 平台持续改善机器人表现。
⸻
Lesson 1 — Designing the 50-Factory System
. Waral dи,
场景
公司已经在 Factory 01 成功部署了一个机器人抓取系统。
.1point3acres
现在管理层提出新的要求:
把它推广到 50 个工厂。
第一反应往往是:
train model
↓
copy model
↓
deploy everywhere
但现实很快会变得复杂。. 1point 3acres
Factory 01 使用 ABB。
Factory 12 使用 Fanuc。
Factory 24 使用 UR。
Factory 37 光照不同。
Factory 42 的摄像头安装角度不同。 ..
Factory 49 的零件材料反光。
因此系统不能被设计成:
. 1point 3 acres
One Robot
+
One Model
+
One Factory
而应该设计为:
Central AI Platform
↓. From 1point 3acres bbs
Factory Edge-baidu 1point3acres
↓-baidu 1point3acres
Robot Abstraction
↓
Physical Robot
Core Architecture
CENTRAL PLATFORM
Data Lake
Dataset Registry
Simulation
Training
Model Registry
Deployment Control
Fleet Observability
│
───────── Secure Network ─────────
│
Factory 01 Factory 02 ... Factory 50
│ │ │
Edge Node Edge Node Edge Node
│ │ │
Robots Robots Robots
Key Principle
Cloud 是 control plane。
Factory 是 execution plane。
机器人必须能够在 cloud 暂时不可用的情况下继续安全运行。
Lab
. 1point3acres
设计:
. 1point3acres
Central Platform
Factory Edge
Robot Runtime
三层 architecture。
.
列出每一层负责什么。
⸻
Lesson 2 — Defining a Robot Episode
. .и
具身智能真正重要的数据单位不是 image。
也不是 video。
而是:. check 1point3acres for more.
Episode
一个 Episode 表示机器人完成一次任务的完整经历。
例如:. 1point3acres
Robot sees object
↓-baidu 1point3acres
Robot selects grasp
↓
Robot moves
↓
Robot grips
↓. 1point 3acres
Robot transports
↓
Robot releases
↓
System evaluates outcome
一个 episode 应该包含:
episode_id
factory_id
robot_id
task_id
camera frames
depth
LiDAR.
joint state
force / torque
gripper state. ----
model version
software version
calibration version
actions. 1point3acres.com
outcome
failure reason
human intervention
safety event
例如:
{
"factory_id": "F17",
"robot_id": "R03",
"task": "pick_part",
"model": "pick-v2.14",
"success": false,
"failure_reason": "grasp_slip",
"human_intervention": true
}
-baidu 1point3acres
Why This Matters
如果半年以后出现事故,仅有一段录像没有多少意义。
你必须知道:
哪个工厂
哪个机器人
哪个模型
哪个 calibration
哪个 task
哪个 environment
Lab
设计自己的 Episode Schema。-baidu 1point3acres
要求必须能够回答:
为什么这个机器人在 Factory 17 失败了?
⸻
Lesson 3 — Edge Data Collection
50 个工厂最大的错误之一是:. check 1point3acres for more.
. ----
Camera Stream
↓ ..
Upload Everything
↓
Cloud. Χ
真实 production 环境里,这会迅速产生巨大的:
network cost
storage cost
processing cost
labeling cost. 1point 3acres
因此每一个 Factory Edge 都应该运行:-baidu 1point3acres
Experience Collector
. 1point3acres.com
结构:
Robot.google и
↓
Episode Recorder
↓
Event Detector
↓
Priority Scorer. 1point3acres.com
↓
Local Buffer. 1point3acres.com
↓
Uploader
Data Selection
正常成功任务:
sample 0.1%–1%
. From 1point 3acres bbs
异常任务:
. 1point3acres
failure 100%. ----
human takeover 100%
safety event 100%
low confidence 100%
OOD 100%
unexpected behavior 100%
. 1point3acres
真正重要的不是:
收集最多数据。
. check 1point3acres for more.
而是:
收集最有学习价值的数据。
Lab
设计一个:
episode_priority_score.1point3acres
例如:
. 1point 3 acres
failure +50-baidu 1point3acres
human intervention +40. ----
OOD +30.
low confidence +20
random novelty +10
决定哪些 episode 上传。
⸻. check 1point3acres for more.
Lesson 4 — Building the Dataset Pipeline
上传的数据不能直接进入训练。. From 1point 3acres bbs
. 需要经过:
Raw Episodes.
↓.1point3acres
Validation
↓
Synchronization
↓
Deduplication
↓
Quality Check
↓
Normalization
↓
Episode Store
↓
Dataset Registry
最终形成:
dataset_pick_v31
Dataset Balance
如果 Factory 01 每天生成 10 倍数据,模型可能最终变成:. From 1point 3acres bbs
Factory 01 specialist.--
因此训练数据需要平衡:
factory
robot
task
object
environment
failure type
例如:
Factory 01 10%
Factory 02 8%. ----
Factory 03 6%. 1point3acres
...
Simulation 30%
. Χ
Lab
给你一个不平衡的数据集:
Factory 01 65%
Factory 02 3%
Factory 03 2%
.... 1point 3 acres
重新设计 sampling strategy。
.
⸻
Lesson 5 — Automated Labeling
工业机器人数据完全靠人工标注通常不可持续。
. .и
很多 label 可以由机器自己产生。
例如:
. From 1point 3acres bbs
gripper closed
+
weight changed
+
source object disappeared
+
destination sensor activated. From 1point 3acres bbs
可以自动判断:
pick_success = true
整个 labeling pipeline:
Episode
↓
Automatic Rules
↓
Vision Model
↓
Confidence Score
↓.--
High Confidence → Accept
Low Confidence → Human Review
Labels
可以包括:
object
pose
grasp
task success
failure type
collision
drop
slip
human intervention
Lab
设计:
. From 1point 3acres bbs
auto-label
human-review
两级 labeling system。
⸻
Lesson 6 — Simulation and Digital Twin
很多危险情况不能在真实工厂故意制造。
例如:
human enters workspace
camera failure
sensor freeze
object falls
robot slips
network delay
PLC timeout
因此需要 Simulation Farm。
. ----
Real Factory. check 1point3acres for more.
↓
Scene Reconstruction.--
↓
Digital Twin.1point3acres
↓
Simulation
↓
Synthetic Episodes. 1point3acres.com
Domain Randomization
simulation 中随机变化:
lighting
camera angle.1point3acres
friction
object weight.--
object texture
background
sensor noise. Χ
robot speed
occlusion
目标是让模型学到:
哪些东西应该保持稳定,哪些东西只是环境变化。
Lab
.1point3acres
给 Factory 17 设计 20 个 simulation variations。
⸻
Lesson 7 — Designing the Model Stack
不要把所有问题塞给一个巨大模型。
. 1point 3acres
更合理的系统:
Vision / VLA Model
↓
Task Planner
↓
Skill Policy
↓
Motion Planning. check 1point3acres for more.
↓
Low-Level Controller
例如:
Instruction:.1point3acres
Pick the red valve.
VLA:. Waral dи,
identify valve
Planner:. 1point 3acres
locate
approach.google и
grasp
lift
move. Χ
place
Controller:-baidu 1point3acres
joint trajectory
servo control
.google и
大型模型负责:
semantic understanding
task planning
high-level decisions
确定性 controller 负责:
motion
velocity
joint control. 1point 3acres
safety envelope
Lab
把一个:
Pick → Inspect → Place
任务拆成不同模型和 controller 层。
⸻
Lesson 8 — Global Model + Factory Adapter
50 个工厂不应该训练 50 个完全独立模型。
一种更加 scalable 的结构:
Global Model. 1point3acres.com
+
Factory Adapter
+
Robot Adapter
+
Task Adapter.--
例如:
. check 1point3acres for more.
Global VLA v7
Factory:
Factory17 Adapter v3
Robot:
Fanuc Adapter v5
Task:
Pick Policy v12
最终:
Runtime Policy =
Global.
+.
Factory. check 1point3acres for more.
+. 1point 3acres
Robot
+
Task
Why
-baidu 1point3acres
这样 Factory 37 学到的新能力,有机会进入 global model。
其他工厂也可以获得改善。
这形成:
Fleet Learning
. ----
Lab. check 1point3acres for more.
-baidu 1point3acres
设计 Factory 01、17 和 42 的 adapter strategy。
⸻
Lesson 9 — Training Pipeline
完整训练过程:.google и
. .и
Dataset Registry. Χ
↓
Training Job
↓
Behavior Cloning
↓
Fine-tuning
↓. .и
Simulation RL
↓
Offline Evaluation
↓
Candidate Model
可能使用:
Behavior Cloning
Imitation Learning
DAgger
Offline RL
Reinforcement Learning
VLA fine-tuning. Χ
LoRA
其中极有价值的一种数据是:
. Χ
AI action
↓
Human overrides
↓
Correct action
Human takeover 本质上就是:
. check 1point3acres for more.
免费获得的 expert demonstration。
Lab
设计一个 Human Intervention Dataset。
⸻.
Lesson 10 — Evaluation Before Deployment
. .и
模型 accuracy 不足以决定机器人能否上线。. 1point3acres.com
真正需要关注:
.--
task success
failure rate
collision
near miss.--
human intervention
cycle time
latency
OOD rate
recovery success
例如:
Task Success 99.4%
Human Intervention 0.6%
Collision 0
P99 Latency 78 ms
Recovery Success 96%
但是更重要的是:
per factory
per robot
per task. 1point 3acres
Factory fleet average 99% 没意义。
因为可能:. 1point 3acres
Factory 01 99.9%
Factory 17 91.2%
Lab
设计一个 Production Readiness Scorecard。
⸻
Lesson 11 — Shadow Deployment
新模型不应该第一次上线就控制机器人。. Χ
先运行:
Production Model
↓
controls robot
.1point3acres
同时:
Candidate Model
↓
observes
↓
makes decisions
↓
does NOT control robot
比较:
. Χ
production_action
candidate_action
如果出现:
Production:. 1point 3acres
pick object A
Candidate:
pick object B. 1point3acres.com
就记录:
-baidu 1point3acres
model disagreement
. 1point 3 acres
并进入 review queue。
.
Lab
.
定义:
. ----
shadow disagreement rate
以及什么情况下允许模型进入下一阶段。
⸻
Lesson 12 — Canary Deployment Across 50 Factories. ----
. 1point3acres
不要:
Model v10
↓
50 factories
应该:
Simulation
↓
Test Cell
↓. 1point3acres
1 Factory. 1point3acres
↓
3 Factories
↓
10 Factories. .и
↓
25 Factories
↓
50 Factories.--
. 1point3acres
这叫 deployment rings。
例如:
Ring 0 Simulation
Ring 1 Test Cell ..
Ring 2 Factory 03. Χ
Ring 3
Factory 03
Factory 11. .и
Factory 27
Ring 4
10 factories
Ring 5
50 factories. Waral dи,
任何阶段失败:
halt rollout. 1point3acres
↓
rollback
Lab
设计 Model v8.4 的 rollout plan。
. 1point3acres⸻
Lesson 13 — Safety Governor
. Waral dи,
AI Policy 不能拥有最终控制权。
Architecture:
AI Policy
↓
Safety Governor
↓
Robot Controller. ----
Safety Governor 应该是 deterministic。
-baidu 1point3acres
例如:
if human_distance < SAFE_DISTANCE:. 1point 3acres
stop()
if target_outside_workspace:. Χ
reject()
if velocity > MAX_SPEED:
reject().1point3acres
if confidence < threshold:
request_human()
模型可以说:
MOVE. 1point 3 acres
Safety Governor 可以回答:
DENIED-baidu 1point3acres
. .и
Lab
设计 10 条不能被 AI model bypass 的 safety rules。
⸻
. Waral dи,
Lesson 14 — Offline Factory Operation
工厂不能因为 cloud unavailable 就停产。
因此:
Internet Down
↓. 1point3acres
Edge Runtime Continues
.--
本地需要拥有:-baidu 1point3acres
current production model. 1point3acres
fallback model
config
calibration
episode buffer
safety system
rollback package
. From 1point 3acres bbs
网络恢复后:. 1point3acres
. 1point 3acres
sync telemetry
sync episodes
sync deployment status
sync model metadata
. .и
Lab.--
设计:.--
Factory 28 断网 12 小时。
. Waral dи,
系统应该如何运行?
⸻
Lesson 15 — Fleet Control Plane
中央平台必须知道整个 fleet 当前状态。
例如:
50 FACTORIES
Robots Online 1,842
Robots Degraded 18
Robots Offline 6
Model v7.3. .и
Production Factories 42. Χ
Model v7.4
Shadow 5.
Model v7.2
Rollback 3
每台机器人需要记录:
hardware. check 1point3acres for more.
firmware. check 1point3acres for more.
model
adapter
calibration
runtime ..
GPU
health
deployment.google и
Lab
设计 Factory Fleet Dashboard。
⸻
Lesson 16 — Physical Observability
普通软件 observability:
. ----
logs
metrics
traces ..
机器人系统还需要:
physical observability
完整 observability:
Software
CPU
GPU
network
latency
Model
confidence.google и
OOD
entropy
disagreement
Robot
joint. check 1point3acres for more.
force
temperature
gripper
Physical
collision
drop
slip
Business. Waral dи,
units/hour
defects
downtime
human interventions
.--
这里开始出现一个非常重要的问题:
模型表现变好了,业务真的变好了吗?
Lab.--
. From 1point 3acres bbs
建立:
Model Metric
→ Robot Metric.google и
→ Business Metric
之间的 mapping。
⸻
Lesson 17 — Closing the Learning Loop
整个系统最后应该形成:
. ----
50 Factories
↓
Robot Experiences
↓
Failure Mining
↓
Novelty Mining
↓
Dataset
↓.--
Simulation
↓. ----
Training
↓
Evaluation
↓
Shadow
↓. check 1point3acres for more.
Canary
↓
Fleet Deployment. .и
↓
50 Factories
Factory 37 遇到一种新的 failure:
reflective packaging
causes bad grasp
Factory 37 上传数据。. Waral dи,
模型重新训练。
.
新版本通过测试。
逐步发布。
最终:
Factory 01–50
都获得这个经验。
这就是:
Fleet Learning
单个机器人学习。
最后变成整个 robot fleet 学习。
. ----
⸻
Final Mission
. check 1point3acres for more.
你现在负责一个:
50 Factory Embodied AI Platform
. From 1point 3acres bbs
公司要求:
99%+ task success
zero uncontrolled collision. 1point 3acres
offline factory operation
centralized fleet management
continuous model improvement. .и
safe rollback
你需要设计:
1. Robot Episode Schema.--
2. Factory Edge Architecture
3. Data Collection Strategy
4. Dataset Pipeline
.1point3acres5. Simulation Pipeline
6. Model Architecture
7. Training Pipeline.--
8. Evaluation Gates
9. Shadow Deployment.--
10. Canary Deployment-baidu 1point3acres
11. Safety Governor
12. Fleet Control Plane
13. Physical Observability
14. Continuous Learning Loop ..
最终 architecture:
┌───────────────────────────┐
│ Embodied AI Platform │
│ │
│ Data │
│ Simulation │
│ Training │
│ Evaluation │
│ Model Registry │
│ Fleet Control │
└─────────────┬─────────────┘
│
Secure Control Plane. 1point 3 acres
│. Waral dи,
┌──────────────────┼──────────────────┐
Factory 01 Factory 17 Factory 50
│ │ │. check 1point3acres for more.
Edge AI Edge AI Edge AI.1point3acres
│ │ │
Robot Robot Robot
└──────────────────┬──────────────────┘
Physical World
│
▼.google и
Episodes
│
▼
Continuous Learning
理解一个核心思想:具身智能真正困难的地方,不只是让机器人“会做”。
更困难的是让数百甚至数千台机器人在真实世界中安全地运行、持续地产生数据、持续改善,并且每一次模型变化都可以被观察、验证、控制和回滚。
这才是从 Robotics Demo 走向 Embodied AI Production System。 |