11 분 소요

0. Introduction

Paper link

Project page

Code and report

Model collection

Vision-language model이 image를 잘 설명하고 spatial question을 잘 푼다고 해서 robot action까지 잘 생성하는 것은 아니다. Embodied agent에는 scene understanding, affordance reasoning, action planning, end-effector motion, state transition prediction이 함께 필요하다.

PhysBrain 1.5는 이 capability를 separate head or separate model로 두지 않는다. Language answer, end-effector trajectory, future RGB-depth-mask를 모두 discrete token sequence로 바꾸고, Qwen3-VL 기반 autoregressive backbone 하나에서 next-token prediction으로 학습한다.

Paper가 말하는 physical loop는 단순하다.

  1. Observation을 본다.
  2. Scene and task를 이해한다.
  3. Action을 생성한다.
  4. Environment가 어떻게 바뀔지 예측한다.
  5. Updated observation을 다음 interaction의 context로 사용한다.

중요한 점은 세 capability가 같은 interface를 쓴다는 것이다. Understanding task는 language and spatial token을 내고, action task는 ActionPiece token을 내며, future-state task는 visual code를 낸다. Loss family는 같고, target token type과 loss mask가 달라진다.

한 줄 요약: PhysBrain 1.5는 language, action, future visual state를 하나의 expanded vocabulary와 masked next-token objective로 통합하고, large-scale human interaction video에서 physical prior를 학습한 뒤 robot and simulation data로 adaptation한 2B and 8B physical foundation model family다.

이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.

  • General VLM을 embodied understanding, action generation, future-state prediction의 common backbone으로 확장한다.
  • Human interaction video를 perception, wrist motion, state transition supervision으로 재구성한다.
  • Continuous end-effector trajectory를 ActionPiece discrete vocabulary로 바꿔 language model interface에 넣는다.
  • RGB, depth, human or robot mask를 spatially aligned visual token sequence로 함께 생성한다.
  • Understanding benchmark의 강한 결과와 control evidence의 현재 한계를 한 논문 안에서 같이 볼 수 있다.

1. Problem Setting

1-1. VLM의 spatial reasoning과 physical control은 다른 문제다

General VLM은 object, relation, text, scene를 설명할 수 있다. 하지만 embodied system에는 추가 조건이 있다.

  • Camera viewpoint and robot embodiment가 바뀐다.
  • Action은 continuous trajectory다.
  • Coordinate convention and scale이 source마다 다르다.
  • 같은 instruction도 current pose에 따라 다른 motion이 필요하다.
  • Action 이후 state transition까지 coherent해야 한다.
  • Wrong prediction은 image caption error가 아니라 physical failure로 이어질 수 있다.

따라서 physical foundation model은 visual question answering score만 높아서는 부족하다. Observation에서 action and future state로 이어지는 interface가 필요하다.

1-2. Capability를 separate model로 나눌 때 생기는 문제

기존 pipeline은 대개 아래처럼 분리된다.

  • VLM: scene understanding과 planning
  • Policy model: continuous action prediction 담당
  • World model: future observation prediction 담당
  • Controller: robot-specific execution 담당

이 modularity는 각 component를 독립적으로 최적화하기 쉽다는 장점이 있다. 반면 representation and data interface가 달라 error가 누적될 수 있다.

예를 들어 VLM이 language plan을 만들고 policy가 이를 다시 parse하면 spatial precision이 손실될 수 있다. World model이 policy가 보는 action representation과 다른 coordinate를 쓰면 predicted future와 execution이 맞지 않는다.

PhysBrain 1.5는 capability를 하나의 token generation problem으로 통일한다. 이는 모든 deployment problem을 해결한다는 뜻이 아니라, pre-training stage에서 shared physical representation을 만들겠다는 선택이다.

1-3. Heterogeneous action data의 문제

Human motion, bimanual robot, single-arm robot, simulation은 action space가 다르다.

  • Coordinate origin and axis convention이 다르다.
  • Unit and scale이 다를 수 있다.
  • Rotation representation이 다르다.
  • Control frequency 정보가 다르다.
  • Gripper semantics가 다르다.
  • 일부 dataset에는 calibration metadata가 부족하다.

모든 source를 canonical global coordinate로 강제 변환하려면 불확실한 assumption이 들어간다. PhysBrain은 source-native local motion을 유지하고, previous action chunk를 context로 넣어 current convention을 추론하게 한다.

이 선택은 data integration을 쉽게 하지만, deployment에서 action history and calibration에 대한 의존성을 만든다.

2. Core Idea

2-1. One vocabulary, one backbone, masked supervision

PhysBrain 1.5는 base VLM vocabulary를 세 영역으로 확장한다.

\[\mathcal{V} = \mathcal{V}_{lang} \cup \mathcal{V}_{act} \cup \mathcal{V}_{vis}\]
  • $\mathcal{V}_{lang}$: language와 spatial response token
  • $\mathcal{V}_{act}$: ActionPiece trajectory token 집합
  • $\mathcal{V}_{vis}$: quantized future-state visual token 집합

같은 autoregressive model이 task format에 따라 다른 token family를 생성한다. Loss는 task-specific position만 supervise하는 masked NTP로 볼 수 있다.

\[\mathcal{L} = - \sum_{t=1}^{T} m_t \log p_{\theta}(x_t \mid x_{<t})\]

$m_t$는 answer, action, or future-state target position에서만 1이 된다.

별도 action regression head, pixel reconstruction head, world-model head를 붙이는 대신 vocabulary and serialization contract로 modality를 구분한다.

2-2. ActionPiece가 continuous trajectory를 token으로 바꾼다

PhysBrain은 left and right wrist end-effector의 short motion chunk를 예측한다. 한 time step은 다음 10개 value로 구성된다.

  • Relative translation 3
  • Relative rotation 6
  • Gripper scalar 1

Chunk horizon은 16 steps이므로 continuous representation은 wrist당 160 values다.

\[a_{t,k} = [ \Delta p_{t,k}^{x}, \Delta p_{t,k}^{y}, \Delta p_{t,k}^{z}, r_{t,k}^{1:6}, g_{t,k} ], \quad k=1,\ldots,16\]

ActionPiece tokenizer는 28.7M trajectory segments에서 discrete vocabulary를 학습한다.

  • Action vocabulary 크기: 512
  • 한 wrist chunk: 32 action tokens
  • Left and right wrist는 shared codebook 사용

이 구조는 continuous control을 classification-like token prediction으로 바꾼다. Language model training infrastructure를 그대로 재사용할 수 있고, action and language token을 same sequence에 넣을 수 있다.

2-3. Previous action chunk를 local convention으로 사용한다

Action generation input에는 current observation, instruction과 함께 immediately preceding action chunk가 들어갈 수 있다.

이 history는 다음 정보를 암묵적으로 전달한다.

  • Motion scale 정보
  • Direction convention 정보
  • Control frequency 정보
  • Local coordinate orientation 정보
  • Current movement trend 정보
  • 특정 robot or human motion pattern

즉 action generation은 isolated command prediction보다 trajectory continuation에 가깝다.

2-4. Future state를 RGB, depth, mask로 함께 예측한다

PhysBrain의 future-state target은 세 modality가 spatially aligned되어 있다.

  • RGB
  • Depth
  • Human or robot mask

각 modality는 shared VQ-VAE로 $128 x 128$ target에서 $16 x 16$ code grid로 바뀐다. Modality당 256 codes이며, RGB-depth-mask token을 같은 spatial location 기준으로 interleave한다.

이 serialization은 같은 pixel location의 appearance, geometry, actor identity를 가까운 autoregressive context에 둔다. Future prediction을 photorealistic RGB 하나로 끝내지 않고, depth and mask consistency를 함께 요구한다.

3. Architecture / Method

3-1. Overview

Item Description
Base model Qwen3-VL-Instruct family
Released sizes 2B and 8B
Unified outputs Language, spatial response, action, future visual state
Action tokenizer ActionPiece, 512-token vocabulary
Action horizon 16 future steps per chunk
Future modalities RGB, depth, human or robot mask
Training objective Task-masked autoregressive next-token prediction
Pre-training source Human interaction video plus general multimodal instruction data
SFT source Human demonstrations, real robot trajectories, simulation

3-2. Understanding output

Understanding task는 ordinary language answer뿐 아니라 pointing, trajectory trace, spatial localization, affordance, multi-view relation을 포함한다.

Paper의 28 benchmarks는 대략 다섯 capability group으로 구성된다.

  1. 기초 visual-spatial perception
  2. Spatial and multi-view understanding 평가
  3. Embodied cognition, reasoning, and planning 평가
  4. Spatial grounding, pointing, and affordance 평가
  5. Visual-trace and trajectory reasoning 평가

이 evaluation이 important한 이유는 control benchmark 전에 observation and reasoning prior가 얼마나 강한지 보여주기 때문이다.

3-3. Action output

Human video에서 wrist pose를 회복하고, current pose에 relative한 short action chunk를 만든다. Robot and simulation data도 same ActionPiece interface로 변환한다.

Action tokenizer가 모든 embodiment를 완전히 canonicalize하는 것은 아니다. Source-native axes and units를 유지하는 경우가 있으며, previous action context가 이를 보완한다.

따라서 model output을 실제 robot command로 쓰려면 아래 layer가 추가로 필요하다.

  • ActionPiece decoder 구성
  • Robot-specific coordinate mapping 구성
  • Safety and collision check 구성
  • Low-level controller 구성
  • Closed-loop observation update 구성

Paper의 token prediction은 이 full stack 중 policy representation layer에 해당한다.

3-4. Future-state output

Future-state token은 language answer와 같은 output head에서 생성된다. Inference 후에는 modality별 code grid로 de-interleave하고 frozen VQ decoder로 복원한다.

이 방식의 장점은 discrete sequence interface다. 단점은 resolution and codebook bottleneck이다. $128 x 128$ future state는 broad layout and motion을 보이기에는 충분하지만, contact detail and fine geometry를 직접 검증하기에는 제한적이다.

3-5. Physical loop as a training interface

PhysBrain의 중요한 contribution은 architecture module보다 training interface다.

  • Perception supervision은 object, geometry, affordance를 가르친다.
  • Action supervision은 state-conditioned motion을 가르친다.
  • Future-state supervision은 action 이후 scene transition을 가르친다.
  • General multimodal data는 original VLM capability를 보존한다.

세 supervision family가 shared backbone에서 만난다.

4. Training / Data / Recipe

4-1. Physical-aware pre-training

Pre-training은 약 30,000 hours의 human interaction video에서 embodied supervision을 만든다.

Supervision family Samples
Physical perception 24.3M
Human action 31.2M
Future state 26.8M
General language and vision-language instruction 14.9M

중요한 점은 robot data가 아닌 human interaction video가 initial physical prior의 중심이라는 것이다.

Physical perception

Caption, VQA, detection, segmentation, depth, pointing, counting, temporal grounding, tracking, multi-view correspondence를 포함한다.

Human action

Human pose and hand motion에서 wrist end-effector trajectory를 회복한다. Pose recovery가 불안정하거나 temporal alignment가 깨진 sample은 제거한다.

Future state

Interaction 전 context와 later observation을 pair한다. RGB frame, estimated depth, human segmentation mask를 spatially align한다.

4-2. Embodied supervised fine-tuning

SFT는 human, robot, simulation domain을 섞는다.

Supervision family Samples
Embodied understanding 6.61M
Action 5.4M
Future state 1.2M
General instruction 1M

Action corpus에는 약 2,620 hours의 real-robot and simulation trajectory와 high-quality human motion이 포함된다. 17 sources across bimanual, single-arm, simulation setting을 사용한다.

이 stage는 physical prior를 executable action distribution에 맞추는 역할을 한다.

4-3. Optimization

Reported setup의 핵심은 다음과 같다.

  • Two stages, 각 stage는 one epoch
  • Precision은 bf16
  • Optimizer는 Distributed Adam
  • Cosine learning-rate decay 사용
  • Sequence packing 사용
  • Maximum sequence length는 32,768
  • Technical summary가 보고한 peak learning rate는 $2 \times 10^{-5}$
  • Gradient clipping은 1.0

Training scale 자체는 개인 연구자가 scratch에서 재현하기 어렵다. Practical reuse는 released checkpoint를 fine-tune하거나, ActionPiece and serialization idea를 smaller model에 적용하는 방향이 더 현실적이다.

4-4. Engineering notes

1) Action coordinates must be logged explicitly

Source-native coordinate를 유지한다면 dataset row마다 unit, axis, origin, handedness, control frequency를 versioning해야 한다. Previous action context가 모든 ambiguity를 자동으로 해결한다고 가정하면 위험하다.

2) Token perplexity와 control success를 분리해야 한다

Action token perplexity가 낮아져도 physical task success, collision, contact stability가 자동으로 좋아지는 것은 아니다. Offline sequence metric and closed-loop control metric을 함께 봐야 한다.

3) Future-state target provenance가 중요하다

Depth and mask가 sensor ground truth인지, perception model pseudo-label인지에 따라 error interpretation이 달라진다. Pseudo-label artifact가 world dynamics로 학습될 수 있다.

4) Action history bootstrap을 설계해야 한다

Inference 첫 step에는 preceding ground-truth chunk가 없다. Zero history, demonstrator history, policy-generated history 중 어떤 initialization을 쓰는지 deployment protocol이 필요하다.

5) Safety layer는 model 밖에 남는다

Unified token generation은 planning interface를 단순화하지만, torque limit, collision avoidance, workspace constraint, emergency stop을 대체하지 않는다.

5. Evaluation

5-1. Embodied understanding

PhysBrain 1.5-8B는 28 embodied benchmarks에서 overall average 72.5를 보고한다.

  • Evaluated open-source models 중 overall first
  • 14 benchmarks에서 open-source first
  • 24 benchmarks에서 open-source top-2
  • PhysBrain 1.5-2B overall 66.6

Overall score는 28 benchmark의 unweighted mean이다. Closed model은 reference로 표시되고 open-source ranking에서 제외된다.

이 result는 understanding capability의 strong evidence다. 다만 heterogeneous benchmark average는 protocol, scale, task count에 영향을 받는다. Category average and per-task failure를 같이 봐야 한다.

5-2. General multimodal retention

Model은 base Qwen3-VL-Instruct 대비 일부 benchmark에서 개선하고, 일부 document, video, screen-understanding task에서 소폭 하락한다.

이는 embodied training이 general capability를 완전히 보존했다기보다 trade-off를 관리했다는 의미다. Average가 비슷하더라도 어떤 domain이 내려갔는지 확인해야 한다.

5-3. Action generation

Paper는 action-token perplexity 개선을 보고한다.

Split Before After
Held-out in-distribution 17.12 5.07
RoboDojo 12.39 6.57

Qualitative trajectory는 object placement, tool use, bimanual manipulation, OOD scene에서 reference trend를 따라가는 사례를 보여준다.

하지만 perplexity는 token distribution fit을 측정한다. Real robot task completion, closed-loop recovery, collision, safety margin을 직접 측정하지 않는다.

5-4. Future-state prediction

Paper는 RGB, depth, mask의 qualitative examples를 제시한다. Short horizon에서 scene layout, object identity, geometry, robot region이 plausible하게 유지되는 것을 보여준다.

현재 evidence는 qualitative 중심이다. FVD, depth error, mask IoU, physical consistency, action-conditioned causality 같은 numerical metric은 제한적이다.

5-5. What really matters in the experiments

1) Evidence strength가 capability마다 다르다

  • Embodied understanding: broad quantitative benchmark
  • Action generation: perplexity와 qualitative trajectory
  • Future state: qualitative visualization
  • Closed-loop control: headline metric 없음

따라서 세 capability를 같은 confidence로 받아들이면 안 된다.

2) Base model comparison은 useful but not sufficient하다

Qwen3-VL 대비 28 benchmarks improvement는 physical-aware training의 효과를 보여준다. 그러나 data scale, task mixture, output format이 동시에 바뀐다. Action and future supervision이 understanding score를 얼마나 독립적으로 올렸는지 ablation이 필요하다.

3) Proprietary model comparison은 protocol symmetry를 확인해야 한다

Closed model의 prompt, image preprocessing, tool access, decoding, refusal policy가 open model과 완전히 같지 않을 수 있다. Overall gap이 작다는 headline보다 evaluation contract를 먼저 봐야 한다.

6. Limitations

  1. Closed-loop robot success가 없다.
    • Understanding score and action perplexity는 strong offline evidence지만 real-world control success를 대체하지 않는다.
    • Repeated interaction에서 compounding error and recovery capability를 확인해야 한다.
  2. Action representation이 local convention에 의존한다.
    • Global canonicalization을 피한 대신 previous action context and source metadata가 중요해졌다.
    • New robot의 scale and axis convention에 대한 calibration protocol이 필요하다.
  3. Human motion supervision에 estimation noise가 있다.
    • Wrist trajectory, depth, mask는 perception pipeline에서 recovery될 수 있다.
    • Human motion과 robot kinematics의 차이도 남는다.
  4. Future-state resolution and evaluation이 제한적이다.
    • $128 x 128$ discrete state는 coarse dynamics에는 유용하지만 contact-rich manipulation을 평가하기에는 낮다.
    • Quantitative physical consistency metric이 더 필요하다.
  5. General capability trade-off가 있다.
    • 일부 multimodal benchmark는 base model보다 낮다.
    • Unified training이 모든 domain의 Pareto improvement를 보장하지 않는다.
  6. Training cost가 매우 크다.
    • 약 30,000 hours human video, tens of millions samples, 8B model two-stage training은 full reproduction barrier가 높다.

7. My Take

7-1. Why this matters for my work

PhysBrain 1.5의 가장 중요한 contribution은 physical task를 token interface로 묶은 점이다. Action and future state를 special head가 아니라 vocabulary extension으로 다루면 standard LLM training stack, packing, masking, decoding infrastructure를 재사용할 수 있다.

이 방식은 multimodal document task에도 힌트를 준다. Bounding box, reading order, layout transition, UI action처럼 continuous or structured output을 discrete token family로 만들고, language reasoning과 same backbone에서 학습할 수 있다.

7-2. Reuse potential

1) Small-scale unified token experiment

Full 8B pre-training을 재현하지 않아도 language plus coordinate token, language plus action token의 interference를 smaller VLM에서 검증할 수 있다.

2) Action history as adaptation context

Robot-specific adapter 없이 previous trajectory를 context로 넣어 convention을 추론하는 idea는 test-time adaptation 관점에서 흥미롭다. 다만 explicit metadata baseline과 비교해야 한다.

3) Future state as auxiliary supervision

Policy output만 학습하는 대신 predicted depth and mask를 auxiliary token target으로 추가하면 representation이 physical transition을 더 잘 담는지 검증할 수 있다.

4) Evaluation contract separation

Understanding benchmark, offline action likelihood, simulator success, real robot success를 하나의 score로 합치지 않고 계층적으로 보고해야 한다.

7-3. Follow-up papers

  • Qwen3-VL Technical Report: base VLM architecture와 visual token interface
  • Vision-Language-Action literature: action-conditioned multimodal policy 연구
  • Action tokenization methods: discrete action vocabulary와 trajectory chunking
  • Video world model literature: action-conditioned future-state prediction 연구
  • RoboDojo and embodied benchmark suites: OOD action과 reasoning evaluation

8. Summary

  • PhysBrain 1.5는 language, action, future visual state를 one autoregressive vocabulary로 통합한다.
  • ActionPiece는 16-step wrist trajectory를 512-entry discrete vocabulary의 token sequence로 바꾼다.
  • Future state는 aligned RGB, depth, human or robot mask code로 생성된다.
  • Pre-training은 약 30,000 hours human interaction video에서 physical prior를 만들고, SFT에서 robot and simulation data를 추가한다.
  • 8B model은 28 embodied understanding benchmarks에서 average 72.5를 보고한다.
  • Understanding evidence는 강하지만 action and future-state evidence는 offline and qualitative 중심이며 closed-loop control은 추가 검증이 필요하다.

댓글남기기