10 분 소요

0. Introduction

Paper link

Project page

Code link

Dataset

Video world model에서 가장 눈에 띄는 결과는 몇 분 or 몇 시간의 rollout이다. 하지만 실제 연구 병목은 generation horizon 하나가 아니다. Source dataset마다 clip length, camera pose, caption style, quality metadata가 다르고, video backbone마다 latent representation, attention structure, objective, runtime environment가 다르다.

SolarWM은 이 문제를 single model proposal보다 open research stack으로 다룬다.

  • Multi-source video를 unified frame-aligned contract로 바꾼다.
  • Data processing과 training mixture design을 분리한다.
  • Wan2.2, LTX-2.5, MiniMax-H3의 native representation을 유지한다.
  • Shared camera-conditioning and training interface를 제공한다.
  • Bidirectional model을 causal few-step world model로 바꾸는 three-stage recipe를 적용한다.
  • Sliding-window inference로 short training horizon보다 훨씬 긴 rollout을 생성한다.

Paper의 headline은 5-second sequence로 학습하고 minute-to-hour rollout을 만든다는 것이다. 하지만 더 오래 남는 기여는 data contract, backbone registry, stage contract, release manifest에 가깝다.

한 줄 요약: SolarWM은 1.43M canonical video clips의 frame-aligned data engine과 backbone-native adaptation layer를 만들고, bidirectional adaptation, teacher-forced AnyFlow, DMD via self-gradient forcing의 three-stage recipe로 5B to 33B interactive video world model을 구성한 open stack이다.

이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.

  • World model 연구에서 data preparation부터 inference까지 끊기기 쉬운 reproducibility chain을 하나로 묶는다.
  • 10 source datasets를 14 independently addressable owners로 나눠 filtering and mixture를 재구성할 수 있게 한다.
  • Four model routes across Wan2.2, LTX-2.5, MiniMax-H3를 same framework에서 다룬다.
  • Camera geometry를 fused-PRoPE로 attention path에 넣는다.
  • 5-second training clip에서 60-minute endpoint까지 이어지는 uninterrupted autoregressive example을 제시한다.
  • 동시에 long-horizon evidence가 qualitative 중심이라는 한계도 분명하다.

1. Problem Setting

1-1. Heterogeneous video data를 단순히 섞을 수 없는 이유

Video world model data는 image dataset보다 contract가 복잡하다.

  • Frame rate와 clip duration
  • Camera intrinsics와 extrinsics
  • Metric scale 정보
  • Scene cut와 transition
  • Dynamic object motion 정보
  • Caption granularity 정보
  • Quality와 compression artifact
  • License와 provenance
  • Selection and rejection reason

Dataset마다 이 field의 availability and convention이 다르다. Raw clip을 한 folder에 넣고 random sampling하면 model이 source-specific artifact, caption style, camera convention을 함께 학습할 수 있다.

SolarWM은 physical corpus와 logical recipe를 분리한다.

  • Physical corpus는 clip and annotation을 한 번 만든다.
  • Recipe는 source weight, quality tier, repeat factor, temporal view를 지정한다.
  • New mixture를 만들 때 expensive source processing을 다시 하지 않는다.

이 설계는 data ablation을 model training과 독립적으로 만들려는 시도다.

1-2. Video backbone마다 adaptation path가 다르다

Wan2.2, LTX-2.5, MiniMax-H3는 parameter count만 다른 model이 아니다.

  • Latent encoder and temporal compression이 다르다.
  • Attention and positional encoding이 다르다.
  • Flow-matching objective and noise schedule이 다르다.
  • Text and image condition path가 다르다.
  • Runtime and distributed environment가 다르다.

모든 backbone을 하나의 wrapper에 억지로 맞추면 pretrained behavior를 깨뜨릴 수 있다. 반대로 route마다 independent codebase를 만들면 data, training, evaluation을 비교하기 어렵다.

SolarWM은 common contract와 backbone-native implementation을 분리한다.

1-3. Long-horizon world model의 exposure bias

Bidirectional video generator는 full clip의 noisy token을 동시에 denoise한다. Interactive world model은 past generated frames를 condition으로 다음 chunk를 autoregressively 만든다.

Training에서는 clean history를 보고 inference에서는 model-generated history를 보면 exposure bias가 생긴다. Small error가 next window의 input이 되고, long horizon에서 drift가 누적된다.

SolarWM의 three-stage recipe는 이 transition을 단계적으로 만든다.

2. Core Idea

2-1. Reconfigurable data engine

SolarWM sample contract를 개념적으로 쓰면 다음과 같다.

\[\mathcal{S} = \{ V, K_{1:T}, P_{1:T}, c, q, s, p \}\]
  • $V$: visual frames or encoded video
  • $K_{1:T}$: frame-aligned camera intrinsics
  • $P_{1:T}$: frame-aligned camera poses
  • $c$: caption 정보
  • $q$: quality와 motion metadata
  • $s$: selection decision과 rejection reason
  • $p$: source provenance 정보

같은 frame index를 기준으로 observation, camera, caption, quality, selection을 묶는다. Dataset reader가 source name을 보고 camera normalization을 추측하지 않도록 contract를 명시한다.

2-2. Ten source datasets and fourteen owners

Paper abstract는 10 source datasets라고 쓰고, repository는 14 dataset owners라고 쓴다. 둘은 counting unit이 다르다.

  • DL3DV는 10-second and 60-second view로 분리된다.
  • MiraData, Sekai-Walking, SpatialVID에는 clean-plate derivative가 추가된다.
  • Original source 기준으로는 10개다.
  • Independently addressable recipe owner 기준으로는 14개다.

Blog에서는 이를 “10 source datasets, 14 logical owners”로 구분하는 편이 안전하다.

2-3. Backbone-native adaptation

SolarWM core는 concrete backbone을 import하지 않는다. Common layer는 아래를 관리한다.

  • 엄격하게 versioning된 config
  • Canonical data index 관리
  • Distributed topology 관리
  • Checkpoint와 EMA policy 관리
  • Train, infer, preencode entry point 제공
  • Manifest와 validation output 생성

Backend plugin은 model-specific math를 책임진다.

  • Backbone-native latent representation
  • Backbone-native noise와 objective
  • Camera condition mapping 구현
  • Attention patch 구현
  • Stage-specific loss 구현

이 boundary 덕분에 shared recipe를 쓰면서도 backbone-specific behavior를 유지한다.

2-4. Fused-PRoPE camera conditioning

SolarWM은 camera pose and intrinsics를 projective rotary positional encoding으로 self-attention path에 넣는다.

기존 video RoPE가 temporal-spatial location을 나타내면, projective transform은 camera geometry를 query, key, value tensor에 반영한다. 별도 control tower or extra attention pass보다 native attention path를 확장하는 방식이다.

중요한 점은 common camera contract는 같지만, 각 backbone이 tensor shape and RoPE convention에 맞게 mapping한다는 것이다.

2-5. Three-stage causal adaptation

Stage0.5: Bidirectional flow matching

Full clip을 bidirectionally 학습해 video, text, camera condition representation을 안정화한다. 대부분의 visual prior and camera control capability를 이 stage에서 만든다.

Stage1: Teacher-forced AnyFlow

Clean history를 condition으로 noisy target chunk를 학습한다. Denoising and finite-step flow map을 함께 학습해 causal rollout의 initializer를 만든다.

Stage2: DMD via self-gradient forcing

Causal student가 자신의 autoregressive history 위에서 rollout한다. Frozen teacher and trainable critic의 distribution signal로 few-step student를 학습한다.

Stage2는 inference distribution을 training 안으로 가져와 exposure bias를 줄이는 역할을 한다.

3. Architecture / Method

3-1. Overview

Axis SolarWM design
Data scale 1,425,694 canonical clips, about 25.85 TB
Source organization 10 original sources, 14 logical owners
Quality tiers high, xhigh, rejected with reasons
Backbones Wan2.2-5B, Wan2.2-14B, LTX-2.5-22B, MiniMax-H3-33B
Camera interface Fused-PRoPE
Training Stage0.5 FM, Stage1 TF-AnyFlow, Stage2 DMD via SGF
Deployment Four-step causal sampling, sliding-window rollout
Runtime Separate environment per backbone, shared SolarWM CLI and manifest

3-2. Data selection as metadata operation

SolarWM은 rejected clip을 버리지 않는다. Released policy 기준으로 approximately 876K clips가 high or xhigh이고, approximately 549K clips가 rejected partition에 남는다.

Rejected sample도 metadata and reason을 유지한다. 사용자는 threshold or source policy를 바꿔 new mixture를 만들 수 있다.

이 design은 quality filtering을 irreversible preprocessing이 아니라 versioned recipe로 바꾼다.

3-3. Clean-plate derivatives

Dynamic people and vehicles는 camera motion과 독립적인 motion을 만든다. Camera-controlled world model에서는 이 motion이 supervision ambiguity가 될 수 있다.

SolarWM data engine은 일부 source에서 dynamic actor를 제거한 clean-plate derivative를 별도 owner로 만든다. Original and clean version 중 하나를 superior data로 간주하지 않고, 서로 다른 supervision type으로 관리한다.

3-4. Sliding-window inference

Causal model은 fixed-length window를 반복한다.

\[\hat{V}_{t:t+H} = F_{\theta} ( \hat{V}_{t-W:t}, C_{t:t+H}, text )\]
  • $W$: history window 크기
  • $H$: 다음 generated chunk 크기
  • $C$: camera trajectory condition

Generated chunk 일부가 next window history로 들어간다. Sequence length는 bounded하지만, total rollout horizon은 반복 횟수만큼 늘어난다.

이 방식은 memory cost를 제한하는 대신 old scene state를 explicit하게 보존하지 않는다. Window 밖 정보가 사라질 수 있어 long-term identity and topology drift가 핵심 risk다.

3-5. Four models, one interface의 실제 범위

Repository release status를 보면 모든 route가 same maturity는 아니다.

  • Wan2.2-5B: Stage0.5, Stage1, Stage2가 available
  • MiniMax-H3: Stage0.5, Stage1, Stage2가 available
  • Wan2.2-14B: later causal stages는 coming soon일 수 있음
  • LTX-2.5: later causal stages는 coming soon일 수 있음

Paper의 model family claim과 currently executable release matrix를 구분해야 한다.

4. Training / Data / Recipe

4-1. Corpus statistics

Property Reported value
Canonical clips 1,425,694
Logical owners 14
Approximate storage 25.85 TB
High tier 471,798
Xhigh tier 404,795
Rejected 549,101

Data source는 real-world, synthetic, game environment를 포함한다. Caption style도 unified contract로 맞춘다.

4-2. Decoupling processing and mixture

Source processing은 아래 expensive operation을 담당한다.

  • Clip extraction 수행
  • Camera geometry recovery 수행
  • Captioning 수행
  • Quality와 motion metric 계산
  • Clean-plate transform 수행
  • Sharding과 provenance 기록

Mixture recipe는 아래 inexpensive operation을 담당한다.

  • Split membership 지정
  • Quality tier 지정
  • Source weight 지정
  • Repeat factor 지정
  • Temporal window 지정
  • Backend-specific view 지정

이 분리 덕분에 data ablation마다 raw video를 다시 처리할 필요가 없다.

4-3. Stage0.5

Bidirectional adaptation은 pretrained video generator를 camera-conditioned backbone으로 만든다.

이 stage가 중요한 이유는 later causal stage가 visual quality and camera controllability를 처음부터 다시 학습하지 않게 하기 때문이다. Paper는 most learning이 bidirectional stage에서 일어나고, causal adaptation and DMD는 shorter convergence path를 가진다고 해석한다.

4-4. Stage1

Teacher forcing은 clean past and noised future를 사용한다. AnyFlow objective는 arbitrary finite-step map을 학습해 별도 ODE distillation initializer 없이 Stage2로 넘어가게 한다.

Practical point는 Stage1이 ordinary autoregressive next-frame prediction이 아니라 denoising and finite-step flow를 함께 학습한다는 것이다.

4-5. Stage2

DMD with self-gradient forcing는 student-generated history를 사용한다.

  • Student가 causal rollout을 생성한다.
  • Frozen teacher가 target distribution signal을 준다.
  • Trainable critic이 student and teacher distribution gap을 추정한다.
  • Few-step student를 autoregressive inference distribution에 맞춘다.

이 stage가 long rollout robustness를 보장하는 것은 아니지만, clean-history-only training보다 exposure bias에 직접 대응한다.

4-6. Engineering notes

1) Resolved config and manifest를 artifact로 남겨야 한다

Backbone, stage, data generation, source weights, camera convention, checkpoint, latent encoder를 하나의 manifest에 묶어야 result를 재현할 수 있다.

2) Raw, latent, annotation release를 구분해야 한다

Public dataset page가 존재한다고 full raw payload가 one-click download라는 뜻은 아니다.

  • Preencoded latent generation은 recipe별 repository에서 받을 수 있다.
  • Annotation package로 upstream source video를 다시 받아 raw-WDS를 rebuild할 수 있다.
  • Prepared raw-WDS는 access form이 필요할 수 있다.

3) Backbone license를 separately audit해야 한다

SolarWM core code는 Apache-2.0이지만 LTX and MiniMax derivative는 별도 community license and territory restriction을 가질 수 있다.

4) Long rollout evaluation log가 필요하다

Sparse frame grid만으로는 drift onset을 찾기 어렵다. Per-window identity, geometry, camera error, motion response, failure time을 log해야 한다.

5) Camera control and object interaction을 분리해야 한다

SolarWM의 strongest evidence는 camera-controllable scene continuation이다. Agent action이 object state를 바꾸는 causal world interaction과 동일하게 읽으면 과도하다.

5. Evaluation

5-1. Long-horizon demonstration

Paper는 real first frame에서 시작해 fixed scene prompt and predetermined camera trajectory를 따라가는 uninterrupted rollout을 제시한다.

  • Training sequence: about 5 seconds
  • Inference 범위: minute-scale to hour-scale
  • Highlight: 60-minute endpoint
  • Later frame: generated history에 autoregressively condition
  • No restart from input image
  • No reference-frame injection
  • No attention sink

이 example은 horizon extrapolation이 가능함을 보여준다. 그러나 sparse sampled frame이 main evidence이며 standard long-video benchmark score는 제한적이다.

5-2. Four-step interaction

Distilled causal models은 reported setup에서 four sampling steps and 16 fps interaction을 목표로 한다. Few-step world model에서 quality and latency를 동시에 보려는 design이다.

실제 real-time 여부는 backbone size, resolution, hardware, sequence parallelism, window size에 따라 달라진다. “Real-time”을 universal latency claim으로 읽으면 안 된다.

5-3. Cross-backbone validation

Same data contract and staged recipe를 5B to 33B route에 적용한 점이 important하다. 이는 method가 one model architecture에만 맞는 hack이 아니라 framework-level abstraction일 가능성을 보여준다.

다만 release matrix and quantitative result의 completeness가 route마다 다르므로, four backbone을 equal evidence로 해석하면 안 된다.

5-4. What really matters in the experiments

1) Long duration and long-term memory는 같은 지표가 아니다

60 minutes를 생성할 수 있다는 것은 runtime continuity를 보여준다. Old object identity, hidden room topology, causal event, inventory state를 정확히 기억한다는 뜻은 아니다.

2) Camera responsiveness and physical validity를 분리해야 한다

Camera trajectory를 따라가며 coherent scene을 유지하는 것과 action-conditioned object physics는 다르다. Current evidence는 전자에 더 가깝다.

3) Sparse frame sampling은 failure duration을 숨길 수 있다

Endpoint가 recognizable해도 중간에 transient collapse, texture reset, object disappearance가 있었을 수 있다. Dense temporal metric and failure timeline이 필요하다.

4) Open stack의 operational completeness가 핵심 평가 대상이다

이 논문의 contribution은 benchmark score만이 아니다. Data contract, config, backend registry, checkpoint, preencoding, license documentation이 실제로 연결되는지 확인하는 것이 중요하다.

6. Limitations

  1. Long-horizon evidence가 qualitative 중심이다.
    • Standardized FVD, VBench, identity-memory, geometry-drift curve가 headline result에 충분히 포함되지 않는다.
    • 60-minute sparse frame는 strong demonstration이지만 complete reliability test는 아니다.
  2. Sliding window는 old state를 explicit하게 보존하지 않는다.
    • Window 밖의 object, topology, event history가 사라질 수 있다.
    • Hour-scale visual coherence와 persistent world state를 구분해야 한다.
  3. Camera-controlled video와 general interactive world model 사이에 gap이 있다.
    • User camera trajectory response는 잘 정의되어 있지만 object manipulation, agent action, contact dynamics는 더 많은 evidence가 필요하다.
  4. Release completeness가 backbone마다 다르다.
    • Full three-stage path가 available한 route와 coming soon인 route가 섞여 있다.
  5. Fully open이라는 표현에는 access and license condition이 있다.
    • Full raw payload는 annotation rebuild or access process가 필요할 수 있다.
    • Upstream media and backbone license는 SolarWM Apache license와 별개다.
  6. Compute barrier가 높다.
    • 25 TB-scale data, video preencoding, 5B to 33B training, multiple runtime environments는 small lab reproduction을 어렵게 한다.

7. My Take

7-1. Why this matters for my work

SolarWM의 핵심은 hour-long video보다 contract engineering이다. World model project는 model code보다 data provenance, camera convention, latent generation, stage transition, checkpoint compatibility에서 더 자주 무너진다.

SolarWM은 research idea를 repeatable pipeline으로 만들기 위해 어떤 boundary가 필요한지 보여준다.

  • Physical data와 logical mixture의 분리
  • Common framework와 backbone-native math의 분리
  • Bidirectional capability와 causal adaptation의 분리
  • Quality model과 distilled deployment model의 분리
  • Public annotation과 prepared raw payload의 분리

7-2. Reuse potential

1) Dataset contract

Video, camera, caption, quality, rejection reason, provenance를 frame index 기준으로 묶는 schema는 다른 world-model project에도 바로 적용할 수 있다.

2) Rejected data retention

Filtered sample을 삭제하지 않고 reason-coded partition에 남기면 threshold ablation and future recipe를 재구성할 수 있다.

3) Backend registry

Model-specific latent and objective는 plugin에 두고, training orchestration and manifest를 common layer에 두는 architecture는 multi-backbone research에 유용하다.

4) Horizon audit

Long rollout은 endpoint showcase보다 time-to-failure, per-window drift, recovery, state persistence를 중심으로 평가해야 한다.

7-3. Follow-up papers

  • LIVE: long-horizon interactive video world modeling 연구
  • Memory Forcing: scene consistency를 위한 explicit spatio-temporal memory
  • H3-World: language-conditioned interactive video world model 연구
  • World Action Models survey: future prediction과 action coupling taxonomy
  • Video diffusion distillation과 DMD 관련 연구

8. Summary

  • SolarWM은 data preparation, training, inference를 묶은 open video world-model stack이다.
  • 10 original sources의 1.43M clips를 하나의 frame-aligned contract 아래 14 logical owners로 노출한다.
  • Wan2.2, LTX-2.5, MiniMax-H3의 native representation을 유지하며 shared framework를 사용한다.
  • Three stages는 bidirectional FM, teacher-forced AnyFlow, DMD via SGF로 구성된다.
  • Five-second training sequence에서 minute-to-hour sliding-window rollout을 제시한다.
  • 가장 중요한 한계는 hour-scale evidence가 qualitative하고, persistent memory and physical interaction metric이 부족하다는 점이다.

댓글남기기