Source-linked AI summary

LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting

Louen Pottier

arXiv:2608.16324v1cs.CVcs.LG

TL;DR

기존 video 기반 scene reconstruction은 상호작용 없이 기록된 motion을 재생하는 반면, 학습된 physical simulator는 sensor나 simulation에서 얻은 generalized-coordinate data를 요구한다. LaGSplat은 latent dissipative Lagrangian과 움직이는 Gaussian primitive를 결합해 monocular video에서 interactive dynamics를 추론하며, 7개 real system에서 관측된 frequency를 재현하고 soft robot에서 baseline보다 우수한 성능을 보인다.

  • 문제

    Video 기반 scene reconstruction은 상호작용 없이 기록된 motion을 재생하는 반면, 학습된 physical simulator는 sensor나 simulation에서 얻은 generalized-coordinate data를 요구한다.

  • 방법

    LaGSplat은 하나의 비지도 low-dimensional latent state를 dissipative Lagrangian의 generalized coordinate이자 움직이는 Gaussian primitive의 conditioning variable로 사용한다.

  • 결과

    7개 real system에서 excitation이 가해졌을 때 natural frequency는 관측값과 몇 퍼센트 이내로 일치하며, LaGSplat은 2-segment soft robot에서 5개 model 중 가장 정확하다.

  • 시사점 및 한계

    재구성된 object에는 image-space 또는 volume-space force를 가할 수 있으며, 이 force는 decoder Jacobian을 통해 learned dynamics로 전달되고 retraining 없이 작동한다.

  • 시사점 및 한계

    이 방법은 generalized coordinate가 적고 kinematics가 smooth하며 dynamics가 dissipative인 system에 적용되며, contact, impact, topology change, plasticity, hysteresis는 제외한다.

Abstract

from arXiv · show

We present LaGSplat (Latent Lagrangian Gaussian Splatting), a framework that infers interactive, physics-governed dynamics from one or a few monocular videos. At inference it lets a user push on the filmed object, rigid or deformable, with an external force that was never measured, annotated, or seen during training. This is possible because a low-dimensional latent state $\mathbf{q} \in \mathbb{R}^d$ plays two roles at once: it is the generalised coordinate of a learned dissipative Lagrangian and the conditioning variable of a Gaussian Splatting decoder. The inductive bias of this decoder, whose primitives are explicit points $μ_i(\mathbf{q})$ that move with the object, is what lets a force $f$ applied in the image pull back into a latent generalised force $J(\mathbf{q})^\top f$ and enter the equations of motion, which pixel-space (CNN) or neural-field (NeRF) decoders cannot do. We validate LaGSplat on test cases of increasing difficulty, from rigid to deformable and from autonomous to forced real systems, combining monocular video and sensor measurements. We further demonstrate interactive use: forces of arbitrary magnitude and direction can be applied to the reconstructed object at any time, its response rendered in real time, in 2D or 3D. Assuming a dissipative Euler-Lagrange equation over a few generalised coordinates trades generality for a bounded, plausible response to unseen forces, where an unconstrained predictor diverges.

1 서론

LaGSplat은 explicit Gaussian primitives와 학습된 latent Lagrangian을 결합해 촬영된 객체를 실시간 physics-governed interactive simulation으로 재구성한다. 이동하는 primitive는 재학습 없이 image-space force를 latent generalized force로 전달하는 데 필요한 Jacobian을 제공한다.

  • 동기: 기록된 클립을 재생하는 time-parameterized dynamic scene reconstruction과 달리, LaGSplat은 재구성된 객체와의 상호작용을 지원한다.Dynamic scene reconstruction은 실시간 photorealistic representation을 제공하지만, time-based parameterization으로는 객체와 상호작용할 수 없다.
  • 기여: LaGSplat은 genuine mechanical law와 객체와 함께 이동하고 객체를 렌더링하는 explicit primitives를 결합해 공백을 메운다.기존 방법은 persistent explicit moving primitives 없는 structured dynamics 또는 genuine mechanical law 없는 explicit reconstruction 중 하나만 제공한다.
  • 방법: low-dimensional latent state q는 학습된 dissipative Lagrangian의 generalized coordinate이자 Gaussian Splatting decoder의 conditioning variable이다.Encoder가 frame을 q로 매핑하고, latent Lagrangian이 equations of motion을 적분하며, decoder가 q를 2D 또는 3D image로 렌더링한다.
  • Interactive forcing: Explicit decoder point μ_i(q)는 latent variation에 따라 이동하므로, virtual work를 통해 image-space force f가 generalized force J(q)^T f가 된다.training video에서 force가 전혀 가해지지 않았더라도 이 force는 재학습 없이 동일한 equations of motion에 들어간다.
  • 기여: 이 pipeline은 monocular video를 식별된 physical model과 2D 또는 3D로 렌더링되는 real-time interactive simulation으로 변환한다.Explicit Gaussian primitives와 differentiable rasterization을 latent Lagrangian 및 generalized-force dynamics와 결합한다.

2 관련 연구

관련 연구는 동적 렌더링, 학습된 mechanics, video-supervised dynamics, force-interactive model을 아우른다. LaGSplat은 latent coordinate 위의 학습된 dissipative Lagrangian과, Jacobian으로 force를 전달하는 명시적 material-point Gaussian decoder를 결합한다.

  • 분야 개요: 문헌은 dynamic novel-view synthesis, analytical mechanics, video-generated physical rendering, 그리고 simulated system의 footage에서 dynamics를 직접 학습하는 방법으로 나뉜다.핵심 구분은 video가 output인지 supervision인지, 그리고 model이 촬영된 system을 재현하는지 아니면 해당 종류에 그럴듯한 scene을 재현하는지에 있다.
  • 렌더링 표현: NeRF-based decoder는 Eulerian 방식이므로 force를 부착할 material point를 제공하지 않지만, 명시적으로 움직이는 primitive는 position과 Jacobian을 제공한다.canonical primitive를 warp하는 neural network는 부착 기준을 만족하지만, motion-rich content에는 direct primitive가 선호된다.
  • Physics와 force control: PhysGaussian [37] [45] [74] 및 VR-GS [29]와 같은 prescribed-physics method는 external dynamics를 사용하는 반면, Force Prompting [18]은 mechanical structure 없이 force response를 학습한다.Force Prompting [18]은 force-video example을 요구하고, 생성된 각 clip에 고정된 force vector를 사용하며, real-time도 아니고 per-frame controllable하지도 않다.
  • Latent mechanics: NeuROK [17] 및 관련 latent-mechanics method는 latent integration과 Jacobian pullback을 공유하지만, 이들의 decoder는 affine Gaussian primitive가 아니라 mesh를 렌더링한다.Reduced-order graphics method도 material-point decoder와 pullback structure를 사용하지만, LaGSplat은 학습된 state 위에서 governing law를 학습한다.
  • LaGSplat의 차별점: LaGSplat은 higher-dimensional Gaussian에서 time을 latent coordinate로 대체하고, dissipative Lagrangian을 학습하며, J^T f를 통해 임의의 unseen force를 전달한다.Wang et al. [69]와 같은 rigid-only force-interactive system과 달리, deformable object를 다루고 어느 순간에도 force를 재정의할 수 있다.

3 방법

LaGSplat은 소산 Euler–Lagrange dynamics와 state-conditioned Gaussian Splatting decoder를 동시에 parameterize하는 연속 latent state를 학습한다. 학습된 kinematics는 image-space force를 latent generalized force로 끌어와 observation geometry와 force-scale ambiguity를 고려하면서, 관측되지 않은 loading에도 bounded interactive response를 가능하게 한다.

  • Architecture와 training: LaGSplat은 encoder, dissipation이 포함된 latent Lagrangian, Gaussian Splatting decoder를 공동 학습한 뒤 autoencoder를 고정하고 encoded trajectory에 맞춰 dynamics를 학습한다.Stage 1에서는 appearance를 맞추고 latent chart를 고정한다. Stage 2에서는 clip을 한 번 encoding한 뒤 Lagrangian을 학습한다.
  • Architecture와 training: Lipschitz-constrained encoder는 dynamics objective가 encoded trajectory를 미분해 q̇와 q̈를 얻기 때문에 temporal continuity를 regularize한다.이 방법은 encoder의 깊이나 표현력을 높이는 대신 Zhu et al. [81]의 continuity regularizer를 사용한다.
  • Gaussian Splatting decoder: Decoder는 space–state volume (x, y, q)을 tile하므로 각 scene state를 한 번만 저장하고 Euler–Lagrange trajectory, 관측되지 않은 initial condition, 또는 applied force에서 render할 수 있다.Gaussian primitive을 q에 conditioning하면 state-dependent center가 생성되고 query한 latent state에서 먼 primitive이 억제된다. Rendering은 n=2 video output 또는 n=3 interactive 3D scene을 지원한다.
  • Latent dynamics: 학습된 latent chart는 whiten하며, error는 decoded pixel motion에 따라 latent direction에 가중치를 부여하는 anisotropic pull-back image metric Ḡ에서 측정한다.이 metric은 evaluation 중 rasterization 없이 decoded image error를 first order로 근사한다.
  • Latent dynamics: Dissipation과 unique minimum을 갖는 potential은 bounded free dynamics를 보장하여 정지 상태로 settle하게 하고, force가 유지되면 loaded equilibrium으로 수렴하게 한다.ICNN은 pendulum, rocking chair, hanging bag에 대해 하나의 convex basin을 모델링한다. 반면 soft continuum robot은 i-ResNet과 결합한 ICNN을 사용해 unique equilibrium을 보존하면서 convex level set을 완화한다.
  • Interactive forcing: Gaussian kinematics Jacobian은 primitive force를 latent generalized force로 pull primitive forces back하며, neighborhood contact aggregation과 O(nd) inference cost로 network differentiation 없이 관측되지 않은 loading을 가능하게 한다.Response direction과 shape은 학습된 kinematics가 결정하지만, physical force magnitude는 하나의 global energy factor κ까지 ambiguity가 남는다. 이는 (L, D)와 (αL, αD)가 동일한 observed trajectory를 생성하기 때문이다.

4 실험

LaGSplat은 자율 rigid 및 deformable system과 압력 측정이 포함된 연속 구동 soft robot의 monocular real-object video에서 평가된다. 이 테스트 전반에서 그럴듯한 dynamics를 복원하며, decoder에서 유도한 Jacobian을 통해 inference 시 image-space 또는 world-space external force를 지원하는 유일한 방법이다.

  • 전체 dynamics 평가: 거의 모든 system에서 비지도 natural-frequency error εfreq는 1–6%이며, one-segment robot은 recording이 natural mode를 가진하게 자극하지 못해 31.4%에 이른다.double pendulum의 εq = 0.947은 전체 clip에서 1.544로 증가하는 반면, hanging-bag error는 두 주기에서 0.150에서 일곱 주기에서 0.396으로 증가한다.
  • Soft continuum robot: LaGSplat은 더 어려운 two-segment robot에서도 정확도를 유지하며, d=4 coordinates로 0.769 × 10−3을 달성하고 해당 system에서 comparison model을 ×0.98만큼 능가한다.이 결과는 다섯 seed의 median이며, 네 baseline은 모두 d=10을 사용한다. one-segment recording은 주로 inertia나 natural frequency가 아니라 quasi-static pressure-to-shape map을 드러낸다.
  • Interactive force 적용: External image force는 retraining 없이 inference 시 f_lat = J^T f로 LaGSplat의 learned dynamics에 들어가며, Gaussian map q 7→μ_i가 유도한 Jacobian을 사용한다.decoder는 image primitive에서 latent generalized coordinate로 force를 끌어오고, training은 autonomous dynamics와 actuation map을 식별한다. 이 capability는 비교 대상 entry에는 없다.
  • Interactive force 적용: 1차원 rainbow rocker에서 동일한 force는 lever arm에 따라 달라지고 부호가 반전되는 latent shift를 생성하며, application point에 따라 Δq = −0.50, −0.76, −0.84 및 +1.22를 포함한다.null Jacobian을 갖는 background Gaussian은 force를 전달하지 않으므로, response가 임의의 image content가 아니라 object motion을 따른다는 점을 보여준다.
  • 3차원 interaction: multi-view synthetic supervision을 사용하면 동일한 transport mechanism이 3D에서도 작동한다. ray-selected Gaussian에 world-space force를 가하면 object를 q = +1.52의 새로운 equilibrium으로 이동시킬 수 있다.시연에는 training motion에 없던 방향으로의 push가 포함되어, 관측된 trajectory를 넘어서는 interactive response를 보여준다.

5 논의

LaGSplat의 physics prior는 구조를 개선하지만 적용 범위를 관측 가능하고 매끄러운 저차원 dissipative system으로 제한하며, force mechanism은 측정된 force를 대상으로 아직 검증되지 않았다. 향후 연구는 force-labeled data, hidden constitutive variable, simulator coupling, 개선된 latent 및 geometric modeling을 대상으로 한다.

  • 한계: Lagrangian prior는 contact, impact, topology change, plasticity, hysteresis를 배제하므로, 방법을 소수의 coordinate를 갖는 매끄러운 dissipative dynamics로 제한한다.이 경계는 representation 자체가 아니라 구체화된 dynamics에서 비롯된다.
  • 한계: Monocular learning은 image-observable state를 요구하므로, out-of-plane motion에는 multiple-view supervision이 필요하고 plasticity에는 숨은 loading-history 정보가 필요하다.시각적으로 일치하는 frame도 서로 다른 plastic history를 인코딩할 수 있어, 이러한 variable은 single-frame supervision으로는 보이지 않는다.
  • 한계: 중심적인 force mechanism은 measured force를 대상으로 검증되지 않았다. fixed-view video, few degrees of freedom, applied point force, resulting deformation을 함께 포함하는 dataset이 없기 때문이다.Section 4.5에서는 대신 energy factor κ에 불변인 Jacobian-derived ratio, sign, sum 및 pullback-kernel property를 점검한다.
  • 한계: Appearance-first fitting은 여전히 논쟁적인 design choice이며, latent dimension d를 수작업으로 선택하면 minimal state size를 과대 추정할 수 있다.Friedl et al. [14]은 jointly trained model이 sequential training보다 우수하다고 보고한 반면, Zhu et al. [81]은 frozen-encoder approach가 jointly trained baseline보다 우수하다고 보고했다. Chen et al. [8]은 minimal latent width를 넘어설 때 reconstruction이 저하된다고 보고했다.
  • 향후 연구: 향후 연구는 measured-force dataset, plasticity와 hysteresis를 위한 internal variable, external simulator와의 coupling, per-primitive density에 대한 개선된 처리를 제안한다.Internal-variable model은 더 큰 d와 multi-frame encoder를 필요로 하며, simulator coupling은 force mechanism의 isolation test를 보존해야 한다.

6 결론

LaGSplat은 영상만으로 공유된 저차원 물리 상태를 학습하고, 이를 일반화 좌표이자 Gaussian decoder의 조건으로 사용한다. 실제 시스템 전반에서 관측된 동역학을 정확히 복원하며, 재학습 없이 이전에 보지 못한 힘에도 제한적이고 그럴듯한 응답을 생성한다.

  • LaGSplat은 비지도 물리 상태 q를 학습하며, 이 상태는 소산 Lagrangian을 매개변수화하는 동시에 명시적 이동점 Gaussian primitive를 조건화한다.영상 자체가 학습 신호를 제공하며, q에는 직접적인 supervision이 없다.
  • 7개의 실제 시스템에서 학습된 고유 진동수는 해당 시스템을 가진 모든 기록에서 영상 관측과 몇 퍼센트 이내로 일치한다.시스템은 자유도 하나에서 네 개까지 아우른다.
  • Krauss et al.’s [35]의 두 세그먼트 soft continuum robot 데이터에서 LaGSplat은 10개 대신 4개의 좌표를 사용하면서 5개 모델 중 가장 정확하다.평가는 저자들이 직접 수집한 데이터, 분할, 프로토콜을 사용한다.
  • 추론 시 영상 또는 재구성된 볼륨의 힘은 재학습 없이 decoder Jacobian을 통해 학습된 방정식에 들어가며, 학습 중 힘 측정값이 없었음에도 가능하다.힘의 방향과 형태는 학습된 kinematics를 따르고, 크기는 하나의 전역 에너지 인자로만 결정된다.
  • Lagrangian prior는 도달 가능한 시스템의 범주를 제한하지만, 보지 못한 힘에 대한 응답을 제한적이고 물리적으로 그럴듯하게 유지한다.

A 유도

이 부록은 학습된 운동학이 힘과 관성을 어떻게 전달하는지 유도하고, 유계 응답 주장을 뒷받침하는 에너지 균형을 설명한다.

  • A.1: 학습된 운동학은 힘과 관성을 모두 전달한다.이 유도는 Section A.1에서 다룬다.
  • A.2: 이 부록은 유계 응답 언급의 근거가 되는 에너지 균형을 유도한다.이 분석은 Section A.2에서 다룬다.

A.1 학습된 운동학에서의 힘과 관성

학습된 운동학은 latent configuration을 primitive position으로 매핑하여 추가 모델링 선택 없이 generalized force와 kinetic energy를 결정한다. Force transport는 smooth latent-coordinate change에 불변이며, normalized primitive opacity는 configuration-dependent mass matrix가 trajectory를 따라 변하도록 한다.

  • A.1 학습된 운동학에서의 힘과 관성: 학습된 configuration-to-primitive map은 표준 analytical mechanics 규칙에 따라 generalized force와 kinetic energy를 모두 결정한다.이는 각 particle의 position 역할을 하므로 map이 고정되면 추가적인 모델링 선택이 필요하지 않다.
  • A.1 학습된 운동학에서의 힘과 관성: primitive에 가해진 force는 primitive Jacobian을 통해 canonical latent generalized force로 transport되며 virtual work를 보존한다.이 구성은 equations of motion에 force를 주입하는 과정이 energy balance와 일관되도록 한다.
  • A.1 학습된 운동학에서의 힘과 관성: transport된 force와 그에 따른 response는 latent coordinate의 임의의 smooth reparameterization에 불변이다.Jacobian transformation과 virtual-displacement transformation이 상쇄되어 latent chart가 달라져도 work는 변하지 않는다.
  • A.1 학습된 운동학에서의 힘과 관성: Normalized latent opacity는 equal-mass primitive에 가중치를 부여하여 configuration-dependent mass matrix를 생성한다.primitive velocity가 dot(μ_i) = J_i dot(q)를 만족하므로 state dependence는 gating weight를 통해 유입된다. 따라서 mass matrix는 trajectory를 따라 변하며 각 J_i가 constant인 경우에도 Coriolis term을 갖는다.

A.2 유계 외력 응답

양의 정부호 관성, coercive potential, 균일한 양의 감쇠가 주어지면 LaGSplat의 continuous-time dynamics는 유계이고 dissipative하게 유지된다. Free rollout은 유일한 potential minimum으로 수렴하며, 유지되는 constant force는 loaded equilibrium으로 유계 수렴한다. symplectic integration은 discrete drift를 제한하지만, 분포 밖의 정량적 오차까지 제한하지는 않는다.

  • 가정: 유계 응답 보장은 positive-definite inertia, coercive potential floor, 그리고 isotropic floor c0 > 0을 갖는 damping에 의존한다.이러한 가정은 운동에너지와 dissipative control을 보장하면서 위치의 무한 drift를 방지한다.
  • Free rollout: coercive potential, positive-definite inertia, uniformly positive damping 하에서 Free rollout은 유계이며 유일한 global potential minimum으로 수렴한다. 이는 LaSalle’s invariance principle 에 따른다.Ė ≤ −c0∥q̇∥2에 따라 에너지가 감소하고, energy sublevel set은 q̇와 q를 모두 유계로 만든다.
  • Maintained force: Maintained constant force는 ∂V/∂q = f를 만족하는 loaded equilibrium으로 수렴하는 유계 응답을 생성한다.Loaded energy E_f = T + V − f^Tq는 주입된 power를 상쇄하며, quadratic potential floor는 이를 coercive하게 유지한다.
  • Discrete rollout: symplectic scheme은 장시간의 discrete rollout을 drift에 맞서 보존하지만, 유계성은 training data에서 멀리 떨어진 경우의 정량적 정확성을 의미하지 않는다.따라서 continuous time에 대해 제시된 보장은 응답 증가를 제한하지만 정확한 extrapolation을 보장하지는 않는다.

B 동일한 예산에서의 Deformation-Field Decoder

동일한 학습 예산에서 LaGSplat의 affine latent-to-primitive decoder는 재구성 성능을 유지하거나 향상시키면서, learned deformation field보다 실질적으로 더 매끄럽고 의미 있는 primitive Jacobian을 생성한다. deformation field의 추가 용량은 이미지 품질을 높이지 못하고, force application과 inertia estimation에 필요한 미분값을 오염시킬 수 있다.

  • Decoder parameterization: LaGSplat은 q를 조건으로 하는 (n+d)D volume에 primitives를 배치하므로 각 μ_i가 q에 대해 affine이고 Jacobian이 constant인 반면, 대안 모델은 nonlinear한 q-dependent warp를 학습한다.deformation family는 임의의 곡선을 따라갈 수 있고 motion 내내 primitives를 object parts에 유지할 가능성이 있지만, straight latent-volume trajectory는 그렇지 않다.
  • Reconstruction: 동일한 reconstruction budget에서 deformation field는 L1 (0.0157 vs 0.0142)과 SSIM loss (0.076 vs 0.063) 모두에서 LaGSplat보다 뒤처지며, parameter 수는 9.8M versus 195k이다.두 decoder는 15,000 primitives, 동일한 losses, regularizers, optimizer, schedule을 사용하고, parameterization만 다르다.
  • The Jacobian field: Figure 11은 LaGSplat의 primitive velocity가 더 매끄러움을 보여주며, median neighbor difference는 deformation field의 6.9 px/s에 비해 1.4 px/s이다.deformation field는 static background와 rocker region 위의 primitives도 이동시키므로, 시각적으로 유사한 rendering에도 불구하고 덜 의미 있는 Jacobian을 나타낸다.
  • Why the Jacobian matters: frame reconstruction은 material tracking이나 primitive velocity를 supervise하지 않으므로, decoder는 clip에 맞으면서도 force pullback과 inertia estimation에 필요한 의미 없는 ∂μ_i/∂q를 산출할 수 있다.LaGSplat에서는 동일한 primitives가 모든 frame을 설명해야 하고, 인접한 state 사이의 작은 shift가 해당 primitives가 덮는 matter를 보존하므로 material tracking이 emergent하게 나타난다.

C 구현 세부사항

LaGSplat은 시스템 전반에서 공유 encoder, decoder, LNN 아키텍처를 사용하며, 시스템별로 latent dimension, resolution, primitive count, run length를 조정한다. 학습은 Gaussian lifecycle management와 latent whitening 절차를 포함한 두 단계의 Adam 최적화로 수행된다.

  • 구현 설정: 모든 시스템에 동일한 아키텍처를 사용하며, latent dimension, image resolution, primitive count, run length는 시스템별로 설정된다.
  • 아키텍처: 연속성을 보존하는 convolutional autoencoder는 12×12 convolution 3개, 두 개의 채널, stride 2, ELU activation, 4×4 average pooling, 그리고 q로의 linear map을 사용한다.약 2 000개의 parameter를 가지며, decoder는 104에서 2 × 10^5개의 parameter를 갖는 1 000–15 000개의 Gaussian을 포함하고, LNN은 15 000개 미만의 parameter를 갖는다.
  • 학습: 두 학습 단계 모두 Adam을 사용하며, Stage 1에서는 약 5 × 10^4 optimizer steps 동안 encoder와 decoder를 공동 학습한다.epoch 수는 clip 크기에 따라 달라진다.
  • 학습: Stage 1에서는 0.08 L1 + 0.02 (1 − SSIM) + 0.01 anisotropy를 최적화하는 동시에, peak opacity가 0.02 미만으로 떨어진 Gaussian을 주기적으로 재활용한다.재활용 과정에서는 full (n+d)-dimensional covariance의 principal eigenvector를 따라 과부하된 Gaussian을 분할하거나 high-error region으로 teleport한다. latent whitening에는 encoded covariance의 eigendecomposition을 사용한다.
Loading 2608.16324v1…