Hi LingBot-VA team, your causal video-action world model is very close to the embodied side of DreamX-World(https://github.com/AMAP-ML/DreamX-World) . DreamX-World currently focuses on camera/event-controlled video worlds rather than robot action spaces, but the long-rollout and memory problems look similar. If useful, we can discuss a small cross-test where action-conditioned robot videos and camera/event-conditioned world videos are evaluated for temporal drift, controllability, and recovery after revisits.
Hi LingBot-VA team, your causal video-action world model is very close to the embodied side of DreamX-World(https://github.com/AMAP-ML/DreamX-World) . DreamX-World currently focuses on camera/event-controlled video worlds rather than robot action spaces, but the long-rollout and memory problems look similar. If useful, we can discuss a small cross-test where action-conditioned robot videos and camera/event-conditioned world videos are evaluated for temporal drift, controllability, and recovery after revisits.