Miao Liu, Haoxiang Huang, Shiji Zhou, Xiang Liu, Sen Cui, Yueqing Sun, Qi Gu, Zhekai Wang, Zhikang Chen
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Object-centric world models forecast future videos by evolving a set of entity slots, but the variables receiving dynamics supervision are often unconstrained visual features. We introduce MOSH-WM, a mask-grounded soft-Hamiltonian world model that makes its position-like state explicitly depend on slot-owned image support. A frozen video-slot encoder produces slots and masks; spatial moments of mask-owned support form a canonical state Q, temporal differences form P , and a learned energy supplies a soft directional bias to a bounded learned increment. Decoder-relevant appearance and identity are stored separately in a causal visual context. A gated composer and bounded residual then combine this context with the propagated phase state to reconstruct decoder-compatible slots. On OBJ3D, given six observed frames and evaluated over the following 30 frames, MOSH-WM reduces LPIPS by 25.0% and spatial MSE by 33.7% relative to the strongest object-centric baseline. On CLEVRER, given six observed frames and evaluated over the following ten frames, the corresponding reductions are 14.5% and 18.7%. Horizon-resolved visual and object-state measurements show that the complete model accumulates error more slowly throughout the 30-frame closed-loop rollout. Project page: https://github.com/moshwm-anon/-moshwm-anon.github.io.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2608.22750")
get_code_for_paper("2608.22750")
have("2608.22750")
Connect an agent — have() is free.