1 Introduction
Recent advances in video generation have improved visual quality, motion realism, and temporal coherence, and have supported the development of video world models built on autoregressive generation [30, 22, 50, 12, 67, 26]. While most text- or image-conditioned systems generate a finite clip from conditions specified before inference, video world models support an ongoing interactive process. They continually predict subsequent observations from observed or generated visual states, allowing users to change camera positions, viewing directions, and control actions throughout a rollout [54, 1]. Video generation thus extends from producing a recording of an event to supporting exploration of an evolving visual world, with applications in interactive content creation, virtual production, game generation, and embodied-agent simulation [33, 75, 14, 61, 47]. In particular, as large pretrained world models acquire stronger visual and motion priors, a key question is how to flexibly extend their controllability without retraining, enabling the same general world model to accept diverse forms of visual control and world information and make fuller use of its pretrained capabilities.
Supporting such diverse controls requires a world model to integrate complementary visual evidence from observations, geometry, and generated history, which differ in representation, spatial coverage, and temporal relevance. As viewpoints and scene states change, generation must remain synchronised with the dynamics specified by the visual evidence while maintaining correct spatial placement and occlusion relationships in the target view. Large viewpoint changes require using the model’s generative prior to complete unobserved scene and object surfaces, while long-horizon revisits require recovering earlier appearance and spatial layout. The challenge is therefore to make relevant evidence available and use it at the appropriate locations and times throughout a rollout.
Existing approaches have explored camera conditioning, source-video rerendering, geometric control, and long-term memory [78, 26, 36, 73, 3, 67, 66]. Many rely on control-specific pathways or representations, so supporting new evidence types or combinations often requires redesign or additional training. This motivates a shared interface through which the same pretrained world model can use heterogeneous current and historical visual evidence without learning a separate pathway for each control, which is the focus of our paper.
Our key observation is that causal video world models equipped with a clean-state cache already have a shared entry point for visual information: their native self-attention. During autoregressive generation, the model reads the initial observation and recently finalised outputs as clean visual states, while camera and temporal encodings specify their viewpoints and temporal positions. This suggests that external control information can be converted into the same representation: clean visual states associated with camera poses, event times, and valid spatial regions, made accessible through the existing attention layers. Instead of adapting the model to each new control, we express different controls in a visual representation that the pretrained model already processes. Under this view, extending world-model control becomes a visual evidence construction and orchestration problem, rather than a model adaptation problem.
Based on this idea, we introduce World in World (WiW), a training-free visual-evidence interface for controlling frozen causal video world models (Figure 1). WiW converts various visual conditions, such as source observations, target-view projections, rendered geometry, and generated history, into clean visual states annotated with camera, temporal, and spatial-validity information. These states can be directly accessed through the model’s native self-attention, allowing different sources of evidence to guide generation where and when they are reliable. Our method is entirely training-free, requiring neither updates to the pretrained world model nor learned control-specific modules. Under this unified interface, flexible world-model control reduces to three complementary problems: constructing useful visual evidence, localising the relevant evidence, and regulating its influence on generation.
To provide the frozen world model with the complementary information required for controllable long-horizon rollouts, we first construct multiple forms of visual evidence for different control requirements. Our core idea is to let each evidence source contribute the information it can provide most reliably, while expressing all of them through the same clean-state interface that the pretrained model already understands. Concretely, source-video observations provide appearance and content references from the recorded event. Since these observations do not explicitly determine where their content should appear under the target camera, target-view scene evidence uses depth-based projection to place observed appearance in the requested view and provide spatial guidance within geometrically supported regions. When large camera motions expose subject surfaces that are absent from the source observations, rendered geometry evidence supplies target-view shape and appearance proposals for the corresponding event state, helping the pretrained model complete newly visible subject surfaces. Finally, because a finite rolling cache eventually removes access to earlier generated states, we maintain a rollout-wide history and retrieve relevant and diverse archived states as generated evidence when the camera returns. By assigning different control requirements to complementary evidence sources, WiW provides the frozen backbone with appearance references, spatial guidance, completion cues, and long-range visual context through a single shared visual interface.
To ensure that the model reads the right evidence and uses it with appropriate strength, we further regulate how visual evidence participates in native self-attention. Our core idea is to separately control where to read and how strongly to use the available evidence, addressing errors in evidence localisation and differences in evidence reliability. For where to read, appearance-based attention can become ambiguous under large viewpoint changes, dynamic motion, and repeated textures. We therefore introduce correspondence-guided attention routing (CGAR), which uses persistent point identities and camera geometry to route current queries towards geometrically corresponding source-video tokens when valid matches are available. For how strongly to use, different evidence sources can vary in reliability across spatial regions and denoising stages. We introduce evidence-wise attention CFG (EWA), which compares native and evidence-conditioned attention responses to independently amplify compatible and complementary information while suppressing excessive guidance. Both mechanisms operate directly within the model’s native self-attention, and EWA reuses responses from the same denoising forward pass, introducing no additional network function evaluations (NFE) for guidance.
Through this formulation, WiW provides a unified, training-free framework for extending the control capabilities of frozen world models through visual evidence, without introducing control-specific training or adaptation. Notably, beyond exploring worlds freely generated by a world model, WiW enables exploration within the world of a given video: users can navigate the recorded dynamic world along new camera trajectories while preserving the appearance and temporal progression of the original event. We instantiate WiW on the publicly released causal-fast checkpoint of LingBot-World 2.0 [12], with all pretrained parameters frozen, and evaluate camera-controlled video rerendering on DAVIS and OpenVid-1M [46, 43] across diverse viewpoint changes. Beyond rerendering, WiW supports applications including bullet-time generation, video stabilization, video editing, and K/V sharing between two generation cases produced by the same frozen model, demonstrating the versatility of the proposed interface.
Our main contributions are:
We introduce WiW, a training-free visual-evidence interface for flexibly extending the control capabilities of frozen causal video world models.
We develop complementary visual evidence that enables event-synchronised, spatially aligned, and long-horizon-consistent exploration of dynamic video worlds.
We introduce correspondence-guided attention routing (CGAR) and evidence-wise attention CFG (EWA) to localise and regulate heterogeneous visual evidence within native self-attention.
We demonstrate flexible exploration of given video worlds across diverse camera trajectories and downstream applications using a single frozen world-model backbone.
2 Related work
2.1 World models
A video world model predicts the visual observations an observer would receive while moving through an environment, given an initial observation and a requested camera or action. In static environments, a key requirement is to maintain consistent scene structure and appearance as the viewpoint changes and previously seen regions are revisited. One line of work uses explicit scene representations, such as point clouds, meshes, or Gaussians, to organize observations in a shared coordinate system and render them from the target camera. These representations provide a clear spatial reference, but require completing unobserved regions and maintaining the scene state as exploration continues [30, 71, 70, 37, 56, 74, 7, 24, 49, 2, 20, 11, 35, 55]. Another line of work directly predicts target-view observations with video models, using learned generative priors to handle viewpoint changes and newly visible regions [53, 76, 57, 79, 16]. When the environment changes over time, the model must also maintain consistency between viewpoint changes, scene structure, and event progression. In dynamic environments, methods need to generate observations that match both the target viewpoint and the current state of the recorded or simulated event. Previous work has studied this problem through dynamic scene generation, source-video re-observation, and continuous world modeling [9, 10, 65, 77, 3, 26, 44, 68, 39, 17, 41]. Autoregressive and self-forcing methods further extend scene generation to continuous rollouts, requiring models to maintain consistent appearance, spatial relationships, and event states over long explorations [4, 23, 22, 42, 6, 1, 12]. Pretrained video models have appearance and motion priors that generalize to new scenes, but their native conditioning interfaces are usually determined by the inputs used during training. The visual-history pathway can itself serve as a control interface: Warp-as-History [60] feeds camera-warped observations as pseudo-history, aligns their temporal positions with target frames. We build on a pretrained causal video model and study how to reuse its native self-attention mechanism for reading visual states, allowing it to accept additional control information while keeping its parameters frozen.
2.2 Conditioning mechanisms for video world models
Existing video world models usually use dedicated conditioning mechanisms to introduce different types of control into generation. For camera control, prior methods represent camera trajectories as rays, poses, positional encodings, projected features, or rendered geometric proxies, and use corresponding conditioning modules to guide generation [27, 32, 64, 31, 40, 63, 13, 29]. Source-video rerendering extends control from camera motion to re-observation of dynamic content. These methods often train additional modules for source-video inputs to bring the original event’s appearance and dynamics into the target view [3, 8, 51, 26, 58]. Similarly, the concurrent work Wonder supports image- and video-conditioned world generation by jointly training a rendered control field, a sparse memory, and a distilled causal student to combine multiple conditions [67]. These methods show that dedicated conditioning pathways can support specific control tasks, but different types of control often use different input interfaces and training procedures. Beyond current observations and external controls, continuous exploration also requires access to scene information generated earlier. Since causal models typically read only a limited recent context, prior work uses explicit geometric memory, persistent states, or historical information retrieval to recover scene content beyond the current context and constrain subsequent generation [34, 59, 62, 21, 66, 72, 28, 69]. Memory therefore helps maintain long-term consistency and can also serve as a visual condition alongside current observations and geometry. WiW follows this idea by including historical information in a unified visual conditioning framework, allowing scene evidence from different sources to jointly guide subsequent generation.
3 Method
We propose World in World (WiW), a unified visual-evidence interface for extending the control capabilities of frozen causal video world models. We illustrate the framework through camera-controlled video rerendering: given a source video of a dynamic event and a target camera trajectory , we aim to rerender the event from the requested viewpoints while preserving its appearance and temporal progression. To this end, we convert source observations, geometry, and generated history into clean visual states with camera, temporal, and spatial-validity information. The frozen model can then read this evidence through native self-attention, without additional training or learned control-specific adapters.
Figure 2 summarizes the overall pipeline. We first define the shared evidence interface and explain how visual evidence participates in native self-attention (Section 3.1). Building on this interface, we project source observations into the target view to provide layout references (Section 3.2), use rendered geometry to guide completion of newly exposed subject surfaces (Section 3.3), and retrieve generated history for consistent long-horizon revisits (Section 3.4). Together, these sources provide complementary references for generation. Attention routing and evidence-wise attention CFG then localize relevant evidence and regulate its influence, respectively (Section 3.5).
3.1 Control through Visual Evidence
When generating the current chunk, our causal video backbone reads the initial observation and recently finalized states through native self-attention. We use this existing pathway to convert multiple forms of visual evidence, including source-video observations, target-view scene projections, rendered geometry, and generated history, into the same type of states that the model already reads. These sources differ in representation, valid spatial extent, and temporal coverage. Their shared representation must therefore specify evidence content, camera and temporal information, spatial support, and channel activation across denoising stages.
To represent this information consistently, at denoising step , converter maps the raw evidence of channel and target camera trajectory to:
| (1) |
Here, denotes the visual content of the evidence; records per-frame camera intrinsics, poses, and event-time indices; specifies the spatial support of evidence token ; and determines whether the channel is active at step .
To turn this representation into attention features that the model can read, we feed the clean visual states corresponding to , together with camera and temporal information , into the frozen video backbone with the diffusion timestep set to . A single network forward pass extracts the keys and values from each attention layer, which we cache as clean K/V. During target-chunk generation, queries from the current denoising states read these cached features through self-attention, allowing external visual evidence to guide generation. All evidence types introduced below use this procedure to obtain their attention features.
These K/V features participate in attention through either the native block or temporary auxiliary blocks. The native block retains the backbone’s 18 latent-frame-state slots: six source anchors, eight recent-history states, and four states in the current chunk. Source-video evidence occupies the source-anchor slots, while other evidence is attached as temporary auxiliary attention blocks and removed once the current chunk is finalized. All blocks reuse the frozen backbone’s query/key/value (Q/K/V) projections, while preserving its native text- and camera-conditioning pathways.
When visual evidence enters attention, its temporal relationship to the current generation states needs to be specified. We therefore set the temporal coordinates of rotary position embeddings (RoPE) and apply the corresponding rotations to attention queries and keys, incorporating relative temporal information into attention.
Spatial support weights regulate the contribution of evidence tokens by scaling their unnormalized attention weights by . Channel activation further determines the denoising stages in which each evidence source participates.
3.2 Target-View Scene Evidence
Source-video observations provide appearance references for the recorded event, but they remain in the source-camera view and do not explicitly determine where their content should appear in the target image. Camera motion changes projected positions and occlusion relationships, while ambiguous cross-view appearance matches can cause visible structures to drift. To provide an explicit spatial layout reference, we project source observations into the target view, allowing the model to read observed appearance aligned with the requested camera.
To construct this reference for target frame and temporally aligned source view , we use source RGB and estimate its depth with DepthCrafter [18]. We back-project source pixels into 3D and reproject them from source camera to target camera , where denotes camera intrinsics and is the absolute camera-to-world pose:
| (2) |
Projection uses a shared coordinate system and scale. The outputs , , and are target-view RGB, a binary visibility mask, and target-camera depth, respectively. Following the shared interface, we encode the projected RGB with target-camera and event-time information, and use visibility to determine the evidence’s spatial support.
Since occlusion and local geometric errors can affect the projection, we determine its spatial support from visibility and geometric reliability, then convert it into token-level support weights . These weights allow the model to use reliable target-view layout references while reducing the influence of uncertain regions.
3.3 Rendered Geometry Evidence
Target-view projection reorganizes content already present in the source observations, but its coverage remains limited by source-view visibility. When camera motion exposes subject surfaces not observed in the source video, these regions lack direct shape and appearance references, and their generation may deviate from the subject’s structure or current pose. To provide evidence for completing such regions, we use a renderable subject representation to produce shape and appearance proposals from the target camera at the corresponding event time. Given subject state at event time and target camera , rendering yields:
| (3) |
Here, , , and denote RGB, a geometric support mask, and target-camera depth, respectively. The rendering is encoded through the shared interface, and its support mask is converted into token-level support weights.
For example, when the subject is human, we represent its state as . Here, is an avatar reconstructed with LHM++ [48] from sharp, minimally occluded full-body source-video crops, and contains the per-frame SMPL-X [45] parameters that drive the avatar. We align the avatar and body model to the projected scene’s coordinate frame and depth scale using depth correspondences in the source view. LHM++ provides target-view RGB and support masks for observed and inferred unseen surfaces, while the aligned SMPL-X geometry provides depth for resolving occlusion between subjects. We also compare this depth with the projected scene depth to remove unreliable background projections near subject boundaries.
3.4 Retrieved Generated Evidence
Scene projections and geometry renderings provide observed and reconstruction-based evidence for the current target view. As generation progresses, previously generated appearance and layout also become important references for subsequent frames. The native rolling cache retains only recent states, so states corresponding to an earlier region may have been evicted when the camera returns, leaving their generated content inaccessible to the model. To recover these references, we maintain a rollout-wide history bank and retrieve finalized states relevant to the current view.
The history bank accumulates evidence as the rolling cache is updated. For each evicted finalized state, we archive its per-layer clean K/V with the associated camera pose and temporal index. Since the bank grows throughout generation, we select a bounded number of relevant states to participate in attention for each target chunk. Specifically, we rank historical states by target-view surface coverage and viewing-direction compatibility, then select a diverse top- subset to reduce redundant references. The model reads the retrieved K/V directly through temporary auxiliary attention blocks, recovering historical appearance and layout beyond the rolling cache to maintain scene consistency.
3.5 Attention Routing and Guidance
The construction and retrieval steps above provide complementary evidence for the source event, target-view layout, subject geometry, and generated history. However, making evidence accessible does not ensure that queries locate the relevant positions within it. Viewpoint changes and dynamic motion alter the appearance of a surface, while repeated textures can make distinct locations look similar, introducing ambiguity into appearance-based attention. We therefore first use geometric correspondences to match current generation tokens with source-video tokens and guide evidence reading, then regulate each channel’s additional influence through its attention response.
Correspondence-guided attention routing. To associate queries from current generation tokens that have valid geometric correspondences with matching source-video locations, we introduce correspondence-guided attention routing (CGAR). We track points with persistent identities in the source video and combine depth and camera information to establish token-level correspondences between the current view and the source video. For a geometrically matched query and source-video key , we add the logarithm of correspondence weight to the routing attention score:
| (4) |
Here, is the attention-head dimension, and controls the contribution of the correspondence. We combine the routing response with the other attention-block responses through joint normalization, allowing current generation queries to read geometrically matched source-video evidence.
Evidence-wise attention CFG. The evidence sources constructed above participate in video generation through self-attention during denoising. Different auxiliary evidence channels provide constraints on target-view layout, subject structure, and generated history, and the information they provide also differs in reliability and importance. Their influence on generation therefore needs to be controlled independently. Using all evidence with the same strength makes it difficult to balance adherence to different conditions with the model’s generative prior. Inspired by classifier-free guidance (CFG) [15] and NAG [5], we propose evidence-wise attention CFG (EWA). EWA independently adjusts the guidance strength of each evidence source at the level of attention responses, strengthening useful control signals while preserving the content quality and naturalness of native generation.
Specifically, at the same attention layer and for the same queries, let denote the response obtained by attending only to the native block, and let denote the jointly normalized response obtained by attending to both the native block and evidence . Directly amplifying would also amplify its component along the native generation direction. We therefore strengthen only the complementary direction introduced by the evidence relative to the native response. We first remove the projection of onto , modulate the remaining component by the nonnegative cosine similarity between the two responses, and use an independent guidance strength to control the correction magnitude. The updated response for evidence is:
| (5) |
Here, is the cosine similarity between the two responses. Removing the projection along the native response direction allows EWA to strengthen the complementary information provided by the evidence without repeatedly amplifying the model’s existing response.
We then add the evidence-specific corrections to the joint response of all active attention blocks and bound the output magnitude using the norm of the native response as a reference, preventing excessive guidance from disrupting generation.
EWA operates directly on attention responses within the same denoising forward pass. Its computation reuses existing attention-block outputs and normalization results, without requiring a separate full denoising-network forward pass for each evidence source. It therefore adds no network function evaluations (NFE) for guidance.
4 Experiments
We evaluate World in World through camera-controlled video rerendering, where a source video serves as visual evidence for exploring the recorded world along a target camera trajectory. We further examine long-horizon revisiting and human-motion transfer, and ablate individual components to assess how the same frozen backbone uses complementary visual evidence.
4.1 Experimental Setup
Implementation. We instantiate World in World on the publicly released causal-fast checkpoint of LingBot-World 2.0 [12], keeping all pretrained parameters frozen. Visual evidence is constructed and attached at inference time, without additional training or learned control-specific adapters.
Baselines. Following prior work on camera-controlled video rerendering, we select videos from DAVIS [46] and OpenVid-1M [43] and compare with ReCamMaster, TrajectoryCrafter, WorldForge, InSpatio-World, UniWorld-View, and CameraAnything [3, 73, 52, 26, 78, 36]. Each baseline uses its official implementation, released checkpoint, and recommended configuration. ReCamMaster and TrajectoryCrafter generate 81 and 49 frames, respectively, as required by their official implementations; the remaining methods use their officially recommended frame counts. All methods receive the same source video and continuous target camera path, with camera parameters converted to the convention required by each method. Before evaluation, we align the generated videos across methods in terms of camera angles and frame count. All methods share the same metric preprocessing pipeline.
Metrics. We evaluate image fidelity using PSNR, SSIM, and LPIPS. To assess video quality, temporal consistency, and dynamics, we use the official VBench implementation [25] to report seven dimensions: Aesthetic Quality, Imaging Quality, Temporal Flickering, Motion Smoothness, Subject Consistency, Background Consistency, and Dynamic Degree. We additionally report Overall, computed as the average score of these seven dimensions.
Following InSpatio-World [26], we evaluate camera control using translation error (TransError) and rotation error (RotError). We independently estimate camera trajectories from the generated videos using Depth Anything 3 and ViPE [38, 19]. For each estimator, we compare the estimated trajectories with the prescribed target trajectories. The reported TransError and RotError are obtained by averaging the corresponding errors across the two estimators.
4.2 Quantitative Comparisons
Table 1 reports quantitative results for camera-controlled video rerendering on DAVIS and OpenVid-1M. Following previous state-of-the-art, we evaluate performance in terms of VBench, camera trajectory errors and image quality.
| VBench | Camera errors | Image quality | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Aesth. | Img. | Flick. | Smooth. | Subj. | Bg. | Dyn. | Overall | TransError | RotError (∘) | PSNR | SSIM | LPIPS |
| ReCamMaster | 54.283 | 66.290 | 95.178 | 97.503 | 88.941 | 92.154 | 92.500 | 83.836 | 0.124289 | 7.8626 | 15.1383 | 0.444346 | 0.377962 |
| TrajectoryCrafter | 51.322 | 61.466 | 94.169 | 97.423 | 87.032 | 90.727 | 98.750 | 82.984 | 0.090812 | 5.3254 | 19.9689 | 0.625112 | 0.217875 |
| WorldForge | 49.915 | 62.356 | 94.262 | 96.934 | 84.683 | 90.017 | 98.524 | 82.384 | 0.079961 | 4.5021 | 19.6494 | 0.589250 | 0.218029 |
| InSpatio-World | 53.162 | 66.076 | 93.622 | 96.764 | 87.070 | 91.188 | 96.250 | 83.447 | 0.082399 | 4.4011 | 20.6855 | 0.613398 | 0.183115 |
| UniWorld-View | 54.237 | 67.534 | 93.444 | 96.627 | 88.946 | 91.777 | 97.500 | 84.295 | 0.068705 | 4.1958 | 22.4735 | 0.770554 | 0.121668 |
| CameraAnything | 54.365 | 67.526 | 94.489 | 96.908 | 87.399 | 91.817 | 88.750 | 83.036 | 0.186443 | 3.1868 | 15.1583 | 0.453712 | 0.342180 |
| Ours | 56.556 | 68.004 | 93.711 | 97.749 | 88.948 | 92.627 | 98.750 | 85.192 | 0.068622 | 2.8326 | 23.1511 | 0.787205 | 0.121664 |
4.3 Qualitative Comparisons
Figure 3 compares camera-controlled rerendering under different camera motion. Each example presents the source video and representative outputs under the same target camera path.
4.4 Additional Applications
Our framework expresses different control requirements through a shared visual-evidence interface. Figure 4 extends this idea beyond rerendering to control over event timing, changes to appearance and motion, and the sharing of generated visual information. In the K/V-sharing example, Model A and Model B denote two instances of the same frozen model, each corresponding to a different generation case. Communication between the instances is implemented by passing cached K/V from one instance to the other as visual evidence to guide generation. Across these applications, different sources of evidence guide generation through the same native attention mechanism, with camera and time information specifying how that evidence relates to the requested output. This allows us to extend the model’s control capabilities while keeping the backbone frozen, without task-specific training.
4.5 Ablation Studies
We examine how the choice and use of visual evidence contribute to controllable generation. Table 2 shows that the full method achieves the lowest camera errors and the best or tied-best results on all reported VBench dimensions. Accurate camera control requires a clear spatial reference for the observed content. Figure 5 examines target-view warping and source-camera Plücker conditioning from this perspective. Removing warping causes the largest degradation across all reported metrics, with rotation and translation errors rising to approximately and their full-method values. Removing source-camera conditioning also increases camera errors. These results support retaining both the target-view layout and the source-view camera information. Figure 6 examines how evidence is used and what happens when current observations provide insufficient support. CGAR and EWA help preserve subject appearance and background structure. Rendered geometry and historical retrieval address two gaps in the available evidence: newly exposed subject surfaces and earlier scene states beyond the rolling cache.
| VBench | Camera errors | ||||||
|---|---|---|---|---|---|---|---|
| Variant | Subj. | Bg. | Smooth. | Aesth. | Flick. | RotError (∘) | TransError |
| w/o EWA and CGAR | 88.076 | 91.747 | 96.054 | 50.688 | 92.857 | 2.0346 | 0.056937 |
| w/o EWA | 88.056 | 91.706 | 96.055 | 50.720 | 92.867 | 1.9944 | 0.056608 |
| w/o CGAR | 88.055 | 91.728 | 96.072 | 50.713 | 92.891 | 1.7828 | 0.053913 |
| w/o Target-View Warp | 83.378 | 89.540 | 94.351 | 50.240 | 90.792 | 6.1158 | 0.581847 |
| w/o Plücker | 88.066 | 91.692 | 96.073 | 50.669 | 92.881 | 1.8811 | 0.056887 |
| Full method | 88.076 | 91.749 | 96.074 | 50.720 | 92.902 | 1.7823 | 0.053291 |
5 Conclusion
We presented World in World (WiW), a training-free visual-evidence interface for extending the controllability of frozen causal video world models. WiW represents source observations, target-view projections, rendered geometry, and generated history as camera- and time-labelled clean visual states that the backbone reads through native self-attention. Correspondence-guided attention routing localizes relevant source-video evidence, while evidence-wise attention CFG regulates auxiliary contributions using responses from the same denoising forward pass, without additional denoising-network evaluations for guidance. On the reported DAVIS and OpenVid-1M evaluations, WiW achieves the highest average score across seven VBench dimensions and the lowest camera trajectory errors among the compared methods. Qualitative results further illustrate several downstream applications through the same interface. These results demonstrate that constructing, selecting, and regulating visual evidence can extend the capabilities of a pretrained world model, enabling exploration of the dynamic world depicted in a given video.
References
- [1] AlayaWorld Team, K. Zhang, C. Li, Y. Zhan, Y. Ge, Y. Yin, J. Tan, K. He, L. Fan, M. Zhai, R. Liu, X. Xu, X. Chu, Z. Li, Z. Lin, Z. Wang, Z. Meng, and Z. Gao (2026) AlayaWorld: interactive long-horizon world modeling—full technical report. arXiv preprint arXiv:2607.18367. External Links: 2607.18367 Cited by: §1, §2.1.
- [2] S. Bahmani, T. Shen, J. Ren, J. Huang, Y. Jiang, H. Turki, A. Tagliasacchi, D. B. Lindell, Z. Gojcic, S. Fidler, H. Ling, J. Gao, and X. Ren (2025) Lyra: generative 3D scene reconstruction via video diffusion model self-distillation. arXiv preprint arXiv:2509.19296. External Links: 2509.19296 Cited by: §2.1.
- [3] J. Bai, M. Xia, X. Fu, X. Wang, L. Mu, J. Cao, Z. Liu, H. Hu, X. Bai, P. Wan, and D. Zhang (2025) ReCamMaster: camera-controlled generative rendering from a single video. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 14834–14844. Cited by: §1, §2.1, §2.2, §4.1.
- [4] B. Chen, D. M. Monsó, Y. Du, M. Simchowitz, R. Tedrake, and V. Sitzmann (2024) Diffusion forcing: next-token prediction meets full-sequence diffusion. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 37. External Links: 2407.01392 Cited by: §2.1.
- [5] D. Chen, H. Bandyopadhyay, K. Zou, and Y. Song (2026) Normalized attention guidance: universal negative guidance for diffusion models. Advances in Neural Information Processing Systems 38, pp. 137412–137446. Cited by: §3.5.
- [6] J. Chen, H. Zhu, X. He, Y. Wang, J. Zhou, W. Chang, Y. Zhou, Z. Li, Z. Fu, J. Pang, and T. He (2025) DeepVerse: 4D autoregressive video generation as a world model. arXiv preprint arXiv:2506.01103. External Links: 2506.01103 Cited by: §2.1.
- [7] L. Chen, Z. Zhou, M. Zhao, Y. Wang, G. Zhang, W. Huang, H. Sun, J. Wen, and C. Li (2025) FlexWorld: progressively expanding 3D scenes for flexiable-view synthesis. arXiv preprint arXiv:2503.13265. External Links: 2503.13265 Cited by: §2.1.
- [8] Y. Chen, Z. Ye, Z. Fang, X. Chen, X. Zhang, J. Liu, N. Wang, G. Zhang, and H. Liu (2025) PostCam: camera-controllable novel-view video generation with query-shared cross-attention. arXiv preprint arXiv:2511.17185. External Links: 2511.17185 Cited by: §2.2.
- [9] Z. Chen, T. Liu, L. Zhuo, J. Ren, Zeng Tao, H. Zhu, F. Hong, L. Pan, and Z. Liu (2025) 4DNeX: feed-forward 4D generative modeling made easy. arXiv preprint arXiv:2508.13154. External Links: 2508.13154 Cited by: §2.1.
- [10] Y. Dai, F. Jiang, C. Wang, M. Xu, and Y. Qi (2025) FantasyWorld: geometry-consistent world modeling via unified video and 3D prediction. arXiv preprint arXiv:2509.21657. External Links: 2509.21657 Cited by: §2.1.
- [11] S. Gao, Z. Wang, Q. Cao, D. Yu, C. Wang, and J. Bian (2026) PixWorld: unifying 3D scene generation and reconstruction in pixel space. arXiv preprint arXiv:2607.05373. External Links: 2607.05373 Cited by: §2.1.
- [12] Z. Gao, Q. Wang, J. Zhu, J. Chen, Z. Liu, Q. Bai, J. Wang, Y. Yuan, H. Wang, Y. Lu, K. L. Cheng, H. Zhang, J. Gao, T. Feng, Y. Liu, Y. Yao, Y. Xu, X. Zhu, Y. Shen, and H. Ouyang (2026) Infinite worlds with versatile interactions. arXiv preprint arXiv:2607.07534. External Links: 2607.07534 Cited by: §1, §1, §2.1, §4.1.
- [13] Z. Gu, R. Yan, J. Lu, P. Li, Z. Dou, C. Si, Z. Dong, Q. Liu, C. Lin, Z. Liu, W. Wang, and Y. Liu (2025) Diffusion as shader: 3D-aware video diffusion for versatile video generation control. arXiv preprint arXiv:2501.03847. External Links: 2501.03847 Cited by: §2.2.
- [14] X. He, C. Peng, Z. Liu, B. Wang, Y. Zhang, Q. Cui, F. Kang, B. Jiang, M. An, Y. Ren, B. Xu, H. Guo, K. Gong, S. Wu, W. Li, X. Song, Y. Liu, Y. Li, and Y. Zhou (2025) Matrix-Game 2.0: an open-source real-time and streaming interactive world model. arXiv preprint arXiv:2508.13009. External Links: 2508.13009, Document, Link Cited by: §1.
- [15] J. Ho and T. Salimans (2022) Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598. External Links: 2207.12598, Link Cited by: §3.5.
- [16] L. Höllein and M. Nießner (2026) World reconstruction from inconsistent views. arXiv preprint arXiv:2603.16736. External Links: 2603.16736 Cited by: §2.1.
- [17] T. Hu, H. Peng, X. Liu, and Y. Ma (2025) EX-4D: extreme viewpoint 4D video synthesis via depth watertight mesh. arXiv preprint arXiv:2506.05554. External Links: 2506.05554 Cited by: §2.1.
- [18] W. Hu, X. Gao, X. Li, S. Zhao, X. Cun, Y. Zhang, L. Quan, and Y. Shan (2025) DepthCrafter: generating consistent long depth sequences for open-world videos. In CVPR, Cited by: §3.2.
- [19] J. Huang, Q. Zhou, H. Rabeti, A. Korovko, H. Ling, X. Ren, T. Shen, J. Gao, D. Slepichev, C. Lin, J. Ren, K. Xie, J. Biswas, L. Leal-Taixé, and S. Fidler (2025) ViPE: video pose engine for 3D geometric perception. arXiv preprint arXiv:2508.10934. External Links: 2508.10934 Cited by: §4.1.
- [20] J. Huang, Y. Yang, B. Yang, L. Ma, Y. Ma, and Y. Liao (2026) Gen3R: 3D scene generation meets feed-forward reconstruction. arXiv preprint arXiv:2601.04090. External Links: 2601.04090 Cited by: §2.1.
- [21] J. Huang, X. Hu, B. Han, S. Shi, Z. Tian, T. He, and L. Jiang (2025) Memory forcing: spatio-temporal memory for consistent scene generation on minecraft. arXiv preprint arXiv:2510.03198. External Links: 2510.03198 Cited by: §2.2.
- [22] S. Huang, J. Wu, Q. Zhou, S. Miao, and M. Long (2025) Vid2World: crafting video diffusion models to interactive world models. arXiv preprint arXiv:2505.14357. External Links: 2505.14357 Cited by: §1, §2.1.
- [23] X. Huang, Z. Li, G. He, M. Zhou, and E. Shechtman (2025) Self forcing: bridging the train-test gap in autoregressive video diffusion. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 38. External Links: 2506.08009 Cited by: §2.1.
- [24] Y. Huang, W. Chen, W. Zheng, X. Tao, P. Wan, J. Zhou, and J. Lu (2025) Terra: explorable native 3D world model with point latents. arXiv preprint arXiv:2510.14977. External Links: 2510.14977 Cited by: §2.1.
- [25] Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al. (2023) VBench: comprehensive benchmark suite for video generative models. arXiv preprint arXiv:2311.17982. External Links: 2311.17982 Cited by: §4.1.
- [26] InSpatio Team, D. Shen, G. Zhang, H. Liu, H. Ji, H. Bao, H. Zhai, J. Liu, J. Guo, N. Wang, S. Pan, W. Pan, W. Xie, X. Liu, X. Xiang, X. Zhang, X. Chen, Y. Wang, Y. Chen, Z. Fan, Z. Le, Z. Ye, and Z. Zhao (2026) INSPATIO-WORLD: a real-time 4D world simulator via spatiotemporal autoregressive modeling. arXiv preprint arXiv:2604.07209. External Links: 2604.07209 Cited by: §1, §1, §2.1, §2.2, §4.1, §4.1.
- [27] W. Jang, S. Liu, S. Sanyal, J. C. Perez, K. W. Ng, S. Agrawal, J. Perez-Rua, Y. Douratsos, and T. Xiang (2026) Rays as pixels: learning a joint distribution of videos and camera trajectories. In International Conference on Machine Learning (ICML), External Links: 2604.09429 Cited by: §2.2.
- [28] M. Joo, D. Park, T. Lee, K. Lee, and H. J. Kim (2026) Retrieve what’s missing: coverage-maximizing retrieval for consistent long video generation. arXiv preprint arXiv:2606.02479. External Links: 2606.02479 Cited by: §2.2.
- [29] G. Kim, J. Han, and S. Cho (2025) VideoFrom3D: 3D scene video generation via complementary image and video diffusion models. arXiv preprint arXiv:2509.17985. External Links: 2509.17985 Cited by: §2.2.
- [30] L. Kong, Y. Yang, J. Mei, Y. Liu, A. Liang, D. Zhu, D. Lu, W. Yin, X. Hu, M. Jia, J. Deng, K. Zhang, Y. Wu, T. Yan, S. Gao, S. Wang, L. Li, L. Pan, Y. Liu, J. Zhu, W. T. Ooi, S. C. H. Hoi, and Z. Liu (2025) 3D and 4D world modeling: a survey. arXiv preprint arXiv:2509.07996. External Links: 2509.07996 Cited by: §1, §2.1.
- [31] J. Lee, J. Jung, J. Han, T. Narihira, K. Fukuda, J. Seo, S. Hong, Y. Mitsufuji, and S. Kim (2025) 3D scene prompting for scene-consistent camera-controllable video generation. arXiv preprint arXiv:2510.14945. External Links: 2510.14945 Cited by: §2.2.
- [32] C. Li, Y. Yang, J. Shao, H. Zhou, K. Schwarz, and Y. Liao (2026) ReRoPE: repurposing RoPE for relative camera control. arXiv preprint arXiv:2602.08068. External Links: 2602.08068 Cited by: §2.2.
- [33] J. Li, J. Tang, Z. Xu, L. Wu, Y. Zhou, S. Shao, T. Yu, Z. Cao, and Q. Lu (2025) Hunyuan-GameCraft: high-dynamic interactive game video generation with hybrid history condition. arXiv preprint arXiv:2506.17201. External Links: 2506.17201, Document, Link Cited by: §1.
- [34] R. Li, P. Torr, A. Vedaldi, and T. Jakab (2025) VMem: consistent interactive video scene generation with surfel-indexed view memory. In Proceedings of the IEEE/CVF International Conference on Computer Vision, External Links: 2506.18903 Cited by: §2.2.
- [35] X. Li, T. Wang, Z. Gu, S. Zhang, C. Guo, and L. Cao (2025) FlashWorld: high-quality 3D scene generation within seconds. arXiv preprint arXiv:2510.13678. External Links: 2510.13678 Cited by: §2.1.
- [36] Y. Li, Y. Zeng, K. L. Cheng, J. Zhu, H. Wang, W. Wang, Y. Meng, H. Ouyang, Q. Wang, Y. Yu, Z. Wang, Y. Zhang, Y. Shen, and D. Lin (2026) CameraAnything: refilming videos with arbitrary camera control. arXiv preprint arXiv:2607.24591. External Links: 2607.24591, Link Cited by: §1, §4.1.
- [37] H. Liang, J. Cao, V. Goel, G. Qian, S. Korolev, D. Terzopoulos, K. N. Plataniotis, S. Tulyakov, and J. Ren (2024) Wonderland: navigating 3D scenes from a single image. arXiv preprint arXiv:2412.12091. External Links: 2412.12091 Cited by: §2.1.
- [38] H. Lin, S. Chen, J. Liew, D. Y. Chen, Z. Li, G. Shi, J. Feng, and B. Kang (2025) Depth Anything 3: recovering the visual space from any views. arXiv preprint arXiv:2511.10647. External Links: 2511.10647 Cited by: §4.1.
- [39] K. H. Lin, Z. Liu, P. Salamanca, Y. Kant, R. Burgert, Y. Xu, K. Namekata, Y. Zhao, B. Zhou, M. Goldblum, P. Debevec, and N. Yu (2026) Vista4D: video reshooting with 4D point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, External Links: 2604.21915 Cited by: §2.1.
- [40] X. Liu, D. Ji, L. Liu, L. Zhu, X. Chen, Q. Xu, P. Shu, H. Yu, J. Jiang, F. Gao, and S. Ma (2026) CamGeo: sparse camera-conditioned image-to-video generation with 3D geometry priors. In Proceedings of the International Conference on Machine Learning, External Links: 2605.30895 Cited by: §2.2.
- [41] D. Lu, A. Liang, T. Huang, X. Fu, Y. Zhao, B. Ma, L. Pan, W. Yin, L. Kong, W. T. Ooi, and Z. Liu (2025) See4D: pose-free 4D generation via auto-regressive video inpainting. arXiv preprint arXiv:2510.26796. Note: Accepted to Eurographics 2026 External Links: 2510.26796 Cited by: §2.1.
- [42] X. Mao, S. Lin, Z. Li, C. Li, W. Peng, T. He, J. Pang, M. Chi, Y. Qiao, and K. Zhang (2025) Yume: an interactive world generation model. arXiv preprint arXiv:2507.17744. External Links: 2507.17744 Cited by: §2.1.
- [43] K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y. Tai (2024) OpenVid-1M: a large-scale high-quality dataset for text-to-video generation. arXiv preprint arXiv:2407.02371. External Links: 2407.02371 Cited by: §1, §4.1.
- [44] A. Paliwal, A. Iyer, S. Yadav, M. A. Afridi, and M. Harikumar (2026) Reshoot-Anything: a self-supervised model for in-the-wild video reshooting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, External Links: 2604.21776 Cited by: §2.1.
- [45] G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black (2019) Expressive body capture: 3D hands, face, and body from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10975–10985. Cited by: §3.3.
- [46] J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbeláez, A. Sorkine-Hornung, and L. Van Gool (2017) The 2017 DAVIS challenge on video object segmentation. arXiv preprint arXiv:1704.00675. External Links: 1704.00675 Cited by: §1, §4.1.
- [47] R. Qian, Z. Wang, J. Zhang, K. Zou, W. Yu, J. Li, Z. Liu, Y. Li, F. Kang, K. Huang, M. An, H. Zhang, B. Jiang, J. Wang, H. Sun, Y. Liu, and Y. Li (2026) Matrix-Game 3.5: enhancing real-time streaming interactive world models with patch memory. arXiv preprint arXiv:2608.29910. External Links: 2608.29910, Document, Link Cited by: §1.
- [48] L. Qiu, P. Li, H. Li, Q. Zuo, X. Gu, Y. Dong, W. Yuan, R. Peng, S. Zhu, X. Han, G. Chen, and Z. Dong (2025) LHM++: an efficient large human reconstruction model for pose-free images to 3D. arXiv preprint arXiv:2506.13766. External Links: 2506.13766, Link Cited by: §3.3.
- [49] X. Ren, T. Shen, J. Huang, H. Ling, Y. Lu, M. Nimier-David, T. Müller, A. Keller, S. Fidler, and J. Gao (2025) GEN3C: 3D-informed world-consistent video generation with precise camera control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, External Links: 2503.03751 Cited by: §2.1.
- [50] Robbyant Team, Z. Gao, Q. Wang, Y. Zeng, J. Zhu, K. L. Cheng, Y. Li, H. Wang, Y. Xu, S. Ma, Y. Chen, J. Liu, Y. Cheng, Y. Yao, J. Zhu, Y. Meng, K. Zheng, Q. Bai, J. Chen, Z. Shen, Y. Yu, X. Zhu, Y. Shen, and H. Ouyang (2026) Advancing open-source world models. arXiv preprint arXiv:2601.20540. External Links: 2601.20540, Document, Link Cited by: §1.
- [51] J. Seo, J. Han, J. Jung, S. Jin, J. Lee, T. Narihira, K. Fukuda, T. Shibuya, D. Ahn, S. Hu, S. Kim, and Y. Mitsufuji (2025) Vid-CamEdit: video camera trajectory editing with generative rendering from estimated geometry. arXiv preprint arXiv:2506.13697. External Links: 2506.13697 Cited by: §2.2.
- [52] C. Song, Y. Yang, T. Zhao, R. Li, and C. Zhang (2026) Taming video models for 3D and 4D generation via zero-shot camera control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 40352–40363. Cited by: §4.1.
- [53] W. Sun, S. Chen, F. Liu, Z. Chen, Y. Duan, J. Zhang, and Y. Wang (2024) DimensionX: create any 3D and 4D scenes from a single image with controllable video diffusion. arXiv preprint arXiv:2411.04928. External Links: 2411.04928 Cited by: §2.1.
- [54] W. Sun, H. Zhang, H. Wang, J. Wu, Z. Wang, Z. Wang, Y. Wang, J. Zhang, T. Wang, and C. Guo (2025) WorldPlay: towards long-term geometric consistency for real-time interactive world modeling. arXiv preprint arXiv:2512.14614. External Links: 2512.14614 Cited by: §1.
- [55] Team HY-World, C. Cao, X. Zuo, Z. Wang, Y. Zhang, J. Wu, Z. Liu, Y. Gong, Y. Liu, B. Yuan, C. Zhang, C. Li, D. Guo, F. Yang, H. Zhang, H. Cao, J. Zhu, J. Lin, J. Xiao, J. Zhang, J. Yu, L. Wang, L. Wang, L. Wang, Linus, M. Chen, P. He, P. Zhao, Q. Chen, R. Chen, R. Shao, S. Liu, W. Qin, X. Niu, X. Yuan, Y. Sun, Y. Tang, Y. Sun, Y. Lian, Y. Tan, Y. Liu, Y. Yin, Z. Min, T. Wang, and C. Guo (2026) HY-World 2.0: a multi-modal world model for reconstructing, generating, and simulating 3D worlds. arXiv preprint arXiv:2604.14268. External Links: 2604.14268 Cited by: §2.1.
- [56] H. Wang, Y. Liu, Z. Liu, W. Wang, Z. Dong, and B. Yang (2024) VistaDream: sampling multiview consistent images for single-view scene reconstruction. arXiv preprint arXiv:2410.16892. External Links: 2410.16892 Cited by: §2.1.
- [57] H. Wang, F. Liu, J. Chi, and Y. Duan (2025) VideoScene: distilling video diffusion model to generate 3D scenes in one step. arXiv preprint arXiv:2504.01956. External Links: 2504.01956 Cited by: §2.1.
- [58] H. Wang, Y. Chen, H. Huang, C. Zhang, and X. Li (2026) Directing the world: fast autoregressive video generation with compositional human-camera control. arXiv preprint arXiv:2606.27964. External Links: 2606.27964 Cited by: §2.2.
- [59] J. Wang, L. Ye, T. Lu, J. Xiao, J. Zhang, Y. Guo, X. Liu, R. Chellappa, C. Peng, A. Yuille, and J. Chen (2025) EvoWorld: evolving panoramic world generation with explicit 3D memory. arXiv preprint arXiv:2510.01183. External Links: 2510.01183 Cited by: §2.2.
- [60] Y. Wang and T. He (2026) Warp-as-History: generalizable camera-controlled video generation from one training video. arXiv preprint arXiv:2605.15182. External Links: 2605.15182 Cited by: §2.1.
- [61] Z. Wang, Z. Liu, J. Li, K. Huang, B. Xu, F. Kang, M. An, P. Wang, B. Jiang, Y. Wei, Y. Xietian, J. Pei, L. Hu, B. Jiang, H. Xue, Z. Wang, H. Sun, W. Li, W. Ouyang, X. He, Y. Liu, Y. Li, and Y. Zhou (2026) Matrix-Game 3.0: real-time and streaming interactive world model with long-horizon memory. arXiv preprint arXiv:2604.08995. External Links: 2604.08995, Document, Link Cited by: §1.
- [62] Z. Wei, X. Guo, X. Li, X. Xiang, M. Wei, Y. Zhu, Q. Wang, X. Wang, P. Wan, X. Hou, and Q. Fan (2026) Geometry-aware implicit memory for video world models. arXiv preprint arXiv:2606.02436. External Links: 2606.02436 Cited by: §2.2.
- [63] H. Wu, D. Wu, T. He, J. Guo, Y. Ye, Y. Duan, and J. Bian (2025) Geometry forcing: marrying video diffusion and 3D representation for consistent world modeling. arXiv preprint arXiv:2507.07982. External Links: 2507.07982 Cited by: §2.2.
- [64] C. Xiang, J. Liu, J. Zhang, X. Yang, Z. Fang, S. Wang, Z. Wang, Y. Zou, H. Su, and J. Zhu (2026) Geometry-aware rotary position embedding for consistent video world model. arXiv preprint arXiv:2602.07854. External Links: 2602.07854 Cited by: §2.2.
- [65] X. Xiang, Z. Duan, Y. Chen, Z. Wei, G. Zhang, Z. Gu, Z. Gao, H. Huang, C. Zhang, Q. Fan, and X. Li (2026) VideoWeave: unlocking geometric consistency in video generation via joint geometry-video modeling. arXiv preprint arXiv:2606.14162. External Links: 2606.14162 Cited by: §2.1.
- [66] Z. Xiao, Y. Lan, Y. Zhou, W. Ouyang, S. Yang, Y. Zeng, and X. Pan (2025) WorldMem: long-term consistent world simulation with memory. arXiv preprint arXiv:2504.12369. External Links: 2504.12369 Cited by: §1, §2.2.
- [67] J. Xu, H. Jiang, Z. Shu, K. Sunkavalli, V. M. Patel, and Y. Mei (2026) Wonder: video world model done better. arXiv preprint arXiv:2607.26037. External Links: 2607.26037 Cited by: §1, §1, §2.2.
- [68] Y. Xu, J. Shi, Z. Wang, W. Song, F. Shao, C. Liang, J. Xiao, and L. Chen (2026) RealCam: real-time novel-view video generation with interactive camera control. arXiv preprint arXiv:2605.06051. External Links: 2605.06051 Cited by: §2.1.
- [69] J. Yi, M. Kim, P. H. Cho, W. Jang, S. Yun, and S. Kim (2026) WorldKV: efficient world memory with world retrieval and compression. arXiv preprint arXiv:2605.22718. External Links: 2605.22718 Cited by: §2.2.
- [70] H. Yu, H. Duan, C. Herrmann, W. T. Freeman, and J. Wu (2024) WonderWorld: interactive 3D scene generation from a single image. arXiv preprint arXiv:2406.09394. External Links: 2406.09394 Cited by: §2.1.
- [71] H. Yu, H. Duan, J. Hur, K. Sargent, M. Rubinstein, W. T. Freeman, F. Cole, D. Sun, N. Snavely, J. Wu, and C. Herrmann (2024) WonderJourney: going from anywhere to everywhere. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6658–6667. Cited by: §2.1.
- [72] J. Yu, J. Bai, Y. Qin, Q. Liu, X. Wang, P. Wan, D. Zhang, and X. Liu (2025) Context as memory: scene-consistent interactive long video generation with memory retrieval. arXiv preprint arXiv:2506.03141. External Links: 2506.03141 Cited by: §2.2.
- [73] M. Yu, W. Hu, J. Xing, and Y. Shan (2025) TrajectoryCrafter: redirecting camera trajectory for monocular videos via diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), External Links: Document Cited by: §1, §4.1.
- [74] W. Yu, J. Xing, L. Yuan, W. Hu, X. Li, Z. Huang, X. Gao, T. Wong, Y. Shan, and Y. Tian (2024) ViewCrafter: taming video diffusion models for high-fidelity novel view synthesis. arXiv preprint arXiv:2409.02048. External Links: 2409.02048 Cited by: §2.1.
- [75] Y. Zhang, C. Peng, B. Wang, P. Wang, Q. Zhu, F. Kang, B. Jiang, Z. Gao, E. Li, Y. Liu, and Y. Zhou (2025) Matrix-Game: interactive world foundation model. arXiv preprint arXiv:2506.18701. External Links: 2506.18701, Document, Link Cited by: §1.
- [76] Y. Zhao, C. Lin, K. Lin, Z. Yan, L. Li, Z. Yang, J. Wang, G. H. Lee, and L. Wang (2024) GenXD: generating any 3D and 4D scenes. arXiv preprint arXiv:2411.02319. External Links: 2411.02319 Cited by: §2.1.
- [77] S. Zheng, M. Yin, W. Hu, X. Li, Y. Shan, and Y. Fu (2026) VerseCrafter: dynamic realistic video world model with 4D geometric control. arXiv preprint arXiv:2601.05138. External Links: 2601.05138 Cited by: §2.1.
- [78] H. Zhou, W. Yu, C. Feng, X. Zhou, Y. Tian, and L. Yuan (2026) UniWorld-View: large-baseline view synthesis via video diffusion models. arXiv preprint arXiv:2608.04701. External Links: 2608.04701 Cited by: §1, §4.1.
- [79] H. Zhu, C. Wang, P. Tu, J. Luo, T. He, X. Jin, and Z. Chen (2026) GTA: advancing image-to-3D world generation via geometry then appearance video diffusion. arXiv preprint arXiv:2605.12957. External Links: 2605.12957 Cited by: §2.1.