Abstract
We present ’RenderFormer-V2’, a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to primitive transport, and a view-dependent stage that transforms the internal neural scene representation into image pixels. Different from RenderFormer, our model employs a novel combined windowed-attention and rendering-informed attention sink in the view-independent stage to improve scalability while maintaining render accuracy. To further improve versatility, RenderFormer-V2 supports heterogeneous scene primitives, including environment maps and participating media, and it employs a material encoding independent of the underlying surface reflectance model that encodes material appearance via a novel neural embedding. We demonstrate the versatility of RenderFormer-V2 on a variety of scenes and perform an extensive ablation of the improved attention mechanism.
Keywords:
Neural Rendering, Transformer, Neural Material Embedding, Windowed Attention, Attention Sink
1 Introduction
Neural rendering aims to visualize virtual scenes without relying on manually encoded rules of light transport, but instead based on relations between geometry, materials, and light learned from data. Many neural rendering solutions offer limited generalizability beyond the training data [12, 11] or rely on per-scene training strategies [30]. Recently, RenderFormer [43] formulated light transport simulation as a regressive sequence-to-sequence translation problem, where an input sequence of triangle tokens is transformed into pixel-patch tokens through a transformer-based two stage pipeline: a view-independent stage that resolves light transport between triangles, and a view-dependent stage that resolves transport from triangles to the camera. Once trained, RenderFormer can render a wide variety of virtual scenes without fine-tuning or further training. Although RenderFormer is more general than prior solutions, it is still far from practical: it is limited to scenes of less than k triangles, it only supports a hard-coded GGX BRDF model [33], and it is limited to scenes with (max. ) triangular diffuse light sources.
In this paper we introduce ’RenderFormer-V2’, a versatile transformer-based neural rendering system that addresses RenderFormer’s limitations via a number of carefully designed architectural innovations (Figure 1). Similar to RenderFormer, RenderFormer-V2 formulates light transport simulation as a regressive sequence-to-sequence translation. RenderFormer’s main bottleneck in supporting larger triangle meshes is the brute-force self-attention between triangle tokens in the view-independent stage. Not only does this have a quadratic complexity with respect to the number of triangles, it also leads RenderFormer to loose focus for very large triangle meshes. Inspired by recent advances in supporting larger context windows for large-language models, RenderFormer-V2 employs a novel sparse-attention variant, consisting of a combination of windowed attention [21] and render-aware attention-sinks [37], tuned for resolving view-independent light transport. Furthermore, to improve generalizability, we allow for other primitives than triangles (e.g., voxels) and employ a simpler positional encoding based on the centroid of the primitive.
RenderFormer-V2 further decouples the material specification from a hard-coded BRDF model, and instead employs a latent material embedding for encoding different BRDF models including transparent and measured materials. A key observation is that the latent material encoding does not need to be invertible to the input (BRDF) parameters, but it only needs to encode the appearance of the material; we rely on RenderFormer-V2 to learn how to map the material appearance into pixel values. Moreover, to model spatially varying materials, we reuse a pretrained VAE encoder [35] to encode texture patches (of latent material properties) per primitive. Similar to the latent material appearance space, we only require the VAE encoder, and let RenderFormer-V2 learn how to interpret the encoded textures during rendering.
Whereas RenderFormer has a dedicated emittance parameter associated with each triangle to model light sources, we leverage RenderFormer-V2’s ability to mix different primitives to embed different lighting types, ranging from triangular light sources to environment maps, into specialized tokens.
We demonstrate the versatility of RenderFormer-V2 by rendering more complex and larger scenes than RenderFormer with a greater variety in lighting and materials. We perform an in-depth ablation study to validate our design decisions. The trained RenderFormer-V2 model and code can be found at: https://renderformer.github.io/v2.
2 Related Work
Neural Rendering
[30] aims to predict the effects of light transport through a virtual scene. Early work in neural rendering employs specially learned neural representations of the scene [11, 12, 42, 14, 48] and thus are overfitted to a single or limited number of scenes. To circumvent the need to learn neural scene representations, image-space neural rendering systems [20, 44, 24] take as input G-buffers of intrinsic components of the scene, and output a shaded image seen from the same viewpoint. Because the G-buffers only capture a portion of the scene, image-space neural rendering methods must necessarily hallucinate (or ignore) transport between visible and non-visible parts of the scene.
Recently, a new class of neural rendering systems leverage attention layers [31] to model light transport between 3D primitives. Xu et al. [38] model diffuse light transport in a point cloud representation of the scene. Closest to our method is RenderFormer [43] which employs a two-stage transformer architecture that models the transport: (1) between triangles and (2) from the triangles to the camera. However, RenderFormer employs a brute-force attention mechanism which does not scale well to large triangle meshes. Moreover, RenderFormer only supports triangles as geometric primitives, diffuse (triangle-shaped) light sources, and a per-triangle hard-coded GGX microfacet BRDF model [33]. In contrast, RenderFormer-V2 employs an efficient sparse attention mechanism to support a large number (k) of primitives, and flexible geometry, lighting, and materials representations and textures.
Long Context Modeling with Transformers
Classic transformers compute attention between all pair-wise token combinations, resulting a quadratic complexity with respect to the number of tokens in the sequence. Moreover, when the sequence grows, attention per-token tends to decrease and be spread over many tokens, and as a consequence the transformer loses focus, resulting in a decreased performance. Addressing both issues is critical for scaling a transformer-based rendering architecture beyond a few thousand tokens. Here, we focus on the most relevant classes of transformer scaling methods, and refer to Tay et al. [29] for a detailed overview.
Windowed attention mechanisms [21, 40, 36, 3] focus on addressing the compute complexity, and built on the observation that in many cases proximity is a good indicator of importance, and hence these mechanism hard-constrain the attention computation to a small window around the target token. Consequently, windowed attention mechanisms ignore long-range interactions which can be important for light transport modeling. More generally, windowed attention mechanisms belong to a class of sparse attention methods that employ static attention patterns [19, 25, 16] and their effectiveness is highly dependent on whether the attention sparsity matches the attention pattern. While light transport through a scene can be sparse, it does not follow a pre-determined sparsity pattern.
Native-Sparse Attention (NSA) [41] dynamically determines the sparseness by employing three different attention streams: (i) a sliding window to capture local attention, (ii) compressed attention that determines the importance of groups of input tokens, and (iii) a fine-grained attention on the groups of tokens identified as important. However, the computational cost of NSA is significantly higher than static sparse attention patterns due to the secondary retrieval stage. Moreover, NSA requires Grouped-Query Attention [1] which lowers the model’s capacity, and thus adversely affects performance.
Hierarchical attention mechanisms (e.g., [39, 50, 45]) leverage the observation that attention tends to be focused near the query, and that the attention variation at distant tokens decreases. Hence, by creating a multi-resolution hierarchy of token and computing attention with the token selected from the hierarchy based on distance, attention can be better focused and more efficiently computed. However, multi-resolution hierarchies implicitly assume that positional distance is proportional to distance in the sequence or image, and thus implicitly assume a (regular) uniform spatial distribution of tokens. This is not the case for 3D scenes, where primitives are clustered at various points in space (i.e., objects). Point Transformer v3 [36] addresses this limitation by (i) serializing the point cloud along space-filling curves, and (ii) grouping and padding to ensure the point cloud is divisible by the target patch size. While, Point Transformer v3 improves speed and memory overhead, its implementation is more complex and is computationally more expensive than sparse attention methods. We employ a less complex and resource intensive strategy for extending the context window using a similar serialization strategy as Point Transformer v3.
Xiao et al. [37] observed that the soft-max operation in the attention computation tends to ’dump’ excess attention in a single token (i.e., attention sink). A similar behavior was also observed in vision transformers [18]. To avoid attention being dumped in a random token, Xiao et al. [37] propose to keep a few dedicated attention sink tokens to model global relations and a local sliding windowed attention to model local relations. This local-global dichotomy has been further refined in follow up work [47, 23]. We also build on this idea, and introduce rendering-relevant semantics for the sinks. First, we place all light sources in the sinks as these are likely to interact with all surfaces. Second, inspired by the compressed tokens in NSA [41], we add to the sink summarization tokens for groups of primitives based on Hilbert space-filling curves.
3 Background - RenderFormer
RenderFormer-V2 builds and improves on RenderFormer [43]. We therefore first review RenderFormer’s architecture before detailing RenderFormer-V2.
RenderFormer is an end-to-end trained transformer-based neural renderer that takes as input a sequence of triangles with GGX BRDF parameters [33] and emittance strength, as well as camera parameters, and it outputs a rendered image of the scene with full global illumination. RenderFormer consist of two stages with a slightly different architecture. The first (i.e., view-independent) stage, consisting of self-attention layers [31], transforms the input sequence of embedded triangle tokens (expressed in world coordinates) to a sequence of per-triangle tokens which encode triangle-to-triangle light transport. The second stage (i.e., view-dependent stage), consisting of 6 repetitions of a cross-attention layer [31] followed by a self-attention layer, operates on view-bundle tokens. A view-bundle token is an embedding of a grid of camera rays expressed in the camera coordinate system. The cross-attention layer computes the attention between the view-bundle tokens and the transformed triangle tokens from the first stage. The view-dependent stage is followed by a dense vision transformer to convert the transformed ray-bundle tokens into pixel values for each ray. The triangles are embedded as the sum of: (a) the per-vertex normal embedding (using NeRF positional encoding with 6 frequencies that is subsequently expanded to the token-length vector through a linear layer), (b) the GGX BRDF parameters (expanded by a linear layer to the token-length vector), and (c) the monochrome emittance (expanded by a linear layer). RenderFormer adapts RoPE [27] to apply a relative positional encoding on the D vector obtained by stacking the D coordinates of the triangle’s vertices. RoPE is applied at each layer in RenderFormer with the vertices expressed in world coordinates in the view-independent stage and in camera coordinates in the view-dependent stage. Hence, only the ray direction of the camera rays is embedded by stacking the view rays and subsequently expanded them to a length ray-bundle embedding via a linear layer. RenderFormer is trained end-to-end, first at a resolution and with scenes containing at most k triangles followed by a second training stage where the output resolution is increased to and the triangle count is increased to k. RenderFormer is trained with a weighted and LPIPS [46] loss on log-transformed reference renders.
4 Overview
Similar to RenderFormer, RenderFormer-V2 features a two-stage transformer-based neural rendering pipeline where the first stage resolves view-independent intra-primitive transport and the second stage transforms view-dependent ray-bundles to output tokens based on the transformed scene primitives from the first stage. However, RenderFormer-V2 deviates from RenderFormer is a number of critical steps: (i) RenderFormer-V2 is not limited to only triangle tokens and it supports a mixture of different scene primitives, including different types of light sources (Section 5), and (ii) RenderFormer-V2 scales better in terms of efficiency and accuracy to a larger number of scene primitives (Section 6). Figure 2 summarizes the RenderFormer-V2 pipeline.
5 Scene Embedding
We represent a virtual scene as a sequence of heterogeneous tokens that encode geometry, material, lighting, and camera information. In contrast to RenderFormer where the positional encoding is tailored to triangles as scene primitives, and which relies on a hard-coded camera-transformation to encode the camera position, we employ a uniform relative positional encoding strategy for all tokens (including ray-bundles):
For tokens representing a concept with a 3D spatial location (e.g., geometric primitive or camera) we use RoPE to encode the centroid of the concept with a RoPE dimension of (i.e., 20 frequencies). RoPE ensures that the scene embedding is invariant to scene translations.
For tokens representing positionless concepts (e.g., environment map) we employ RoPE using the centroid of the whole scene to ensure that translating the scene does not affect the relative attention computation between tokens from both categories.
As the different concepts are defined by different parameters, we employ a separate embedding for each token type, and rely on the training process to enable RenderFormer-V2 to differentiate between the different primitive embeddings.
5.1 Triangle & Material Embedding
To embed a triangle, we first embed the different components (vertices, normals, and materials) and combine them via addition into the final token.
Vertex Embedding
We stack the 3 positions of vertices (minus the centroid of the triangle) in a D vector, and apply (NeRF) positional encoding [22] with frequencies exponentially spaced between and . Finally, we apply a (trainable) linear layer to expand to a token-length (i.e., ) vector followed by RMS-normalization.
Normal Embedding
We apply the same process as for vertex embedding to encode the per-vertex normals. Note, the vertex and normal embedding use separately trainable linear layers for expansion.
Material Embedding
We desire an embedding of material appearance that is not tied to a particular BRDF model. Inspired by prior work on learning a latent embedding for BRDFs [28, 17, 49, 13, 10, 26], we also learn a material appearance embedding. A key advantage of RenderFormer-V2’s transformer architecture is that it is powerful enough to directly learn how to evaluate the embedded material appearance (given the view and lighting) without the need to rely on a pretrained reverse mapping from latent code to material appearance. As we are interested in encoding the appearance rather than the exact BRDF, we follow an encoding inspired by Serrano et al.’s [26] perceptual material similarity metric and embed rendered images of a sphere under the Uffizi Gallery light probe. We opt for a sphere for it simplicity and the Uffizi Gallery light probe because it is color neural and it contains a good mix of low and high frequency lighting features [4]. Practically, we employ a CNN-based auto-encoder with a D latent feature vector at the bottleneck. We pretrain this encoder with an L1 loss on images rendered with Blender Cycles of randomly generated materials with the Principled BRDF model [5]. Furthermore, to encourage a coherent manifold, we apply a smoothness regularization term [9], and a activation to constrain the values in the embedding to . Figure 3 visualizes the learned latent material appearance space. While we currently use the Principled BRDF model for generating training data, this can easily be extended to include other analytical BRDF models or measured BRDFs. For efficiency, we also train an additional MLP for each analytical BRDF model (after the latent space is trained) to map its parameters directly into the latent space to bypass the need to render a sphere; for measured BRDFs we render the material and use the pretrained encoder.
Texture Embedding
To support spatially varying materials, we embed all materials in a texture. For each triangle, we first project the BRDF parameters into the learned latent space and subsequently rasterize the per-triangle texture ( channels), local normal map ( channels), and a displacement map (as a channel height offset) in image patches, which we subsequently encode with a pretrained VAE [35] into an -channel latent feature map. Finally, we compress the feature map via a single learnable linear layer to a -length vector and add it to the token embedding.
5.2 Voxel Embedding
To demonstrate RenderFormer-V2’s ability to handle heterogenous geometric primitives, we also encode voxels filled with a scattering medium into a separate token-type.
Rotation and Scale Embedding
For each voxel we encode the rotation matrix and scale vector that describes the voxel’s relative rotation and per-axis scale in world coordinates. Similar to the normal encoding for triangles, we employ (NeRF) positional encoding with frequencies, which is subsequently expanded via a linear layer to a token-length vector.
Scattering and Absorption
As (RGB) scattering and (RGB) absorption coefficients are optical parameters of scattering media, and thus model independent, we opt to directly encode them. In addition, we also encode the anisotropy coefficient of the scattering function, yielding a D vector. To support spatially-varying scattering, we encode a volumetric texture of the D scattering feature vector, and linearly project the feature vector into a token-length vector that is added to the rotation and scale embedding.
5.3 Triangular Light Source Embedding
Similar to RenderFormer, we embed triangular light sources with homogeneous diffuse emittance. In contrast to RenderFormer, we store colored RGB emittance as an explicit light source token (instead of combining it with geometry tokens) which is expanded via a separate linear layer to the token-length. Similar to the triangle primitives, we add the vertex positions and per-vertex normals embedding to the token.
5.4 Environment Lighting Embedding
We follow an environment encoding similar to DiffusionRenderer [20].
Texture Embedding
We store both an LDR (clamped to ) and (log-encoded) HDR version (normalized by the log maximum value) of the environment map at resolution, and encode each map using a pretrained VAE [35] yielding a latent feature map for each. This feature map is too large to store in a single token. Hence, we opt to split the latent feature map in patches (of size ) that are compressed by a linear layer into two token-length vectors; hence yielding separate LDR and HDR environment map tokens.
Direction Embedding
While a environment map token does not have a position, each pixel in the pixel patch does correspond to a lighting direction. To make RenderFormer-V2 aware of the exact bundle of rays that correspond to the pixels in the patch per token, we create a direction map that, for each pixel, stores the corresponding (normalized) 3D direction vector in world coordinates. This direction map is projected via a linear layer to a token-length vector and added to the environment lighting embedding.
Lighting Strength Embedding
During texture embedding we normalized the log-encoded HDR by the log maximum value. To retain this information, we expand this value to a token-length vector and add it to the final embedding.
5.5 Camera Embedding
Similar to RenderFormer, we embed the (normalized) ray directions in an map, and project it to a token-length vector using a single linear layer. Whereas in RenderFormer the ray directions are expressed in camera coordinates, we encode them directly in world coordinates; the origin of the ray bundle is already taken care of via the uniform relative positional encoding strategy.
6 RenderFormer-V2 Architecture
While RenderFormer-V2 follows RenderFormer’s two stage design, each stage differs in how attention is computed. We first discuss the modification to the second stage (i.e., view-dependent), follow by the more significant modifications to the first stage (i.e., view-independent).
6.1 View-dependent Stage
Unlike RenderFormer’s view-dependent stage, RenderFormer-V2’s view-dependent stage operates in world coordinates (rather than camera coordinates). Moreover, we replace the full self-attention layers by SWIN windowed attention layers [21] with a window size of and a shift of . While the computation cost changes modestly from to ( is the number of scene tokens (the dominating factor), the number of ray-bundles, and the window size), its main advantage is the ability to scale to larger resolutions more easily because the context window size is now resolution independent.
6.2 View-independent Stage
The view-independent stage is the main bottle neck when increasing the number of tokens as computation cost scales quadratically with the number of primitives. Moreover, when the sequence grows large, naive attention tends to loose focus, resulting in a loss of render fidelity. While often the dominating factor, primitive-to-primitive light transport is not solely a local phenomena; a subset of global factors such as illumination and large-scale occlusion can significantly affect light transport through the scene. Hence, we combine attention from local primitives with attention from global tokens, while mitigating focus-loss on large sequences.
Sorting & Serialization
Unlike ray-bundles, scene primitives are not regularly spaced, making it more challenging to exploit locality while retaining efficient GPU-computation. Inspired by Point Transformer v3 [36], we sort and linearize geometry tokens using Hilbert curves such that nearby (in the 1D sequence) tokens are likely close in 3D space too. This allows us to model local attention with sliding window attention [3]; we take tokens before and after the target token, yielding a computational complexity independent of the number of primitives.
Attention Sinks for Rendering
To model global light transport, we adapt attention sinks [37]. The tokens in the attention sink are always included in the attention calculations. We strategically assign three different token types with rendering relevant semantics to the sink:
Global Register Tokens: we add register tokens [6] to the input sequence for storing storing global information.
Light Source Tokens: it is likely that the light transport on most geometric primitives is directly affected by the light sources in the scene. Hence, by placing the light sources in the sink we ensure that each light source is taken in account for each primitive regardless of distance. Conceptually, the attention computation with respect to the light sources is analogous to importance sampling the light sources in path tracing.
Summarization Tokens: while long range light transport is important, we argue that the precise details of distant geometry matter less. Therefore, we model long-range interactions with a coarser geometry representation. Specifically, we perform mean pooling over every consecutive geometry tokens to form a summarization token that we add to the attention sink.
7 Training
Training Data
Similarly to RenderFormer, we generate scenes by placing to randomly selected objects from the ObjaVerse dataset in one of four randomly selected template scenes (a ground plane with one, two or three walls). Different from RenderFormer’s scene generation process, we assign randomly generated materials (using the Principled BRDF model [5] mapped into our material appearance latent) from the following five categories: (a) homogeneous opaque diffuse+specular material ( chance; with the sum of the diffuse and specular albedo restricted to , and roughness log-sampled in ), (b) a homogeneous opaque material with metallic+roughness parameters (), (c) a diffuse+specular SVBRDF randomly sampled from the MatSynth SVBRDF dataset [32] (), (d) a metallic+roughness SVBRDF from MatSynth (), or (e) a homogeneous transparent metallic+roughness material (; with metallic sampled in , roughness log-sampled in , and each RGB channel of the base color sampled in ). In of the scenes, we add one (1/2) or two (1/2) volumes with randomly chosen scattering and absorption. The camera is randomly placed and aimed at the scene with a FOV chosen between and . Unlike RenderFormer, we allow the camera to be placed inside the scene. Up to triangular light sources are randomly placed outside the scene with a random (RGB) intensity between and . Furthermore, for of the scenes, we add environment lighting randomly selected from PolyHaven. To avoid color bias, we randomly swap the color channels in the environment map. Each scene is rendered offline using Blender Cycles with samples per pixel. For efficiency, we pre-generate a training dataset of M randomly sampled scenes spanning resolutions – with k–k primitives, totaling TB.
Loss Function
We employ the same loss as RenderFormer:
| (1) |
where the log encoding of the image serves to avoid specular reflections dominating the L1 error, and the LPIPS loss [46] minimizes perceptual differences.
Training Process & Refinement
We employ a five-stage training regime to first focus on learning the principles of light transport, before adding scene-complexity:
In stage 1, we decimate the generated scenes to k primitives and render the scene at resolution. To prioritize learning accurate coarse-scale transport, we employ full-attention at this stage. Furthermore, we gradually increase the complexity of the scene by first training exclusively on homogeneous materials ( day on A100 GPUs). Next we include SVBRDFs ( additional day), followed by the inclusion of environment lighting ( day; to focus training on the lighting effects, we mask out pixels that directly see the environment map), and finally adding volumetric objects ( day).
In stage 2 we increase the primitive budget to k as well as the resolution to , and continue to train for days on the same setup.
In stage 3, we keep the scene parameters the same, but switch from full-attention to our combined attention sink and windowed attention ( days).
In stage 4, we increase the primitive budget to k ( week).
Finally, in stage 5 we increase the primitive budget to k as well as the resolution to for another days of training.
Yielding a total training time of days on A100 GPUs. While the training cost is significant, we emphasize that once pretrained, no fine-tuning or training is needed for rendering a new scene.
8 Results
Capabilities
Figures 1 and 5 demonstrate the capabilities of RenderFormer-V2 on a wide variety of scenes. Figure 1 demonstrates that RenderFormer-V2 can render, with full global light transport, scenes containing refractive surfaces (1st) including caustics, environment lighting (2nd), volumetric scattering (3rd), and textures (4th). Figure 5 demonstrates displacement mapping (1st), volumetric scattering (2nd), and complex scenes with a large number of primitives (3rd and 4th). Prior neural rendering methods either require per-scene fine-tuning and/or cannot support all of these effects. Compared to path tracing, RenderFormer-V2 does not require specialized code to handle displacement mapping or volumetric scattering, and it can learn such effects solely by example.
RenderFormer-V2 employs a flexible material appearance latent space for assigning material properties to surface in the scene. Figure 4 demonstrates that, even though our latent space is trained on materials modeled with the Disney principled BRDF model [5], it can also model materials not part of the training set (e.g., in this case selected measured materials from the RGL BRDF dataset [8]). Moreover, as our latent space is based on encoding rendered images, it can easily be retrained to encompass material appearances currently not covered (e.g., anisotropic and color changing materials).
Comparison to RenderFormer
RenderFormer-V2 is closely related to RenderFormer [43]; both use a similar two-stage transformer-based pipeline. Compared to RenderFormer, RenderFormer-V2 achieves better accuracy when rendering scenes with a large number of triangles thanks to its sparse attention mechanism (Figure 6). Figure 9 qualitatively compares render quality for a scene with varying number of triangles; the darkening at k triangles when rendered with RenderFormer is a direct consequence of loss of attention. We used the publicly available version of RenderFormer which is trained on triangles; we found that training RenderFormer for larger triangle meshes is unstable. Moreover, RenderFormer-V2 is also considerably more efficient as shown in Figure 7 and Figure 8 thanks to the render-informed sparse attention in the view-independent stage and the SWIN-attention in the view-dependent stage respectively (all timings are measured on a single NVIDIA A100). Empirically, we found that rendering time is approximately equally distributed over both stages (i.e., view-dependent vs. view-independent). For reference, we also include timings of Blender Cycles on the same scenes using adaptive samples per pixel (i.e., the same setting as used for the training images).
RenderFormer-V2 not only supports larger triangle meshes, it can also be more easily fine-tuned for higher image resolution. Figure 10 compares two scenes rendered at and . An interesting observation is that despite both scenes containing the same number of triangles, that at higher resolution RenderFormer-V2 is able to more faithfully render fine detailed geometry (e.g., the dragon’s claws and Lucy’s face and hand).
| Model Variant | PSNR | SSIM | LPIPS | HDR-FLIP |
|---|---|---|---|---|
| RenderFormer-V2 | 28.25 | 0.8982 | 0.0997 | 0.4200 |
| w/o AS | 27.36 | 0.8841 | 0.1247 | 0.4359 |
| w/o SW | 26.92 | 0.8715 | 0.1289 | 0.4517 |
| AS w/o L | 27.40 | 0.8813 | 0.1081 | 0.4250 |
| AS w/o S | 27.09 | 0.8754 | 0.1210 | 0.4485 |
| AS w/o L, S | 26.13 | 0.8474 | 0.1514 | 0.4959 |
| More S (#seq / 32) | 27.78 | 0.8923 | 0.1072 | 0.4231 |
| Less S (#seq / 128) | 26.48 | 0.8690 | 0.1346 | 0.4655 |
| Less S (#seq / 256) | 26.93 | 0.8663 | 0.1316 | 0.4580 |
| Full Attention | 27.91 | 0.8953 | 0.1027 | 0.4287 |
Ablation
We ablate the different components of RenderFormer-V2’s sparse attention mechanism. Table 1 shows average PSNR, SSIM [34], LPIPS [46], and HDR-FLIP [2] over randomly generated scenes (using the same distribution as the training data, but with a different set of environment maps, SVBRDFs, and shapes than used for training) for different ablation variants: without sliding windowed attention, without attention sink, and leaving out lighting and/or summarization token from the attention sink – each ablation variant is trained up to stage 4 (k primitives). Only when all components are included, we achieve the highest accuracy in rendering. Surprisingly, our sparse attention outperforms full-attention, which we attribute to the full-attention model having to distribute its attention over too many tokens (i.e., loss of focus). Figure 11 further qualitatively demonstrates the difference in quality between the different ablation variants. We refer to the supplemental material for additional results and comparisons.
Limitations
While RenderFormer-V2 addresses many shortcomings of RenderFormer, it is not without limitations. First, RenderFormer-V2 is still limited to maximum light sources per scene, a constraint inherited from its training data. However, we argue that more complex lighting conditions are more efficiently modeled with the addition of an environment map. Furthermore, similar to RenderFormer, our model is trained on single frames, and it does not explicitely enforce temporal coherence. While RenderFormer-V2 supports textures/SVBRDFs, the resolution is fixed per triangle primitive (). Consequently, significant texture-quality degradation can occur for large triangles. However, this can easily be resolved by subdividing triangles to impose a maximum triangle size. By supporting heterogeneous primitives and using a material appearance parameterization independent of a hard-coded BRDF model, RenderFormer-V2 can easily be extended to support new primitives. However, training is most effective when new primitives are added in the first training stage. Consequently, adding a new primitive typically requires significant retraining. Improving extensibility without requiring a full retraining is an interesting avenue for future research.
9 Conclusion
In this paper we presented RenderFormer-V2, a transformer-based neural rendering model that takes as input a sequence of primitives (i.e., textured triangles, volumetric elements, light sources, environment maps, and camera) and outputs a rendered image of the scene with full global illumination. RenderFormer-V2 employs a rendering-aware sparse-attention to resolve intra-primitive light transport which enables RenderFormer-V2 to scale better to larger scenes, both for training as well as inference. Furthermore, we employ a flexible latent material appearance space to specify surface reflectance. RenderFormer-V2 is trained end-to-end, thereby avoiding the need for specialized code to handle displacement mapping or volumetric scattering.
Acknowledgment
Chong Zeng was supported by the Stanford Graduate Fellowship. This work was partially supported by the Brown Institute for Media Innovation at Stanford University.
References
- [1] J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai (2023) GQA: training generalized multi-query transformer models from multi-head checkpoints. In EMNPL, pp. 4895–4901. External Links: Link, Document Cited by: §2.
- [2] P. Andersson, J. Nilsson, T. Akenine-Möller, M. Oskarsson, K. Åström, and M. D. Fairchild (2020) FLIP: A Difference Evaluator for Alternating Images. Proceedings of the ACM on Computer Graphics and Interactive Techniques 3 (2), pp. 15:1–15:23. Cited by: §8, Appendix 0.B.
- [3] I. Beltagy, M. E. Peters, and A. Cohan (2020) Longformer: the long-document transformer. External Links: 2004.05150, Link Cited by: §2, §6.2.
- [4] J. Bieron and P. Peers (2020) An adaptive brdf fitting metric. Computer Graphics Forum 39 (4). External Links: Document Cited by: §5.1.
- [5] B. Burley (2012) Physically Based Shading at Disney. In SIGGRAPH 2012 Course Notes, Cited by: §5.1, §7, §8.
- [6] T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski (2024) Vision transformers need registers. In International Conference on Learning Representations (ICLR), Cited by: item 1.
- [7] J. Dong, B. FENG, D. Guessous, Y. Liang, and H. He (2025) FlexAttention: a programming model for generating fused attention variants.. In Eighth Conference on Machine Learning and Systems, External Links: Link Cited by: Appendix 0.A.
- [8] J. Dupuy and W. Jakob (2018) An adaptive parameterization for efficient material acquisition and rendering. Transactions on Graphics (Proceedings of SIGGRAPH Asia) 37 (6), pp. 274:1–274:18. External Links: Document Cited by: §8.
- [9] D. Gao, X. Li, Y. Dong, P. Peers, K. Xu, and X. Tong (2019) Deep inverse rendering for high-resolution svbrdf estimation from an arbitrary number of images. ACM Trans. Graph. 38 (4). External Links: ISSN 0730-0301, Link, Document Cited by: §5.1.
- [10] F. Gokbudak, A. Sztrajman, C. Zhou, F. Zhong, R. Mantiuk, and C. Oztireli (2024) Hypernetworks for generalizable brdf representation. In ECCV, pp. 73–89. External Links: Link, Document Cited by: §5.1.
- [11] J. Granskog, F. Rousselle, M. Papas, and J. Novák (2020) Compositional neural scene representations for shading inference. ACM Trans. Graph. 39 (4). Cited by: §1, §2.
- [12] J. Granskog, T. N. Schnabel, F. Rousselle, and J. Novák (2021) Neural scene graph rendering. ACM Trans. Graph. 40 (4). Cited by: §1, §2.
- [13] J. Guo, Z. Li, X. He, B. Wang, W. Li, Y. Guo, and L. Yan (2023) MetaLayer: a meta-learned bsdf model for layered materials. ACM Trans. Graph. 42 (6). External Links: Link, Document Cited by: §5.1.
- [14] A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa (2023) Instruct-nerf2nerf: editing 3d scenes with instructions. In CVPR, pp. 19740–19750. Cited by: §2.
- [15] A. Henry, P. R. Dachapally, S. Pawar, and Y. Chen (2020) Query-key normalization for transformers. External Links: 2010.04245, Link Cited by: Appendix 0.A.
- [16] J. Ho, N. Kalchbrenner, D. Weissenborn, and T. Salimans (2019) Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180. Cited by: §2.
- [17] B. Hu, J. Guo, Y. Chen, M. Li, and Y. Guo (2020) DeepBRDF: a deep representation for manipulating measured brdf. Computer Graphics Forum 39 (2), pp. 157–166. External Links: Document, Link Cited by: §5.1.
- [18] S. Kang, J. Kim, J. Kim, and S. J. Hwang (2025) See what you are told: visual attention sink in large multimodal models. In ICLR, Cited by: §2.
- [19] X. Li, M. Li, T. Cai, H. Xi, S. Yang, Y. Lin, L. Zhang, S. Yang, J. Hu, K. Peng, M. Agrawala, I. Stoica, K. Keutzer, and S. Han (2025) Radial attention: sparse attention with energy decay for long video generation. arXiv preprint arXiv:2506.19852. Cited by: §2.
- [20] R. Liang, Z. Gojcic, H. Ling, J. Munkberg, J. Hasselgren, C. Lin, J. Gao, A. Keller, N. Vijaykumar, S. Fidler, et al. (2025) Diffusion renderer: neural inverse and forward rendering with video diffusion models. In CVPR, pp. 26069–26080. Cited by: §2, §5.4.
- [21] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In ICCV, Vol. , pp. 9992–10002. External Links: Document Cited by: §1, §2, §6.1.
- [22] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng (2021) NeRF: representing scenes as neural radiance fields for view synthesis. Commun. ACM 65 (1), pp. 99–106. External Links: Document Cited by: §5.1.
- [23] T. Munkhdalai, M. Faruqui, and S. Gopal (2024) Leave no context behind: efficient infinite context transformers with infini-attention. External Links: 2404.07143, Link Cited by: §2.
- [24] O. Nalbach, E. Arabadzhiyska, D. Mehta, H.-P. Seidel, and T. Ritschel (2017) Deep shading: convolutional neural networks for screen space shading. Comput. Graph. Forum 36 (4), pp. 65–78. External Links: Link, Document Cited by: §2.
- [25] N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran (2018) Image Transformer. In ICML, pp. 4052–4061. Cited by: §2.
- [26] A. Serrano, B. Chen, C. Wang, M. Piovarči, H. Seidel, P. Didyk, and K. Myszkowski (2021) The effect of shape and illumination on material perception: model and applications. ACM Transactions on Graphics (TOG) 40 (4), pp. 1–16. Cited by: §5.1.
- [27] J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu (2024) RoFormer: enhanced transformer with rotary position embedding. Neurocomputing 568, pp. 127063. Cited by: §3.
- [28] A. Sztrajman, G. Rainer, T. Ritschel, and T. Weyrich (2021) Neural brdf representation and importance sampling. In Computer Graphics Forum, Vol. 40, pp. 332–346. Cited by: §5.1.
- [29] Y. Tay, M. Dehghani, D. Bahri, and D. Metzler (2022) Efficient transformers: a survey. ACM Comput. Surv. 55 (6). External Links: Link, Document Cited by: §2.
- [30] A. Tewari, J. Thies, B. Mildenhall, P. Srinivasan, E. Tretschk, W. Yifan, C. Lassner, V. Sitzmann, R. Martin-Brualla, S. Lombardi, T. Simon, C. Theobalt, M. Niessner, J. T. Barron, G. Wetzstein, M. Zollhöfer, and V. Golyanik (2022) Advances in neural rendering. Comp. Graph. Forum 41 (2), pp. 703–735. Cited by: §1, §2.
- [31] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. In NeurIPS, pp. 6000–6010. Cited by: §2, §3.
- [32] G. Vecchio and V. Deschaintre (2024) MatSynth: a modern pbr materials dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22109–22118. Cited by: §7.
- [33] B. Walter, S. R. Marschner, H. Li, and K. E. Torrance (2007) Microfacet models for refraction through rough surfaces. In Proceedings of the 18th Eurographics Conference on Rendering Techniques, EGSR’07, pp. 195–206. Cited by: §1, §2, §3.
- [34] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), pp. 600–612. Cited by: §8.
- [35] C. Wu, J. Li, J. Zhou, J. Lin, K. Gao, K. Yan, S. Yin, S. Bai, X. Xu, Y. Chen, Y. Chen, Z. Tang, Z. Zhang, Z. Wang, A. Yang, B. Yu, C. Cheng, D. Liu, D. Li, H. Zhang, H. Meng, H. Wei, J. Ni, K. Chen, K. Cao, L. Peng, L. Qu, M. Wu, P. Wang, S. Yu, T. Wen, W. Feng, X. Xu, Y. Wang, Y. Zhang, Y. Zhu, Y. Wu, Y. Cai, and Z. Liu (2025) Qwen-image technical report. External Links: 2508.02324, Link Cited by: §1, §5.1, §5.4.
- [36] X. Wu, L. Jiang, P. Wang, Z. Liu, X. Liu, Y. Qiao, W. Ouyang, T. He, and H. Zhao (2024) Point transformer v3: simpler, faster, stronger. In CVPR, Cited by: §2, §2, §6.2.
- [37] G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis (2024) Efficient streaming language models with attention sinks. External Links: 2309.17453, Link Cited by: §1, §2, §6.2.
- [38] B. Xu, C. Wang, T. Li, L. Wu, B. Wronski, R. Ramamoorthi, M. Salvi, et al. (2025) A generalizable light transport 3d embedding for global illumination. arXiv preprint arXiv:2510.18189. Cited by: §2.
- [39] J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao (2021) Focal attention for long-range interactions in vision transformers. In Neurips, Cited by: §2.
- [40] Y. Yang, Y. Guo, J. Xiong, Y. Liu, H. Pan, P. Wang, X. Tong, and B. Guo (2023) Swin3D: a pretrained transformer backbone for 3d indoor scene understanding. External Links: 2304.06906 Cited by: §2.
- [41] J. Yuan, H. Gao, D. Dai, J. Luo, L. Zhao, Z. Zhang, Z. Xie, Y. Wei, L. Wang, Z. Xiao, Y. Wang, C. Ruan, M. Zhang, W. Liang, and W. Zeng (2025) Native sparse attention: hardware-aligned and natively trainable sparse attention. In ACL, External Links: Document Cited by: §2, §2.
- [42] Y. Yuan, Y. Sun, Y. Lai, Y. Ma, R. Jia, and L. Gao (2022) Nerf-editing: geometry editing of neural radiance fields. In CVPR, pp. 18353–18364. Cited by: §2.
- [43] C. Zeng, Y. Dong, P. Peers, H. Wu, and X. Tong (2025) RenderFormer: transformer-based neural rendering of triangle meshes with global illumination. In Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, SIGGRAPH Conference Papers ’25. External Links: Link, Document Cited by: §1, §2, §3, §8, Figure 6, Figure 6, Appendix 0.C, Appendix 0.D.
- [44] Z. Zeng, V. Deschaintre, I. Georgiev, Y. Hold-Geoffroy, Y. Hu, F. Luan, L. Yan, and M. Hašan (2024) RGBx: image decomposition and synthesis using material- and lighting-aware diffusion models. In ACM SIGGRAPH 2024 Conference Papers, SIGGRAPH ’24. External Links: Link, Document Cited by: §2.
- [45] P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao (2021) Multi-scale vision longformer: a new vision transformer for high-resolution image encoding. In ICCV, pp. 2978–2988. External Links: Document Cited by: §2.
- [46] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, Cited by: §3, §7, §8.
- [47] Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Re, C. Barrett, Z. Wang, and B. Chen (2023) H2O: heavy-hitter oracle for efficient generative inference of large language models. In Neurips, External Links: Link Cited by: §2.
- [48] C. Zheng, Y. Huo, H. Huang, H. Sheng, J. Huang, R. Tang, H. Zhu, R. Wang, and H. Bao (2024) Neural global illumination via superposed deformable feature fields. In SIGGRAPH Asia 2024 Conference Papers, pp. 1–11. Cited by: §2.
- [49] C. Zheng, R. Zheng, R. Wang, S. Zhao, and H. Bao (2021) A compact representation of measured brdfs using neural processes. ACM Transactions on Graphics (TOG) 41 (2), pp. 1–15. Cited by: §5.1.
- [50] Z. Zhu and R. Soricut (2021) H-transformer-1d: fast one-dimensional hierarchical attention for sequences. External Links: 2107.11906 Cited by: §2.
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
Appendix 0.A Additional Implementation Details
Table 1summarizes the key hyper-parameters of RenderFormer-V2’s two-stage transformer pipeline. Both stages share the same model dimension (768), number of attention heads (6), and Feed-Forward Network (FFN) dimension (3072). The sparse attention in the view-independent stage, consisting of a local sliding window (total size 512) combined with global attention sink tokens, is implemented using PyTorch’s flex_attention API [7]. We also employ QK-Norm [15] to stabilize the attention mechanism. The complete RenderFormer-V2 model comprises approximately 207M parameters in total.
| View-indep. | View-dep. | |
| Layers | 12 | 6 |
| Model Dimension | 768 | 768 |
| Attention Heads | 6 | 6 |
| Attention Type | Sparse Self-Attn. | Cross-Attn. + Swin Self-Attn. |
| FFN Dimension | 3072 | 3072 |
| FFN Activation | SwiGLU | |
| Normalization | RMSNorm | |
| RenderFormer-V2 | Reference | Diff () | FLIP | Metrics |
![]() | ![]() | ![]() | ![]() | #Tokens: 4481 PSNR: 26.50 SSIM: 0.9481 LPIPS: 0.0485 FLIP: 0.1189 |
![]() | ![]() | ![]() | ![]() | #Tokens: 3009 PSNR: 22.24 SSIM: 0.9265 LPIPS: 0.0398 FLIP: 0.1773 |
![]() | ![]() | ![]() | ![]() | #Tokens: 490 PSNR: 36.51 SSIM: 0.9861 LPIPS: 0.0373 FLIP: 0.0825 |
![]() | ![]() | ![]() | ![]() | #Tokens: 17653 PSNR: 30.99 SSIM: 0.9653 LPIPS: 0.0320 FLIP: 0.1168 |
| RenderFormer-V2 | Reference | Diff () | FLIP | Metrics |
![]() | ![]() | ![]() | ![]() | #Tokens: 5633 PSNR: 30.97 SSIM: 0.9297 LPIPS: 0.0319 FLIP: 0.1845 |
![]() | ![]() | ![]() | ![]() | #Tokens: 36353 PSNR: 27.05 SSIM: 0.9633 LPIPS: 0.0648 FLIP: 0.1605 |
![]() | ![]() | ![]() | ![]() | #Tokens: 29276 PSNR: 24.34 SSIM: 0.9171 LPIPS: 0.0712 FLIP: 0.1992 |
![]() | ![]() | ![]() | ![]() | #Tokens: 3993 PSNR: 26.41 SSIM: 0.9243 LPIPS: 0.0474 FLIP: 0.1031 |
Appendix 0.B Qualitative and Quantitative Comparison
Figure 1and 2 show qualitative comparisons for all the scenes from the main submission with respect to reference Blender Cycles path-traced renderings. For each scene we also show a difference image (scaled to better show the differences) and a FLIP error image [2]. In addition, we list the total number of tokens per scene, and the PSNR, SSIM, LPIPS and FLIP errors.
Transparent Torus: (Figure 1, 1st row) the differences are mainly visible on high curvature areas of the refractive torus, as well as in the intensity of the caustic on the ground plane.
Environment Lit Spheres: (2nd row) the differences are mainly due to differences at high-frequency edges in the image at texture / environment map pixel edges. Furthermore, we can also observe a slight overall brightness difference.
Smoky Bunny: (3rd row) Due to differences in how anti-aliasing is handled (Blender Cycles uses adaptive filtering), larger differences are visible at high frequency edges in the rendered images.
Three Teapot Scene: (4th row) Similar to the previous scene, differences are mainly concentrated at high frequency edges in the image. Another area of difference is the inter-reflection below the painting; RenderFormer-V2 assumes the back of the painting is also textured and reflected onto the wall.
Displacement Mapped Cornell Cube: (Figure 2, 1st row). Again most differences are in high-frequency areas, due to (1) differences in filtering, and (2) RenderFormer-V2 sometimes misses small details.
Spaceship in Smoke: (2nd row) The main differences are due to minor inaccuracies in reflected directions (i.e., shifted or missing highlights).
Dinner Scene: (3rd row) Again, the main differences are at high frequency edges in the rendering, as well as slightly more blurred shadows.
Cube Pile Scene: (4th row) Similar as in prior scenes; the main differences are at high frequency edges.
Appendix 0.C Comparisons with RenderFormer
RenderFormer-V2 shares some architectural similarities with RenderFormer [43]. To better assess the differences and improvements, we perform an in-depth comparison.
0.C.1 Architecture
Figure 3 contrasts the RenderFormer architecture with RenderFormer-V2’s architecture. The key differences are:
Input Tokens. RenderFormer encodes all primitives (including light sources) in a homogeneous triangle token. RenderFormer-V2 supports heterogeneous tokens, each with its own encoding procedure, supporting a wide range of primitives ranging from triangles, volume elements, and environment maps. In addition, RenderFormer-V2 adds support for spatially varying materials (with displacement mapping) and uses a BRDF-model-agnostic material specification.
View-independent Attention. Whereas RenderFormer utilizes a full (dense) self-attention, RenderFormer-V2 uses a sliding windowed attention combined with attention sinks. The sinks have rendering-aware semantics and include: global registers, light sources, and summarization tokens to model long-range light transport.
View-dependent Attention. RenderFormer uses full self-attention between all ray-bundle tokens. This is costly and it is unlikely that pixels far away in the image will be meaningfully related to nearby pixels. Therefore, RenderFormer-V2 uses SWIN attention instead.
By extending the supported token types and by adopting a novel sparse attention architecture, RenderFormer-V2 enables more accurate rendering of more complex scenes with richer visual effects.
0.C.2 Qualitative Comparison
Figure 4 qualitatively compares three scenes rendered with RenderFormer-V2 vs. RenderFormer vs. a path-tracer reference rendering (difference images are scaled to better highlight discrepancies). In all cases, we observe that RenderFormer-V2 more accurately renders fine geometrical details such as the corrugated structures and the narrow gaps between cuboids in the first row, as well as the high-frequency geometric patterns in the second and third rows. The difference images further show that RenderFormer-V2 yields smaller rendering errors overall.
| Ours | Diff () | RenderFormer | Diff () | Reference |
![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() |
| Model | PSNR | SSIM | LPIPS | FLIP |
|---|---|---|---|---|
| RenderFormer | 26.2423 | 0.9473 | 0.0886 | 0.1332 |
| RenderFormer-V2 | 28.0810 | 0.9648 | 0.0503 | 0.1172 |
| RenderFormer-V2 (finetuned) | 30.7304 | 0.9742 | 0.0244 | 0.0952 |
| Reference | RenderFormer | RenderFormer-V2 | RF2 (fine-tuned) | |
|---|---|---|---|---|
| 512 | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ||
| 1024 | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ||
| 2048 | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() |
0.C.3 Resolution Scalability
Thanks to the locality of the SWIN attention in the view-dependent stage, RenderFormer-V2 exhibits better scalability with respect to changes in render-resolution. We validate this quantitatively (Table 2) and qualitatively (Figure 5) by comparing the accuracy of RenderFormer vs. RenderFormer-V2 vs. a resolution fine-tune of RenderFormer-V2. RenderFormer shows noticeably higher errors when rendering at higher resolutions ( and ). Moreover, fine-tuning RenderFormer-V2 at resolution further improves RenderFormer-V2 ’s ability to recover fine details under high-resolution settings.
| Tri Count | Model | PSNR | SSIM | LPIPS | FLIP |
|---|---|---|---|---|---|
| 4K | RenderFormer-V2 | 33.84 | 0.9776 | 0.0198 | 0.1025 |
| RenderFormer | 34.51 | 0.9823 | 0.0244 | 0.0838 | |
| 8K | RenderFormer-V2 | 33.00 | 0.9682 | 0.0305 | 0.1230 |
| RenderFormer | 32.89 | 0.9708 | 0.0358 | 0.1079 | |
| 16K | RenderFormer-V2 | 29.96 | 0.9457 | 0.0502 | 0.1932 |
| RenderFormer | 29.36 | 0.9418 | 0.0604 | 0.1555 | |
| 32K | RenderFormer-V2 | 28.07 | 0.9205 | 0.0729 | 0.2387 |
| RenderFormer | 24.17 | 0.8451 | 0.1500 | 0.3080 | |
| 64K | RenderFormer-V2 | 26.79 | 0.9020 | 0.0873 | 0.2609 |
| RenderFormer | 18.72 | 0.6651 | 0.3070 | 0.5313 | |
| 128K | RenderFormer-V2 | 25.82 | 0.8847 | 0.0999 | 0.3220 |
| RenderFormer | 15.88 | 0.4604 | 0.4772 | 0.7466 |
0.C.4 Full Detailed Metrics
Finally, for completeness, in Table 3 we show the full rendering quality comparison with RenderFormer which was summarized in the main submission in the error plot (Fig. 6).
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() |
Appendix 0.D Additional Results
Figure 6 shows scenes used by Zeng et al. [43] to demonstrate the capabilities of RenderFormer, which RenderFormer-V2 can also handle without issues.
Finally, we show additional complex room scenes in Figure 7 with fine geometrical details, complex light transport, and textures.
We refer to the supplementary video for additional results demonstrating RenderFormer-V2’s capabilities as well as its stability to changes in scene and camera parameters.
| Pred | Reference | Diff () | FLIP | Metrics |
![]() | ![]() | ![]() | ![]() | PSNR: 29.07 SSIM: 0.9285 LPIPS: 0.04980 FLIP: 0.2044 |
![]() | ![]() | ![]() | ![]() | PSNR: 30.92 SSIM: 0.9353 LPIPS: 0.04512 FLIP: 0.1125 |
![]() | ![]() | ![]() | ![]() | PSNR: 34.99 SSIM: 0.9506 LPIPS: 0.06065 FLIP: 0.1719 |
![]() | ![]() | ![]() | ![]() | PSNR: 28.92 SSIM: 0.8827 LPIPS: 0.1240 FLIP: 0.2515 |
Appendix 0.E Ablation Test Scenes
Figure 8 shows example scenes from our test set used for the ablation studies. These procedurally generated scenes contain diverse textures, a variety of test objects, large scene token counts, as well as volumetric effects, area lighting, and environment lighting.
Appendix 0.F Limitations
Figure 9 illustrates the limitations of RenderFormer-V2 in rendering high-resolution textures on large triangles. Since each triangle is associated with a fixed-resolution texture embedding (), the effective texture resolution becomes low when the triangle covers a large surface area. By subdividing the mesh to increase the number of triangles and thus reduce triangle size, the texture quality can be significantly improved.
| Reference | No Subdiv. | |
![]() | ![]() | ![]() |
![]() | ![]() | ![]() |
References
- [1] J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai (2023) GQA: training generalized multi-query transformer models from multi-head checkpoints. In EMNPL, pp. 4895–4901. External Links: Link, Document Cited by: §2.
- [2] P. Andersson, J. Nilsson, T. Akenine-Möller, M. Oskarsson, K. Åström, and M. D. Fairchild (2020) FLIP: A Difference Evaluator for Alternating Images. Proceedings of the ACM on Computer Graphics and Interactive Techniques 3 (2), pp. 15:1–15:23. Cited by: §8, Appendix 0.B.
- [3] I. Beltagy, M. E. Peters, and A. Cohan (2020) Longformer: the long-document transformer. External Links: 2004.05150, Link Cited by: §2, §6.2.
- [4] J. Bieron and P. Peers (2020) An adaptive brdf fitting metric. Computer Graphics Forum 39 (4). External Links: Document Cited by: §5.1.
- [5] B. Burley (2012) Physically Based Shading at Disney. In SIGGRAPH 2012 Course Notes, Cited by: §5.1, §7, §8.
- [6] T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski (2024) Vision transformers need registers. In International Conference on Learning Representations (ICLR), Cited by: item 1.
- [7] J. Dong, B. FENG, D. Guessous, Y. Liang, and H. He (2025) FlexAttention: a programming model for generating fused attention variants.. In Eighth Conference on Machine Learning and Systems, External Links: Link Cited by: Appendix 0.A.
- [8] J. Dupuy and W. Jakob (2018) An adaptive parameterization for efficient material acquisition and rendering. Transactions on Graphics (Proceedings of SIGGRAPH Asia) 37 (6), pp. 274:1–274:18. External Links: Document Cited by: §8.
- [9] D. Gao, X. Li, Y. Dong, P. Peers, K. Xu, and X. Tong (2019) Deep inverse rendering for high-resolution svbrdf estimation from an arbitrary number of images. ACM Trans. Graph. 38 (4). External Links: ISSN 0730-0301, Link, Document Cited by: §5.1.
- [10] F. Gokbudak, A. Sztrajman, C. Zhou, F. Zhong, R. Mantiuk, and C. Oztireli (2024) Hypernetworks for generalizable brdf representation. In ECCV, pp. 73–89. External Links: Link, Document Cited by: §5.1.
- [11] J. Granskog, F. Rousselle, M. Papas, and J. Novák (2020) Compositional neural scene representations for shading inference. ACM Trans. Graph. 39 (4). Cited by: §1, §2.
- [12] J. Granskog, T. N. Schnabel, F. Rousselle, and J. Novák (2021) Neural scene graph rendering. ACM Trans. Graph. 40 (4). Cited by: §1, §2.
- [13] J. Guo, Z. Li, X. He, B. Wang, W. Li, Y. Guo, and L. Yan (2023) MetaLayer: a meta-learned bsdf model for layered materials. ACM Trans. Graph. 42 (6). External Links: Link, Document Cited by: §5.1.
- [14] A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa (2023) Instruct-nerf2nerf: editing 3d scenes with instructions. In CVPR, pp. 19740–19750. Cited by: §2.
- [15] A. Henry, P. R. Dachapally, S. Pawar, and Y. Chen (2020) Query-key normalization for transformers. External Links: 2010.04245, Link Cited by: Appendix 0.A.
- [16] J. Ho, N. Kalchbrenner, D. Weissenborn, and T. Salimans (2019) Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180. Cited by: §2.
- [17] B. Hu, J. Guo, Y. Chen, M. Li, and Y. Guo (2020) DeepBRDF: a deep representation for manipulating measured brdf. Computer Graphics Forum 39 (2), pp. 157–166. External Links: Document, Link Cited by: §5.1.
- [18] S. Kang, J. Kim, J. Kim, and S. J. Hwang (2025) See what you are told: visual attention sink in large multimodal models. In ICLR, Cited by: §2.
- [19] X. Li, M. Li, T. Cai, H. Xi, S. Yang, Y. Lin, L. Zhang, S. Yang, J. Hu, K. Peng, M. Agrawala, I. Stoica, K. Keutzer, and S. Han (2025) Radial attention: sparse attention with energy decay for long video generation. arXiv preprint arXiv:2506.19852. Cited by: §2.
- [20] R. Liang, Z. Gojcic, H. Ling, J. Munkberg, J. Hasselgren, C. Lin, J. Gao, A. Keller, N. Vijaykumar, S. Fidler, et al. (2025) Diffusion renderer: neural inverse and forward rendering with video diffusion models. In CVPR, pp. 26069–26080. Cited by: §2, §5.4.
- [21] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In ICCV, Vol. , pp. 9992–10002. External Links: Document Cited by: §1, §2, §6.1.
- [22] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng (2021) NeRF: representing scenes as neural radiance fields for view synthesis. Commun. ACM 65 (1), pp. 99–106. External Links: Document Cited by: §5.1.
- [23] T. Munkhdalai, M. Faruqui, and S. Gopal (2024) Leave no context behind: efficient infinite context transformers with infini-attention. External Links: 2404.07143, Link Cited by: §2.
- [24] O. Nalbach, E. Arabadzhiyska, D. Mehta, H.-P. Seidel, and T. Ritschel (2017) Deep shading: convolutional neural networks for screen space shading. Comput. Graph. Forum 36 (4), pp. 65–78. External Links: Link, Document Cited by: §2.
- [25] N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran (2018) Image Transformer. In ICML, pp. 4052–4061. Cited by: §2.
- [26] A. Serrano, B. Chen, C. Wang, M. Piovarči, H. Seidel, P. Didyk, and K. Myszkowski (2021) The effect of shape and illumination on material perception: model and applications. ACM Transactions on Graphics (TOG) 40 (4), pp. 1–16. Cited by: §5.1.
- [27] J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu (2024) RoFormer: enhanced transformer with rotary position embedding. Neurocomputing 568, pp. 127063. Cited by: §3.
- [28] A. Sztrajman, G. Rainer, T. Ritschel, and T. Weyrich (2021) Neural brdf representation and importance sampling. In Computer Graphics Forum, Vol. 40, pp. 332–346. Cited by: §5.1.
- [29] Y. Tay, M. Dehghani, D. Bahri, and D. Metzler (2022) Efficient transformers: a survey. ACM Comput. Surv. 55 (6). External Links: Link, Document Cited by: §2.
- [30] A. Tewari, J. Thies, B. Mildenhall, P. Srinivasan, E. Tretschk, W. Yifan, C. Lassner, V. Sitzmann, R. Martin-Brualla, S. Lombardi, T. Simon, C. Theobalt, M. Niessner, J. T. Barron, G. Wetzstein, M. Zollhöfer, and V. Golyanik (2022) Advances in neural rendering. Comp. Graph. Forum 41 (2), pp. 703–735. Cited by: §1, §2.
- [31] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. In NeurIPS, pp. 6000–6010. Cited by: §2, §3.
- [32] G. Vecchio and V. Deschaintre (2024) MatSynth: a modern pbr materials dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22109–22118. Cited by: §7.
- [33] B. Walter, S. R. Marschner, H. Li, and K. E. Torrance (2007) Microfacet models for refraction through rough surfaces. In Proceedings of the 18th Eurographics Conference on Rendering Techniques, EGSR’07, pp. 195–206. Cited by: §1, §2, §3.
- [34] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), pp. 600–612. Cited by: §8.
- [35] C. Wu, J. Li, J. Zhou, J. Lin, K. Gao, K. Yan, S. Yin, S. Bai, X. Xu, Y. Chen, Y. Chen, Z. Tang, Z. Zhang, Z. Wang, A. Yang, B. Yu, C. Cheng, D. Liu, D. Li, H. Zhang, H. Meng, H. Wei, J. Ni, K. Chen, K. Cao, L. Peng, L. Qu, M. Wu, P. Wang, S. Yu, T. Wen, W. Feng, X. Xu, Y. Wang, Y. Zhang, Y. Zhu, Y. Wu, Y. Cai, and Z. Liu (2025) Qwen-image technical report. External Links: 2508.02324, Link Cited by: §1, §5.1, §5.4.
- [36] X. Wu, L. Jiang, P. Wang, Z. Liu, X. Liu, Y. Qiao, W. Ouyang, T. He, and H. Zhao (2024) Point transformer v3: simpler, faster, stronger. In CVPR, Cited by: §2, §2, §6.2.
- [37] G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis (2024) Efficient streaming language models with attention sinks. External Links: 2309.17453, Link Cited by: §1, §2, §6.2.
- [38] B. Xu, C. Wang, T. Li, L. Wu, B. Wronski, R. Ramamoorthi, M. Salvi, et al. (2025) A generalizable light transport 3d embedding for global illumination. arXiv preprint arXiv:2510.18189. Cited by: §2.
- [39] J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao (2021) Focal attention for long-range interactions in vision transformers. In Neurips, Cited by: §2.
- [40] Y. Yang, Y. Guo, J. Xiong, Y. Liu, H. Pan, P. Wang, X. Tong, and B. Guo (2023) Swin3D: a pretrained transformer backbone for 3d indoor scene understanding. External Links: 2304.06906 Cited by: §2.
- [41] J. Yuan, H. Gao, D. Dai, J. Luo, L. Zhao, Z. Zhang, Z. Xie, Y. Wei, L. Wang, Z. Xiao, Y. Wang, C. Ruan, M. Zhang, W. Liang, and W. Zeng (2025) Native sparse attention: hardware-aligned and natively trainable sparse attention. In ACL, External Links: Document Cited by: §2, §2.
- [42] Y. Yuan, Y. Sun, Y. Lai, Y. Ma, R. Jia, and L. Gao (2022) Nerf-editing: geometry editing of neural radiance fields. In CVPR, pp. 18353–18364. Cited by: §2.
- [43] C. Zeng, Y. Dong, P. Peers, H. Wu, and X. Tong (2025) RenderFormer: transformer-based neural rendering of triangle meshes with global illumination. In Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, SIGGRAPH Conference Papers ’25. External Links: Link, Document Cited by: §1, §2, §3, §8, Figure 6, Figure 6, Appendix 0.C, Appendix 0.D.
- [44] Z. Zeng, V. Deschaintre, I. Georgiev, Y. Hold-Geoffroy, Y. Hu, F. Luan, L. Yan, and M. Hašan (2024) RGBx: image decomposition and synthesis using material- and lighting-aware diffusion models. In ACM SIGGRAPH 2024 Conference Papers, SIGGRAPH ’24. External Links: Link, Document Cited by: §2.
- [45] P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao (2021) Multi-scale vision longformer: a new vision transformer for high-resolution image encoding. In ICCV, pp. 2978–2988. External Links: Document Cited by: §2.
- [46] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, Cited by: §3, §7, §8.
- [47] Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Re, C. Barrett, Z. Wang, and B. Chen (2023) H2O: heavy-hitter oracle for efficient generative inference of large language models. In Neurips, External Links: Link Cited by: §2.
- [48] C. Zheng, Y. Huo, H. Huang, H. Sheng, J. Huang, R. Tang, H. Zhu, R. Wang, and H. Bao (2024) Neural global illumination via superposed deformable feature fields. In SIGGRAPH Asia 2024 Conference Papers, pp. 1–11. Cited by: §2.
- [49] C. Zheng, R. Zheng, R. Wang, S. Zhao, and H. Bao (2021) A compact representation of measured brdfs using neural processes. ACM Transactions on Graphics (TOG) 41 (2), pp. 1–15. Cited by: §5.1.
- [50] Z. Zhu and R. Soricut (2021) H-transformer-1d: fast one-dimensional hierarchical attention for sequences. External Links: 2107.11906 Cited by: §2.




































































































