Vladislav Bargatin
AI Center, Lomonosov MSU, Moscow, Russia
Alexander Yakovenko
AI Center, Lomonosov MSU, Moscow, Russia
Lomonosov Moscow State University, Moscow, Russia
Khaled Abud
AI Center, Lomonosov MSU, Moscow, Russia
Lomonosov Moscow State University, Moscow, Russia
MSU Institute for Artificial Intelligence, Moscow, Russia
https://github.com/msu-video-group/freeflow
{vladislav.bargatin, alexander.yakovenko}@graphics.cs.msu.ru
Dmitriy Vatolin
{khaled.abud, dmitriy}@graphics.cs.msu.ru
MSU Institute for Artificial Intelligence, Moscow, Russia
https://github.com/msu-video-group/freeflow
{vladislav.bargatin, alexander.yakovenko}@graphics.cs.msu.ru
Abstract
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefined heuristics, which can limit its expressivity and lead to more complex pipelines and additional computational cost. We present FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder–decoder. FreeFlow combines three attention variants: window attention for local processing, shifted-window attention for cross-window information exchange, and a global attention operating at a reduced resolution. The resulting architecture scales naturally with model capacity, enabling a consistent accuracy gain from small to large variants. Despite the absence of standard inductive biases, FreeFlow achieves state-of-the-art results on major benchmarks, including Sintel (0.68/1.48 EPE on Clean/Final), KITTI-2015 (3.23 Fl-all), and Spring (3.192 1px), while remaining memory efficient at 1080p inference.
Keywords:
Optical flow Vision Transformers High-resolution
1 Introduction
Optical flow estimation (the dense per-pixel motion between frames) is a fundamental task in low-level vision, with applications ranging from video understanding[32, 44, 62] and tracking to video restoration and synthesis[12, 24, 57, 6].
Early optical flow methods posed estimation as variational optimization[27, 10], later accumulating stronger regularization [41], coarse-to-fine schemes, and hand-crafted descriptors[53] at the cost of growing complexity.
Deep learning initially simplified optical flow, as FlowNet[9] showed that a feed-forward network can regress flow directly, but later work reintroduced classical biases in highly effective yet hard-wired architectures. PWC-Net[42] combined pyramids, warping, and local correlation volumes, while RAFT[45] popularized iterative refinement over an all-pairs correlation volume and convex upsampling; many follow-ups then build upon these foundations with additional modules for occlusions[14], temporal cues[38, 7], and training/inference refinements[50]. Which improve accuracy, but lead to complex pipelines that are harder to modify, scale, and repurpose beyond optical flow.
In parallel, the broader computer vision literature has moved in the opposite direction: vision transformers[8] increasingly replace bespoke pipelines in detection[5], segmentation[18, 16], depth prediction[58, 59], and 3D tasks[47, 46, 15], driven by the observation that generic, data-driven feed-forward models can learn the required structure from scale and supervision. This trend motivates revisiting optical flow through the same lens: can we omit flow-specific components and still obtain a state-of-the-art approach?
Several recent approaches move toward more generic architectures, but still retain flow-specific structure or incur practical limitations. DDVM[37] casts optical flow prediction in a diffusion framework, but the prediction is still produced through iterative process. CroCo[52] is close to a pure transformer, yet it operates at a relatively small fixed resolution and typically requires tiling for higher-resolution inputs, while still relying on a large convolutional decoder. Most recently, WAFT[49] and GeoViT[54] explicitly aim for generality, but still rely on iterations and require warping (of features and input frames, respectively). Additionally, these approaches are trained at relatively small, sub-megapixel resolutions, which might become a limiting factor for accuracy at high-resolution inference [1].
We therefore propose FreeFlow, a bias-free hierarchical transformer for optical flow estimation. FreeFlow is designed for high-resolution processing, which recent work has shown to be beneficial for optical flow[1] and depth estimation[3]. At high resolution, accurate flow requires both strong local reasoning (to preserve fine structures and motion boundaries) and effective long-range information flow (to resolve large displacements). FreeFlow addresses this with a hierarchical attention design that combines tiled processing with repeated local–global feature interaction: local attention focuses on within-region detail, cross-region exchange propagates information between neighboring tiles, and global mixing enables long-range correspondence. FreeFlow sets a new state-of-the-art on Sintel[4] (EPE clean: 0.68; EPE final: 1.48), KITTI-2015[30] (Fl-all: 3.23), and Spring[29] (1px: 3.192).
Our key contributions are:
Bias-free optical flow transformer. We introduce FreeFlow, a hierarchical transformer for optical flow built without common flow-specific inductive biases and modules (e.g., explicit cost/correlation volumes, warping-based update pipelines, and specialized upsampling), using a single end-to-end trainable architecture.
Hierarchical local–global feature interaction for high-resolution processing. We propose a tiled transformer design with repeated local and long-range feature interaction, enabling accurate flow at high resolutions by combining local detail modeling with global information flow.
State-of-the-art performance across benchmarks. FreeFlow achieves state-of-the-art results on Sintel, KITTI-2015, and Spring, all while being memory-efficient and scalable to smaller parameter counts.
2 Related Works
2.1 Inductive Biases of Optical Flow
Early optical flow methods formulated estimation as variational optimization under photometric constancy and smoothness regularization[27, 10, 41, 53]. Learning-based approaches later treated optical flow as supervised dense prediction[9], and subsequent advances largely introduced explicit architectural priors tailored to correspondence estimation. Representative flow-specific inductive biases include multi-scale pyramids[42, 31], warping-based alignment[42, 49, 54], explicit correlation volumes for large-displacement matching[45, 1], iterative refinement, and convex upsampling[45, 50, 1]. While these choices are effective, they have contributed to increasingly complex pipelines; recent work has started to remove individual priors[17] and move toward more generic formulations[37], but existing methods typically retain some of these biases rather than eliminating them entirely.
| Method | Correlation Volume | Feature Warping | Pyramid Refinement | Iterative Refinement | Convex Upsampling | Inference time Tiling | Flow Head |
| PWC-Net [42] | |||||||
| RAFT [45] | |||||||
| FlowFormer [11] | |||||||
| UniMatch [56] | |||||||
| TransFlow [26] | |||||||
| DPFlow [31] | |||||||
| MEMFOF [1] | |||||||
| WAFT [49] | |||||||
| Geo-VIT [54] | |||||||
| CroCo-Flow [52] | DPT head | ||||||
| Win-Win [20] | DPT head | ||||||
| FreeFlow (ours) | Simple Conv. |
2.2 Vision Transformers in Dense Prediction
Vision transformers[8] (ViTs) are now widely used for dense prediction and geometry-oriented tasks, enabled by large-scale data and a general architecture that can learn the dependencies from supervision rather than relying on hand-crafted priors. This has led to strong results across depth[33, 58, 59, 3], segmentation[18, 16], and 3D settings[47, 46, 15], where feed-forward transformer backbones provide a common foundation for dense outputs and geometric reasoning. Importantly, many of these systems remain architecturally simple. A common pattern is a ViT backbone combined with a convolutional prediction head or decoder for producing dense maps[33, 34]. Other approaches further reduce decoder structure and rely on minimal output token processing, suggesting that heavy convolutional decoders are not necessary to obtain competitive dense predictions [16, 15]. However, scaling such models to high resolution is challenging; common workarounds such as downsampling or tiling can harm accuracy and limit long-range interaction. Swin[25, 22] addresses the issue by using local-window self-attention and shifted windows to pass information across window boundaries, while Hiera[36] provides a simple hierarchical multi-scale ViT backbone for large images. Alternatively, DepthPro does high-resolution processing via a multi-scale pyramid design with late feature fusion[3].
2.3 Transformers in Optical Flow Estimation
Recent transformer-based optical flow methods differ mainly in which flow-specific inductive biases they retain. Some methods keep explicit matching as a central operation via correlation/cost-based similarity: UniMatch and TransFlow follow this direction [56, 26], while FlowFormer explicitly constructs and processes a 4D cost volume with a transformer-style decoder [11].
While others move closer to generic transformer formulations, they unfortunately still leave some biases in place. CroCo-Flow was among the first transformer approaches with competitive accuracy, leveraging binocular pretraining for dense matching, but it relies on slow dense tiling at high resolution [52]. Win-Win builds on CroCo to enable FullHD training and inference without tiling, but does not improve over CroCo-Flow in accuracy [20]. WAFT and GeoViT remove cost volumes but retain iterative warping-based updates (warping features and input images, respectively) [49, 54]. Overall, transformers have often been incorporated by incrementally replacing parts of established optical flow pipelines, which can improve performance yet further diversify and complicate the set of design choices. We summarize these choices in Tab. 1 using common inductive biases and compare to our bias-free approach.
3 Method
In this section, we present FreeFlow, an inductive-bias-free transformer for optical flow estimation (Fig. 3). We design this approach with two goals in mind. First, we aim to demonstrate that state-of-the-art optical flow can be achieved without the flow-specific architectural modules that dominate modern pipelines, such as correlation volumes, feature warping, iterative refinement, etc. (summarized in Tab. 1). Second, we seek a design that remains practical at high resolutions and ensures repeated information exchange across processing scales.
To this end, we study existing high-resolution prediction strategies, identify their limitations, and derive a principled alternative (Fig. 2). A straightforward solution is inference-time tiling in CroCo-/FlowFormer-style pipelines (Fig. 2a), but tiles interact only through late-stage averaging, so information does not flow across tile borders during feature extraction, and the number of forward passes grows with resolution. An alternative is multi-scale decoding with Late Feature Fusion as in DepthPro-like designs (Fig. 2b), which reduces the number of forward passes, but keeps information exchange delayed and tied to fixed-resolution assumptions, and is unreliable in the binocular setting when objects cross tile boundaries. FreeFlow resolves these issues by using a fixed and efficient tiling scheme while enabling dense feature exchange throughout the network (Fig. 2c), combining local processing with cross-tile and global interactions.
We next give a functional description of the architecture and then formalize the attention variants used by FreeFlow.
3.1 Approach
We adopt the CroCo/DUSt3R/MASt3R encoder–decoder high-level architecture for binocular reasoning, i.e., a Siamese ViT encoder followed by a decoder that alternates self- and cross-attention between the two views.
Given two input images , we embed each image into a sequence of non-overlapping patches (with ) using a standard patch projection, producing token sequences where . Both sequences are processed by a Siamese transformer encoder with shared weights to obtain feature representations
| (1) |
with .
A transformer decoder then produces flow tokens
| (2) |
where each decoder block combines self-attention over the current tokens with cross-attention from (queries) to (keys/values), enabling repeated information exchange between the two views. The output is reshaped into a spatial feature map and mapped to a dense flow and confidence field by a prediction head.
Attention Block Types.
To predict highly-detailed globally consistent flow fields, FreeFlow uses a hierarchical attention design that mixes local processing, cross-window exchange, and global context. Each block follows the CroCo-style transformer structure: encoder blocks apply self-attention, while decoder blocks additionally use cross-attention to utilize tokens from the other view. We use three attention variants (Fig. 4):
Window (Win) Attention Block. The token map of size is partitioned into a grid of non-overlapping windows, each containing tokens. Attention is computed independently within each window.
Shifted-Window (Swin) Attention Block. To enable information flow across window (and tile) boundaries, we apply a half-window shift by tokens horizontally and tokens vertically, partition into the same windows, perform window attention, and shift back.
Global Attention Block. To incorporate global context at controlled cost, we downsample by with a stride-2 convolution, apply full attention at the reduced resolution, and upsample by with a stride-2 transposed convolution. We apply normalization after the residual connection.
Each encoder/decoder layer applies the blocks in the fixed order: Win, Swin, and Global. Since attention cost scales quadratically with the number of tokens, the downsampling in the Global block keeps its attention cost comparable to Win/Swin at the original resolution. To encode token positional information within the image, we rely on Rotary Positional Embedding (RoPE) [40].
Attention Scale Factor.
Following prior work[7] on resolution-adaptive attention scaling, we multiply the attention logits by a logarithmic factor of the token count, which improves generalization when running inference at resolutions higher than those seen during training. For a token map of size we use
| (3) | ||||
| (4) |
Unlike prior formulations that normalize the factor to be at the training token count, we use the unnormalized variant and found it to work well in practice.
Flow Head.
Most high-performing optical flow pipelines rely on flow-specific prediction machinery, such as iterative update stages and convex upsampling. In addition, bias-free transformer baselines often use DPT-style heads from [33] that aggregate features from multiple layers (and, in CroCo-style designs, may also reuse encoder features) to form the final prediction. In contrast, FreeFlow predicts flow directly from the final decoded patch map using a simple three-layer head, without iterative refinement or convex upsampling.
Given , we apply a convolution to expand channels to , followed by a convolution to channels, and a transposed convolution with kernel and stride to upsample to . The head outputs , where the first two channels represent optical flow and the remaining three parameterize the uncertainty terms used by the mixture-of-Laplace loss (following SEA-RAFT [50]). Finally, we multiply the flow channels by (the patch size) to obtain flow in pixel units.
4 Experiments
We first detail our training pipeline, consisting of cross-view completion pretraining followed by finetuning for optical flow. We evaluate our method on three popular optical flow benchmarks: Spring [29] (high-resolution real-world sequences), Sintel [4] (synthetic scenes with complex motion and rendering effects), and KITTI-2015 [30] (real driving scenes). Finally, we provide ablations of key design choices, including the attention configuration, pretraining masking ratio, and model scaling.
| Stage | Weights | Datasets | Scale | Crop size | LR | WD | Batch | Steps |
| Pretrain | — | ARKitScenes [2] | 1x | [224, 224] | 8e-4 | 5e-2 | 2048 | 346k |
| MegaDepth [21] | ||||||||
| 3DStreetView [60] | ||||||||
| TaTSKH | Pretrain | TA+T+S+K+H | 2x | 10880 tok. | 4e-5 | 1e-2 | 32 | 450k |
| TaTSKH-hq | TaTSKH | TA+T+S+K+H | 2x | 32640 tok. | 1e-5 | 1e-5 | 32 | 90k |
| Sintel-ft | TaTSKH-hq | S | 2x | [872, 2048] | 1e-5 | 1e-5 | 32 | 12.5k |
| KITTI-ft | TaTSKH-hq | K | 2x | [750, 2484] | 1e-5 | 1e-5 | 32 | 2.5k |
| Spring-ft | TaTSKH-hq | Spring [29] | 1x | [1080, 1920] | 1e-5 | 1e-5 | 32 | 60k |
4.1 Training Details
Following CroCo [51, 52], we first pre-train our model on the cross-view completion task and then finetune the resulting weights with a new head for optical flow estimation. In cross-view completion, a large fraction of patches in one view is replaced by a learned token , and the model reconstructs the missing content conditioned on the second view. Due to its two-image nature, this objective encourages learning dense long-range correspondences, which is well aligned with the downstream task of binocular matching. Training details and datasets are summarized in Tab. 2; please refer to the supplementary for additional details.
Pre-train.
We follow the pre-training stage protocol from CroCo with minor adjustments. During cross-view completion training, a model takes as input two images that represent different views of the same scene: one view is severely masked, and the model has to predict the masked regions using the information from the second view. We sample image pairs from ARKitScenes [2], MegaDepth [21], and 3DStreetView [60], resulting in 3.7M data samples in total. We use fixed-size crops of 224224 resolution and pretrain the model for 346k steps.
The only significant difference from the CroCo setup is when the learned masked patch representation is introduced. In CroCo, masked tokens are removed from the first view and the corresponding tokens are added only at the decoder input. Here, due to the hierarchical nature of our model, we replace masked patches with at the encoder input, while keeping the completion objective unchanged.
Optical Flow Finetune.
The finetuning stage protocol is inspired by MEMFOF [1]; specifically, we adopt their upsampling of training frames, which better matches the motion distribution of FullHD inputs and improves high-resolution performance. Unlike MEMFOF and other curriculum-based training recipes that use multiple sequential stages, we use a single main dataset mixture, denoted TaTSKH in Tab. 2, for simplicity. For additional speed, we split this finetuning into a low- and high-token-count stage (TaTSKH and TaTSKH-hq in Tab. 2).
Instead of using a fixed crop size, to avoid unnecessary padding and to expose the model to a wider motion range, we use variable-resolution training with a fixed token budget per minibatch. More specifically, for each sample we randomly choose one spatial dimension (height or width), sample its value, and set the other dimension to the largest value such that the resulting token count does not exceed the prescribed budget. This produces rectangular crops with varying aspect ratios while keeping compute and memory controlled. The implementation is straightforward as we use a batch size of 1 sample/GPU during the finetuning stage. For benchmark submissions, we further finetune with fixed crop sizes (Sintel-ft, KITTI-ft, Spring-ft in Tab. 2). Following SEA-RAFT[50], we use the Mixture-of-Laplace loss.
In total, it takes from 4 to 5 days to pre-train and around 3 days to finetune our largest model on 32 GPUs.
4.2 Results
We adopt four widely used metrics from established benchmarks in this study: endpoint error (EPE), 1-pixel outlier rate (1px), Fl-score, and WAUC error. Please refer to [35, 4, 29, 31, 30] or the supplementary for their definitions.
| Method | Inf. Cost (1080p) | Params (M) | Spring | ||||
| Memory (GB) | Time (ms) | 1px | EPE | Fl | WAUC | ||
| FlowNet2 [13] | 3.01 | 110 | 162.52 | 6.710∗ | 1.040∗ | 2.823∗ | 90.907∗ |
| PWC-Net [42] | 0.57 | 48 | 9.37 | 82.265∗ | 2.288∗ | 4.889∗ | 45.670∗ |
| RAFT [45] | 7.97 | 406 | 5.26 | 6.790∗ | 1.476∗ | 3.198∗ | 90.920∗ |
| GMA [14] | 11.81 | 830 | 5.88 | 7.074∗ | 0.914∗ | 3.079∗ | 90.722∗ |
| GMFlow [55] | 8.22 | 8024 | 4.72 | 10.355∗ | 0.945∗ | 2.952∗ | 82.337∗ |
| FlowFormer [11] | 1.90 | 2084† | 16.17 | 6.510∗ | 0.723∗ | 2.384∗ | 91.679∗ |
| SEA-RAFT (M) [50] | 8.12 | 198 | 19.67 | 3.686 | 0.363 | 1.347 | 94.534 |
| DPFlow [31] | 4.26 | 401 | 10.02 | 3.442 | 0.340 | 1.311 | 94.980 |
| MemFlow(MF)[7] | 8.06 | 754 | 6.27 | 4.482 | 0.471 | 1.416 | 93.855 |
| StreamFlow(MF)[43] | 18.61 | 898 | 14.25 | 4.152 | 0.467 | 1.424 | 94.404 |
| MEMFOF(MF)[1] | 1.90 | 262 | 75.78 | 3.289 | 0.355 | 1.238 | 95.186 |
| ARFlow(MF)[23] | — | — | 76.50 | 3.265 | 0.353 | 1.212 | 95.283 |
| CroCo-Flow [52] | 2.73 | 3266 | 447.47 | 4.565 | 0.498 | 1.508 | 93.660 |
| Win-Win [20] | 3.82† | 305† | 229.77† | 5.371 | 0.475 | 1.621 | 92.720 |
| WAFT-DAv2-a2 [49] | 20.58 | 489 | 56.93 | 3.298 | 0.304 | 1.197 | 94.990 |
| WAFT-DINOv3-a2 [49] | 18.96 | 408 | 56.47 | 3.182 | 0.325 | 1.246 | 95.051 |
| FreeFlow-S (ours) | 1.02 | 144 | 34.58 | 5.087 | 0.533 | 1.452 | 90.196 |
| FreeFlow-M (ours) | 1.66 | 325 | 102.46 | 3.392 | 0.346 | 1.171 | 94.919 |
| FreeFlow-L (ours) | 2.58 | 607 | 230.72 | 3.192 | 0.278 | 1.048 | 95.235 |
Results on Spring.
FreeFlow achieves state-of-the-art performance on Spring. FreeFlow-L sets the best EPE and Fl among all compared methods (Tab. 3), while remaining on par with the strongest approaches in 1px and WAUC; in particular, it attains the best WAUC among two-frame methods. Compared to WAFT-DAv2-a2, FreeFlow-L improves EPE by 9% and reduces Fl by 14%. Owing to tiling-free native 1080p inference, FreeFlow preserves fine detail while maintaining global motion consistency (Fig. 5). We further highlight that native 1080p processing is possible within a low inference memory budget (Fig. 1).
| Method | Sintel | KITTI-15 | |
| Clean | Final | Fl-all | |
| FlowNet2 [13] | 4.16 | 5.74 | 10.41 |
| PWC-Net [42] | 3.90 | 5.04 | 9.60 |
| RAFT [45] | 1.61 | 2.86 | 5.10 |
| GMA [14] | 1.39 | 2.47 | 5.15 |
| GMFlow+ [56] | 1.03 | 2.37 | 4.49 |
| FlowFormer [11] | 1.16 | 2.09 | 4.68 |
| FlowFormer++ [39] | 1.07 | 1.94 | 4.52 |
| TransFlow [26] | 1.06 | 2.08 | 4.32 |
| SEA-RAFT (L) [50] | 1.31 | 2.60 | 4.30 |
| DPFlow [31] | 1.05 | 1.98 | 3.56 |
| VideoFlow-BOF(MF)[38] | 1.00 | 1.71 | 4.44 |
| VideoFlow-MOF(MF)[38] | 0.99 | 1.65 | 3.65 |
| MEMFOF(MF)[1] | 0.99 | 1.94 | 2.94 |
| MEMFOF-XL(MF)[1] | 0.93 | 1.89 | — |
| ARFlow(MF)[23] | 0.96 | 1.79 | 2.85 |
| DDVM[37] | 1.75 | 2.48 | 3.26 |
| CroCo-Flow[52] | 1.09 | 2.44 | 3.64 |
| Win-Win[20] | 1.15 | 2.34 | — |
| WAFT-DAv2-a2[49] | 0.94 | 2.33 | 3.31 |
| WAFT-DINOv3-a2[49] | 0.95 | 2.02 | 3.56 |
| GeoVIT [54] | 0.79 | 1.88 | 3.79 |
| FreeFlow-S (ours) | 1.03 | 1.99 | 4.06 |
| FreeFlow-M (ours) | 0.80 | 1.77 | 3.33 |
| FreeFlow-L (ours) | 0.68 | 1.48 | 3.23 |
Results on Sintel and KITTI.
Following MEMFOF, we finetune on upsampled frames. Accordingly, for Sintel and KITTI submissions we upscale input images by and downscale the predicted flow by . FreeFlow-L ranks first on Sintel on both Clean and Final (Tab. 4), improving over GeoVIT [54] by 14% on Clean (0.790.68) and 10% over VideoFlow-MOF [38] on Final (1.651.48). Notably, FreeFlow-M is already highly competitive: it is second only to FreeFlow-L on Clean, and ranks fourth on Final, surpassed only by the 3- and 5-frame VideoFlow variants. On KITTI-2015, FreeFlow-L achieves 3.23 Fl-all, outperforming all non-stereo and non-multiframe methods on KITTI-15 (Tab. 4). Qualitative results show that the model captures complex motion patterns using only two frames (Fig. 6). Additional visual comparisons and zero-shot evaluations are provided in the supplementary material.
| Mask. Ratio | Subblock type | Spring (sub-val) | |||
| Win | Swin | Global | 1px | EPE | |
| 0.9 | 1.133 | 0.229 | |||
| 0.9 | 0.801 | 0.190 | |||
| 0.9 | 0.659 | 0.167 | |||
| 0.9 | 0.688 | 0.170 | |||
| 0.925 | 0.685 | 0.166 | |||
| 0.95 | 0.658 | 0.166 | |||
| 0.95 | 0.624 | 0.157 | |||
| 0.975 | 0.656 | 0.174 | |||
4.3 Ablation Study
Unless stated otherwise, all ablations use a scaled-down FreeFlow configuration with 8 encoder and 8 decoder layers, width 256, and 8 attention heads, and are trained with the same training recipe. Following SEA-RAFT and WAFT, we report results on the Spring sub-validation split (scenes 0045 and 0047) after finetuning on the remaining training data. Additional experiments in the supplementary isolate the architecture from the training procedure and study the effect of adding iterative flow-specific biases back into FreeFlow.
Architecture and Masking Ratio Ablation.
We study the interaction between the cross-view completion masking ratio and the hierarchical attention design of FreeFlow (Tab. 5). The masking ratio controls the fraction of patches in the target view that are replaced by the learned token during pretraining, while the attention configuration determines which subblock types (Win/Swin/Global) are present in the encoder and decoder.
With all three subblocks enabled, a masking ratio of 0.95 performs best, improving over the CroCo default 0.9 by 9.3% in 1px and 7.6% in EPE. This is consistent with FreeFlow using smaller patches than CroCo ( vs. ), which reduces the distance to visible regions and makes a higher mask rate beneficial as it makes the pretraining task sufficiently challenging.
We also observe an interaction between masking ratio and attention configuration. At the CroCo default ratio (0.9), removing Global does not hurt and can even slightly improve the sub-validation metrics, whereas at 0.95 Global becomes important for the best performance. This indicates that the masking ratio is an important hyperparameter that should be chosen jointly with the model architecture. Shifted-Window attention consistently contributes to accuracy, and removing it leads to a clear drop on this split. Additionally, for both masking ratios, qualitative comparisons (Fig. 7) show that removing the Global block can introduce obvious motion inconsistencies, even when the sub-validation metrics change only marginally.
| Name | Patch Sizes | Width | Attn. Count | Attn. Heads | MLP Count | Params (M) |
| GeoViT | 16 | 1024 | 246 | 16 | 246 | 377 |
| WAFT | 16 | 384384 | 12125 | 66 | 12+125 | 56 |
| CroCo-Flow | 16 | 1024768 | 2424 | 1612 | 2412 | 447 |
| Win-Win | 16 | 768768 | 1224 | 12 | 1212 | 230 |
| GMFlow+ | 8/4 | 0128 | 0122 | 01 | 062 | 4.7 |
| FreeFlow-S | 8 | 256256 | 1224 | 4+4 | 1212 | 35 |
| FreeFlow-M | 8 | 384384 | 1836 | 6+6 | 1818 | 102 |
| FreeFlow-L | 8 | 512512 | 2448 | 8+8 | 2424 | 231 |
Model Scaling.
A benefit of FreeFlow’s simple, uniform encoder–decoder design is that it can be scaled in a straightforward manner by adjusting depth and width. We therefore study how performance evolves across model sizes (S/M/L) on common optical flow benchmarks. We follow common ViT scaling [61] best practices by co-scaling depth, width, and the number of attention heads while keeping the overall architecture and patch size fixed. The resulting configurations are summarized in Tab. 6.
As shown in Fig. 8, FreeFlow scales predictably: larger variants yield consistent accuracy gains, while smaller variants retain strong performance, indicating that the approach remains effective even when scaled down. Memory scaling with model size is reported in Fig. 1(a).
5 Conclusion
In this work, we introduced FreeFlow, a bias-free hierarchical transformer for optical flow estimation that removes conventional flow-specific design components such as correlation volumes, feature warping, and iterative refinement. Instead, FreeFlow relies on a simple feed-forward encoder–decoder architecture that combines window, shifted-window, and reduced-resolution global attention to capture motion across multiple spatial scales. This design provides a flexible and scalable framework that improves consistently with model capacity while maintaining efficient high-resolution inference.
We show that these standard optical-flow inductive biases are not required to achieve top performance. FreeFlow reaches state-of-the-art results on all popular benchmarks, including Sintel, KITTI-2015, and Spring, demonstrating that strong motion estimation can be obtained from a general-purpose transformer architecture without specialized flow modules. We hope that this work encourages further exploration of simpler and more general architectures for motion estimation and related dense correspondence tasks.
Acknowledgements
The work of Vladislav Bargatin, Alexander Yakovenko and Khaled Abud was supported by the The Ministry of Economic Development of the RussianFederation in accordance with the subsidy agreement (agreement identifier000000C313925P4H0002; grant No 139-15-2025-012). The research was carried out using the MSU-270 supercomputer of Lomonosov Moscow State University.
References
- [1] V. Bargatin, E. Chistov, A. Yakovenko, and D. Vatolin (2025) MEMFOF: high-resolution training for memory-efficient multi-frame optical flow estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8187–8196. External Links: Document Cited by: §1, §1, §2.1, Table 1, Figure 6, Figure 6, §4.1, Table 2, Table 2, Table 3, Table 4, Table 4.
- [2] G. Baruch, Z. Chen, A. Dehghan, Y. Feigin, P. Fu, T. Gebauer, D. Kurz, T. Dimry, B. Joffe, A. Schwartz, and E. Shulman (2021) ARKitScenes: a diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Vol. 1, pp. . Cited by: §4.1, Table 2.
- [3] A. Bochkovskiy, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. Richter, and V. Koltun (2025) Depth Pro: sharp monocular metric depth in less than a second. In International Conference on Learning Representations, Vol. 2025, pp. 75602–75637. Cited by: §1, §2.2.
- [4] D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black (2012) A naturalistic open source movie for optical flow evaluation. In Computer Vision – ECCV 2012, Berlin, Heidelberg, pp. 611–625. External Links: ISBN 978-3-642-33783-3, Document Cited by: §1, Figure 6, Figure 6, §4.2, Table 2, Table 2, Table 4, Table 4, §4.
- [5] N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko (2020) End-to-end object detection with transformers. In Computer Vision – ECCV 2020, Cham, pp. 213–229. External Links: ISBN 978-3-030-58452-8, Document Cited by: §1.
- [6] K. C.K. Chan, X. Wang, K. Yu, C. Dong, and C. C. Loy (2021) BasicVSR: the search for essential components in video super-resolution and beyond. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4945–4954. External Links: Document Cited by: §1.
- [7] Q. Dong and Y. Fu (2024) MemFlow: optical flow estimation and prediction with memory. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 19068–19078. External Links: Document Cited by: §1, §3.1, Table 3.
- [8] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations, Cited by: §1, §2.2.
- [9] A. Dosovitskiy, P. Fischer, E. Ilg, P. Häusser, C. Hazirbas, V. Golkov, P. v. d. Smagt, D. Cremers, and T. Brox (2015) FlowNet: learning optical flow with convolutional networks. In 2015 IEEE International Conference on Computer Vision (ICCV), Vol. , pp. 2758–2766. External Links: Document Cited by: §1, §2.1.
- [10] B. K.P. Horn and B. G. Schunck (1981) Determining optical flow. Artificial Intelligence 17 (1-3), pp. 185–203. External Links: ISSN 0004-3702, Document Cited by: §1, §2.1.
- [11] Z. Huang, X. Shi, C. Zhang, Q. Wang, K. C. Cheung, H. Qin, J. Dai, and H. Li (2022) FlowFormer: a transformer architecture for optical flow. In Computer Vision – ECCV 2022, Cham, pp. 668–685. External Links: ISBN 978-3-031-19790-1, Document Cited by: §2.3, Table 1, Table 3, Table 4.
- [12] Z. Huang, T. Zhang, W. Heng, B. Shi, and S. Zhou (2022) Real-time intermediate flow estimation for video frame interpolation. In Computer Vision – ECCV 2022, Cham, pp. 624–642. External Links: ISBN 978-3-031-19781-9, Document Cited by: §1.
- [13] E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox (2017) FlowNet 2.0: evolution of optical flow estimation with deep networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 1647–1655. External Links: Document Cited by: Table 3, Table 4.
- [14] S. Jiang, D. Campbell, Y. Lu, H. Li, and R. Hartley (2021) Learning to estimate hidden motions with global motion aggregation. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 9752–9761. External Links: Document Cited by: §1, Table 3, Table 4.
- [15] H. Jin, H. Jiang, H. Tan, K. Zhang, S. Bi, T. Zhang, F. Luan, N. Snavely, and Z. Xu (2025) LVSM: a large view synthesis model with minimal 3d inductive bias. In International Conference on Learning Representations, Vol. 2025, pp. 60001–60021. Cited by: §1, §2.2.
- [16] T. Kerssies, N. Cavagnero, A. Hermans, N. Norouzi, G. Averta, B. Leibe, G. Dubbelman, and D. De Geus (2025) Your vit is secretly an image segmentation model. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 25303–25313. External Links: Document Cited by: §1, §2.2.
- [17] S. Kiefhaber, S. Roth, and S. Schaub-Meyer (2025) Removing cost volumes from optical flow estimators. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 79–89. External Links: Document Cited by: §2.1.
- [18] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, P. Dollár, and R. Girshick (2023) Segment anything. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 3992–4003. External Links: Document Cited by: §1, §2.2.
- [19] D. Kondermann, R. Nair, K. Honauer, K. Krispin, J. Andrulis, A. Brock, B. Güssefeld, M. Rahimimoghaddam, S. Hofmann, C. Brenner, and B. Jähne (2016) The HCI Benchmark Suite: stereo and flow ground truth with uncertainties for urban autonomous driving. In 2016 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vol. , pp. 19–28. External Links: Document Cited by: Table 2, Table 2.
- [20] V. Leroy, J. Revaud, T. Lucas, and P. Weinzaepfel (2024) Win-Win: training high-resolution vision transformers from two windows. In International Conference on Learning Representations, Vol. 2024, pp. 48749–48767. Cited by: §2.3, Table 1, Figure 5, Figure 5, Figure 6, Figure 6, Table 3, Table 4.
- [21] Z. Li and N. Snavely (2018) MegaDepth: learning single-view depth prediction from internet photos. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , pp. 2041–2050. External Links: Document Cited by: §4.1, Table 2.
- [22] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte (2021) SwinIR: image restoration using swin transformer. In 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Vol. , pp. 1833–1844. External Links: Document Cited by: §2.2.
- [23] J. Liu, M. Liu, S. Zhu, Y. Zhang, J. Li, M. Y. Yang, F. Nex, H. Cheng, and H. Wang (2026) ARFlow: auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting. In International Conference on Learning Representations, Vol. 2026. Cited by: Table 3, Table 4.
- [24] X. Liu, H. Liu, and Y. Lin (2020) Video frame interpolation via optical flow estimation with image inpainting. International Journal of Intelligent Systems 35 (12), pp. 2087–2102. External Links: Document Cited by: §1.
- [25] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin Transformer: hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 9992–10002. External Links: Document Cited by: §2.2.
- [26] Y. Lu, Q. Wang, S. Ma, T. Geng, Y. V. Chen, H. Chen, and D. Liu (2023) TransFlow: transformer as flow learner. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 18063–18073. External Links: Document Cited by: §2.3, Table 1, Table 4.
- [27] B. D. Lucas and T. Kanade (1981) An Iterative Image Registration Technique with an Application to Stereo Vision. In IJCAI’81: 7th international joint conference on Artificial intelligence, Vol. 2, Vancouver, Canada, pp. 674–679. Cited by: §1, §2.1.
- [28] N. Mayer, E. Ilg, P. Häusser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox (2016) A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4040–4048. External Links: Document Cited by: Table 2, Table 2.
- [29] L. Mehl, J. Schmalfuss, A. Jahedi, Y. Nalivayko, and A. Bruhn (2023) Spring: a high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4981–4991. External Links: Document Cited by: §1, Figure 5, Figure 5, §4.2, Table 2, §4.
- [30] M. Menze and A. Geiger (2015) Object scene flow for autonomous vehicles. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 3061–3070. External Links: Document Cited by: §1, §4.2, Table 2, Table 2, Table 4, Table 4, §4.
- [31] H. Morimitsu, X. Zhu, R. M. Cesar, X. Ji, and X. Yin (2025) DPFlow: adaptive optical flow estimation with a dual-pyramid framework. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 17810–17820. External Links: Document Cited by: §2.1, Table 1, §4.2, Table 3, Table 4.
- [32] A. Piergiovanni and M. S. Ryoo (2019) Representation flow for action recognition. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 9937–9945. External Links: Document Cited by: §1.
- [33] R. Ranftl, A. Bochkovskiy, and V. Koltun (2021) Vision transformers for dense prediction. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 12159–12168. External Links: Document Cited by: §2.2, Table 1, Table 1, §3.1.
- [34] N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C. Wu, R. Girshick, P. Dollar, and C. Feichtenhofer (2025) SAM 2: segment anything in images and videos. In International Conference on Learning Representations, Vol. 2025, pp. 28085–28128. Cited by: §2.2.
- [35] S. R. Richter, Z. Hayder, and V. Koltun (2017) Playing for benchmarks. In 2017 IEEE International Conference on Computer Vision (ICCV), Vol. , pp. 2232–2241. External Links: Document Cited by: §4.2.
- [36] C. Ryali, Y. Hu, D. Bolya, C. Wei, H. Fan, P. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffman, J. Malik, Y. Li, and C. Feichtenhofer (2023) Hiera: a hierarchical vision transformer without the bells-and-whistles. In Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, pp. 29441–29454. Cited by: §2.2.
- [37] S. Saxena, C. Herrmann, J. Hur, A. Kar, M. Norouzi, D. Sun, and D. J. Fleet (2023) The surprising effectiveness of diffusion models for optical flow and monocular depth estimation. In Advances in Neural Information Processing Systems, Vol. 36, pp. 39443–39469. Cited by: §1, §2.1, Table 4.
- [38] X. Shi, Z. Huang, W. Bian, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li (2023) VideoFlow: exploiting temporal cues for multi-frame optical flow estimation. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 12435–12446. External Links: Document Cited by: §1, §4.2, Table 4, Table 4.
- [39] X. Shi, Z. Huang, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li (2023) FlowFormer++: masked cost volume autoencoding for pretraining optical flow estimation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 1599–1610. External Links: Document Cited by: Table 4.
- [40] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu (2024) RoFormer: enhanced transformer with rotary position embedding. Neurocomputing 568, pp. 127063. External Links: ISSN 0925-2312, Document Cited by: §3.1.
- [41] D. Sun, S. Roth, and M. J. Black (2010) Secrets of optical flow estimation and their principles. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Vol. , pp. 2432–2439. External Links: Document Cited by: §1, §2.1.
- [42] D. Sun, X. Yang, M. Liu, and J. Kautz (2018) PWC-Net: cnns for optical flow using pyramid, warping, and cost volume. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , pp. 8934–8943. External Links: Document Cited by: §1, §2.1, Table 1, Table 3, Table 4.
- [43] S. Sun, J. Liu, H. Li, G. Liu, T. H. Li, and W. Gao (2024) StreamFlow: streamlined multi-frame optical flow estimation for video sequences. In Advances in Neural Information Processing Systems, Vol. 37, pp. 9205–9228. External Links: Document Cited by: Table 3.
- [44] S. Sun, Z. Kuang, L. Sheng, W. Ouyang, and W. Zhang (2018) Optical flow guided feature: a fast and robust motion representation for video action recognition. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , pp. 1390–1399. External Links: Document Cited by: §1.
- [45] Z. Teed and J. Deng (2020) RAFT: recurrent all-pairs field transforms for optical flow. In Computer Vision – ECCV 2020, Cham, pp. 402–419. External Links: ISBN 978-3-030-58536-5, Document Cited by: §1, §2.1, Table 1, Table 3, Table 4.
- [46] J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny (2025) VGGT: visual geometry grounded transformer. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 5294–5306. External Links: Document Cited by: §1, §2.2.
- [47] S. Wang, V. Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud (2024) DUSt3R: geometric 3d vision made easy. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 20697–20709. External Links: Document Cited by: §1, §2.2.
- [48] W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer (2020) TartanAir: a dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vol. , pp. 4909–4916. External Links: Document Cited by: Table 2, Table 2.
- [49] Y. Wang and J. Deng (2026) WAFT: warping-alone field transforms for optical flow. In International Conference on Learning Representations, Vol. 2026. Cited by: §1, §2.1, §2.3, Table 1, Figure 5, Figure 5, Figure 6, Figure 6, Table 3, Table 3, Table 4, Table 4.
- [50] Y. Wang, L. Lipson, and J. Deng (2025) SEA-RAFT: simple, efficient, accurate raft for optical flow. In Computer Vision – ECCV 2024, Cham, pp. 36–54. External Links: ISBN 978-3-031-72667-5, Document Cited by: §1, §2.1, §3.1, §4.1, Table 2, Table 2, Table 3, Table 4.
- [51] P. Weinzaepfel, V. Leroy, T. Lucas, R. Brégier, Y. Cabon, V. Arora, L. Antsfeld, B. Chidlovskii, G. Csurka, and J. Revaud (2022) CroCo: self-supervised pre-training for 3d vision tasks by cross-view completion. In Advances in Neural Information Processing Systems, Vol. 35, pp. 3502–3516. Cited by: §4.1.
- [52] P. Weinzaepfel, T. Lucas, V. Leroy, Y. Cabon, V. Arora, R. Brégier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud (2023) CroCo v2: improved cross-view completion pre-training for stereo matching and optical flow. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 17923–17934. External Links: Document Cited by: §1, §2.3, Table 1, Figure 5, Figure 5, Figure 6, Figure 6, §4.1, Table 3, Table 4.
- [53] P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid (2013) DeepFlow: large displacement optical flow with deep matching. In 2013 IEEE International Conference on Computer Vision, Vol. , pp. 1385–1392. External Links: Document Cited by: §1, §2.1.
- [54] H. Wu, K. Cheng, S. Lin, and Z. Wu (2026) A study of finetuning video transformers for multi-view geometry tasks. Proceedings of the AAAI Conference on Artificial Intelligence. External Links: Document Cited by: §1, §2.1, §2.3, Table 1, Figure 6, Figure 6, §4.2, Table 4.
- [55] H. Xu, J. Zhang, J. Cai, H. Rezatofighi, and D. Tao (2022) GMFlow: learning optical flow via global matching. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8111–8120. External Links: Document Cited by: Table 3.
- [56] H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger (2023) Unifying flow, stereo and depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), pp. 13941–13958. External Links: Document Cited by: §2.3, Table 1, Table 4.
- [57] X. Xu, L. Siyao, W. Sun, Q. Yin, and M. Yang (2019) Quadratic video interpolation. In Advances in Neural Information Processing Systems, Vol. 32, pp. 1645–1654. Cited by: §1.
- [58] L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao (2024) Depth Anything: unleashing the power of large-scale unlabeled data. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 10371–10381. External Links: Document Cited by: §1, §2.2.
- [59] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao (2024) Depth Anything V2. In Advances in Neural Information Processing Systems, Vol. 37, pp. 21875–21911. External Links: Document Cited by: §1, §2.2.
- [60] A. R. Zamir, T. Wekel, P. Agrawal, C. Wei, J. Malik, and S. Savarese (2016) Generic 3d representation via pose estimation and matching. In Computer Vision – ECCV 2016, Cham, pp. 535–553. External Links: ISBN 978-3-319-46487-9, Document Cited by: §4.1, Table 2.
- [61] X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer (2022) Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12104–12113. External Links: Document Cited by: §4.3.
- [62] Y. Zhao, K. L. Man, J. Smith, K. Siddique, and S. Guan (2020) Improved two-stream model for human action recognition. EURASIP Journal on Image and Video Processing 2020 (1), pp. 24. External Links: ISSN 1687-5281, Document Cited by: §1.
FreeFlow: A Bias-free Hierarchical
Transformer for Optical Flow Estimation Supplementary Material
This supplementary provides additional details and context on training and evaluation protocols, metric and loss definitions, and extended qualitative and ablation results, and is organized as follows:
Appendix 0.A formally defines evaluation metrics and the loss function;
Appendix 0.B discusses Dense Feature Fusion and the effectiveness of FreeFlow;
Appendix 0.C provides additional ablations and results;
Appendix 0.D shows more qualitative examples.
Appendix 0.A Definitions
In the following sections, is the target flow vector at position , is the predicted flow vector at position , is the number of valid pixels in the target flow field, and is the Iverson bracket. Sums of the form are calculated only over valid pixels.
0.A.1 Metrics
Endpoint Error.
Endpoint error (EPE) is defined as:
| EPE | (5) |
and ranges from at worst to at best. It is adopted as the main metric for the Sintel benchmark.
One-pixel Outlier Rate.
The 1-pixel outlier rate (1px) is defined as:
| 1px | (6) |
and ranges from at worst to at best. It is adopted as the main metric for the Spring benchmark.
Fl-all Outliers Metric.
The Fl-all outliers metric (Fl-all score) is defined as:
| Fl-all | (7) |
and ranges from at worst to at best. Is is adopted as the main metric for the KITTI-15 benchmark due to the noisy nature of the collected real-world data.
Weighted Area Under the Curve.
The weighted area under the curve (WAUC) is formally defined as:
| WAUC | (8) |
and ranges from at worst to at best. In practice, this integral is usually approximated with bins. WAUC can be also viewed as a generalization of 1px score.
0.A.2 Mixture-of-Laplace Loss
For a single flow vector coordinate, the Mixture-of-Laplace (MoL) in SEA-RAFT is defined as:
| (9) |
where is the target flow coordinate, is the predicted flow coordinate, is the predicted mixing coefficient, and is the predicted scale parameter, clamped to the range . In practice and as is used in the official implementation, and are calculated as the softmax of two outputs and . The final MoL loss is then defined as:
| (10) |


Appendix 0.B Architecture Discussion
To isolate the role of feature fusion, we implemented a DepthPro-style Late Feature Fusion baseline by replacing the monocular ViT backbone with a pretrained CroCo encoder–decoder and fusing pyramid features in a DPT-decoder (which we call CroCo-Pro). As shown in Fig. 9(a), this design does not reliably propagate information across scales and tile boundaries: coarse-scale cues are weakly utilized and have limited influence on the final full-resolution prediction. The resulting flow can preserve fine local structure, yet exhibits reduced global coherence, especially when displacements span multiple tiles or when texture is ambiguous.
FreeFlow addresses this limitation by exchanging features throughout the network rather than only at the final decoding stage (Fig. 9(b)). Shifted-window blocks provide repeated cross-tile communication, while reduced-resolution global attention injects long-range context, allowing coarse and fine signals to reinforce each other during decoding. Beyond accuracy, this design is practical at high resolution: all interaction mechanisms are expressed via standard attention operators (windowed, shifted-window, and global attention on a downsampled map), whose token counts are balanced so that their computational cost is comparable. As a result, FreeFlow directly benefits from off-the-shelf efficient attention kernels that have been heavily optimized in recent years, whereas flow-specific modules such as correlation volume indexing, feature warping, or iterative update pipelines require task-specific engineering to reach similar efficiency.
Appendix 0.C Additional results
0.C.1 Training Details
We use AdamW with and . For pretraining, we adopt a linear warmup followed by a cosine learning-rate decay. For optical flow finetuning, we use a linear warmup followed by a linear decay schedule, which is standard in optical flow training. We train with the Mixture-of-Laplace (Sec. 0.A.2) loss from SEA-RAFT.
| Patch size | Pre-train crop | Mask. ratio | Width | Inf. Cost (1080p) | Params (M) | Spring (sub-val) | ||
| Memory (GB) | Time (ms) | 1px | EPE | |||||
| 88 | [224, 224] | 0.95 | 256 | 1.02 | 144 | 34.58 | 0.181 | 0.709 |
| 1616 | [256, 256] | 0.9 | 768 | 2.91 | 120 | 278.88 | 0.174 | 0.798 |
0.C.2 Patch Size Ablation
We investigate the role of the patch size on our method’s performance. As the base model, we take FreeFlow-S. For the model, to make for a fair comparision, we increase the width from to and the number of attention heads from to , resulting in similar execution time.
Tab. 7 shows that the variant is only marginally better in EPE (0.174 vs. 0.181), while being worse in 1px (0.798 vs. 0.709) and substantially more expensive in parameters and memory (279M and 2.91 GB vs. 35M and 1.02 GB). Overall, this supports using smaller patches in our setting, as they provide comparable accuracy at a substantially lower memory and parameter cost.
0.C.3 Comparison to Vanilla ViT
To study the performance of the proposed architecture separately from the used training procedure, we trained a variant in which FreeFlow encoder/decoder modules are substituted for vanilla ViTs. This model closely resembles Croco-Flow’s and Win-Win’s designs (except for the DPT-head used in the postprocessing stage). The results are presented in Tab. 8 (rows 1–3). Even within a larger computational budget (270–410 ms vs 131 ms), the ViT-based model results in a larger prediction error, while the quality difference between the ViT and FreeFlow models with similar parameter counts is negligible (with the ViT model being almost 6 times slower). This confirms that the proposed Local-Global attention design, not the training recipe alone, improves on the quality-performance tradeoff of a standard ViT-based approach.
0.C.4 Injecting Biases back into FreeFlow
We modified FreeFlow to include explicit optical flow biases, testing a GeoViT-like iterative warping procedure as it is straightforward to integrate and currently among the most effective iterative transformer-based approaches. We do not change the model during pretraining, only introducing warping and the recursive module (the same ConvGRU as in GeoViT) during the optical flow fine-tuning stage. Results are provided in Tab. 8 (rows 6–8). The GeoViT-like variant shows slightly better prediction quality within the same parameter budget in some configurations, but at the cost of slightly slower inference due to its iterative nature. This is consistent with expectations: the inductive bias offloads task-specific knowledge into an explicit operator, increasing the effective capacity of the model. This further illustrates that flow-specific biases are not necessary to reach SOTA performance, but can be added to FreeFlow to improve accuracy at the cost of speed and architectural generality.
| Type | Layers | Iters | Time (ms) (1080p) | Params (M) | Spring (sub-val) | |
| 1px | EPE | |||||
| ViT | 4 | — | 278+13 | 15.36 | 1.021 | 0.217 |
| 6 | — | 415+13 | 19.05 | 0.809 | 0.216 | |
| 12 | — | 829+13 | 30.11 | 0.698 | 0.192 | |
| FreeFlow-S | 4 | — | 131+13 | 34.58 | 0.709 | 0.181 |
| FreeFlow-S+ | 8 | — | 256+13 | 60.90 | 0.624 | 0.157 |
| FreeFlow-S with GeoViT warping | 4 | 1 | 131+40 | 33.12 | 0.660 | 0.165 |
| 2 | 2 | 131+62 | 19.96 | 0.719 | 0.171 | |
| 4 | 2 | 262+53 | 33.12 | 0.583 | 0.148 | |
| Method | Biases | Pre-train | TA | Sintel (train) | KITTI-15 (train) | |||||
| IR | CV | W | Name | Size | Clean | Final | Fl-epe | Fl-all | ||
| RAFT | — | — | 1.43 | 2.71 | 5.04 | 17.4 | ||||
| GMA | — | — | 1.30 | 2.74 | 4.69 | 17.1 | ||||
| FlowFormer | ImageNet-1K | 1.3M | 1.01 | 2.40 | 4.09 | 14.7 | ||||
| SEA-RAFT (S) | ImageNet-1K | 1.3M | 1.27 | 3.74 | 4.43 | 15.1 | ||||
| SEA-RAFT (M) | ImageNet-1K | 1.3M | 1.21 | 4.04 | 4.29 | 14.2 | ||||
| SEA-RAFT (L) | ImageNet-1K | 1.3M | 1.19 | 4.11 | 3.62 | 12.9 | ||||
| DPFlow | — | — | 1.02 | 2.26 | 3.37 | 11.1 | ||||
| VideoFlow-BOF(MF) | ImageNet-1K | 1.3M | 1.03 | 2.19 | 3.96 | 15.3 | ||||
| VideoFlow-MOF(MF) | ImageNet-1K | 1.3M | 1.18 | 2.56 | 3.89 | 14.2 | ||||
| MemFlow(MF) | — | — | 0.93 | 2.08 | 3.88 | 13.7 | ||||
| MemFlow-T(MF) | ImageNet-1K | 1.3M | 0.85 | 2.06 | 3.38 | 12.8 | ||||
| StreamFlow(MF) | ImageNet-1K | 1.3M | 0.87 | 2.11 | 3.85 | 12.6 | ||||
| MEMFOF(MF) | ImageNet-1K | 1.3M | 1.10 | 2.70 | 3.31 | 10.1 | ||||
| MEMFOF(MF) | ImageNet-1K | 1.3M | 1.20 | 3.91 | 2.93 | 9.9 | ||||
| ARFlow(MF) | ImageNet-1K | 1.3M | 0.88 | 2.07 | 2.86 | 9.2 | ||||
| WAFT-Twins-a2 | ImageNet-1K | 1.3M | 1.02 | 2.46 | 2.98 | 9.9 | ||||
| WAFT-DAv2-a2 | DAv2 | 63M | 1.01 | 2.49 | 3.28 | 10.9 | ||||
| WAFT-DINOv3-a2 | LVD-1689M | 1.7B | 1.28 | 2.56 | 3.49 | 12.9 | ||||
| GeoViT | Kinetics-400 | 59M | 0.69 | 1.78 | 3.15 | 11.5 | ||||
| CroCo-Flow | CroCo v2 | 15M | 1.28 | 2.58 | — | — | ||||
| FreeFlow-S (ours) | CroCo v2 | 7.4M | 0.91 | 3.16 | 3.41 | 10.4 | ||||
| FreeFlow-M (ours) | CroCo v2 | 7.4M | 1.01 | 3.12 | 5.89 | 14.6 | ||||
| FreeFlow-L (ours) | CroCo v2 | 7.4M | 1.04 | 2.30 | 4.77 | 12.9 | ||||
0.C.5 Zero-shot Performance
We evaluate zero-shot transfer on the Sintel and KITTI training sets by finetuning only on TartanAir (TA) and FlyingThings3D. Specifically, we remove the Sintel, KITTI, and HD1K portions from the TaTSKH / TaTSKH-hq stages and report performance after completing the “hq” stage (Tab. 9).
FreeFlow attains reasonable zero-shot performance on Sintel and KITTI under TA+Things finetuning, but does not improve monotonically with model size. This behavior is expected for a bias-free model in a data-limited regime: transfer is primarily constrained by the motion and appearance coverage of the available training signal, and increasing capacity alone does not guarantee gains.
This trend is reflected by the methods that use broader pretraining datasets. GeoViT, pretrained on the large Kinetics-400 video dataset, achieves the strongest zero-shot performance on Sintel, consistent with exposure to substantially more varied motion. In contrast, CroCo-Flow, pretrained only with CroCo-style data and without the use of TA during training, exhibits weaker transfer. Finally, the WAFT variants suggest that large-scale monocular pretraining alone is not always sufficient for zero-shot optical flow: despite substantially larger pretraining datasets, their transfer does not match methods pretrained on binocular or video data, indicating the importance of multi-view and motion-centric pretraining signals.
0.C.6 Additional Model Analysis
We show visualizations for different attention types in Fig. 10: global attention helps FreeFlow capture large displacements, while local attention processes small shifts. On shifted-window attention in the decoder, the window partition is shifted identically in both frames so that corresponding regions remain co-located across the pair.
Appendix 0.D Additional Qualitative Comparisons
We provide additional qualitative samples for Sintel (Fig. 11) and KITTI-15 (Fig. 12) datasets.