跳到正文
HuggingFace Daily Papers(社区热门论文)
48AI 编辑部评分,满分 100

FreeFlow:无光流专用偏置的层次化 Transformer,刷新 Sintel/KITTI-2015/Spring 基准

2026-09-10 08:00· 1天前
AI 导读

莫斯科国立大学团队提出 FreeFlow,一个不含相关体积、特征 warping 等光流专用组件的层次化 Transformer,仅用单一前馈编码器-解码器,组合窗口注意力、移位窗口注意力和低分辨率全局注意力。

Vladislav Bargatin

AI Center, Lomonosov MSU, Moscow, Russia

Alexander Yakovenko

AI Center, Lomonosov MSU, Moscow, Russia

Lomonosov Moscow State University, Moscow, Russia

Khaled Abud

AI Center, Lomonosov MSU, Moscow, Russia

Lomonosov Moscow State University, Moscow, Russia

MSU Institute for Artificial Intelligence, Moscow, Russia

https://github.com/msu-video-group/freeflow

E-mail

{vladislav.bargatin, alexander.yakovenko}@graphics.cs.msu.ru

Dmitriy Vatolin

E-mail

{khaled.abud, dmitriy}@graphics.cs.msu.ru

MSU Institute for Artificial Intelligence, Moscow, Russia

https://github.com/msu-video-group/freeflow

E-mail

{vladislav.bargatin, alexander.yakovenko}@graphics.cs.msu.ru

Abstract

Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefined heuristics, which can limit its expressivity and lead to more complex pipelines and additional computational cost. We present FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder–decoder. FreeFlow combines three attention variants: window attention for local processing, shifted-window attention for cross-window information exchange, and a global attention operating at a reduced resolution. The resulting architecture scales naturally with model capacity, enabling a consistent accuracy gain from small to large variants. Despite the absence of standard inductive biases, FreeFlow achieves state-of-the-art results on major benchmarks, including Sintel (0.68/1.48 EPE on Clean/Final), KITTI-2015 (3.23 Fl-all), and Spring (3.192 1px), while remaining memory efficient at 1080p inference.

Keywords: 

Optical flow Vision Transformers High-resolution

1 Introduction

Optical flow estimation (the dense per-pixel motion between frames) is a fundamental task in low-level vision, with applications ranging from video understanding[32, 44, 62] and tracking to video restoration and synthesis[12, 24, 57, 6].

Early optical flow methods posed estimation as variational optimization[27, 10], later accumulating stronger regularization [41], coarse-to-fine schemes, and hand-crafted descriptors[53] at the cost of growing complexity.

Deep learning initially simplified optical flow, as FlowNet[9] showed that a feed-forward network can regress flow directly, but later work reintroduced classical biases in highly effective yet hard-wired architectures. PWC-Net[42] combined pyramids, warping, and local correlation volumes, while RAFT[45] popularized iterative refinement over an all-pairs correlation volume and convex upsampling; many follow-ups then build upon these foundations with additional modules for occlusions[14], temporal cues[38, 7], and training/inference refinements[50]. Which improve accuracy, but lead to complex pipelines that are harder to modify, scale, and repurpose beyond optical flow.

In parallel, the broader computer vision literature has moved in the opposite direction: vision transformers[8] increasingly replace bespoke pipelines in detection[5], segmentation[18, 16], depth prediction[58, 59], and 3D tasks[47, 46, 15], driven by the observation that generic, data-driven feed-forward models can learn the required structure from scale and supervision. This trend motivates revisiting optical flow through the same lens: can we omit flow-specific components and still obtain a state-of-the-art approach?

Several recent approaches move toward more generic architectures, but still retain flow-specific structure or incur practical limitations. DDVM[37] casts optical flow prediction in a diffusion framework, but the prediction is still produced through iterative process. CroCo[52] is close to a pure transformer, yet it operates at a relatively small fixed resolution and typically requires tiling for higher-resolution inputs, while still relying on a large convolutional decoder. Most recently, WAFT[49] and GeoViT[54] explicitly aim for generality, but still rely on iterations and require warping (of features and input frames, respectively). Additionally, these approaches are trained at relatively small, sub-megapixel resolutions, which might become a limiting factor for accuracy at high-resolution inference [1].

媒体内容 · 前往原文查看
(a) Spring EPE () versus 1080p inference memory (GB), with FreeFlow variants (S/M/L) illustrating model scaling.
Refer to caption
(b) Qualitative comparison of predicted flow fields on a crop of a visually challenging scene with heavy blur (best viewed zoomed in).
Figure 1: FreeFlow overview. Our method achieves state-of-the-art accuracy on Spring with low 1080p inference memory and the sharpest motion borders among strong baselines, as evidenced by (a) the accuracy–memory scatter plot and (b) the qualitative crop gallery.

We therefore propose FreeFlow, a bias-free hierarchical transformer for optical flow estimation. FreeFlow is designed for high-resolution processing, which recent work has shown to be beneficial for optical flow[1] and depth estimation[3]. At high resolution, accurate flow requires both strong local reasoning (to preserve fine structures and motion boundaries) and effective long-range information flow (to resolve large displacements). FreeFlow addresses this with a hierarchical attention design that combines tiled processing with repeated local–global feature interaction: local attention focuses on within-region detail, cross-region exchange propagates information between neighboring tiles, and global mixing enables long-range correspondence. FreeFlow sets a new state-of-the-art on Sintel[4] (EPE clean: 0.68; EPE final: 1.48), KITTI-2015[30] (Fl-all: 3.23), and Spring[29] (1px: 3.192).

Our key contributions are:

  • Bias-free optical flow transformer. We introduce FreeFlow, a hierarchical transformer for optical flow built without common flow-specific inductive biases and modules (e.g., explicit cost/correlation volumes, warping-based update pipelines, and specialized upsampling), using a single end-to-end trainable architecture.

  • Hierarchical local–global feature interaction for high-resolution processing. We propose a tiled transformer design with repeated local and long-range feature interaction, enabling accurate flow at high resolutions by combining local detail modeling with global information flow.

  • State-of-the-art performance across benchmarks. FreeFlow achieves state-of-the-art results on Sintel, KITTI-2015, and Spring, all while being memory-efficient and scalable to smaller parameter counts.

2 Related Works

2.1 Inductive Biases of Optical Flow

Early optical flow methods formulated estimation as variational optimization under photometric constancy and smoothness regularization[27, 10, 41, 53]. Learning-based approaches later treated optical flow as supervised dense prediction[9], and subsequent advances largely introduced explicit architectural priors tailored to correspondence estimation. Representative flow-specific inductive biases include multi-scale pyramids[42, 31], warping-based alignment[42, 49, 54], explicit correlation volumes for large-displacement matching[45, 1], iterative refinement, and convex upsampling[45, 50, 1]. While these choices are effective, they have contributed to increasingly complex pipelines; recent work has started to remove individual priors[17] and move toward more generic formulations[37], but existing methods typically retain some of these biases rather than eliminating them entirely.

媒体内容 · 前往原文查看
Table 1: Architectural Inductive Biases in Optical Flow Methods. We summarize common design choices used to inject task-specific structure compared to our approach, which has no flow-specific inductive biases and uses a simple decoder head. "DPT head" refers to the decoder head from Dense Prediction Transformer (DPT) [33].
Method Correlation Volume Feature Warping Pyramid Refinement Iterative Refinement Convex Upsampling Inference time Tiling Flow Head
PWC-Net [42] × × × ×
RAFT [45] × × × ×
FlowFormer [11] × × ×
UniMatch [56] × × ×
TransFlow [26] × × × ×
DPFlow [31] × × ×
MEMFOF [1] × × × ×
WAFT [49] × × × ×
Geo-VIT [54] × × ×
CroCo-Flow [52] × × × × × DPT head
Win-Win [20] × × × × × × DPT head
FreeFlow (ours) × × × × × × Simple Conv.

2.2 Vision Transformers in Dense Prediction

Vision transformers[8] (ViTs) are now widely used for dense prediction and geometry-oriented tasks, enabled by large-scale data and a general architecture that can learn the dependencies from supervision rather than relying on hand-crafted priors. This has led to strong results across depth[33, 58, 59, 3], segmentation[18, 16], and 3D settings[47, 46, 15], where feed-forward transformer backbones provide a common foundation for dense outputs and geometric reasoning. Importantly, many of these systems remain architecturally simple. A common pattern is a ViT backbone combined with a convolutional prediction head or decoder for producing dense maps[33, 34]. Other approaches further reduce decoder structure and rely on minimal output token processing, suggesting that heavy convolutional decoders are not necessary to obtain competitive dense predictions [16, 15]. However, scaling such models to high resolution is challenging; common workarounds such as downsampling or tiling can harm accuracy and limit long-range interaction. Swin[25, 22] addresses the issue by using local-window self-attention and shifted windows to pass information across window boundaries, while Hiera[36] provides a simple hierarchical multi-scale ViT backbone for large images. Alternatively, DepthPro does high-resolution processing via a multi-scale pyramid design with late feature fusion[3].

2.3 Transformers in Optical Flow Estimation

Recent transformer-based optical flow methods differ mainly in which flow-specific inductive biases they retain. Some methods keep explicit matching as a central operation via correlation/cost-based similarity: UniMatch and TransFlow follow this direction [56, 26], while FlowFormer explicitly constructs and processes a 4D cost volume with a transformer-style decoder [11].

While others move closer to generic transformer formulations, they unfortunately still leave some biases in place. CroCo-Flow was among the first transformer approaches with competitive accuracy, leveraging binocular pretraining for dense matching, but it relies on slow dense tiling at high resolution [52]. Win-Win builds on CroCo to enable FullHD training and inference without tiling, but does not improve over CroCo-Flow in accuracy [20]. WAFT and GeoViT remove cost volumes but retain iterative warping-based updates (warping features and input images, respectively) [49, 54]. Overall, transformers have often been incorporated by incrementally replacing parts of established optical flow pipelines, which can improve performance yet further diversify and complicate the set of design choices. We summarize these choices in Tab. 1 using common inductive biases and compare to our bias-free approach.

3 Method

Refer to caption
Figure 2: Comparison of high-resolution prediction techniques. (a) No Feature Fusion: per-tile prediction without inter-tile feature exchange (limited global coherence; long-range motions spanning multiple tiles not captured). (b) Late Feature Fusion: multi-scale features are fused only in the decoder (global context arrives late and may not propagate to full-resolution features; with small overlap and no information flow across tile borders, seams may remain visible; see supplementary). (c) Dense Feature Fusion (ours): repeated cross-tile/global feature exchange across the network (information flow is not restricted to the decoder).

In this section, we present FreeFlow, an inductive-bias-free transformer for optical flow estimation (Fig. 3). We design this approach with two goals in mind. First, we aim to demonstrate that state-of-the-art optical flow can be achieved without the flow-specific architectural modules that dominate modern pipelines, such as correlation volumes, feature warping, iterative refinement, etc. (summarized in Tab. 1). Second, we seek a design that remains practical at high resolutions and ensures repeated information exchange across processing scales.

To this end, we study existing high-resolution prediction strategies, identify their limitations, and derive a principled alternative (Fig. 2). A straightforward solution is inference-time tiling in CroCo-/FlowFormer-style pipelines (Fig. 2a), but tiles interact only through late-stage averaging, so information does not flow across tile borders during feature extraction, and the number of forward passes grows with resolution. An alternative is multi-scale decoding with Late Feature Fusion as in DepthPro-like designs (Fig. 2b), which reduces the number of forward passes, but keeps information exchange delayed and tied to fixed-resolution assumptions, and is unreliable in the binocular setting when objects cross tile boundaries. FreeFlow resolves these issues by using a fixed and efficient tiling scheme while enabling dense feature exchange throughout the network (Fig. 2c), combining local processing with cross-tile and global interactions.

We next give a functional description of the architecture and then formalize the attention variants used by FreeFlow.

Refer to caption
Figure 3: Method overview. Given an image pair (I1,I2), we patchify each input into 8×8 tokens and extract features using two shared-weight encoders composed of Window, Shifted-Window, and Global attention blocks (Fig. 4), producing F1 and F2. A transformer decoder combines self-attention with cross-attention between F1 and F2 to form flow tokens, which are depatchified and mapped to a dense optical flow field by a lightweight prediction head.

3.1 Approach

We adopt the CroCo/DUSt3R/MASt3R encoder–decoder high-level architecture for binocular reasoning, i.e., a Siamese ViT encoder followed by a decoder that alternates self- and cross-attention between the two views.

Given two input images I1,I2H×W×3, we embed each image into a sequence of non-overlapping P×P patches (with P=8) using a standard patch projection, producing token sequences X1,X2N×D where N=HPWP. Both sequences are processed by a Siamese transformer encoder with shared weights to obtain feature representations

F1=Encoder(X1),F2=Encoder(X2), (1)

with F1,F2N×D.

A transformer decoder then produces flow tokens

Z=Decoder(F1,F2), (2)

where each decoder block combines self-attention over the current tokens with cross-attention from F1 (queries) to F2 (keys/values), enabling repeated information exchange between the two views. The output ZN×D is reshaped into a spatial feature map ZmapHP×WP×D and mapped to a dense flow and confidence field UH×W×5 by a prediction head.

媒体内容 · 前往原文查看
Figure 4: Attention block variants. Our architecture uses three attention patterns: (i) Window attention over a 4×4 partition (16 non-overlapping windows), (ii) Shifted-Window attention with a half-window offset to exchange information across window boundaries, and (iii) Global attention applied at 2× lower spatial resolution via down/up-sampling. See Fig. 2 for the partitioning visualization.

Attention Block Types.

To predict highly-detailed globally consistent flow fields, FreeFlow uses a hierarchical attention design that mixes local processing, cross-window exchange, and global context. Each block follows the CroCo-style transformer structure: encoder blocks apply self-attention, while decoder blocks additionally use cross-attention to utilize tokens from the other view. We use three attention variants (Fig. 4):

  • Window (Win) Attention Block. The token map of size H8×W8 is partitioned into a 4×4 grid of 16 non-overlapping windows, each containing H32×W32 tokens. Attention is computed independently within each window.

  • Shifted-Window (Swin) Attention Block. To enable information flow across window (and tile) boundaries, we apply a half-window shift by W64 tokens horizontally and H64 tokens vertically, partition into the same 4×4 windows, perform window attention, and shift back.

  • Global Attention Block. To incorporate global context at controlled cost, we downsample by 2× with a stride-2 convolution, apply full attention at the reduced resolution, and upsample by 2× with a stride-2 transposed convolution. We apply normalization after the residual connection.

Each encoder/decoder layer applies the blocks in the fixed order: Win, Swin, and Global. Since attention cost scales quadratically with the number of tokens, the 2× downsampling in the Global block keeps its attention cost comparable to Win/Swin at the original resolution. To encode token positional information within the image, we rely on Rotary Positional Embedding (RoPE) [40].

Attention Scale Factor.

Following prior work[7] on resolution-adaptive attention scaling, we multiply the attention logits by a logarithmic factor of the token count, which improves generalization when running inference at resolutions higher than those seen during training. For a token map of size H8×W8 we use

SelfAttention(Q,K,V) =Softmax(log(H/8×W/8)D×QKT)×V (3)
CrossAttention(Q1,K2,V2) =Softmax(log(H/8×W/8)D×Q1K2T)×V2 (4)

Unlike prior formulations that normalize the factor to be 1 at the training token count, we use the unnormalized variant and found it to work well in practice.

Refer to caption
Figure 5: Qualitative comparison on the Spring benchmark [29]. Examples of error-maps of Win-Win [20], CroCo-Flow [52], WAFT [49], and FreeFlow predictions; the colorbar represents endpoint error. Our approach combines high level of detail (note the hole in the staff that only our method captures and the thin object bounds) with high motion consistency (note the staff’s end in the top example and the low error in the background of the bottom example). Crops are sourced from official leaderboard submissions.

Flow Head.

Most high-performing optical flow pipelines rely on flow-specific prediction machinery, such as iterative update stages and convex upsampling. In addition, bias-free transformer baselines often use DPT-style heads from [33] that aggregate features from multiple layers (and, in CroCo-style designs, may also reuse encoder features) to form the final prediction. In contrast, FreeFlow predicts flow directly from the final decoded patch map using a simple three-layer head, without iterative refinement or convex upsampling.

Given FH8×W8×D, we apply a 3×3 convolution to expand channels to 4D, followed by a 1×1 convolution to 4096 channels, and a transposed convolution with kernel and stride 8 to upsample to H×W. The head outputs H×W×5, where the first two channels represent optical flow and the remaining three parameterize the uncertainty terms used by the mixture-of-Laplace loss (following SEA-RAFT [50]). Finally, we multiply the flow channels by 8 (the patch size) to obtain flow in pixel units.

4 Experiments

We first detail our training pipeline, consisting of cross-view completion pretraining followed by finetuning for optical flow. We evaluate our method on three popular optical flow benchmarks: Spring [29] (high-resolution real-world sequences), Sintel [4] (synthetic scenes with complex motion and rendering effects), and KITTI-2015 [30] (real driving scenes). Finally, we provide ablations of key design choices, including the attention configuration, pretraining masking ratio, and model scaling.

媒体内容 · 前往原文查看
Table 2: Training procedure details. Dataset abbreviations: TA: TartanAir [48], T: Things [28], S: Sintel [4], K: KITTI-2015 [30], H: HD1K [19]. Inspired by SEA-RAFT [50] and MEMFOF [1], the dataset distributions for the TaTSKH stages are TA (0.23), S(0.25), T(0.24), K(0.09), H(0.19).
Stage Weights Datasets Scale Crop size LR WD Batch Steps
Pretrain ARKitScenes [2] 1x [224, 224] 8e-4 5e-2 2048 346k
MegaDepth [21]
3DStreetView [60]
TaTSKH Pretrain TA+T+S+K+H 2x 10880 tok. 4e-5 1e-2 32 450k
TaTSKH-hq TaTSKH TA+T+S+K+H 2x 32640 tok. 1e-5 1e-5 32 90k
Sintel-ft TaTSKH-hq S 2x [872, 2048] 1e-5 1e-5 32 12.5k
KITTI-ft TaTSKH-hq K 2x [750, 2484] 1e-5 1e-5 32 2.5k
Spring-ft TaTSKH-hq Spring [29] 1x [1080, 1920] 1e-5 1e-5 32 60k

4.1 Training Details

Following CroCo [51, 52], we first pre-train our model on the cross-view completion task and then finetune the resulting weights with a new head for optical flow estimation. In cross-view completion, a large fraction of patches in one view is replaced by a learned token emask, and the model reconstructs the missing content conditioned on the second view. Due to its two-image nature, this objective encourages learning dense long-range correspondences, which is well aligned with the downstream task of binocular matching. Training details and datasets are summarized in Tab. 2; please refer to the supplementary for additional details.

Pre-train.

We follow the pre-training stage protocol from CroCo with minor adjustments. During cross-view completion training, a model takes as input two images that represent different views of the same scene: one view is severely masked, and the model has to predict the masked regions using the information from the second view. We sample image pairs from ARKitScenes [2], MegaDepth [21], and 3DStreetView [60], resulting in 3.7M data samples in total. We use fixed-size crops of 224×224 resolution and pretrain the model for 346k steps.

The only significant difference from the CroCo setup is when the learned masked patch representation emask is introduced. In CroCo, masked tokens are removed from the first view and the corresponding emask tokens are added only at the decoder input. Here, due to the hierarchical nature of our model, we replace masked patches with emask at the encoder input, while keeping the completion objective unchanged.

Optical Flow Finetune.

The finetuning stage protocol is inspired by MEMFOF [1]; specifically, we adopt their 2× upsampling of training frames, which better matches the motion distribution of FullHD inputs and improves high-resolution performance. Unlike MEMFOF and other curriculum-based training recipes that use multiple sequential stages, we use a single main dataset mixture, denoted TaTSKH in Tab. 2, for simplicity. For additional speed, we split this finetuning into a low- and high-token-count stage (TaTSKH and TaTSKH-hq in Tab. 2).

Instead of using a fixed crop size, to avoid unnecessary padding and to expose the model to a wider motion range, we use variable-resolution training with a fixed token budget per minibatch. More specifically, for each sample we randomly choose one spatial dimension (height or width), sample its value, and set the other dimension to the largest value such that the resulting token count does not exceed the prescribed budget. This produces rectangular crops with varying aspect ratios while keeping compute and memory controlled. The implementation is straightforward as we use a batch size of 1 sample/GPU during the finetuning stage. For benchmark submissions, we further finetune with fixed crop sizes (Sintel-ft, KITTI-ft, Spring-ft in Tab. 2). Following SEA-RAFT[50], we use the Mixture-of-Laplace loss.

In total, it takes from 4 to 5 days to pre-train and around 3 days to finetune our largest model on 32 GPUs.

4.2 Results

We adopt four widely used metrics from established benchmarks in this study: endpoint error (EPE), 1-pixel outlier rate (1px), Fl-score, and WAUC error. Please refer to [35, 4, 29, 31, 30] or the supplementary for their definitions.

媒体内容 · 前往原文查看
Table 3: Spring benchmark results. "*" indicates that it was submitted by the Spring team without being finetuned on the provided training set. Speed (runtime) and peak GPU memory consumption were measured on a Nvidia RTX 3090 GPU (24 GB) with automatic mixed precision and without memory efficient correlation volumes, where "—" denotes no available data or implementation for a given model and "†" denotes our best estimate based on the authors description. "MF" indicates that the method uses multiple frames (three or more). The best results are indicated in bold, second-best are underlined, third best are indicated in italic. Method configurations are taken from submissions to the Spring benchmark if present, and from submissions to the Sintel benchmark otherwise.
Method Inf. Cost (1080p) Params (M) Spring
Memory (GB) Time (ms) 1px EPE Fl WAUC
FlowNet2 [13] 3.01 110 162.52 6.710 1.040 2.823 90.907
PWC-Net [42] 0.57 48 9.37 82.265 2.288 4.889 45.670
RAFT [45] 7.97 406 5.26 6.790 1.476 3.198 90.920
GMA [14] 11.81 830 5.88 7.074 0.914 3.079 90.722
GMFlow [55] 8.22 8024 4.72 10.355 0.945 2.952 82.337
FlowFormer [11] 1.90 2084 16.17 6.510 0.723 2.384 91.679
SEA-RAFT (M) [50] 8.12 198 19.67 3.686 0.363 1.347 94.534
DPFlow [31] 4.26 401 10.02 3.442 0.340 1.311 94.980
MemFlow(MF)[7] 8.06 754 6.27 4.482 0.471 1.416 93.855
StreamFlow(MF)[43] 18.61 898 14.25 4.152 0.467 1.424 94.404
MEMFOF(MF)[1] 1.90 262 75.78 3.289 0.355 1.238 95.186
ARFlow(MF)[23] 76.50 3.265 0.353 1.212 95.283
CroCo-Flow [52] 2.73 3266 447.47 4.565 0.498 1.508 93.660
Win-Win [20] 3.82 305 229.77 5.371 0.475 1.621 92.720
WAFT-DAv2-a2 [49] 20.58 489 56.93 3.298 0.304 1.197 94.990
WAFT-DINOv3-a2 [49] 18.96 408 56.47 3.182 0.325 1.246 95.051
FreeFlow-S (ours) 1.02 144 34.58 5.087 0.533 1.452 90.196
FreeFlow-M (ours) 1.66 325 102.46 3.392 0.346 1.171 94.919
FreeFlow-L (ours) 2.58 607 230.72 3.192 0.278 1.048 95.235

Results on Spring.

FreeFlow achieves state-of-the-art performance on Spring. FreeFlow-L sets the best EPE and Fl among all compared methods (Tab. 3), while remaining on par with the strongest approaches in 1px and WAUC; in particular, it attains the best WAUC among two-frame methods. Compared to WAFT-DAv2-a2, FreeFlow-L improves EPE by 9% and reduces Fl by 14%. Owing to tiling-free native 1080p inference, FreeFlow preserves fine detail while maintaining global motion consistency (Fig. 5). We further highlight that native 1080p processing is possible within a low inference memory budget (Fig. 1).

媒体内容 · 前往原文查看
Table 4: Sintel [4] and KITTI-15 [30] benchmark results. Sintel uses EPE as it’s metric for both splits, while KITTI-15 uses the Fl-all outliers metric. "—" indicates no published results and "MF" indicates that the method used multiple frames (three or more) to generate it’s submissions.
Method Sintel KITTI-15
Clean Final Fl-all
FlowNet2 [13] 4.16 5.74 10.41
PWC-Net [42] 3.90 5.04 9.60
RAFT [45] 1.61 2.86 5.10
GMA [14] 1.39 2.47 5.15
GMFlow+ [56] 1.03 2.37 4.49
FlowFormer [11] 1.16 2.09 4.68
FlowFormer++ [39] 1.07 1.94 4.52
TransFlow [26] 1.06 2.08 4.32
SEA-RAFT (L) [50] 1.31 2.60 4.30
DPFlow [31] 1.05 1.98 3.56
VideoFlow-BOF(MF)[38] 1.00 1.71 4.44
VideoFlow-MOF(MF)[38] 0.99 1.65 3.65
MEMFOF(MF)[1] 0.99 1.94 2.94
MEMFOF-XL(MF)[1] 0.93 1.89
ARFlow(MF)[23] 0.96 1.79 2.85
DDVM[37] 1.75 2.48 3.26
CroCo-Flow[52] 1.09 2.44 3.64
Win-Win[20] 1.15 2.34
WAFT-DAv2-a2[49] 0.94 2.33 3.31
WAFT-DINOv3-a2[49] 0.95 2.02 3.56
GeoVIT [54] 0.79 1.88 3.79
FreeFlow-S (ours) 1.03 1.99 4.06
FreeFlow-M (ours) 0.80 1.77 3.33
FreeFlow-L (ours) 0.68 1.48 3.23
Refer to caption
Figure 6: Qualitative comparison on Sintel. Examples of FreeFlow-L, GeoViT [54], WAFT-DAv2-a2 [49], CroCo-Flow [52], Win-Win [20], and MEMFOF-XL [1] outputs on the Sintel [4] benchmark. Sourced from the official leaderboard submissions.

Results on Sintel and KITTI.

Following MEMFOF, we finetune on 2× upsampled frames. Accordingly, for Sintel and KITTI submissions we upscale input images by 2× and downscale the predicted flow by 2×. FreeFlow-L ranks first on Sintel on both Clean and Final (Tab. 4), improving over GeoVIT [54] by 14% on Clean (0.790.68) and 10% over VideoFlow-MOF [38] on Final (1.651.48). Notably, FreeFlow-M is already highly competitive: it is second only to FreeFlow-L on Clean, and ranks fourth on Final, surpassed only by the 3- and 5-frame VideoFlow variants. On KITTI-2015, FreeFlow-L achieves 3.23 Fl-all, outperforming all non-stereo and non-multiframe methods on KITTI-15 (Tab. 4). Qualitative results show that the model captures complex motion patterns using only two frames (Fig. 6). Additional visual comparisons and zero-shot evaluations are provided in the supplementary material.

媒体内容 · 前往原文查看
Table 5: Masking ratio and attention-block ablation. Spring sub-validation results for FreeFlow variants with different cross-view completion masking ratios and active attention subblocks (Win/Swin/Global). Gray marks the configuration selected for the main model; best and second-best results are shown in bold and underlined, respectively.
Mask. Ratio Subblock type Spring (sub-val)
Win Swin Global 1px EPE
0.9 × × 1.133 0.229
0.9 × 0.801 0.190
0.9 × 0.659 0.167
0.9 0.688 0.170
0.925 0.685 0.166
0.95 × 0.658 0.166
0.95 0.624 0.157
0.975 0.656 0.174
Refer to caption
Figure 7: Qualitative ablation comparison of FreeFlow models with and without Global Attention block and pretrained with different masking ratios on the Spring benchmark. Zoom in for better view.

4.3 Ablation Study

Unless stated otherwise, all ablations use a scaled-down FreeFlow configuration with 8 encoder and 8 decoder layers, width 256, and 8 attention heads, and are trained with the same training recipe. Following SEA-RAFT and WAFT, we report results on the Spring sub-validation split (scenes 0045 and 0047) after finetuning on the remaining training data. Additional experiments in the supplementary isolate the architecture from the training procedure and study the effect of adding iterative flow-specific biases back into FreeFlow.

Architecture and Masking Ratio Ablation.

We study the interaction between the cross-view completion masking ratio and the hierarchical attention design of FreeFlow (Tab. 5). The masking ratio controls the fraction of patches in the target view that are replaced by the learned token emask during pretraining, while the attention configuration determines which subblock types (Win/Swin/Global) are present in the encoder and decoder.

With all three subblocks enabled, a masking ratio of 0.95 performs best, improving over the CroCo default 0.9 by 9.3% in 1px and 7.6% in EPE. This is consistent with FreeFlow using smaller patches than CroCo (8×8 vs. 16×16), which reduces the distance to visible regions and makes a higher mask rate beneficial as it makes the pretraining task sufficiently challenging.

We also observe an interaction between masking ratio and attention configuration. At the CroCo default ratio (0.9), removing Global does not hurt and can even slightly improve the sub-validation metrics, whereas at 0.95 Global becomes important for the best performance. This indicates that the masking ratio is an important hyperparameter that should be chosen jointly with the model architecture. Shifted-Window attention consistently contributes to accuracy, and removing it leads to a clear drop on this split. Additionally, for both masking ratios, qualitative comparisons (Fig. 7) show that removing the Global block can introduce obvious motion inconsistencies, even when the sub-validation metrics change only marginally.

媒体内容 · 前往原文查看
Table 6: Model configurations. Architectural hyperparameters of FreeFlow variants and our best-effort estimates for transformer-based baselines. “+” denotes separate encoder and decoder widths (or counts), while “×” denotes repeated recurrent decoder calls.
Name Patch Sizes Width Attn. Count Attn. Heads MLP Count Params (M)
GeoViT 16 1024 24×6 16 24×6 377
WAFT 16 384+384 12+12×5 6+6 12+12×5 56
CroCo-Flow 16 1024+768 24+24 16+12 24+12 447
Win-Win 16 768+768 12+24 12 12+12 230
GMFlow+ 8/4 0+128 0+12×2 0+1 0+6×2 4.7
FreeFlow-S 8 256+256 12+24 4+4 12+12 35
FreeFlow-M 8 384+384 18+36 6+6 18+18 102
FreeFlow-L 8 512+512 24+48 8+8 24+24 231
媒体内容 · 前往原文查看
Figure 8: Model scaling on Sintel. Sintel Final EPE () versus parameter count (M), ×106 for FreeFlow variants and transformer-based baselines.

Model Scaling.

A benefit of FreeFlow’s simple, uniform encoder–decoder design is that it can be scaled in a straightforward manner by adjusting depth and width. We therefore study how performance evolves across model sizes (S/M/L) on common optical flow benchmarks. We follow common ViT scaling [61] best practices by co-scaling depth, width, and the number of attention heads while keeping the overall architecture and patch size fixed. The resulting configurations are summarized in Tab. 6.

As shown in  Fig. 8, FreeFlow scales predictably: larger variants yield consistent accuracy gains, while smaller variants retain strong performance, indicating that the approach remains effective even when scaled down. Memory scaling with model size is reported in Fig. 1(a).

5 Conclusion

In this work, we introduced FreeFlow, a bias-free hierarchical transformer for optical flow estimation that removes conventional flow-specific design components such as correlation volumes, feature warping, and iterative refinement. Instead, FreeFlow relies on a simple feed-forward encoder–decoder architecture that combines window, shifted-window, and reduced-resolution global attention to capture motion across multiple spatial scales. This design provides a flexible and scalable framework that improves consistently with model capacity while maintaining efficient high-resolution inference.

We show that these standard optical-flow inductive biases are not required to achieve top performance. FreeFlow reaches state-of-the-art results on all popular benchmarks, including Sintel, KITTI-2015, and Spring, demonstrating that strong motion estimation can be obtained from a general-purpose transformer architecture without specialized flow modules. We hope that this work encourages further exploration of simpler and more general architectures for motion estimation and related dense correspondence tasks.

Acknowledgements

The work of Vladislav Bargatin, Alexander Yakovenko and Khaled Abud was supported by the The Ministry of Economic Development of the RussianFederation in accordance with the subsidy agreement (agreement identifier000000C313925P4H0002; grant No 139-15-2025-012). The research was carried out using the MSU-270 supercomputer of Lomonosov Moscow State University.

References

  • [1] V. Bargatin, E. Chistov, A. Yakovenko, and D. Vatolin (2025) MEMFOF: high-resolution training for memory-efficient multi-frame optical flow estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8187–8196. External Links: Document Cited by: §1, §1, §2.1, Table 1, Figure 6, Figure 6, §4.1, Table 2, Table 2, Table 3, Table 4, Table 4.
  • [2] G. Baruch, Z. Chen, A. Dehghan, Y. Feigin, P. Fu, T. Gebauer, D. Kurz, T. Dimry, B. Joffe, A. Schwartz, and E. Shulman (2021) ARKitScenes: a diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Vol. 1, pp. . Cited by: §4.1, Table 2.
  • [3] A. Bochkovskiy, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. Richter, and V. Koltun (2025) Depth Pro: sharp monocular metric depth in less than a second. In International Conference on Learning Representations, Vol. 2025, pp. 75602–75637. Cited by: §1, §2.2.
  • [4] D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black (2012) A naturalistic open source movie for optical flow evaluation. In Computer Vision – ECCV 2012, Berlin, Heidelberg, pp. 611–625. External Links: ISBN 978-3-642-33783-3, Document Cited by: §1, Figure 6, Figure 6, §4.2, Table 2, Table 2, Table 4, Table 4, §4.
  • [5] N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko (2020) End-to-end object detection with transformers. In Computer Vision – ECCV 2020, Cham, pp. 213–229. External Links: ISBN 978-3-030-58452-8, Document Cited by: §1.
  • [6] K. C.K. Chan, X. Wang, K. Yu, C. Dong, and C. C. Loy (2021) BasicVSR: the search for essential components in video super-resolution and beyond. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4945–4954. External Links: Document Cited by: §1.
  • [7] Q. Dong and Y. Fu (2024) MemFlow: optical flow estimation and prediction with memory. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 19068–19078. External Links: Document Cited by: §1, §3.1, Table 3.
  • [8] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations, Cited by: §1, §2.2.
  • [9] A. Dosovitskiy, P. Fischer, E. Ilg, P. Häusser, C. Hazirbas, V. Golkov, P. v. d. Smagt, D. Cremers, and T. Brox (2015) FlowNet: learning optical flow with convolutional networks. In 2015 IEEE International Conference on Computer Vision (ICCV), Vol. , pp. 2758–2766. External Links: Document Cited by: §1, §2.1.
  • [10] B. K.P. Horn and B. G. Schunck (1981) Determining optical flow. Artificial Intelligence 17 (1-3), pp. 185–203. External Links: ISSN 0004-3702, Document Cited by: §1, §2.1.
  • [11] Z. Huang, X. Shi, C. Zhang, Q. Wang, K. C. Cheung, H. Qin, J. Dai, and H. Li (2022) FlowFormer: a transformer architecture for optical flow. In Computer Vision – ECCV 2022, Cham, pp. 668–685. External Links: ISBN 978-3-031-19790-1, Document Cited by: §2.3, Table 1, Table 3, Table 4.
  • [12] Z. Huang, T. Zhang, W. Heng, B. Shi, and S. Zhou (2022) Real-time intermediate flow estimation for video frame interpolation. In Computer Vision – ECCV 2022, Cham, pp. 624–642. External Links: ISBN 978-3-031-19781-9, Document Cited by: §1.
  • [13] E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox (2017) FlowNet 2.0: evolution of optical flow estimation with deep networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 1647–1655. External Links: Document Cited by: Table 3, Table 4.
  • [14] S. Jiang, D. Campbell, Y. Lu, H. Li, and R. Hartley (2021) Learning to estimate hidden motions with global motion aggregation. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 9752–9761. External Links: Document Cited by: §1, Table 3, Table 4.
  • [15] H. Jin, H. Jiang, H. Tan, K. Zhang, S. Bi, T. Zhang, F. Luan, N. Snavely, and Z. Xu (2025) LVSM: a large view synthesis model with minimal 3d inductive bias. In International Conference on Learning Representations, Vol. 2025, pp. 60001–60021. Cited by: §1, §2.2.
  • [16] T. Kerssies, N. Cavagnero, A. Hermans, N. Norouzi, G. Averta, B. Leibe, G. Dubbelman, and D. De Geus (2025) Your vit is secretly an image segmentation model. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 25303–25313. External Links: Document Cited by: §1, §2.2.
  • [17] S. Kiefhaber, S. Roth, and S. Schaub-Meyer (2025) Removing cost volumes from optical flow estimators. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 79–89. External Links: Document Cited by: §2.1.
  • [18] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, P. Dollár, and R. Girshick (2023) Segment anything. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 3992–4003. External Links: Document Cited by: §1, §2.2.
  • [19] D. Kondermann, R. Nair, K. Honauer, K. Krispin, J. Andrulis, A. Brock, B. Güssefeld, M. Rahimimoghaddam, S. Hofmann, C. Brenner, and B. Jähne (2016) The HCI Benchmark Suite: stereo and flow ground truth with uncertainties for urban autonomous driving. In 2016 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vol. , pp. 19–28. External Links: Document Cited by: Table 2, Table 2.
  • [20] V. Leroy, J. Revaud, T. Lucas, and P. Weinzaepfel (2024) Win-Win: training high-resolution vision transformers from two windows. In International Conference on Learning Representations, Vol. 2024, pp. 48749–48767. Cited by: §2.3, Table 1, Figure 5, Figure 5, Figure 6, Figure 6, Table 3, Table 4.
  • [21] Z. Li and N. Snavely (2018) MegaDepth: learning single-view depth prediction from internet photos. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , pp. 2041–2050. External Links: Document Cited by: §4.1, Table 2.
  • [22] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte (2021) SwinIR: image restoration using swin transformer. In 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Vol. , pp. 1833–1844. External Links: Document Cited by: §2.2.
  • [23] J. Liu, M. Liu, S. Zhu, Y. Zhang, J. Li, M. Y. Yang, F. Nex, H. Cheng, and H. Wang (2026) ARFlow: auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting. In International Conference on Learning Representations, Vol. 2026. Cited by: Table 3, Table 4.
  • [24] X. Liu, H. Liu, and Y. Lin (2020) Video frame interpolation via optical flow estimation with image inpainting. International Journal of Intelligent Systems 35 (12), pp. 2087–2102. External Links: Document Cited by: §1.
  • [25] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin Transformer: hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 9992–10002. External Links: Document Cited by: §2.2.
  • [26] Y. Lu, Q. Wang, S. Ma, T. Geng, Y. V. Chen, H. Chen, and D. Liu (2023) TransFlow: transformer as flow learner. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 18063–18073. External Links: Document Cited by: §2.3, Table 1, Table 4.
  • [27] B. D. Lucas and T. Kanade (1981) An Iterative Image Registration Technique with an Application to Stereo Vision. In IJCAI’81: 7th international joint conference on Artificial intelligence, Vol. 2, Vancouver, Canada, pp. 674–679. Cited by: §1, §2.1.
  • [28] N. Mayer, E. Ilg, P. Häusser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox (2016) A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4040–4048. External Links: Document Cited by: Table 2, Table 2.
  • [29] L. Mehl, J. Schmalfuss, A. Jahedi, Y. Nalivayko, and A. Bruhn (2023) Spring: a high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4981–4991. External Links: Document Cited by: §1, Figure 5, Figure 5, §4.2, Table 2, §4.
  • [30] M. Menze and A. Geiger (2015) Object scene flow for autonomous vehicles. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 3061–3070. External Links: Document Cited by: §1, §4.2, Table 2, Table 2, Table 4, Table 4, §4.
  • [31] H. Morimitsu, X. Zhu, R. M. Cesar, X. Ji, and X. Yin (2025) DPFlow: adaptive optical flow estimation with a dual-pyramid framework. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 17810–17820. External Links: Document Cited by: §2.1, Table 1, §4.2, Table 3, Table 4.
  • [32] A. Piergiovanni and M. S. Ryoo (2019) Representation flow for action recognition. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 9937–9945. External Links: Document Cited by: §1.
  • [33] R. Ranftl, A. Bochkovskiy, and V. Koltun (2021) Vision transformers for dense prediction. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 12159–12168. External Links: Document Cited by: §2.2, Table 1, Table 1, §3.1.
  • [34] N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C. Wu, R. Girshick, P. Dollar, and C. Feichtenhofer (2025) SAM 2: segment anything in images and videos. In International Conference on Learning Representations, Vol. 2025, pp. 28085–28128. Cited by: §2.2.
  • [35] S. R. Richter, Z. Hayder, and V. Koltun (2017) Playing for benchmarks. In 2017 IEEE International Conference on Computer Vision (ICCV), Vol. , pp. 2232–2241. External Links: Document Cited by: §4.2.
  • [36] C. Ryali, Y. Hu, D. Bolya, C. Wei, H. Fan, P. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffman, J. Malik, Y. Li, and C. Feichtenhofer (2023) Hiera: a hierarchical vision transformer without the bells-and-whistles. In Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, pp. 29441–29454. Cited by: §2.2.
  • [37] S. Saxena, C. Herrmann, J. Hur, A. Kar, M. Norouzi, D. Sun, and D. J. Fleet (2023) The surprising effectiveness of diffusion models for optical flow and monocular depth estimation. In Advances in Neural Information Processing Systems, Vol. 36, pp. 39443–39469. Cited by: §1, §2.1, Table 4.
  • [38] X. Shi, Z. Huang, W. Bian, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li (2023) VideoFlow: exploiting temporal cues for multi-frame optical flow estimation. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 12435–12446. External Links: Document Cited by: §1, §4.2, Table 4, Table 4.
  • [39] X. Shi, Z. Huang, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li (2023) FlowFormer++: masked cost volume autoencoding for pretraining optical flow estimation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 1599–1610. External Links: Document Cited by: Table 4.
  • [40] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu (2024) RoFormer: enhanced transformer with rotary position embedding. Neurocomputing 568, pp. 127063. External Links: ISSN 0925-2312, Document Cited by: §3.1.
  • [41] D. Sun, S. Roth, and M. J. Black (2010) Secrets of optical flow estimation and their principles. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Vol. , pp. 2432–2439. External Links: Document Cited by: §1, §2.1.
  • [42] D. Sun, X. Yang, M. Liu, and J. Kautz (2018) PWC-Net: cnns for optical flow using pyramid, warping, and cost volume. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , pp. 8934–8943. External Links: Document Cited by: §1, §2.1, Table 1, Table 3, Table 4.
  • [43] S. Sun, J. Liu, H. Li, G. Liu, T. H. Li, and W. Gao (2024) StreamFlow: streamlined multi-frame optical flow estimation for video sequences. In Advances in Neural Information Processing Systems, Vol. 37, pp. 9205–9228. External Links: Document Cited by: Table 3.
  • [44] S. Sun, Z. Kuang, L. Sheng, W. Ouyang, and W. Zhang (2018) Optical flow guided feature: a fast and robust motion representation for video action recognition. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , pp. 1390–1399. External Links: Document Cited by: §1.
  • [45] Z. Teed and J. Deng (2020) RAFT: recurrent all-pairs field transforms for optical flow. In Computer Vision – ECCV 2020, Cham, pp. 402–419. External Links: ISBN 978-3-030-58536-5, Document Cited by: §1, §2.1, Table 1, Table 3, Table 4.
  • [46] J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny (2025) VGGT: visual geometry grounded transformer. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 5294–5306. External Links: Document Cited by: §1, §2.2.
  • [47] S. Wang, V. Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud (2024) DUSt3R: geometric 3d vision made easy. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 20697–20709. External Links: Document Cited by: §1, §2.2.
  • [48] W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer (2020) TartanAir: a dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vol. , pp. 4909–4916. External Links: Document Cited by: Table 2, Table 2.
  • [49] Y. Wang and J. Deng (2026) WAFT: warping-alone field transforms for optical flow. In International Conference on Learning Representations, Vol. 2026. Cited by: §1, §2.1, §2.3, Table 1, Figure 5, Figure 5, Figure 6, Figure 6, Table 3, Table 3, Table 4, Table 4.
  • [50] Y. Wang, L. Lipson, and J. Deng (2025) SEA-RAFT: simple, efficient, accurate raft for optical flow. In Computer Vision – ECCV 2024, Cham, pp. 36–54. External Links: ISBN 978-3-031-72667-5, Document Cited by: §1, §2.1, §3.1, §4.1, Table 2, Table 2, Table 3, Table 4.
  • [51] P. Weinzaepfel, V. Leroy, T. Lucas, R. Brégier, Y. Cabon, V. Arora, L. Antsfeld, B. Chidlovskii, G. Csurka, and J. Revaud (2022) CroCo: self-supervised pre-training for 3d vision tasks by cross-view completion. In Advances in Neural Information Processing Systems, Vol. 35, pp. 3502–3516. Cited by: §4.1.
  • [52] P. Weinzaepfel, T. Lucas, V. Leroy, Y. Cabon, V. Arora, R. Brégier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud (2023) CroCo v2: improved cross-view completion pre-training for stereo matching and optical flow. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 17923–17934. External Links: Document Cited by: §1, §2.3, Table 1, Figure 5, Figure 5, Figure 6, Figure 6, §4.1, Table 3, Table 4.
  • [53] P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid (2013) DeepFlow: large displacement optical flow with deep matching. In 2013 IEEE International Conference on Computer Vision, Vol. , pp. 1385–1392. External Links: Document Cited by: §1, §2.1.
  • [54] H. Wu, K. Cheng, S. Lin, and Z. Wu (2026) A study of finetuning video transformers for multi-view geometry tasks. Proceedings of the AAAI Conference on Artificial Intelligence. External Links: Document Cited by: §1, §2.1, §2.3, Table 1, Figure 6, Figure 6, §4.2, Table 4.
  • [55] H. Xu, J. Zhang, J. Cai, H. Rezatofighi, and D. Tao (2022) GMFlow: learning optical flow via global matching. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8111–8120. External Links: Document Cited by: Table 3.
  • [56] H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger (2023) Unifying flow, stereo and depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), pp. 13941–13958. External Links: Document Cited by: §2.3, Table 1, Table 4.
  • [57] X. Xu, L. Siyao, W. Sun, Q. Yin, and M. Yang (2019) Quadratic video interpolation. In Advances in Neural Information Processing Systems, Vol. 32, pp. 1645–1654. Cited by: §1.
  • [58] L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao (2024) Depth Anything: unleashing the power of large-scale unlabeled data. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 10371–10381. External Links: Document Cited by: §1, §2.2.
  • [59] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao (2024) Depth Anything V2. In Advances in Neural Information Processing Systems, Vol. 37, pp. 21875–21911. External Links: Document Cited by: §1, §2.2.
  • [60] A. R. Zamir, T. Wekel, P. Agrawal, C. Wei, J. Malik, and S. Savarese (2016) Generic 3d representation via pose estimation and matching. In Computer Vision – ECCV 2016, Cham, pp. 535–553. External Links: ISBN 978-3-319-46487-9, Document Cited by: §4.1, Table 2.
  • [61] X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer (2022) Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12104–12113. External Links: Document Cited by: §4.3.
  • [62] Y. Zhao, K. L. Man, J. Smith, K. Siddique, and S. Guan (2020) Improved two-stream model for human action recognition. EURASIP Journal on Image and Video Processing 2020 (1), pp. 24. External Links: ISSN 1687-5281, Document Cited by: §1.

FreeFlow: A Bias-free Hierarchical
Transformer for Optical Flow Estimation Supplementary Material

This supplementary provides additional details and context on training and evaluation protocols, metric and loss definitions, and extended qualitative and ablation results, and is organized as follows:

Appendix 0.A Definitions

In the following sections, μgt(u,v) is the target flow vector at position (u,v), μ is the predicted flow vector at position (u,v), N is the number of valid pixels in the target flow field, and [] is the Iverson bracket. Sums of the form u,v are calculated only over valid pixels.

0.A.1 Metrics

Endpoint Error.

Endpoint error (EPE) is defined as:

EPE =1Nu,vμgt(u,v)μ(u,v)2. (5)

and ranges from + at worst to 0 at best. It is adopted as the main metric for the Sintel benchmark.

One-pixel Outlier Rate.

The 1-pixel outlier rate (1px) is defined as:

1px =100Nu,v[μgt(u,v)μ(u,v)2>1]. (6)

and ranges from 100 at worst to 0 at best. It is adopted as the main metric for the Spring benchmark.

Fl-all Outliers Metric.

The Fl-all outliers metric (Fl-all score) is defined as:

Fl-all =100Nu,v[μgt(u,v)μ(u,v)2>max(0.05μgt(u,v),3)], (7)

and ranges from 100 at worst to 0 at best. Is is adopted as the main metric for the KITTI-15 benchmark due to the noisy nature of the collected real-world data.

Weighted Area Under the Curve.

The weighted area under the curve (WAUC) is formally defined as:

WAUC =2505(100Nu,v[μgt(u,v)μ(u,v)2x])5x5dx, (8)

and ranges from 0 at worst to 100 at best. In practice, this integral is usually approximated with 100 bins. WAUC can be also viewed as a generalization of 1px score.

0.A.2 Mixture-of-Laplace Loss

For a single flow vector coordinate, the Mixture-of-Laplace (MoL) in SEA-RAFT is defined as:

MixLap(μgt,α,β,μ) =log(α2e|μgtμ|+1α2eβe|μgtμ|eβ), (9)

where μgt is the target flow coordinate, μ is the predicted flow coordinate, α is the predicted mixing coefficient, and β is the predicted scale parameter, clamped to the range [0,10]. In practice and as is used in the official implementation, α and 1α are calculated as the softmax of two outputs α1 and α2. The final MoL loss is then defined as:

MoL=12Nu,vd{x,y}MixLap(μgt(u,v)d,α(u,v),β(u,v),μ(u,v)d). (10)
Refer to caption
(a) CroCo-Pro (Late Feature Fusion)
Refer to caption
(b) FreeFlow-L (Dense Feature Fusion)
Figure 9: Feature fusion ablation. We compare a DepthPro-style Late Feature Fusion baseline (CroCo backbone with multi-scale decoding) to FreeFlow with Dense Feature Fusion. Late fusion does not reliably propagate information between scales and tiles: the model tends to rely on the finest-level tiles and under-utilize coarse-scale features, leading to globally inconsistent flow despite sharp local detail (left). Dense fusion exchanges features throughout the network, yielding predictions that are both locally detailed and globally consistent (right).

Appendix 0.B Architecture Discussion

To isolate the role of feature fusion, we implemented a DepthPro-style Late Feature Fusion baseline by replacing the monocular ViT backbone with a pretrained CroCo encoder–decoder and fusing pyramid features in a DPT-decoder (which we call CroCo-Pro). As shown in Fig. 9(a), this design does not reliably propagate information across scales and tile boundaries: coarse-scale cues are weakly utilized and have limited influence on the final full-resolution prediction. The resulting flow can preserve fine local structure, yet exhibits reduced global coherence, especially when displacements span multiple tiles or when texture is ambiguous.

FreeFlow addresses this limitation by exchanging features throughout the network rather than only at the final decoding stage (Fig. 9(b)). Shifted-window blocks provide repeated cross-tile communication, while reduced-resolution global attention injects long-range context, allowing coarse and fine signals to reinforce each other during decoding. Beyond accuracy, this design is practical at high resolution: all interaction mechanisms are expressed via standard attention operators (windowed, shifted-window, and global attention on a downsampled map), whose token counts are balanced so that their computational cost is comparable. As a result, FreeFlow directly benefits from off-the-shelf efficient attention kernels that have been heavily optimized in recent years, whereas flow-specific modules such as correlation volume indexing, feature warping, or iterative update pipelines require task-specific engineering to reach similar efficiency.

Appendix 0.C Additional results

0.C.1 Training Details

We use AdamW with β1=0.9 and β2=0.95. For pretraining, we adopt a linear warmup followed by a cosine learning-rate decay. For optical flow finetuning, we use a linear warmup followed by a linear decay schedule, which is standard in optical flow training. We train with the Mixture-of-Laplace (Sec. 0.A.2) loss from SEA-RAFT.

媒体内容 · 前往原文查看
Table 7: Patch size ablation. We compare FreeFlow-S with a ×16 variant at matched 1080p inference time. To account for the larger image area represented by each token, we increase the variant’s width and number of attention heads. The base model still achieves better 1px accuracy with much lower memory and parameter cost.
Patch size Pre-train crop Mask. ratio Width Inf. Cost (1080p) Params (M) Spring (sub-val)
Memory (GB) Time (ms) 1px EPE
8×8 [224, 224] 0.95 256 1.02 144 34.58 0.181 0.709
16×16 [256, 256] 0.9 768 2.91 120 278.88 0.174 0.798

0.C.2 Patch Size Ablation

We investigate the role of the patch size on our method’s performance. As the base ×8 model, we take FreeFlow-S. For the ×16 model, to make for a fair comparision, we increase the width from 256 to 768 and the number of attention heads from 4 to 12, resulting in similar execution time.

Tab. 7 shows that the ×16 variant is only marginally better in EPE (0.174 vs. 0.181), while being worse in 1px (0.798 vs. 0.709) and substantially more expensive in parameters and memory (279M and 2.91 GB vs. 35M and 1.02 GB). Overall, this supports using smaller patches in our setting, as they provide comparable accuracy at a substantially lower memory and parameter cost.

0.C.3 Comparison to Vanilla ViT

To study the performance of the proposed architecture separately from the used training procedure, we trained a variant in which FreeFlow encoder/decoder modules are substituted for vanilla ViTs. This model closely resembles Croco-Flow’s and Win-Win’s designs (except for the DPT-head used in the postprocessing stage). The results are presented in  Tab. 8 (rows 1–3). Even within a larger computational budget (270–410 ms vs 131 ms), the ViT-based model results in a larger prediction error, while the quality difference between the ViT and FreeFlow models with similar parameter counts is negligible (with the ViT model being almost 6 times slower). This confirms that the proposed Local-Global attention design, not the training recipe alone, improves on the quality-performance tradeoff of a standard ViT-based approach.

0.C.4 Injecting Biases back into FreeFlow

We modified FreeFlow to include explicit optical flow biases, testing a GeoViT-like iterative warping procedure as it is straightforward to integrate and currently among the most effective iterative transformer-based approaches. We do not change the model during pretraining, only introducing warping and the recursive module (the same ConvGRU as in GeoViT) during the optical flow fine-tuning stage. Results are provided in  Tab. 8 (rows 6–8). The GeoViT-like variant shows slightly better prediction quality within the same parameter budget in some configurations, but at the cost of slightly slower inference due to its iterative nature. This is consistent with expectations: the inductive bias offloads task-specific knowledge into an explicit operator, increasing the effective capacity of the model. This further illustrates that flow-specific biases are not necessary to reach SOTA performance, but can be added to FreeFlow to improve accuracy at the cost of speed and architectural generality.

媒体内容 · 前往原文查看
Table 8: Comparison with close alternative approaches, results on Spring (sub-val) after Spring (sub-train). All models have a width of 256 with 4 attention heads (except for FreeFlow-S+, which has 8), used a masking ratio of 0.95 during pre-training, and followed the training procedure described above. The "iters" column indicates the number of iterative refinements used in GeoViT warping (during both training and inference). "+" splits time into the transformer and the postprocessing head parts.
Type Layers Iters Time (ms) (1080p) Params (M) Spring (sub-val)
1px EPE
ViT 4 278+13 15.36 1.021 0.217
6 415+13 19.05 0.809 0.216
12 829+13 30.11 0.698 0.192
FreeFlow-S 4 131+13 34.58 0.709 0.181
FreeFlow-S+ 8 256+13 60.90 0.624 0.157
FreeFlow-S with GeoViT warping 4 1 131+40 33.12 0.660 0.165
2 2 131+62 19.96 0.719 0.171
4 2 262+53 33.12 0.583 0.148
媒体内容 · 前往原文查看
Table 9: Zero-shot comparison. We report zero-shot evaluation results on the Sintel and KITTI-15 training sets. By default, all methods are trained or fine-tuned for optical flow estimation on (FlyingChairs +) FlyingThings3D, with the "TA" column indicating that TartanAir was additionally used. Method biases abbreviations: IR: Iterative Refinement, CV: Correlation Volume, W: Warping. The pre-train column indicates whether some part of the method was not trained from scratch, size is reported as the total number of frames in the dataset (for CroCo, this is double the number of image pairs, for Kinetics, this is the total number of frames in all videos). "MF" indicates that the method used multiple frames (three or more) to generate it’s submissions.
Method Biases Pre-train TA Sintel (train) KITTI-15 (train)
IR CV W Name Size Clean Final Fl-epe Fl-all
RAFT × × 1.43 2.71 5.04 17.4
GMA × × 1.30 2.74 4.69 17.1
FlowFormer × ImageNet-1K 1.3M × 1.01 2.40 4.09 14.7
SEA-RAFT (S) × ImageNet-1K 1.3M 1.27 3.74 4.43 15.1
SEA-RAFT (M) × ImageNet-1K 1.3M × 1.21 4.04 4.29 14.2
SEA-RAFT (L) × ImageNet-1K 1.3M × 1.19 4.11 3.62 12.9
DPFlow × × 1.02 2.26 3.37 11.1
VideoFlow-BOF(MF) × ImageNet-1K 1.3M × 1.03 2.19 3.96 15.3
VideoFlow-MOF(MF) × ImageNet-1K 1.3M × 1.18 2.56 3.89 14.2
MemFlow(MF) × × 0.93 2.08 3.88 13.7
MemFlow-T(MF) × ImageNet-1K 1.3M × 0.85 2.06 3.38 12.8
StreamFlow(MF) × ImageNet-1K 1.3M × 0.87 2.11 3.85 12.6
MEMFOF(MF) × ImageNet-1K 1.3M × 1.10 2.70 3.31 10.1
MEMFOF(MF) × ImageNet-1K 1.3M 1.20 3.91 2.93 9.9
ARFlow(MF) × ImageNet-1K 1.3M 0.88 2.07 2.86 9.2
WAFT-Twins-a2 × ImageNet-1K 1.3M 1.02 2.46 2.98 9.9
WAFT-DAv2-a2 × DAv2 63M 1.01 2.49 3.28 10.9
WAFT-DINOv3-a2 × LVD-1689M 1.7B 1.28 2.56 3.49 12.9
GeoViT × Kinetics-400 59M × 0.69 1.78 3.15 11.5
CroCo-Flow × × × CroCo v2 15M × 1.28 2.58
FreeFlow-S (ours) × × × CroCo v2 7.4M 0.91 3.16 3.41 10.4
FreeFlow-M (ours) × × × CroCo v2 7.4M 1.01 3.12 5.89 14.6
FreeFlow-L (ours) × × × CroCo v2 7.4M 1.04 2.30 4.77 12.9
Refer to caption
Figure 10: Visualization of the softmax logits for different parts of the proposed local-global attention. The red point represents the query token.

0.C.5 Zero-shot Performance

We evaluate zero-shot transfer on the Sintel and KITTI training sets by finetuning only on TartanAir (TA) and FlyingThings3D. Specifically, we remove the Sintel, KITTI, and HD1K portions from the TaTSKH / TaTSKH-hq stages and report performance after completing the “hq” stage (Tab. 9).

FreeFlow attains reasonable zero-shot performance on Sintel and KITTI under TA+Things finetuning, but does not improve monotonically with model size. This behavior is expected for a bias-free model in a data-limited regime: transfer is primarily constrained by the motion and appearance coverage of the available training signal, and increasing capacity alone does not guarantee gains.

This trend is reflected by the methods that use broader pretraining datasets. GeoViT, pretrained on the large Kinetics-400 video dataset, achieves the strongest zero-shot performance on Sintel, consistent with exposure to substantially more varied motion. In contrast, CroCo-Flow, pretrained only with CroCo-style data and without the use of TA during training, exhibits weaker transfer. Finally, the WAFT variants suggest that large-scale monocular pretraining alone is not always sufficient for zero-shot optical flow: despite substantially larger pretraining datasets, their transfer does not match methods pretrained on binocular or video data, indicating the importance of multi-view and motion-centric pretraining signals.

0.C.6 Additional Model Analysis

We show visualizations for different attention types in Fig. 10: global attention helps FreeFlow capture large displacements, while local attention processes small shifts. On shifted-window attention in the decoder, the window partition is shifted identically in both frames so that corresponding regions remain co-located across the pair.

Appendix 0.D Additional Qualitative Comparisons

We provide additional qualitative samples for Sintel (Fig. 11) and KITTI-15 (Fig. 12) datasets.

Refer to caption
Figure 11: Additional qualitative samples on Sintel. Across the board FreeFlow models produce the sharpest details and effectively separate the wooden structure from the background. Images are sourced from the official leaderboard webpages. Viewer is advised to zoom in.
Refer to caption
Figure 12: Additional qualitative samples on KITTI-15. FreeFlow has significantly more details on complex objects such as the wheels of bicycles, side-view mirrors of cars and tree foliage. Images are sourced from the official leaderboard webpages. Viewer is advised to zoom in.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org