跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 2 天前AI 评分37

TT-VidT:解耦时间轴的高效运动中心视频预训练

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

AI 导读

TT-VidT 通过将 DINOv3 初始化的 ViT-B/16 逐帧空间路径与紧凑 Temporal Transfer Layer 结合,用 Diff Compression 从首帧外观锚点和逐帧运动 token 重建目标帧。

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org