We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
DF26 基准:现有检测手段已难以分辨 AI 生成视频与真实视频
AI 导读
DF26 基准针对单人物公开演讲场景评测 AI 生成视频检测,包含 271 个真实视频和由七个现代视频模型生成的 2,420 个合成视频。结果显示人类和最先进的 deepfake 检测器准确率都接近随机水平,暴露出当前评测协议的局限,并提出了对现代生成模型分布偏移保持鲁棒性的基准需求。
HuggingFace Daily Papers(社区热门论文)
54
AI 编辑部评分,满分 100DF26 基准:现有检测手段已难以分辨 AI 生成视频与真实视频
DF26 基准针对单人物公开演讲场景评测 AI 生成视频检测,包含 271 个真实视频和由七个现代视频模型生成的 2,420 个合成视频。结果显示人类和最先进的 deepfake 检测器准确率都接近随机水平,暴露出当前评测协议的局限,并提出了对现代生成模型分布偏移保持鲁棒性的基准需求。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org