三秒钟的盗窃:为什么 AI 语音诈骗跑赢了所有防御手段

Sharon Brightwell 听到电话那头传来女儿的哭声,那一刻,她可能构筑的任何防线都崩塌了。那声音属于 April,至少每一种本能都这样坚称:同样的音色,同样是一个身处困境的年轻女子那种断续的节奏。那声音说她一边开车一边发短信,撞上了一名孕妇,手机被警方扣了。
随后一名男子接过电话,自称是 April 的律师,解释说保释需要一万五千美元现金。他警告 Brightwell 不要告诉银行这笔钱的用途,因为这可能会损害她女儿的信用。不到一个小时,这位来自佛罗里达州多佛的退休老人就取出了钱,交给了一名她以为与法院有关的快递员。
直到她联系上真正的 April——她整个上午都在上班,从未靠近过任何车祸现场——她才明白,那通电话并不是她女儿打的。根本不是人打的。那哭声是从一段音频片段合成出来的,她以为自己正在营救的女儿,只存在于别人机器里的一串数字模式之中。
Brightwell 的损失,在 2025 年夏天被美国各地的地方新闻报道,如今已成为美国最普通不过的犯罪之一。它同时也是技术含量最高的犯罪之一。这两个事实的碰撞——一场需要机器学习绝对前沿技术的诈骗,竟能针对一位在自家厨房里的普通祖母,大规模实施,且成本几乎为零——正是执法部门、银行、电信公司和监管机构两年来始终未能遏制的一个问题的核心特征。
问题已不再是这项技术是否管用。它管用得出奇地好。问题是,当攻击的精密度与目标人群的认知之间的差距不是以月计而是以年计时,有意义的防护究竟需要什么。
一本二十六年账本中的新增一行
2026 年 4 月,FBI 的互联网犯罪投诉中心发布了关于上一年网络犯罪的年度报告,在这份报告二十六年来的历史中,首次将人工智能驱动的诈骗作为一个独立类别单独列出。数字触目惊心。该局记录了超过 22,000 起与 AI 相关的投诉,调整后损失超过 8.93 亿美元。
其中,报告将 3.52 亿美元的损失归因于 60 岁及以上的受害者,使老年人成为 AI 驱动的金融犯罪中受 targeting 最严重的人群。AI 相关数字置身于一个远为庞大的总量之中:全美网络犯罪损失在一年内上升了 26%,达到 209 亿美元,其中 60 岁及以上的美国人占了 77 亿美元——较上一年跃升约 60%。
FBI 坦承,即便是这些数字也低估了问题的严重性。报告中对 AI 的归因只反映了受害者识别并报案的部分,而大多数克隆语音电话的受害者根本无从得知背后有机器参与。他们相信——正如 Sharon Brightwell 起初也相信的那样——自己是在和自己的亲生孩子说话。
因此,8.93 亿美元最好被理解为一个下限,而非上限——它只是这个类别中可见的那一部分,而这类骗局在本质上就是设计成对受害者保持隐形的。FBI 竟然觉得有必要专门设立这一类别,这本身就是一种信号。犯罪统计数据向来是保守的工具;机构不会为了一时的风潮去重新划定沿用二十六年的报案分类体系。
账本上新增的这一行,等于承认了一种三年前在消费级形态下几乎不存在的工具,如今已成为主流的盗窃手段。
在国际层面,情况更为庞大且持续恶化。2026 年 3 月,INTERPOL 发布了第二版《全球金融欺诈威胁评估》,估计 2025 年全球金融欺诈损失达 4420 亿美元——这一数额相当于丹麦全年的经济产出。该组织将威胁趋势评定为不断升级,并描述了其所称的“欺诈工业化”:诈骗从机会主义的个体行为,演变为有组织的跨国运作,并与人口贩运和网络犯罪相互交织。
关键在于,INTERPOL 发现,AI 增强型欺诈的获利能力大约是传统欺诈的四倍半,而且所谓的智能体 AI 系统如今能够自主策划并执行完整的欺诈行动,从前期侦察一直到勒索赎金。换言之,经济逻辑已经反转。工业化规模的欺骗行为,首次变得几乎零成本制造,却能带来巨额回报。
三秒足矣
祖父母诈骗的核心技术能力,描述起来简单得近乎残酷。一套现代 AI 语音克隆系统只需短短三秒音频,就能生成一段合成语音,在实际效果上与真人原声几乎无法区分。三秒,不过是一条语音留言问候的长度,一段播客的片段,或是发布在公开 Instagram 账号上的生日视频里的音频。
这些原材料并非从安全数据库中窃取,而是人们日复一日通过记录生活的寻常行为主动交出的。一个仅在一段 TikTok 短视频中露过面的孙辈,就已经为诈骗者提供了制造其自导自演绑架案所需的一切。
让这一威胁变得尤为严峻的,不仅是克隆技术本身可行,更在于实现克隆的工具廉价、泛滥,且几乎完全不受监管。2025 年 3 月,Consumer Reports 评估了六家公司——Descript、ElevenLabs、Lovo、PlayHT、Resemble AI 和 Speechify——的语音克隆产品,结论是其中大多数缺乏任何有意义的防欺诈或防滥用保障措施。
该组织发现,其中四款产品仅要求用户勾选一个复选框,声明自己拥有克隆相关语音的合法权利。这四款产品中,没有任何一款采用技术机制来确认说话者是否真正同意,也没有任何一款将克隆限制在用户本人的声音范围内。六家公司中有四家仅需一个姓名或一个电子邮箱地址即可开设账户。
这项调查得出的直白结论——经 NBC News 和 The Register 放大传播——是:这个行业造出了一款能够冒充任何人的工具,然后把它放在一个自我声明的复选框后面。
ElevenLabs 是最知名的提供商之一,它指出了一套多层安全方案:一项禁止冒用他人身份的使用政策、一个公开的 AI 语音分类器(可识别很可能源自其系统的音频)、可将生成内容追溯回制作它的账户的可追溯机制,以及在选举周期前后阻止克隆某些受保护人物的“禁区声音”保障措施。
这些并非无足轻重的举措,而且比若干竞争对手提供的还要多。但它们有一个共同的结构性弱点:几乎所有这些措施都是在事后才发挥作用。它们能帮助调查人员在欺诈已经发生、受害者已经损失积蓄之后确定来源。它们对于从一开始就阻止那三秒克隆音频被生成几乎毫无作用,因为真正能阻止它的东西——对被克隆者已表示同意的强有力、强制性验证——恰恰是一个竞争激烈、快速变化的市场不愿强加于自身的那种摩擦。
当一项保障措施会损害公司的转化率、却只保护竞争对手的客户时,市场不会自愿提供它。事实也确实没有。
失明的取证权威
如果说有哪一个瞬间能说明为什么基于检测的防御正在失效,那便是《纽约时报》于 2026 年 6 月刊登的一篇人物特写。报道的主角是 Hany Farid,加州大学伯克利分校教授,被广泛公认为全球深度伪造取证领域的头号权威。
二十多年来,Farid 凭借辨别真实与合成内容的能力建立了自己的职业生涯,承接来自政府、人权组织、记者和执法部门的请求。据《纽约时报》报道,近来他开始无法通过自己的测试。“我觉得自己正在失明,”他说。地球上最有能力分辨真实录音与 AI 生成录音的人,再也无法可靠地做到这一点。
这一坦白理应终结某一类讨论。多年来,针对合成媒体的应对措施所隐含的承诺一直是:检测会与生成保持同步——每出现一个更逼真的伪造品,就会有一个更灵敏的检测器,而这场军备竞赛虽然令人不安,至少是可以打赢的。Farid 的坦白证明,至少在音频领域,这场竞赛已经输了。
当该领域最顶尖的检测器退化到如同掷硬币,在伪造品被制造并传播之后再将其抓出的策略就根本算不上策略。它只是一种希望。而一场依赖二十分钟恐慌的诈骗,不会给受害者或其银行留下二十分钟去跑一遍连 Hany Farid 都不再信任的取证分析。
这是有意义的防护要求我们接受的第一件、也是最重要的一件事:检测不能成为承重的防线。一位接到电话、听到泣不成声的祖母,不能指望她去完成连该领域顶尖专家都已实质上放弃的取证分析。任何最终依赖于目标对象或任何其他人能够分辨真实声音与克隆声音的方案,都已经过时了。
其影响远不止电话诈骗。如果世界上合成音频检测领域的权威都无法信任自己的判断,那么每一个默默假设人类可以作为兜底验证者的下游系统——被要求"自行斟酌"的银行柜员、被催促"仔细听有没有哪里不对劲"的亲属——都建立在一个早已崩塌的基础之上。
脆弱性的架构
将老年人作为目标归因于天真,这种想法既诱人又错误。促成这篇文章的简报揭示了一个更令人不安的真相:使老年人格外脆弱的那些特征,并非智力上的缺陷,而是一生好好生活所留下的印记。他们往往拥有更高的平均储蓄余额,这是数十年工作积累的产物,这使他们成为高效的目标——一次成功的通话所能带来的收益,远远超过针对年轻人的一次通话。
他们成长于、并且至今仍运作于基于信任的既定沟通模式之中,在这种模式里,来自一位陷入困境的亲属的电话会被当作真实的紧急情况来回应,而不是被当作潜在攻击来审视。他们并非自己的过错,却相对不熟悉 AI 语音合成的存在,因为他们一生中大部分时间所处的世界里,电话那头的声音按定义就是电话那头的人。
而且,像每一位父母和祖父母一样,他们也暴露在家庭紧急情况这一特定情感架构之下——在这种情境中,保护孩子的本能会压倒一切更缓慢、更具怀疑精神的官能。
学术研究已开始将这一现象正式化。2026 年 6 月发表的一篇 arXiv 论文直言不讳地指出,“老年人仍然不成比例地容易受到 AI 增强型诈骗的侵害。”另一项由 Yixin Zou 领导的团队开展、同样于 2026 年初发表的研究,考察了在 AI 日益复杂化背景下专为老年人设计的欺诈干预措施,开发了一款名为 ROLESafe 的基于角色的模拟工具,该工具让参与者通过扮演受害者或帮助者的角色来学习,而非被动观察,从而提升了他们识别欺诈的能力。
还有第三篇论文,来自 Charm Security 公司的研究人员,提出了一个“人类脆弱性与利用框架”——一个结构化的目录,仿照软件安全领域的漏洞数据库,用于对欺诈系统所利用的认知和社会机制进行分类。该框架的前提本身就是一种无声的控诉:安全行业花费了数十年对机器的弱点进行编目和修补,却让人的弱点始终未被记录、未被管理。
祖父母诈骗之所以得逞,是因为它攻击的是系统中从未有人为其编写过补丁的那一部分。
这正是为什么针对老年人的宣传活动虽然必要,却远远不够。骗局所利用的情感机制,并不是一份传单就能填补的知识空白;它是某个人对孙辈的爱,被武器化了。你可以一百遍地告诉一个人声音可以被伪造,但当克隆的声音尖叫着求救的那一刻,这份知识不会及时到来。
2026 年 6 月,《海峡时报》和《马尼拉时报》转载的法新社通讯稿引用了语音安全公司 Pindrop 的阿米特·古普塔对此事的精准概括:“目标不是完美的声音复制。目标是制造足够的情感不确定性和紧迫感,让受害者在核实之前就采取行动。”
一种建立在“受害者会去核实”这一假设之上的防御,恰恰是针对攻击所设计绕过的那个弱点而建立的防御。
那篇通讯稿中最令人不寒而栗的证词并非来自一位年长的受害者,而是来自一位律师。加里·希尔德霍恩是费城的一名律师,他本人就曾是克隆声音骗局的目标,他说即便事后回想、带着职业性的怀疑,他也无法摆脱自己所听到之物的那种确定性:“我会到死都发誓那是你的声音。”
任何倾向于相信警惕就是答案的人,都应该读一读这句话。希尔德霍恩是一位受过训练的辩护人,他的职责就是审视证据、不相信看似合理的故事,而克隆声音击败他的彻底程度,与击败一位惊慌失措的祖母别无二致。这一骗局所利用的脆弱性,并不仅仅存在于老年人、轻信者或技术文盲身上。
它存在于人类听觉系统本身——这个系统经过数千年的演化,将识别出的声音视为识别出的人的证明——而如今,在这漫长历史中第一次,它系统性地、可被利用地出错了。
关于老年人的数据支持而非反驳这一重新框定。FTC 在 2025 年 12 月提交国会的报告发现,60 岁及以上人群报告的欺诈损失总额在 2020 年至 2024 年间大约翻了两番,达到约 24 亿美元,其中 68% 归因于单笔 10 万美元或以上的个人损失。
该机构自己对真实年度成本的估算——考虑到羞耻和尴尬必然导致的长期低报——高达 815 亿美元。这些数字不属于一个轻信的少数群体被一点点骗走零花钱的情形。这些数字属于一代人积累的储蓄,正通过一种专门针对他们的信任模式、财务状况以及他们在家庭情感中心位置而精心调校的机制被榨干。
不对称性,量化呈现
安全公司 Adaptive Security 的首席执行官 Brian Long 用一句话为 AFP 提炼了这种新经济逻辑:“一个人待在一个房间里,有一块键盘,就能制造出无限多的攻击者。”这就是最纯粹形式的不对称。一边是一个自动化系统,能在几秒内生成一个以假乱真的克隆,拨打数千个号码,用合成情感进行每一次对话,而边际成本趋近于零。
另一边是一个个体的人,往往是老年人,孤独一人,只有大约一通惊慌失措的电话的时间来组织防御——而这是世界上最顶尖的法医科学家都无法做到的。
国际刑警组织发现,AI 增强型诈骗的获利能力是传统诈骗的四倍半,这正是这种失衡在财务层面的体现。当一次攻击既变得更廉价、又变得更有利可图时,攻击的数量不会线性增长,而是会爆炸式增长。美国网络犯罪损失在单一年份跃升 26%,六十岁以上人群的损失几乎翻倍,这就是那场爆炸在一国账本上的样子。
而法新社的报道还指出了另一点加剧伤害的因素:羞耻感。布法罗的母亲 Liz Benz 曾经历她所称的“足足二十分钟的恐惧”——一个克隆声音告诉她,她十六岁的儿子被绑架了——她说,在她公开此事之后,其他受害者的信息如潮水般涌来,其中许多人选择匿名,因为被骗的羞辱让他们保持沉默。
在此,低报并不是一个统计脚注。它是一种结构性特征:这种犯罪被设计成让受害者觉得自己太愚蠢而不愿站出来,而这反过来又让数据、起诉和政策应对失去了它们所需的证据。一种能让自己的证人噤声的犯罪,是一种会以复利不断累积的犯罪。
为什么这一负担不能落在家庭身上
在整个 2026 年,被 FBI、美国银行家协会和消费者权益倡导者反复传播的一条最广为流传的建议,是约定一个家庭“安全词”——一个只有亲属知道的秘密短语,在任何紧急来电中都要被要求说出。如果对方的声音说不出它,就挂断,并用已知号码回拨。这条建议是合理的。但作为一种系统性防御,它也极其不够,值得把原因说清楚。
安全词只有在家庭的每一位成员都采用它、记住它,并且在那个被精心设计来摧毁冷静判断力的时刻仍有镇定去要求对方说出它时,才会奏效。它把击败一个工业化、自动化、耗资数十亿美元的犯罪机器的全部重担,压在了一个惊恐万分的个体在其一周中最糟糕时刻的认知自律上。
它假定那位八十岁的老人——根据 FTC 于 2025 年底发布的数据,其报告损失的中位数超过 1600 美元——在听到孙辈尖叫的同时,还能冷静地回忆起一套协议并加以执行。有些人做得到。而许多人,从设计上就注定做不到。一种只有在目标于极限压力下表现完美时才奏效的防御,不是防御;它是一种事后把责任归咎于受害者的方式。
这是对把保护交到家庭和个人手中更深层的反对理由。它把技术和金融系统失灵的责任,转嫁到最无力承担的人身上,然后当他们失败时,又把他们的失败当作个人失败来对待。声音克隆工具是由公司制造并销售的。通话是由电信网络承载的。资金是通过银行流转的。
这些各方都运营在那些本可以大规模拦截诈骗的关键节点上。而在厨房里的那位祖母并不掌握这些节点。有意义的保护要求把负担从链条的末端——它目前所在的位置——转移到它本该属于的中间环节。一个面对工业化威胁却只是向其最脆弱的成员发布更好建议的社会,已经把发布指南与提供保护混为一谈。
拦截真正可能发生的地方
接下来依次考察这三个制度性瓶颈,因为每一个都既体现了结构性防御的前景,也暴露了它当下的失败。
第一个是电话网络。在美国,STIR/SHAKEN 框架本意是解决来电显示欺骗问题,它允许发端运营商对通话进行加密签名以证明其合法性,并让终端运营商在通话到达手机之前验证该签名。2025 年 12 月,FCC 有线竞争局在其三年一度的有效性报告中得出结论:该框架在正确应用时确实能有效验证来电显示。
这个限定条件承担着极其重要的作用。犯罪分子很早就发现,通过较老的非 IP 网络路由通话可以完全绕过该系统,而 FCC 在 2025 年和 2026 年的大部分时间里都在试图堵上这一漏洞,并推动 Rich Call Data,即在手机上显示经过验证的来电者姓名和标志。
但 STIR/SHAKEN 验证的是号码,不是人,更不是声音。它能告诉你一通电话确实来自某条线路。它无法告诉你那条线路上哭泣的女儿是一台机器。面对来自伪造号码或仅仅是陌生号码的克隆声音,这套框架几乎毫无用处。同一个 FCC 在 2024 年 2 月宣布,根据《电话消费者保护法》,机器人电话中由 AI 生成的声音是非法的——这是一份有意义的意向声明,但它只管束大规模自动拨号,而管不到定义"祖父母骗局"的那种有针对性的、一对一的紧急电话。
第二个瓶颈是承载克隆工具的平台。在这里,最有前景的结构性干预手段是来源溯源而非检测——这一尝试体现在 C2PA 标准中,旨在为一段媒体附加一份防篡改的加密记录,描述生成它的设备或模型。2025 年,该标准被批准为 ISO 规范,并与 Google 的 SynthID、Meta 的 AudioSeal 等水印方案一道,成为互联网事实上的来源溯源语言。
其逻辑是合理的:与其追问一段录音看起来或听起来是否伪造,不如追问它是否携带了证明其由真实麦克风录制的凭证。但针对这一特定犯罪,来源溯源自身存在一个致命的不对称性。缺失凭证并不能证明是伪造,因为大量合法音频从未被签名,而且元数据经常被社交平台、截图和重新编码所剥离。
更根本的是,一个拨打诈骗电话的骗子没有任何义务传输 C2PA 清单,而模拟信号缺口——通过电话线播放合成音频——在传输过程中就会摧毁任何数字水印。来源溯源可以在事后帮助确认一段病毒式传播的视频是合成的。但对于正在进行的实时克隆语音通话,它几乎无能为力。
第三个卡点——也是最有希望的一个——是银行。这里是资金真正流动的地方,因此也是拦截在机制上最具杠杆效应之处。英国提供了最清晰的自然实验。2024 年 10 月,支付系统监管机构(Payment Systems Regulator)将授权推送支付欺诈的赔付设为强制要求:当受害者被诱骗授权向诈骗者转账时,转出银行和接收银行现在必须向其赔付,双方各承担一半责任,最高 85,000 英镑,须在五个工作日内完成,且对弱势客户明确禁止适用消费者过失例外条款。
账户名称核查服务 Confirmation of Payee 如今已运行于数十亿笔交易之上。PSR 自己的仪表盘显示,在截至 2025 年 12 月底的十五个月里,此类诈骗造成的资金损失中,有 89%——约 2.43 亿英镑——获得了赔付,而规则生效前这一比例为 65%。
强制赔付的意义不仅仅在于让受害者得到完整补偿,尽管这一点很重要。它更深层的目的是重新定位经济激励。一旦银行对损失负有责任,它们就有了强大的动力去构建那些真正能在欺诈转账完成之前将其阻止的摩擦机制、异常检测和干预协议——暂缓支付、对老年客户大额取现设置冷静期、由支行人工致电询问为什么一位退休老人突然要把她的积蓄全部转给一个快递员。
一位老奶奶无法对自己施加的转账摩擦,银行可以对她的账户施加,而责任制度给了银行这样做的理由。批评者,包括《Electronic Payments International》的评论员,警告说如果不与预防措施相配合,仅靠赔付有可能变成对欺诈者的补贴,而且这一策略必须超越仅仅向受害者退款。
这一批评是正确的,它指向的是正确答案而非偏离答案:责任是迫使预防发生的杠杆,而不是预防的替代品。英国实验的教训不是赔付能解决欺诈——它不能——而是它改变了欺诈是谁的问题,而一个被迫承担某个问题的机构,最终会设计出针对该问题的解决方案。
真正有意义的保护究竟需要什么
把这些线索串联起来,就能勾勒出一幅清晰的图景,说明什么才是真正有效的,而非仅仅听起来令人安心。
这首先要求放弃将检测作为主要防线。Hany Farid 的盲区并非一个靠更好的分类器就能解决的暂时性挫折;它是这样一个世界的永久结构性状况:在这个世界里,生成已经跑赢了辨别。任何最终保障措施是寄望于某个人、在某个地方分辨真假的做法,都已经失败了。
这其次要求对武器的供应进行监管。《消费者报告》发现语音克隆工具仅靠一个自我声明的勾选框来把关,这是一种政策选择,而非自然法则。在克隆某人的声音之前强制要求可验证的同意——Descript 和 Resemble AI 已部分实现,而其他公司尚未做到——在技术上可行,且不会废除该技术的正当用途。
欧盟的《人工智能法案》,其对通用目的模型的义务已于 2025 年和 2026 年陆续开始适用,以及田纳西州的 ELVIS 法案等州法规——该法案要求克隆声音须获得书面同意——都是将语音合成视为其已然成为的受监管能力的早期姿态。
它们仍远远领先于执法,又远远落后于威胁。
第三,它要求把拦截的责任压到占据咽喉要道的机构身上——而在它们不愿主动作为的地方,通过法律责任强制其行动。英国的分担赔偿制度并不完美、也不完整,但它展示了这一机制:当银行承担损失时,银行就会构建摩擦。从转接电话中获利的电信运营商,应当承担相应的义务去验证来电,并在可能的情况下对其加以标记。
销售克隆技术的平台,应当承担核实同意的义务。共同的原则是:责任应当落在既有能力防止伤害、又从造成伤害的活动中获取商业利益的一方身上——而在每一种情况下,这一方都是某个机构,在任何情况下都不是一位接电话的八十岁老人。
第四,它要求把人的脆弱性当作一件需要管理的事,而不是一件需要被指责的事。Charm Security 框架的洞见——欺诈所利用的认知与情感机制,值得像软件缺陷一样被系统地编目——应当指导银行如何设计其干预措施,指导运营商如何设计其警告,指导公共机构如何设计教育:这种教育要超越传单,达到 ROLESafe 研究发现真正能改变行为的那种基于角色的演练。
宣传活动和家庭安全暗语并非毫无价值。它们只是最后、也最薄弱的一道防线,仅作为承担真正工作的结构性防御的兜底而有用。
防御差距以年而非月来衡量,其最深层的根源在于:攻击是一个技术问题,而防御是一个制度问题。克隆一段声音只需三秒,而且每月都在进步。而通过一项报销法规、重构一家银行系统的欺诈控制、在一个行业内强制实施同意验证、堵上国家电话网络中非 IP 的漏洞,都需要数年时间,因为它们需要法律、协调、资金,以及克服每一个从现状中获利的商业利益。这种不对称不仅仅是技术上的;它是机器能被建造出来的速度与社会能作出回应的速度之间的不对称。
Sharon Brightwell 最终拿回了一部分钱,靠的是调查人员和她的银行的尽职努力。许多人则拿不回来。FBI 归因于 AI 欺诈老年受害者的 3.52 亿美元,以及它无法看到的远大于此的金额,代表着一种财富转移——从最承受不起损失的人,转移到迄今所设计出的最高效的犯罪产业。
弥合这一差距,不会靠教祖母们去怀疑孙辈声音的真伪。它将来自以法律和工程为准绳作出决定:站在克隆声音与现金之间的那些机构,应对经手之物负责。在这一决定作出之前,这场三秒盗窃仍将是世界上最容易实施的重罪,也是受害者最难被相信的重罪——因为那些证据,从设计上听起来,恰恰就像他们所爱的某个人。
参考文献
- 美国联邦调查局,《加密货币与 AI 诈骗令美国人损失数十亿美元》,2026 年 4 月。https://www.fbi.gov/news/press-releases/cryptocurrency-and-ai-scams-bilk-americans-of-billions
- 互联网犯罪投诉中心(FBI),《2025 年 IC3 年度报告》,2026 年 4 月。https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf
- ABA Banking Journal,《FBI:2025 年网络犯罪损失增长 26%》,2026 年 4 月。https://bankingjournal.aba.com/2026/04/fbi-cybercrime-losses-increased-26-in-2025/
- SecureWorld,《FBI:2025 年 AI 驱动的欺诈金额突破 8.93 亿美元——实际损失可能远高于此》,2026 年 4 月。https://www.secureworld.io/industry-news/ai-enabled-fraud-topped-893m-fbi
- AARP,《2025 年老年人遭受欺诈重创》,2026 年 4 月。https://www.aarp.org/money/scams-fraud/fbi-ftc-report-2025-losses/
- HousingWire,《FBI:2025 年老年人因网络犯罪损失 77.5 亿美元——激增 59%》,2026 年 4 月。https://www.housingwire.com/articles/fbi-seniors-cybercrime-2025/
- INTERPOL,《INTERPOL 报告警告全球金融欺诈威胁日益复杂精密》,2026 年 3 月。https://www.interpol.int/en/News-and-Events/News/2026/INTERPOL-report-warns-of-increasingly-sophisticated-global-financial-fraud-threat
- Help Net Security,《全球欺诈损失攀升至 4420 亿美元》,2026 年 3 月 18 日。https://www.helpnetsecurity.com/2026/03/18/online-fraud-victims-losses-interpol-report/
- ICLG,《国际刑警组织警告全球欺诈正走向"工业化"》,2026 年 3 月。https://iclg.com/news/23665-interpol-warns-of-industrialisation-of-global-fraud
- 《纽约时报》,《在 AI 时代,全球顶尖的深度伪造专家不再相信自己的眼睛》,2026 年 6 月。(经由 beSpacific 报道。)https://www.bespacific.com/the-worlds-leading-deepfake-expert-no-longer-trusts-his-own-eyes/
- FOX 13 Tampa Bay,《多佛一名女子在诈骗者利用人工智能冒充其女儿后损失 1.5 万美元》,2025 年 7 月。https://www.fox13news.com/news/dover-woman-loses-15k-after-scammers-used-artificial-intelligence-impersonate-daughter
- 法新社(经由《马尼拉时报》),《"20 分钟的恐怖":AI 助推美国语音冒充诈骗》,2026 年 6 月 4 日。https://www.manilatimes.net/2026/06/04/news/world/20-minutes-of-terror-ai-boosts-us-voice-impersonation-scams/2358259
- 《消费者报告》,《消费者报告对 AI 语音克隆产品的评估》,2025 年 3 月。https://www.consumerreports.org/media-room/press-releases/2025/03/consumer-reports-assessment-of-ai-voice-cloning-products/
- NBC News,《AI 能窃取你的声音,而你几乎无能为力》,2025 年 3 月。https://www.nbcnews.com/tech/security/ai-voice-cloning-software-flimsy-guardrails-report-finds-rcna195131
- The Register,《Consumer Reports 指出 AI 语音克隆防护措施薄弱》,2025 年 3 月 10 日。https://www.theregister.com/2025/03/10/ai_voice_cloning_safeguards/
- ElevenLabs,《安全》,2026 年。https://elevenlabs.io/safety
- Ben, A., Rahav, T., Illaev, D., Nahon, A. 和 Grushka, A.(Charm Security),《人类脆弱性与漏洞利用(HVE)框架》,arXiv:2606.10083,2026 年 6 月。https://arxiv.org/abs/2606.10083
- Deng, Y., Chen, X., Liao, J., Li, B. 和 Zou, Y.,《经历者、帮助者还是旁观者:基于角色模拟的老年人网络诈骗干预》,arXiv:2601.12324,2026 年 1 月。https://arxiv.org/abs/2601.12324
- 美国联邦通信委员会,《FCC 将机器人电话中的 AI 生成语音定为非法》,2024 年 2 月。https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal
- 美国联邦通信委员会,《通过主叫号码认证(STIR/SHAKEN)打击伪造机器人电话》,2026 年更新。https://www.fcc.gov/call-authentication
- EyeSift,《C2PA 深度伪造检测 2026:AI 图像水印与 SynthID》,2026 年。https://www.eyesift.com/ai-image-detection-2026-c2pa-content-credentials-synthid-watermarks-diffusion-fingerprints-deepfake/
- 支付系统监管局,《PS25/5 综合政策声明:APP 诈骗赔付要求》,2025 年 5 月。https://www.psr.org.uk/media/rhelv4op/ps25-5-app-scams-reimbursement-consolidated-policy-statement-may-2025.pdf
- 《国际电子支付》,"为什么英国的诈骗治理策略必须超越赔付",2025 年。https://www.electronicpaymentsinternational.com/comment/why-uk-scam-strategy-must-move-beyond-reimbursement/
- 美国联邦贸易委员会,《保护老年消费者 2024–2025:美国联邦贸易委员会报告》,2025 年 12 月。https://www.ftc.gov/system/files/ftc_gov/pdf/P144400-OlderAdultsReportDec2025.pdf
- 美国律师协会,《AI 克隆语音诈骗的兴起》,2025 年 9 月。https://www.americanbar.org/groups/senior_lawyers/resources/voice-of-experience/2025-september/ai-cloned-voice-scam/
The Three-Second Theft: Why AI Voice Fraud Outruns Every Defence

Sharon Brightwell heard her daughter crying down the line, and that was the end of any defence she might have mounted. The voice belonged to April, or so every instinct insisted: the same timbre, the same broken rhythm of a young woman in distress. The voice said she had been texting while driving, that she had hit a pregnant woman, that her phone had been seized by police. A man then took over the call, identifying himself as April's attorney, and explained that bail would cost fifteen thousand dollars in cash. He warned Brightwell not to tell the bank what the money was for, because it might damage her daughter's credit.
Within the hour, the retiree from Dover, Florida had withdrawn the money and handed it to a courier she believed was connected to the courts. Only when she reached the real April, who had spent the morning at work and never been near a car accident, did she understand that her daughter had not made the call. No human had. The crying had been synthesised from a fragment of audio, and the daughter she thought she was rescuing existed only as a pattern of numbers in someone else's machine.
Brightwell's loss, reported across American local news in the summer of 2025, is now one of the most ordinary crimes in the United States. It is also one of the most technically advanced. The collision of those two facts — that a fraud requiring the absolute frontier of machine learning can be perpetrated against an ordinary grandmother in her kitchen, at scale, for the price of nothing — is the defining feature of a problem that law enforcement, banks, telecoms companies and regulators have spent two years failing to contain. The question is no longer whether the technology works. It works appallingly well. The question is what meaningful protection requires when the gap between the sophistication of the attack and the awareness of the target is measured not in months but in years.
A New Line in a Twenty-Six-Year Ledger
In April 2026, the FBI's Internet Crime Complaint Center published its annual report on the previous year's online crime, and for the first time in the report's twenty-six-year history it broke out artificial-intelligence-enabled fraud as a distinct category. The numbers were stark. The bureau logged more than 22,000 complaints with an AI nexus and adjusted losses exceeding 893 million dollars. Of that sum, the report attributed 352 million dollars in losses to victims aged sixty and over, making older adults the single most heavily targeted demographic in AI-enabled financial crime. The AI figure sat inside a far larger total: cybercrime losses across the United States rose 26 per cent in a single year to 20.9 billion dollars, with Americans aged sixty and older accounting for 7.7 billion of that — a roughly 60 per cent jump on the previous year.
The FBI was candid that even these figures understate the problem. AI attribution in the report reflects only what victims recognised and reported, and most victims of a cloned-voice call never learn that a machine was involved at all. They believe, as Sharon Brightwell initially believed, that they spoke to their own child. The 893 million dollars is therefore best read as a floor, not a ceiling — the visible portion of a category that is, by its nature, designed to remain invisible to the people it harms. That the FBI felt compelled to create the category at all is itself a signal. Crime statistics are conservative instruments; agencies do not redraw twenty-six-year-old reporting taxonomies for a passing fashion. The new line in the ledger is an admission that a tool which barely existed in consumer form three years ago has become a mainstream instrument of theft.
Internationally, the picture is larger and worsening. In March 2026, INTERPOL published the second edition of its Global Financial Fraud Threat Assessment, estimating worldwide losses to financial fraud at 442 billion dollars in 2025 — a sum comparable to the entire annual economic output of Denmark. The organisation rated the threat trajectory as escalating and described what it called the “industrialisation of fraud”: the migration of scamming from opportunistic individuals to organised, transnational operations that intersect with human trafficking and cybercrime. Crucially, INTERPOL found that AI-enhanced fraud is roughly four and a half times more profitable than its traditional equivalent, and that so-called agentic AI systems can now autonomously plan and execute entire fraud campaigns, from reconnaissance through to the ransom demand. The economics, in other words, have inverted. For the first time, deception at industrial scale costs almost nothing to manufacture and returns a fortune.
Three Seconds Is All It Takes
The technical capability at the centre of the grandparent scam is brutally simple to describe. A modern AI voice-cloning system requires as little as three seconds of audio to produce a synthetic voice that is, for practical purposes, indistinguishable from the original. Three seconds is the length of a voicemail greeting, a snatch of a podcast, the audio under a birthday video posted to a public Instagram account. The raw material is not stolen from a secure database; it is volunteered, every day, by the ordinary act of living a recorded life. A grandchild who appears in a single TikTok clip has supplied everything a fraudster needs to manufacture their own kidnapping.
What makes the threat acute is not merely that the cloning works but that the tools to do it are cheap, abundant and almost entirely unpoliced. In March 2025, Consumer Reports assessed the voice-cloning products of six companies — Descript, ElevenLabs, Lovo, PlayHT, Resemble AI and Speechify — and concluded that a majority lacked any meaningful safeguard against fraud or misuse. Four of the products, the organisation found, required only that a user tick a box affirming they had the legal right to clone the voice in question. None of those four employed any technical mechanism to confirm that the speaker had actually consented, or to restrict cloning to the user's own voice. Four of the six companies required nothing more than a name or an email address to open an account. The investigation's blunt conclusion, amplified by NBC News and The Register, was that the industry had built a tool capable of impersonating anyone and then placed it behind a self-attestation checkbox.
ElevenLabs, one of the most prominent providers, points to a multi-layered safety programme: a prohibited-use policy that bans impersonation, a public AI speech classifier that can identify audio likely to have originated from its system, traceability that links generated content back to the account that produced it, and “no-go voices” safeguards that block the cloning of certain protected figures around election cycles. These are not trivial measures, and they are more than several competitors offer. But they share a structural weakness: almost all of them operate after the fact.
They help investigators establish provenance once a fraud has already occurred and a victim has already lost their savings. They do very little to prevent the three-second clone from being generated in the first place, because the thing that would prevent it — robust, mandatory verification that the person being cloned has consented — is precisely the friction that a competitive, fast-moving market is reluctant to impose on itself. When a safeguard costs a company conversions and protects only the customers of its rivals, the market will not supply it voluntarily.
It has not.
The Forensic Authority Who Went Blind
If there is a single moment that captures why detection-based defences are failing, it arrived in a New York Times profile published in June 2026. Its subject was Hany Farid, the University of California, Berkeley professor who is, by broad consensus, the world's foremost authority on deepfake forensics. For more than two decades Farid had built a career on the ability to separate the real from the synthetic, fielding requests from governments, human-rights organisations, journalists and law enforcement. Lately, the Times reported, he had begun failing his own tests. “I feel like I'm going blind,” he said. The man best equipped on Earth to distinguish a genuine recording from an AI-generated one could no longer reliably do so.
That admission ought to end a certain kind of conversation. For years, the implicit promise of the response to synthetic media has been that detection would keep pace with generation — that for every more convincing fake, there would be a more sensitive detector, and that the arms race, though uncomfortable, was at least winnable. Farid's confession is evidence that, in the audio domain at least, the race has been lost. When the foremost detector in the field is reduced to a coin-toss, the strategy of catching fakes after they have been made and circulated is not a strategy at all. It is a hope. And a fraud that depends on twenty minutes of panic does not give a victim, or their bank, twenty minutes to run a forensic analysis that even Hany Farid would no longer trust.
This is the first and most important thing that meaningful protection requires us to accept: detection cannot be the load-bearing defence. A grandmother on the phone with a sobbing voice cannot be expected to perform forensic analysis that the discipline's leading expert has effectively abandoned. Any plan that ultimately rests on the target, or anyone else, being able to tell the difference between a real voice and a cloned one is already obsolete. The implication runs deeper than telephone fraud. If the world's authority on detecting synthetic audio cannot trust his own judgement, then every downstream system that quietly assumes a human can serve as a fallback verifier — the bank teller who is told to “use discretion,” the relative urged to “listen carefully for anything off” — rests on a foundation that has already crumbled.
The Architecture of Vulnerability
It is tempting, and wrong, to attribute the targeting of older adults to naivety. The brief that prompts this article identifies a more uncomfortable truth: the characteristics that make older people disproportionately vulnerable are not deficiencies of intelligence but features of a life well lived. They tend to hold higher average savings balances, the accumulated product of decades of work, which makes them efficient targets — a single successful call can yield far more than one aimed at a younger person. They were raised in, and still operate within, established patterns of trust-based communication, in which a phone call from a distressed relative is answered as a genuine emergency rather than interrogated as a potential attack.
They are, through no fault of their own, relatively unfamiliar with the existence of AI voice synthesis, having spent most of their lives in a world where a voice on the line was definitionally a person on the line. And they are exposed, like every parent and grandparent, to the particular emotional architecture of the family-emergency scenario, in which the instinct to protect a child overrides every slower, more sceptical faculty.
Academic research has begun to formalise this. An arXiv paper published in June 2026 noted plainly that “older adults remain disproportionately vulnerable to AI-enhanced scams.” A separate study from a team led by Yixin Zou, also published in early 2026, examined fraud interventions designed specifically for older adults amid escalating AI sophistication, developing a role-based simulation tool called ROLESafe that improved participants' ability to identify fraud when they learned by playing the part of victim or helper rather than passive observer.
And a third paper, from researchers at the firm Charm Security, proposed a Human Vulnerabilities and Exploits Framework — a structured catalogue, modelled on the software-security world's vulnerability databases, for classifying the cognitive and social mechanisms that fraud systems exploit. The framework's premise is itself a quiet indictment: the security industry has spent decades cataloguing and patching the weaknesses of machines while leaving the weaknesses of people undocumented and unmanaged. The grandparent scam succeeds because it attacks the one part of the system for which no patch has ever been written.
This is why awareness campaigns aimed at older adults, while necessary, cannot be sufficient. The emotional mechanism the scam exploits is not a gap in knowledge that a leaflet can fill; it is the love a person has for their grandchild, weaponised. You can tell someone a hundred times that voices can be faked, and in the moment a cloned voice screams for help, the knowledge will not arrive in time. The AFP wire story carried by The Straits Times and the Manila Times in June 2026 quoted Amit Gupta of the voice-security firm Pindrop putting the matter precisely: “The objective is not perfect voice replication. The objective is creating enough emotional uncertainty and urgency that the victim acts before verifying.” A defence built around the assumption that victims will verify is a defence built against the very weakness the attack is engineered to bypass.
The most chilling testimony in that wire story came not from an elderly victim but from a lawyer. Gary Schildhorn, a Philadelphia attorney who was himself targeted by a cloned-voice scam, said that even with hindsight and professional scepticism he could not shake the certainty of what he had heard: “I will go to my grave swearing that it was your voice.” That sentence ought to be read by anyone tempted to believe that vigilance is the answer. Schildhorn is a trained advocate, paid to interrogate evidence and disbelieve plausible stories, and the clone defeated him as completely as it defeated a panicked grandmother. The vulnerability the fraud exploits is not located only in the elderly, or the credulous, or the technologically illiterate. It is located in the human auditory system itself, which evolved over millennia to treat a recognised voice as proof of a recognised person — and which is now, for the first time in that long history, systematically and exploitably wrong.
The data on older adults reinforces rather than contradicts this reframing. The FTC's December 2025 report to Congress found that total fraud losses reported by people aged sixty and over had roughly quadrupled between 2020 and 2024, reaching about 2.4 billion dollars, with 68 per cent of that sum attributable to individual losses of 100,000 dollars or more. The agency's own estimate of the true annual cost, accounting for the chronic underreporting that shame and embarrassment guarantee, ranged as high as 81.5 billion dollars. These are not the numbers of a credulous minority being separated from pocket money. They are the numbers of a generation's accumulated savings being drained through a mechanism specifically calibrated to their patterns of trust, their financial position and their place at the emotional centre of a family.
The Asymmetry, Quantified
Brian Long, the chief executive of the security firm Adaptive Security, distilled the new economics for AFP in a single sentence: “One guy in a room with a keyboard can make an infinite number of attackers.” That is the asymmetry in its purest form. On one side stands an automated system that can generate a convincing clone in seconds, dial thousands of numbers, and conduct each conversation with synthesised emotion, for a marginal cost approaching zero. On the other stands an individual human being, often elderly, alone, and given roughly the length of a panicked phone call to mount a defence that the world's leading forensic scientist could not.
INTERPOL's finding that AI-enhanced fraud is four and a half times more profitable than the traditional kind is the financial expression of this imbalance. When an attack becomes both cheaper to mount and more lucrative to complete, the volume of attacks does not rise linearly; it explodes. The 26 per cent single-year jump in American cybercrime losses, and the near-doubling of losses among the over-sixties, are what that explosion looks like in a national ledger. And the AFP wire noted something else that compounds the harm: shame. The Buffalo mother Liz Benz, who endured what she called “a good twenty minutes of terror” when a cloned voice told her that her sixteen-year-old son had been taken hostage, said that after she went public she was flooded with messages from other victims — many of whom chose to stay anonymous, because the humiliation of having been fooled kept them silent.
Underreporting is not a statistical footnote here. It is a structural feature of a crime designed to make its victims feel too foolish to come forward, which in turn starves the data, the prosecutions and the policy response of the evidence they need. A crime that silences its own witnesses is a crime that compounds at interest.
Why the Burden Cannot Sit With Families
The most widely circulated piece of advice, repeated by the FBI, the American Bankers Association and consumer advocates throughout 2026, is to agree a family “safe word” — a secret phrase known only to relatives, to be demanded in any emergency call. If the voice cannot produce it, hang up and call back on a known number. The advice is sound. It is also, as a systemic defence, hopelessly inadequate, and it is worth being clear about why.
A safe word works only if every member of a family adopts it, remembers it, and has the presence of mind to demand it in a moment engineered to obliterate presence of mind. It places the entire burden of defeating an industrial, automated, billion-dollar criminal apparatus on the cognitive discipline of a frightened individual at the worst moment of their week. It assumes that the eighty-year-old whose median reported loss, according to FTC data released in late 2025, exceeds 1,600 dollars will, while hearing her grandchild scream, calmly recall a protocol and execute it. Some will. Many, by design, will not. A defence that works only when the target performs flawlessly under maximum stress is not a defence; it is a way of allocating blame to the victim after the fact.
This is the deeper objection to placing protection in the hands of families and individuals. It transfers responsibility for a failure of the technological and financial system onto the people least equipped to bear it, and then, when they fail, treats their failure as a personal one. The voice-cloning tools were built and sold by companies. The calls are carried by telecommunications networks. The money moves through banks. Each of those parties operates at the chokepoints where the fraud could actually be interdicted at scale. The grandmother in her kitchen does not. Meaningful protection requires moving the burden from the end of the chain, where it currently sits, to the points in the middle where it belongs. A society that responds to an industrialised threat by issuing better advice to its most vulnerable members has confused the publication of guidance with the provision of protection.
Where Interdiction Could Actually Happen
Consider the three institutional chokepoints in turn, because each illustrates both the promise and the present failure of structural defence.
The first is the telephone network. In the United States, the STIR/SHAKEN framework was meant to address caller-ID spoofing by allowing originating carriers to cryptographically sign a call as legitimate and terminating carriers to verify that signature before it reaches a handset. In December 2025, the FCC's Wireline Competition Bureau concluded in its triennial efficacy report that the framework does authenticate caller ID effectively when properly applied. The qualification is doing enormous work. Criminals discovered early that routing calls through older, non-IP networks could evade the system entirely, and the FCC spent much of 2025 and 2026 trying to close that gap and pushing towards Rich Call Data, which would display a verified caller name and logo on the handset.
But STIR/SHAKEN authenticates the number, not the human, and certainly not the voice. It can tell you that a call genuinely originated from a given line. It cannot tell you that the sobbing daughter on that line is a machine. Against a cloned voice arriving from a spoofed or simply unfamiliar number, the framework is close to irrelevant. The same FCC declared in February 2024 that AI-generated voices in robocalls were illegal under the Telephone Consumer Protection Act — a meaningful statement of intent that nonetheless governs only mass automated dialling, not the targeted, one-to-one emergency call that defines the grandparent scam.
The second chokepoint is the platform that hosts the cloning tool. Here the most promising structural intervention is provenance rather than detection — the attempt, embodied in the C2PA standard, to attach a tamper-evident cryptographic record to a piece of media describing the device or model that produced it. In 2025 the standard was ratified as an ISO specification and, alongside watermarking schemes such as Google's SynthID and Meta's AudioSeal, became the de facto provenance language of the internet. The logic is sound: rather than asking whether a recording looks or sounds fake, ask whether it carries a credential proving it was made by a real microphone.
But provenance has a fatal asymmetry of its own for this specific crime. Missing credentials are not proof of fakery, because vast quantities of legitimate audio were never signed and because the metadata is routinely stripped by social platforms, screenshots and re-encoding. More fundamentally, a fraudster placing a phone call is under no obligation to transmit a C2PA manifest, and the analogue gap — playing synthetic audio down a telephone line — destroys any digital watermark in the act of transmission. Provenance can help establish, afterwards, that a viral video was synthetic.
It does almost nothing to stop a live cloned-voice call in progress.
The third chokepoint — and the most promising — is the bank. This is where the money actually moves, and therefore where interdiction has the greatest mechanical leverage. The United Kingdom offers the clearest natural experiment. In October 2024, the Payment Systems Regulator made reimbursement for authorised push payment fraud mandatory: where a victim is tricked into authorising a transfer to a fraudster, the sending and receiving banks must now reimburse them, splitting the liability fifty-fifty, up to 85,000 pounds, within five business days, with the consumer-negligence exception explicitly barred for vulnerable customers. Confirmation of Payee, the account-name-checking service, now runs on billions of transactions. The PSR's own dashboard showed that in the fifteen months to the end of December 2025, 89 per cent of money lost to such scams — some 243 million pounds — was reimbursed, against a 65 per cent rate before the rules took effect.
The point of mandatory reimbursement is not merely to make victims whole, though that matters. Its deeper purpose is to relocate the financial incentive. Once banks are liable for the losses, they acquire a powerful reason to build the friction, the anomaly detection and the intervention protocols that actually stop a fraudulent transfer before it completes — the held payment, the cooling-off period on a large cash withdrawal by an older customer, the human call from the branch asking why a retiree is suddenly emptying her savings to a courier.
The transfer friction that a grandmother cannot impose on herself, a bank can impose on her account, and a liability regime gives it the reason to do so. Critics, including commentators in Electronic Payments International, warn that reimbursement alone risks becoming a subsidy to fraudsters if it is not paired with prevention, and that the strategy must move beyond simply paying victims back. That critique is correct, and it points towards the right answer rather than away from it: liability is the lever that forces prevention, not a substitute for it.
The lesson of the British experiment is not that reimbursement solves fraud — it does not — but that it changes whose problem fraud is, and an institution made to own a problem will, eventually, engineer against it.
What Meaningful Protection Actually Requires
Pull these threads together and a coherent picture emerges of what would actually work, as distinct from what merely sounds reassuring.
It requires, first, abandoning detection as the primary line of defence. Hany Farid's blindness is not a temporary setback to be solved by a better classifier; it is a permanent structural condition of a world in which generation has outrun discrimination. Any plan whose final safeguard is someone, somewhere, telling the real from the fake has already failed.
It requires, second, regulating the supply of the weapon. The Consumer Reports finding that voice-cloning tools sit behind a self-attestation checkbox is a policy choice, not a law of nature. Mandatory, verifiable consent before a voice can be cloned — of the kind Descript and Resemble AI have partially implemented and the others have not — is technically feasible and would not abolish the legitimate uses of the technology. The European Union's AI Act, whose obligations for general-purpose models began applying through 2025 and 2026, and state statutes such as Tennessee's ELVIS Act, which requires written consent to clone a voice, are early gestures towards treating voice synthesis as the regulated capability it has become. They remain far ahead of enforcement and far behind the threat.
It requires, third, placing the burden of interdiction on the institutions that occupy the chokepoints — and, where they will not act voluntarily, compelling them through liability. The British reimbursement regime is imperfect and incomplete, but it demonstrates the mechanism: when banks own the loss, banks build the friction. Telecoms carriers that profit from carrying calls should bear a corresponding duty to authenticate and, where possible, flag them. Platforms that sell cloning should bear a duty to verify consent. The common principle is that responsibility ought to sit with the party that has both the capability to prevent the harm and the commercial benefit from the activity that causes it — which, in every case, is an institution, and in no case is an eighty-year-old answering her phone.
It requires, fourth, treating the human vulnerability as a thing to be managed rather than a thing to be blamed. The Charm Security framework's insight — that the cognitive and emotional mechanisms exploited by fraud deserve the same systematic cataloguing as software flaws — should inform how banks design their interventions, how carriers design their warnings, and how public bodies design education that goes beyond leaflets to the kind of role-based rehearsal that the ROLESafe research found actually changes behaviour. Awareness campaigns and family safe words are not worthless. They are simply the last and weakest line, useful only as a backstop to structural defences that do the real work.
The deepest reason the defensive gap is measured in years rather than months is that the attack is a technology problem and the defence is an institutional one. Cloning a voice takes three seconds and improves monthly. Passing a reimbursement regulation, rewiring a banking system's fraud controls, mandating consent verification across an industry, and closing the non-IP loophole in a national telephone network take years, because they require law, coordination, money and the overcoming of every commercial interest that profits from the status quo. The asymmetry is not merely technical; it is the asymmetry between how fast a machine can be built and how slowly a society can respond.
Sharon Brightwell got some of her money back, eventually, through the diligence of investigators and her bank. Many do not. The 352 million dollars the FBI attributes to older victims of AI fraud, and the far larger sum it cannot see, represent a transfer of wealth from the people who can least afford to lose it to the most efficient criminal enterprise yet devised. Closing the gap will not come from teaching grandmothers to doubt the sound of their grandchildren's voices. It will come from deciding, as a matter of law and engineering, that the institutions standing between the clone and the cash are responsible for what passes through their hands. Until that decision is made, the three-second theft will remain the easiest serious crime in the world to commit, and the hardest for its victims to be believed about — because the evidence, by design, sounds exactly like someone they love.
References
- Federal Bureau of Investigation, “Cryptocurrency and AI Scams Bilk Americans of Billions”, April 2026. https://www.fbi.gov/news/press-releases/cryptocurrency-and-ai-scams-bilk-americans-of-billions
- Internet Crime Complaint Center (FBI), “2025 IC3 Annual Report”, April 2026. https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf
- ABA Banking Journal, “FBI: Cybercrime losses increased 26% in 2025”, April 2026. https://bankingjournal.aba.com/2026/04/fbi-cybercrime-losses-increased-26-in-2025/
- SecureWorld, “FBI: AI-Enabled Fraud Topped $893M in 2025—Real Toll Likely Far Higher”, April 2026. https://www.secureworld.io/industry-news/ai-enabled-fraud-topped-893m-fbi
- AARP, “Older Adults Hit Hard by Fraud in 2025”, April 2026. https://www.aarp.org/money/scams-fraud/fbi-ftc-report-2025-losses/
- HousingWire, “FBI: Seniors lost $7.75B to cybercrime in 2025 — a 59% jump”, April 2026. https://www.housingwire.com/articles/fbi-seniors-cybercrime-2025/
- INTERPOL, “INTERPOL report warns of increasingly sophisticated global financial fraud threat”, March 2026. https://www.interpol.int/en/News-and-Events/News/2026/INTERPOL-report-warns-of-increasingly-sophisticated-global-financial-fraud-threat
- Help Net Security, “Global fraud losses climb to $442 billion”, 18 March 2026. https://www.helpnetsecurity.com/2026/03/18/online-fraud-victims-losses-interpol-report/
- ICLG, “Interpol warns of 'industrialisation' of global fraud”, March 2026. https://iclg.com/news/23665-interpol-warns-of-industrialisation-of-global-fraud
- The New York Times, “In the Age of A.I., the World's Leading Deepfake Expert No Longer Trusts His Own Eyes”, June 2026. (Reported via beSpacific.) https://www.bespacific.com/the-worlds-leading-deepfake-expert-no-longer-trusts-his-own-eyes/
- FOX 13 Tampa Bay, “Dover woman loses $15K after scammers used artificial intelligence to impersonate daughter”, July 2025. https://www.fox13news.com/news/dover-woman-loses-15k-after-scammers-used-artificial-intelligence-impersonate-daughter
- Agence France-Presse (via The Manila Times), “'20 minutes of terror': AI boosts US voice impersonation scams”, 4 June 2026. https://www.manilatimes.net/2026/06/04/news/world/20-minutes-of-terror-ai-boosts-us-voice-impersonation-scams/2358259
- Consumer Reports, “Consumer Reports' Assessment of AI Voice Cloning Products”, March 2025. https://www.consumerreports.org/media-room/press-releases/2025/03/consumer-reports-assessment-of-ai-voice-cloning-products/
- NBC News, “AI can steal your voice, and there's not much you can do about it”, March 2025. https://www.nbcnews.com/tech/security/ai-voice-cloning-software-flimsy-guardrails-report-finds-rcna195131
- The Register, “Consumer Reports calls out poor AI voice-cloning safeguards”, 10 March 2025. https://www.theregister.com/2025/03/10/ai_voice_cloning_safeguards/
- ElevenLabs, “Safety”, 2026. https://elevenlabs.io/safety
- Ben, A., Rahav, T., Illaev, D., Nahon, A., and Grushka, A. (Charm Security), “The Human Vulnerabilities & Exploits (HVE) Framework”, arXiv:2606.10083, June 2026. https://arxiv.org/abs/2606.10083
- Deng, Y., Chen, X., Liao, J., Li, B., and Zou, Y., “Experiencer, Helper, or Observer: Online Fraud Intervention for Older Adults Through Role-based Simulation”, arXiv:2601.12324, January 2026. https://arxiv.org/abs/2601.12324
- Federal Communications Commission, “FCC Makes AI-Generated Voices in Robocalls Illegal”, February 2024. https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal
- Federal Communications Commission, “Combating Spoofed Robocalls with Caller ID Authentication (STIR/SHAKEN)”, updated 2026. https://www.fcc.gov/call-authentication
- EyeSift, “C2PA Deepfake Detection 2026: AI Image Watermarks & SynthID”, 2026. https://www.eyesift.com/ai-image-detection-2026-c2pa-content-credentials-synthid-watermarks-diffusion-fingerprints-deepfake/
- Payment Systems Regulator, “PS25/5 Consolidated policy statement: APP scams reimbursement requirement”, May 2025. https://www.psr.org.uk/media/rhelv4op/ps25-5-app-scams-reimbursement-consolidated-policy-statement-may-2025.pdf
- Electronic Payments International, “Why the UK's scam strategy must move beyond reimbursement”, 2025. https://www.electronicpaymentsinternational.com/comment/why-uk-scam-strategy-must-move-beyond-reimbursement/
- Federal Trade Commission, “Protecting Older Consumers 2024–2025: A Report of the Federal Trade Commission”, December 2025. https://www.ftc.gov/system/files/ftc_gov/pdf/P144400-OlderAdultsReportDec2025.pdf
- American Bar Association, “The Rise of the AI-Cloned Voice Scam”, September 2025. https://www.americanbar.org/groups/senior_lawyers/resources/voice-of-experience/2025-september/ai-cloned-voice-scam/