这是关于中国开源社区自 2025 年 1 月“DeepSeek 时刻”以来历史性进展的三部分系列博客中的第三篇,也是最后一篇。第一篇关于战略变化与开放产物增长,可在此处查看
此处
,第二篇关于架构与硬件转变,可在此处查看
此处
在这第三篇文章中,我们考察中国知名 AI 组织的路径与轨迹,并提出开源的未来方向。
对于为开源生态做出贡献并依赖开源生态的 AI 研究人员和开发者,以及理解这一快速变化环境的政策制定者而言,由于组织内部与全球社区的收益,开源在近期内是中国 AI 组织的主导且流行的路径。从模型到论文再到部署基础设施,公开分享各类产物,对应着一种以大规模部署和集成为目标的战略。
中国自发形成的开源 AI 生态
在考察了自 DeepSeek 的 R1 以来的战略与架构变化之后,我们首次得以一窥一个自发形成的开源 AI 生态正在中国成形。一批强大参与者的汇聚——有些已在开源领域立足,有些是新玩家,还有些则彻底改变方向以投身这一新的开放文化——表明这种开放协作的方式是互利的。
这一合作正在跨越国界;Hugging Face 上关注度最高的组织是 DeepSeek,关注度第四高的是 Qwen。
除了模型之外,公开分享科学与技术不仅为其他 AI 组织提供了信息,也惠及了整个开源社区。Hugging Face 上最受欢迎的论文大多来自中国组织,即字节跳动、DeepSeek、腾讯和 Qwen。
来源:https://huggingface.co/spaces/evijit/PaperVerse
老牌巨头
阿里巴巴将开源定位为一种生态与基础设施战略。Qwen并未被塑造成单一旗舰模型,而是持续扩展为一个覆盖多种规模、任务和模态的家族,并在 Hugging Face 及其自有平台 ModelScope 上频繁更新。其影响力并未集中在任何一个单一版本上。相反,它作为组件在不同场景中被反复复用,逐渐承担起通用 AI 基础的角色。到 2025 年年中,Qwen 成为 Hugging Face 上衍生模型最多的模型,超过 113k 个模型以 Qwen 为基础,超过 200k 个模型仓库标注了 Qwen,远超 Meta 的 Llama 的 27k 或 DeepSeek 的 6k。从组织整体来看,阿里巴巴拥有的衍生模型数量几乎与 Google 和 Meta 两家之和相当。
与此同时,阿里巴巴将模型开发与云及硬件基础设施对齐,把模型、芯片、平台和应用整合为统一的工程栈。
腾讯也迈出了从借用走向自研的重要一步。作为 R1 发布后首批将 DeepSeek 集成到核心面向消费者产品中的大公司之一,腾讯最初并未将开源作为公开叙事。相反,它通过插件式集成引入成熟模型,进行了大规模内部验证,之后才开始发布自身能力。从 2025 年 5 月起,腾讯在自身已有优势的领域加速开源发布,例如视觉、视频和 3D,并以自有品牌腾讯混元(现为 Tencent HY)命名,这些模型迅速在社区中获得采用。
字节跳动则遵循其“AI 应用工厂”的思路,开始选择性地开源高价值组件,同时将竞争重心保持在产品入口和大规模使用上。在此背景下,字节跳动 Seed 团队贡献了若干值得关注的开源成果,包括用于多模态 UI 理解的 UI-TARS-1.5、以数据为中心的代码建模 **Seed-Coder **,以及用于系统性推理评估的 SuperGPQA 数据集。尽管开源存在感相对低调,字节跳动在中国 AI 市场已实现显著规模,其 AI 应用 豆包在 2025 年 12 月 DAU 突破 1 亿。
其中最引人注目的变化来自百度,其 CEO 公开对开源持保留态度,但也开始了转型:在多年优先发展闭源模型之后,它通过免费访问和开放发布重新进入生态,例如Ernie 4.5 系列。这一转变伴随着对其开源框架PaddlePaddle以及其自研 AI 芯片昆仑芯的重新投入,后者于 2026 年 1 月 1 日宣布 IPO。通过在更开放的系统中将模型、芯片和 PaddlePaddle 连接起来,百度可以降低成本、吸引开发者并影响标准,同时在算力、成本和监管的共同约束下保持战略控制。
“DeepSeek 时刻”的常态化
在初创公司中,月之暗面、Z.ai和MiniMax迅速做出调整,并在 R1 发布后的数月内为开源社区带来了新的动力。Kimi K2、GLM-4.5和MiniMax M2等模型都跻身 AI-World 开源模型里程碑排行榜。2025 年底,Z.ai 和 MiniMax 发布了各自迄今最先进的开源模型,随后又紧接相继宣布了 IPO 计划。
Kimi K2 的开源被广泛描述为社区迎来的"又一个 DeepSeek 时刻"。尽管月之暗面尚未宣布 IPO,但市场报告显示,该公司在 2025 年底前筹集了约 5 亿美元融资,并将 AGI 和基于智能体的系统定位为其主要商业化目标。
那些应用优先的公司,如小红书、哔哩哔哩、小米和美团,此前只专注于应用层,如今也开始训练和发布自己的模型。凭借在真实使用场景和领域数据方面的原生优势,一旦强大的推理能力通过开源以低成本变得可用,构建自研模型就变得切实可行。它围绕自身特定业务来调优 AI,而不是受制于外部供应商的成本结构或限制。
如果说商业世界抓住了 ROI 为正的增长机遇,那么研究机构和更广泛的社区则更加欣然地欢迎这一转变。BAAI和上海人工智能实验室等机构将更多精力转向工具链、评估体系、数据平台和部署基础设施,如 FlagOpen、OpenDataLab 和 OpenCompass 等项目。这些努力并非追逐单一模型的性能,而是夯实生态系统的长期基础。
未来的基础
新生态的决定性特征不在于模型数量更多,而在于一整条链条已经形成。模型可以被开源并扩展;部署可以被复用并规模化;软件与硬件可以协同并互换;治理能力可以被嵌入并审计。这是一次从孤立突破向一个真正能在现实世界中运行的系统的转变。
这一生态并非一夜之间出现。它建立在自 2017 年以来多年积累的基础设施“顺风”之上。过去几年间,中国定期投资于数据中心和算力中心,逐步形成了以“东数西算”战略为核心的全国一体化算力布局。这一国家规划确立了 8 大算力枢纽和 10 大数据中心集群,引导算力需求从东部向中西部地区转移。
公开信息表明,中国打算持续投资于能源产能的增长。截至 2025 年,中国总算力规模约为 1590 EFLOPS,位居全球前列。中国方面的消息来源称,专为 AI 训练和部署量身打造的智能算力预计将同比增长约 43%,远超通用算力的增速。与此同时,数据中心平均电能利用效率(PUE)降至约 1.46,表明效能有所提升,为规模化 AI 提供了坚实的硬件基础。能源显然是一个关键焦点。
如果说 2017 年的《新一代人工智能发展规划》主要是在定方向和打基础,那么 2025 年 8 月的“人工智能+”行动方案则明显将重心转向了规模化部署与深度融合。这标志着一条与 AGI 方向不同的追求路径。R1 的出现,在工程与生态层面补上了缺失的那股“升力”。它是催化剂,系统性地激活了此前已经建好的算力、能源和数据基础设施。
因此,在 R1 发布后的一年里,中国 AI 的发展沿着两条主要路径加速推进。第一,AI 更深入地嵌入产业流程,从聊天机器人迈向智能体和工作流。第二,更加重视自主可控的 AI 系统,这体现在更灵活的训练路径和日益本地化的部署策略上。
回过头看,真正的转折点不是模型数量的增长,而是开源模型使用方式的根本性变化。开源从一种可选选择,变成了系统设计中的默认假设。模型成为更大工程系统中可复用、可组合的组件。
回望过去,展望未来
从 DeepSeek 到“人工智能+”,中国在 2025 年的路径并非追逐性能巅峰,而是构建一条围绕开源、工程效率和可规模化交付展开的务实路径,这条路径已经开始自行运转。
资源约束并未限制中国 AI 的发展。在某些方面,它们重塑了中国 AI 的发展轨迹。DeepSeek R1 的发布成为一个催化事件,在国内产业中引发了一系列连锁反应,并加速了一个结构更加有机的生态系统的形成。与此同时,这一转变也为国内持续的研发创造了一个关键窗口期。随着这一生态系统日趋成熟,其更长期的影响——以及全球 AI 社区将如何与一个日益自给自足的中国 AI 生态系统互动——将成为未来讨论中的重要问题。
This is the third and final blog in a three-part series on China's open source community's historical advancements since January 2025's "DeepSeek Moment." The first blog on strategic changes and open artifact growth is available
here
, and the second blog on architectural and hardware shifts is available
here
In this third article, we examine paths and trajectories of prominent Chinese AI organizations, and posit future directions for open source.
For AI researchers and developers contributing to and relying on the open source ecosystem and for policymakers understanding the rapidly changing environment, due to intraorganizational and global community gains, open source is the dominant and popular approach for Chinese AI organizations for the near future. Openly sharing artifacts from models to papers to deployment infrastructure maps to a strategy with the goal of large-scale deployment and integration.
China's Organic Open Source AI Ecosystem
Having examined strategic and architectural changes since DeepSeek's R1, we get a glimpse for the first time at how an organic open source AI ecosystem is taking shape in China. A culmination of powerful players, some established in open source, some new players, and some changing course entirely to contribute to the new open culture, signal that the open collaborative approach is mutually beneficial.
This collaboration is reaching beyond national boundaries; the most followed organization on Hugging Face is DeepSeek, and the fourth most followed is Qwen.
In addition to models, openly sharing science and techniques has not only informed other AI organizations, but also the entire open source community. The most popular papers on Hugging Face largely come from Chinese organizations, namely ByteDance, DeepSeek, Tencent, and Qwen.
source: https://huggingface.co/spaces/evijit/PaperVerse
The Established
Alibaba positioned open source as an ecosystem and infrastructure strategy. Qwen was not shaped as a single flagship model, but continuously expanded into a family covering multiple sizes, tasks, and modalities, with frequent updates on Hugging Face and their own platform ModelScope. Its influence did not concentrate on any single version. Instead, it was repeatedly reused as a component across different scenarios, gradually taking on the role of a general AI foundation. By mid-2025, Qwen became the model with most derivatives on Hugging Face, with over 113k models using Qwen as a base, and over 200k model repositories tagging Qwen, far exceeding Meta's Llama's 27k or DeepSeek's 6k. Organization-wide, Alibaba boasts the most derivatives almost as much as both Google and Meta combined.
At the same time, Alibaba aligned model development with cloud and hardware infrastructure, integrating models, chips, platforms, and applications into a single engineering stack.
Tencent also made a significant move from borrowing to building. As one of the first major companies to integrate DeepSeek into core consumer-facing products after R1's release, Tencent did not initially frame open source as a public narrative. Instead, it brought mature models in through plug-in style integration, ran large-scale internal validation, and only later began to release its own capabilities. From May 2025 onward, Tencent accelerated open releases in areas where it already had strengths, such as vision, video, and 3D with its own brand named Tencent Hunyuan (it's now Tencent HY), and these models quickly gained adoption in the community.
ByteDance, by following its "AI application factory" approach, starts selectively open sourcing high value components while keeping its competitive focus on product entry points and large scale usage. In this context, the ByteDance Seed team has contributed several notable open-source artifacts, including UI-TARS-1.5 for multimodal UI understanding, **Seed-Coder **for data-centric code modeling, and the SuperGPQA dataset for systematic reasoning evaluation. Despite a relatively low-profile open-source presence, ByteDance has achieved significant scale in China's AI market, with its AI application Doubao surpassing 100 million DAU in December 2025.
The most notable change within Baidu, whose CEO openly calls short on open source, also started its shift: after years of prioritizing closed models, it re-entered the ecosystem through free access and open release, such as the Ernie 4.5 series. This shift was accompanied by renewed investment in its open-source framework, PaddlePaddle, as well as its own AI chip Kunlunxin, which announced an IPO on January 1, 2026. By connecting models, chips, and PaddlePaddle within a more open system, Baidu can lower costs, attract developers, and influence standards, while maintaining strategic control under shared constraints of compute, cost, and regulation.
The Normalcy of "DeepSeek Moments"
Among startups, Moonshot, Z.ai, and MiniMax adjusted rapidly and brought new momentum to the open-source community within months after R1. Models such as Kimi K2, GLM-4.5, and MiniMax M2 all earned places on AI-World's open-model milestone rankings.At the end of 2025, Z.ai and MiniMax released their most advanced open-source models to date and subsequently announced their IPO plans in close succession.
The open-sourcing of Kimi K2 was widely described as a "Another DeepSeek moment" for the community. Although Moonshot has not announced an IPO, market reports indicate that the company raised approximately $500M in funding by the end of 2025, with AGI and agent-based systems positioned as its primary commercialization objectives.
Those application-first companies such as Xiaohongshu, Bilibili, Xiaomi, and Meituan, previously focused only on the application layer, began training and releasing their own models. With their native advantage in real usage scenarios and domain data, once strong reasoning became available at low cost through open source, building in-house models became practical. It tunes AI around their specific businesses, rather than being constrained by the cost structures or limits of external providers.
If the business world seized the ROI positive opportunity for growth, research institutions and the broader community welcomed the shift even more willingly. Organizations such as BAAI and Shanghai AI Lab redirected more effort toward toolchains, evaluation systems, data platforms, and deployment infrastructure, projects like FlagOpen, OpenDataLab, and OpenCompass. These efforts did not chase single-model performance, but instead strengthened the long-term foundations of the ecosystem.
Foundations for the Future
The defining feature of the new ecosystem is not that there are more models, but that an entire chain has formed. Models can be open-sourced and extended; deployments can be reused and scaled; software and hardware can be coordinated and swapped; and governance capabilities can be embedded and audited. This is a shift from isolated breakthroughs to a system that can actually run in the real world.
This ecosystem did not appear overnight. It is built on years of accumulated infrastructure "tailwind" since 2017. Over the past several years, China has periodically invested in data centers and compute centers, gradually forming a nationwide, integrated compute layout centered on the "East Data, West Compute" strategy. The national plan established 8 major compute hubs and 10 data center clusters, guiding compute demand from the east toward the central and western regions.
Public information suggests China intends to invest in continual growth in energy capacity. China's total compute capacity is around 1590 EFLOPS as of the year 2025, ranking among the top globally. Sources in China assert that intelligent compute capacity, tailored for AI training and deployment, is expected to grow by roughly 43% year over year, far outpacing general-purpose compute. At the same time,the average data center power usage effectiveness (PUE) fell to around 1.46, indicating better effectiveness and providing a solid hardware foundation for AI at scale. Energy is a clear key focus.
If the 2017 "New Generation AI Development Plan" was mainly about setting direction and building foundations, then the August 2025 "AI+" action plan clearly shifted focus toward large-scale deployment and deep integration. This marks a directionally different pursuit from AGI. The emergence of R1 provided the missing "lift" at the engineering and ecosystem level. It was the catalyst that systematically activated compute, energy, and data infrastructure that had already been built.
As a result, in the year following R1's release, China's AI development accelerated along two main paths. First, AI became more deeply embedded in industrial processes, moving beyond chatbots toward agents and workflows. Second, greater emphasis was placed on autonomous and controllable AI systems, reflected in more flexible training pathways and increasingly localized deployment strategies.
Looking back, the real turning point was not the growth in the number of models, but a fundamental change in how open-source models are used. Open source moved from an optional choice to a default assumption in system design. Models became reusable and composable components within larger engineering systems.
Looking Back to Look Forward
From DeepSeek to "AI+", China's path in 2025 was not about chasing peak performance. It was about building a practical path organized around open source, engineering efficiency, and scalable delivery, a path that has already begun to run on its own.
Resource constraints did not limit China's AI development. In some respects, they reshaped its trajectory. The release of DeepSeek R1 acted as a catalytic event, triggering a chain of responses across the domestic industry and accelerating the formation of a more organically structured ecosystem. At the same time, this shift created a critical window for continued domestic research and development. As this ecosystem matures, its longer-term impact---and how the global AI community may engage with an increasingly self-sustaining AI ecosystem in China---will become important questions for future discussion.
