# Artificial Analysis 评测蚂蚁集团 Ling-3.0-flash-Fin 金融模型

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-09-16 21:49
- AIHOT 分数：54
- AIHOT 链接：https://aihot.news/items/cmu45sg1e0aggro4w1miibl4n
- 原文链接：https://x.com/ArtificialAnlys/status/2100220634433466398

## AI 摘要

Artificial Analysis 评测蚂蚁集团发布的金融开源权重模型 Ling-3.0-flash-Fin，Intelligence Index 得 23 分，Finance & Accounting Index 得 24 分。

## 正文

Ling-3.0-flash-Fin, Ant Group’s new finance-focused open weights model, scores 23 on the Artificial Analysis Intelligence Index and 24 on the Finance & Accounting Index, and is on the Intelligence vs. Active Parameter Pareto Frontier

@AntLingAGI has released Ling-3.0-flash-Fin, a finance-focused model built on Ling-3.0-flash. Ant Group announced that it developed the model with financial institutions and industry experts to support financial research, including checking sources, building valuation spreadsheets and writing reports. This text-only model comes after their release of their image and video input-capable model Ling-3.0-flash-VL, which scored 25 on the Intelligence Index.

Key results:

➤ Ling-3.0-flash-Fin matches MiniMax-M2.7’s Intelligence Index score with roughly half the active parameters. Both score 23, while Flash-Fin activates 5.1B parameters per token compared with MiniMax-M2.7’s 10B.

➤ Ling-3.0-flash-Fin matches Ling-3.0-flash-VL at 24 on the Artificial Analysis Finance & Accounting Index. Fin has higher business knowledge accuracy than VL (17% vs. 11%), but also higher business knowledge hallucination (33% vs. 19%)

➤ Ling-3.0-flash-Fin scores slightly below Ling-3.0-flash-VL on professional knowledge work. It scores 1171 Elo on GDPval-AA v2 and 967 on AA-Briefcase, compared with 1225 and 986 respectively for Ling-3.0-flash-VL. Both benchmarks test agents on professional tasks such as producing documents and spreadsheets.

➤ Difficult agentic tasks remain a challenge for Ling-3.0-flash-Fin. It scores 7% on AutomationBench-AA, which tests workflows across business apps while respecting guardrails, compared with 16% for the flash-VL model. Both models score 0% on Terminal-Bench v4.0, which tests difficult terminal-use tasks.

➤ Ling-3.0-flash-Fin uses more output tokens than the flash-VL model and MiniMax-M2.7. It averages ~67k output tokens per Intelligence Index task, about 34% more than VL (~50k) and 3.2x MiniMax-M2.7 (~21k).

Additional model details:

➤ Type: Open weights reasoning model.

➤ Size: 124B total parameters, 5.1B active per token (MoE).

➤ Context window: 256K tokens.

➤ Modalities: Text input and output.

➤ API availability: Available through @OpenRouter, including a rate-limited free endpoint.

➤ License: MIT.
