AI 安全需要从上下文到行动的溯源链
AI Security Needs a Chain of Provenance From Context to Action
生产级 AI 系统由检索、模型推理、工具调用、API 访问与跨系统操作等多层串联而成,安全风险往往出现在层与层之间的接缝处,因此关键不在于单个组件是否安全,而在于组织能否在转换过程中保留溯源信息。
Most AI security discussions focus on the model: can it be jailbroken, can it leak sensitive data, can an attacker manipulate its prompt? Those questions matter, but production AI systems are increasingly made of several connected layers. A user request can trigger retrieval from documents, model inference, a tool call, access to an API, and an action in another system before the workflow is complete. The important security property is therefore not only whether each component is safe. It is whether the organisation can preserve provenance across the transformations between them: where information came from, which identity introduced it, what the model was allowed to do with it, what policy governed the next step, and what action ultimately occurred.
AI risk often appears at the seams
A model can behave correctly while the application around it creates risk. A retrieval system may insert untrusted instructions into context. A tool may accept arguments the model should never be able to generate. An application may authorise an action based on the user’s broad account permissions even though the specific task did not require them.
These are seam failures. The individual components can all appear legitimate, yet the transition from one component to another changes the meaning of the data. A retrieved document is data to the application, but language to the model. Model output is a suggestion to the application, but it can become a command when passed directly to an execution layer. An OAuth token is authentication evidence to a service, but it can also become delegated authority for an agent that did not exist when the token was first granted.
This is why AI systems need a chain of provenance rather than isolated controls. Provenance gives each transition a security context. Without it, legitimate components can combine into an unsafe chain.
Provenance begins before the prompt
The first question is not simply “What did the user type?” It is “Who or what initiated this task, under which identity, and with which authority?” That origin should remain visible throughout the workflow.
Enterprise AI increasingly receives instructions from more than humans. A scheduled process can call an agent. Another agent can delegate a subtask. An application can invoke a model automatically after a business event. Each origin creates a different trust context.
Security metadata should travel with the request: initiating identity, device or workload context, data classification, approved purpose, and the maximum authority available to the workflow. If that context disappears after the first model call, downstream controls must reconstruct intent from incomplete evidence. Preserving that metadata prevents downstream controls from losing the original trust context.
Retrieved context needs its own trust label
RAG systems make AI more useful by grounding responses in enterprise data. They also create a difficult security condition because retrieved content can contain instructions, outdated permissions, manipulated text, or information that the user is not entitled to combine with other sources. A secure retrieval layer should preserve source identity, access rights, freshness, and trust level alongside the content. The model should not receive a flattened block of text that erases where each piece came from.
That matters for more than prompt injection. Provenance allows downstream controls to answer whether a conclusion was based on authoritative data, whether sensitive information crossed a policy boundary, and whether the source was still valid at the time of the decision. NIST’s 2026 Cyber AI Profile workshop report emphasises the need to address AI-specific attack surfaces while integrating AI risks into broader cybersecurity governance. A provenance model is useful because it connects those concerns to evidence that conventional security teams already know how to manage: identity, source, policy, access, and change history.
Model output should not inherit trust automatically
One of the most dangerous shortcuts in AI application design is treating structured output as trusted output. A model can return valid JSON, a correct function name, or arguments that satisfy a schema and still propose an unsafe action. Validation proves structure. It does not prove authorisation.
That distinction is central to AI security. The application should treat model output as untrusted input to a separate decision layer. Before a tool call executes, trusted code should verify whether the initiating identity can perform the action, whether the target resource is in scope, whether the data involved is permitted, and whether the action exceeds a risk threshold requiring confirmation.
The model can recommend. It should not define its own authority. That separation keeps authority outside the model even when its output looks well formed.
Recent coverage of emerging AI transparency requirements under the EU AI Act. Transparency and security increasingly intersect here. If organisations are expected to explain how AI-mediated outcomes are produced, provenance becomes both a governance mechanism and a security control. A system that cannot show which source, model, policy, and tool contributed to an action is difficult to audit and difficult to defend.
Tool calls create a new evidence boundary
Agentic systems make provenance even more important because the result of one model call may create the state that influences the next one. A single task identifier should follow that state change. Otherwise, the evidence fragments as the workflow grows.
Imagine an agent that reads a support ticket, retrieves customer records, updates a CRM entry, drafts an email, and sends it. The security record should not be five unrelated logs. It should preserve one execution chain showing the initiating request, data sources, authorisation decisions, tool invocations, user confirmations, and resulting changes.
This lets investigators distinguish legitimate automation from manipulated automation. It also enables policy enforcement mid-workflow. If the agent’s task changes from reading to modifying, or from internal processing to external communication, the security requirements can change with it.
Provenance makes revocation practical
AI systems are dynamic. Documents are corrected, permissions change, models are updated, plugins are replaced, and agents receive new tools. A decision that was valid yesterday may rely on a source or permission that is no longer trusted today.
With strong provenance, organisations can identify which outputs or actions depended on a compromised source, vulnerable component, or revoked permission. Without it, the blast radius becomes difficult to calculate. That traceability is what makes targeted remediation possible.
This is analogous to software supply-chain security. Knowing that an application contains a vulnerable dependency is useful because teams can trace which systems are affected. AI systems need comparable traceability for context and actions.
The security architecture should preserve causality
The goal is not to record hidden model reasoning. Security does not require access to private chain-of-thought. It requires observable causality around the system: what entered, what sources were used, which identity was acting, what tools were available, which policy decision occurred, and what changed as a result.
That creates a practical chain: request identity → retrieved context → model interaction → proposed action → authorisation decision → tool execution → resulting state. Each transition should preserve enough evidence for the next layer to make an informed decision. The chain is useful only if that security context survives every handoff.
As AI moves from generating content to changing business systems, provenance becomes a security primitive. The organisations best positioned to scale AI safely will not be those that assume every component can be made perfectly trustworthy. They will be those that can trace how trust is transformed from context to action and stop the chain when one link no longer deserves it.
来源:Artificial Intelligence News(网页) · artificialintelligence-news.com