跳到正文
O'Reilly Radar· Monika Dvorackova·· 2 小时前AI 评分42

AI 栈中的保存缺口:为什么能力进步快于可重建性

The Preservation Gap in the AI Stack: Why Capability Is Advancing Faster Than Reconstructability

AI 导读

O'Reilly Radar 文章提出 AI 系统存在"保存缺口":运行时依赖(模型端点、检索数据、策略、工具契约等)在执行时才确定,现有基础设施只保存模型、代码、部署和观测记录,不保存某次执行的实际绑定关系。作者为此提出参考架构 Orrery,用于保存重建 AI 系统历史执行所需的运行时绑定与依赖状态,并区分了"可复现性"与"可重建性"两个概念。

正文

Artificial intelligence systems are difficult to reproduce because their behavior depends on nondeterministic models and on data, configurations, policies, and external services that continue to change. But exact reproducibility is not always what operators, investigators, or auditors need. They often need something different: historical reconstructability. The central claim is that reconstructability is a system property established during execution, not inferred later from surviving artifacts. To make that property concrete, this article introduces Orrery, a reference architecture for preserving the runtime bindings and historical dependency states needed to reconstruct an AI system’s past execution.

Imagine an AI-assisted decision that must be examined six months after it was made. Rerunning the system may produce a different answer, but an investigator may still need to determine which model endpoint, policy, retrieved data, tool contract, and runtime configuration were used at the time. That is the problem addressed here.

The AI industry remains much better at evaluating capability than at preserving the knowledge required to understand past executions. Evaluation pipelines track benchmark scores, task success rates, retrieval precision and recall, tool-call success rates, latency, and cost, and they grow more elaborate with every release. In practice, this machinery often reduces to one acceptance question: What can this system do? That question captures capability, but not whether a past execution will remain understandable after the system changes.

That missing property is historical reconstructability: the ability to establish which dependencies were used and under which conditions a particular execution occurred. It is related to, but distinct from, reproducibility. Reproducibility asks whether a result can be produced again under equivalent conditions. Reconstructability asks whether the conditions of the original execution can still be identified, even when repeating that execution would not produce an identical result.

This distinction matters when a consequential AI execution completes successfully today and is challenged six months later. The model may still exist in a registry. The application revision may still exist in Git. The trace may still identify the request and the services it crossed. Yet none of these records necessarily reveals which model endpoint served the request, which input transformation and runtime configuration were applied, which data, feature state, or retrieved context entered the computation, or which policy state governed it. In an agentic system, the missing history may also include memory, tool contracts, delegation, and external effects. Even when a dependency identifier was recorded, the historical state to which it referred may no longer be available or resolvable.

The problem is that an AI runtime may use dependencies whose historical identities and relationships are not preserved. This is an architectural problem, not merely a logging problem. The gap lies between what the runtime uses and what the surrounding infrastructure is required to preserve. I call this the preservation gap. The AI stack preserves models, code, deployments, and observations, but it does not require the effective historical configuration of a particular execution to remain available. The components may remain while the relations that made them one execution disappear. An AI runtime must assemble today’s execution; it is not necessarily required to remember exactly what it assembled yesterday. Reconstructability must therefore be established during execution, before later mutation erases the bindings that gave the execution its effective configuration.

Hermetic builds illustrate the opposite design pattern. They close the dependency graph before execution begins by pinning inputs, versions, and external dependencies. AI runtimes face the reverse situation: Part of the dependency graph remains open until execution, as model endpoints may resolve through mutable aliases, data and feature state evolve, runtime configuration and policy change, and retrieved context is selected only when a request is processed. Agentic systems extend the graph further through memory, tools, delegation, and external effects. In a hermetic build, dependency closure is an input. In a modern AI runtime, part of it may be an output of the run.

Version control records source evolution, transaction logs record state transitions, and distributed tracing records execution relationships. These mechanisms preserve different views of a run but not its materially relevant execution-specific configuration. A trace may retain execution flow, a deployment manifest the declared application state, and a model registry the model revision. What is often missing is the record of which artifacts and states were bound together in a particular execution.

The dependency graph no longer closes at deployment

A deployment records the configuration known before execution begins. AI systems are harder to reconstruct because materially relevant dependencies are selected during execution rather than fixed at deployment. At inference time, the system may resolve a model endpoint, input transformation, data or feature state, retrieved context, runtime configuration, safety controls, and policy. Agentic systems extend this composition through memory, tools, delegation, and external effects. These execution-specific selections are runtime bindings: facts about what the execution actually used.

The following example makes this distinction concrete. Consider a hypothetical agentic execution identified as E891. During this execution, the system retrieves two documents, evaluates a policy, invokes a tool, and produces an external effect. The deployment may identify application revision 8f31c2 and model v17, while the execution itself binds prompt h31, retrieval index r42, entities d182@17 and d761@4, tool contract h42, policy v8, and ultimately effect e3. The deployment system could not have known this entire set in advance, because some of these dependencies were selected only as the execution proceeded.

The difference between what a deployment declared and what an execution actually used becomes consequential when dependencies change. A model alias may resolve differently, a feature or data source may change, preprocessing and runtime configuration may evolve, and policy v8 may become v9. In agentic systems, retrieved context, memory, and tool contracts introduce further independent mutation. Deployment identity describes what was declared, but the historical record of an execution must describe what was actually used. Unless those binding relations are captured when they occur, the effective configuration of E891 cannot be recovered reliably from later system state.

Figure 1. A deployment captures the initial configuration, while its execution-scoped dependency closure emerges as additional dependencies are bound during execution.
Figure 1. A deployment captures the initial configuration, while its execution-scoped dependency closure emerges as additional dependencies are bound during execution. (Diagram by the author.)

Current state is a lossy projection

Once runtime bindings are treated as part of the execution, the limitation of current-state records becomes clear. The execution-scoped dependency closure is the combination of the declared state and the runtime bindings that formed the effective configuration of a particular execution.

As dependencies mutate, different historical executions can leave behind the same observable evidence. The current state may show which artifacts still exist without revealing which combination a particular execution used. Once the distinguishing bindings are gone, the history cannot be reconstructed from what remains. A system can retain its artifacts while losing reconstructability: An identifier proves neither that a dependency was used nor that the historical state it denotes remains available. Reconstructability therefore requires durable binding relations and continued access to the historical states they identify.

Existence is not usage

The existence of an artifact and its use in a particular execution are different facts. Concurrent versions, caches, retries, and asynchronous changes make it unreliable to infer whether an execution used an artifact merely because that artifact existed at the relevant time.

W3C PROV already provides the relevant semantic distinction: An activity can use an entity. The missing primitive is therefore not a new provenance vocabulary, but a runtime requirement to record which dependency a particular execution used and to preserve an identity that can be resolved later.

This requirement also exposes the limit of observability. OpenTelemetry can transport dependency identifiers through attributes, links, context, and baggage, but encoding an identifier does not preserve its historical meaning. A trace may retain execution topology while the identities that explain it disappear. Instrumentation can carry preservation metadata, but it cannot guarantee durable storage or continued resolution of the referenced states.

Historical reconstructability therefore requires more than retained artifacts. For a defined class of executions and a specified retention period, the system must preserve two properties: binding integrity, which records which dependency versions and states an execution actually used, and resolution integrity, which keeps those recorded identities resolvable to the historical states they denote.

Preservation requires failure semantics

Once reconstructability is treated as a system property, the architecture must define what happens when the records required to reconstruct an execution fail to reach durable storage. Suppose policy v8 authorizes execution E891 to invoke tool contract h42, and the tool commits effect e3. If the usage record fails to reach durable storage, the action succeeds while the historical links among the effect, its policy, the tool contract, and the execution context are lost. The external effect remains, but the historical record no longer reliably explains which policy and tool contract authorized it. The external world and the historical record have diverged.

Figure 2. Execution E891 branches into a committed external effect and a failed usage record.
Figure 2. Execution E891 branches into a committed external effect and a failed usage record. (Diagram by the author.)

Figure 2 shows why this condition cannot be treated as ordinary telemetry loss. When an external effect has been committed but the corresponding usage record is missing, the architecture does not guarantee historical reconstructability. It provides only best-effort historical evidence.

Reconstructability becomes an engineering guarantee only when the architecture defines when preservation records become durable and what the system does if they cannot be stored.

A system may require a durable usage record before committing the effect, atomically persist the effect intent and preservation record through a transactional outbox, quarantine the execution for reconciliation, or trigger a compensating action. The implementation is application-specific, but preservation failure must have explicit semantics rather than disappear into telemetry.

Existing technologies provide the building blocks for this layer. Provenance and lineage models represent relations, tracing propagates execution context, registries identify versions, and content-addressed stores preserve artifacts. None, however, creates a runtime obligation to preserve the execution-scoped dependency closure required for reconstruction.

This layer is a preservation plane: the contract and machinery that keep an execution’s dependency closure identifiable as its surrounding systems change. It defines what to record, how records become durable, and how referenced states remain resolvable. Orrery applies this principle as a reference architecture in which materially relevant runtime bindings are captured when they occur and retained together with resolvable identities for the historical states they reference.

AIGov Core: reconstructability as a tested property

AIGov Core provides a limited implementation of this principle. Its evidence ledger stores hash-chained events, while its replay engine reconstructs a governance verdict from a stored export without querying live state. A companion check labels the export Ready, Partial, or NotReady according to unresolved lineage, missing policy artifacts, version drift, or replay failure. This demonstrates that reconstructability can be tested, but it does not solve the complete preservation problem. The system can replay only what was recorded; it cannot recover a runtime binding that was never captured.

The preservation gap is the mismatch between what an execution depends on and what the system persists. Closing it requires a runtime obligation: Whenever an execution binds materially relevant state, the system must record that binding, preserve a durable identity for the dependency, and retain access to the historical state it denotes for the defined retention period. Reconstructability is therefore not a by-product of capability or observability. It is an architectural property that must be designed, tested, and enforced during execution.

来源:O'Reilly Radar · oreilly.com