跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 28 天前AI 评分41

提出 agentic data cracking:让推理智能体自适应结构化非结构化数据以降低 token 成本

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

AI 导读

论文提出 agentic data cracking,智能体在推理时以边际成本派生 cracking 子智能体,自适应且投机地把网页、报告、合同、PDF 等非结构化数据抽取为结构化形式。在 FanOutQA 基准上仅扩展每题一个相关问题,该方法在保持准确率的同时将成本降低 53%,且理想预结构化存储下推理比原方案便宜 28X。作者称这是面向智能体推理的下一代数据基础设施的第一步。

正文

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this today, but at prohibitive cost. Each question repeatedly opens large documents to recover scattered evidence, consuming up to a million tokens. However, if the data were already structured, the same question would reduce to a cheap database lookup. For example, on FanOutQA benchmark, reasoning over an ideal pre-structured store is 28X cheaper, and the gap grows to orders of magnitude as questions fan out over more documents. Yet structuring everything in advance is not viable: documents hold vastly more possible structure than any workload will use, and the useful structure and documents are unknown until queries arrive. We propose agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself. Structuring is adaptive because observed queries decide when it happens and what matters, and speculative because it goes beyond the current question. Whenever the agent opens a document to answer, a cracking sub-agent forks from the already-loaded context at marginal cost and extracts grounded structure likely to serve related future queries. Over time, an increasing share of queries is fully covered by structured data and answered without opening a document, keeping agentic accuracy at close to RAG cost. On FanOutQA, extended with merely one related question per test question, cracking cuts cost by 53% while preserving accuracy. Agentic data cracking is a first step toward next-generation data infrastructure for agentic reasoning over unstructured data: a shared substrate beneath the model where knowledge that reasoning already paid to uncover accumulates.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org