【文章标题】:Agentic Context Management: Memory and Cost as Architecture Problems 【文章标题】:智能体上下文管理:将记忆与成本视为架构问题

【文章正文】: Computer Science > Artificial Intelligence [Submitted on 23 Jul 2026] Title:Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems View PDF HTML (experimental) 【文章正文】: 计算机科学 > 人工智能 [提交于 2026年7月23日] 标题:智能体上下文管理:将智能体记忆与成本视为生命周期和架构问题来解决 查看 PDF HTML(实验性)

        Abstract:Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing is too narrow. Actively managing what an agent holds in mind is a lifecycle, not merely a store: it spans deciding what to remember, extracting and structuring it, choosing the right store per data type, consolidating and forgetting while preserving provenance, deciding what is relevant now, anticipating what is needed next, and compacting context to a budget without losing what matters. In serious production this operates not over a single user but across an organizational scope hierarchy. We name this discipline Agentic Context Management (ACM) and decompose it into five primitives: architecting, ingesting, scoping, anticipating, and compacting & consolidation. We then make the economic case: naive context accumulation grows token cost quadratically in conversation length, crude summarization buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity. We describe a reference implementation, Maximem Synap, that realizes the five primitives as a multi-tenant service and reports 92% on LongMemEval and 93.2% on LoCoMo under the configuration detailed in Section 6. We close with dimensions existing benchmarks do not yet capture, latency, token efficiency, and context-rot resistance, and the frontier of decision-level and organization-level context the category points toward.
        摘要:生产环境中的AI智能体发生故障,较少是因为推理能力不足,更多是因为它们无法管理其推理上下文中的内容:对话历史、大型提示词、庞大的工具定义以及急剧膨胀的工具输出。智能体淹没在自己不断积累的历史记录中,同时支付着每轮都在增长的Token成本,导致在对话内部和跨对话中出现回忆缺失。现有的应对方法将此视为一个存储与检索问题。我们认为这种框架过于狭隘。主动管理智能体脑海中所持的内容是一个生命周期,而不仅仅是一个存储库:它涵盖了决定记住什么、提取并结构化这些信息、为每种数据类型选择合适的存储、在保留来源的同时进行整合与遗忘、决定当前什么是相关的、预测接下来需要什么,以及在不丢失重要信息的前提下将上下文压缩到预算范围内。在严肃的生产环境中,这不仅针对单个用户运行,而是跨越组织范围层级运行。我们将这一学科命名为智能体上下文管理(ACM),并将其分解为五个基本操作:架构设计、数据摄取、范围界定、需求预测、压缩与整合。接着,我们从经济学角度进行论证:简单的上下文累积会使Token成本随对话长度呈二次方增长,粗糙的摘要方法以准确性断崖式下降为代价换取线性成本,只有经过验证的压缩方法才能在保持保真度的同时实现线性成本。我们描述了一个参考实现Maximem Synap,它将这五个基本操作实现为多租户服务,并在第6节详述的配置下,在LongMemEval上达到92%的准确率,在LoCoMo上达到93.2%的准确率。最后,我们提出了现有基准测试尚未涵盖的维度:延迟、Token效率和抗上下文腐化能力,以及该类别所指向的决策级和组织级上下文的前沿方向。

References & Citations

Loading...

参考文献与引用

加载中...

Bibliographic and Citation Tools Bibliographic Explorer (What is the Explorer?)

        Connected Papers (What is Connected Papers?)
      
    
        Litmaps (What is Litmaps?)
      
    
        scite Smart Citations (What are Smart Citations?)

文献与引用工具 文献浏览器(什么是浏览器?)

        Connected Papers(什么是Connected Papers?)
      
    
        Litmaps(什么是Litmaps?)
      
    
        scite 智能引用(什么是智能引用?)
      
    Code, Data and Media Associated with this Article
        alphaXiv (What is alphaXiv?)
      
    
        CatalyzeX Code Finder for Papers (What is CatalyzeX?)
      
    
        DagsHub (What is DagsHub?)
      
    
        Gotit.pub (What is GotitPub?)
      
    
        Hugging Face (What is Huggingface?)
      
    
        ScienceCast (What is ScienceCast?)
      
    与本文相关的代码、数据和媒体
        alphaXiv(什么是alphaXiv?)
      
    
        CatalyzeX 论文代码查找器(什么是CatalyzeX?)
      
    
        DagsHub(什么是DagsHub?)
      
    
        Gotit.pub(什么是GotitPub?)
      
    
        Hugging Face(什么是Huggingface?)
      
    
        ScienceCast(什么是ScienceCast?)
      
    Demos

Recommenders and Search Tools Influence Flower (What are Influence Flowers?)

          CORE Recommender (What is CORE?)
      
    演示

推荐与搜索工具 Influence Flower(什么是Influence Flowers?)

          CORE 推荐器(什么是CORE?)
        
      arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.

      arXivLabs:与社区合作者共同开展的实验性项目

arXivLabs 是一个框架,允许合作者直接在我们的网站上开发和分享新的 arXiv 功能。 与 arXivLabs 合作的个人和组织都拥抱并接受了我们的开放、社区、卓越和用户数据隐私价值观。arXiv 致力于践行这些价值观,并仅与遵守这些价值观的合作伙伴开展合作。 有能为 arXiv 社区创造价值的项目想法?了解更多关于 arXivLabs 的信息。