多彩编程 多彩编程MZPH · CODE BLOG
ARTICLE DETAIL

文章详情

深耕前端与后端开发技术的一线实战笔记与踩坑复盘。

INMS 论文笔记:给 LLM Agents 搭一套 Memory Sharing 骨架,TaoToken 统一 Key 接入

INMS 论文笔记:给 LLM Agents 搭一套 Memory Sharing 骨架,TaoToken 统一 Key 接入 1. 从 INMS 论文说起多任务 Agent 为什么需要 Memory Sharing如果你最近在折腾 LLM Agents大概率会遇到一个很别扭的问题单个 Agent 跑单任务时表现还行一旦让它同时处理多个任务或者让多个 Agent 协作记忆就开始打架。任务 A 的上下文污染了任务 B 的推理或者两个 Agent 明明在做同一件事却各自维护一份重复的对话历史token 烧得飞快。INMSMemory Sharing for Large Language Model based Agents这篇论文想解决的就是这件事。它的核心主张不复杂Agent 的记忆不应该是一个大而全的全局变量而应该是一套可共享、可隔离、可复用的骨架。共享的部分让多个 Agent 复用同一份事实性知识隔离的部分保证每个任务有自己的私有上下文互不串味。我读完论文后的直观感受是它把 Agent 记忆拆成了两层一层是跨任务共享的公共记忆池存放那些与具体任务无关的稳定信息比如工具用法、环境配置、领域常识另一层是任务私有的工作记忆只服务于当前这条 Agent Loop。共享层负责省 token、省重复推理私有层负责保证推理的独立性和可追溯性。这套思路落地到工程上最直接的问题就是多个 Agent、多个任务、多个模型调用怎么统一管理 Key 和 API 通道如果每个 Agent 各自配一套凭证配置会迅速失控。这也是我为什么在搭这套骨架时选择用 TaoToken 做统一接入层——一个 Key 覆盖多个模型通道Agent 侧只需要关心记忆读写逻辑不用管底层凭证分发。下面我会把 INMS 的 Memory Sharing 机制拆成可跑的配置骨架给出 config.toml 和 settings.json 两份片段再带你跑一次记忆共享的读写验证最后把常见的坑列出来。2. TaoToken 前置统一 Key 与 API 通道怎么准备在动手写配置之前先把接入层理清楚。INMS 骨架里会有多个 Agent 实例每个实例在每一轮 Loop 里都可能发起模型调用。如果每个 Agent 都单独配 Key你会遇到三个麻烦Key 轮换时要改多处、不同模型走不同通道时配置分散、调试时很难定位是哪条通道出的问题。TaoToken 在这里扮演的角色是统一入口。你只需要在它那边生成一个 API Key然后所有 Agent 的模型调用都指向同一个 API 地址。官网入口在 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content API 基地址是 https://taotoken.net/api 注意这个地址后面不加任何 UTM 参数。具体操作路径是这样先到控制台 https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite 创建项目然后在 API Keys 页面 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 生成一个 Key。这个 Key 就是后面 config.toml 里要填的凭证。有一点要提醒Key 生成后只显示一次复制下来存到本地环境变量或者配置文件里别直接提交到 git。我习惯用环境变量TAOTOKEN_API_KEY来存配置文件里引用变量名而不是明文。如果你还没决定用哪个模型可以先到模型对话页面 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 试一下确认通道通了再写进 Agent 配置。对于长期跑编码类 Agent 的场景Coding Plan https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 会更合适因为它的额度模型更贴合高频 Loop 调用。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 里面写了不同语言 SDK 的调用方式。如果你用的是 Claude Code 这类工具Anthropic 兼容通道的说明在 https://taotoken.net/claudecode-anthropic?utm_sourcetaotoken_aicg_blog_endutm_contentclaudecode-anthropicutm_campaignrewrite 。3. 可复制配置config.toml 与 settings.json 骨架现在进入正题。INMS 的 Memory Sharing 骨架我把它拆成两个配置文件config.toml 管 Agent 运行时的记忆策略和模型通道settings.json 管共享记忆池的存储结构和隔离规则。先看 config.toml。这份配置定义了三个 Agent 实例它们共享一个公共记忆池但各自有独立的私有记忆空间。# config.toml - INMS Memory Sharing 骨架配置 [api] # TaoToken 统一接入 base_url https://taotoken.net/api api_key_env TAOTOKEN_API_KEY timeout_seconds 60 max_retries 3 [memory] # 共享记忆池配置 shared_pool_enabled true shared_pool_path ./memory/shared_pool.jsonl shared_pool_max_entries 5000 # 私有记忆配置 private_memory_enabled true private_memory_path ./memory/private private_memory_ttl_hours 72 # 隔离策略task / agent / hybrid isolation_mode hybrid [memory.sharing_rules] # 哪些类型的记忆允许跨 Agent 共享 shareable_types [tool_usage, env_config, domain_fact] # 哪些类型强制隔离 isolated_types [task_context, user_intent, intermediate_reasoning] # 共享记忆的写入权限 write_requires_confirm false [agents.planner] model claude-sonnet role task_planner memory_scope shared_read_private_write [agents.executor] model claude-sonnet role task_executor memory_scope shared_read_private_write [agents.reviewer] model claude-sonnet role result_reviewer memory_scope shared_read_only [loop] max_iterations 20 stop_on_final true history_window 10这份配置里几个关键点值得展开。isolation_mode hybrid是 INMS 的核心思路落地共享层和私有层同时存在按记忆类型决定走哪条路。shareable_types里放的是工具用法、环境配置、领域事实这类跨任务稳定的信息isolated_types里放的是任务上下文、用户意图、中间推理过程这类一旦串味就会出问题的内容。memory_scope控制每个 Agent 对记忆的读写权限。planner 和 executor 可以读共享池、写私有记忆reviewer 只读共享池不产生新的私有记忆。这样设计的原因是 reviewer 的职责是校验不应该引入新的状态。再看 settings.json它管的是共享记忆池的存储结构和读写接口。{ memory_schema: { version: 1.0, entry_fields: { id: string, type: string, content: object, source_agent: string, task_id: string, created_at: timestamp, ttl: integer, shareable: boolean } }, shared_pool: { storage: jsonl, index_fields: [type, source_agent, task_id], dedup_key: [type, content_hash], max_entry_size_kb: 64 }, private_memory: { storage: jsonl, per_agent_dir: true, auto_compact_threshold: 200 }, read_write_api: { read_shared: GET /memory/shared?type{type}limit{n}, write_shared: POST /memory/shared, read_private: GET /memory/private/{agent_id}, write_private: POST /memory/private/{agent_id} }, conflict_resolution: { strategy: latest_wins, merge_on_type_match: true } }dedup_key用type加content_hash做去重避免多个 Agent 把同一条工具用法重复写进共享池。conflict_resolution用 latest_wins因为共享池里存的多是事实性信息新版本覆盖旧版本是合理的。把这两份配置放到项目根目录后目录结构大概是这样project/ ├── config.toml ├── settings.json ├── memory/ │ ├── shared_pool.jsonl │ └── private/ │ ├── planner.jsonl │ └── executor.jsonl └── agent_runner.py4. 验证请求跑一次记忆共享读写闭环配置写好了接下来要验证共享记忆真的能跨 Agent 读写。我写了一个最小验证脚本模拟 planner 写入一条工具用法到共享池然后 executor 从共享池读出来。# verify_memory_sharing.py import json import os import hashlib from datetime import datetime from pathlib import Path SHARED_POOL Path(./memory/shared_pool.jsonl) PRIVATE_DIR Path(./memory/private) def content_hash(content: dict) - str: raw json.dumps(content, sort_keysTrue) return hashlib.sha256(raw.encode()).hexdigest()[:16] def write_shared(entry: dict): entry[id] content_hash(entry[content]) entry[created_at] datetime.utcnow().isoformat() entry[shareable] True with open(SHARED_POOL, a, encodingutf-8) as f: f.write(json.dumps(entry, ensure_asciiFalse) \n) print(f[write_shared] id{entry[id]} type{entry[type]}) def read_shared(mem_type: str, limit: int 10): results [] if not SHARED_POOL.exists(): return results with open(SHARED_POOL, r, encodingutf-8) as f: for line in f: if not line.strip(): continue entry json.loads(line) if entry.get(type) mem_type: results.append(entry) return results[-limit:] def write_private(agent_id: str, entry: dict): PRIVATE_DIR.mkdir(parentsTrue, exist_okTrue) path PRIVATE_DIR / f{agent_id}.jsonl entry[created_at] datetime.utcnow().isoformat() with open(path, a, encodingutf-8) as f: f.write(json.dumps(entry, ensure_asciiFalse) \n) print(f[write_private] agent{agent_id} type{entry[type]}) if __name__ __main__: # 1. planner 写入一条共享的工具用法 write_shared({ type: tool_usage, content: { tool: shell, command: npm install, note: Node 项目缺依赖时使用 }, source_agent: planner, task_id: task-001, ttl: 86400 }) # 2. planner 写入一条私有任务上下文 write_private(planner, { type: task_context, content: {goal: 给项目补 README, step: 1}, task_id: task-001 }) # 3. executor 从共享池读取工具用法 entries read_shared(tool_usage) print(f[read_shared] found {len(entries)} entries) for e in entries: print(f - {e[content][tool]}: {e[content][command]}) # 4. 验证隔离executor 读不到 planner 的私有记忆 planner_private PRIVATE_DIR / planner.jsonl executor_private PRIVATE_DIR / executor.jsonl print(f[isolation] planner private exists: {planner_private.exists()}) print(f[isolation] executor private exists: {executor_private.exists()})跑这个脚本之前先确认环境变量已经设好export TAOTOKEN_API_KEY你的Key python verify_memory_sharing.py预期输出是这样[write_shared] ida3f8c2d1e4b5 typetool_usage [write_private] agentplanner typetask_context [read_shared] found 1 entries - shell: npm install [isolation] planner private exists: True [isolation] executor private exists: False看到found 1 entries说明共享池读写通了executor private exists: False说明隔离生效了——executor 没有自己的私有记忆文件它只能读共享池不能碰 planner 的私有上下文。如果你想把模型调用也串进来可以在 Agent Loop 里加一段每轮构造 Prompt 时先从共享池按shareable_types拉取相关记忆拼进 system prompt再把私有记忆作为 history 的一部分。这样模型每一轮看到的就是「共享事实 私有上下文」的组合。5. 本篇常见错排查配置跑不通的时候问题通常集中在几个地方。我把踩过的坑列一下。第一个坑是共享池文件路径不对。config.toml 里写的是./memory/shared_pool.jsonl这是相对路径取决于你从哪个目录启动脚本。如果你在项目根目录跑没问题但从子目录跑就会找不到文件。解决办法是用绝对路径或者在脚本开头os.chdir(Path(__file__).parent)。第二个坑是 Key 没设进环境变量。config.toml 里写的是api_key_env TAOTOKEN_API_KEY意思是运行时从环境变量读。如果你直接跑脚本没 export模型调用会报 401。验证方法是echo $TAOTOKEN_API_KEY空的话就补上。第三个坑是共享池去重没生效。如果你发现同一条工具用法被写了多次检查dedup_key里的content_hash是不是对 content 做了排序序列化。字典顺序不一致会导致 hash 不同去重就失效了。上面脚本里用了sort_keysTrue就是为了避免这个问题。第四个坑是隔离模式配错。如果你把isolation_mode设成task那所有记忆都按 task_id 隔离共享池就形同虚设。INMS 的 hybrid 模式才是共享和隔离并存的正确姿势。第五个坑是 TTL 过期没清理。私有记忆设了ttl_hours 72但如果你没写清理逻辑过期条目会一直堆着。建议在 Agent 启动时加一个 compact 步骤把超过 TTL 的条目删掉。第六个坑是模型通道超时。Agent Loop 里每轮都要调模型如果timeout_seconds设太短复杂任务会频繁超时。我一般设 60 秒配合max_retries 3基本够用。如果还是超时检查一下是不是共享池拉取的数据量太大把history_window调小试试。6. 接入与排障统一 Key 的后续动作把上面的骨架跑通之后下一步就是把它接到真实的 Agent 工作流里。如果你在排障阶段遇到 API 调用问题先去 API Keys 页面 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 确认 Key 状态再对照接入文档 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 检查请求格式。如果你还在选模型阶段想先验证共享记忆在不同模型下的表现可以到模型对话页面 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 手动试几轮观察同一份共享记忆喂给不同模型时的输出差异。对于长期跑编码类 Agent、需要高频 Loop 调用的场景Coding Plan https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 的额度模型更贴合不用每次调用都担心额度碎片化。如果你用的是 Claude Code 这类工具链Anthropic 兼容通道的配置说明在 https://taotoken.net/claudecode-anthropic?utm_sourcetaotoken_aicg_blog_endutm_contentclaudecode-anthropicutm_campaignrewrite 。最后说一个我自己的习惯共享池里的条目我会定期导出做一次人工 review把那些已经过时的工具用法和领域事实清掉。INMS 的共享机制省的是重复推理的 token但如果共享池本身被污染了省下来的 token 会以更隐蔽的方式还回去。保持共享池干净比一味往里塞东西更重要。
返回列表