 注入指南)
人工智能AI AgentAgent 记忆RAG后端MCP 服务【免费下载链接】honchoMemory library for building stateful agents项目地址https://gitcode.com/gh_mirrors/hon/honcho点击查看免费下载本文围绕 Honcho开源 Memory library for building stateful agents的 Agent 记忆访问层展开讲解如何让 AI 在对话中按需读取用户上下文。你将掌握三种核心接入模式——将 Dialectic 对话封装为工具调用、在每次 LLM 调用前预取固定属性、以及用context()一键注入历史与用户表征——并了解各自适用场景与底层实现能够在 Python 与 TypeScript 项目中直接落地。模式总览与选型Honcho 的记忆由后台推理引擎异步构建你通过add_messages()喂入对话消息系统抽取前提、生成结论最终沉淀为对每个参与者Peer的表征Representation。Agent 读取这份记忆有两条主要通道peer.chat(query)是实时推理问答数秒延迟session.context()则是近乎实时的格式化读取。本指南对应的三种接入模式建立在客户端/Peer/Session 基础搭建之上基础样板请先阅读 核心接入模式。模式触发方式延迟典型场景Pattern ADialectic Chat 作为工具调用Agent 自行决定何时调用数秒实时推理自主 Agent、需要按需深挖用户背景/偏好/目标Pattern B预取固定属性每次 LLM 调用前查询数秒多次chat()简单集成、prompt 需要固定字段Pattern Ccontext()注入每轮对话读取近即时每轮对话的持续 grounding、对话历史注入速度提示chat()执行实时的 Dialectic 推理耗时数秒Pattern A 和 B 都会调用它context()Pattern C是近即时读取。每轮对话 grounding 优先用context()只有确实需要“经过推理的答案”时才调用chat()。在开始编码前可用 Honcho CLI 快速验证连通性honcho init持久化 API Key 与 URL、honcho doctor检查连接、配置、workspace 健康、honcho peer chat交互式测试 Dialectic 端点具体见 集成技能指南。Pattern A将 Dialectic Chat 封装为 Agent 工具调用这是推荐给 Agent 化系统的模式把 Honcho 的 chat 端点暴露为 AI Agent 的一个工具Tool让模型在需要理解用户背景、偏好、历史交互或目标时自行决定查询。核心收益是“按需召回”——模型只在真正需要时才付出数秒的推理延迟而不是每轮都付。Python 实现OpenAI Function Callingimport os import json import openai from honcho import Honcho honcho Honcho(workspace_idmy-app, api_keyos.environ[HONCHO_API_KEY]) # 定义 Agent 的工具 honcho_tool { type: function, function: { name: query_user_context, description: Query Honcho to retrieve relevant context about the user based on their history and preferences. Use this when you need to understand the users background, preferences, past interactions, or goals., parameters: { type: object, properties: { query: { type: string, description: A natural language question about the user, e.g. What are this users main goals? or What communication style does this user prefer? } }, required: [query] } } } def handle_honcho_tool_call(user_id: str, query: str) - str: 执行 Honcho chat 工具调用。 peer honcho.peer(user_id) return peer.chat(query) # 在 Agent 循环中使用 def run_agent(user_id: str, user_message: str): messages [{role: user, content: user_message}] response openai.chat.completions.create( modelgpt-4, messagesmessages, tools[honcho_tool] ) # 处理工具调用 if response.choices[0].message.tool_calls: for tool_call in response.choices[0].message.tool_calls: if tool_call.function.name query_user_context: args json.loads(tool_call.function.arguments) result handle_honcho_tool_call(user_id, args[query]) # 将工具结果继续回传对话...TypeScript 实现OpenAI Function Callingimport OpenAI from openai; import { Honcho } from honcho-ai/sdk; const honcho new Honcho({ workspaceId: my-app, apiKey: process.env.HONCHO_API_KEY }); const honchoTool: OpenAI.ChatCompletionTool { type: function, function: { name: query_user_context, description: Query Honcho to retrieve relevant context about the user based on their history and preferences., parameters: { type: object, properties: { query: { type: string, description: A natural language question about the user } }, required: [query] } } }; async function handleHonchoToolCall(userId: string, query: string): Promisestring { const peer await honcho.peer(userId); return await peer.chat(query); }深入chat()的底层能力从源码看Python SDK 的Peer.chat()通过POST /peers/{peer_id}/chat调用 Dialectic 端点请求体默认带{query: query, stream: False}并支持以下可选参数可用于增强 Pattern A 的工具设计target指定目标 Peer查询“本 Peer 对目标 Peer 的局部表征”local representation而非全局表征session/sessions把查询限定到单个 Session 或一组 Session 的白名单内跨 Session 推理出的结论不会出现在sessions白名单召回中因其出处无法证明落在白名单内scope用命名作用域Scope限定查询来源与session/sessions互斥且要求 workspace 级 API Keyreasoning_level推理深度取值为minimal/low/medium/high/max不传默认lowresponse_format传入 Pydantic 模型类可拿到解析后的实例传入 JSON Schema dictroot type 为 object则返回 JSON 字符串——适合让工具结果结构化输出include_evidence为True时返回ChatResponse同时携带答案与 Dialectic 读取时看到的证据证据由 Agent 自身读取整理比单纯引用列表更宽泛timeout单次 HTTP 尝试的超时秒数须大于 0缺省时使用客户端全局超时重试会延长总耗时。对应 TypeScript 一侧的peer.chat与peer.chatStream位于 TypeScript SDK 的 peer 模块参数语义一致。这意味着 Pattern A 的工具函数可以做得更精细例如按 Session 限定召回范围或要求结构化 JSON 输出便于下游程序消费。Pattern B预取上下文Targeted Pre-fetch Queries对于不需要 Agent 自主决策的简单集成可在每次 LLM 调用之前用一组预定义的固定问题批量拉取用户关键属性拼进 System Prompt。这类“固定字段”模式实现直观、可预测但代价是每次调用前都要付出数次chat()的推理延迟。Python 实现def get_user_context_for_prompt(user_id: str) - dict: 通过定向 Honcho 查询拉取关键用户属性。 peer honcho.peer(user_id) return { communication_style: peer.chat(What communication style does this user prefer? Be concise.), expertise_level: peer.chat(What is this users technical expertise level? Be concise.), current_goals: peer.chat(What are this users current goals or priorities? Be concise.), preferences: peer.chat(What key preferences should I know about this user? Be concise.) } def build_system_prompt(user_context: dict) - str: return fYou are a helpful assistant. Heres what you know about this user: Communication style: {user_context[communication_style]} Expertise level: {user_context[expertise_level]} Current goals: {user_context[current_goals]} Key preferences: {user_context[preferences]} Tailor your responses accordingly.TypeScript 实现async function getUserContextForPrompt(userId: string): PromiseRecordstring, string { const peer await honcho.peer(userId); const [style, expertise, goals, preferences] await Promise.all([ peer.chat(What communication style does this user prefer? Be concise.), peer.chat(What is this users technical expertise level? Be concise.), peer.chat(What are this users current goals or priorities? Be concise.), peer.chat(What key preferences should I know about this user? Be concise.) ]); return { communicationStyle: style, expertiseLevel: expertise, currentGoals: goals, preferences: preferences }; }工程要点查询措辞建议追加 “Be concise.”避免推理答案冗长挤占 prompt 预算TypeScript 一侧可用Promise.all并发执行多个查询以减少总等待时间四个chat()并行整体延迟约等于单次推理若部分字段并非每轮都需要可考虑对查询做缓存如按用户 时间窗口缓存或直接改选 Pattern C 的context()以换取近即时读取在 集成技能指南 的 Phase 2 访谈中选择 Pre-fetch 的用户会被进一步询问关心的上下文维度沟通风格、专业水平、目标/优先级、偏好、近期活动摘要、自定义查询本文的字段组可直接作为默认选项。Pattern C用context()为 LLM 注入会话上下文context()是每轮对话 grounding 的首选它一次性返回会话内近期消息、两级会话摘要short/long以及目标 Peer 的表征全部经过 token 预算裁剪并可直接转换为 OpenAI / Anthropic 的消息数组。Python 实现import openai session honcho.session(conversation-123) user honcho.peer(user-123) assistant honcho.peer(assistant) # 获取已按 LLM 要求格式化的上下文 context session.context( tokens2000, peer_targetuser.id, # 包含该用户的表征 summaryTrue # 包含会话摘要 ) # 转换为 OpenAI 格式 messages context.to_openai(assistantassistant) # 或 Anthropic 格式 # messages context.to_anthropic(assistantassistant) # 追加新的用户消息 messages.append({role: user, content: What should I focus on today?}) response openai.chat.completions.create( modelgpt-4, messagesmessages ) # 存储本轮对话供 Honcho 继续推理 session.add_messages([ user.message(What should I focus on today?), assistant.message(response.choices[0].message.content) ])TypeScript 实现import OpenAI from openai; const session await honcho.session(conversation-123); const user await honcho.peer(user-123); const assistant await honcho.peer(assistant); // 获取已按 LLM 要求格式化的上下文 const context await session.context({ tokens: 2000, peerTarget: user.id, // 包含该用户的表征 summary: true // 包含会话摘要 }); // 转换为 OpenAI 格式 const messages context.toOpenAI(assistant); // 或 Anthropic 格式 // const messages context.toAnthropic(assistant); // 追加新的用户消息 messages.push({ role: user, content: What should I focus on today? }); const openai new OpenAI(); const response await openai.chat.completions.create({ model: gpt-4, messages }); // 存储本轮对话 await session.addMessages([ user.message(What should I focus on today?), assistant.message(response.choices[0].message.content!) ]);context()返回什么从 Python SDK 的SessionContext模型 可以看到session.context()打包了一个可直接投喂给 LLM 调用的“会话本地视图”近期消息Recent messages来自当前 Session 的消息列表按tokens预算裁剪token 计数基于 tiktoken结果与 OpenAI 模型兼容见 Session.context() 源码说明会话摘要Conversation summaries当summaryTrue时附带采用short / long 两级摘要SessionSummaries中的short_summary与long_summary字段让更早的轮次仍然“算数”却不必占满 token 预算目标 Peer 的表征与卡片传入peer_target时Honcho 会把对该用户综合理解的表征peer_representation与 Peer Cardpeer_card折叠进上下文不传则只得到会话本地上下文没有跨会话记忆。值得一提的进阶参数Session.context()还支持peer_perspective以某 Peer 的视角读取目标表征必须与peer_target搭配、scope以某 Scope 的观察视角读取、sessions表征白名单此时 Peer Card 会被省略因为派生结论与卡片无法证明逐会话出处、limit_to_session仅保留本 Session 内的结论以及search_query/search_top_k/search_max_distance/include_most_frequent/max_conclusions等语义检索与结论数量控制参数详见 Session.context() 签名。to_openai()/to_anthropic()的格式细节转换辅助方法把上述内容组织成目标提供商的messages数组你传入的assistantPeer 决定角色归属OpenAI 格式Pythonto_openai来自assistantPeer 的消息标记为role: assistant其余标记为role: user并附带name: peer_id表征、Peer Card、摘要分别以peer_representation、peer_card、summaryXML 标签包裹作为system消息置于最前Anthropic 格式Pythonto_anthropicassistant消息原样输出其余消息以{peer_id}: {content}形式标记为user表征/卡片/摘要同样以标签包裹此处为user角色。源码注释提示未来版本可能实现 Anthropic 要求的角色交替role alternation。TypeScript 一侧的SessionContext.toOpenAI/toAnthropic在 TypeScript SDK 的 session_context 模块 中实现输出结构与 Python 完全对应两个 SDK 的行为保持一致。流式响应Streaming当推理答案较长、希望以流式方式返回给终端用户时两种通道都支持流式读取stream peer.chat_stream(What do we know about this user?) for chunk in stream: print(chunk, end, flushTrue)const stream await peer.chatStream(What do we know about this user?); for await (const chunk of stream) { process.stdout.write(chunk); }从源码看PythonPeer.chat_stream()与chat()共享同一套参数体系target、session、scope、sessions、reasoning_level、response_format、include_evidence请求体仅将stream置为True底层通过 SSE 流解析SSEStreamParser逐块产出文本。若开启了include_evidence证据在流结束时随最后一个事件返回因此DialecticStreamResponse的evidence属性只有在流被完整消费后才可用。response_format在流式模式下保持为原始文本累积成 JSON 字符串需要自行解析如Model.model_validate_json。模式组合与落地建议三种模式并非互斥实战中常组合使用Agent 化系统会自主决策是否查询选Pattern A必要时在工具函数内用session/scope/reasoning_level精细控制召回范围与推理成本每轮对话都要 groundingPattern C的context()近即时、token 可控适合注入历史 表征在完整 示例 中可以看到session.context(summaryTrue, tokens50)的极简用法该示例将 token 上限设得极低以仅保留少量小消息简单集成、字段固定选Pattern B预取固定属性若字段随时间变化不敏感可加缓存降低推理调用频次。无论采用哪种模式都请遵守基础规范全应用使用单一 workspace为所有实体用户与 AI 助手创建 Peer对确定性脚本机器人设置observe_meFalseAI 助手可保留观察为可选优化每次对话后调用add_messages()喂给推理引擎——没有消息就没有推理没有推理就没有记忆消息处理是异步的不要轮询等待推理完成详见 核心接入模式 与 集成技能指南 的常见错误清单。赞分享人工智能AI AgentAgent 记忆RAG后端MCP 服务【免费下载链接】honchoMemory library for building stateful agents项目地址https://gitcode.com/gh_mirrors/hon/honcho点击查看免费下载相关推荐AI 助手联网搜索总答非所问Exa、Tavily、Brave 三款 MCP 搜索引擎选型指南AI 助手联网搜索总答非所问Exa、Tavily、Brave 三款 MCP 搜索引擎选型指南 问 AI 助手一句最近发生了什么它要么用训练数据里的旧闻搪教程AI AgentHoncho TypeScript SDK 实战指南用 honcho-ai/sdk 为有状态 Agent 接入会话记忆Honcho TypeScript SDK 实战指南用 honcho ai/sdk 为有状态 Agent 接入会话记忆 本指南以 honcho ai/sd人工智能AI AgentAgent 记忆RAG后端MCP 服务Honcho实战为CrewAI和LangGraph Agent系统注入持久化记忆能力Honcho实战为CrewAI和LangGraph Agent系统注入持久化记忆能力 Honcho 是一款为 Agent 提供 持久化记忆 的基础设施mem人工智能AI AgentAgent 记忆RAG后端MCP 服务创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考