多彩编程 多彩编程MZPH · CODE BLOG
ARTICLE DETAIL

文章详情

深耕前端与后端开发技术的一线实战笔记与踩坑复盘。

了解你的事实:Elasticsearch AI Indices 让 agents 无需通读内容,就能保留答案 - part 2

了解你的事实:Elasticsearch AI Indices 让 agents 无需通读内容,就能保留答案 - part 2 作者来自 Elastic Kathleen DeRusso, Matt Nowzari, Apostolos Matsagkas, Peter Pišljar一篇技术实践指南介绍如何将事实预先计算到 Elasticsearch AI Index 中让 agents 通过一次 ES|QL 查询直接获得答案而无需阅读完整文档从而减少 token 消耗并降低延迟。Elasticsearch 与业界领先的 Gen AI 工具和提供商进行了原生集成。观看我们的网络研讨会了解如何超越 RAG 基础或者使用 Elastic vector database 构建可用于生产环境的应用。要为你的使用场景构建最佳搜索解决方案现在可以开始免费 cloud 试用或者立即在你的本地计算机上试用 Elastic。为了回答一个问题而将完整文档加载到 agent 的上下文中成本很高而且每次检索失败都会进一步增加成本。在这篇实践指南中我们采用的是预先计算事实的方法。Kibana workflow 将每个文档提炼成事实级别的 Knowledge IndicatorKI存储在 Elasticsearch AI Index 中并通过一次 Elasticsearch Query LanguageES|QL查询进行检索。针对同一个问题与阅读原始文档相比从 KI 中获取答案的 agent 使用更少的 token 和更低的延迟就能得到相同的、有依据的答案而且无需将任何完整文档加载到上下文中。这些事实只需预先计算一次然后就可以存储起来供未来遇到类似查询的 agents 使用。这是我们关于使用 AI indices 构建上下文系列文章的第 2 部分第 1 部分介绍了如何将 agents 路由到正确的 index。上下文管理依赖于良好的检索。与其让 agents 为每个问题重新发现相同的内容不断重复类似的步骤并消耗 tokenElastic 的 agentic AI 功能可以让我们预先计算这些细节并将它们以结构化、可搜索的形式存储起来同时让 agents 直接加载这些上下文。我们将这种预先计算的上下文单元称为 Knowledge Indicator。默认的 agentic retrieval augmented generationRAG模式恰恰相反。它在查询时检索完整文档并将这些文档直接放入模型的上下文中因此每个问题都要为这种检索付出 token 和延迟成本。将答案预先计算为 KI可以把这部分成本移出关键路径只需执行一次即可。它是如何工作的AI Index、Kibana Workflows 和query-kiskill通过 AI indices 构建上下文主要包含三个部分AI Index一个用于存储 KI 的特殊 Elasticsearch index、用于创建 KI 的 Kibana Workflows以及帮助 agents 使用 ES|QL 直接查询 KI 的query-kiskill这篇博客文章与第 1 部分类似因为我们使用的是相同的核心构建模块。但在这篇文章中我们展示的是一个非常不同的使用场景。我们不是预先计算 index 元数据而是从已建立索引的文档中提炼出特定的事实这些事实可以直接用于回答 agents 的问题而无需后续搜索。如果你希望在阅读这些示例的过程中端到端地自行创建相同的 KI我们还提供了一个 notebook。前提条件Elasticsearch Serverless 和 LLM API key本教程假设你已经具备一个 Elasticsearch Serverless project。如果你还没有可以注册试用。一个用于访问 Elasticsearch project 的 API key。一个兼容 OpenAI 的大型语言模型LLMAPI key用于通过 Deep Agents scripts 访问 AI indices。将 BrowseComp-Plus 示例语料库加载到 Elasticsearch首先我们需要一些数据源。数据源可以是已经存在于 Elasticsearch indices 中的数据也可以是通过 connectors 或 ES|QL data sources 访问的外部数据。在这篇博客中我们将创建一个名为browsecomp-plus的 index用于存放示例数据其 mappings 如下{ browsecomp-plus: { mappings: { _meta: { description: BrowseComp-Plus corpus: ~100k human-verified web documents (news articles, Wikipedia entries, institutional pages) used as a reasoning-intensive browsing/QA retrieval benchmark. BM25-only index. }, properties: { docid: { type: keyword, meta: { description: Stable corpus document id. } }, text: { type: text, meta: { description: Full document text: title, date, and body content. } }, title: { type: text, meta: { description: Document title (from the documents front matter). } }, url: { type: keyword, meta: { description: Source URL the document was crawled from. } } } } } }并通过 _bulk API 填充一小部分 BrowseComp-Plus 数据。你可以使用配套的 notebook将这些示例数据加载到你的 project 中。创建用于存储 KI 的 AI Index与第 1 部分类似第一步是创建一个 AI IndexPUT ai-index-idx-my-corpus该 index 已预先配置好与第 1 部分中列出的相同必需 mappings。这里我们直接使用 semantic_text 开箱即用地执行混合搜索。Agents 如何使用 ES|QL 检索 KIKI 是 AI Index 中的一个文档。KI 之所以有用在于检索也就是查询 AI Index 以找到正确的内容。这个查询被封装在一个小型、可移植且与 harness 无关的 skill 中可以在任何 agent harness 中运行。下面是一个query-kiskill 示例--- name: query-ki description: - Retrieve Knowledge Indicators (precomputed context) from the Elasticsearch AI Index before answering. Use it to find which index to search (routing profiles) or to look up precomputed facts without reading source documents. Trigger on any question that depends on specific facts, names, dates, or on choosing a data source. allowed-tools: esql_query --- # Retrieving Knowledge Indicators Knowledge Indicators (KIs) live in Elasticsearch indices named ai-index-* . Retrieve them by calling the esql_query tool with the query below. Substitute the users question for query , and corpus_entry as the ki_type for facts. esql FROM ai-index-idx-* METADATA _id, _index, _score | WHERE type ki_type | FORK (WHERE MATCH(content, query) OR MATCH(description, query) | SORT _score DESC | LIMIT 20) (WHERE MATCH(content.semantic, query) OR MATCH(description.semantic, query) | SORT _score DESC | LIMIT 20) | FUSE | SORT _score DESC | KEEP title, content, description, tags | LIMIT 5 Ground your answer in what the query returns, and cite the KI titles you used. If nothing relevant comes back, say so rather than guessing.将其保存为skills/query-ki/SKILL.md。下面是这个 skill 所执行的操作我们将corpus_entry定义为 KI 使用场景。我们在 AI indices 上执行混合 ES|QL 搜索按照适当的type进行过滤并使用 reciprocal rank fusionRRF作为融合结果的默认方法。在判断哪些事实与用户查询相关时KI 结果将直接为 agent 的答案提供依据。当我们说 AI indices 和 KI 与harness 无关时是因为这个 skill 本质上只是指令加查询。它可以在 Elastic Agent Builder、Kibana workflow agent、Claude Code 或任何其他 harness 中运行。我们将使用 Deep Agents 来演示如何在 Kibana 生态系统之外查询它。由于 AI Index 的核心就是一个 Elasticsearch index因此你也可以直接探索其中的数据。为 agentic RAG 将事实预先计算为 KI在这个示例中我们提取实际事实让 agents 无需消耗完整文档就能检索答案。我们为选定的每个文档生成一个基于事实的 KI不过你实际生成的 KI 数量和结构完全可以自定义。我们将使用 BrowseComp-Plus 语料库的一个样本将其建立索引到browsecomp-plusindex 中其中包含docid、url、title和text字段。基线使用 RRF 检索完整文档首先这是一个简单的 RRF 查询POST /_query?formattxt { query: FROM browsecomp-plus METADATA _score, _id, _index | FORK (WHERE match(title, What was the actress who played Torvi from Vikings also known for?) | SORT _score DESC | LIMIT 100) (WHERE match(text, What was the actress who played Torvi from Vikings also known for?) | SORT _score DESC | LIMIT 100) | FUSE // uses RRF by default | SORT _score DESC | KEEP _id, title, text | LIMIT 10 }这会将数百个单词的原始正文内容直接放入模型的上下文中。它可能有效但成本很高而且每次检索失败都会进一步增加成本。构建 Kibana workflow下面的 workflow 使用一次 ES|QL 查询读取一批文档并将每个文档生成一个事实级别的 KI写入 AI Index。每次迭代执行两个步骤generate_ki将原始文档提炼为结构化 KI而sink_ki则以docid作为键将其写入 AI Index从而确保重复运行具有幂等性。将下面的 YAML 复制并粘贴到 Elastic Workflows 编辑器中version: 1 name: browsecomp-plus-doc-ki description: Query the BrowseComp-Plus corpus with ES|QL, generate a KI per doc with an AI agent, and bulk-write each into the AI Index as a corpus_entry. enabled: true tags: - precomputed-context - browsecomp-plus triggers: - type: manual steps: - name: query_corpus type: elasticsearch.esql.query with: # WHERE drops empty bodies and restricts to the curated KI_DOCIDS -- the # specific documents this examples question depends on -- so the workflow # generates only a handful of KIs instead of one per corpus document. # SUBSTRING keeps the prompt bounded (a full body would blow the context window). # Column order drives the foreach.item[N] indices: # item[0]docid item[1]title item[2]url item[3]text query: FROM browsecomp-plus | WHERE text IS NOT NULL AND docid IN (11589, 50639, 64501, 41758, 57766, 84983, 82008) | KEEP docid, title, url, text | EVAL text SUBSTRING(text, 1, 12000) - name: loop_corpus_docs type: foreach foreach: {{ steps.query_corpus.output.values }} steps: # Turn the raw doc into a retrieval-optimized Knowledge Indicator. - name: generate_ki type: ai.agent timeout: 300s with: message: You are a knowledge engineer building a Knowledge Indicator (KI) for an enterprise document-retrieval corpus. A KI is a compact, high-signal record that a hybrid (BM25 semantic) search engine and an AI agent use to FIND and JUDGE the source document without reading it in full. Read the document below and extract a faithful, richly structured KI. Follow these rules strictly: - Be 100% grounded: never state anything not supported by the text. - Prefer concrete, named specifics (people, organizations, products, dates, places, figures) over vague phrasing. - Write for retrieval, not prose flourish. No marketing language. - If a field cannot be determined from the text, return an empty string or empty array rather than guessing. Document ID: {{ foreach.item[0] }} Original Title: {{ foreach.item[1] }} Source URL: {{ foreach.item[2] }} Document Body: {{ foreach.item[3] }} schema: type: object properties: title: type: string description: A concise, specific, human-readable title ( 12 words). summary: type: string description: A dense 3-5 sentence factual summary capturing the documents main claims, named entities, and conclusions. PRIMARY semantic search surface. answers_questions: type: array items: type: string description: 2-5 natural-language questions this document can authoritatively answer. key_entities: type: array items: type: string description: 3-10 salient named entities (people, organizations, products, places, dates) explicitly mentioned in the text. topics: type: array items: type: string description: 3-8 short topic/category labels. tagline: type: string description: A single ultra-short phrase ( 6 words) as a quick-reference label. required: - title - summary - answers_questions - key_entities - topics # Direct bulk write to the AI Index. The explicit index action row sets # _id docid so re-runs upsert in place (idempotent). index: in with # supplies the default target index for the bulk request. - name: sink_ki type: elasticsearch.bulk with: index: ai-index-idx-my-corpus operations: - index: _id: {{ foreach.item[0] }} - timestamp: {{ execution.startedAt | date: %Y-%m-%dT%H:%M:%S.%LZ }} type: corpus_entry title: {{ foreach.item[1] | default: steps.generate_ki.output.structured_output.title }} tags: - browsecomp-plus references: uri: {{ foreach.item[2] }} attributes: docid: {{ foreach.item[0] }} url: {{ foreach.item[2] }} source_index: browsecomp-plus tagline: {{ steps.generate_ki.output.structured_output.tagline }} topics: {{ steps.generate_ki.output.structured_output.topics | json }} answers_questions: {{ steps.generate_ki.output.structured_output.answers_questions | json }} key_entities: {{ steps.generate_ki.output.structured_output.key_entities | json }} content: SOURCE / PROVENANCE Backing Elasticsearch index: browsecomp-plus Document ID (docid): {{ foreach.item[0] }} Source URL: {{ foreach.item[2] }} Retrieve the full original document with ES|QL: FROM browsecomp-plus | WHERE docid {{ foreach.item[0] }} KNOWLEDGE INDICATOR {{ steps.generate_ki.output.structured_output.summary }} Questions this document answers: {{ steps.generate_ki.output.structured_output.answers_questions | join: | }} Key entities: {{ steps.generate_ki.output.structured_output.key_entities | join: , }} description: {{ steps.generate_ki.output.structured_output.tagline }}. Topics: {{ steps.generate_ki.output.structured_output.topics | join: , }}. Entities: {{ steps.generate_ki.output.structured_output.key_entities | join: , }}.下面是这个 workflow 所执行的操作query_corpus针对browsecomp-plusindex 运行 ES|QL 查询并应用一些规则例如丢弃正文为空的文档并将每个正文截断到 12,000 个字符以确保 agent prompt 保持在上下文窗口范围内。注意在这个示例中我们会挑选一些具体的 KI ID因为如果为 index 中的每个文档都生成 KI会花费很长时间而我们希望跟着示例操作的人能够在较短时间内完成这个练习。loop_corpus_docs遍历所有返回的文档并针对每个文档依次运行以下两个步骤generate_ki读取文档并调用 LLM 生成一个严格基于文档内容、有结构的 KI。sink_ki将每个 KI 批量写入 AI Indexai-index-idx-my-corpus并将其作为corpus_entry类型的 KI。它强制将_id设置为与文档的docid相同从而确保重复运行 workflow 时具有幂等性。总而言之这个 workflow 会将每个原始语料库文档转换为紧凑、可搜索的元数据记录让 agents 无需将完整源文档加载到上下文窗口中就能找到这些记录并判断其内容。这个 workflow 仅用于示例同样适用第 1 部分中提到的foreach注意事项。如果需要进行大规模处理请使用 workflow.executeAsync 或原生并行支持。cheat sheet 对优化 Workflows 很有帮助。在生产环境中通过使用 ai.prompt 或选择不同的模型来创建 KI也可能进一步降低成本并提高效率。检查 AI Index 中的 KIWorkflow 运行后你可以查询 AI Index查看写入其中的内容FROM ai-index-idx-* | WHERE type corpus_entry | KEEP title, description, attributes, tags | LIMIT 25下面是其中一个 KI 文档的示例{ _index: ai-index-idx-my-corpus, _id: 57766, _version: 1, _seq_no: 0, _primary_term: 1, found: true, _source: { timestamp: 2026-08-05T20:35:39.034Z, type: corpus_entry, title: Vikings (TV series) - Wikipedia, tags: [ browsecomp-plus ], references: { uri: https://en.wikipedia.org/wiki/Vikings_%28TV_series%29 }, attributes: { docid: 57766, url: https://en.wikipedia.org/wiki/Vikings_%28TV_series%29, source_index: browsecomp-plus, tagline: Ragnar Lothbroks rise and legacy, topics: [Historical drama television,Viking Age,Norse mythology and sagas,Canadian-Irish co-production,Television cast and production,Medieval Scandinavia], answers_questions: [When did the Vikings TV series premiere and on which network?,Who created and wrote the Vikings TV series?,Where was the Vikings TV series filmed?,Who are the main cast members of Vikings?,What historical and literary sources inspired the Vikings TV series?], key_entities: [Michael Hirst,Travis Fimmel,Katheryn Winnick,History Channel,Amazon Prime Video,Ashford Studios,County Wicklow, Ireland,Vikings: Valhalla,Ragnar Lodbrok,Wardruna] }, content: SOURCE / PROVENANCE Backing Elasticsearch index: browsecomp-plus Document ID (docid): 57766 Source URL: https://en.wikipedia.org/wiki/Vikings_%28TV_series%29 Retrieve the full original document with ES|QL: FROM browsecomp-plus | WHERE docid 57766 KNOWLEDGE INDICATOR Vikings is a historical drama television series created and written by Michael Hirst, co-produced between Canada and Ireland, that premiered on the History Channel on March 3, 2013, and concluded on March 3, 2021, after 6 seasons and 89 episodes. The series is inspired by the sagas of legendary Norse hero Ragnar Lodbrok — drawing on 13th-century texts Ragnars saga Loðbrókar and Ragnarssona þáttr, as well as Saxo Grammaticus Gesta Danorum — and follows Ragnars rise from farmer to Scandinavian king, then the exploits of his sons across England, Scandinavia, Kievan Rus, the Mediterranean, and North America. Principal cast includes Travis Fimmel as Ragnar Lothbrok, Katheryn Winnick as Lagertha, Gustaf Skarsgård as Floki, and Alexander Ludwig as Bjorn Ironside, among many others. The series was filmed entirely in Ireland at Ashford Studios and County Wicklow, with additional location shoots in Iceland, Morocco, Norway, and Canada; the first season budget was US$40 million. A sequel series, Vikings: Valhalla, premiered on Netflix on February 25, 2022. Questions this document answers: When did the Vikings TV series premiere and on which network? | Who created and wrote the Vikings TV series? | Where was the Vikings TV series filmed? | Who are the main cast members of Vikings? | What historical and literary sources inspired the Vikings TV series? Key entities: Michael Hirst, Travis Fimmel, Katheryn Winnick, History Channel, Amazon Prime Video, Ashford Studios, County Wicklow, Ireland, Vikings: Valhalla, Ragnar Lodbrok, Wardruna , description: Ragnar Lothbroks rise and legacy. Topics: Historical drama television, Viking Age, Norse mythology and sagas, Canadian-Irish co-production, Television cast and production, Medieval Scandinavia. Entities: Michael Hirst, Travis Fimmel, Katheryn Winnick, History Channel, Amazon Prime Video, Ashford Studios, County Wicklow, Ireland, Vikings: Valhalla, Ragnar Lodbrok, Wardruna. } }从 LangChain Deep Agents 查询 KI我们将使用 LangChain Deep Agents 和一个兼容 OpenAI 的 key来展示 AI indices 和 KI 可以与任何 agent harness 配合使用无论是在 Kibana 的 Agent Builder 生态系统内部还是外部。首先让我们创建facts_baseline_agent.py在应用 KI 之前测量我们的基线# Example question: What was the actress who played Torvi from Vikings also known for? import os import sys import time from elasticsearch import Elasticsearch from langchain_core.messages import AIMessage from langchain_core.tools import tool from langchain_openai import ChatOpenAI from deepagents import create_deep_agent if len(sys.argv) 2: sys.exit(fUsage: python {sys.argv[0]} your question) es Elasticsearch(os.environ[ES_URL], api_keyos.environ[ES_API_KEY]) tool def esql_query(query: str) - list[dict] | str: Execute an ES|QL query against Elasticsearch and return the matching rows. Args: query: A complete ES|QL query string, e.g. FROM browsecomp-plus | LIMIT 5. Full-text search syntax: WHERE MATCH(field, value) — not field MATCH value. try: resp es.esql.query(queryquery, formatjson) cols [c[name] for c in resp[columns]] return [dict(zip(cols, row)) for row in resp[values]] except Exception as e: return fES|QL error: {e} tool def get_mapping(index: str) - dict: Return the field mapping for an Elasticsearch index or pattern. return es.indices.get_mapping(indexindex).body baseline_agent create_deep_agent( modelChatOpenAI( # any OpenAI-compatible endpoint; configure via LLM_* env vars base_urlos.environ.get(LLM_BASE_URL, https://openrouter.ai/api/v1), modelos.environ.get(LLM_MODEL, anthropic/claude-sonnet-4.5), api_keyos.environ[LLM_API_KEY], ), tools[esql_query, get_mapping], # no query-ki skill system_prompt( You are a research assistant answering questions about a document corpus stored in the Elasticsearch index browsecomp-plus (fields: docid, url, title, text). You have NOT memorized the corpus. Answer by querying the raw index directly with ES|QL via the esql_query tool. Full-text search syntax: WHERE MATCH(field, \value\) — never use field MATCH \value\. Use get_mapping if you are unsure of field names. Ground your answer strictly in the rows returned, and cite the docid or url you used. ), ) start time.perf_counter() result baseline_agent.invoke( { messages: [ { role: user, content: sys.argv[1], } ] } ) latency time.perf_counter() - start print(\n--- Tool calls ---) for m in result[messages]: if isinstance(m, AIMessage) and m.tool_calls: for tc in m.tool_calls: print(f [{tc[name]}] {str(tc[args])[:120]}) total sum( len(m.tool_calls) for m in result[messages] if isinstance(m, AIMessage) and m.tool_calls ) print(fTotal: {total}\n) print(--- Usage ---) input_tokens sum( (m.usage_metadata or {}).get(input_tokens, 0) for m in result[messages] if isinstance(m, AIMessage) and m.usage_metadata ) output_tokens sum( (m.usage_metadata or {}).get(output_tokens, 0) for m in result[messages] if isinstance(m, AIMessage) and m.usage_metadata ) print(fTokens: {input_tokens output_tokens} (input {input_tokens}, output {output_tokens})) print(fLatency: {latency:.2f}s\n) print(--- Answer ---) print(result[messages][-1].content)当我运行这个基线查询《Vikings》中饰演 Torvi 的女演员还因什么而闻名/What was the actress who played Torvi from Vikings also known for?时它输出了以下内容% python facts_baseline_agent.py What was the actress who played Torvi from Vikings also known for? --- Tool calls --- [esql_query] {query: FROM browsecomp-plus | WHERE MATCH(text, Torvi Vikings actress) | LIMIT 10} [esql_query] {query: FROM browsecomp-plus | WHERE MATCH(text, Georgia Hirst) | LIMIT 10} [read_file] {file_path: /large_tool_results/toolu_bdrk_01SVEwHdtc9uE1zdfgCwchb5, offset: 0, limit: 50} [esql_query] {query: FROM browsecomp-plus | WHERE MATCH(text, Torvi Vikings TV) | LIMIT 10} [esql_query] {query: FROM browsecomp-plus | WHERE MATCH(text, Vikings series cast actress) | LIMIT 10} [esql_query] {query: FROM browsecomp-plus | WHERE MATCH(text, Michael Hirst Vikings) | LIMIT 10} [esql_query] {query: FROM browsecomp-plus | WHERE MATCH(text, Torvi) | LIMIT 20} [read_file] {file_path: /large_tool_results/toolu_bdrk_01PS3ryzML5BstqMXnV4c65N, offset: 0, limit: 100} Total: 8 --- Usage --- Tokens: 386187 (input 384940, output 1247) Latency: 44.86s --- Answer --- Perfect! I found the information. Based on the document from docid 11589 (Georgia Hirsts Wikipedia page), I now have the answer to the question. **Georgia Hirst**, the actress who played Torvi in the History Channel series Vikings (2014-2020), is also known for: 1. **Her role as Becky in the zombie horror movie Ravers (2018)** - This film had its world premiere at the 2018 London FrightFest Film Festival. 2. **Being the daughter of Michael Hirst**, the creator and writer of Vikings. This is a notable connection thats frequently mentioned about her. 3. **Her role in the short film Dungeness (2018)**, directed by Stephen Gates, where she was the sole actor. 4. **Her role as Grace in the interactive rom-com Five Dates (2020)**. 5. **Her advocacy work** - In 2017, she made an appearance on Sky News urging young women to get smear tests after being diagnosed with precancerous cells at age 22. 6. **Her philanthropy** - She joined The Celebs in 2021 to record a cover of The Beatles Let It Be in support of the Mind charity. So while shes primarily known for her role as Torvi in Vikings (where she appeared in 63 episodes across seasons 2-6), shes also recognized for her work in horror films like Ravers, her health advocacy, and being part of the Hirst family that created the show.注意Deep Agents 会自动添加read_file工具来处理分页的工具结果因此它会出现在输出中。接下来让我们创建一个知道如何使用query-kiskill 的 agent即facts_ki_agent.py# Example question: What was the actress who played Torvi from Vikings also known for? import os import sys import time from elasticsearch import Elasticsearch from langchain_core.messages import AIMessage from langchain_core.tools import tool from langchain_openai import ChatOpenAI from deepagents import create_deep_agent from deepagents.backends.filesystem import FilesystemBackend if len(sys.argv) 2: sys.exit(fUsage: python {sys.argv[0]} your question) es Elasticsearch(os.environ[ES_URL], api_keyos.environ[ES_API_KEY]) tool def esql_query(query: str) - list[dict] | str: Execute an ES|QL query against Elasticsearch and return the matching rows. Args: query: A complete ES|QL query string, e.g. FROM ai-index-idx-* | LIMIT 5. try: resp es.esql.query(queryquery, formatjson) cols [c[name] for c in resp[columns]] return [dict(zip(cols, row)) for row in resp[values]] except Exception as e: return fES|QL error: {e} # FilesystemBackend loads skills from disk, relative to root_dir. backend FilesystemBackend(root_dir., virtual_modeFalse) agent create_deep_agent( modelChatOpenAI( # any OpenAI-compatible endpoint; configure via LLM_* env vars base_urlos.environ.get(LLM_BASE_URL, https://openrouter.ai/api/v1), modelos.environ.get(LLM_MODEL, anthropic/claude-sonnet-4.5), api_keyos.environ[LLM_API_KEY], ), tools[esql_query], skills[skills], backendbackend, system_prompt( You are a research assistant answering questions about a document corpus. You have NOT memorized the corpus. When a question depends on specific facts, names, dates, or events, use the query-ki skill to retrieve Knowledge Indicators before answering. Ground your answer strictly in what it returns, and cite the KI titles you used. ), ) start time.perf_counter() result agent.invoke( { messages: [ { role: user, content: sys.argv[1], } ] } ) latency time.perf_counter() - start print(\n--- Tool calls ---) for m in result[messages]: if isinstance(m, AIMessage) and m.tool_calls: for tc in m.tool_calls: print(f [{tc[name]}] {str(tc[args])[:120]}) total sum( len(m.tool_calls) for m in result[messages] if isinstance(m, AIMessage) and m.tool_calls ) print(fTotal: {total}\n) print(--- Usage ---) input_tokens sum( (m.usage_metadata or {}).get(input_tokens, 0) for m in result[messages] if isinstance(m, AIMessage) and m.usage_metadata ) output_tokens sum( (m.usage_metadata or {}).get(output_tokens, 0) for m in result[messages] if isinstance(m, AIMessage) and m.usage_metadata ) print(fTokens: {input_tokens output_tokens} (input {input_tokens}, output {output_tokens})) print(fLatency: {latency:.2f}s\n) print(--- Answer ---) print(result[messages][-1].content)这个 agent 会查询 KI indices 来生成答案而不是将每个文档都加载到上下文中。当我使用相同的查询运行这个 agent 时得到的输出如下% python facts_ki_agent.py What was the actress who played Torvi from Vikings also known for? --- Tool calls --- [read_file] {file_path: ./query-ki/SKILL.md, limit: 1000} [esql_query] {query: FROM ai-index-idx-* METADATA _id, _index, _score\n| WHERE type corpus_entry\n| FORK\n (WHERE MATCH(co Total: 2 --- Usage --- Tokens: 27625 (input 27037, output 588) Latency: 15.22s --- Answer --- Based on the retrieved information, **Georgia Hirst** played Torvi in Vikings. She is also known for: 1. **Ravers (2018)** - A zombie horror film where she played the character Becky. This was her debut film and premiered at the London FrightFest Film Festival. 2. **Being the daughter of Michael Hirst** - The creator and writer of Vikings itself, making her connection to the show a notable family affair. Her older half-sister Maude Hirst also appeared in Vikings as Helga. 3. **Cervical cancer awareness advocacy** - She has publicly advocated for cervical cancer screening after being diagnosed with precancerous cells at age 22 and successfully recovering through treatment. 4. **Charity work** - In 2021, she participated in a celebrity cover of The Beatles Let It Be in support of the Mind charity (a mental health organization), alongside Anne Hegerty, Ivan Kaye, Eunice Olumide, and Shona McGarty. **Sources cited:** Georgia Hirst and Georgia Hirst - Wikipedia Knowledge Indicators from the AI Index.预先计算事实可以减少多少 agent 的 token 使用量两个 agents 得出了相似的结论但它们采取了截然不同的路径对于相同的问题和相同的有依据的答案从 KI 中获取答案时token 数量减少了 93%工具调用次数也从 8 次减少到 2 次。指标基线无 AI Index使用 AI Index工具调用总数82read_file调用次数21esql_query调用次数6全部针对browsecomp-plusindex1来自ai-index-idx-*消耗的 token386,18727,625延迟44.86 秒15.22 秒答案有依据、正确有依据、正确具体的工具调用次数、延迟和答案会因运行情况以及使用的 agents 不同而有所变化。两个 agents 都生成了可靠且有依据的答案。区别在于成本。从 AI Index 查询 KI 将 token 使用量降低了 93%并将延迟降低了大约三分之二。下面是两条路径的并排对比这次过程比较深入但它展示了 AI indices 和 Workflows 结合使用能够实现什么相同的答案只需消耗一小部分 token。在 Elasticsearch Serverless 中构建预先计算的上下文本实践指南展示了如何基于文档化的事实生成更复杂的 KI并使用 Elasticsearch 原生功能查询这些 KI以满足知识检索使用场景。在 agentic search 系统中上下文管理至关重要。而从本质上来说上下文管理是一个检索问题。AI indices 可以帮助你在 Elastic Stack 中管理上下文。现在就可以在 Serverless 中进行尝试并通过我们的 Discuss forums 或 Community Slack 中的#stack-kibanachannel 告诉我们你的想法。我们也很希望听听你有哪些希望通过 AI indices 解决的使用场景。原文Agentic RAG without reading the documents | Elasticsearch Labs
返回列表