<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[集成 - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[集成 - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/cn/search-labs/blog/category/integrations</link>
    </image>
    <link>https://www.elastic.co/cn/search-labs/blog/category/integrations</link>
    <atom:link href="https://www.elastic.co/cn/search-labs/rss/category/integrations.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[cn]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 22:10:44 GMT</lastBuildDate>
  <item>
    <title><![CDATA[Kibana 仪表板 API：为每种面板类型提供稳定的 API 规范，正式发布前已经过 50 多个团队测试]]></title>
    <description><![CDATA[以代码形式管理 Kibana 仪表板：提交至 Git、跨环境发布，并借助 Kibana API 和 Terraform 实现部署自动化。]]></description>
    <content:encoded><![CDATA[<p><a href="https://dashboardsapispec.kibana.dev/dashboards#tag/Dashboards">Kibana 仪表板和可视化 API</a> 在 Elastic 9.5 中已可用于生产环境，所有订阅级别均可使用，并提供完全向后兼容性。以 JSON 格式定义仪表板，将其提交到 Git，然后使用持续集成和持续部署 (CI/CD) 管道、<a href="https://registry.terraform.io/providers/elastic/elasticstack/latest/docs/resources/kibana_dashboard">Terraform</a> 或您现有的任何工具，将其部署到不同环境。在 <a href="https://www.elastic.co/search-labs/blog/kibana-dashboards-as-code-terraform-api">9.4 技术预览期间</a>，50 多个团队测试了该 API，其中一些团队已将其用于生产环境。9.5 版本还新增了用于<a href="https://dashboardsapispec.kibana.dev/tags.html">标签</a>的终端（技术预览阶段）；<a href="https://dashboardsapispec.kibana.dev/markdowns.html">Markdown</a> 和<a href="https://dashboardsapispec.kibana.dev/links.html#tag/Links">链接</a>面板终端现已在 Elastic Cloud Serverless 中提供，并将在 9.6 中推出。</p><h2>Kibana 仪表板 API 中的向后兼容性意味着什么</h2><p>在技术预览期间，API 结构可能会随版本发生变化。[1]现在已不再如此。正式发布 (GA) 意味着：</p><ul><li><p><strong>完全向后兼容。</strong>新字段和面板类型将陆续添加，但现有字段和行为保持不变。未来如需引入任何破坏性变更，都会经过慎重评估，并且只会在新的 Elastic Stack 主版本中引入。</p></li><li><p><strong>可用于生产环境，并提供全面支持。</strong>该 API 享有 Elastic 提供的全面支持保障。您可以放心地在生产环境中使用该 API，进行自动化部署、跨环境发布以及以编程方式管理仪表板。</p></li></ul><h2>Kibana 新增用于标签、Markdown 和链接面板的 API 终端</h2><p>Elastic 9.5 还新增了一个用于<a href="https://dashboardsapispec.kibana.dev/tags.html"><strong>标签</strong></a>的独立终端，让您可以对仪表板进行分类和筛选。现在，您可以通过专用 CRUD 终端以编程方式管理这些标签，从而更轻松地跨环境大规模组织仪表板。	</p><p>新的 <a href="https://dashboardsapispec.kibana.dev/markdowns.html"><strong>Markdown</strong></a> 和<a href="https://dashboardsapispec.kibana.dev/links.html#tag/Links"><strong>链接</strong></a>面板终端现已在 Serverless 中提供，并将在下一个 Elastic Stack 版本 (9.6) 中推出。</p><h2>Kibana 仪表板 API 支持哪些面板类型？</h2><p>仪表板 API 支持 9.5 中的所有<em>按值</em>面板（即直接在仪表板中定义的面板，而非保存以供重复使用的库面板）。每种受支持的面板类型都有一个类型化且经过验证的架构。</p><p><strong>面板类型</strong></p><p><strong>状态</strong></p><p>XY 图表</p><p>支持</p><p>指标</p><p>支持</p><p>饼图</p><p>支持</p><p>仪表盘图</p><p>支持</p><p>热图</p><p>支持</p><p>数据表</p><p>支持</p><p>树状图</p><p>支持</p><p>Discover 会话</p><p>支持</p><p>控件</p><p>支持</p><p>Markdown</p><p>支持</p><p>链接</p><p>支持</p><p>ML 面板</p><p>支持</p><p>Observability 面板</p><p>支持</p><p>Maps</p><p>即将推出</p><p>Vega</p><p>即将推出</p><h2>如何以代码形式管理 Kibana 仪表板</h2><p>仪表板 API 支持一套完整的以代码形式管理仪表板的工作流：将仪表板导出为简洁、可进行差异比较的 JSON，将其提交到 Git 并作为唯一可信来源，在拉取请求中审查更改，然后将同一定义部署到开发、预发布和生产环境。一旦以代码形式管理仪表板，就应将 Git 作为唯一可信来源：下次部署会覆盖直接在 UI 中所做的更改。</p><p>在空间、集群或阶段之间迁移仪表板时，主要挑战在于仪表板会按 ID 引用 Data view、库可视化等对象。由于这些 ID 是自动生成的，且不同环境中的 ID 各不相同，因此从一个环境导出的仪表板可能会引用另一个环境中不存在的对象。有三种方法可以处理这个问题，以下按自动化程度从高到低列出：</p><ul><li><p><strong>使用 Terraform。</strong><a href="https://registry.terraform.io/providers/elastic/elasticstack/latest/docs/resources/kibana_dashboard">Elastic Stack Terraform 提供程序</a>会跟踪每项资源，并自动映射各环境中的 ID，因此当您将仪表板从开发环境发布到生产环境时，引用可保持一致。</p></li><li><p><strong>定义按值的 </strong><a href="https://www.elastic.co/docs/explore-analyze/visualize/esorql"><strong>Elasticsearch 查询语言 (ES|QL) 面板</strong></a><strong>。</strong>构建面板时，可移植性最高的方法是在仪表板中直接使用 ES|QL 定义其可视化。<a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql-kibana">ES|QL</a> 查询会从查询中指定的索引读取数据，因此面板不会包含对 Data view 或库对象的外部引用。这样便可获得一个完全自包含、可移植的仪表板。</p></li><li><p><strong>分配一致的 ID。</strong>如果您引用 Data view、库可视化等已保存对象，请使用 PUT (upsert) 并指定 ID 来创建这些对象，而不要使用会自动生成 ID 的 POST。使用易读的 ID（例如 logs-prod），这样更便于在不同环境中重复使用和识别。</p></li></ul><p><a href="https://www.elastic.co/docs/explore-analyze/dashboards/manage-dashboards-as-code#dashboards-as-code-portability">有关这些可移植性模式以及完整的以代码形式管理仪表板的工作流的详细介绍，请参阅“以代码形式管理仪表板”文档。</a></p><h3>使用 PUT 通过仪表板 API 创建 Kibana 仪表板</h3><p>下面是一个简单示例：使用 PUT 而非 POST 创建包含指标面板的仪表板，并以仪表板名称 (service-health-overview) 作为自定义 ID。同样的逻辑也适用于创建保存到库中的独立可视化。</p>PUT kbn:/api/dashboards/service-health-overview
{
  "title": "服务运行状况概览",
  "description": "通过 API 管理的关键服务指标",
  "tags": [
    "production",
    "sre-team"
  ],
  "panels": [
    {
      "type": "vis",
      "grid": {
        "x": 0,
        "y": 0,
        "w": 12,
        "h": 8
      },
      "config": {
        "title": "错误率 (5xx)",
        "type": "metric",
        "data_source": {
          "type": "esql",
          "query": "FROM logs-* | WHERE http.response.status_code &gt;= 500 | STATS error_rate=count(*) BY host.name"
        },
        "metrics": [
          {
            "type": "primary",
            "column": "count"
          }
        ]
      }
    }
  ]
}<h2>Kibana 仪表板 API 路线图：Maps、Vega 和独立终端</h2><p>我们正在积极扩展 API 的功能范围。下一步将支持 Maps 和 Vega 面板，并为其添加类型化架构。我们还在为 Discover 会话（除了现有的仪表板面板支持之外）、Vega、Maps 和注释构建独立的 CRUD 终端，使其与仪表板生命周期解耦。</p><p>有关完整的架构定义，请参阅<a href="https://dashboardsapispec.kibana.dev/dashboards#tag/Dashboards">仪表板 API 文档</a>。对于 Terraform 用户，<a href="https://registry.terraform.io/providers/elastic/elasticstack/latest/docs/resources/kibana_dashboard">Elastic Stack Terraform 提供程序</a>支持正式发布 (GA) 的仪表板 API。</p><h2>注意</h2><ol><li><p>核心终端自技术预览以来未发生变化。如果您基于 9.4 构建了集成，这些集成在 9.5 中也可正常运行。仅有两项轻微的破坏性变更，分别影响仪表板列表和时长单位格式，详见<a href="https://www.elastic.co/docs/release-notes/kibana/breaking-changes">此处</a>。</p></li></ol>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/dashboards-as-code-kibana-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/dashboards-as-code-kibana-api</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[开发者体验]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Teresa Alvarez Soler]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8ed7e33de291f255/6a730619c8b7ac02b251f9d3/image1.png" length="0" type="image/png"/>
    <pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[不到 5 分钟即可完成本地部署：Jina 嵌入模型现已支持本地部署]]></title>
    <description><![CDATA[所有 28 个 Jina AI 模型（包括重排序器）均作为可随时部署的 Docker 容器提供，零遥测且无许可证服务器。即插即用兼容 OpenAI、Cohere、Voyage AI 和 Elastic 推理服务 API。]]></description>
    <content:encoded><![CDATA[<p>所有 28 个 Jina AI 嵌入和重排序模型现已作为完全离线的 Docker 容器提供，用于本地部署，其中包括 <a href="https://www.elastic.co/cn/search-labs/blog/jina-embeddings-v5-omni-all-media-one-index">jina-embeddings-v5-omni</a><a href="https://www.elastic.co/cn/search-labs/blog/jina-embeddings-v5-omni-all-media-one-index"></a> 和 <a href="https://www.elastic.co/cn/search-labs/tutorials/jina-tutorial/jina-reranker-v3">jina-reranker-v3</a>。下载一个，将其传输到本地隔离的或受防火墙保护的系统中，本地推理即可在五分钟内运行起来。这些容器是完全独立的，不会建立任何外部连接。不会对 Hugging Face 或任何模型注册表进行调用。同时也没有许可证服务器、遥测或日志记录终端。对于受监管行业、有数据主权要求或互联网访问不可靠或根本无法使用的环境，这消除了对第三方 AI 服务的依赖。Jina 本地部署支持 Elastic 推理服务 (EIS)、OpenAI、Cohere、Voyage AI 和 Gemini API 模式，因此现有应用程序无需修改代码即可运行。</p><p>最强大的 AI 模型在远程云安装上运行，通过 Web API 进行访问，这意味着您必须信任您的 AI 服务提供商的安全、服务可用性和稳定的价格。您很难兼顾对可靠性、隐私、可控成本和良好数据治理的合理需求，与日益强大、复杂且资源密集型的 AI 使用。</p><p>政府监管、法院裁决以及出于他人利益考虑而做出的商业决策，最近都导致对特定服务的访问受限。而且，即使您可以切换到其他服务，AI 模型也不是那种只要您想换就能随时替换的组件。使用语义嵌入的应用程序依赖于在查询时和数据摄取时能够访问相同的模型。无法访问嵌入模型意味着您的搜索系统将陷入瘫痪。</p><p>AI 定价模式加剧了这种风险。主要 AI 供应商最近披露的财务信息，让客户有充分理由对潜在的价格上涨感到担忧。依赖成本不可预测的产品，会给可能无法产生明确回报的资本密集型 AI 投资增加更多风险。</p><p>Jina 本地部署是 Elastic 应对这些挑战的解决方案。</p><h2>谁需要本地部署 AI？</h2><p>对您的 AI 模型进行本地托管和直接控制，支持各种技术需求、行业要求和商业利益。</p><p>本地安装可以减少您支付给 AI 服务提供商的费用，但会将硬件和可靠访问的成本转嫁给您的组织。根据您的使用量，它可能确实更便宜。但还有其他紧迫的原因促使您考虑自行运行 AI。如果下面描述的任何问题涉及您的企业，请考虑采用像 Jina 本地部署这样的本地 AI 解决方案。此列表并非详尽无遗。</p><p>用例</p><p>为什么选择本地部署</p><p>示例</p><p>隔离/高安全</p><p>无出站数据传输；完全网络隔离</p><p>国防、情报、机密研究</p><p>监管合规性</p><p>数据主权；无跨境传输或第三方暴露</p><p>医疗保健（《健康保险流通与责任法案》[HIPAA]）、金融、欧盟企业（《通用数据保护条例》[GDPR]）</p><p>延迟关键型</p><p>零网络依赖；对连接故障零容忍</p><p>机器人技术、边缘计算、车辆、船舶</p><p>成本可预测性</p><p>固定基础架构成本与按代币定价且未来费率不确定</p><p>高容量连续推理工作负载</p><p>责任降低</p><p>无第三方数据泄露；维护法律特权与注意义务</p><p>律师事务所、政府机构</p><h3>为什么隔离的和受防火墙保护的系统需要本地部署 AI</h3><p>隔离的和受防火墙保护的系统无法使用外部 AI API。Jina 本地部署完全在您的基础架构内运行，无出站连接。</p><p>对于管理特别敏感数据的组织而言，安全和隐私考量至关重要。如果您将敏感数据随意提供给可能缺乏足够安全措施或可能受制于外国政府要求的远程第三方，那么在保护敏感数据方面的投入就毫无意义。</p><p>处理敏感数据的组织中的员工通常会接受一些安全数据处理方面的培训，但是当他们在处理数据时，他们的 Web 浏览器可能会打开互联网上的任何页面，因此这种培训效果并不理想。无论是通过隔离网络还是限制极为严格的防火墙，隔离都是最有效的安全措施，但这使得使用任何类型的外部服务都变得困难。</p><h3>适用于对延迟敏感和高可用性系统的本地部署 AI</h3><p>软件即服务和云计算代表了一种折衷方案，既能保证在自己的计算机上提供高度便捷、可靠的服务，又能降低将问题外包给其他人的成本。但它们也存在延迟不稳定、服务中断以及出现故障时完全失控等问题。AI 服务也不例外。如果无法访问嵌入模型时搜索系统离线，那么这可能不再是一个好的折衷方案。</p><p>此外，依赖外部 AI 始终会带来一些难以预见或管理的风险。由于政治事件、恶劣天气或船舶拖锚压住水下光纤电缆等原因，互联网访问和网络延迟可能会在毫无预警的情况下下降。政府可以且近期已经通过出口禁令突然封锁对 AI 模型的访问。AI 服务提供商有时会撤回模型，以促使您切换到较新的模型。必须在外部服务的灵活性和可控成本与依赖风险之间进行权衡。</p><h3>满足 GDPR、HIPAA 和数据主权合规要求的本地部署 AI</h3><p>收集个人数据的组织受到越来越严格的监管，这些监管在不同司法管辖区之间往往有所不同，并且可能存在相互矛盾的要求。值得注意的是，<a href="https://www.hhs.gov/hipaa/for-professionals/privacy/laws-regulations/index.html">HIPAA 规则</a>对美国医疗保健提供商施加了非常严格的数据保护要求，而<a href="https://laws-lois.justice.gc.ca/eng/acts/p-8.6/">加拿大</a>、<a href="https://gdpr-info.eu/">欧盟</a>以及<a href="https://www.japaneselawtranslation.go.jp/en/laws/view/4241">许多亚洲司法管辖区</a>强大的通用数据保护法律则要求所有处理个人信息的企业都必须安全地进行操作，并限制将此类数据传输给其他方或其他司法管辖区。如果外国实体在这些司法管辖区内有任何客户，这些规则甚至可以对其施加义务。金融机构通常会受到更为严格的规则约束，并且对于信息安全承担着与防范其他形式犯罪活动相同的直接责任。</p><p>监管合规性可能与第三方 AI 服务不兼容，尤其是在使用这些服务涉及跨境数据传输的情况下。</p><p>此外，近期发生的事件表明，当国际云运营商面临外国政府压力时，限制数据存储物理位置的规定可能并非可靠的保护手段。不同司法管辖区之间的本地法律可能会发生冲突，这不仅要求在本地进行数据存储和处理，还可能导致无法使用第三方服务。在某些情况下，唯一的解决方案是将您流程的所有部分（包括您的 AI 系统）都转为内部自主运营。</p><h3>来自第三方数据传输的 AI 责任风险</h3><p>数据保护法和公认的对敏感数据的注意义务通常会产生法律责任，有时甚至会产生非常严重的责任。您可能需要对第三方服务提供商处理您数据的行为承担责任。虽然法院和法律程序可能对不安全的服务提供商提供一些追溯保护，但这些补救措施对国家安全机构、执法部门或犯罪黑客来说既不可用，也通常无效。</p><p>对于政府而言，已经出现过跨境云服务提供商向外国行为体泄露敏感国家信息的情况。</p><p>但即使你不担心外国政府或黑客，而且你的外部 AI 服务提供商本身也是安全的，仅仅是他们是外部人员这一事实就可能造成责任风险。</p><p>例如，在大多数司法管辖区，律师与客户之间的沟通享有特殊的法律保护，律师事务所在记录或存储此类信息时承担着严格的责任。在美国，这种“律师与客户交流保密特权”非常著名，甚至成为了电影和电视剧情节的核心。但失去特权的一种途径是与没有特权的人交流信息，而最近的发展表明，外部 AI 服务提供商可能符合失去特权的条件。</p><p>至少在美国，仅仅通过互联网 API 使用第三方 AI 服务（例如嵌入提供索引服务的模型）就可能违反关键的保密规则。即使没有发生安全漏洞，律师事务所也可能仅仅因为使用外部托管软件而被起诉、受到处分或被吊销执照。</p><h3>适用于离线、边缘和物理隔离系统的本地部署 AI</h3><p>计算机系统隔离不仅仅是出于安全原因。例如，行驶中的车辆不能在任何关键功能上依赖互联网连接。船舶和飞机都拥有非常庞大的机载计算机系统，这些系统必须在没有互联网连接的情况下运行，因此无法使用外部 AI 服务。海上平台、荒野地区的远程设施，以及北极、南极和小岛上缺乏与全球网络充分物理连接的计算机服务，都是受益于在本地托管其所需全部服务的安装实例。随着 AI 在企业计算中的作用日益增强，解决这些局限性变得更加重要。</p><p>AI 在物理系统中的新兴应用（机器人技术和其他空间受限或以外部世界为中心的用例，如物流管理系统甚至超市收银台）可能会连接到全球互联网，但它们无法容忍连接故障或延迟峰值。如果它们依赖 AI 系统来运行，该 AI 系统就需要尽可能做到本地化和可靠。</p><h2>谁不需要本地部署 AI？</h2><p>远程软件服务和场外 AI 确实具有优势。运行 AI 模型可能需要昂贵且支持大量电力、以寿命极短而闻名的处理器。由于市场因素和外部经济冲击，目前获取高质量硬件尤为困难。在这种情况下，使用外部 API 支付代币费用，而不是承担本地 AI 高昂的资本成本，可能是明智之举。</p><p>对于偶尔使用的用户来说，外部 API 是最合理的选择。如果您使用 AI 模型主要是为了批量处理数据以进行分析，而不是运行一个必须时刻在线的搜索系统，那么投资资本密集型硬件和本地安装就没有什么意义。</p><p>此外，如果您的数据处理已经基于云，例如，出于可靠性和可访问性的原因，电子商务网站托管在云端，那么使用位于同一云基础架构中的 AI 服务可能比引入您自己的授权 AI 模型部署更具性价比。您已经依赖于云服务提供商，因此依赖其 AI 服务并不会增加太多的风险。</p><p>如果您的用例符合此描述，您可以在 <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">EIS</a>、<a href="https://aws.amazon.com/marketplace/seller-profile?id=seller-stch2ludm6vgy">AWS Marketplace</a> 和 <a href="https://console.cloud.google.com/marketplace/browse?q=jina">Google Cloud Platform</a> 上获取 Jina AI 模型，以专门满足您的需求。</p><p>下表总结了关键因素。您的答案取决于您的数据、基础架构以及使用模式。</p><p>因素</p><p>首选本地部署</p><p>首选云 API</p><p>使用模式</p><p>连续或高容量推理</p><p>间歇性处理或批量处理</p><p>数据敏感性</p><p>受监管、主权或机密</p><p>无跨境或第三方限制</p><p>网络环境</p><p>隔离、受防火墙保护或不可靠</p><p>稳定、始终在线的互联网</p><p>现有基础架构</p><p>拥有或能够采购 GPU 硬件</p><p>已在云端托管，并配备同地部署的 AI</p><p>成本模型</p><p>固定硬件 + 许可；规模化生产可预测</p><p>按代币定价；前期投入较低，长期收益可变</p><p>延迟容忍度</p><p>无（机器人技术、边缘、实时）</p><p>网络波动是可接受的</p><p>运营责任</p><p>您的团队负责管理硬件和可用性</p><p>提供商负责管理硬件和更新；您负责管理集成</p><p>您必须根据自身具体情况和用例来考虑成本和收益，同时考虑到前一节中强调的适用于您的问题。成本效益分析无疑会随时间而变化。我们甚至无法预测 AI 行业的未来或硬件价格的短期走势。</p><h2>推出 Jina 本地部署</h2><p>对于能够从本地 AI 服务中受益的用户，我们推出了 <a href="https://github.com/jina-ai/jina-on-prem/wiki/">Jina 本地部署</a>，这是一个完全独立的安装套件，用于 Jina AI 的高性能模型。</p><p>Jina AI 的模型准确度可与<a href="https://mteb-leaderboard.hf.space/benchmark/MTEB(Multilingual%2C%20v2)">规模大得多的</a>嵌入式模型相媲美，从而降低了计算成本、内存占用和硬件要求。这使得它们成为希望或需要保持 AI 本地部署的用户的理想选择。针对各种规模的用例提供商业许可以及按比例定价的可扩展解决方案。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt190865fb3ebde472/6a6a33d0065b162508701ff9/02559ceca556a26c53eb703ae87d421452b27251-1374x1400.png" alt="MMTEB Multilingual v2 leaderboard showing Jina AI embedding model rankings: jina-embeddings-v5-omni-small and jina-embeddings-v5-text-small ranked 13th, jina-embeddings-v5-omni-nano and jina-embeddings-v5-text-nano ranked 19th, competing against models from Microsoft, Google, Tencent, NVIDIA and Qwen" /><h3>Jina 本地部署支持哪些 API 模式？</h3><ul><li><p>可以作为完整的依赖集合进行本地安装，也可以作为 <a href="https://www.docker.com/">Docker 容器</a>进行安装和运行，几分钟即可完成。</p></li><li><p>Jina 本地部署安装<em>不会</em>调用外部系统。</p><ul><li><p>不会调用 Hugging Face Hub 或任何模型注册表（已内置 HF_HUB_OFFLINE=1 和 TRANSFORMERS_OFFLINE=1）。</p></li><li><p>没有许可证服务器。</p></li><li><p>没有遥测或日志记录终端。</p></li></ul></li><li><p>支持 CPU 和 GPU 硬件，并支持 GPU 自动检测。</p></li><li><p>所有 28 个 Jina AI 模型均已提供，包括最新的 <a href="https://www.elastic.co/cn/search-labs/blog/jina-embeddings-v5-omni-all-media-one-index">jina-embeddings-v5-omni</a> 多模态嵌入模型和 <a href="https://www.elastic.co/cn/search-labs/tutorials/jina-tutorial/jina-reranker-v3">jina-reranker-v3</a>。</p></li><li><p>通过标准 AI API 模式访问：<a href="https://jina.ai/api-dashboard">Jina API</a>、OpenAI、Cohere、Voyage AI 和 Gemini。Jina 本地部署是一个即插即用的解决方案，适用于基于这些模式构建的应用程序。</p></li><li><p>可直接替代 <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">EIS</a> 提供的模型。Jina 本地部署可直接与<a href="https://www.elastic.co/cn/blog/deploy-elastic-air-gapped-disconnected-environments">隔离的 Elastic 部署</a>集成。</p></li></ul><h2>Jina AI 本地部署模型的硬件要求</h2><p>不同的 Jina 模型所需的硬件要求各有不同。下表展示了使用 GPU 设置的最新模型的推荐配置。您不需要比 NVIDIA L4 GPU 更强大的设备，不过对于 v5 嵌入模型，建议使用 A100。我们最新的嵌入模型目前至少需要 8 GB 的 VRAM。</p><p>模型</p><p>最小 VRAM</p><p>推荐 GPU</p><p>jina-embeddings-v5-text-nano</p><p>2 GB</p><p>T4 / L4</p><p>jina-embeddings-v5-text-small</p><p>3 GB</p><p>L4 / A10G</p><p>jina-embeddings-v5-omni-small</p><p>8 GB</p><p>L4 / A10G / A100</p><p>jina-reranker-v3</p><p>3 GB</p><p>L4</p><p>jina-clip-v2</p><p>4 GB</p><p>L4</p><p>jina-code-embeddings-1.5b</p><p>4 GB</p><p>L4</p><p>ReaderLM-v2</p><p>4 GB</p><p>L4</p><p>如果同时使用多个模型，则 VRAM 需求将会增加。有关更多信息，请参阅<a href="https://github.com/jina-ai/jina-on-prem/wiki/Sizing-And-Hardware">规模和硬件页面</a>。</p><h2>如何使用 Docker 安装 Jina 本地部署</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt20265d09e2d8d0f4/6a6a33d1065b162105701ffd/ada9881af407168298b1940f8537ad71a5411c89-1999x1200.png" alt="" /><p>最快的开始方法是<a href="https://www.docker.com/get-started/">安装 Docker</a>（如果您尚未安装），并按照 <a href="https://github.com/jina-ai/jina-on-prem/wiki/QuickStart">Jina 本地部署快速入门</a>页面上的说明进行操作。</p><p>所有 28 个 Jina 模型均提供预构的 Docker 容器。下载其中一个并将其传输至您的安装目标，即可在不到 5 分钟的时间内运行 Jina AI 模型。</p><p>对于多模态或自定义构建，或者要下载在容器外部安装所需的完整依赖集，请按照<a href="https://github.com/jina-ai/jina-on-prem/wiki/Bundling-Guide">打包指南</a>中概述的步骤进行操作。</p><p>您的 Jina 本地部署安装支持所有 Jina API 和 EIS 功能以及通过 OpenAI、Cohere、Voyage AI 和 Gemini API 生成嵌入，因此可以使用标准接口集成到已有应用程序中。有关更多信息，请参阅 <a href="https://github.com/jina-ai/jina-on-prem/wiki/API-Reference">API 文档</a>。</p><p>Jina 模型（包括通过 Jina 本地部署安装的模型）按多种许可条款提供，其中最新的模型可在 <a href="https://creativecommons.org/licenses/by-nc/4.0/deed.en">CC BY-NC 4.0</a> 许可证下免费用于非商业用途。如需获得 Jina 本地部署的商业许可，请联系 <a href="https://www.elastic.co/cn/contact">Elastic 销售</a>。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/on-prem-ai-jina-embedding-models</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/on-prem-ai-jina-embedding-models</guid>
    <category><![CDATA[Jina AI]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Scott Martens]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt17731ab0c6ec66f6/6a6a33d140a4941014ca5c9a/09bc6dac4e6a86c7877f8ed78d68f5d581aeffa9-1999x1200.png" length="0" type="image/png"/>
    <pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[为 Elasticsearch 注入活力：添加对原生 Prometheus API 的支持]]></title>
    <description><![CDATA[通过原生 PromQL、发现及元数据终端，从兼容 Prometheus 的客户端直接查询 Elasticsearch。使用 Prometheus Remote Write 将数据发送到 Elasticsearch。]]></description>
    <content:encoded><![CDATA[<p>将任何与 Prometheus 兼容的客户端指向 Elasticsearch，直接对照现有指标运行 PromQL。Elasticsearch 正在以技术预览版形式添加原生 Prometheus 查询、发现及元数据终端，这些终端适用于通过 Prometheus Remote Write、OpenTelemetry 或批量 API 采集的指标。API 运行在 Elasticsearch 的时序数据流 (TSDS) 之上，因此无需额外维护独立的 Prometheus 专用存储层。</p><p>本文解释了查询、发现和元数据终端如何基于先前的摄取和查询工作，共同构成该 API 接口表面。多篇配套文章对各个部分进行了更深入的探讨：</p><ul><li><p><a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">ES|QL 中的原生 PromQL 支持</a>涵盖了如何将 PromQL 查询转换为 ES|QL 执行计划。</p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch">使用 Remote Write 将 Prometheus 指标发送到 Elasticsearch</a>，涵盖了数据摄取设置。</p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">How Prometheus Remote Write Ingestion Works in Elasticsearch</a> 一文介绍了 Prometheus 远程写入的摄取内部机制。</p></li></ul><p>这项工作仍在进行中。以下各节列出了目前支持的部分以及仍在发展的部分。</p><h2>API 接口表面</h2><p>如今，与 Prometheus 兼容的 API 接口表面分为三类。</p><h3>查询终端</h3><p>查询终端允许 Prometheus 兼容客户端评估 PromQL 表达式：</p><ul><li><p><code>GET /_prometheus/api/v1/query_range</code> 在一个时间窗口内（矩阵结果）评估 PromQL 表达式。</p></li><li><p><code>GET /_prometheus/api/v1/query</code> 在单一时间点评估（向量结果）。当前实现为返回最后一个样本的短范围查询。</p></li></ul><p>目前查询终端仅支持 GET 方法。某些客户端默认使用 POST 请求，因此您可能需要将其配置为使用 GET 请求。Prometheus 的 POST 约定使用 <code>application/x-www-form-urlencoded</code> 主体，Elasticsearch 的 HTTP 层会在请求到达处理程序之前将其拒绝，以防止 CSRF 攻击。</p><p>有关 PromQL 的完整覆盖状态，请参阅 <a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">ES|QL 中有关 PromQL 的配套文章</a>。</p><h3>元数据终端</h3><p>元数据终端提供客户端进行自动补全、变量下拉列表和指标浏览所需的发现信息。</p><p>系列、标签和标签值终端均接受 <code>match[]</code> 选择器和时间范围（<code>start</code>/<code>end</code>）。参数 <code>match[]</code> 接受一个 Prometheus 系列选择器（如 <code>http_requests_total{job="api"}</code>），并将响应限制为匹配的时间序列。这确保了在具有大量指标的集群上，响应速度快且相关性高。例如：</p>GET /_prometheus/api/v1/series?match[]=http_requests_total{job="api"}GET /_prometheus/api/v1/labels?match[]=http_requests_totalGET /_prometheus/api/v1/label/instance/values?match[]=http_requests_total{job="api"}<p>第一个返回 <code>http_requests_total</code> 中符合 <code>job="api"</code> 条件的所有系列及其完整标签集。第二个仅返回 <code>http_requests_total</code> 系列中存在的标签名称。第三个仅返回匹配序列中出现的 <code>instance</code> 值。</p><p><code>GET /_prometheus/api/v1/metadata</code> 有所不同：它返回每个指标的类型和单位，并可通过 <code>metric</code> 参数按名称筛选。</p>GET /_prometheus/api/v1/metadata?metric=http_requests_total<p>它不接受 <code>match[]</code> 选择器或时间范围。在 Prometheus 中，元数据是从活动抓取目标（它们公开的 <code>HELP</code>、<code>TYPE</code> 和 <code>UNIT</code> 行）收集的，因此响应不涉及数据扫描。Elasticsearch 没有那样的专用元数据存储，因此当前实现通过访问最近 24 小时的时间序列数据来发现指标元数据。这样可以保持快速查询，而无需进行完整的索引扫描。该 24 小时回看窗口是固定的：Prometheus 元数据 API 不公开 <code>start</code> 或 <code>end</code> 参数，Elasticsearch 无法利用这些参数来实现用户可调。</p><p>元数据终端的工作原理，包括支持它们的 <code>TS_INFO</code> 和 <code>METRICS_INFO</code> 命令，已在<a href="https://www.elastic.co/search-labs/blog//elasticsearch-native-prometheus-api#ts-info-and-metrics-info">下方</a>详细说明。</p><h3>索引预过滤</h3><p>所有查询和元数据终端都接受 <code>/_prometheus/</code> 之后的可选 <code>{index}</code> 路径段：</p>GET /_prometheus/metrics-prod-*/api/v1/query_range?query=up&amp;start=...&amp;end=...<p>这限制了在任何表达式评估开始之前，查询针对哪些 Elasticsearch 索引运行。在跨团队或环境拥有大量数据流的集群中，这避免了扫描无关索引，并能显著降低查询延迟。您可以为每个索引模式配置单独的数据源，以便团队获得对其自身指标的限定访问权限。</p><h3>关于远程写入的说明</h3><p>对于数据摄取，Elasticsearch 还公开了标准的 Prometheus Remote Write 终端：</p><ul><li><p><code>POST /_prometheus/api/v1/write</code> 通过 Prometheus Remote Write v1 协议摄取时间序列。v2 暂不支持。</p></li></ul><p>Remote Write 将数据写入 Elasticsearch 的现有时序数据流 (TSDS)，而不是一个单独的 Prometheus 特定存储层。Prometheus 标签变成 TSDS 维度，而度量名称变成索引映射中的字段。<a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">远程写入架构文章</a>详细介绍了完整的映射，包括如何推断指标类型以及如何使用 <code>labels.</code> 前缀存储标签。</p><h3>运作方式</h3><p>在底层，所有终端的工作方式相同：解析传入的 HTTP 参数，构建 ES|QL 查询计划，针对时序数据流执行，并将列式结果转换回 Prometheus 客户端期望的 JSON 格式。</p><h2>TS_INFO 和 METRICS_INFO</h2><p>元数据终端需要回答诸如“存在哪些标签？”或“定义了哪些指标类型？”等问题，要在可能数百万个时间序列中完成，而无需扫描每个数据点。</p><p>在内部实现上，Prometheus 元数据终端会通过围绕两个新的处理命令 <code>METRICS_INFO</code> 和 <code>TS_INFO</code> 构建 ES|QL 计划，来回答这些问题。您无需直接使用这些命令即可使用 Prometheus API，但它们是元数据响应背后的核心执行原语。两者都通过每个时间序列仅访问一个文档来提取其元数据，而不是扫描所有样本。这意味着它们的成本与不同时间序列的数量成正比，而不是与数据点的数量成正比。</p><p><code>METRICS_INFO</code> 为每个不同的指标返回一行，包含其名称、类型、单位和关联的维度字段。<code>TS_INFO</code> 更加细粒度：为每个（指标，时间序列）组合返回一行，将实际的维度值作为一个 JSON 对象包含在内。</p><p>一篇关于 <code>TS_INFO</code> 和 <code>METRICS_INFO</code> 的专门博客文章即将推出，内容将涵盖两阶段执行模型、它们的扩展方式，以及如何在 ES|QL 查询中直接使用它们，而不仅仅是通过 Prometheus API。</p><h3>元数据终端如何使用它们</h3><p>每个元数据终端都会都以这些命令之一为核心构建 ES|QL 计划。</p><p><code>/api/v1/labels</code> <code>/api/v1/series</code> 使用 <code>TS_INFO</code>，因为它们需要按时间序列的详细信息（存在哪些标签，每个系列由哪些维度值标识）。<code>/api/v1/metadata</code>和<code>/api/v1/label/__name__/values</code>使用<code>METRICS_INFO</code>，因为它们只需要每个指标的信息（指标名称、类型、单位）。</p><p><code>/api/v1/label/{name}/values</code> 对于常规标签（除 <code>__name__</code> 之外的任何标签），则不使用这两个命令。如 <code>job</code> 或 <code>instance</code>）的常规标签是索引中的实际维度字段，因此终端可以直接使用分组聚合来查询它们。当提供 <code>match[]</code> 选择器时，它们将转换为 <code>WHERE</code> 子句，在聚合运行之前对时序进行筛选。</p><p><code>__name__</code> 标签需要一种不同的策略，因为它并不总是作为维度字段存在。Prometheus Remote Write 确实存储 <code>labels.__name__</code>，但通过其他路径（OpenTelemetry、批量 API）摄取的指标则不会。指标名称被编码在字段名称本身中（例如 <code>metrics.http_requests_total</code>）。您可以通过查看索引映射来枚举字段名称，但仅凭映射无法告诉您哪个指标具有哪些维度，也无法通过 <code>match[]</code> 选择器的标签值进行筛选。<code>METRICS_INFO</code> 可以同时执行两项操作：跨索引枚举指标名称，同时使用上游 <code>WHERE</code> 筛选器。</p><p>在所有情况下，API 层都会将转换处理回 Prometheus 约定：去除 <code>labels.</code> 和 <code>metrics.</code> 存储前缀，并为缺少它的非 Prometheus 指标合成 <code>__name__</code>。</p><h2>结语</h2><p>结果：任何兼容 Prometheus 的客户端都可以通过其已经理解的终端来查询和探索 Elasticsearch 指标。Remote Write 指标、OpenTelemetry 指标以及通过其他途径索引的指标，都将通过同一 API 呈现，并由相同的 TSDS 索引提供支持。</p><p>本文提及的所有 Prometheus API 目前在 Elasticsearch Serverless 中以技术预览版形式提供。对于自管型集群和 Elastic Cloud Hosted 部署，除 <code>GET /_prometheus/api/v1/metadata</code> 外，其他 API 在 Elasticsearch 9.4 中以技术预览版形式提供。要在本地进行实验，请使用 <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">start-local</a>。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api</guid>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt12b4e100d5bbb7f0/6a16f7a22b835ff747f4afdd/c7b333bd73e8a1f4e18486b2d692ba742788dcfd-1376x768.jpg" length="0" type="image/jpeg"/>
    <pubDate>Mon, 11 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 TypeScript 构建 Elasticsearch MCP 服务器]]></title>
    <description><![CDATA[学习如何使用 TypeScript 和 Claude Desktop 创建 Elasticsearch MCP 服务器。]]></description>
    <content:encoded><![CDATA[<p>在 Elasticsearch 中处理大型知识库时，找到信息只是成功的一半。工程师通常还需要综合多个文档的结果，生成摘要，并追溯答案的来源。模型上下文协议 (MCP) 提供了一种标准化的方式，可将 Elasticsearch 与大语言模型 (LLM) 驱动的应用程序连接起来，以实现上述目标。虽然 Elastic 提供官方解决方案，例如 Elastic Agent Builder（其功能包括 <a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">MCP 终端</a>），但构建自定义 MCP 服务器可让您完全掌控搜索逻辑、结果格式，以及如何将检索到的内容传递给 LLM，以用于综合分析、生成摘要和提供引用。</p><p>本文将探讨构建自定义 Elasticsearch MCP 服务器的优势，并展示如何使用 TypeScript 创建该服务器，以将 Elasticsearch 连接到 LLM 驱动的应用程序。</p><h2>为什么要构建自定义 Elasticsearch MCP 服务器？</h2><p>Elastic 为 <a href="https://www.elastic.co/docs/solutions/search/mcp">MCP 服务器</a>提供了一些替代方案：</p><ul><li><p><a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">Elastic Agent Builder MCP 服务器，适用于 Elasticsearch 9.2 及以上版本</a></p></li><li><p><a href="https://github.com/elastic/mcp-server-elasticsearch?tab=readme-ov-file#elasticsearch-mcp-server">适用于旧版本的 Elasticsearch MCP 服务器（Python）</a></p></li></ul><p>如果您需要更好地控制 MCP 服务器与 Elasticsearch 的交互，构建自己的自定义服务器可以让您灵活地根据自身需求进行定制。例如，Agent Builder 的 MCP 终端仅限于 Elasticsearch 查询语言 (ES|QL) 查询，而自定义服务器允许您使用完整的查询 DSL。在将结果传递给 LLM 之前，您还可以控制结果的格式，并可以集成其他处理步骤，例如我们将在本教程中实现的由 OpenAI 驱动的摘要功能。</p><p>通过阅读本文，您将学会使用 TypeScript 创建 MCP 服务器，该服务器可搜索存储在 Elasticsearch 索引中的信息，对其进行总结并提供引用。我们将使用 Elasticsearch 进行检索，使用 OpenAI 的 <code>gpt-4o-mini</code> 模型提炼摘要并生成引用，并使用 Claude Desktop 作为 MCP 客户端和 UI 来接收用户查询并提供回复。最终我们将得到一个内部知识助手，帮助工程师在整个组织的技术文档中发现并综合最佳实践。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltad9133cb083ad352/6a170c19b0367d411e72bd5b/ec5771a874cf9740d4cac6888622cbe8cd6aede7-1999x1133.png" alt="使用 TypeScript 和 Claude Desktop 构建 Elastic MCP 服务器。" /><h2>准备工作：</h2><ul><li><p>Node.js 20 +</p></li><li><p>Elasticsearch</p></li><li><p>OpenAI API 密钥</p></li><li><p>Claude Desktop</p></li></ul><h3>什么是 MCP？</h3><p><a href="https://www.elastic.co/what-is/mcp">MCP</a> 是由 <a href="https://www.anthropic.com/news/model-context-protocol">Anthropic</a> 创建的开放标准，提供大型语言模型与外部系统（如 Elasticsearch）之间的安全双向连接。您可以在<a href="https://www.elastic.co/search-labs/blog/mcp-current-state">这篇文章</a>中了解更多关于 MCP 现状的信息。</p><p>MCP 的发展<a href="https://www.elastic.co/search-labs/blog/mcp-current-state#mcp-project-updates:-transport,-elicitation,-and-structured-tooling">每天都在变化</a>，服务器的使用范围越来越广。此外，构建自定义 MCP 服务器也非常简单，我们将在本文中进行演示。</p><h3>MCP 客户端</h3><p><a href="https://modelcontextprotocol.io/clients">可用的 MCP 客户端</a>由很多，每个客户端都有自己的特点和局限性。为了简化和普及，我们将使用 <a href="https://claude.ai/download">Claude Desktop</a> 作为演示中的 MCP 客户端。它将作为聊天界面，用户可以用自然语言提问，它还将自动调用我们的 MCP 服务器提供的工具来搜索文档和生成摘要。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt06fd7a02042094e1/6a170c1b14b2700024e3c651/66eb0b11473347b6cf2d85718251eeac38d6249d-1999x1491.png" alt="Claude 4.5 十四行诗页面，附有“到喝咖啡和用 Claude 时间了？今天我能为您做什么？”" /><h2>创建 Elasticsearch MCP 服务器</h2><p>通过使用 <a href="https://github.com/modelcontextprotocol/typescript-sdk">TypeScript 软件开发工具包</a>，我们可以轻松创建一个能够根据用户查询输入来查询 Elasticsearch 数据的服务器。</p><p>本文将介绍将 Elasticsearch MCP 服务器与 Claude Desktop 客户端集成的步骤：</p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#configure-mcp-server-for-elasticsearch">为 Elasticsearch 配置 MCP 服务器。</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#load-the-mcp-server-into-claude-desktop">将 MCP 服务器加载到 Claude Desktop。</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#test-it-out">测试一下。</a></p></li></ol><h3>为 Elasticsearch 配置 MCP 服务器。</h3><p>首先，我们来初始化一个 Node 应用程序：</p>npm init -y<p>这将会创建一个 <code>package.json</code> 文件，有了它，我们就可以开始安装该应用程序所需的依赖项。</p>npm install @elastic/elasticsearch @modelcontextprotocol/sdk openai zod &amp;&amp; npm install --save-dev ts-node @types/node typescript<ul><li><p><strong>@elastic/elasticsearch</strong> 将使我们能够访问 Elasticsearch Node.js 库。</p></li><li><p><strong>@modelcontextprotocol/sdk</strong> 提供核心工具来创建和管理 MCP 服务器、注册工具以及处理与 MCP 客户端的通信。</p></li><li><p><strong>openai</strong> 允许与 OpenAI 模型进行交互以生成摘要或自然语言响应。</p></li><li><p><a href="https://zod.dev/"><strong>zod</strong></a>帮助定义和验证每个工具中输入和输出数据的结构化模式。</p></li></ul><p><code>ts-node</code>，<code>@types/node</code> 和 <code>typescript</code> 将在开发过程中用于键入代码和编译脚本。</p><h4>配置数据集</h4><p>为了提供 Claude Desktop 可以使用我们的 MCP 服务器进行查询的数据，我们将使用模拟的<a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/dataset.json">内部知识库数据集</a>。来自该数据集的文档是这样子的：</p>{
    "id": 5,
    "title": "Logging Standards for Microservices",
    "content": "Consistent logging across microservices helps with debugging and tracing. Use structured JSON logs and include request IDs and timestamps. Avoid logging sensitive information. Centralize logs in Elasticsearch or a similar system. Configure log rotation to prevent storage issues and ensure logs are searchable for at least 30 days.",
    "tags": ["logging", "microservices", "standards"]
}<p>为了摄取数据，我们准备了一个脚本，该脚本在 Elasticsearch 中创建一个索引并将数据集加载到其中。您可以<a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/setup.ts">在这里</a>找到它。</p><h4>MCP 服务器</h4><p>创建一个名为 <a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/index.ts"><code>index.ts</code></a> 的文件，并添加以下代码来导入依赖项并处理环境变量：</p>// index.ts
import { z } from "zod";
import { Client } from "@elastic/elasticsearch";
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import OpenAI from "openai";

const ELASTICSEARCH_ENDPOINT =
  process.env.ELASTICSEARCH_ENDPOINT ?? "http://localhost:9200";
const ELASTICSEARCH_API_KEY = process.env.ELASTICSEARCH_API_KEY ?? "";
const OPENAI_API_KEY = process.env.OPENAI_API_KEY ?? "";
const INDEX = "documents";<p>此外，让我们初始化客户端以处理 Elasticsearch 和 OpenAI 的调用：</p>const openai = new OpenAI({
  apiKey: OPENAI_API_KEY,
});

const _client = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: {
    apiKey: ELASTICSEARCH_API_KEY,
  },
});<p>为了使我们的实现更加稳健，并确保输入和输出结构化，我们将使用 <a href="https://zod.dev/"><code>zod</code></a> 定义模式。这使我们能够在运行时验证数据，及早发现错误，并使工具响应更容易以编程方式进行处理：</p>const DocumentSchema = z.object({
  id: z.number(),
  title: z.string(),
  content: z.string(),
  tags: z.array(z.string()),
});

const SearchResultSchema = z.object({
  id: z.number(),
  title: z.string(),
  content: z.string(),
  tags: z.array(z.string()),
  score: z.number(),
});

type Document = z.infer&lt;typeof DocumentSchema&gt;;
type SearchResult = z.infer&lt;typeof SearchResultSchema&gt;;<p>请在<a href="https://www.elastic.co/search-labs/blog/structured-outputs-elasticsearch-guide">此处</a>了解更多关于结构化输出的信息。</p><p>现在让我们初始化 MCP 服务器：</p>const server = new McpServer({
  name: "Elasticsearch RAG MCP",
  description:
    "A RAG server using Elasticsearch. Provides tools for document search, result summarization, and source citation.",
  version: "1.0.0",
});<h4>定义 MCP 工具</h4><p>完成所有配置后，我们就可以开始编写将由 MCP 服务器公开的工具了。此服务器公开两种工具：</p><ul><li><p><strong><code>search_docs</code></strong><strong>：</strong>使用全文本搜索在 Elasticsearch 中搜索文档。</p></li><li><p><strong><code>summarize_and_cite</code></strong><strong>：</strong>汇总和综合先前检索到的文档中的信息，以回答用户的问题。该工具还可添加引用源文档的引文。</p></li></ul><p>这两个工具共同构成了一个简单的“检索后总结”工作流，其中一个工具获取相关文档，另一个工具使用这些文档生成汇总的引用回复。</p><h4>工具响应格式</h4><p>每个工具都可以接受任意输入参数，但必须以以下结构作出响应：</p><ul><li><p><strong>内容：</strong>这是工具以非结构化格式做出的响应。该字段通常用于返回文本、图像、音频、链接或嵌入内容。在本应用程序中，它将用于返回包含工具生成的信息的格式化文本。</p></li><li><p><strong>结构化内容： </strong>这是一个可选返回，用于以结构化格式提供每个工具的结果。这对程序化用途非常有用。虽然本 MCP 服务器没有使用它，但如果您想开发其他工具或以编程方式处理结果，它可能会很有用。</p></li></ul><p>基于这个结构，让我们详细探讨每个工具。</p><h4>Search_docs 工具</h4><p>此工具在 Elasticsearch 索引中执行 <a href="https://www.elastic.co/docs/solutions/search/full-text">全文本搜索</a>，以根据用户查询检索最相关的文档。它突出显示关键匹配项，并快速提供相关性评分概述。</p>server.registerTool(
  "search_docs",
  {
    title: "Search Documents",
    description:
      "Search for documents in Elasticsearch using full-text search. Returns the most relevant documents with their content, title, tags, and relevance score.",
    inputSchema: {
      query: z
        .string()
        .describe("The search query terms to find relevant documents"),
      max_results: z
        .number()
        .optional()
        .default(5)
        .describe("Maximum number of results to return"),
    },
    outputSchema: {
      results: z.array(SearchResultSchema),
      total: z.number(),
    },
  },
  async ({ query, max_results }) =&gt; {
    if (!query) {
      return {
        content: [
          {
            type: "text",
            text: "Query parameter is required",
          },
        ],
        isError: true,
      };
    }

    try {
      const response = await _client.search({
        index: INDEX,
        size: max_results,
        query: {
          bool: {
            must: [
              {
                multi_match: {
                  query: query,
                  fields: ["title^2", "content", "tags"],
                  fuzziness: "AUTO",
                },
              },
            ],
            should: [
              {
                match_phrase: {
                  title: {
                    query: query,
                    boost: 2,
                  },
                },
              },
            ],
          },
        },
        highlight: {
          fields: {
            title: {},
            content: {},
          },
        },
      });

      const results: SearchResult[] = response.hits.hits.map((hit: any) =&gt; {
        const source = hit._source as Document;

        return {
          id: source.id,
          title: source.title,
          content: source.content,
          tags: source.tags,
          score: hit._score ?? 0,
        };
      });

      const contentText = results
        .map(
          (r, i) =&gt;
            `[${i + 1}] ${r.title} (score: ${r.score.toFixed(
              2,
            )})\n${r.content.substring(0, 200)}...`,
        )
        .join("\n\n");

      const totalHits =
        typeof response.hits.total === "number"
          ? response.hits.total
          : (response.hits.total?.value ?? 0);

      return {
        content: [
          {
            type: "text",
            text: `Found ${results.length} relevant documents:\n\n${contentText}`,
          },
        ],
        structuredContent: {
          results: results,
          total: totalHits,
        },
      };
    } catch (error: any) {
      console.log("Error during search:", error);

      return {
        content: [
          {
            type: "text",
            text: `Error searching documents: ${error.message}`,
          },
        ],
        isError: true,
      };
    }
  }
);<p><em>我们将 fuzziness : “AUTO” 配置</em><em>为根据被分析的词元的长度具有可变的拼写错误容忍度。我们还设置了</em> <em><code>title^2</code></em> <em>来提高标题字段匹配的文档的分数。</em></p><h4>摘要和引用工具</h4><p>该工具根据上一次搜索中检索到的文档生成摘要。它使用 OpenAI 的 <code>gpt-4o-mini</code> 模型来综合最相关的信息，提供直接来自搜索结果的响应，以回答用户的问题。除了摘要之外，它还返回所使用源文档的引用元数据。</p>server.registerTool(
  "summarize_and_cite",
  {
    title: "Summarize and Cite",
    description:
      "Summarize the provided search results to answer a question and return citation metadata for the sources used.",
    inputSchema: {
      results: z
        .array(SearchResultSchema)
        .describe("Array of search results from search_docs"),
      question: z.string().describe("The question to answer"),
      max_length: z
        .number()
        .optional()
        .default(500)
        .describe("Maximum length of the summary in characters"),
      max_docs: z
        .number()
        .optional()
        .default(5)
        .describe("Maximum number of documents to include in the context"),
    },
    outputSchema: {
      summary: z.string(),
      sources_used: z.number(),
      citations: z.array(
        z.object({
          id: z.number(),
          title: z.string(),
          tags: z.array(z.string()),
          relevance_score: z.number(),
        })
      ),
    },
  },
  async ({ results, question, max_length, max_docs }) =&gt; {
    if (!results || results.length === 0 || !question) {
      return {
        content: [
          {
            type: "text",
            text: "Both results and question parameters are required, and results must not be empty",
          },
        ],
        isError: true,
      };
    }

    try {
      const used = results.slice(0, max_docs);

      const context = used
        .map(
          (r: SearchResult, i: number) =&gt;
            `[Document ${i + 1}: ${r.title}]\\n${r.content}`
        )
        .join("\n\n---\n\n");

      // Generate summary with OpenAI
      const completion = await openai.chat.completions.create({
        model: "gpt-4o-mini",
        messages: [
          {
            role: "system",
            content:
              "You are a helpful assistant that answers questions based on provided documents. Synthesize information from the documents to answer the user's question accurately and concisely. If the documents don't contain relevant information, say so.",
          },
          {
            role: "user",
            content: `Question: ${question}\\n\\nRelevant Documents:\\n${context}`,
          },
        ],
        max_tokens: Math.min(Math.ceil(max_length / 4), 1000),
        temperature: 0.3,
      });

      const summaryText =
        completion.choices[0]?.message?.content ?? "No summary generated.";

      const citations = used.map((r: SearchResult) =&gt; ({
        id: r.id,
        title: r.title,
        tags: r.tags,
        relevance_score: r.score,
      }));

      const citationText = citations
        .map(
          (c: any, i: number) =&gt;
            `[${i + 1}] ID: ${c.id}, Title: "${c.title}", Tags: ${c.tags.join(
              ", ",
            )}, Score: ${c.relevance_score.toFixed(2)}`,
        )
        .join("\n");

      const combinedText = `Summary:\\n\\n${summaryText}\\n\\nSources used (${citations.length}):\\n\\n${citationText}`;

      return {
        content: [
          {
            type: "text",
            text: combinedText,
          },
        ],
        structuredContent: {
          summary: summaryText,
          sources_used: citations.length,
          citations: citations,
        },
      };
    } catch (error: any) {
      return {
        content: [
          {
            type: "text",
            text: `Error generating summary and citations: ${error.message}`,
          },
        ],
        isError: true,
      };
    }
  }
);<p>最后，我们需要用 <a href="https://github.com/modelcontextprotocol/typescript-sdk?tab=readme-ov-file#stdio">stdio</a> 启动服务器。这意味着 MCP 客户端将通过读取和写入其标准输入和输出流与我们的服务器进行通信。stdio 是最简单的传输选项，适用于客户端作为子进程启动的本地 MCP 服务器。在文件末尾添加以下代码：</p>const transport = new StdioServerTransport();
server.connect(transport);<p>现在请您使用以下命令编译该项目：</p>npx tsc index.ts --target ES2022 --module node16 --moduleResolution node16 --outDir ./dist --strict --esModuleInterop<p>这将创建一个 <code>dist</code> 文件夹，并在其中创建一个 <code>index.js</code> 文件。</p><h3>将 MCP 服务器加载到 Claude Desktop。</h3><p>请按照<a href="https://modelcontextprotocol.io/docs/develop/connect-local-servers">本指南</a>配置 MCP 服务器和 Claude Desktop。在 Claude 配置文件中，我们需要设置以下值：</p>{
  "mcpServers": {
    "elasticsearch-rag-mcp": {
      "command": "node",
      "args": [   "/Users/user-name/app-dir/dist/index.js"
      ],
      "env": {
        "ELASTICSEARCH_ENDPOINT": "your-endpoint-here",
        "ELASTICSEARCH_API_KEY": "your-api-key-here",
        "OPENAI_API_KEY": "your-openai-key-here"
      }
    }
  }
}<p><code>args</code> 的值应指向 <code>dist</code> 文件夹中的编译后文件。您还需要在配置文件中设置环境变量，使其名称与代码中定义的名称完全一致。</p><h3>测试一下</h3><p>在执行每个工具前，点击<strong>搜索和工具</strong>，确保这些工具已启用。您还可以在这里启用或禁用每个工具：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt395a7337021f9820/6a170c1c67045bb74d45c228/172981c2a54adabc70d5819013c3007670935605-1999x1002.png" alt="Claude 4.5 十四行诗页面，附有“下午好，Jeff。今天我能为您做什么？”" /><p>最后，让我们从 Claude Desktop 聊天中测试 MCP 服务器，并开始提问：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf4ac458dc0206271/6a170c1e66c4f91328f8c072/03654c0f8c53c714f801fba8b25747071179209b-1999x1353.png" alt="用户在 Claude Desktop 聊天中请求搜索有关身份验证方法和基于角色的访问控制的文档，以及 Claude 的回复。" /><p>针对问题“<strong>搜索有关身份验证方法和基于角色的访问控制的文档</strong>”，已执行 <code>search_docs</code> 工具并返回以下结果：</p>Most Relevant Documents:
Access Control and Role Management (highest relevance) - This document covers role-based access control (RBAC) principles, including ensuring users only have necessary permissions, regular auditing of user roles, revoking inactive accounts, and implementing just-in-time access for sensitive operations.
User Authentication with OAuth 2.0 - This document explains OAuth 2.0 authentication, which enables secure delegated access without credential sharing. It covers configuring identity providers, token management with limited scope and lifetime, and secure storage of refresh tokens.
Container Security Guidelines - While primarily about container security, this document touches on access control aspects like running containers as non-root users and avoiding embedded credentials.
Incident Response Playbook - This mentions role assignment during incidents (incident commander, communications lead, etc.), which relates to access control in emergency scenarios.
Logging Standards for Microservices - This document includes guidance on avoiding logging sensitive information, which is relevant to authentication security.<p>回复是：“太好了！我找到了 5 份关于身份验证方法和基于角色的访问控制的相关文档。以下是找到的内容：”</p><p>该工具调用会将源文档作为其响应有效负载的一部分返回，这些文档随后用于生成引用。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbaf48a9468770ce2/6a170c21509168bffee1bb14/25ff4c7e9563d99752f95540dafdc7fd211a66e3-800x530.gif" alt="Claude 4.5 Sonnet 页面，滚动的响应包括有关身份验证方法和基于角色的访问控制的五个相关文档。" /><p>您也可以在一次交互中串联使用多个工具。在这种情况下，Claude Desktop 会分析用户的问题，并确定首先需要调用 <code>search_docs</code> 来检索相关文档，然后将这些结果传递给 <code>summarize_and_cite</code> 以生成最终回答，所有这些都无需用户单独提示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta46ff45274e64192/6a170c230c4857a91501aac1/e6a8a46acb4236e77058f18bcd2f0737b5882c05-1999x1101.png" alt="Claude Desktop 聊天，附有“Jeff 回来了”，外加一个新的用户问题：“在我们的系统中改进身份验证和访问控制的主要建议是什么？附上参考文献。”" /><p>在这种情况下，对于“<strong>在我们的系统中改进身份验证和访问控制的主要建议是什么？附上参考文献</strong>”，我们得到了以下结果：</p>Based on the documentation, here are the main recommendations to improve authentication and access control across your systems:

Key Recommendations

1. Implement Role-Based Access Control (RBAC)
Ensure users have only the permissions necessary for their job functions. This principle of least privilege minimizes security risks by limiting access to what's truly needed. [1]

2. Regular Access Audits
Conduct regular audits of user roles and promptly revoke access for inactive accounts. This helps prevent unauthorized access from dormant or compromised accounts. [1]

3. Just-in-Time (JIT) Access
Implement just-in-time access for sensitive operations, where permissions are granted only when needed and for a limited duration. This reduces the attack surface by minimizing standing privileges. [1]

4. OAuth 2.0 for Secure Authentication
Use OAuth 2.0 to enable secure delegated access without sharing user credentials. This protocol allows third-party applications to access resources on behalf of users without exposing passwords. [2]

5. Token Security and Management
Configure your identity provider to issue tokens with limited scope and lifetime

Store refresh tokens securely
Validate access tokens consistently to maintain security [2]
References

Access Control and Role Management (Tags: security, access-control)
User Authentication with OAuth 2.0 (Tags: authentication, oauth)
These recommendations work together to create a defense-in-depth approach, where multiple security layers protect your systems from unauthorized access.<p>与上一步一样，我们可以看到每个工具对该问题的响应：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8f633c518e708a99/6a170c25ab7f082991db9ed6/cb606d356b2f7d5e4878a5eff71bc881869ac0ee-800x585.gif" alt="Claude 桌面聊天页面，包含滚动文本，其中包括每个工具对问题“在我们的系统中改进身份验证和访问控制的主要建议是什么？附上参考文献。”的响应。" /><p><em>注意：如果出现子菜单询问是否批准使用每个工具，请选择</em><em><strong>“始终允许”</strong></em><em>或</em><em><strong>“允许一次”</strong></em><em>。</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6627ee0bff1862df/6a170c266f7f040f6f91488c/aea942ba9b0037526ea215bec65690f1a5c3099c-1522x250.png" alt="Claude Desktop 的 &quot;始终允许&quot; 和 &quot;允许一次&quot; 选项，供用户选择。" /><h2>结论</h2><p>MCP 服务器代表了本地和远程应用中 LLM 工具标准化的重要一步。虽然完全兼容仍在开发中，但我们正朝这个方向快速推进。</p><p>在本文中，我们学习了如何用 TypeScript 构建一个自定义 MCP 服务器，将 Elasticsearch 连接到基于 LLM 的应用。我们的服务器公开了两个工具：<code>search_docs</code> 用于使用查询 DSL 检索相关文档；<code>summarize_and_cite</code> 用于通过 OpenAI 模型和 Claude Desktop 作为客户端 UI 生成带引用的摘要。</p><p>不同客户端和服务器提供商之间的兼容性前景看起来一片光明。下一步包括为您的智能体添加更多功能和灵活性。这里有一篇实用的<a href="https://www.elastic.co/search-labs/blog/llm-functions-elasticsearch-intelligent-query">文章</a>介绍了如何使用搜索模板参数化查询，以获得精确性和灵活性。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude</guid>
    <category><![CDATA[智能体 AI]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5600198cb47666a5/6a170c28509168ce3ae1bb18/0bb24c05fff391f42070c2883182ea6fe9cb9680-1280x720.png" length="0" type="image/png"/>
    <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 Elasticsearch 推理 API 以及 Hugging Face 模型]]></title>
    <description><![CDATA[了解如何使用推理终端将 Elasticsearch 连接到 Hugging Face 模型，并利用语义搜索和聊天补全功能构建多语言博客推荐系统。]]></description>
    <content:encoded><![CDATA[<p>在最近的更新中，Elasticsearch 引入了原生集成，用于连接到托管在 <a href="https://endpoints.huggingface.co/">Hugging Face Inference Service</a> 上的模型。在本文中，我们将探讨如何配置此集成，并使用大型语言模型 (LLM) 通过简单的 API 调用执行推理。我们将使用 <a href="https://huggingface.co/HuggingFaceTB/SmolLM3-3B">SmolLM3-3B</a>，这是一款轻量级通用模型，在资源使用和答案质量之间取得了良好的平衡。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9094997548bd70f8/6a170d6a839dfa0ad6dcff54/7ddadf1976421a860a7d62087239adb9150d808b-1999x1388.png" alt="散点图显示了几个小型语言模型，其 X 轴表示模型大小（以十亿为单位的参数），Y 轴表示胜率（百分比）。SmolLM3-3B 在效率趋势中名列前茅，其胜率高于其他类似大小的模型。" /><h2>准备工作</h2><ul><li><p><strong>Elasticsearch 9.3 或 Elastic Cloud Serverless：</strong>您可以按照<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">这些说明</a>创建云部署，或者改用 <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart#local-dev-quick-start"><code>start-local</code></a> 快速入门。</p></li><li><p><strong>Python 3.12：</strong><a href="https://www.python.org/">在此处</a>下载 Python。</p></li><li><p><strong>Hugging Face </strong><a href="https://huggingface.co/docs/hub/en/security-tokens">访问令牌</a>。</p></li></ul><h2>使用 Hugging Face 推理终端完成聊天</h2><p>首先，我们将构建一个实用示例，将 Elasticsearch 连接到 Hugging Face <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put">推理终端</a>，以从博客文章集合中生成 AI 驱动的推荐。对于应用知识库，我们将使用公司博客文章数据集，其中包含有价值但通常难以查找的信息。</p><p>通过这个终端，<a href="https://www.elastic.co/docs/solutions/search/semantic-search">语义搜索</a>可以检索与给定查询最相关的文章，而 Hugging Face LLM 则会根据这些结果生成简短的上下文推荐。</p><p>让我们来看看我们将要构建的信息流的高级概述：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf217b7b7db4e1e6c/6a170d6ca929cf8022ae0a3b/1dfbc2323438feaaa42e13ab242dd1f7166f74aa-1200x676.png" alt="流程图展示了 Elasticsearch 索引将语义搜索结果输入到推理终端，该终端会返回文章推荐。" /><p>在本文中，我们将测试 <strong>SmolLM3-3B</strong> 是否能将其紧凑的大小与强大的多语言推理和工具调用能力相结合。根据搜索查询，我们将把所有匹配的内容（英语和西班牙语）发送到 LLM，以生成一份推荐文章列表，并根据搜索查询和结果提供自定义描述。</p><p>以下是具备 AI 推荐生成系统的文章网站用户界面可能的外观。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt20e69b9a06fecd65/6a170d6e839dfa6f97dcff58/8d3b86b212f28ff279f2da67a33e6134039f0e4e-1999x949.png" alt="具备 AI 推荐生成系统的文章网站的用户界面，列出了三个示例，文本为英语，标题为英语或西班牙语。" /><p>您可以在已链接的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/notebook.ipynb">笔记本</a>中找到此应用程序的完整实现。</p><h3>配置 Elasticsearch 推理终端</h3><p>要使用 Elasticsearch <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put-hugging-face">Hugging Face 推理终端</a>，我们需要两个重要元素：Hugging Face API 密钥和正在运行的 Hugging Face 终端 URL。它应该如下所示：</p>PUT _inference/chat_completions/hugging-face-smollm3-3b
{
    "service": "hugging_face",
    "service_settings": {
        "api_key": "hugging-face-access-token", 
        "url": "url-endpoint" 
    }
}<p>Elasticsearch 中的 Hugging Face 推理终端支持不同的任务类型：<code>text_embedding</code>、<code>completion</code>、<code>chat_completion</code> 和 <code>rerank</code>。在这篇博客文章中，我们使用 <code>chat_completion</code> 是因为我们需要模型根据搜索结果和系统提示生成对话式推荐。此终端允许我们使用 Elasticsearch API 以简单的方式直接从 Elasticsearch 执行聊天完成：</p>POST _inference/chat_completion/hugging-face-smollm3-3b/_stream
{
  "messages": [
      { "role": "user", "content": "&lt;user prompt&gt;" }
  ]
}<p>这将作为应用程序的核心，接收通过模型传递的提示和搜索结果。有了理论基础，我们就开始实施应用程序。</p><h4>在 Hugging Face 上设置推理终端</h4><p>要部署 Hugging Face 模型，我们将使用 <a href="https://huggingface.co/inference-endpoints/dedicated">Hugging Face 一键式部署</a>，这是一种用于部署模型终端的简单快速的服务。请记住，这是一项付费服务，使用它可能会产生额外费用。此步骤将创建用于生成文章推荐的模型实例。</p><p>您可以从一键目录中选择一个模型：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta7bdfa43d6766324/6a170d6fb339d59e5476a039/b816e9fba1fe172687bf58f5143fb1f838c1077f-549x331.png" alt="接口视图显示了一个已筛选为“smoll3”的模型目录，其中显示了一个名为“smollm3‑3b”的模型，具有文本生成、vLLM、GPU 1× NVIDIA L4，标价 0.8 美元，并附有一条建议将搜索范围扩展至所有 Hugging Face 模型的提示。" /><p>让我们选择 <strong>SmolLM3-3B</strong> 模型：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdb0a2e6ffd7deb20/6a170d710c48574b7401aafc/610d3aba0429f3666c2df3616d513eb6a4397c0c-502x478.png" alt="用于创建 SmolLM3‑3B 模型终端的接口，显示模型名称、“已由 Hugging Face 验证”注释、终端名称字段、每个运行副本每小时 0.80 美元的成本、cURL 选项和“创建终端”按钮。" /><p>从此处获取 Hugging Face 终端 URL：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt25714021711ed6ff/6a170d72c1e8a54853f88336/025094ddb2cfbd1f0f216a5ec4e119b0f4fa2c42-646x328.png" alt="名为“smollm3‑3b‑pnz”的 Hugging Face 推理终端的仪表板视图，显示绿色运行状态、一个活跃副本、过去一小时内的零请求、导航选项卡和显示的终端 URL。" /><p>正如在 Elasticsearch <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put-hugging-face">Hugging Face 推理终端文档</a>中提到的，文本生成需要一个与 OpenAI API 兼容的模型。因此，我们需要将 <code>/v1/chat/completions</code> 子路径附加到 Hugging Face 终端 URL。最终结果将如下所示：</p>https://j2g31h0futopfkli.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions<p>有了这个，我们就可以在 Python 笔记本中开始编码了。</p><h4>生成 Hugging Face API 密钥</h4><p>创建 <a href="https://huggingface.co/join">Hugging Face 账户</a>，并按照<a href="https://huggingface.co/docs/hub/en/security-tokens#user-access-tokens">以下说明</a>获取 API 令牌。您可以选择三种令牌类型：<em>细粒度</em>（推荐用于生产，因为它仅提供对特定资源的访问）、<em>读取</em>（适用于只读访问）或<em>写入</em>（适用于读取和写入访问）。在本教程中，读取令牌就足够了，因为我们只需要调用推理终端。请保存此密钥以备下一步使用。</p><h4>设置 Elasticsearch 推理终端</h4><p>首先，让我们声明一个 Elasticsearch Python 客户端：</p>os.environ["ELASTICSEARCH_API_KEY"] = "your-elasticsearch-api-key"
os.environ["ELASTICSEARCH_URL"] = "https://xxxx.us-central1.gcp.cloud.es.io:443"

es_client = Elasticsearch(
    os.environ["ELASTICSEARCH_URL"], api_key=os.environ["ELASTICSEARCH_API_KEY"]
)<p>接下来，我们创建一个使用 Hugging Face 模型的 Elasticsearch 推理终端。此终端将允许我们基于博客文章和传递给模型的提示来生成响应。</p>INFERENCE_ENDPOINT_ID = "smollm3-3b-pnz"

os.environ["HUGGING_FACE_INFERENCE_ENDPOINT_URL"] = (
 "https://j2g31h0futopfkli.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions"
)
os.environ["HUGGING_FACE_API_KEY"] = "hf_xxxxx"

resp = es_client.inference.put(
        task_type="chat_completion",
        inference_id=INFERENCE_ENDPOINT_ID,
        body={
            "service": "hugging_face",
            "service_settings": {
                "api_key": os.environ["HUGGING_FACE_API_KEY"],
                "url": os.environ["HUGGING_FACE_INFERENCE_ENDPOINT_URL"],
            },
        },
    )<h3>数据集</h3><p>该数据集包含将要查询的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/dataset.json">博客文章</a>，代表整个工作流中使用的多语言内容集：</p>// Articles dataset document example: 
{
    "id": "6",
    "title": "Complete guide to the new API: Endpoints and examples",
    "author": "Tomas Hernandez",
    "date": "2025-11-06",
    "category": "tutorial",
    "content": "This guide describes in detail all endpoints of the new API v2. It includes code examples in Python, JavaScript, and cURL for each endpoint. We cover authentication, resource creation, queries, updates, and deletion. We also explain error handling, rate limiting, and best practices. Complete documentation is available on our developer portal."
  }<h4>Elasticsearch 映射</h4><p>定义数据集后，我们需要创建一个适合博客文章结构的数据模式。以下<a href="https://www.elastic.co/docs/manage-data/data-store/mapping">索引映射</a>将用于在 Elasticsearch 中存储数据：</p>INDEX_NAME = "blog-posts"

mapping = {
    "mappings": {
        "properties": {
            "id": {"type": "keyword"},
            "title": {
                "type": "object",
                "properties": {
                    "original": {
                        "type": "text",
                        "copy_to": "semantic_field",
                        "fields": {"keyword": {"type": "keyword"}},
                    },
                    "translated_title": {
                        "type": "text",
                        "fields": {"keyword": {"type": "keyword"}},
                    },
                },
            },
            "author": {"type": "keyword", "copy_to": "semantic_field"},
            "category": {"type": "keyword", "copy_to": "semantic_field"},
            "content": {"type": "text", "copy_to": "semantic_field"},
            "date": {"type": "date"},
            "semantic_field": {"type": "semantic_text"},
        }
    }
}


es_client.indices.create(index=INDEX_NAME, body=mapping)<p>在这里，我们可以更清楚地看到数据的结构。我们将使用语义搜索来检索基于自然语言的结果，同时使用 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/copy-to"><code>copy_to</code></a> 属性将字段内容复制到 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_text</code></a> 字段中。此外，<code>title</code> 字段包含两个子字段：<code>original</code> 子字段根据文章的原始语言存储英语或西班牙语标题；而 <code>translated_title</code> 子字段仅存在于西班牙语文章中，并包含原始标题的英语翻译。</p><h3>采集数据</h3><p>以下代码片段使用<a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript/bulk_examples">批量 API</a> 将博客文章数据集摄取到 Elasticsearch 中：</p>def build_data(json_file, index_name):
    with open(json_file, "r") as f:
        data = json.load(f)

    for doc in data:
        action = {"_index": index_name, "_source": doc}
        yield action


try:
    success, failed = helpers.bulk(
        es_client,
        build_data("dataset.json", INDEX_NAME),
    )
    print(f"{success} documents indexed successfully")

    if failed:
        print(f"Errors: {failed}")
except Exception as e:
    print(f"Error: {str(e)}")<p>现在，我们已将文章摄取到 Elasticsearch 中，我们需要创建一个能够针对 <code>semantic_text</code> 字段进行搜索的函数：</p>def perform_semantic_search(query_text, index_name=INDEX_NAME, size=5):
    try:
        query = {
            "query": {
                "match": {
                    "semantic_field": {
                        "query": query_text,
                    }
                }
            },
            "size": size,
        }

        response = es_client.search(index=index_name, body=query)
        hits = response["hits"]["hits"]

        return hits
    except Exception as e:
        print(f"Semantic search error: {str(e)}")
        return []<p>我们还需要一个调用推理终端的函数。在这种情况下，我们将使用 <strong><code>chat_completion</code></strong>任务类型调用终端，以获取流式响应：</p>def stream_chat_completion(messages: list, inference_id: str = INFERENCE_ENDPOINT_ID):
    url = f"{ELASTICSEARCH_URL}/_inference/chat_completion/{inference_id}/_stream"
    payload = {"messages": messages}
    headers = {
        "Authorization": f"ApiKey {ELASTICSEARCH_API_KEY}",
        "Content-Type": "application/json",
    }

    try:
        response = requests.post(url, json=payload, headers=headers, stream=True)
        response.raise_for_status()

        for line in response.iter_lines(decode_unicode=True):
            if line:
                line = line.strip()

                if line.startswith("event:"):
                    continue

                if line.startswith("data: "):
                    data_content = line[6:]

                    if not data_content.strip() or data_content.strip() == "[DONE]":
                        continue

                    try:
                        chunk_data = json.loads(data_content)

                        if "choices" in chunk_data and len(chunk_data["choices"]) &gt; 0:
                            choice = chunk_data["choices"][0]
                            if "delta" in choice and "content" in choice["delta"]:
                                content = choice["delta"]["content"]
                                if content:
                                    yield content

                    except json.JSONDecodeError as json_err:
                        print(f"\nJSON decode error: {json_err}")
                        print(f"Problematic data: {data_content}")
                        continue

    except requests.exceptions.RequestException as e:
        yield f"Error: {str(e)}"<p>现在，我们可以编写一个函数，调用语义搜索函数以及 <code>chat_completions</code> 推理终端和建议终端，以生成将分配到卡片中的数据：</p>def recommend_articles(search_query, index_name=INDEX_NAME, max_articles=5):
    print(f"\n{'='*80}")
    print(f"🔍 Search Query: {search_query}")
    print(f"{'='*80}\n")

    articles = perform_semantic_search(search_query, index_name, size=max_articles)

    if not articles:
        print("❌ No relevant articles found.")
        return None, None

    print(f"✅ Found {len(articles)} relevant articles\n")

    # Build context with found articles
    context = "Available blog articles:\n\n"
    for i, article in enumerate(articles, 1):
        source = article.get("_source", article)
        context += f"Article {i}:\n"
        context += f"- Title: {source.get('title', 'N/A')}\n"
        context += f"- Author: {source.get('author', 'N/A')}\n"
        context += f"- Category: {source.get('category', 'N/A')}\n"
        context += f"- Date: {source.get('date', 'N/A')}\n"
        context += f"- Content: {source.get('content', 'N/A')}\n\n"

    system_prompt = """You are an expert content curator that recommends blog articles.

    Write recommendations in a conversational style starting with phrases like:
    - "If you're interested in [topic], this article..."
    - "This post complements your search with..."
    - "For those looking into [topic], this article provides..."


    FORMAT REQUIREMENTS:
    - Return ONLY a JSON array
    - Each element must have EXACTLY these three fields: "article_number", "title", "recommendation"
    - If the original title is in spanish, use the "translated_title" subfield in the "title" field

    Keep each recommendation concise (2-3 sentences max) and focused on VALUE to the reader.

    EXAMPLE OF CORRECT FORMAT:
    [
        {"article_number": 1, "title": "Article title in english", "recommendation": "If you are interested in [topic], this article provides..."},
        {"article_number": 2, "title": "Article title in english", "recommendation": " for those looking into [topic], this article provides..."}
    ]

    Return ONLY the JSON array following this exact structure."""

    user_prompt = f"""Search query: "{search_query}"

    Generate recommendations for the following articles: {context}
    """

    messages = [
        {"role": "system", "content": "/no_think"},
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_prompt},
    ]

    # LLM generation
    print(f"{'='*80}")
    print("🤖 Generating personalized recommendations...\n")

    full_response = ""

    for chunk in stream_chat_completion(messages):
        print(chunk, end="", flush=True)
        full_response += chunk

    return context, articles, full_response<p>最后，我们需要提取信息并将其格式化以便打印：</p>def display_recommendation_cards(articles, recommendations_text):
    print("\n" + "=" * 100)
    print("📇 RECOMMENDED ARTICLES".center(100))
    print("=" * 100 + "\n")

    # Parse JSON recommendations - clean tags and extract JSON
    recommendations_list = []
    try:

        # Clean up &lt;think&gt; tags
        cleaned_text = re.sub(
            r"&lt;think&gt;.*?&lt;/think&gt;", "", recommendations_text, flags=re.DOTALL
        )
        # Remove markdown code blocks ( ... ``` or ``` ... ```)
        cleaned_text = re.sub(r"```(?:json)?", "", cleaned_text)
        cleaned_text = cleaned_text.strip()

        parsed = json.loads(cleaned_text)

        # Extract recommendations from list format
        for item in parsed:
            article_number = item.get("article_number")
            title = item.get("title", "")
            rec_text = item.get("recommendation", "")

            if article_number and rec_text:
                recommendations_list.append(
                    {
                        "article_number": article_number,
                        "title": title,
                        "recommendation": rec_text,
                    }
                )
    except json.JSONDecodeError as e:
        print(f"⚠️  Could not parse recommendations as JSON: {e}")
        return

    for i, article in enumerate(articles, 1):
        source = article.get("_source", article)

        # Card border
        print("┌" + "─" * 98 + "┐")

        # Find recommendation and title for this article number
        recommendation = None
        title = None
        for rec in recommendations_list:
            if rec.get("article_number") == i:
                recommendation = rec.get("recommendation")
                title = rec.get("title")
                break

        # Print title
        title_lines = textwrap.wrap(f"📌 {title}", width=94)
        for line in title_lines:
            print(f"│  {line}".ljust(99) + "│")

        # Card border
        print("├" + "─" * 98 + "┤")

        # Print recommendation
        if recommendation:
            recommendation_lines = textwrap.wrap(recommendation, width=94)
            for line in recommendation_lines:
                print(f"│  {line}".ljust(99) + "│")

        # Card bottom
        print("└" + "─" * 98 + "┘")<p>让我们通过询问一个有关安全博客文章的问题来测试一下：</p>search_query = "Security and vulnerabilities"

context, articles, recommendations = recommend_articles(search_query)

print("\nElasticsearch context:\n", context)

# Display visual cards
display_recommendation_cards(articles, recommendations)<p>如下所示，我们可以看到工作流在控制台中生成的卡片：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4aa221a08a51aeb3/6a170d7460084be1413c45d6/730d35212594bb3db30447c3ea7e2a92857287b7-1999x1515.png" alt="标题为“推荐文章”的部分显示了五篇方框式文章摘要，包括身份验证系统漏洞、迁移风险、REST API v2 性能和身份验证改进、通知系统更改以及新 API 的完整指南等主题。" /><p>您可以在<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/results.md">此文件</a>中查看全部结果，包括所有点击和 LLM 响应。</p><p>我们正在征集与“安全与漏洞”相关的文章。此问题将用作针对 Elasticsearch 中存储的文档的搜索查询。然后将检索到的结果传递给模型，该模型根据这些结果的内容生成推荐。我们可以看到，该模型出色地生成了引人入胜的短文本，能够激发读者点击的欲望。</p><h2>结论</h2><p>本示例展示了如何将 Elasticsearch 和 Hugging Face 结合起来，为 AI 应用程序创建一个快速高效的集中式系统。由于 Hugging Face 拥有丰富的模型目录，这种方法不仅减少了人工操作，还具有灵活性。通过使用 SmolLM3-3B，我们特别看到了紧凑的多语言模型在与语义搜索搭配使用时，仍能提供有意义的推理和内容生成。这些工具共同为构建智能内容分析和多语言应用程序提供了可扩展且高效的基础。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/hugging-face-elasticsearch-inference-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/hugging-face-elasticsearch-inference-api</guid>
    <category><![CDATA[智能体 AI]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f961af4cb26ec97/6a170d767d8d6790c770e790/1417d6ff033712206c9bd4bcc22074ee3437ce96-1999x1125.png" length="0" type="image/png"/>
    <pubDate>Mon, 23 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[适用 Elasticsearch 的 Gemini CLI 扩展及工具和技能]]></title>
    <description><![CDATA[Elastic 推出了适用于谷歌 Gemini CLI 的扩展，用于在开发人员和智能体工作流中搜索、检索和分析 Elasticsearch 数据。
]]></description>
    <content:encoded><![CDATA[<p>我们很高兴地宣布， Elastic 发布了适用于 Google 的 Gemini CLI 扩展，将 <a href="https://www.elastic.co/elasticsearch">Elasticsearch</a> 和 <a href="https://www.elastic.co/elasticsearch/agent-builder">Elastic Agent Builder</a> 的全部功能直接引入您的 AI 开发工作流。此扩展还提供几种最近开发的智能体技能，用于与 Elasticsearch 交互。</p><p>该扩展以开源项目的形式在<a href="https://github.com/elastic/gemini-cli-elasticsearch">此处</a>提供。</p><h2>Gemini CLI 是什么？如何安装？</h2><p><a href="https://geminicli.com/">Gemini CLI</a> 是一个开源的 AI 智能体，它可将 Google 的 Gemini 模型直接引入命令行。它允许开发人员从终端与 AI 进行交互，以执行诸如生成代码、编辑文件、运行 shell 命令和从网上检索信息等任务。</p><p>与典型的聊天界面不同，Gemini CLI 可与您的本地开发环境集成，这意味着它可以直接在终端内理解项目上下文、修改文件、运行构建或测试，以及自动化工作流。这对于想要在不离开命令行工作流的情况下进行 AI 辅助编码和自动化的开发人员、网站可靠性工程师 (SREs) 和工程师来说非常有用。</p><p>Gemini CLI 可通过多个软件包管理器安装。最常用的方法是通过 npm 安装：</p>npm install -g @google/gemini-cli<p>如要了解其他安装选项，请参阅<a href="https://geminicli.com/docs/get-started/installation/">官方安装页面</a>。</p><p>安装完成后，运行以下命令启动 CLI：</p>gemini<p>您会看到一个屏幕，如图 1 所示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4ff43abe550941b4/6a17072ba29299fc2ad00fa4/6dfcec4a77b3dc83bf0d974417bf2e211abb1f4f-876x468.png" alt="Gemini CLI 的屏幕截图。" /><h2>配置 Elasticsearch</h2><p>我们需要运行一个 Elasticsearch 实例。如要使用模型上下文协议 (MCP) 服务器，您还需要安装 Kibana 9.3+。如要使用下面描述的 Elasticsearch 查询语言 (ES|QL) 技能 (<code>esql</code>)，则不需要 Kibana。</p><p>您可以在 <a href="https://www.elastic.co/cloud">Elastic Cloud</a> 上激活免费试用版，或使用 <a href="https://github.com/elastic/start-local"><code>start-local</code></a> 脚本在本地安装：</p>curl -fsSL https://elastic.co/start-local | sh<p>这将在您的计算机上安装 Elasticsearch 和 Kibana，并生成一个用于配置 Gemini CLI 的 API 密钥。</p><p>API 密钥将显示为上一条命令的输出，并存储在 <strong>.env</strong> 文件中，该文件位于 <strong><code>elastic-start-local</code></strong> 文件夹。</p><p>如果您使用的是本地部署的 Elasticsearch（例如使用 <code>start-local</code>），并且您想将 Elastic Agent Builder 与 MCP 一起使用，那么您还需要连接一个大型语言模型 (LLM)。您可以阅读<a href="https://www.elastic.co/docs/explore-analyze/ai-features/llm-guides/llm-connectors">此文档页面</a>以了解不同的选项。</p><p>如果您使用的是 Elastic Cloud（或无服务器架构），那么您已经预先建立了 LLM 连接。</p><h2>安装 Elasticsearch 扩展</h2><p>您可以使用以下命令为 Gemini CLI 安装 Elasticsearch 扩展：</p>gemini extensions install https://github.com/elastic/gemini-cli-elasticsearch<p>您可以通过打开 Gemini 并执行以下命令来检查扩展程序是否已成功安装：</p>/extensions list<p>您应该看到 Elasticsearch 扩展可用。</p><p>如要使用 MCP 集成，您需要安装 Elasticsearch 9.3 或更高版本。您需要从 <a href="https://www.elastic.co/kibana">Kibana</a> 获取您的 MCP 服务器 URL：</p><ul><li><p>从智能体处获取 MCP 服务器 URL &gt; 查看所有工具 &gt; 管理 MCP &gt; 复制 MCP 服务器 URL。</p></li><li><p>URL 将如下所示：https://your-kibana-instance/api/agent_builder/mcp</p></li></ul><p>您需要 Elasticsearch 终端 URL。这通常显示在 Kibana Elasticsearch 页面的顶部。如果您使用 <code>start-local</code> 运行 Elasticsearch，那么您已经在 <code>start-local</code>.env 文件的<code>ES_LOCAL_URL</code> 密钥中拥有了终端。</p><p>您还需要一个 API 密钥。如果您使用 <code>start-local</code> 运行 Elasticsearch，那么您已经在 <code>start-local</code> .env 文件中拥有了 <code>ES_LOCAL_API_KEY</code>。否则，您可以使用 Kibana 界面创建 API 密钥，详见<a href="https://www.elastic.co/docs/deploy-manage/api-keys/elasticsearch-api-keys">此处</a>：</p><ul><li><p>在 Kibana 中：Stack Management &gt; Security &gt; API 密钥 &gt; 创建 API 密钥。</p></li><li><p>我们建议仅设置 API 密钥的读取权限，并启用 <code>feature_agentBuilder.read</code> 权限，详见<a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/permissions#grant-access-with-roles">此处</a>。</p></li><li><p>复制已编码的 API 密钥值。</p></li></ul><p>在您的 shell 中设置所需的环境变量：</p>export ELASTIC_URL="your-elasticsearch-url"
export ELASTIC_MCP_URL="your-elasticsearch-mcp-url"
export ELASTIC_API_KEY="your-encoded-api-key"<h2>安装示例数据集</h2><p>您可以安装 Kibana 提供的<strong>电子商务订单</strong>数据集。它包含一个名为 <strong><code>kibana_sample_data_ecommerce</code></strong> 的单个索引，其中包含来自一家电子商务网站的 4675 个订单的信息。对于每笔订单，我们都有以下信息：</p><ul><li><p>客户信息（姓名、ID 号码、出生日期、电子邮件等）。</p></li><li><p>订单日期。</p></li><li><p>订单编号。</p></li><li><p>产品（包含价格、数量、ID、类别、折扣和其他详情的所有产品列表）</p></li><li><p>SKU。</p></li><li><p>总价（不含税，含税）。</p></li><li><p>总数量。</p></li><li><p>地理信息（城市、国家、洲、位置、地区）。</p></li></ul><p>如要安装示例数据，请在 Kibana 中打开<strong>集成</strong>页面（在顶部搜索栏中搜索“集成”），然后安装<strong>示例数据</strong>。更多详情请参阅<a href="https://www.elastic.co/docs/explore-analyze/#gs-get-data-into-kibana">此处</a>的文档。</p><p>本文旨在展示如何轻松配置 Gemini CLI 以连接到 Elasticsearch 并与 <strong><code>kibana_sample_data_ecommerce</code></strong> 索引交互。</p><h2>如何使用 Elasticsearch MCP（模型上下文协议）</h2><p>您可以在 Gemini 中使用以下命令检查连接：</p>/mcp list<p>您应该会看到 <strong><code>elastic-agent-builder</code></strong> 已启用，如图 2 所示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt52b85e7255360f3b/6a17072da929cf33d3ae08f5/1508423bc1d1bc3c04a1cb01e2d59495a3516ed1-1465x844.png" alt="`elastic-agent-builder` MCP 服务器及其工具列表。" /><p>Elasticsearch 提供了一组默认工具。请参阅<a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools/builtin-tools-reference">此处</a>的描述。</p><p>使用这些工具，您可以与 Elasticsearch 进行交互，提出类似以下的问题：</p><ul><li><p><code>Give me the list of all the indexes available in Elasticsearch.</code></p></li><li><p><code>How many customers are based in the USA in the kibana_sample_data_ecommerce index of Elasticsearch?</code></p></li></ul><p>根据问题的不同，Gemini 会使用一个或多个可用工具来尝试回答问题。</p><h2>/elastic 命令</h2><p>在 Gemini CLI 的 Elasticsearch 扩展中，我们还添加了<strong><code>/elastic</code></strong> 命令。</p><p>如果执行 <strong><code>/help</code></strong> 命令，您将看到所有可用的 <code>/elastic</code> 选项（图 3）：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt741c7451ecab10d2/6a17072ea6c2b9ccd6e79643/5b2a0727ce7a04354878dd048253d3f4d062324b-1983x230.png" alt="可用的 `/elastic` 命令。" /><p>这些命令在您想直接执行 <code>elastic-agent-builder</code> MCP 服务器的特定工具时会很有用。例如，使用以下命令可以获取 <code>kibana_sample_data_ecommerce</code> 的映射：</p>/elastic:get-mapping kibana_sample_data_ecommerce<p>这些命令本质上是执行特定工具的快捷方式，而不是依赖 Gemini 模型来确定应该调用哪个工具。</p><h2>如何使用 Elasticsearch 的技能？</h2><p>该扩展还附带了 ES|QL 的<a href="https://github.com/elastic/gemini-cli-elasticsearch/tree/main/skills/esql">代理技能，ES|QL</a> 是 Elasticsearch 中提供的 <a href="https://www.elastic.co/docs/explore-analyze/discover/try-esql">Elasticsearch 查询语言</a>。<a href="https://agentskills.io/home">Agent Skills</a> 是一种开放格式，为 AI 编码智能体（如 Gemini CLI）提供特定任务的自定义指令。它们使用一种称为<em>渐进式披露</em>的概念，即在系统初始提示中只添加对技能的简要说明。当您要求智能体执行任务时，比如查询 Elasticsearch，它会将请求与相关技能匹配，并动态加载详细说明。这是一种高效管理词元预算的方法，同时为 AI 提供所需的准确上下文。</p><p><strong><code>esql</code></strong><strong>技能</strong>旨在让 Gemini CLI 直接针对集群编写和执行 ES|QL 查询。ES|QL 是一种功能强大的管道化查询语言，能非常直观地进行数据探索、日志分析和聚合。启用该技能后，您无需查找 ES|QL 语法；只需用自然语言向 Gemini CLI 提出有关数据的问题，智能体会处理剩下的问题。</p><p>执行操作是通过在终端中运行简单的 <a href="https://curl.se/">curl</a> 命令来完成的。之所以能做到这一点，是因为 Elasticsearch 提供了一套丰富的 REST API，可轻松用于将系统集成到任何架构中。</p><p><strong> esql </strong><strong> 技能的功能：</strong></p><ul><li><p><strong>发现索引和模式：</strong>智能体可以使用该技能的内置工具列出可用索引并获取字段映射。例如，在为电子商务数据集编写查询之前，智能体可以在 <strong><code>kibana_sample_data_ecommerce</code></strong> 上运行模式检查，以了解可用的字段，如 <strong><code>taxful_total_price</code></strong> 或 <strong><code>category</code></strong>。</p></li><li><p><strong>无缝自然语言翻译：</strong>该技能不仅仅为智能体提供了一个简单的参考手册；它还提供了一个专门的指南，用于解读用户意图。当您用自然语言输入请求（如“按服务分组显示平均响应时间”）时，智能体会使用技能捆绑的模式匹配功能，将您的文字立即转换为正确的 ES|QL 聚合、筛选条件和命令。</p></li><li><p><strong>自我纠正：</strong>如果查询失败（例如由于类型不匹配或语法错误），该技能会返回生成的查询以及确切的 Elasticsearch 错误信息，使智能体能够立即修复查询并再次尝试，而无需您进行干预。</p></li></ul><p>由于 <code>esql</code> 技能也作为工具在 <code>elastic-agent-builder</code> MCP 服务器中可用，我们需要暂时禁用该服务器。您可以使用以下命令禁用它：</p>/mcp disable elastic-agent-builder<p>然后，您只需在 Gemini CLI 中输入类似这样的提示：</p>Find the top 5 product categories by total sales revenue in the kibana_sample_data_ecommerce index<p>代理将：</p><ul><li><p>认识到 <code>esql</code> 技能的必要性。</p></li><li><p>检查 <strong><code>kibana_sample_data_ecommerce</code></strong> 的模式。</p></li><li><p>构建一个 ES|QL 查询，比如：<code>FROM kibana_sample_data_ecommerce | STATS total_revenue = SUM(taxful_total_price) BY category.keyword | SORT total_revenue DESC | LIMIT 5</code>。</p></li><li><p>针对 Elasticsearch API 执行查询。</p></li><li><p>直接在终端中向您展示最终答案。</p></li></ul><p>在这里，我们分享了一个 Gemini 对先前提示的响应示例：</p>───────────────────────────────────────────────────────────
&gt; Find the top 5 product categories by total sales revenue in the kibana_sample_data_ecommerce index   ───────────────────────────────────────────────────────────
✓  Activate Skill "esql": Interact with Elasticsearch using ES|QL and curl. Use when querying, managing indices,  …
Skill esql activated. Resources loaded from /home/user/.gemini/extensions/elasticsearch/skills/esql: 
Showing up to 200 items (files + folders).                                                                     /home/user/.gemini/extensions/elasticsearch/skills/esql/                                                            ├───SKILL.md 
└───references/
├───esql-reference.md 
├───generation-tips.md 
└───query-patterns.md

✦ I will start by exploring the mapping of the kibana_sample_data_ecommerce index to identify the fields for product category and sales revenue. Then, I will construct and execute an ES|QL query to determine the top 5 product categories by total sales revenue.

✓  Shell curl -s "${ELASTIC_URL%/}/kibana_sample_data_ecommerce/_mapping" -H "Authorization: ApiKey $(printenv ELAS…                                                                                                                       {  
   "kibana_sample_data_ecommerce": { 
     "mappings": { 
       "properties": {
         "category": { 
           "type": "text", 
           "fields": {
             "keyword": {
               "type": "keyword"
             }
           }
         }, 
         "currency": {
           "type": "keyword"
         }, 

✦ Okay, I have the mapping. The category is products.category.keyword and sales revenue is taxful_total_price. I'll now execute the ES|QL query to get the top 5 categories by revenue.

✓  Shell curl -s -X POST "${ELASTIC_URL%/}/_query" \  -H "Authorization: ApiKey $(printenv ELASTIC_API_KEY)" \  -H …
{ 
   "took": 62,
   "is_partial": false, 
   "completion_time_in_millis": 1772805644164,
   "documents_found": 4675,
   "values_loaded": 12084, 
   "start_time_in_millis": 1772805644102, 
   "expiration_time_in_millis": 1773237644087,  
   "columns": [ 
     { 
       "name": "total_revenue",    
       "type": "double"  
     }, 
     {   
       "name": "products.category.keyword",

✦ The query for top 5 product categories by revenue was successful. I'll format the returned JSON values into a   user-friendly list.

 ✓  Shell 
                             
 Top 5 product categories by total sales revenue: 
 1. Men's Clothing: 149393.91
 2. Women's Clothing: 135099.91
 3. Women's Shoes: 105479.17
 4. Men's Shoes: 91797.92
 5. Women's Accessories: 60830.31 

✦ I've found the answer; now I'll format it for clarity.

 ✓  Shell 

 Here are the top 5 product categories by total sales revenue:                                                        
                                                                      
 | Rank | Category              | Total Revenue | 
 |------|-----------------------|---------------| 
 | 1    | Men's Clothing        | 149,393.91    |
 | 2    | Women's Clothing      | 135,099.91    |
 | 3    | Women's Shoes         | 105,479.17    |
 | 4    | Men's Shoes           | 91,797.92     | 
 | 5    | Women's Accessories   | 60,830.31     | 

✦ I've formatted and presented the answer. I'm finished.
 ───────────────────────────────────────────────────────────<p>值得注意的是，Gemini 模型如何通过显示其遵循的所有步骤来生成最终响应。在这里，您可以清楚地看到技能对模型推理过程的影响。当模型第一次意识到需要使用某项技能或执行 shell 命令时，它会采用人机协同的方式请求权限。</p><p>通过处理模式发现、查询生成和执行的繁重工作，<code>esql</code> 技能可以让您完全专注于答案，而不是获得答案的机制。您将获得所需的数据，格式正确且直接存储在终端中，无需写一行语法或切换到其他应用。</p><h2>结论</h2><p>在本文中，我们介绍了我们最近发布的适用于 Gemini CLI 的 Elasticsearch 扩展。此扩展让您可以使用 Gemini 和 Elastic Agent Builder 提供的 Elasticsearch MCP 服务器（从 9.3.0 版本开始提供）以及 <code>/elastic</code> 命令与您的 Elasticsearch 实例进行交互。</p><p>此外，该扩展还包含一项 <code>esql</code> 技能，可以将用户的自然语言请求转换为 ES|QL 查询。这种技能在无法使用 MCP 服务器时特别有用，因为底层通信是由在终端中执行的简单 curl 命令驱动的。Elasticsearch 提供了一套丰富的 REST API，可以轻松集成到任何项目中。这在开发智能体 AI 应用时尤为有用。</p><p>有关 Gemini CLI 扩展的更多信息，请访问<a href="https://github.com/elastic/gemini-cli-elasticsearch">此处</a>的项目库。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/gemini-cli-extension-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/gemini-cli-extension-elasticsearch</guid>
    <category><![CDATA[集成]]></category>
    <category><![CDATA[智能体 AI]]></category>
    <dc:creator><![CDATA[Walter Rafelsberger,Enrico Zimuel]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4ff43abe550941b4/6a17072ba29299fc2ad00fa4/6dfcec4a77b3dc83bf0d974417bf2e211abb1f4f-876x468.png" length="0" type="image/png"/>
    <pubDate>Tue, 17 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 Mastra 和 Elasticsearch 构建具有语义召回功能的知识代理]]></title>
    <description><![CDATA[了解如何使用 Mastra 和 Elasticsearch 作为记忆和信息检索的向量存储，构建具有语义调用功能的知识代理。]]></description>
    <content:encoded><![CDATA[<p>在构建可靠的人工智能代理和架构方面，<a href="https://www.elastic.co/search-labs/blog/context-engineering-overview">情境工程</a>正变得越来越重要。随着模型越来越完善，其有效性和可靠性已不再依赖于训练有素的数据，而更多地取决于模型在正确环境中的立足程度。能够在正确的时间检索和应用最相关信息的代理更有可能产生准确和可信的输出结果。</p><p>在本博客中，我们将使用<a href="https://mastra.ai/">Mastra</a>构建一个知识代理，它能记住用户所说的话，并能在稍后调用相关信息，使用 Elasticsearch 作为记忆和检索后端。您可以轻松地将这一概念扩展到现实世界的使用案例中，例如，支持代理可以记住过去的对话和解决方案，使他们能够根据先前的上下文为特定用户定制响应或更快地提供解决方案。</p><p>在这里，您将看到如何一步一步地建造它。如果你迷失了方向，或者只是想运行一个已完成的示例，请<a href="https://github.com/jdarmada/getting-started-mastra-elastic/tree/main">点击此处</a>查看软件仓库。</p><h2>什么是 Mastra？</h2><p>Mastra 是一个开源的 TypeScript 框架，用于构建具有可交换推理、内存和工具部分的人工智能代理。它的<a href="https://mastra.ai/docs/memory/semantic-recall">语义调用</a>功能通过将信息作为嵌入信息存储在向量数据库中，使代理能够记住和检索过去的互动。这样，代理就能保留长期对话的上下文和连续性。Elasticsearch 支持高效的密集矢量搜索，是实现这一功能的绝佳矢量存储工具。当触发语义调用时，代理会将过去的相关信息拉入模型的上下文窗口，使模型能够将检索到的上下文作为其推理和响应的基础。</p><h2>入门必备</h2><ul><li><p>节点 v18+</p></li><li><p>Elasticsearch（8.15 或更新版本）</p></li><li><p>Elasticsearch API 密钥</p></li><li><p><a href="https://help.openai.com/en/articles/4936850-where-do-i-find-my-openai-api-key">OpenAI API 密钥</a></p></li></ul><p>注意：您需要这个是因为演示使用了 OpenAI 提供商，但 Mastra 支持其他人工智能 SDK 和社区模型提供商，因此您可以根据自己的设置轻松更换。</p><h2>构建 Mastra 项目</h2><p>我们将使用 Mastra 内置的 CLI 为我们的项目提供脚手架。运行该命令：</p>npm create mastra@latest<p>您将收到一组提示，首先是</p><p>1.为项目命名。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt87f941f654d03827/6a16f7af67045b214d45bfa1/2b9fe559e0276140dd539e24f916a73c60870405-620x84.png" alt="在 Mastra 应用程序中命名提示符" /><p>2.我们可以保留默认值，也可以不填。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7dbb3d4f27435cac/6a16f7b0cdacbf29497d27de/e04729eb03bce8499e973e18c28642402340d0e5-852x68.png" alt="告诉 mastra 保存提示文件的位置" /><p>3.在本项目中，我们将使用 OpenAI 提供的模型。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1f654f6cb9397e94/6a16f7b2964cea899a08b942/a86596a469a71bdf8bd99cbaf528d0f0cf7272c0-436x222.png" alt="在 Mastra 中选择 OpenAI 提供的模型" /><p>4.选择 "暂时跳过 "选项，因为我们将把所有环境变量存储在一个".env "文件中，稍后再进行配置。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltff106117521a3519/6a16f7b3c1e8a5031af880d8/02b19ccc34af0bdacf52fd94b519d036540ca2e6-426x114.png" alt="为 OpenAI 密钥选择暂时跳过" /><p>5.我们也可以跳过该选项。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcda7d9c51c3878d7/6a16f7b450916809dbe1b892/b3fe63d19d270bc2e0de1dd92033bf8b26750819-990x208.png" alt="" /><p>初始化完成后，我们就可以进入下一步。</p><h3>安装依赖项</h3><p>接下来，我们需要安装一些依赖项：</p>npm install ai @ai-sdk/openai @elastic/elasticsearch dotenv<ul><li><p><code>ai</code> - 核心人工智能 SDK 软件包，提供用于在 JavaScript/TypeScript 中管理人工智能模型、提示和工作流程的工具。Mastra 是在 Vercel 的<a href="https://ai-sdk.dev/">人工智能 SDK</a>基础上构建的，因此我们需要依赖它来实现模型与代理的交互。</p></li><li><p><code>@ai-sdk/openai</code> - 将 AI SDK 连接到 OpenAI 模型（如 GPT-4、GPT-4o 等）的插件，可使用 OpenAI API 密钥进行 API 调用。</p></li><li><p><code>@elastic/elasticsearch</code> -<a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript">Node.js 的官方 Elasticsearch 客户端</a>、用于连接到弹性云或本地集群，以进行索引、搜索和矢量操作。</p></li><li><p><code>dotenv</code> - 从 .env 文件中加载环境变量文件到 process.env 文件中、允许您安全地注入 API 密钥和 Elasticsearch 端点等凭证。</p></li></ul><h3>配置环境变量</h3><p>如果还没有<code>.env</code> 文件，请在项目根目录下创建该文件。或者，你也可以复制并重命名我在<a href="https://github.com/jdarmada/getting-started-mastra-elastic/blob/main/.env.example">软件仓库</a>中提供的<code>.env</code> 示例。在该文件中，我们可以添加以下变量：</p>ELASTICSEARCH_ENDPOINT="your-endpoint-here"
ELASTICSEARCH_API_KEY="your-key-here"
OPENAI_API_KEY="your-key-here"<p>基本设置到此结束。从这里，您就可以开始构建和协调代理。我们将更进一步，添加 Elasticsearch 作为存储和矢量搜索层。</p><h2>添加 Elasticsearch 作为向量存储</h2><p>新建一个名为<code>stores</code> 的文件夹，并在其中添加此<a href="https://github.com/jdarmada/getting-started-mastra-elastic/blob/main/src/mastra/stores/elastic-store.ts">文件</a>。在 Mastra 和 Elastic 正式推出 Elasticsearch 向量存储集成之前，<a href="https://github.com/abhiaiyer91">Abhi Aiyer</a>（Mastra 首席技术官）分享了名为<code>ElasticVector</code> 的早期原型类。简单地说，它将 Mastra 的内存抽象与 Elasticsearch 的密集向量功能连接起来，因此开发人员可以将 Elasticsearch 作为其代理的向量数据库。</p><p>让我们深入了解整合的重要部分：</p><h3>输入 Elasticsearch 客户端</h3><p>本节定义了<code>ElasticVector</code> 类，并设置了 Elasticsearch 客户端连接，同时支持标准部署和无服务器部署。</p>export interface ElasticVectorConfig extends ClientOptions {
    /**
     * Explicitly specify if connecting to Elasticsearch Serverless.
     * If not provided, will be auto-detected on first use.
     */
    isServerless?: boolean;
    
    /**
     * Maximum documents to count accurately when describing indices.
     * Higher values provide accurate counts but may impact performance on large indices.
     * 
     * @default 10000
     */
    maxCountAccuracy?: number;
}

export class ElasticVector extends MastraVector {
    private client: Client;
    private isServerless: boolean | undefined;
    private deploymentChecked: boolean = false;
    private readonly maxCountAccuracy: number;

    constructor(config: ElasticVectorConfig) {
        super();
        this.client = new Client(config);
        this.isServerless = config.isServerless;
        this.maxCountAccuracy = config.maxCountAccuracy ?? 10000;
    }
}<ul><li><p><code>ElasticVectorConfig extends ClientOptions</code>:这将创建一个新的配置接口，继承所有 Elasticsearch 客户端选项（如<code>node</code>,<code>auth</code>,<code>requestTimeout</code> ）并添加我们的自定义属性。这意味着用户可以通过任何有效的 Elasticsearch 配置和我们的无服务器特定选项。</p></li><li><p><code>extends MastraVector</code>:这样，<code>ElasticVector</code> 就可以继承 Mastra 的基础<code>MastraVector</code> 类，这是所有矢量存储集成都要遵守的通用接口。这可以确保从代理的角度来看，Elasticsearch 的行为与其他任何 Mastra 向量后端一样。</p></li><li><p><code>private client: Client</code>:这是一个私有属性，用于保存 Elasticsearch JavaScript 客户端的实例。这样，班级就可以直接与群集对话。</p></li><li><p><code>isServerless</code> 和<code>deploymentChecked</code> ：这些属性共同作用，以检测和缓存我们连接的是无服务器还是标准 Elasticsearch 部署。首次使用时会自动检测，也可以明确配置。</p></li><li><p><code>constructor(config: ClientOptions)</code>:该构造函数接收一个配置对象（包含 Elasticsearch 凭据和可选的无服务器设置），并使用它在<code>this.client = new Client(config)</code> 行中初始化客户端。</p></li><li><p><code>super()</code>:它调用 Mastra 的基本构造函数，因此继承了日志记录、验证助手和其他内部钩子。</p></li></ul><p>此时，Mastra 知道有一个名为 <code>ElasticVector</code></p><h3>检测部署类型</h3><p>在创建索引之前，适配器会自动检测您使用的是标准 Elasticsearch 还是 Elasticsearch Serverless。这一点很重要，因为无服务器部署不允许手动配置分片。</p>private async detectServerless(): Promise&lt;boolean&gt; {
    // Return cached result if already detected
    if (this.deploymentChecked) {
        return this.isServerless ?? false;
    }

    // Use explicit configuration if provided
    if (this.isServerless !== undefined) {
        this.deploymentChecked = true;
        this.logger?.info(
            `Using explicit deployment type: ${this.isServerless ? 'Serverless' : 'Standard'}`
        );
        return this.isServerless;
    }

    try {
        const info = await this.client.info();
        
        // Primary detection: build flavor (most reliable)
        const isBuildFlavorServerless = info.version?.build_flavor === 'serverless';
        
        // Secondary detection: tagline (fallback)
        const isTaglineServerless = info.tagline?.toLowerCase().includes('serverless') ?? false;
        
        this.isServerless = isBuildFlavorServerless || isTaglineServerless;
        this.deploymentChecked = true;
        
        this.logger?.info(
            `Auto-detected ${this.isServerless ? 'Serverless' : 'Standard'} Elasticsearch deployment`,
            { 
                buildFlavor: info.version?.build_flavor, 
                version: info.version?.number,
                detectionMethod: isBuildFlavorServerless ? 'build_flavor' : 'tagline'
            }
        );
        
        return this.isServerless;
    } catch (error) {
        this.logger?.warn(
            'Could not auto-detect deployment type, assuming Standard Elasticsearch. ' +
            'Set isServerless: true explicitly in config if using Serverless.',
            { error: error instanceof Error ? error.message : String(error) }
        );
        this.isServerless = false;
        this.deploymentChecked = true;
        return false;
    }
}<p>发生了什么？</p><ul><li><p>首先检查您是否在配置中明确设置了<code>isServerless</code> （跳过自动检测）。</p></li><li><p>调用 Elasticsearch 的<code>info()</code> API 获取群集信息</p></li><li><p>检查<code>build_flavor field</code> （无服务器部署返回<code>serverless</code>)</p></li><li><p>如果没有 "构建味道"，则退回到检查标语阶段</p></li><li><p>缓存结果，避免重复调用应用程序接口</p></li><li><p>如果检测失败，则默认为标准部署</p></li></ul><p> 使用示例</p>// Option 1: Auto-detect (recommended)
const vector = new ElasticVector({
    node: 'https://your-cluster.es.cloud',
    auth: { apiKey: 'your-api-key' }
});
// Detection happens automatically on first index operation

// Option 2: Explicit configuration (faster startup)
const vector = new ElasticVector({
    node: 'https://your-serverless.es.cloud',
    auth: { apiKey: 'your-api-key' },
    isServerless: true  // Skips auto-detection
});<h3>在 Elasticsearch 中创建 "内存 "存储</h3><p>下面的函数设置了一个 Elasticsearch 索引，用于存储嵌入式内容。它会检查索引是否已经存在。如果没有，它就会用下面的映射创建一个，其中包含一个<code>dense_vector</code> 字段，用于存储嵌入和自定义相似度度量。</p><p>有些事情需要注意：</p><ul><li><p><code>dimension</code> 参数是每个嵌入向量的长度，这取决于你使用的嵌入模型。在我们的例子中，我们将使用 OpenAI 的<code>text-embedding-3-small</code> 模型生成嵌入，该模型输出大小为<code>1536</code> 的向量。我们将以此作为默认值。</p></li><li><p>下面的映射中使用的<code>similarity</code> 变量是由辅助函数 c<code>onst similarity = this.mapMetricToSimilarity(metric)</code> 定义的，该函数接收<code>metric</code> 参数的值，并将其转换为与 Elasticsearch 兼容的关键字，用于所选的距离度量。</p><ul><li><p>例如Mastra 使用<code>cosine</code>,<code>euclidean</code>, 和<code>dotproduct</code> 等一般术语来表示向量相似性。如果我们直接将度量<code>euclidean</code> 传递到 Elasticsearch 映射中，就会出现错误，因为 Elasticsearch 希望关键字<code>l2_norm</code> 代表欧氏距离。</p></li></ul></li><li><p>无服务器兼容性：代码会自动省略无服务器部署的分片和副本设置，因为 Elasticsearch Serverless 会自动管理这些设置。</p></li></ul>async createIndex(params: CreateIndexParams): Promise&lt;void&gt; {
    const { indexName, dimension = 1536, metric = 'cosine' } = params;

    try {
        const exists = await this.client.indices.exists({ index: indexName });

        if (exists) {
            try {
                await this.validateExistingIndex(indexName, dimension, metric);
                this.logger?.info(`Index "${indexName}" already exists and is valid`);
                return;
            } catch (validationError) {
                throw new Error(
                    `Index "${indexName}" exists but does not match the required configuration: ${
                        validationError instanceof Error ? validationError.message : String(validationError)
                    }`
                );
            }
        }

        const isServerless = await this.detectServerless();
        const similarity = this.mapMetricToSimilarity(metric);

        const indexConfig: any = {
            index: indexName,
            mappings: {
                properties: {
                    vector: {
                        type: 'dense_vector',
                        dims: dimension,
                        index: true,
                        similarity: similarity,
                    },
                    metadata: {
                        type: 'object',
                        enabled: true,
                        dynamic: true, // Allows flexible metadata structures
                    },
                },
            },
        };

        // Only configure shards/replicas for non-serverless deployments
        // Serverless manages infrastructure automatically
        if (!isServerless) {
            indexConfig.settings = {
                number_of_shards: 1,
                number_of_replicas: 0, // Increase for production HA deployments
            };
        }

        await this.client.indices.create(indexConfig);

        this.logger?.info(
            `Created ${isServerless ? 'Serverless' : 'Standard'} Elasticsearch index "${indexName}"`,
            { dimension, metric, similarity }
        );
    } catch (error) {
        const errorMessage = error instanceof Error ? error.message : String(error);
        this.logger?.error(`Failed to create index "${indexName}": ${errorMessage}`);
        throw new Error(`Failed to create index "${indexName}": ${errorMessage}`);
    }
}<h3>互动后存储新的记忆或笔记</h3><p>该函数接收每次交互后生成的新嵌入以及元数据，然后使用 Elastic 的<code>bulk</code> API 将其插入或更新到索引中。<code>bulk</code> API 将多个写入操作合并为一个请求；索引性能的提升确保了在代理内存不断增长的情况下，更新仍能保持高效。</p>async upsert(params: UpsertVectorParams): Promise&lt;string[]&gt; {
    const { indexName, vectors, metadata = [], ids } = params;

    try {
        // Generate unique IDs if not provided
        const vectorIds = ids || vectors.map((_, i) =&gt; 
            `vec_${Date.now()}_${i}_${Math.random().toString(36).substr(2, 9)}`
        );

        const operations = vectors.flatMap((vec, index) =&gt; [
            { index: { _index: indexName, _id: vectorIds[index] } },
            {
                vector: vec,
                metadata: metadata[index] || {},
            },
        ]);

        const response = await this.client.bulk({
            refresh: true,
            operations,
        });

        if (response.errors) {
            const erroredItems = response.items.filter((item: any) =&gt; item.index?.error);
            const erroredIds = erroredItems.map((item: any) =&gt; item.index?._id);
            const errorDetails = erroredItems.slice(0, 3).map((item: any) =&gt; ({
                id: item.index?._id,
                error: item.index?.error?.reason || item.index?.error,
                type: item.index?.error?.type
            }));
            
            const errorMessage = `Failed to upsert ${erroredIds.length}/${vectors.length} vectors`;
            console.error(`${errorMessage}. Sample errors:`, JSON.stringify(errorDetails, null, 2));
            this.logger?.error(errorMessage, { 
                failedCount: erroredIds.length, 
                totalCount: vectors.length,
                sampleErrors: errorDetails 
            });
            
            // Still return successfully inserted IDs
            const successfulIds = vectorIds.filter((id, idx) =&gt; 
                !erroredIds.includes(id)
            );
            
            if (successfulIds.length === 0) {
                throw new Error(`${errorMessage}. All operations failed. See logs for details.`);
            }
            
            return successfulIds;
        }

        this.logger?.info(`Successfully upserted ${vectors.length} vectors to "${indexName}"`);
        return vectorIds;
    } catch (error) {
        const errorMessage = error instanceof Error ? error.message : String(error);
        this.logger?.error(`Failed to upsert vectors to "${indexName}": ${errorMessage}`);
        throw new Error(`Failed to upsert vectors to "${indexName}": ${errorMessage}`);
    }
}<h3>查询相似向量以实现语义召回</h3><p>该功能是语义召回功能的核心。代理使用向量搜索，在我们的索引中找到类似的存储嵌入。</p>async query(params: QueryVectorParams&lt;any&gt;): Promise&lt;QueryResult[]&gt; {
    const { indexName, queryVector, topK = 10, filter, includeVector = false } = params;

    try {
        const knnQuery: any = {
            field: 'vector',
            query_vector: queryVector,
            k: topK,
            num_candidates: Math.max(topK * 10, 100), // Search more candidates for better recall
        };

        // Apply metadata filters if provided
        if (filter) {
            knnQuery.filter = this.buildElasticFilter(filter);
        }

        const sourceFields = ['metadata'];
        if (includeVector) {
            sourceFields.push('vector');
        }

        const response = await this.client.search({
            index: indexName,
            knn: knnQuery,
            size: topK,
            _source: sourceFields,
        });

        const results = response.hits.hits.map((hit: any) =&gt; ({
            id: hit._id,
            score: hit._score || 0,
            metadata: hit._source?.metadata || {},
            vector: includeVector ? hit._source?.vector : undefined,
        }));

        this.logger?.debug(`Query returned ${results.length} results from "${indexName}"`);
        return results;
    } catch (error) {
        const errorMessage = error instanceof Error ? error.message : String(error);
        this.logger?.error(`Failed to query vectors from "${indexName}": ${errorMessage}`);
        throw new Error(`Failed to query vectors from "${indexName}": ${errorMessage}`);
    }
}<p>引擎盖下</p><ul><li><p>使用 Elasticsearch 中的<code>knn</code> API 运行<a href="https://www.elastic.co/docs/solutions/search/vector/knn">kNN</a>（k-近邻）查询。</p></li><li><p>检索与输入查询向量最相似的 K 个向量。</p></li><li><p>可选择应用元数据过滤器来缩小搜索结果范围（例如，仅在特定类别或时间范围内进行搜索）</p></li><li><p>返回结构化结果，包括文档 ID、相似性得分和存储的元数据。</p></li></ul><h2>创建知识代理</h2><p>现在，我们已经通过<code>ElasticVector</code> 集成看到了 Mastra 和 Elasticsearch 之间的连接，让我们来创建知识代理本身。</p><p>在<code>agents</code> 文件夹中，创建一个名为<code>knowledge-agent.ts</code> 的文件。我们可以从连接环境变量和初始化 Elasticsearch 客户端开始。</p>import { Agent } from '@mastra/core/agent';
import { Memory } from '@mastra/memory';
import { openai } from '@ai-sdk/openai';
import { Client } from '@elastic/elasticsearch';
import { ElasticVector } from '../stores/elastic-store';
import dotenv from "dotenv";

dotenv.config();

const ELASTICSEARCH_ENDPOINT = process.env.ELASTICSEARCH_ENDPOINT;
const ELASTICSEARCH_API_KEY = process.env.ELASTICSEARCH_API_KEY;

//Error check for undefined credentials
if (!ELASTICSEARCH_ENDPOINT || !ELASTICSEARCH_API_KEY) {
  throw new Error('Missing Elasticsearch credentials');
}

//Check to see if a connection can be established
const testClient = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: { 
    apiKey: ELASTICSEARCH_API_KEY 
  },
});

try {
  await testClient.ping();
  console.log('Connected to Elasticsearch successfully');
} catch (error: unknown) {
  if (error instanceof Error) {
    console.error('Failed to connect to Elasticsearch:', error.message);
  } else {
    console.error('Failed to connect to Elasticsearch:', error);
  }
  process.exit(1);
}
//Initialize the Elasticsearch vector store
const vectorStore = new ElasticVector({
  node: ELASTICSEARCH_ENDPOINT,
  auth: {
    apiKey: ELASTICSEARCH_API_KEY,
  },
//Optional: Explicitly set to true if using Elasticsearch Serverless to skip auto-detection and improve startup time
//isServerless: true,
});<p>在这里，我们</p><ul><li><p>使用<code>dotenv</code> 从<code>.env</code> 文件中加载变量。</p></li><li><p>检查 Elasticsearch 凭据是否被正确注入，我们是否能成功建立与客户端的连接。</p></li><li><p>在<code>ElasticVector</code> 构造函数中输入 Elasticsearch 端点和 API 密钥，以创建我们之前定义的向量存储实例。</p></li><li><p>如果使用 Elasticsearch Serverless，可选择指定<code>isServerless: true</code> 。这样可以跳过自动检测步骤，缩短启动时间。如果省略，适配器将在首次使用时自动检测您的部署类型。</p></li></ul><p>接下来，我们可以使用 Mastra 的<code>Agent</code> 类来定义代理。</p>export const knowledgeAgent = new Agent({
    name: 'KnowledgeAgent',
    instructions: 'You are a helpful knowledge assistant.',
    model: openai('gpt-4o'),
    memory: new Memory({

        vector: vectorStore,

        //embedder used to create embeddings for each message
        embedder: 'openai/text-embedding-3-small',

        //set semantic recall options
        options: {
            semanticRecall: {
                topK: 3, // retrieve 3 similar messages
                messageRange: 2, // include 2 messages before/after each match
                scope: 'resource',
            },
        },
    }),
});<p>我们可以定义的字段有</p><ul><li><p><code>name</code> 和<code>instructions</code> ：赋予其特性和主要功能。</p></li><li><p><code>model</code>:我们通过<code>@ai-sdk/openai</code> 软件包使用 OpenAI 的<code>gpt-4o</code> 。</p></li><li><p><code>memory</code>:</p><ul><li><p><code>vector</code>:指向我们的 Elasticsearch 存储库，因此嵌入式会从那里存储和检索。</p></li><li><p><code>embedder</code>:使用哪种模型生成嵌入模型</p></li><li><p><code>semanticRecall</code> 选项决定召回如何进行：</p><ul><li><p><code>topK</code>:检索多少条语义相似的信息。</p></li><li><p><code>messageRange</code>:每场比赛应包括多少对话内容。</p></li><li><p><code>scope</code>:定义内存边界。</p></li></ul></li></ul></li></ul><p>快好了我们只需将新创建的代理添加到 Mastra 配置中。在名为<a href="http://index.ts/"><code>index.ts</code></a> 的文件中，导入知识代理并将其插入<code>agents</code> 字段。</p>export const mastra = new Mastra({
  agents: { knowledgeAgent },
  storage: new LibSQLStore({
    // stores observability, scores, ... into memory storage, if it needs to persist, change to file:../mastra.db
    url: ":memory:",
  }),
  logger: new PinoLogger({
    name: 'Mastra',
    level: 'info',
  }),
  telemetry: {
    // Telemetry is deprecated and will be removed in the Nov 4th release
    enabled: false, 
  },
  observability: {
    // Enables DefaultExporter and CloudExporter for AI tracing
    default: { enabled: true }, 
  },
});<p>其他领域包括</p><ul><li><p><code>storage</code>:这是 Mastra 的内部数据存储，用于存储运行历史、可观察性指标、分数和缓存。有关 Mastra 存储的更多信息，请访问<a href="https://mastra.ai/docs/server-db/storage">此处</a>。</p></li><li><p><code>logger</code>:Mastra 使用<a href="https://github.com/pinojs/pino">Pino</a>，这是一个轻量级结构化 JSON 日志记录器。它可捕捉代理启动和停止、工具调用和结果、错误以及 LLM 响应时间等事件。</p></li><li><p><code>observability</code>:控制人工智能跟踪和代理执行的可见性。它可以跟踪</p><ul><li><p>每个推理步骤的开始/结束。</p></li><li><p>使用了哪种模式或工具。</p></li><li><p>输入和输出。</p></li><li><p>分数和评估</p></li></ul></li></ul><h3>使用 Mastra Studio 测试代理</h3><p>祝贺你如果您已经到达这里，那么您就可以运行这个代理，测试它的语义回忆能力了。幸运的是，Mastra 提供了一个内置的聊天用户界面，这样我们就不必自己创建了。</p><p>要启动 Mastra 开发服务器，请打开终端并运行以下命令：</p>npm run dev<p>在初始捆绑和启动服务器后，它应该会为你提供一个 Playground 的地址。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte5f857fddc74ffc9/6a16f7b6a6c2b995d5e794c0/8b045f70008d26aec4d2e6b59d61085555b9c5b2-686x116.png" alt="Playground 服务器地址" /><p>将此地址粘贴到浏览器中，您将看到 Mastra Studio。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc7fdda6ce46ce068/6a16f7b7b0367d4f7672bacf/69bc80fe8486edd9e0cf91d87b39f465aeb23111-1600x438.png" alt="粘贴 Playground 地址以访问 Mastra Studio" /><p>选择<code>knowledgeAgent</code> ，然后开始聊天。</p><p>为了快速测试一切接线是否正确，请给它提供一些信息，如 "团队宣布 10 月份的销售业绩增长了 12% ，主要是由企业续订驱动的。下一步是扩大对中端市场客户的拓展"。下一步，开始新的聊天，并提问："我们说过下一步需要关注哪个客户群？知识代理应该能够回忆起您在第一次聊天中提供给它的信息。您应该会看到类似的回复：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfec3266e81a7213b/6a16f7b92b835f6f70f4afe2/da8ebddad89874023ed440a8f1ad2cb04ed043f4-1070x288.png" alt="在 Mastra Studio 中与知识代理聊天--代理可以调用信息" /><p>看到这样的响应，意味着代理成功地将我们之前的信息以嵌入的形式存储在 Elasticsearch 中，并在稍后使用向量搜索进行检索。</p><h3>检查代理的长期记忆存储</h3><p>在 Mastra Studio 的代理配置中，前往<code>memory</code> 选项卡。这可以让您了解您的代理随着时间的推移学到了什么。嵌入并存储在 Elasticsearch 中的每一条消息、响应和交互都会成为长期记忆的一部分。您可以对过去的交互进行语义搜索，以快速找到代理之前了解到的信息或上下文。这与代理在语义回想时使用的机制基本相同，但在这里你可以直接检查它。在下面的示例中，我们搜索 "销售 "一词，并返回所有包含销售内容的互动。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte428134d7bf2a43a/6a16f7bbb0367d185872bad3/3decaa0c332d288c5ae0b11c25f592c7d50c2f0f-1104x1320.png" alt="如何检查知识代理的长期记忆存储" /><h2>结论</h2><p>通过连接 Mastra 和 Elasticsearch，我们可以为代理提供内存，这是上下文工程的关键层。有了语义记忆功能，代理可以随着时间的推移建立上下文，将他们的反应建立在所学知识的基础上。这意味着更准确、更可靠、更自然的互动。</p><p>早期的整合只是一个起点。同样的模式可以让支持代理记住过去的票单，让内部机器人检索相关文档，或者让人工智能助理在对话中回忆起客户的详细信息。我们还在努力实现与 Mastra 的正式集成，以便在不久的将来使这种搭配更加完美。</p><p>我们很期待看到您的下一个作品。试试吧，探索<a href="https://mastra.ai/">Mastra</a>及其内存功能，并随时与社区分享您的发现。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/knowledge-agent-semantic-recall-mastra-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/knowledge-agent-semantic-recall-mastra-elasticsearch</guid>
    <category><![CDATA[智能体 AI]]></category>
    <category><![CDATA[开发者体验]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[JD Armada]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt09afdbff05603865/6a16f7bd839dfabbf2dcfcb5/b8d51c2726d5573385c9246a7821d12ade4f1b0e-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 06 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>