<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Jeffrey Rengifo - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Jeffrey Rengifo - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/cn/search-labs/author/jeffrey-rengifo</link>
    </image>
    <link>https://www.elastic.co/cn/search-labs/author/jeffrey-rengifo</link>
    <atom:link href="https://www.elastic.co/cn/search-labs/rss/author/jeffrey-rengifo.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[cn]]></language>
    <lastBuildDate>Tue, 29 Sep 2026 05:55:28 GMT</lastBuildDate>
  <item>
    <title><![CDATA[如何衡量和提升 Elasticsearch 搜索召回率：通过混合搜索将召回率从 0.43 提升至 0.75]]></title>
    <description><![CDATA[了解如何通过将 BM25 词汇搜索与 Jina AI 向量嵌入相结合来测量和提高 Elasticsearch 中的搜索召回率，并使用 rank_eval API 以实际数据验证改进效果。]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/docs/solutions/search/full-text">词汇搜索</a>使用 <a href="https://www.elastic.co/blog/practical-bm25-part-1-how-shards-affect-relevance-scoring-in-elasticsearch">BM25 排序算法</a>，对于各种查询来说成本低、速度快且非常有效。但它有一个盲点：无法处理与文档没有共同标记的查询。在本文中，您将准确衡量 BM25 的不足之处。我们将使用 Elasticsearch 的<a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval">排名评估 API</a> (<code>rank_eval</code>)，并通过添加 <a href="https://www.elastic.co/search-labs/es/blog/jina-embeddings-v3-elastic-inference-service">Jina AI 嵌入</a>，通过 <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic 推理服务</a> (EIS) 来缩小这一差距。您会看到召回分数从 <code>0.43</code> 提升到 <code>0.75</code>，并理解其原因。</p><h2>什么是召回？</h2><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#k-recall">召回率</a> 以 <code>0</code> 到 <code>1</code> 的范围来衡量用户真正想要的文档有多少出现在搜索结果中。如果某个查询应显示三个产品，而您的搜索结果仅有两个进入前 10 名，则该查询的得分为 <code>recall@10 = 0.67</code>。这是一个基于集合的指标：它并不关心相关文档在这 <em>k</em> 个结果中的位置。位置 10 的相关文档与位置 1 的相关文档具有同等效力。高召回率意味着您不会丢失相关结果。</p><p>
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ffd147b13705680/6a170a6fe8fbce11a539fc22/b13af2a5d0ca055535d8bfe3dfe4b3d1093ee6da-1457x796.png" alt="维恩图展示了如何计算 Recall@10，通过显示所有相关文档与 BM25 检索出的前 10 个结果的重叠情况，得出 Recall@10 得分为 0.40。" /><p>该图表显示了两组文档：所有相关文档（左侧）和 BM25 实际检索到的文档（前 10 个，右侧）。只有交集部分才计入召回率，找到了 <code>prod_1</code> 和 <code>prod_2</code>，而 <code>prod_3</code>、<code>prod_4</code> 和 <code>prod_6</code> 则完全遗漏。结果：<code>Recall@10 = 2/5 = </code><strong><code>0.40</code></strong>。</p><h2>准备工作</h2><p>让我们言归正传，更好地了解召回的工作原理。本演示使用 Python。您可以在配套笔记本 (<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/relevance-tuning-improving-recall-adding-vectors/notebook.ipynb">notebook.ipynb</a>) 中跟着操作，其中每个代码块都是一个可直接运行的单元。</p><p>提供的代码使用以下内容：</p><ul><li><p>Elasticsearch 9.3+</p></li><li><p>Python 3.10+</p></li></ul>pip install elasticsearch pandas plotly python-dotenv<ul><li><p>包含 Elasticsearch 凭据的 <code>.env</code> 文件</p></li></ul>ELASTICSEARCH_URL=https://your-cluster-url
ELASTICSEARCH_API_KEY=your-api-key<h2>该数据集</h2><p>我们将使用包含 1,000 种产品的产品目录，涵盖鞋类、电子产品、工具等多个类别。</p><p>每份文档有四个字段：</p><p>字段</p><p>类型</p><p>“标题”</p><p>文本</p><p>“描述”</p><p>文本</p><p>“品牌”</p><p>关键字</p><p>`类别`</p><p>关键字</p><p>该数据集加载自 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/relevance-tuning-improving-recall-adding-vectors/dataset.csv"><code>dataset.csv</code></a>。</p><h2>词汇搜索的支持和局限性</h2><p>BM25 是 Elasticsearch 和大多数搜索引擎的默认排名算法。它根据查询词在文档中的出现频率对其进行评分，并根据文档长度和这些词在整个索引中的出现频率进行调整。在此基础上，您还可以获得<a href="https://www.elastic.co/docs/reference/text-analysis/analyzer-reference">分析器</a>：小写规范化、词干提取和停用词消除。查询“跑步鞋”将匹配“跑步鞋”，也可能匹配“跑步”。</p><p>这对很多查询都很有效：</p><ul><li><p>“跑鞋”会立即匹配标题中包含这些确切标记的产品。</p></li><li><p>“蓝牙扬声器”会显示便携式音频产品，因为这些词语是逐字匹配的。</p></li></ul><p>搜索结果具有确定性和可解释性：文档排名靠前，是因为查询词出现在其中。调试相关性很简单。</p><h3>出现问题的地方</h3><p>现在，让我们针对同一目录尝试这些查询：</p><ul><li><p><strong>“护肤流程”：</strong>在任何产品标题中都没有出现“流程”这个词。BM25 能够部分匹配“护肤”这一词，但面部精华液、身体精油和保湿霜等产品是用“维生素 C”、“视黄醇”或“提亮”等术语来描述的，这些术语与查询词都没有重叠。构成完整护肤流程的产品分散在索引中，没有任何共同的令牌将其关联起来。</p></li></ul>ID: B06XX6DS3P, Score: 9.0552, Title: Replenix Retinol Smooth + Tighten Body Lotion - Collagen-Boosting, Regenerating Anti-Aging Body Cream, Reduces Appearance of Stretch Marks, 6.7 oz.

  ID: B08XMPKJ1L, Score: 5.2699, Title: Bio-Oil Skincare Body Oil (Natural) Serum for Scars and Stretchmarks, Face and Body Moisturizer Hydrates Skin, with Organic Jojoba Oil and Vitamin E, For All Skin Types, 6.7 oz

  ID: B01CY764KQ, Score: 5.0057, Title: Nike Up Or Down Men Deodorant - Pack of 2 | Long-Lasting Fragrance, Body Spray Combo for Men | Deodorant for Active Living | Nike Men's Deo Set | Ultimate Odor Protection | Grooming Essentials | Signature Nike Scent | High-Performance Men's Deodorant<ul><li><p><strong>“宠物旅行配件”：</strong>这是一个用例分组，而非产品类别。宠物狗背带、宠物汽车座椅和旅行笼都与此相关，但它们的描述侧重于便携性、安全性和舒适性，而非“旅行配件”。BM25 与“宠物”大致匹配，但无法区分旅行专用产品与宠物目录中的其他产品。</p></li></ul>ID: B0BVV7BKTW, Score: 7.4371, Title: Large Foldable Travel Duffel Bag with Shoes Compartment

ID: B07TNPHYNV, Score: 6.6455, Title: 40 Pieces Christmas Bronze Jingle Bells Craft Small Bells

ID: B08R8FRW53, Score: 6.6335, Title: CUBY Dog and Cat Sling Carrier
ID: B08QMCQYGM, Score: 6.5259, Title: YTFGGY Whiteboard Pinstripe Tape 6 Rolls 1/8"
ID: B0CP3LQSWM, Score: 6.2994, Title: Portable Dog Water Bottle 32 Oz<p>这是一个<strong>召回问题</strong>。相关文档已存在于您的索引中。BM25 无法找到它们，因为用户的用词和文档中的词语匹配度不够高。</p><p>添加同义词有助于处理已知情况。但您无法枚举用户表达某种意图的所有方式。这就是向量发挥作用的地方。</p><h2>为何要测量召回率</h2><p>在解决问题之前，需要先对问题进行量化。</p><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#k-recall"><strong>Recall@k</strong></a> 衡量有多少用户真正想要的文档出现在搜索结果中。正式来说：</p>Recall@k = (relevant documents found in top k) / (total relevant documents)<p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#k-precision"><strong>Precision@k</strong></a> 衡量前 k 个结果，以及其中有多少是实际相关的：</p>Precision@k = (relevant documents in top k) / k<p>高精度意味着您返回的结果质量较高。在电子商务领域，缺少相关产品（召回率低）通常比显示稍有瑕疵的结果（精度较低）更糟糕，因为隐藏的产品意味着销售损失。</p><p>Elasticsearch 的 <code>rank_eval</code> API 允许您系统地测量两者。您提供一系列查询，每个查询都有一组已评分的文档，Elasticsearch 会为您计算所有查询的指标。</p><h2>设置评估</h2><p><code>rank_eval</code> API 需要一个<strong>评级数据集</strong>：查询与每个查询相关的文档之间的映射，以及相关性等级（0＝不相关，1＝相关，2＝高度相关）。</p><p>在笔记本中，这是<a href="https://www.elastic.co/docs/solutions/search/ranking/learning-to-rank-ltr#learning-to-rank-judgement-list">判断列表</a>：</p>judgments = [
    # Query 1: "running shoes" BM25 handles well (tokens appear in product titles) 
    {"query_id": "q1", "doc_id": "B09NQJFRW6", "grade": 2, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B08JMD4LMM", "grade": 2, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B08VRJ6F2Q", "grade": 2, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B07S8NRRWR", "grade": 2, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B01HD620I8", "grade": 2, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B07DX86321", "grade": 2, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B0968YVLQ8", "grade": 1, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B093QJ39ZS", "grade": 1, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B096FGSC39", "grade": 1, "query": "running shoes"},
    {"query_id": "q1", "doc_id": "B01GVQWVV2", "grade": 1, "query": "running shoes"},

    # Query 2: "skincare routine" intent-based, "routine" never appears in product titles
    {"query_id": "q2", "doc_id": "B08XMPKJ1L", "grade": 2, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B0BN3WQB92", "grade": 2, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B0BT7B7P5T", "grade": 2, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B00NPA2WEY", "grade": 2, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B06XX6DS3P", "grade": 1, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B07PDRD1KT", "grade": 1, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B074J7869B", "grade": 1, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B08JV31QW4", "grade": 1, "query": "skincare routine"},
    {"query_id": "q2", "doc_id": "B00K3TVJMQ", "grade": 1, "query": "skincare routine"},

    # Query 3: "study desk setup" intent-based, products are desks/stands/organizers
    {"query_id": "q3", "doc_id": "B08CS35J2T", "grade": 2, "query": "study desk setup"},
    {"query_id": "q3", "doc_id": "B09B3LFDXJ", "grade": 2, "query": "study desk setup"},
    {"query_id": "q3", "doc_id": "B07W58LMND", "grade": 1, "query": "study desk setup"},
    {"query_id": "q3", "doc_id": "B0CHYDX91L", "grade": 1, "query": "study desk setup"},

    # Query 4: "pet travel accessories" use-case grouping, products are carriers/crates/seats
    {"query_id": "q4", "doc_id": "B08R8FRW53", "grade": 2, "query": "pet travel accessories"},
    {"query_id": "q4", "doc_id": "B01MYUYX33", "grade": 2, "query": "pet travel accessories"},
    {"query_id": "q4", "doc_id": "B003C5RKE4", "grade": 2, "query": "pet travel accessories"},
    {"query_id": "q4", "doc_id": "B09GF8GBF6", "grade": 1, "query": "pet travel accessories"},
    {"query_id": "q4", "doc_id": "B0CP3LQSWM", "grade": 1, "query": "pet travel accessories"},
]<p>这种混合是有意为之：<code>q1</code> 是 BM25 可以很好处理的查询（产品标题中的精确标记），而 <code>q2</code>、<code>q3</code> 和 <code>q4</code> 是基于意图的查询，用户的意图是以概念而非具体产品关键词来表达的。</p><h2>测量 BM25 基线召回率</h2><p>首先，设置 Elasticsearch 客户端，并对原始文本数据建立索引：</p>import os
import json
import pandas as pd
import plotly.graph_objects as go
from elasticsearch import Elasticsearch, helpers
from dotenv import load_dotenv

load_dotenv()

es = Elasticsearch(
    os.getenv("ELASTICSEARCH_URL"),
    api_key=os.getenv("ELASTICSEARCH_API_KEY")
)

INDEX_NAME = "ecommerce-products"<p>现在为 BM25 构建 <code>rank_eval</code> 请求。列表中的每个请求都将会查询及其评分结合起来：</p>judgments_df = pd.DataFrame(judgments)

bm25_requests = []
for query_id, query_text in (
    judgments_df[["query_id", "query"]].drop_duplicates().values
):
    relevant_docs = judgments_df[judgments_df["query_id"] == query_id]
    ratings = [
        {"_index": INDEX_NAME, "_id": row["doc_id"], "rating": row["grade"]}
        for _, row in relevant_docs.iterrows()
    ]

    bm25_requests.append({
        "id": query_id,
        "request": {
            "query": {
                "multi_match": {
                    "query": query_text,
                    "fields": ["title", "description"]
                }
            }
        },
        "ratings": ratings,
    })

bm25_eval = {
    "requests": bm25_requests,
    "metric": {"recall": {"k": 10, "relevant_rating_threshold": 1}},
}

bm25_result = es.rank_eval(index=INDEX_NAME, body=bm25_eval)
print("BM25 Recall@10:", bm25_result.body["metric_score"])<p>结果：</p>BM25 Recall@10: 0.43<p><code>0.43</code> 这意味着在所有四个查询中，BM25 只找到了它应该找到的文档的 43%。这种不足集中体现在基于意图的查询中：“护肤流程”漏掉了面部精华液和身体精油，因为“流程”一词从未出现在产品标题中；而“宠物旅行配件”则检索出了一些不相关的宠物产品，却遗漏了那些以便携性和安全性而非“旅行配件”来描述的宠物笼和宠物箱。</p><p>这就是我们的基准。现在我们有了一个要超越的数字。</p><h2>使用 Jina 嵌入添加向量搜索</h2><p><a href="https://www.elastic.co/docs/solutions/search/vector"><code>Vector search</code></a> 将文档和查询编码为高维向量，这是一种由数百甚至数千个数值组成的向量，每个数值都对它所代表的数据的特定特征进行编码。意义相似的文档最终会在向量空间中靠近，即使它们没有共同的词汇。“健身器材”和“哑铃套装”会放在一起，因为这两个概念是相关的。我选择 Elasticsearch 作为我的向量数据库，是因为它支持混合搜索，让我既能理解语义，又能精确查找关键字。</p><p><a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">EIS</a> 包括通过其<a href="https://www.elastic.co/docs/api/doc/elasticsearch/group/endpoint-inference">推理 API</a> 嵌入模型的开箱即用支持。</p><h3>步骤 1：使用 Jina 嵌入 v5 作为推理终端</h3>INFERENCE_ENDPOINT_ID = ".jina-embeddings-v5-text-small"<p>如果您的集群具有 GPU 资源（在 Elastic Cloud 和 Elasticsearch 9.3+ 中可用），嵌入将在 GPU 上生成，这比 CPU 推理快得多，并消除了历史上使向量在扩展时变得昂贵的性能权衡。</p><p>为什么要特别选用 Jina 嵌入？<a href="https://www.elastic.co/search-labs/blog/jina-embeddings-v5-text">jina-embeddings-v5-text</a> 是一种多语言模型（支持 119 种以上语言），具有 32,000 个标记的上下文窗口，并支持特定任务的<a href="https://arxiv.org/abs/2106.09685">低秩自适应 (LoRA) 适配器</a>。它适用于开箱即用的简短产品描述。<a href="https://huggingface.co/jinaai/jina-embeddings-v5-text-small">点击此处</a>了解有关 <code>jina-embeddings-v5-text</code> 模型的更多信息。</p><h3>步骤 2：创建具有语义字段的索引</h3>index_mappings = {
    "mappings": {
        "properties": {
            "title": {"type": "text", "copy_to": "semantic_field"},
            "description": {"type": "text", "copy_to": "semantic_field"},
            "brand": {"type": "keyword"},
            "category": {"type": "keyword"},
            "semantic_field": {
                "type": "semantic_text",
                "inference_id": INFERENCE_ENDPOINT_ID,
            },
        }
    }
}

if not es.indices.exists(index=INDEX_NAME):
    es.indices.create(index=INDEX_NAME, body=index_mappings)
    print(f"Created index: {INDEX_NAME}")<p>这里的关键在于 <a href="https://www.elastic.co/docs/solutions/search/semantic-search/semantic-search-semantic-text"><code>semantic_text</code></a> 字段类型。这是对 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/dense-vector"><code>dense_vector</code></a> 的更高级别的抽象：您将其指向一个推理终端，Elasticsearch 会自动生成嵌入。</p><p><code>title</code> 和<code>description</code> 上的 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/copy-to"><code>copy_to</code></a> 属性意味着这两个字段的内容都会流入 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_field</code></a> 进行嵌入，因此单个向量就能捕获完整的产品表示。</p><h3>步骤 3：为产品编制索引</h3>def bulk_index(products, index_name):
    actions = []
    for product in products:
        doc_id = product.get("_id")
        source = {k: v for k, v in product.items() if k != "_id"}
        action = {"_index": index_name, "_source": source}
        if doc_id:
            action["_id"] = doc_id
        actions.append(action)

    success, failed = helpers.bulk(es, actions, raise_on_error=False)
    if failed:
        for error in failed:
            print(f"Error: {error}")
    else:
        print(f"Successfully indexed {success} documents")

bulk_index(products, INDEX_NAME)<p>索引时，Elasticsearch 会调用每个文档的推理端点，并将生成的嵌入存储在 <code>semantic_field</code> 中。您无需编写任何额外代码。</p><h2>混合搜索：将 BM25 与向量结合并采用 RRF</h2><p>添加向量可以提高召回率，但仅使用向量可能会在精确匹配查询中失去精度；“跑鞋”仍应将逐字匹配的结果排在首位。混合搜索则保留词汇成分，以保持这种精确性。</p><p>使用<a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion">倒数排序融合</a> (RRF) 的混合搜索可以保持两者的优点：</p><ul><li><p>BM25 可以高精度处理精确和近似精确的查询。</p></li><li><p>语义搜索能以高召回率处理基于意图和多语言的查询。</p></li><li><p>RRF 将两份排名表合并为一份排名表。</p></li></ul><p>RRF 公式根据每个文档在每个结果列表中的排名，为每个文档分配分数：</p>score = sum(1 / (rank_constant + rank))<p>在两个列表中均排名靠前的文档将获得更高的综合得分。<code>rank_constant</code>用于控制排名较低的文档获得的权重大小。</p>hybrid_requests = []

for query_id, query_text in (
    judgments_df[["query_id", "query"]].drop_duplicates().values
):
    relevant_docs = judgments_df[judgments_df["query_id"] == query_id]
    ratings = [
        {"_index": INDEX_NAME, "_id": row["doc_id"], "rating": row["grade"]}
        for _, row in relevant_docs.iterrows()
    ]

    hybrid_requests.append({
        "id": query_id,
        "request": {
            "retriever": {
                "rrf": {
                    "retrievers": [
                        {
                            "standard": {
                                "query": {
                                    "multi_match": {
                                        "query": query_text,
                                        "fields": ["title", "description"],
                                    }
                                }
                            }
                        },
                        {
                            "standard": {
                                "query": {
                                    "match": {
                                        "semantic_field": {"query": query_text}
                                    }
                                }
                            }
                        },
                    ],
                    "rank_window_size": 50,
                    "rank_constant": 5,
                }
            }
        },
        "ratings": ratings,
    })

hybrid_eval = {
    "requests": hybrid_requests,
    "metric": {"recall": {"k": 10, "relevant_rating_threshold": 1}},
}

hybrid_result = es.rank_eval(index=INDEX_NAME, body=hybrid_eval)
print("Hybrid Recall@10:", hybrid_result.body["metric_score"])<p>结果：</p>Hybrid Recall@10: 0.75<p>混合搜索在 BM25 (<code>0.43</code>) 的基础上有了显著提升，并为“跑鞋”等精确匹配查询保留了精确度。</p><h2>结果：前后结果对比</h2><p>以下是所有三种方法的完整对比：</p>methods = {
    "BM25 (Lexical)": bm25_requests,
    "Hybrid (BM25 + Vectors)": hybrid_requests,
}

recall_metric = {"recall": {"k": 10, "relevant_rating_threshold": 1}}

comparison_data = []
for method_name, requests in methods.items():
    result = es.rank_eval(
        index=INDEX_NAME,
        body={"requests": requests, "metric": recall_metric}
    )
    comparison_data.append({
        "method": method_name,
        "recall@10": result.body["metric_score"]
    })

comparison_df = pd.DataFrame(comparison_data)
print(comparison_df.to_string(index=False))<p>结果：</p><p>方法</p><p>Recall@10</p><p>BM25（词法）</p><p>0.43</p><p>混合型（BM25 + 向量）</p><p>0.75</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5a1d72b57056fe64/6a170a71c1e8a56c58f882ab/e49f6c10516b0a48a0ad75962c6590ee07311407-700x500.png" alt="条形图比较了 BM25 词汇搜索和 BM25 与向量相结合的混合搜索的 Recall@10，结果显示混合搜索的召回率明显更高。" /><p>按查询细分：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt871347f754c866d0/6a170a73839dfa40abdcfeb4/40e36dcb7b34cbf4649c512bcb60cef60f1778a6-700x500.png" alt="分组条形图比较了四个产品查询中 BM25 词法搜索和混合搜索的 Recall@10，显示混合搜索在每个查询中始终优于词法搜索。" /><h2>结论</h2><p>在这篇文章中，我们看到，当用户键入精确的查询时，BM25 词汇搜索是可靠的，但当他们根据意图而非关键词进行搜索时，其召回率就会下降。借助 <code>rank_eval</code>，我们建立了一个可重复的基线，用真实数据来衡量这一差距。在此基础上，我们添加了一个由 Jina 嵌入提供支持的 <code>semantic_text</code> 字段，并再次运行了评估。结果：混合搜索将召回率从 <code>0.43</code> 提高到 <code>0.75</code>，同时保留了精确匹配查询的精确度，但实际幅度取决于您的查询组合。</p><p>该模式可扩展至本示例之外：从用户的实际查询中收集判断，以 <code>rank_eval</code> 作为基准运行，添加 <code>semantic_text</code>，然后再次进行测量。您将确切了解改进了哪些方面以及改进了多少。</p><h2>后续步骤</h2><ul><li><p>深入了解召回与向量搜索：《<a href="https://www.elastic.co/search-labs/blog/recall-vector-search-quantization">召回与向量搜索量化</a>》，作者：Jeff Vestal</p></li><li><p>添加重排序功能，以进一步提升前几条结果的精准度</p></li><li><p>探索 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/rrf.html">Elasticsearch 混合搜索文档</a></p></li><li><p>阅读有关 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-rank-eval.html"><code>rank_eval</code></a> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-rank-eval.html">API</a> 的更多信息</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-relevance-tuning-improve-recall</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-relevance-tuning-improve-recall</guid>
    <category><![CDATA[混合搜索]]></category>
    <category><![CDATA[向量数据库]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt37c9d2971b5a2db3/6a170a75cf4f254223b2d149/492c9b5432a2b9e40cebb3b60f0df019a8c7bf6d-1280x720.png" length="0" type="image/png"/>
    <pubDate>Mon, 04 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 TypeScript 构建 Elasticsearch MCP 服务器]]></title>
    <description><![CDATA[学习如何使用 TypeScript 和 Claude Desktop 创建 Elasticsearch MCP 服务器。]]></description>
    <content:encoded><![CDATA[<p>在 Elasticsearch 中处理大型知识库时，找到信息只是成功的一半。工程师通常还需要综合多个文档的结果，生成摘要，并追溯答案的来源。模型上下文协议 (MCP) 提供了一种标准化的方式，可将 Elasticsearch 与大语言模型 (LLM) 驱动的应用程序连接起来，以实现上述目标。虽然 Elastic 提供官方解决方案，例如 Elastic Agent Builder（其功能包括 <a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">MCP 终端</a>），但构建自定义 MCP 服务器可让您完全掌控搜索逻辑、结果格式，以及如何将检索到的内容传递给 LLM，以用于综合分析、生成摘要和提供引用。</p><p>本文将探讨构建自定义 Elasticsearch MCP 服务器的优势，并展示如何使用 TypeScript 创建该服务器，以将 Elasticsearch 连接到 LLM 驱动的应用程序。</p><h2>为什么要构建自定义 Elasticsearch MCP 服务器？</h2><p>Elastic 为 <a href="https://www.elastic.co/docs/solutions/search/mcp">MCP 服务器</a>提供了一些替代方案：</p><ul><li><p><a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">Elastic Agent Builder MCP 服务器，适用于 Elasticsearch 9.2 及以上版本</a></p></li><li><p><a href="https://github.com/elastic/mcp-server-elasticsearch?tab=readme-ov-file#elasticsearch-mcp-server">适用于旧版本的 Elasticsearch MCP 服务器（Python）</a></p></li></ul><p>如果您需要更好地控制 MCP 服务器与 Elasticsearch 的交互，构建自己的自定义服务器可以让您灵活地根据自身需求进行定制。例如，Agent Builder 的 MCP 终端仅限于 Elasticsearch 查询语言 (ES|QL) 查询，而自定义服务器允许您使用完整的查询 DSL。在将结果传递给 LLM 之前，您还可以控制结果的格式，并可以集成其他处理步骤，例如我们将在本教程中实现的由 OpenAI 驱动的摘要功能。</p><p>通过阅读本文，您将学会使用 TypeScript 创建 MCP 服务器，该服务器可搜索存储在 Elasticsearch 索引中的信息，对其进行总结并提供引用。我们将使用 Elasticsearch 进行检索，使用 OpenAI 的 <code>gpt-4o-mini</code> 模型提炼摘要并生成引用，并使用 Claude Desktop 作为 MCP 客户端和 UI 来接收用户查询并提供回复。最终我们将得到一个内部知识助手，帮助工程师在整个组织的技术文档中发现并综合最佳实践。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltad9133cb083ad352/6a170c19b0367d411e72bd5b/ec5771a874cf9740d4cac6888622cbe8cd6aede7-1999x1133.png" alt="使用 TypeScript 和 Claude Desktop 构建 Elastic MCP 服务器。" /><h2>准备工作：</h2><ul><li><p>Node.js 20 +</p></li><li><p>Elasticsearch</p></li><li><p>OpenAI API 密钥</p></li><li><p>Claude Desktop</p></li></ul><h3>什么是 MCP？</h3><p><a href="https://www.elastic.co/what-is/mcp">MCP</a> 是由 <a href="https://www.anthropic.com/news/model-context-protocol">Anthropic</a> 创建的开放标准，提供大型语言模型与外部系统（如 Elasticsearch）之间的安全双向连接。您可以在<a href="https://www.elastic.co/search-labs/blog/mcp-current-state">这篇文章</a>中了解更多关于 MCP 现状的信息。</p><p>MCP 的发展<a href="https://www.elastic.co/search-labs/blog/mcp-current-state#mcp-project-updates:-transport,-elicitation,-and-structured-tooling">每天都在变化</a>，服务器的使用范围越来越广。此外，构建自定义 MCP 服务器也非常简单，我们将在本文中进行演示。</p><h3>MCP 客户端</h3><p><a href="https://modelcontextprotocol.io/clients">可用的 MCP 客户端</a>由很多，每个客户端都有自己的特点和局限性。为了简化和普及，我们将使用 <a href="https://claude.ai/download">Claude Desktop</a> 作为演示中的 MCP 客户端。它将作为聊天界面，用户可以用自然语言提问，它还将自动调用我们的 MCP 服务器提供的工具来搜索文档和生成摘要。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt06fd7a02042094e1/6a170c1b14b2700024e3c651/66eb0b11473347b6cf2d85718251eeac38d6249d-1999x1491.png" alt="Claude 4.5 十四行诗页面，附有“到喝咖啡和用 Claude 时间了？今天我能为您做什么？”" /><h2>创建 Elasticsearch MCP 服务器</h2><p>通过使用 <a href="https://github.com/modelcontextprotocol/typescript-sdk">TypeScript 软件开发工具包</a>，我们可以轻松创建一个能够根据用户查询输入来查询 Elasticsearch 数据的服务器。</p><p>本文将介绍将 Elasticsearch MCP 服务器与 Claude Desktop 客户端集成的步骤：</p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#configure-mcp-server-for-elasticsearch">为 Elasticsearch 配置 MCP 服务器。</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#load-the-mcp-server-into-claude-desktop">将 MCP 服务器加载到 Claude Desktop。</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#test-it-out">测试一下。</a></p></li></ol><h3>为 Elasticsearch 配置 MCP 服务器。</h3><p>首先，我们来初始化一个 Node 应用程序：</p>npm init -y<p>这将会创建一个 <code>package.json</code> 文件，有了它，我们就可以开始安装该应用程序所需的依赖项。</p>npm install @elastic/elasticsearch @modelcontextprotocol/sdk openai zod &amp;&amp; npm install --save-dev ts-node @types/node typescript<ul><li><p><strong>@elastic/elasticsearch</strong> 将使我们能够访问 Elasticsearch Node.js 库。</p></li><li><p><strong>@modelcontextprotocol/sdk</strong> 提供核心工具来创建和管理 MCP 服务器、注册工具以及处理与 MCP 客户端的通信。</p></li><li><p><strong>openai</strong> 允许与 OpenAI 模型进行交互以生成摘要或自然语言响应。</p></li><li><p><a href="https://zod.dev/"><strong>zod</strong></a>帮助定义和验证每个工具中输入和输出数据的结构化模式。</p></li></ul><p><code>ts-node</code>，<code>@types/node</code> 和 <code>typescript</code> 将在开发过程中用于键入代码和编译脚本。</p><h4>配置数据集</h4><p>为了提供 Claude Desktop 可以使用我们的 MCP 服务器进行查询的数据，我们将使用模拟的<a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/dataset.json">内部知识库数据集</a>。来自该数据集的文档是这样子的：</p>{
    "id": 5,
    "title": "Logging Standards for Microservices",
    "content": "Consistent logging across microservices helps with debugging and tracing. Use structured JSON logs and include request IDs and timestamps. Avoid logging sensitive information. Centralize logs in Elasticsearch or a similar system. Configure log rotation to prevent storage issues and ensure logs are searchable for at least 30 days.",
    "tags": ["logging", "microservices", "standards"]
}<p>为了摄取数据，我们准备了一个脚本，该脚本在 Elasticsearch 中创建一个索引并将数据集加载到其中。您可以<a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/setup.ts">在这里</a>找到它。</p><h4>MCP 服务器</h4><p>创建一个名为 <a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/index.ts"><code>index.ts</code></a> 的文件，并添加以下代码来导入依赖项并处理环境变量：</p>// index.ts
import { z } from "zod";
import { Client } from "@elastic/elasticsearch";
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import OpenAI from "openai";

const ELASTICSEARCH_ENDPOINT =
  process.env.ELASTICSEARCH_ENDPOINT ?? "http://localhost:9200";
const ELASTICSEARCH_API_KEY = process.env.ELASTICSEARCH_API_KEY ?? "";
const OPENAI_API_KEY = process.env.OPENAI_API_KEY ?? "";
const INDEX = "documents";<p>此外，让我们初始化客户端以处理 Elasticsearch 和 OpenAI 的调用：</p>const openai = new OpenAI({
  apiKey: OPENAI_API_KEY,
});

const _client = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: {
    apiKey: ELASTICSEARCH_API_KEY,
  },
});<p>为了使我们的实现更加稳健，并确保输入和输出结构化，我们将使用 <a href="https://zod.dev/"><code>zod</code></a> 定义模式。这使我们能够在运行时验证数据，及早发现错误，并使工具响应更容易以编程方式进行处理：</p>const DocumentSchema = z.object({
  id: z.number(),
  title: z.string(),
  content: z.string(),
  tags: z.array(z.string()),
});

const SearchResultSchema = z.object({
  id: z.number(),
  title: z.string(),
  content: z.string(),
  tags: z.array(z.string()),
  score: z.number(),
});

type Document = z.infer&lt;typeof DocumentSchema&gt;;
type SearchResult = z.infer&lt;typeof SearchResultSchema&gt;;<p>请在<a href="https://www.elastic.co/search-labs/blog/structured-outputs-elasticsearch-guide">此处</a>了解更多关于结构化输出的信息。</p><p>现在让我们初始化 MCP 服务器：</p>const server = new McpServer({
  name: "Elasticsearch RAG MCP",
  description:
    "A RAG server using Elasticsearch. Provides tools for document search, result summarization, and source citation.",
  version: "1.0.0",
});<h4>定义 MCP 工具</h4><p>完成所有配置后，我们就可以开始编写将由 MCP 服务器公开的工具了。此服务器公开两种工具：</p><ul><li><p><strong><code>search_docs</code></strong><strong>：</strong>使用全文本搜索在 Elasticsearch 中搜索文档。</p></li><li><p><strong><code>summarize_and_cite</code></strong><strong>：</strong>汇总和综合先前检索到的文档中的信息，以回答用户的问题。该工具还可添加引用源文档的引文。</p></li></ul><p>这两个工具共同构成了一个简单的“检索后总结”工作流，其中一个工具获取相关文档，另一个工具使用这些文档生成汇总的引用回复。</p><h4>工具响应格式</h4><p>每个工具都可以接受任意输入参数，但必须以以下结构作出响应：</p><ul><li><p><strong>内容：</strong>这是工具以非结构化格式做出的响应。该字段通常用于返回文本、图像、音频、链接或嵌入内容。在本应用程序中，它将用于返回包含工具生成的信息的格式化文本。</p></li><li><p><strong>结构化内容： </strong>这是一个可选返回，用于以结构化格式提供每个工具的结果。这对程序化用途非常有用。虽然本 MCP 服务器没有使用它，但如果您想开发其他工具或以编程方式处理结果，它可能会很有用。</p></li></ul><p>基于这个结构，让我们详细探讨每个工具。</p><h4>Search_docs 工具</h4><p>此工具在 Elasticsearch 索引中执行 <a href="https://www.elastic.co/docs/solutions/search/full-text">全文本搜索</a>，以根据用户查询检索最相关的文档。它突出显示关键匹配项，并快速提供相关性评分概述。</p>server.registerTool(
  "search_docs",
  {
    title: "Search Documents",
    description:
      "Search for documents in Elasticsearch using full-text search. Returns the most relevant documents with their content, title, tags, and relevance score.",
    inputSchema: {
      query: z
        .string()
        .describe("The search query terms to find relevant documents"),
      max_results: z
        .number()
        .optional()
        .default(5)
        .describe("Maximum number of results to return"),
    },
    outputSchema: {
      results: z.array(SearchResultSchema),
      total: z.number(),
    },
  },
  async ({ query, max_results }) =&gt; {
    if (!query) {
      return {
        content: [
          {
            type: "text",
            text: "Query parameter is required",
          },
        ],
        isError: true,
      };
    }

    try {
      const response = await _client.search({
        index: INDEX,
        size: max_results,
        query: {
          bool: {
            must: [
              {
                multi_match: {
                  query: query,
                  fields: ["title^2", "content", "tags"],
                  fuzziness: "AUTO",
                },
              },
            ],
            should: [
              {
                match_phrase: {
                  title: {
                    query: query,
                    boost: 2,
                  },
                },
              },
            ],
          },
        },
        highlight: {
          fields: {
            title: {},
            content: {},
          },
        },
      });

      const results: SearchResult[] = response.hits.hits.map((hit: any) =&gt; {
        const source = hit._source as Document;

        return {
          id: source.id,
          title: source.title,
          content: source.content,
          tags: source.tags,
          score: hit._score ?? 0,
        };
      });

      const contentText = results
        .map(
          (r, i) =&gt;
            `[${i + 1}] ${r.title} (score: ${r.score.toFixed(
              2,
            )})\n${r.content.substring(0, 200)}...`,
        )
        .join("\n\n");

      const totalHits =
        typeof response.hits.total === "number"
          ? response.hits.total
          : (response.hits.total?.value ?? 0);

      return {
        content: [
          {
            type: "text",
            text: `Found ${results.length} relevant documents:\n\n${contentText}`,
          },
        ],
        structuredContent: {
          results: results,
          total: totalHits,
        },
      };
    } catch (error: any) {
      console.log("Error during search:", error);

      return {
        content: [
          {
            type: "text",
            text: `Error searching documents: ${error.message}`,
          },
        ],
        isError: true,
      };
    }
  }
);<p><em>我们将 fuzziness : “AUTO” 配置</em><em>为根据被分析的词元的长度具有可变的拼写错误容忍度。我们还设置了</em> <em><code>title^2</code></em> <em>来提高标题字段匹配的文档的分数。</em></p><h4>摘要和引用工具</h4><p>该工具根据上一次搜索中检索到的文档生成摘要。它使用 OpenAI 的 <code>gpt-4o-mini</code> 模型来综合最相关的信息，提供直接来自搜索结果的响应，以回答用户的问题。除了摘要之外，它还返回所使用源文档的引用元数据。</p>server.registerTool(
  "summarize_and_cite",
  {
    title: "Summarize and Cite",
    description:
      "Summarize the provided search results to answer a question and return citation metadata for the sources used.",
    inputSchema: {
      results: z
        .array(SearchResultSchema)
        .describe("Array of search results from search_docs"),
      question: z.string().describe("The question to answer"),
      max_length: z
        .number()
        .optional()
        .default(500)
        .describe("Maximum length of the summary in characters"),
      max_docs: z
        .number()
        .optional()
        .default(5)
        .describe("Maximum number of documents to include in the context"),
    },
    outputSchema: {
      summary: z.string(),
      sources_used: z.number(),
      citations: z.array(
        z.object({
          id: z.number(),
          title: z.string(),
          tags: z.array(z.string()),
          relevance_score: z.number(),
        })
      ),
    },
  },
  async ({ results, question, max_length, max_docs }) =&gt; {
    if (!results || results.length === 0 || !question) {
      return {
        content: [
          {
            type: "text",
            text: "Both results and question parameters are required, and results must not be empty",
          },
        ],
        isError: true,
      };
    }

    try {
      const used = results.slice(0, max_docs);

      const context = used
        .map(
          (r: SearchResult, i: number) =&gt;
            `[Document ${i + 1}: ${r.title}]\\n${r.content}`
        )
        .join("\n\n---\n\n");

      // Generate summary with OpenAI
      const completion = await openai.chat.completions.create({
        model: "gpt-4o-mini",
        messages: [
          {
            role: "system",
            content:
              "You are a helpful assistant that answers questions based on provided documents. Synthesize information from the documents to answer the user's question accurately and concisely. If the documents don't contain relevant information, say so.",
          },
          {
            role: "user",
            content: `Question: ${question}\\n\\nRelevant Documents:\\n${context}`,
          },
        ],
        max_tokens: Math.min(Math.ceil(max_length / 4), 1000),
        temperature: 0.3,
      });

      const summaryText =
        completion.choices[0]?.message?.content ?? "No summary generated.";

      const citations = used.map((r: SearchResult) =&gt; ({
        id: r.id,
        title: r.title,
        tags: r.tags,
        relevance_score: r.score,
      }));

      const citationText = citations
        .map(
          (c: any, i: number) =&gt;
            `[${i + 1}] ID: ${c.id}, Title: "${c.title}", Tags: ${c.tags.join(
              ", ",
            )}, Score: ${c.relevance_score.toFixed(2)}`,
        )
        .join("\n");

      const combinedText = `Summary:\\n\\n${summaryText}\\n\\nSources used (${citations.length}):\\n\\n${citationText}`;

      return {
        content: [
          {
            type: "text",
            text: combinedText,
          },
        ],
        structuredContent: {
          summary: summaryText,
          sources_used: citations.length,
          citations: citations,
        },
      };
    } catch (error: any) {
      return {
        content: [
          {
            type: "text",
            text: `Error generating summary and citations: ${error.message}`,
          },
        ],
        isError: true,
      };
    }
  }
);<p>最后，我们需要用 <a href="https://github.com/modelcontextprotocol/typescript-sdk?tab=readme-ov-file#stdio">stdio</a> 启动服务器。这意味着 MCP 客户端将通过读取和写入其标准输入和输出流与我们的服务器进行通信。stdio 是最简单的传输选项，适用于客户端作为子进程启动的本地 MCP 服务器。在文件末尾添加以下代码：</p>const transport = new StdioServerTransport();
server.connect(transport);<p>现在请您使用以下命令编译该项目：</p>npx tsc index.ts --target ES2022 --module node16 --moduleResolution node16 --outDir ./dist --strict --esModuleInterop<p>这将创建一个 <code>dist</code> 文件夹，并在其中创建一个 <code>index.js</code> 文件。</p><h3>将 MCP 服务器加载到 Claude Desktop。</h3><p>请按照<a href="https://modelcontextprotocol.io/docs/develop/connect-local-servers">本指南</a>配置 MCP 服务器和 Claude Desktop。在 Claude 配置文件中，我们需要设置以下值：</p>{
  "mcpServers": {
    "elasticsearch-rag-mcp": {
      "command": "node",
      "args": [   "/Users/user-name/app-dir/dist/index.js"
      ],
      "env": {
        "ELASTICSEARCH_ENDPOINT": "your-endpoint-here",
        "ELASTICSEARCH_API_KEY": "your-api-key-here",
        "OPENAI_API_KEY": "your-openai-key-here"
      }
    }
  }
}<p><code>args</code> 的值应指向 <code>dist</code> 文件夹中的编译后文件。您还需要在配置文件中设置环境变量，使其名称与代码中定义的名称完全一致。</p><h3>测试一下</h3><p>在执行每个工具前，点击<strong>搜索和工具</strong>，确保这些工具已启用。您还可以在这里启用或禁用每个工具：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt395a7337021f9820/6a170c1c67045bb74d45c228/172981c2a54adabc70d5819013c3007670935605-1999x1002.png" alt="Claude 4.5 十四行诗页面，附有“下午好，Jeff。今天我能为您做什么？”" /><p>最后，让我们从 Claude Desktop 聊天中测试 MCP 服务器，并开始提问：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf4ac458dc0206271/6a170c1e66c4f91328f8c072/03654c0f8c53c714f801fba8b25747071179209b-1999x1353.png" alt="用户在 Claude Desktop 聊天中请求搜索有关身份验证方法和基于角色的访问控制的文档，以及 Claude 的回复。" /><p>针对问题“<strong>搜索有关身份验证方法和基于角色的访问控制的文档</strong>”，已执行 <code>search_docs</code> 工具并返回以下结果：</p>Most Relevant Documents:
Access Control and Role Management (highest relevance) - This document covers role-based access control (RBAC) principles, including ensuring users only have necessary permissions, regular auditing of user roles, revoking inactive accounts, and implementing just-in-time access for sensitive operations.
User Authentication with OAuth 2.0 - This document explains OAuth 2.0 authentication, which enables secure delegated access without credential sharing. It covers configuring identity providers, token management with limited scope and lifetime, and secure storage of refresh tokens.
Container Security Guidelines - While primarily about container security, this document touches on access control aspects like running containers as non-root users and avoiding embedded credentials.
Incident Response Playbook - This mentions role assignment during incidents (incident commander, communications lead, etc.), which relates to access control in emergency scenarios.
Logging Standards for Microservices - This document includes guidance on avoiding logging sensitive information, which is relevant to authentication security.<p>回复是：“太好了！我找到了 5 份关于身份验证方法和基于角色的访问控制的相关文档。以下是找到的内容：”</p><p>该工具调用会将源文档作为其响应有效负载的一部分返回，这些文档随后用于生成引用。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbaf48a9468770ce2/6a170c21509168bffee1bb14/25ff4c7e9563d99752f95540dafdc7fd211a66e3-800x530.gif" alt="Claude 4.5 Sonnet 页面，滚动的响应包括有关身份验证方法和基于角色的访问控制的五个相关文档。" /><p>您也可以在一次交互中串联使用多个工具。在这种情况下，Claude Desktop 会分析用户的问题，并确定首先需要调用 <code>search_docs</code> 来检索相关文档，然后将这些结果传递给 <code>summarize_and_cite</code> 以生成最终回答，所有这些都无需用户单独提示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta46ff45274e64192/6a170c230c4857a91501aac1/e6a8a46acb4236e77058f18bcd2f0737b5882c05-1999x1101.png" alt="Claude Desktop 聊天，附有“Jeff 回来了”，外加一个新的用户问题：“在我们的系统中改进身份验证和访问控制的主要建议是什么？附上参考文献。”" /><p>在这种情况下，对于“<strong>在我们的系统中改进身份验证和访问控制的主要建议是什么？附上参考文献</strong>”，我们得到了以下结果：</p>Based on the documentation, here are the main recommendations to improve authentication and access control across your systems:

Key Recommendations

1. Implement Role-Based Access Control (RBAC)
Ensure users have only the permissions necessary for their job functions. This principle of least privilege minimizes security risks by limiting access to what's truly needed. [1]

2. Regular Access Audits
Conduct regular audits of user roles and promptly revoke access for inactive accounts. This helps prevent unauthorized access from dormant or compromised accounts. [1]

3. Just-in-Time (JIT) Access
Implement just-in-time access for sensitive operations, where permissions are granted only when needed and for a limited duration. This reduces the attack surface by minimizing standing privileges. [1]

4. OAuth 2.0 for Secure Authentication
Use OAuth 2.0 to enable secure delegated access without sharing user credentials. This protocol allows third-party applications to access resources on behalf of users without exposing passwords. [2]

5. Token Security and Management
Configure your identity provider to issue tokens with limited scope and lifetime

Store refresh tokens securely
Validate access tokens consistently to maintain security [2]
References

Access Control and Role Management (Tags: security, access-control)
User Authentication with OAuth 2.0 (Tags: authentication, oauth)
These recommendations work together to create a defense-in-depth approach, where multiple security layers protect your systems from unauthorized access.<p>与上一步一样，我们可以看到每个工具对该问题的响应：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8f633c518e708a99/6a170c25ab7f082991db9ed6/cb606d356b2f7d5e4878a5eff71bc881869ac0ee-800x585.gif" alt="Claude 桌面聊天页面，包含滚动文本，其中包括每个工具对问题“在我们的系统中改进身份验证和访问控制的主要建议是什么？附上参考文献。”的响应。" /><p><em>注意：如果出现子菜单询问是否批准使用每个工具，请选择</em><em><strong>“始终允许”</strong></em><em>或</em><em><strong>“允许一次”</strong></em><em>。</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6627ee0bff1862df/6a170c266f7f040f6f91488c/aea942ba9b0037526ea215bec65690f1a5c3099c-1522x250.png" alt="Claude Desktop 的 &quot;始终允许&quot; 和 &quot;允许一次&quot; 选项，供用户选择。" /><h2>结论</h2><p>MCP 服务器代表了本地和远程应用中 LLM 工具标准化的重要一步。虽然完全兼容仍在开发中，但我们正朝这个方向快速推进。</p><p>在本文中，我们学习了如何用 TypeScript 构建一个自定义 MCP 服务器，将 Elasticsearch 连接到基于 LLM 的应用。我们的服务器公开了两个工具：<code>search_docs</code> 用于使用查询 DSL 检索相关文档；<code>summarize_and_cite</code> 用于通过 OpenAI 模型和 Claude Desktop 作为客户端 UI 生成带引用的摘要。</p><p>不同客户端和服务器提供商之间的兼容性前景看起来一片光明。下一步包括为您的智能体添加更多功能和灵活性。这里有一篇实用的<a href="https://www.elastic.co/search-labs/blog/llm-functions-elasticsearch-intelligent-query">文章</a>介绍了如何使用搜索模板参数化查询，以获得精确性和灵活性。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude</guid>
    <category><![CDATA[智能体 AI]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5600198cb47666a5/6a170c28509168ce3ae1bb18/0bb24c05fff391f42070c2883182ea6fe9cb9680-1280x720.png" length="0" type="image/png"/>
    <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 Elasticsearch 推理 API 以及 Hugging Face 模型]]></title>
    <description><![CDATA[了解如何使用推理终端将 Elasticsearch 连接到 Hugging Face 模型，并利用语义搜索和聊天补全功能构建多语言博客推荐系统。]]></description>
    <content:encoded><![CDATA[<p>在最近的更新中，Elasticsearch 引入了原生集成，用于连接到托管在 <a href="https://endpoints.huggingface.co/">Hugging Face Inference Service</a> 上的模型。在本文中，我们将探讨如何配置此集成，并使用大型语言模型 (LLM) 通过简单的 API 调用执行推理。我们将使用 <a href="https://huggingface.co/HuggingFaceTB/SmolLM3-3B">SmolLM3-3B</a>，这是一款轻量级通用模型，在资源使用和答案质量之间取得了良好的平衡。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9094997548bd70f8/6a170d6a839dfa0ad6dcff54/7ddadf1976421a860a7d62087239adb9150d808b-1999x1388.png" alt="散点图显示了几个小型语言模型，其 X 轴表示模型大小（以十亿为单位的参数），Y 轴表示胜率（百分比）。SmolLM3-3B 在效率趋势中名列前茅，其胜率高于其他类似大小的模型。" /><h2>准备工作</h2><ul><li><p><strong>Elasticsearch 9.3 或 Elastic Cloud Serverless：</strong>您可以按照<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">这些说明</a>创建云部署，或者改用 <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart#local-dev-quick-start"><code>start-local</code></a> 快速入门。</p></li><li><p><strong>Python 3.12：</strong><a href="https://www.python.org/">在此处</a>下载 Python。</p></li><li><p><strong>Hugging Face </strong><a href="https://huggingface.co/docs/hub/en/security-tokens">访问令牌</a>。</p></li></ul><h2>使用 Hugging Face 推理终端完成聊天</h2><p>首先，我们将构建一个实用示例，将 Elasticsearch 连接到 Hugging Face <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put">推理终端</a>，以从博客文章集合中生成 AI 驱动的推荐。对于应用知识库，我们将使用公司博客文章数据集，其中包含有价值但通常难以查找的信息。</p><p>通过这个终端，<a href="https://www.elastic.co/docs/solutions/search/semantic-search">语义搜索</a>可以检索与给定查询最相关的文章，而 Hugging Face LLM 则会根据这些结果生成简短的上下文推荐。</p><p>让我们来看看我们将要构建的信息流的高级概述：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf217b7b7db4e1e6c/6a170d6ca929cf8022ae0a3b/1dfbc2323438feaaa42e13ab242dd1f7166f74aa-1200x676.png" alt="流程图展示了 Elasticsearch 索引将语义搜索结果输入到推理终端，该终端会返回文章推荐。" /><p>在本文中，我们将测试 <strong>SmolLM3-3B</strong> 是否能将其紧凑的大小与强大的多语言推理和工具调用能力相结合。根据搜索查询，我们将把所有匹配的内容（英语和西班牙语）发送到 LLM，以生成一份推荐文章列表，并根据搜索查询和结果提供自定义描述。</p><p>以下是具备 AI 推荐生成系统的文章网站用户界面可能的外观。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt20e69b9a06fecd65/6a170d6e839dfa6f97dcff58/8d3b86b212f28ff279f2da67a33e6134039f0e4e-1999x949.png" alt="具备 AI 推荐生成系统的文章网站的用户界面，列出了三个示例，文本为英语，标题为英语或西班牙语。" /><p>您可以在已链接的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/notebook.ipynb">笔记本</a>中找到此应用程序的完整实现。</p><h3>配置 Elasticsearch 推理终端</h3><p>要使用 Elasticsearch <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put-hugging-face">Hugging Face 推理终端</a>，我们需要两个重要元素：Hugging Face API 密钥和正在运行的 Hugging Face 终端 URL。它应该如下所示：</p>PUT _inference/chat_completions/hugging-face-smollm3-3b
{
    "service": "hugging_face",
    "service_settings": {
        "api_key": "hugging-face-access-token", 
        "url": "url-endpoint" 
    }
}<p>Elasticsearch 中的 Hugging Face 推理终端支持不同的任务类型：<code>text_embedding</code>、<code>completion</code>、<code>chat_completion</code> 和 <code>rerank</code>。在这篇博客文章中，我们使用 <code>chat_completion</code> 是因为我们需要模型根据搜索结果和系统提示生成对话式推荐。此终端允许我们使用 Elasticsearch API 以简单的方式直接从 Elasticsearch 执行聊天完成：</p>POST _inference/chat_completion/hugging-face-smollm3-3b/_stream
{
  "messages": [
      { "role": "user", "content": "&lt;user prompt&gt;" }
  ]
}<p>这将作为应用程序的核心，接收通过模型传递的提示和搜索结果。有了理论基础，我们就开始实施应用程序。</p><h4>在 Hugging Face 上设置推理终端</h4><p>要部署 Hugging Face 模型，我们将使用 <a href="https://huggingface.co/inference-endpoints/dedicated">Hugging Face 一键式部署</a>，这是一种用于部署模型终端的简单快速的服务。请记住，这是一项付费服务，使用它可能会产生额外费用。此步骤将创建用于生成文章推荐的模型实例。</p><p>您可以从一键目录中选择一个模型：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta7bdfa43d6766324/6a170d6fb339d59e5476a039/b816e9fba1fe172687bf58f5143fb1f838c1077f-549x331.png" alt="接口视图显示了一个已筛选为“smoll3”的模型目录，其中显示了一个名为“smollm3‑3b”的模型，具有文本生成、vLLM、GPU 1× NVIDIA L4，标价 0.8 美元，并附有一条建议将搜索范围扩展至所有 Hugging Face 模型的提示。" /><p>让我们选择 <strong>SmolLM3-3B</strong> 模型：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdb0a2e6ffd7deb20/6a170d710c48574b7401aafc/610d3aba0429f3666c2df3616d513eb6a4397c0c-502x478.png" alt="用于创建 SmolLM3‑3B 模型终端的接口，显示模型名称、“已由 Hugging Face 验证”注释、终端名称字段、每个运行副本每小时 0.80 美元的成本、cURL 选项和“创建终端”按钮。" /><p>从此处获取 Hugging Face 终端 URL：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt25714021711ed6ff/6a170d72c1e8a54853f88336/025094ddb2cfbd1f0f216a5ec4e119b0f4fa2c42-646x328.png" alt="名为“smollm3‑3b‑pnz”的 Hugging Face 推理终端的仪表板视图，显示绿色运行状态、一个活跃副本、过去一小时内的零请求、导航选项卡和显示的终端 URL。" /><p>正如在 Elasticsearch <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put-hugging-face">Hugging Face 推理终端文档</a>中提到的，文本生成需要一个与 OpenAI API 兼容的模型。因此，我们需要将 <code>/v1/chat/completions</code> 子路径附加到 Hugging Face 终端 URL。最终结果将如下所示：</p>https://j2g31h0futopfkli.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions<p>有了这个，我们就可以在 Python 笔记本中开始编码了。</p><h4>生成 Hugging Face API 密钥</h4><p>创建 <a href="https://huggingface.co/join">Hugging Face 账户</a>，并按照<a href="https://huggingface.co/docs/hub/en/security-tokens#user-access-tokens">以下说明</a>获取 API 令牌。您可以选择三种令牌类型：<em>细粒度</em>（推荐用于生产，因为它仅提供对特定资源的访问）、<em>读取</em>（适用于只读访问）或<em>写入</em>（适用于读取和写入访问）。在本教程中，读取令牌就足够了，因为我们只需要调用推理终端。请保存此密钥以备下一步使用。</p><h4>设置 Elasticsearch 推理终端</h4><p>首先，让我们声明一个 Elasticsearch Python 客户端：</p>os.environ["ELASTICSEARCH_API_KEY"] = "your-elasticsearch-api-key"
os.environ["ELASTICSEARCH_URL"] = "https://xxxx.us-central1.gcp.cloud.es.io:443"

es_client = Elasticsearch(
    os.environ["ELASTICSEARCH_URL"], api_key=os.environ["ELASTICSEARCH_API_KEY"]
)<p>接下来，我们创建一个使用 Hugging Face 模型的 Elasticsearch 推理终端。此终端将允许我们基于博客文章和传递给模型的提示来生成响应。</p>INFERENCE_ENDPOINT_ID = "smollm3-3b-pnz"

os.environ["HUGGING_FACE_INFERENCE_ENDPOINT_URL"] = (
 "https://j2g31h0futopfkli.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions"
)
os.environ["HUGGING_FACE_API_KEY"] = "hf_xxxxx"

resp = es_client.inference.put(
        task_type="chat_completion",
        inference_id=INFERENCE_ENDPOINT_ID,
        body={
            "service": "hugging_face",
            "service_settings": {
                "api_key": os.environ["HUGGING_FACE_API_KEY"],
                "url": os.environ["HUGGING_FACE_INFERENCE_ENDPOINT_URL"],
            },
        },
    )<h3>数据集</h3><p>该数据集包含将要查询的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/dataset.json">博客文章</a>，代表整个工作流中使用的多语言内容集：</p>// Articles dataset document example: 
{
    "id": "6",
    "title": "Complete guide to the new API: Endpoints and examples",
    "author": "Tomas Hernandez",
    "date": "2025-11-06",
    "category": "tutorial",
    "content": "This guide describes in detail all endpoints of the new API v2. It includes code examples in Python, JavaScript, and cURL for each endpoint. We cover authentication, resource creation, queries, updates, and deletion. We also explain error handling, rate limiting, and best practices. Complete documentation is available on our developer portal."
  }<h4>Elasticsearch 映射</h4><p>定义数据集后，我们需要创建一个适合博客文章结构的数据模式。以下<a href="https://www.elastic.co/docs/manage-data/data-store/mapping">索引映射</a>将用于在 Elasticsearch 中存储数据：</p>INDEX_NAME = "blog-posts"

mapping = {
    "mappings": {
        "properties": {
            "id": {"type": "keyword"},
            "title": {
                "type": "object",
                "properties": {
                    "original": {
                        "type": "text",
                        "copy_to": "semantic_field",
                        "fields": {"keyword": {"type": "keyword"}},
                    },
                    "translated_title": {
                        "type": "text",
                        "fields": {"keyword": {"type": "keyword"}},
                    },
                },
            },
            "author": {"type": "keyword", "copy_to": "semantic_field"},
            "category": {"type": "keyword", "copy_to": "semantic_field"},
            "content": {"type": "text", "copy_to": "semantic_field"},
            "date": {"type": "date"},
            "semantic_field": {"type": "semantic_text"},
        }
    }
}


es_client.indices.create(index=INDEX_NAME, body=mapping)<p>在这里，我们可以更清楚地看到数据的结构。我们将使用语义搜索来检索基于自然语言的结果，同时使用 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/copy-to"><code>copy_to</code></a> 属性将字段内容复制到 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_text</code></a> 字段中。此外，<code>title</code> 字段包含两个子字段：<code>original</code> 子字段根据文章的原始语言存储英语或西班牙语标题；而 <code>translated_title</code> 子字段仅存在于西班牙语文章中，并包含原始标题的英语翻译。</p><h3>采集数据</h3><p>以下代码片段使用<a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript/bulk_examples">批量 API</a> 将博客文章数据集摄取到 Elasticsearch 中：</p>def build_data(json_file, index_name):
    with open(json_file, "r") as f:
        data = json.load(f)

    for doc in data:
        action = {"_index": index_name, "_source": doc}
        yield action


try:
    success, failed = helpers.bulk(
        es_client,
        build_data("dataset.json", INDEX_NAME),
    )
    print(f"{success} documents indexed successfully")

    if failed:
        print(f"Errors: {failed}")
except Exception as e:
    print(f"Error: {str(e)}")<p>现在，我们已将文章摄取到 Elasticsearch 中，我们需要创建一个能够针对 <code>semantic_text</code> 字段进行搜索的函数：</p>def perform_semantic_search(query_text, index_name=INDEX_NAME, size=5):
    try:
        query = {
            "query": {
                "match": {
                    "semantic_field": {
                        "query": query_text,
                    }
                }
            },
            "size": size,
        }

        response = es_client.search(index=index_name, body=query)
        hits = response["hits"]["hits"]

        return hits
    except Exception as e:
        print(f"Semantic search error: {str(e)}")
        return []<p>我们还需要一个调用推理终端的函数。在这种情况下，我们将使用 <strong><code>chat_completion</code></strong>任务类型调用终端，以获取流式响应：</p>def stream_chat_completion(messages: list, inference_id: str = INFERENCE_ENDPOINT_ID):
    url = f"{ELASTICSEARCH_URL}/_inference/chat_completion/{inference_id}/_stream"
    payload = {"messages": messages}
    headers = {
        "Authorization": f"ApiKey {ELASTICSEARCH_API_KEY}",
        "Content-Type": "application/json",
    }

    try:
        response = requests.post(url, json=payload, headers=headers, stream=True)
        response.raise_for_status()

        for line in response.iter_lines(decode_unicode=True):
            if line:
                line = line.strip()

                if line.startswith("event:"):
                    continue

                if line.startswith("data: "):
                    data_content = line[6:]

                    if not data_content.strip() or data_content.strip() == "[DONE]":
                        continue

                    try:
                        chunk_data = json.loads(data_content)

                        if "choices" in chunk_data and len(chunk_data["choices"]) &gt; 0:
                            choice = chunk_data["choices"][0]
                            if "delta" in choice and "content" in choice["delta"]:
                                content = choice["delta"]["content"]
                                if content:
                                    yield content

                    except json.JSONDecodeError as json_err:
                        print(f"\nJSON decode error: {json_err}")
                        print(f"Problematic data: {data_content}")
                        continue

    except requests.exceptions.RequestException as e:
        yield f"Error: {str(e)}"<p>现在，我们可以编写一个函数，调用语义搜索函数以及 <code>chat_completions</code> 推理终端和建议终端，以生成将分配到卡片中的数据：</p>def recommend_articles(search_query, index_name=INDEX_NAME, max_articles=5):
    print(f"\n{'='*80}")
    print(f"🔍 Search Query: {search_query}")
    print(f"{'='*80}\n")

    articles = perform_semantic_search(search_query, index_name, size=max_articles)

    if not articles:
        print("❌ No relevant articles found.")
        return None, None

    print(f"✅ Found {len(articles)} relevant articles\n")

    # Build context with found articles
    context = "Available blog articles:\n\n"
    for i, article in enumerate(articles, 1):
        source = article.get("_source", article)
        context += f"Article {i}:\n"
        context += f"- Title: {source.get('title', 'N/A')}\n"
        context += f"- Author: {source.get('author', 'N/A')}\n"
        context += f"- Category: {source.get('category', 'N/A')}\n"
        context += f"- Date: {source.get('date', 'N/A')}\n"
        context += f"- Content: {source.get('content', 'N/A')}\n\n"

    system_prompt = """You are an expert content curator that recommends blog articles.

    Write recommendations in a conversational style starting with phrases like:
    - "If you're interested in [topic], this article..."
    - "This post complements your search with..."
    - "For those looking into [topic], this article provides..."


    FORMAT REQUIREMENTS:
    - Return ONLY a JSON array
    - Each element must have EXACTLY these three fields: "article_number", "title", "recommendation"
    - If the original title is in spanish, use the "translated_title" subfield in the "title" field

    Keep each recommendation concise (2-3 sentences max) and focused on VALUE to the reader.

    EXAMPLE OF CORRECT FORMAT:
    [
        {"article_number": 1, "title": "Article title in english", "recommendation": "If you are interested in [topic], this article provides..."},
        {"article_number": 2, "title": "Article title in english", "recommendation": " for those looking into [topic], this article provides..."}
    ]

    Return ONLY the JSON array following this exact structure."""

    user_prompt = f"""Search query: "{search_query}"

    Generate recommendations for the following articles: {context}
    """

    messages = [
        {"role": "system", "content": "/no_think"},
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_prompt},
    ]

    # LLM generation
    print(f"{'='*80}")
    print("🤖 Generating personalized recommendations...\n")

    full_response = ""

    for chunk in stream_chat_completion(messages):
        print(chunk, end="", flush=True)
        full_response += chunk

    return context, articles, full_response<p>最后，我们需要提取信息并将其格式化以便打印：</p>def display_recommendation_cards(articles, recommendations_text):
    print("\n" + "=" * 100)
    print("📇 RECOMMENDED ARTICLES".center(100))
    print("=" * 100 + "\n")

    # Parse JSON recommendations - clean tags and extract JSON
    recommendations_list = []
    try:

        # Clean up &lt;think&gt; tags
        cleaned_text = re.sub(
            r"&lt;think&gt;.*?&lt;/think&gt;", "", recommendations_text, flags=re.DOTALL
        )
        # Remove markdown code blocks ( ... ``` or ``` ... ```)
        cleaned_text = re.sub(r"```(?:json)?", "", cleaned_text)
        cleaned_text = cleaned_text.strip()

        parsed = json.loads(cleaned_text)

        # Extract recommendations from list format
        for item in parsed:
            article_number = item.get("article_number")
            title = item.get("title", "")
            rec_text = item.get("recommendation", "")

            if article_number and rec_text:
                recommendations_list.append(
                    {
                        "article_number": article_number,
                        "title": title,
                        "recommendation": rec_text,
                    }
                )
    except json.JSONDecodeError as e:
        print(f"⚠️  Could not parse recommendations as JSON: {e}")
        return

    for i, article in enumerate(articles, 1):
        source = article.get("_source", article)

        # Card border
        print("┌" + "─" * 98 + "┐")

        # Find recommendation and title for this article number
        recommendation = None
        title = None
        for rec in recommendations_list:
            if rec.get("article_number") == i:
                recommendation = rec.get("recommendation")
                title = rec.get("title")
                break

        # Print title
        title_lines = textwrap.wrap(f"📌 {title}", width=94)
        for line in title_lines:
            print(f"│  {line}".ljust(99) + "│")

        # Card border
        print("├" + "─" * 98 + "┤")

        # Print recommendation
        if recommendation:
            recommendation_lines = textwrap.wrap(recommendation, width=94)
            for line in recommendation_lines:
                print(f"│  {line}".ljust(99) + "│")

        # Card bottom
        print("└" + "─" * 98 + "┘")<p>让我们通过询问一个有关安全博客文章的问题来测试一下：</p>search_query = "Security and vulnerabilities"

context, articles, recommendations = recommend_articles(search_query)

print("\nElasticsearch context:\n", context)

# Display visual cards
display_recommendation_cards(articles, recommendations)<p>如下所示，我们可以看到工作流在控制台中生成的卡片：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4aa221a08a51aeb3/6a170d7460084be1413c45d6/730d35212594bb3db30447c3ea7e2a92857287b7-1999x1515.png" alt="标题为“推荐文章”的部分显示了五篇方框式文章摘要，包括身份验证系统漏洞、迁移风险、REST API v2 性能和身份验证改进、通知系统更改以及新 API 的完整指南等主题。" /><p>您可以在<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/results.md">此文件</a>中查看全部结果，包括所有点击和 LLM 响应。</p><p>我们正在征集与“安全与漏洞”相关的文章。此问题将用作针对 Elasticsearch 中存储的文档的搜索查询。然后将检索到的结果传递给模型，该模型根据这些结果的内容生成推荐。我们可以看到，该模型出色地生成了引人入胜的短文本，能够激发读者点击的欲望。</p><h2>结论</h2><p>本示例展示了如何将 Elasticsearch 和 Hugging Face 结合起来，为 AI 应用程序创建一个快速高效的集中式系统。由于 Hugging Face 拥有丰富的模型目录，这种方法不仅减少了人工操作，还具有灵活性。通过使用 SmolLM3-3B，我们特别看到了紧凑的多语言模型在与语义搜索搭配使用时，仍能提供有意义的推理和内容生成。这些工具共同为构建智能内容分析和多语言应用程序提供了可扩展且高效的基础。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/hugging-face-elasticsearch-inference-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/hugging-face-elasticsearch-inference-api</guid>
    <category><![CDATA[智能体 AI]]></category>
    <category><![CDATA[集成]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f961af4cb26ec97/6a170d767d8d6790c770e790/1417d6ff033712206c9bd4bcc22074ee3437ce96-1999x1125.png" length="0" type="image/png"/>
    <pubDate>Mon, 23 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 LangGraph.js 和 Elasticsearch 构建金融 AI 搜索工作流。]]></title>
    <description><![CDATA[学习如何将 LangGraph.js 与 Elasticsearch 结合使用，构建一个 AI 驱动的金融搜索工作流，将自然语言查询转换为动态的条件过滤器，用于投资和市场分析。]]></description>
    <content:encoded><![CDATA[<p>构建 AI 搜索应用并不轻松：多重任务、数据拉取与抽取都需要紧密配合，才能形成流畅连贯的工作流。LangGraph 通过节点式结构让开发者轻松编排 AI 代理，从而大幅简化了整个流程。在本文中，我们将运用 <a href="https://langchain-ai.github.io/langgraphjs/">LangGraph.js</a> 构建一个面向金融场景的 AI 搜索解决方案。</p><h2>什么是 LangGraph</h2><p><a href="https://langchain-ai.github.io/langgraphjs/">LangGraph</a> 是一个用于构建 AI 代理，并将其编排进工作流，从而打造 AI 辅助应用的框架。LangGraph 采用节点式架构，我们可以声明代表不同任务的函数，并将这些函数指定为工作流中的节点。多个节点相互作用后形成的便是一个图结构。LangGraph 是更广泛的 <a href="https://js.langchain.com/docs/introduction/">LangChain</a> 生态系统的一部分，该生态为构建模块化、可组合的 AI 系统提供了丰富的工具。</p><p>为了更直观地理解 LangGraph 有何用处，我们不妨用它来解决一个真实的业务难题。</p><h2>解决方案概述</h2><p>在一家风险投资公司中，投资人可以访问一个带有大量筛选条件的大型数据库，但一旦需要组合多重条件，查询就会变得既繁琐又缓慢。这可能会导致一些本应纳入投资视野的优质初创公司被漏掉。结果就是，团队要耗费大量时间去筛选最佳标的，甚至因此错失投资机会。</p><p>借助 LangGraph 和 Elasticsearch，我们能够使用自然语言进行过滤搜索，从而无需用户手动构建包含数十个筛选器的复杂请求。为了提高灵活性，工作流会根据用户输入在两种查询类型之间自动选择：</p><ul><li><p><strong>聚焦投资维度的查询</strong>：这类查询专注于初创公司的财务与融资维度，例如<a href="https://www.investopedia.com/articles/personal-finance/102015/series-b-c-funding-what-it-all-means-and-how-it-works.asp">融资轮次</a>、估值或<a href="https://www.investopedia.com/terms/r/revenue.asp">营收</a>等指标。<em>示例：</em>“查找已完成 A 轮或 B 轮融资、融资额在 800 万至 2,500 万美元之间且月收入超过 50 万美元的初创公司。”</p></li><li><p><strong>聚焦市场维度的查询</strong>：这类查询侧重于<a href="https://en.wikipedia.org/wiki/Vertical_market">行业垂直领域</a>、<a href="https://en.wikipedia.org/wiki/Target_market">目标市场</a>或<a href="https://www.investopedia.com/terms/b/businessmodel.asp">商业模式</a>，帮助识别特定领域或地区中的投资机会。<em>示例：</em>“查找位于旧金山、纽约或波士顿的金融科技和医疗健康领域初创公司。”</p></li></ul><p>为了让查询更稳健，我们会让 LLM 生成<a href="https://www.elastic.co/docs/solutions/search/search-templates">搜索模板</a>，而不是直接构造完整的 <a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/querydsl">DSL 查询</a>。通过这种方式，你获得的始终是预期的查询结果，LLM 只需填入参数，而不必每次从头构建整条查询。</p><h2>开始前的准备工作</h2><ul><li><p>Elasticsearch API密钥</p></li><li><p>OpenAPI API密钥</p></li><li><p>Node 18 或更高版本</p></li></ul><h2>分步操作指南</h2><p>在本节中，我们先来看一下这个应用的外观。为此，我们将使用 <a href="https://www.typescriptlang.org/">TypeScript</a>，这是 JavaScript 的一个超集，添加了静态类型，使代码更可靠且更易维护，并能更早发现错误，同时又与现有 JavaScript 完全兼容。</p><p>节点的流程将如下所示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt90db8f03f372608c/6a170986dc55de6e16e00d93/b47d7f238c4964a6febc0de7fe5e68b186f539c3-363x555.png" alt="" /><p>上图由 LangGraph 生成，直观地呈现了工作流结构，包括各节点的执行顺序和它们之间的条件分支关系：</p><ul><li><p><strong>decideStrategy：</strong>使用 LLM 分析用户查询，在“聚焦投资维度”与“聚焦市场维度”这两种专门搜索策略之间做出选择。</p></li><li><p><strong>PrepareInvestmentSearch：</strong>从查询中提取筛选值并构建一个强调财务和资金相关参数的预定义模板。</p></li><li><p><strong>prepareMarketSearch</strong>：同样会提取筛选条件，但重点是围绕市场、行业和地域背景，动态生成相应的搜索参数。</p></li><li><p><strong>ExecuteSearch：</strong>通过搜索模板将构建好的查询发送到 Elasticsearch，检索并返回所有匹配的初创公司文档。</p></li><li><p><strong>visualizeResults：</strong>将最终结果整理成清晰易读的摘要，呈现融资、行业、营收等关键创业公司属性。</p></li></ul><p>该流程包括一个<a href="https://langchain-ai.github.io/langgraphjs/how-tos/branching/?h=conditional#how-to-create-branches-for-parallel-node-execution">条件分支</a>，相当于一条“if”语句，可根据用户输入决定使用投资还是市场搜索路径。这种由 LLM 驱动的决策机制让工作流具备自适应和上下文感知能力，后续章节将对这一机制进行更详细的说明。</p><h3>LangGraph 状态</h3><p>在查看各个节点之前，我们需要先理解节点之间的通信和数据共享方式。为此，LangGraph 可支持定义工作流状态。这个状态就是在各个节点之间传递的共享状态。</p><p>该状态相当于一个共享容器，在整个工作流中保存中间数据：从最开始的用户自然语言查询，到选定的搜索策略、为 Elasticsearch 准备好的参数、检索到的搜索结果，一直到最后的格式化输出，都会依次写入其中。</p><p>这种结构让每个节点都能读取和更新状态，确保从用户输入到可视化实现顺畅一致的信息流动。</p>const VCState = Annotation.Root({
  input: Annotation&lt;string&gt;(), // User's natural language query
  searchStrategy: Annotation&lt;string&gt;(), // Search strategy chosen by LLM
  searchParams: Annotation&lt;any&gt;(), // Prepared search parameters
  results: Annotation&lt;any[]&gt;(), // Search results
  final: Annotation&lt;string&gt;(), // Final formatted response
});<h3>设置应用程序</h3><p>本节所有代码均可在 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch">elasticsearch-labs 仓库</a> 中找到。</p><p>在应用所在的文件夹中打开终端，并通过以下命令初始化一个 Node.js 应用：</p>npm init -y<p>现在我们可以为这个项目安装必要的依赖项：</p>npm install @elastic/elasticsearch @langchain/langgraph @langchain/openai @langchain/core dotenv zod &amp;&amp; npm install --save-dev @types/node tsx typescript<ul><li><p><strong><code>@elastic/elasticsearch</code></strong>：帮助我们处理 Elasticsearch 请求，例如数据摄取和检索。</p></li><li><p><strong><code>@langchain/langgraph</code></strong>：用于提供所有 LangGraph 工具的 JS 依赖项。</p></li><li><p><strong><code>@langchain/openai</code></strong>：适用于 LangChain 的 OpenAI LLM 客户端。</p></li><li><p>@langchain/core：为 LangChain 应用提供基础构建模块，包括提示模板。</p></li><li><p><strong><code>dotenv</code></strong>：在 JavaScript 中使用环境变量所需的依赖项。</p></li><li><p><strong><code>zod</code></strong>: 对类型数据的依赖。</p></li></ul><p><code>@types/node</code> <code>tsx</code> <code>typescript</code> 允许我们编写和运行 TypeScript 代码。</p><p>现在创建以下文件：</p><ul><li><p><code>elasticsearchSetup</code><a href="http://ingest.ts/"><code>.ts</code></a>：将创建索引映射，从 JSON 文件加载数据集，并将数据摄取到 Elasticsearch。</p></li><li><p><a href="http://main.ts/"><code>main.ts</code></a>：将包含 LangGraph 应用。</p></li><li><p><code>.env</code>：用于存储环境变量的文件</p></li></ul><p>在 <code>.env</code> 文件中，我们添加以下环境变量：</p>ELASTICSEARCH_ENDPOINT="your-endpoint-here"
ELASTICSEARCH_API_KEY="your-key-here"
OPENAI_API_KEY="your-key-here"<p>OpenAPI APIKey 不会直接在代码中使用，而是由 <code>@langchain/openai</code> 库在内部调用。</p><p>所有关于映射创建、搜索模板创建和数据集摄取的逻辑都可以在 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts"><code>elasticsearchSetup.ts</code></a> 文件中找到。在接下来的步骤中，我们将重点关注 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/main.ts"><code>main.ts</code></a> 文件。此外，您可以查看该数据集，以便更好地理解 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/dataset.json"><code>dataset.json</code></a> 中数据结构。</p><h3>LangGraph 应用程序</h3><p>在 <code>main.ts</code> 文件中，我们导入一些必要的依赖项来构建整个 LangGraph 应用。在此文件中，您还必须定义各个节点函数以及工作流状态的声明。在后续步骤中，我们会在 <code>main</code> 方法中完成这个图结构的声明。<code>elasticsearchSetup.ts</code> 文件中包含一组 Elasticsearch 辅助函数，我们会在后续步骤的各个节点中使用这些函数。</p>import { writeFileSync } from "node:fs";
import { StateGraph, Annotation, START, END } from "@langchain/langgraph";
import { ChatOpenAI } from "@langchain/openai";
import { z } from "zod";
import {
  esClient,
  ingestDocuments,
  createSearchTemplates,
  INDEX_NAME,
  INVESTMENT_FOCUSED_TEMPLATE,
  MARKET_FOCUSED_TEMPLATE,
  createIndex,
} from "./elasticsearchSetup.js";

const llm = new ChatOpenAI({ model: "gpt-4o-mini" });<p>如前所述，LLM 客户端将根据用户的问题生成 Elasticsearch 搜索模板参数。</p>async function saveGraphImage(app: any): Promise&lt;void&gt; {
  try {
    const drawableGraph = app.getGraph();
    const image = await drawableGraph.drawMermaidPng();
    const arrayBuffer = await image.arrayBuffer();

    const filePath = "./workflow_graph.png";
    writeFileSync(filePath, new Uint8Array(arrayBuffer));
    console.log(`📊 Workflow graph saved as: ${filePath}`);
  } catch (error: any) {
    console.log("⚠️  Could not save graph image:", error.message);
  }
}<p>上面的方法会生成一张 png 格式的图结构图像，并在后台调用 <a href="https://mermaid.ink/">Mermaid.ink API</a>。当你希望通过一张带有样式的可视化图来直观了解应用中各个节点之间的交互时，这个功能就会非常有用。</p><h3>LangGraph 节点</h3><p>现在让我们看看每个节点的详细信息：</p><h3>decideSearchStrategy 节点</h3><p><code>decideSearchStrategy</code> 节点分析用户输入，并确定是执行投资聚焦搜索还是市场聚焦搜索。它使用具有结构化输出模式（用 Zod 定义）的 LLM 对查询类型进行分类。在做出决策之前，它会通过聚合从索引中检索可用的筛选条件，确保模型掌握最新的行业、地域和融资等上下文信息。</p><p>为了提取过滤器可能的值并将其发送到 LLM，让我们使用<a href="https://www.elastic.co/docs/explore-analyze/query-filter/aggregations">聚合</a>查询直接从 Elasticsearch 索引中检索它们。这个逻辑被分配到一个名为 <code>getAvailableFilters</code> 的方法中：</p>async function getAvailableFilters() {
  try {
    const response = await esClient.search({
      index: INDEX_NAME,
      size: 0,
      aggs: {
        industries: {
          terms: { field: "industry", size: 100 },
        },
        locations: {
          terms: { field: "location", size: 100 },
        },
        funding_stages: {
          terms: { field: "funding_stage", size: 20 },
        },
        business_models: {
          terms: { field: "business_model", size: 10 },
        },
        lead_investors: {
          terms: { field: "lead_investor", size: 100 },
        },
        funding_amount_stats: {
          stats: { field: "funding_amount" },
        },
      },
    });

    return response.aggregations;
  } catch (error) {
    console.error("❌ Error getting available filters:", error);
    return {};
  }
}<p>通过上述聚合查询，我们得到以下结果：</p>{
  "industries": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "logistics",
        "doc_count": 5
      },
      ...
    ]
  },
  "locations": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "San Francisco, CA",
        "doc_count": 4
      },
      {
        "key": "New York, NY",
        "doc_count": 3
      },
      ...
    ]
  },
  "funding_stages": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "Series A",
        "doc_count": 8
      },
      ...
    ]
  },
  "business_models": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "B2B",
        "doc_count": 13
      },
      ...
    ]
  },
  "lead_investors": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "Battery Ventures",
        "doc_count": 1
      },
      {
        "key": "Benchmark Capital",
        "doc_count": 1
      },
      ...
    ]
  },
  "funding_amount_stats": {
    "count": 20,
    "min": 4500000,
    "max": 35000000,
    "avg": 14075000,
    "sum": 281500000
  }
}<p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/responses/aggregationsResponse.json">点击此处</a>查看所有结果。</p><p>对于这两种策略，我们将使用混合搜索来检测问题的结构化部分（过滤器）和主观部分（语义）。以下是使用<a href="https://www.elastic.co/docs/solutions/search/search-templates">搜索模板</a>的两个查询示例：</p>await esClient.putScript({
      id: INVESTMENT_FOCUSED_TEMPLATE,
      script: {
        lang: "mustache",
        source: `{
          "size": 5,
          "retriever": {
            "rrf": {
              "retrievers": [
                {
                  "standard": {
                    "query": {
                      "semantic": {
                        "field": "semantic_field",
                        "query": "{{query_text}}"
                      }
                    }
                  }
                },
                {
                  "standard": {
                    "query": {
                      "bool": {
                        "filter": [
                          {"terms": {"funding_stage": {{#join}}{{#toJson}}funding_stage{{/toJson}}{{/join}}}},
                          {"range": {"funding_amount": {"gte": {{funding_amount_gte}}{{#funding_amount_lte}},"lte": {{funding_amount_lte}}{{/funding_amount_lte}}}}},
                          {"terms": {"lead_investor": {{#join}}{{#toJson}}lead_investor{{/toJson}}{{/join}}}},
                          {"range": {"monthly_revenue": {"gte": {{monthly_revenue_gte}}{{#monthly_revenue_lte}},"lte": {{monthly_revenue_lte}}{{/monthly_revenue_lte}}}}}
                        ]
                      }
                    }
                  }
                }
              ],
              "rank_window_size": 100,
              "rank_constant": 20
            }
          }
        }`,
      },
    });<p>查看 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts#L119"><code>elasticsearchSetup.ts</code></a> 文件中详细的查询。在接下来的节点中，将决定使用这两个查询中的哪一个：</p>// Node 1: Decide search strategy using LLM
async function decideSearchStrategy(state: typeof VCState.State) {
  // Zod schema for specialized search strategy decision
  const SearchDecisionSchema = z.object({
    search_type: z
      .enum(["investment_focused", "market_focused"])
      .describe("Type of specialized search strategy to use"),
    reasoning: z
      .string()
      .describe("Brief explanation of why this search strategy was chosen"),
  });

  const decisionLLM = llm.withStructuredOutput(SearchDecisionSchema);

  // Get dynamic filters from Elasticsearch
  const availableFilters = await getAvailableFilters();

  const prompt = `Query: "${state.input}"
    Available filters: ${JSON.stringify(availableFilters, null, 2)}

    Choose between two specialized search strategies:
    
    - investment_focused: For queries about funding stages, funding amounts, monthly revenue, lead investors, financial performance
    
    - market_focused: For queries about industries, locations, business models, market segments, geographic markets
    
    Analyze the query intent and choose the most appropriate strategy.
  `;

  try {
    const result = await decisionLLM.invoke(prompt);
    console.log(
      `🤔 Search strategy: ${result.search_type} - ${result.reasoning}`
    );

    return {
      searchStrategy: result.search_type,
    };
  } catch (error: any) {
    console.error("❌ Error in decideSearchStrategy:", error.message);
    return {
      searchStrategy: "investment_focused",
    };
  }
}<h3>prepareInvestmentSearch 和 prepareMarketSearch 节点</h3><p>两个节点都使用共享的辅助函数 <code>extractFilterValues</code>，该函数利用 LLM 来识别用户输入中提到的相关过滤器，例如行业、地点、资金阶段、商业模式等。我们正在使用这个架构来构建我们的 <a href="https://www.elastic.co/docs/solutions/search/search-templates">搜索模板</a>。</p>// Extract all possible filter values from user input
async function extractFilterValues(input: string) {
  const FilterValuesSchema = z.object({
    // Investment-focused filters
    funding_stage: z
      .array(z.string())
      .default([])
      .describe("Funding stage values mentioned in query"),
    funding_amount_gte: z
      .number()
      .default(0)
      .describe("Minimum funding amount in USD"),
    funding_amount_lte: z
      .number()
      .default(100000000)
      .describe("Maximum funding amount in USD"),
    lead_investor: z
      .array(z.string())
      .default([])
      .describe("Lead investor values mentioned in query"),
    monthly_revenue_gte: z
      .number()
      .default(0)
      .describe("Minimum monthly revenue in USD"),
    monthly_revenue_lte: z
      .number()
      .default(10000000)
      .describe("Maximum monthly revenue in USD"),
    industry: z
      .array(z.string())
      .default([])
      .describe("Industry values mentioned in query"),
    location: z
      .array(z.string())
      .default([])
      .describe("Location values mentioned in query"),
    business_model: z
      .array(z.string())
      .default([])
      .describe("Business model values mentioned in query"),
  });

  const extractorLLM = llm.withStructuredOutput(FilterValuesSchema);
  const availableFilters = await getAvailableFilters();

  const extractPrompt = `Extract ALL relevant filter values from: "${input}"
    Available options: ${JSON.stringify(availableFilters, null, 2)}
    Extract only values explicitly mentioned in the query. Leave fields empty if not mentioned.`;

  return await extractorLLM.invoke(extractPrompt);
}<p>根据检测到的意图，工作流会选择以下两种路径之一：</p><p><strong>PrepareInvestmentSearch：</strong>构建以财务为导向的搜索参数，包括融资阶段、融资金额、投资者以及营收相关信息。您可以在 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts"><code>elasticsearchSetup.ts</code></a> 文件中找到整个查询模板：</p>// Node 2A: Prepare Investment-Focused Search Parameters 
async function prepareInvestmentSearch(state: typeof VCState.State) {
  console.log(
    "💰 Preparing INVESTMENT-FOCUSED search parameters with financial emphasis..."
  );

  try {
    // Extract all filter values from input
    const values = await extractFilterValues(state.input);

    let searchParams: any = {
      template_id: INVESTMENT_FOCUSED_TEMPLATE,
      query_text: state.input,
      ...values,
    };

    return { searchParams };
  } catch (error) {
    console.error("❌ Error preparing investment-focused params:", error);
    return {
      searchParams: {},
    };
  }
}<p><strong>prepareMarketSearch：</strong>创建以行业、地域和商业模式为重点的市场驱动参数。在 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts"><code>elasticsearchSetup.ts</code></a> 文件中查看完整查询：</p>// Node 2B: Prepare Market-Focused Search Parameters
async function prepareMarketSearch(state: typeof VCState.State) {
  console.log(
    "🔍 Preparing MARKET-FOCUSED search parameters with market emphasis..."
  );

  try {
    // Extract all filter values from input
    const values = await extractFilterValues(state.input);

    let searchParams: any = {
      template_id: MARKET_FOCUSED_TEMPLATE,
      query_text: state.input,
      ...values,
    };

    return { searchParams };
  } catch (error) {
    console.error("❌ Error preparing market-focused params:", error);
    return {};
  }
}<h3>executeSearch 节点</h3><p>该节点从状态中获取生成的搜索参数，首先将其发送到 Elasticsearch，使用<a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-render-search-template">_render API</a>来可视化查询以便调试，然后发送请求以检索结果。</p>// Node 3: Execute Search
async function executeSearch(state: typeof VCState.State) {
  const { searchParams } = state;

  try {
    // getting formed query from template for debugging
    const renderedTemplate = await esClient.renderSearchTemplate({
      id: searchParams.template_id,
      params: searchParams,
    });

    console.log(
      "📋 Complete query:",
      JSON.stringify(renderedTemplate.template_output, null, 2)
    );

    const results = await esClient.searchTemplate({
      index: INDEX_NAME,
      id: searchParams.template_id,
      params: searchParams,
    });

    return {
      results: results.hits.hits.map((hit: any) =&gt; hit._source),
    };
  } catch (error: any) {
    console.error(`❌ ${state.searchParams.search_type} search error:`, error);
    return { results: [] };
  }
}<h3>visualizeResults 节点</h3><p>最后，此节点显示 Elasticsearch 结果。</p>// Node 4: Visualize results
async function visualizeResults(state: typeof VCState.State) {
  const results = state.results || [];

  let formattedResults = `🎯 Found ${results.length} startups matching your criteria:\n\n`;

  results.forEach((startup: any, index: number) =&gt; {
    formattedResults += `${index + 1}. **${startup.company_name}**\n`;
    formattedResults += `   📍 ${startup.location} | 🏢 ${startup.industry} | 💼 ${startup.business_model}\n`;
    formattedResults += `   💰 ${startup.funding_stage} - $${(
      startup.funding_amount / 1000000
    ).toFixed(1)}M\n`;
    formattedResults += `   👥 ${startup.employee_count} employees | 📈 $${(
      startup.monthly_revenue / 1000
    ).toFixed(0)}K MRR\n`;
    formattedResults += `   🏦 Lead: ${startup.lead_investor}\n`;
    formattedResults += `   📝 ${startup.description}\n\n`;
  });

  return {
    final: formattedResults,
  };
}<p>从程序角度来看，整个图结构如下所示：</p>  const workflow = new StateGraph(VCState)
    // Register nodes - these are the processing functions
    .addNode("decideStrategy", decideSearchStrategy)
    .addNode("prepareInvestment", prepareInvestmentSearch)
    .addNode("prepareMarket", prepareMarketSearch)
    .addNode("executeSearch", executeSearch)
    .addNode("visualizeResults", visualizeResults)
    // Define execution flow with conditional branching
    .addEdge(START, "decideStrategy") // Start with strategy decision
    .addConditionalEdges(
      "decideStrategy",
      (state: typeof VCState.State) =&gt; state.searchStrategy, // Conditional function
      {
        investment_focused: "prepareInvestment", // If investment focused -&gt; RRF template preparation
        market_focused: "prepareMarket", // If market focused -&gt; dynamic query preparation
      }
    )
    .addEdge("prepareInvestment", "executeSearch") // Investment prep -&gt; execute
    .addEdge("prepareMarket", "executeSearch") // Market prep -&gt; execute
    .addEdge("executeSearch", "visualizeResults") // Execute -&gt; visualize
    .addEdge("visualizeResults", END); // End workflow<p>正如你所见，我们有一个条件边，应用在此决定接下来运行哪个“路径”或节点。当工作流需要分支逻辑时，例如在多个工具之间进行选择或包含人机交互步骤，此功能非常有用。</p><p>了解了 LangGraph 的核心功能后，我们可以设置代码运行的应用程序：</p><p>将所有内容在 <code>main</code> 方法中整合起来，在名为 workflow 的变量中声明这个包含所有元素的图结构：</p>async function main() {
  await createIndex();
  await createSearchTemplates();
  await ingestDocuments();

  // Create the workflow graph with shared state
  const workflow = new StateGraph(VCState)
    // Register nodes - these are the processing functions
    .addNode("decideStrategy", decideSearchStrategy)
    .addNode("prepareInvestment", prepareInvestmentSearch)
    .addNode("prepareMarket", prepareMarketSearch)
    .addNode("executeSearch", executeSearch)
    .addNode("visualizeResults", visualizeResults)
    // Define execution flow with conditional branching
    .addEdge(START, "decideStrategy") // Start with strategy decision
    .addConditionalEdges(
      "decideStrategy",
      (state: typeof VCState.State) =&gt; state.searchStrategy, // Conditional function
      {
        investment_focused: "prepareInvestment", // If investment focused -&gt; RRF template preparation
        market_focused: "prepareMarket", // If market focused -&gt; dynamic query preparation
      }
    )
    .addEdge("prepareInvestment", "executeSearch") // Investment prep -&gt; execute
    .addEdge("prepareMarket", "executeSearch") // Market prep -&gt; execute
    .addEdge("executeSearch", "visualizeResults") // Execute -&gt; visualize
    .addEdge("visualizeResults", END); // End workflow


  const app = workflow.compile();

  await saveGraphImage(app);

  const query =
    "Find startups with Series A or Series B funding between $8M-$25M and monthly revenue above $500K";

  const marketResult = await app.invoke({ input: query });
  console.log(marketResult.final);
}<p>查询变量用来模拟用户在一个虚拟搜索框中输入的内容：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltba7189d5f4e63403/6a1709880e2e49cc3041a076/e8d76909eb2bc1bb62f3ca9a8b3e4b85fcec2893-1600x164.png" alt="" /><p>系统会从这句自然语言“查找已完成 A 轮或 B 轮、融资额在 800 万至 2,500 万美元之间且月收入高于 50 万美元的初创公司”中，自动抽取出所有筛选条件。</p><p>最后，调用主方法：</p>main().catch(console.error);<h3>实施结果</h3>🔍 Checking if index exists...
🏗️ Creating index...
✅ Index created successfully!
Ingesting documents...
✅ Documents ingested successfully!
✅ Investment-focused template created successfully!
✅ Market-focused template created successfully!

📊 Workflow graph saved as: ./workflow_graph.png

🔍 Query: "Find startups with Series A or Series B funding between $8M-$25M and monthly revenue above $500K"

🤔 Search strategy: investment_focused - The query specifically seeks profitable fintech startups with defined funding amounts and high monthly revenue, which aligns closely with financial performance metrics and investment-related criteria.

💰 Preparing INVESTMENT-FOCUSED search parameters with financial emphasis...

📋 Complete query: {
  "size": 5,
  "retriever": {
    "rrf": {
      "retrievers": [
        {
          "standard": {
            "query": {
              "semantic": {
                "field": "semantic_field",
                "query": "Find startups with Series A or Series B funding between $8M-$25M and monthly revenue above $500K"
              }
            }
          }
        },
        {
          "standard": {
            "query": {
              "bool": {
                "filter": [
                  {
                    "terms": {
                      "funding_stage": [
                        "Series A",
                        "Series B"
                      ]
                    }
                  },
                  {
                    "range": {
                      "funding_amount": {
                        "gte": 8000000,
                        "lte": 25000000
                      }
                    }
                  },
                  {
                    "terms": {
                      "lead_investor": []
                    }
                  },
                  {
                    "range": {
                      "monthly_revenue": {
                        "gte": 500000,
                        "lte": 0
                      }
                    }
                  }
                ]
              }
            }
          }
        }
      ],
      "rank_window_size": 100,
      "rank_constant": 20
    }
  }
}
🎯 Found 5 startups matching your criteria:

1. **TechFlow**
   📍 San Francisco, CA | 🏢 logistics | 💼 B2B
   💰 Series A - $8.0M
   👥 45 employees | 📈 $500K MRR
   🏦 Lead: Sequoia Capital
   📝 TechFlow optimizes supply chain operations using AI-powered route optimization and real-time tracking. Founded in 2023, shows remarkable growth with $500K monthly revenue.

2. **DataViz**
   📍 New York, NY | 🏢 enterprise software | 💼 B2B
   💰 Series A - $10.0M
   👥 42 employees | 📈 $450K MRR
   🏦 Lead: Battery Ventures
   📝 DataViz creates intuitive data visualization tools for enterprise customers. No-code platform allows business users to create dashboards without technical expertise.

3. **FinanceAI**
   📍 San Francisco, CA | 🏢 fintech | 💼 B2C
   💰 Series C - $25.0M
   👥 120 employees | 📈 $1200K MRR
   🏦 Lead: Tiger Global Management
   📝 FinanceAI provides AI-powered investment advisory services to retail investors. Uses machine learning to analyze market trends with over 100,000 active users.

4. **UrbanMobility**
   📍 New York, NY | 🏢 logistics | 💼 B2B2C
   💰 Series B - $15.0M
   👥 78 employees | 📈 $750K MRR
   🏦 Lead: Kleiner Perkins
   📝 UrbanMobility revolutionizes urban transportation through autonomous delivery drones and smart logistics hubs. Partners with major retailers for same-day delivery across Manhattan and Brooklyn.

5. **HealthTech Solutions**
   📍 Boston, MA | 🏢 healthcare | 💼 B2B
   💰 Series B - $18.0M
   👥 95 employees | 📈 $900K MRR
   🏦 Lead: General Catalyst
   📝 HealthTech Solutions develops medical devices and software for remote patient monitoring. Comprehensive telehealth platform reducing hospital readmissions by 30%.

✨  Done in 18.80s.<p>对于这条输入，应用会选择<strong>聚焦投资维度</strong>的路径，由此我们可以看到 LangGraph 工作流生成的 Elasticsearch 查询，它会从用户输入中抽取出各类数值与区间。此外，我们还能看到应用了这些提取参数后实际发送到 Elasticsearch 的查询，以及最后由 <code>visualizeResults</code> 节点格式化输出的结果。</p><p>现在，我们再用这条查询来测试<strong>聚焦市场维度</strong>的节点：“查找位于旧金山、纽约或波士顿的金融科技和医疗健康初创公司”：</p>...

🔍 Query: Find fintech and healthcare startups in San Francisco, New York, or Boston

🤔 Search strategy: market_focused - The query is focused on finding fintech startups in San Francisco that are disrupting traditional banking and payment systems, which pertains to specific industries (fintech) and locations (San Francisco). Thus, a market-focused strategy is more appropriate.

🔍 Preparing MARKET-FOCUSED search parameters with market emphasis...

📋 Complete query: {
  "size": 5,
  "retriever": {
    "rrf": {
      "retrievers": [
        {
          "standard": {
            "query": {
              "semantic": {
                "field": "semantic_field",
                "query": "Find fintech and healthcare startups in San Francisco, New York, or Boston"
              }
            }
          }
        },
        {
          "standard": {
            "query": {
              "bool": {
                "filter": [
                  {
                    "terms": {
                      "industry": [
                        "fintech",
                        "healthcare"
                      ]
                    }
                  },
                  {
                    "terms": {
                      "location": [
                        "San Francisco, CA",
                        "New York, NY",
                        "Boston, MA"
                      ]
                    }
                  },
                  {
                    "terms": {
                      "business_model": []
                    }
                  }
                ]
              }
            }
          }
        }
      ],
      "rank_window_size": 50,
      "rank_constant": 10
    }
  }
}
🎯 Found 5 startups matching your criteria:

1. **FinanceAI**
   📍 San Francisco, CA | 🏢 fintech | 💼 B2C
   💰 Series C - $25.0M
   👥 120 employees | 📈 $1200K MRR
   🏦 Lead: Tiger Global Management
   📝 FinanceAI provides AI-powered investment advisory services to retail investors. Uses machine learning to analyze market trends with over 100,000 active users.

2. **CryptoWallet**
   📍 Miami, FL | 🏢 fintech | 💼 B2C
   💰 Series B - $16.0M
   👥 73 employees | 📈 $820K MRR
   🏦 Lead: Coinbase Ventures
   📝 CryptoWallet provides secure digital wallet solutions for cryptocurrency trading and storage. Multi-chain support with enterprise-grade security features.

...

✨  Done in 7.41s.<h2>学习经验</h2><p>在写作过程中我学到了：</p><ul><li><p>我们必须向 LLM 提供筛选器的精确取值，否则就要完全依赖用户输入这些值。对于低基数，这种方法很好，但当基数很高时，我们需要通过一些机制来过滤结果</p></li><li><p>使用搜索模板比让大语言模型编写 Elasticsearch 查询能使结果更加一致，而且速度也更快</p></li><li><p>条件边是一种强大的机制，用于构建具有多个变体和分支路径的应用程序。</p></li><li><p>结构化输出在使用大型语言模型生成信息时非常有用，因为它能强制执行可预测且类型安全的响应。这不仅提高了整体可靠性，还减少了对提示词的误解。</p></li></ul><p>通过混合检索结合语义和结构化搜索，可以产生更好、更相关的结果，在精确性和上下文理解之间取得平衡。</p><h2>结论</h2><p>在这个例子中，我们将 LangGraph.js 与 Elasticsearch 结合，创建一个动态工作流，能够解释自然语言查询并决定使用金融或市场聚焦的搜索策略。这种方法减少了手工查询的复杂性，同时提升了风险投资分析师的灵活性和准确性。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-agent-workflow-finance-langgraph-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-agent-workflow-finance-langgraph-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[智能体 AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt013eba5d152f11f3/6a1709892b835f6784f4b1a6/12b6057d84c6356267cd178a3c6c1a5c61123ece-2000x1256.png" length="0" type="image/png"/>
    <pubDate>Fri, 05 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[人工智能驱动的仪表盘：从设想到 Kibana]]></title>
    <description><![CDATA[使用 LLM 生成仪表盘，处理图像并将其转化为 Kibana 仪表盘。
]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/kibana/kibana-lens">Kibana Lens</a>让仪表盘的拖放变得非常简单，但当你需要几十个面板时，点击次数就会增加。如果你能勾画出一个仪表盘，截图后让法律硕士为你完成整个过程，那会怎么样？</p><p>在本文中，我们将实现这一目标。我们将创建一个应用程序，它可以获取仪表盘的图像，分析映射，然后生成仪表盘，而无需接触 Kibana！</p><p><strong>步骤</strong>：</p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#background-&amp;-application-workflow">后台&amp; 应用程序工作流程</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#prepare-data">准备数据</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#llm-configuration">LLM 配置</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#application-functions">应用功能</a></p></li></ol><h2>后台&amp; 应用程序工作流程</h2><p>我首先想到的是让 LLM 生成整个 NDJSON 格式的 Kibana<a href="https://www.elastic.co/docs/explore-analyze/find-and-organize/saved-objects">保存对象</a>，然后将它们导入 Kibana。</p><p>我们尝试了几种型号：</p><ul><li><p>双子座 2.5 pro</p></li><li><p>GPT o3 / o4-mini-high / 4.1</p></li><li><p>克劳德 4 号十四行诗</p></li><li><p>Grok 3</p></li><li><p>Deepseek (Deepthink R1)</p></li></ul><p>至于提示语，我们从最简单的开始：</p>You are an Elasticsearch Saved-Object generator (Kibana 9.0).
INPUTS
=====
1. PNG screenshot of a 4-panel dashboard (attached).
2. Index mapping (below) – trimmed down to only the fields present in the screenshot.
3. Example NDJSON of *one* metric visualization (below) for reference.

TASK
====
Return **only** a valid NDJSON array that recreates the dashboard exactly:
* 2 metric panels (Visits, Unique Visitors)
* 1 pie chart (Most used OS)
* 1 vertical bar chart (State Geo Dest)
* Use index pattern `kibana_sample_data_logs`.
* Preserve roughly the same layout (2×2 grid).
* Use `panelIndex` values 1-4 and random `id` strings.
* Kibana version: 9.0<p>尽管我们看了<a href="https://www.elastic.co/search-labs/blog/function-calling-with-elastic#:~:text=Few%2Dshot%20prompting%20involves%20providing%20examples%20of%20the%20types%20of%20queries%20you%20want%20it%20to%20return%2C%20which%20helps%20in%20increasing%20consistency.">一些简单的示例</a>，并详细解释了如何建立每种可视化，但我们还是一无所获。如果您对这项实验感兴趣，请<a href="https://gist.github.com/TomasMurua/a78dc283e115624731beffc98984b70b">点击此处</a>了解详情。</p><p>采用这种方法的结果是，在尝试将 LLM 生成的文件上传到 Kibana 时看到了这些信息：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9ea005966a783057/6a1707d266c4f90e4ef8bf88/2b599443b5613c9f0fc3235581614add5b4b3900-891x98.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5e5632d6d95b998c/6a1707d3a6c2b9441de79661/d87ccfc033bc00ee8188c5cae18043fbca22784c-741x233.png" alt="" /><p>这意味着生成的 JSON 无效或格式不当。最常见的问题是 LLM 生成不完整的 NDJSON、产生参数幻觉，或者返回普通 JSON 而非 NDJSON，无论我们如何努力去执行其他操作。</p><p>受<a href="https://www.elastic.co/search-labs/blog/llm-functions-elasticsearch-intelligent-query">这篇文章</a>的启发--<a href="https://www.elastic.co/docs/solutions/search/search-templates">搜索模板</a>比 LLM 自由式更有效--我们决定给 LLM 提供模板，而不是要求它生成完整的 NDJSON 文件，然后我们在代码中使用 LLM 给出的参数来创建适当的可视化。</p><p>申请工作流程如下：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f7738a4c7ddd0cd/6a1707d52b835f0a25f4b166/52c587cf0cf3517fdd4ee7ab95581dd4f2bce030-725x668.png" alt="" /><p></p><p><em>为简单起见，我们将省略一些代码，但您可以在 </em><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/from-image-idea-to-kibana-dashboard-using-ai/from-image-idea-to-kibana-dashboard-using-ai.ipynb"><em><strong>本</strong></em></a><em> 笔记本</em>上找到完整应用程序的工作代码  。</p><h2>准备工作</h2><p>在开始开发之前，您需要具备以下条件：</p><ol><li><p>Python 3.8 或更高版本</p></li><li><p><a href="https://docs.python.org/3/library/venv.html">Venv</a>Python 环境</p></li><li><p>运行的 Elasticsearch 实例及其端点和 API 密钥</p></li><li><p>存储在环境变量 OPENAI_API_KEY 下的 OpenAI API 密钥：</p></li></ol>export OPENAI_API_KEY="your-openai-api-key"<h2>准备数据</h2><p>在数据方面，我们将保持简单，使用 Elastic 样本网络日志。您可以<a href="https://www.elastic.co/docs/manage-data/ingest/sample-data#add-sample-data-sets">在此</a>了解如何将这些数据导入群集。</p><p>每份文档都包含向应用程序发出请求的主机的详细信息，以及请求本身及其响应状态的信息。下面是一个文件示例：</p>{
    "agent": "Mozilla/5.0 (X11; Linux i686) AppleWebKit/534.24 (KHTML, like Gecko) Chrome/11.0.696.50 Safari/534.24",
    "bytes": 8509,
    "clientip": "70.133.115.149",
    "extension": "css",
    "geo": {
        "srcdest": "US:IT",
        "src": "US",
        "dest": "IT",
        "coordinates": {
            "lat": 38.05134111,
            "lon": -103.5106908
        }
    },
    "host": "cdn.elastic-elastic-elastic.org",
    "index": "kibana_sample_data_logs",
    "ip": "70.133.115.149",
    "machine": {
        "ram": 5368709120,
        "os": "osx"
    },
    "memory": null,
    "message": "70.133.115.149 - - [2018-08-30T23:35:31.492Z] \"GET /styles/semantic-ui.css HTTP/1.1\" 200 8509 \"-\" \"Mozilla/5.0 (X11; Linux i686) AppleWebKit/534.24 (KHTML, like Gecko) Chrome/11.0.696.50 Safari/534.24\"",
    "phpmemory": null,
    "referer": "http://twitter.com/error/john-phillips",
    "request": "/styles/semantic-ui.css",
    "response": 200,
    "tags": [
        "success",
        "info"
    ],
    "@timestamp": "2025-07-03T23:35:31.492Z",
    "url": "https://cdn.elastic-elastic-elastic.org/styles/semantic-ui.css",
    "utc_time": "2025-07-03T23:35:31.492Z",
    "event": {
        "dataset": "sample_web_logs"
    },
    "bytes_gauge": 8509,
    "bytes_counter": 51201128
}<p>现在，让我们抓取刚刚加载的索引的映射，<code>kibana_sample_data_logs</code> ：</p>INDEX_NAME = "kibana_sample_data_logs"

es_client = Elasticsearch(
    [os.getenv("ELASTICSEARCH_URL")],
    api_key=os.getenv("ELASTICSEARCH_API_KEY"),
)

result = es_client.indices.get_mapping(index=INDEX_NAME)
index_mappings = result[list(result.keys())[0]]["mappings"]["properties"]<p>我们将把映射与稍后加载的图像一起传递。</p><h2>LLM 配置</h2><p>让我们对 LLM 进行配置，使其使用<a href="https://python.langchain.com/docs/concepts/structured_outputs/">结构化输出</a>来输入图像，并接收包含我们需要传递给函数的信息的 JSON，以生成 JSON 对象。</p><p>我们安装依赖项：</p>pip install elasticsearch pydantic langchain langchain-openai -q<p>Elasticsearch 将帮助我们检索<a href="https://www.elastic.co/docs/manage-data/data-store/mapping">索引映射</a>。Pydantic 允许我们在 Python 中定义模式，然后要求 LLM 遵循这些模式，而<a href="https://www.elastic.co/search-labs/integrations/langchain">LangChain</a>框架则有助于更轻松地调用 LLM 和人工智能工具。</p><p>我们将创建一个 Pydantic 模式，以定义我们希望从 LLM 得到的输出。我们需要从图片中了解图表类型、字段、可视化标题和仪表盘标题：</p>class Visualization(BaseModel):
    title: str = Field(description="The dashboard title")
    type: List[Literal["pie", "bar", "metric"]]
    field: str = Field(
        description="The field that this visualization use based on the provided mappings"
    )


class Dashboard(BaseModel):
    title: str = Field(description="The dashboard title")
    visualizations: List[Visualization]<p>对于图像输入，我们将发送一个我刚刚画好的仪表盘：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7870f6421986d11d/6a1707d78b73cb3408189fa3/36441d7b5dc1f3ff2ac2a30710208d57ad41c716-1600x898.jpg" alt="" /><p>现在我们声明 LLM 模型调用和图像加载。该函数将接收 Elasticsearch 索引的映射和我们要生成的仪表盘图像。</p><p>通过<code>with_structured_output</code> ，我们可以使用 Pydantic<code>Dashboard</code> 模式作为 LLM 生成的响应对象。通过<a href="https://docs.pydantic.dev/latest/">Pydantic</a>，我们可以定义带有验证功能的数据模型，从而确保 LLM 输出与预期结构相匹配。</p><p>要将图像转换为 base64 并作为输入发送，可以使用<a href="https://www.base64-image.de/">在线转换器</a> <a href="https://www.geeksforgeeks.org/python-convert-image-to-string-and-vice-versa/">或用代码</a>完成。</p>prompt = f"""
    You are an expert in analyzing Kibana dashboards from images for the version 9.0.0 of Kibana.

    You will be given a dashboard image and an Elasticsearch index mapping.

    Below are the index mappings for the index that the dashboard is based on.
    Use this to help you understand the data and the fields that are available.

    Index Mappings:
    {index_mappings}

    Only include the fields that are relevant for each visualization, based on what is visible in the image.
    """

message = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": prompt},
            {
                "type": "image",
                "source_type": "base64",
                "data": image_base64,
                "mime_type": "image/png",
            },
        ],
    }
]


try:
    llm = init_chat_model("gpt-4.1-mini")
    llm = llm.with_structured_output(Dashboard)
    dashboard_values = llm.invoke(message)

    print("Dashboard values generated by the LLM successfully")
    print(dashboard_values)
except Exception as e:
    print(f"Failed to analyze image and match fields: {str(e)}")<p>LLM 已经掌握了 Kibana 面板的上下文，因此我们不需要在提示中解释所有内容，只需提供一些细节，确保它不会忘记自己正在使用 Elasticsearch 和 Kibana。</p><p>让我们来分析一下提示：</p><p>部门</p><p>原因</p><p>您是根据 Kibana 9.0.0 版本的图像分析 Kibana 仪表板的专家。</p><p>通过强化 Elasticsearch 和 Elasticsearch 版本，我们降低了 LLM 产生旧参数/无效参数的可能性。</p><p>您将获得一个仪表盘图像和一个 Elasticsearch 索引映射。</p><p>我们解释说，图片是关于仪表盘的，以避免法律硕士做出任何错误的解释。</p><p>下面是仪表盘所基于的索引的索引映射，使用它可以帮助你理解数据和可用字段。索引映射： {index_mappings}</p><p>提供映射至关重要，这样 LLM 才能动态选择有效字段。否则，我们就可能在这里硬编码映射，这太死板了，或者依靠图像包含正确的字段名，这也不可靠。</p><p>根据图像中可见的内容，只包含与每个可视化相关的字段。</p><p>我们必须添加这一增强功能，因为有时它会尝试添加与图像无关的字段。</p><p>这将返回一个包含要显示的可视化数组的对象：</p>"Dashboard values generated by the LLM successfully
title=""Client, Extension, OS, and Response Keyword Analysis""visualizations="[
   "Visualization(title=""Count of Client IP",
   "type="[
      "metric"
   ],
   "field=""clientip"")",
   "Visualization(title=""Extension Keyword Distribution",
   "type="[
      "pie"
   ],
   "field=""extension.keyword"")",
   "Visualization(title=""Most Used OS",
   "type="[
      "bar"
   ],
   "field=""machine.os.keyword"")",
   "Visualization(title=""Response Keyword Distribution",
   "type="[
      "bar"
   ],
   "field=""response.keyword"")"
]<h2>处理 LLM 答复</h2><p>我们在上创建了一个 2x2 面板仪表盘示例，然后使用 "<a href="https://www.elastic.co/docs/api/doc/kibana/operation/operation-get-dashboards-dashboard">获取仪表盘 API "</a>将其导出为 JSON 格式，然后将面板存储为可视化模板（饼状、条状、度量），在这些模板中，我们可以替换部分参数，根据问题创建带有不同字段的新可视化。</p><p>您可以<a href="https://github.com/Delacrobix/elasticsearch-labs/tree/supporting-blog-content/from-image-idea-to-kibana-dashboard-using-ai/supporting-blog-content/from-image-idea-to-kibana-dashboard-using-ai/templates"><strong>在此处</strong></a>查看模板 JSON 文件。请注意我们是如何用 {<code>variable_name</code>} 更改我们稍后要替换的对象值的。
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc55d69d84a08e668/6a1707d8a2929903acd00fb8/ec7e1ac0cd8b470df13e60940162b56778acb386-315x234.png" alt="" /><p>根据 LLM 提供的信息，我们可以决定使用哪个模板，替换哪些值。</p><p><code>fill_template_with_analysis</code> 将接收单个面板的参数，包括可视化的 JSON 模板、标题、字段和可视化在网格上的坐标。</p><p>然后，它会替换模板的值，并返回最终的 JSON 可视化。</p>def fill_template_with_analysis(
    template: Dict[str, Any],
    visualization: Visualization,
    grid_data: Dict[str, Any],
):
    template_str = json.dumps(template)
    replacements = {
	 "{visualization_id}": str(uuid.uuid4()),
        "{title}": visualization.title,
        "{x}": grid_data["x"],
        "{y}": grid_data["y"],
    }

    if visualization.field:
        replacements["{field}"] = visualization.field

    for placeholder, value in replacements.items():
        template_str = template_str.replace(placeholder, str(value))

    return json.loads(template_str)<p>为了简单起见，我们将为 LLM 决定创建的面板分配静态坐标，并生成如上图所示的 2x2 网格仪表盘。</p># Filling templates fields
panels = []    
grid_data = [
    {"x": 0, "y": 0},
    {"x": 12, "y": 0},
    {"x": 0, "y": 12},
    {"x": 12, "y": 12},
]


i = 0

for vis in dashboard_values.visualizations:
    for vis_type in vis.type:
        template = templates.get(vis_type, templates.get("bar", {}))
        filled_panel = fill_template_with_analysis(template, vis, grid_data[i])
        panels.append(filled_panel)
        i += 1<p>根据 LLM 决定的可视化类型，我们将选择一个 JSON 文件模板，并使用<code>fill_template_with_analysis</code> 替换相关信息，然后将新面板追加到稍后用于创建仪表盘的数组中。</p><p>仪表盘准备就绪后，我们将使用<a href="https://www.elastic.co/docs/api/doc/kibana/operation/operation-post-dashboards-dashboard-id"> 创建 仪表盘 API</a> 将新的 JSON 文件推送到 Kibana 以生成仪表盘：
</p>try:
    dashboard_id = str(uuid.uuid4())

    # post request to create the dashboard endpoint
    url = f"{os.getenv('KIBANA_URL')}/api/dashboards/dashboard/{dashboard_id}"

    dashboard_config = {
        "attributes": {
            "title": dashboard_values.title,
            "description": "Generated by AI",
            "timeRestore": True,
            "panels": panels,  # Visualizations with the values generated by the LLM
            "timeFrom": "now-7d/d",
            "timeTo": "now",
        },
    }

    headers = {
        "Content-Type": "application/json",
        "kbn-xsrf": "true",
        "Authorization": f"ApiKey {os.getenv('ELASTICSEARCH_API_KEY')}",
    }

    requests.post(
        url,
        headers=headers,
        json=dashboard_config,
    )

    # Url to the generated dashboard
    dashboard_url = f"{os.getenv('KIBANA_URL')}/app/dashboards#/view/{dashboard_id}"

    print("Dashboard URL: ", dashboard_url)
    print("Dashboard ID: ", dashboard_id)

except Exception as e:
    print(f"Failed to create dashboard: {str(e)}")<p>要执行脚本并生成仪表盘，请在控制台中运行以下命令：</p>python &lt;file_name&gt;.py<p>最终结果将是这样的</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ceffed004153a4f/6a1707d9a929cf9147ae0901/e909afbf0e47d9a6e0f7bd07dfb2efcfa5cf06ac-921x715.png" alt="" /><h2>结论</h2><p>在将文本转化为代码或将图像转化为代码时，LLM 展示了其强大的视觉能力。仪表盘 API 还能将 JSON 文件转化为仪表盘，而通过 LLM 和一些代码，我们就能将图片转化为 Kibana 仪表盘。</p><p>下一步是通过使用不同的网格设置、仪表盘大小和位置来提高仪表盘视觉效果的灵活性。此外，为更复杂的可视化和可视化类型提供支持也是对该应用程序的有益补充。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-powered-dashboards</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-powered-dashboards</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo,Tomás Murúa]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt41727cbee6155a68/6a1707dbb0367dd2fd72bc86/eb60ceb2fbc3941745b21ae3357cbb6ea8fab18c-1443x811.png" length="0" type="image/png"/>
    <pubDate>Wed, 16 Jul 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[正确使用 JavaScript 的 Elasticsearch，第二部分]]></title>
    <description><![CDATA[了解生产环境最佳实践，并学习如何在 Serverless 环境中运行 Elasticsearch Node.js 客户端，以减少代码错误。 ]]></description>
    <content:encoded><![CDATA[<p>这是 Elasticsearch in JavaScript 系列的第二部分。在<a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i"> 第一部分 中 ，</a> 我们学习了如何正确设置环境、配置 Node.js 客户端、索引数据和搜索。在第二部分中，我们将学习如何实施生产最佳实践，并在无服务器环境中运行 Elasticsearch<a href="http://node.js">Node.js</a>客户端。</p><p>我们将审查</p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#production-best-practices">生产最佳实践</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#error-handling">错误处理能力</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#testing">测试</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#serverless-environments">无服务器环境</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#running-the-client-on-elastic-serverless">在 Elastic Serverless 上运行客户端</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#running-the-client-on-function-as-a-service-environment">在功能即服务环境中运行客户端</a></p></li></ul></li></ul><p><em>您可以 </em><a href="https://github.com/Delacrobix/JS-client-best-practices_article"><em><strong>在这里</strong></em></a>查看示例的源代码 <em><strong>。</strong></em></p><h2>生产最佳实践</h2><h3>Elasticsearch 中的错误处理</h3><p>Node.js 中 Elasticsearch 客户端的一个有用功能是，它为 Elasticsearch 中可能出现的错误提供了对象，因此您可以用不同的方式验证和处理这些错误。</p><p>要<a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript/connecting#client-error-handling">查看全部内容</a>，请执行此操作： </p>const { errors } = require('@elastic/elasticsearch')
console.log(errors)<p>让我们回到搜索示例，处理一些可能出现的错误：</p>app.get("/search/lexic", async (req, res) =&gt; {
 ....
  } catch (error) {
    if (error instanceof errors.ResponseError) {
      let errorMessage =
        "Response error!, query malformed or server down, contact the administrator!";

      if (error.body.error.type === "parsing_exception") {
        errorMessage = "Query malformed, make sure mappings are set correctly";
      }

      res.status(error.meta.statusCode).json({
        erroStatus: error.meta.statusCode,
        success: false,
        results: null,
        error: errorMessage,
      });
    }

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p><code>ResponseError</code> 尤其是当答案为<code>4xx</code> 或<code>5xx</code> 时，即表示请求不正确或服务器不可用。</p><p>我们可以通过生成错误查询来测试这类错误，比如尝试<strong>在文本类型字段上进行术语查询：</strong></p><p>默认错误：</p> {
    "success": false,
    "results": null,
    "error": "parsing_exception\n\tRoot causes:\n\t\tparsing_exception: [terms] query does not support [visit_details]"
}<p>定制错误： </p>{
    "erroStatus": 400,
    "success": false,
    "results": null,
    "error": "Response error!, query malformed or server down; contact the administrator!"
}<p>我们还可以以某种方式捕捉和处理每种类型的错误。例如，我们可以在<code>TimeoutError</code> 中添加重试逻辑。</p>app.get("/search/semantic", async (req, res) =&gt; {
    try {
  ...
  } catch (error) {
    if (error instanceof errors.TimeoutError) {


     // Retry logic...

      res.status(error.meta.statusCode).json({
        erroStatus: error.meta.statusCode,
        success: false,
        results: null,
        error:
          "The request took more than 10s after 3 retries. Try again later.",
      });
    }
  }
});<h3>测试</h3><p>测试是保证应用程序稳定性的关键。为了以一种与 Elasticsearch 隔离的方式测试代码，我们可以在创建集群时使用<a href="https://github.com/elastic/elasticsearch-js-mock">elasticsearch-js-mock</a>库。</p><p>通过该库，我们可以实例化一个与真实客户端非常相似的客户端，但只需将客户端的 HTTP 层替换为模拟层，其他部分与原始客户端保持一致，就能满足我们的配置要求。</p><p>我们将安装 mocks 库和用于自动测试的<a href="https://github.com/avajs/ava">AVA</a>。</p><p><code>npm install @elastic/elasticsearch-mock</code></p><p><code>npm install --save-dev ava</code></p><p>我们将配置<code>package.json</code> 文件以运行测试。确保它看起来是这样的：</p>"type": "module",
	"scripts": {
		"test": "ava"
	},
	"devDependencies": {
		"ava": "^5.0.0"
	}<p>现在，让我们创建<code>test.js</code> 文件并安装我们的模拟客户端：</p>const { Client } = require('@elastic/elasticsearch')
const Mock = require('@elastic/elasticsearch-mock')

const mock = new Mock()
const client = new Client({
  node: 'http://localhost:9200',
  Connection: mock.getConnection()
})<p>现在，为语义搜索添加一个模拟：</p>function createSemanticSearchMock(query, indexName) {
  mock.add(
    {
      method: "POST",
      path: `/${indexName}/_search`,
      body: {
        query: {
          semantic: {
            field: "semantic_field",
            query: query,
          },
        },
      },
    },
    () =&gt; {
      return {
        hits: {
          total: { value: 2, relation: "eq" },
          hits: [
            {
              _id: "1",
              _score: 0.9,
              _source: {
                owner_name: "Alice Johnson",
                pet_name: "Buddy",
                species: "Dog",
                breed: "Golden Retriever",
                vaccination_history: ["Rabies", "Parvovirus", "Distemper"],
                visit_details:
                  "Annual check-up and nail trimming. Healthy and active.",
              },
            },
            {
              _id: "2",
              _score: 0.7,
              _source: {
                owner_name: "Daniel Kim",
                pet_name: "Mochi",
                species: "Rabbit",
                breed: "Mixed",
                vaccination_history: [],
                visit_details:
                  "Nail trimming and general health check. No issues.",
              },
            },
          ],
        },
      };
    }
  );
}<p>现在我们可以为代码创建一个测试，确保 Elasticsearch 部分始终返回相同的结果：</p>import test from 'ava';

test("performSemanticSearch must return formatted results correctly", async (t) =&gt; {
  const indexName = "vet-visits";
  const query = "Which pets had nail trimming?";

  createSemanticSearchMock(query, indexName);

  async function performSemanticSearch(esClient, q, indexName = "vet-visits") {
    try {
      const result = await esClient.search({
        index: indexName,
        body: {
          query: {
            semantic: {
              field: "semantic_field",
              query: q,
            },
          },
        },
      });

      return {
        success: true,
        results: result.hits.hits,
      };
    } catch (error) {
      if (error instanceof errors.TimeoutError) {
        return {
          success: false,
          results: null,
          error: error.body.error.reason,
        };
      }

      return {
        success: false,
        results: null,
        error: error.message,
      };
    }
  }

  const result = await performSemanticSearch(esClient, query, indexName);

  t.true(result.success, "The search must be successful");
  t.true(Array.isArray(result.results), "The results must be an array");

  if (result.results.length &gt; 0) {
    t.true(
      "_source" in result.results[0],
      "Each result must have a _source property"
    );
    t.true(
      "pet_name" in result.results[0]._source,
      "Results must include the pet_name field"
    );
    t.true(
      "visit_details" in result.results[0]._source,
      "Results must include the visit_details field"
    );
  }
});<p>让我们进行测试。</p><p><code>npm run test</code></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt36304e286146f362/6a170559d7c02237b2de638f/42feae845ae8eae03c37ad7ad114e8db35984812-1186x302.png" alt="" /><p>完成！从现在起，我们就可以测试我们的应用程序，100% 专注于代码而不是外部因素。</p><h2>无服务器环境</h2><h3>如何在 Elastic Serverless 上运行客户端</h3><p>我们介绍了在云端或内部运行 Elasticsearch 的情况；不过，Node.js 客户端也支持与<a href="https://www.elastic.co/guide/en/serverless/current/intro.html">Elastic Cloud Serverless</a> 的连接。</p><p>Elastic Cloud Serverless 允许您创建一个项目，在这个项目中，您无需担心基础设施问题，因为 Elastic 会在内部处理这些问题，您只需担心您想索引的数据以及您想在多长时间内访问这些数据。</p><p>从使用角度来看，Serverless 将计算与存储分离，为<a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-tier-autoscaling">搜索</a>和<a href="https://www.elastic.co/search-labs/blog/elasticsearch-ingest-autoscaling">索引</a>提供了自动扩展功能。这样，您就可以只增长实际需要的资源。</p><p>客户端会进行以下调整，以连接到无服务器：</p><ul><li><p>关闭嗅探，忽略任何与嗅探相关的选项</p></li><li><p>忽略配置中传递的除第一个节点外的所有节点，并忽略任何节点过滤和选择选项</p></li><li><p>启用压缩和 "TLSv1_2_method"（与为弹性云配置时相同）</p></li><li><p>为所有请求添加 "elastic-api-version "HTTP 头信息</p></li><li><p>默认使用 "云连接池"，而不是 "加权连接池</p></li><li><p>关闭卖方 "内容类型 "和 "接受 "标头，转而使用标准 MIME 类型</p></li></ul><p>要连接无服务器项目，需要使用参数 serverMode：serverless。</p>const { Client } = require('@elastic/elasticsearch')
const client = new Client({
  node: 'ELASTICSEARCH_ENDPOINT',
  auth: { apiKey: 'ELASTICSEARCH_API_KEY' },
  serverMode: "serverless",
});<h3>如何在函数即服务环境中运行客户端</h3><p>在示例中，我们使用了 Node.js 服务器，但您也可以使用功能即服务环境连接 AWS lambda、GCP Run 等功能。</p>'use strict'

const { Client } = require('@elastic/elasticsearch')

const client = new Client({
  // client initialisation
})

exports.handler = async function (event, context) {
  // use the client
}<p>另一个例子是连接像 Vercel 这样的服务，它也是无服务器的。您可以查看这个<a href="https://github.com/elastic/elasticsearch-js/blob/main/docs/examples/proxy/README.md">完整的示例</a>，了解如何做到这一点，但<a href="https://github.com/elastic/elasticsearch-js/blob/main/docs/examples/proxy/api/search.js">搜索端点</a>最相关的部分如下所示：</p>const response = await client.search(
  {
    index: INDEX,
    // You could directly send from the browser
    // the Elasticsearch's query DSL, but it will
    // expose you to the risk that a malicious user
    // could overload your cluster by crafting
    // expensive queries.
    query: {
      match: { field: req.body.text },
    },
  },
  {
    headers: {
      Authorization: `ApiKey ${token}`,
    },
  }
);<p>该端点位于 /api 文件夹中，从服务器端运行，因此客户端只能控制与搜索词相对应的 "文本 "参数。</p><p>使用 "功能即服务 "的意义在于，与全天候运行的服务器不同，功能只启动运行该功能的机器，一旦完成，机器就会进入休息模式，以减少资源消耗。</p><p>如果应用程序没有收到太多请求，这种配置会很方便；否则，成本会很高。您还需要考虑<a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtime-environment.html">函数的生命周期</a>和运行时间（在某些情况下可能只有几秒钟）。</p><h2>结论</h2><p>在本文中，我们学习了如何处理错误，这在生产环境中至关重要。我们还介绍了在模拟 Elasticsearch 服务的过程中测试应用程序的方法，无论集群的状态如何，这种方法都能提供可靠的测试，让我们专注于我们的代码。</p><p>最后，我们演示了如何通过配置 Elastic Cloud Serverless 和 Vercel 应用程序来启动完全无服务器堆栈。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii</guid>
    <category><![CDATA[Javascript]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc58be329ffebcd60/6a17043e47d49c0bc62d88ab/70fb0ff949f6db9ac9b8a28ecb4329ab915ebf46-720x420.png" length="0" type="image/png"/>
    <pubDate>Mon, 19 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[正确使用 JavaScript 的 Elasticsearch，第一部分]]></title>
    <description><![CDATA[讲解如何用 JavaScript 创建可投入生产的 Elasticsearch 后端。  

探索如何使用 JavaScript 与 Elasticsearch，遵循客户端/服务器最佳实践，搭建包含多个搜索端点的服务器，用于查询 Elasticsearch 文档。]]></description>
    <content:encoded><![CDATA[<p>本文是系列文章的第一篇，介绍如何使用 JavaScript 使用 Elasticsearch。在本系列中，您将学习如何在 JavaScript 环境中使用 Elasticsearch 的基础知识，并回顾创建搜索应用程序的最相关功能和最佳实践。最后，您将了解使用 JavaScript 运行 Elasticsearch 所需的一切。</p><p>在第一部分中，我们将回顾</p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#environment">环境</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#frontend,-backend,-or-serverless?">前端、后端还是无服务器？</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#connecting-the-client">连接客户端</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#indexing-documents">编制文件索引</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#elasticsearch-client">Elasticsearch 客户端</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#semantic-mappings">语义映射</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#bulk-helper">批量助手</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#searching-data">搜索数据</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#lexical-query-(/search/lexic?q=%3Cquery-term%3E)">词法查询</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#semantic-query-(/search/semantic?q=%3Cquery-term%3E)">语义查询</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#hybrid-query-(/search/hybrid?q=%3Cquery-term%3E)">混合查询</a></p></li></ul></li></ul><p><em>您可以 </em><a href="https://github.com/Delacrobix/JS-client-best-practices_article"><em><strong>在这里</strong></em></a>查看示例的源代码 <em><strong>。</strong></em></p><h3>什么是 Elasticsearch Node.js 客户端？</h3><p><a href="https://www.elastic.co/guide/en/elasticsearch/client/javascript-api/current/index.html">Elasticsearch Node.js 客户端</a>是一个 JavaScript 库，它将 Elasticsearch API 的 HTTP REST 调用放到了 JavaScript 中。这样就能更轻松地处理和使用帮助程序，简化批量编制文档索引等任务。</p><h2>环境</h2><h3>前端、后端还是无服务器？</h3><p>要使用 JavaScript 客户端创建搜索应用程序，我们至少需要两个组件：Elasticsearch 集群和运行客户端的 JavaScript 运行时。</p><p>JavaScript 客户端支持所有 Elasticsearch 解决方案（云、on-prem 和 Serverless），它们之间没有重大区别，因为客户端内部会处理所有变化，所以你不必担心使用哪一种。</p><p>不过，JavaScript 运行时必须从<strong>服务器</strong>运行，而<strong>不能直接从浏览器</strong>运行。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd3ec469c83e3a71a/6a17e3d5445de91da44d00b6/92ce6cfd923c8008fa44f617a58193642d9d5879-661x410.png" alt="在 JavaScript 环境中使用 Elasticsearch。" /><p>这是因为从浏览器调用 Elasticsearch 时，用户可能会获得敏感信息，如集群 API 密钥、主机或查询本身。Elasticsearch 建议<strong>永远不要将集群直接暴露在互联网上 </strong>，而是使用一个中间层来抽象所有这些信息，这样用户只能看到参数。您可以<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/es-security-principles.html#security-protect-cluster-traffic">在这里</a>了解更多相关信息。</p><p>我们建议使用这样的模式：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d7f215f2e70230a/6a17e3d6fbc5f83de6491a13/a08769f08ec73fe57bf2e961cfdfbb1cdd57919d-972x429.png" alt="设置 Elasticsearch Node.js 客户端。" /><p>在这种情况下，客户端只向服务器发送搜索条件和验证密钥，而服务器则完全控制查询和与 Elasticsearch 的通信。</p><h3>连接客户端</h3><p>首先，按照<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">以下步骤</a>创建一个 API 密钥。</p><p>按照前面的示例，我们将创建一个简单的 Express 服务器，并使用 Node.JS 服务器的客户端连接到该服务器。</p><p>我们将使用 NPM 初始化项目，并安装 Elasticsearch 客户端和<a href="https://expressjs.com/">Express。</a>后者是一个在 Node.js 中调用服务器的库。使用 Express，我们可以通过 HTTP 与后端交互。</p><p>让我们初始化项目：</p><p><code>npm init -y</code></p><p>安装依赖项：</p><p><code>npm install @elastic/elasticsearch express split2 dotenv</code></p><p>让我来为你分析一下：</p><ul><li><p><a href="https://www.npmjs.com/package/@elastic/elasticsearch"><em><strong>@elastic/elasticsearch</strong></em></a>：它是 Node.js 的官方客户端</p></li><li><p><a href="https://www.npmjs.com/package/express"><em><strong>快递</strong></em></a>：它将使我们能够运行一个轻量级的 nodejs 服务器，以暴露 Elasticsearch</p></li><li><p><a href="https://www.npmjs.com/package/split2"><em><strong>split2</strong></em></a>： 将文本行分割成数据流。每次处理一行 ndjson 文件时非常有用</p></li><li><p><a href="https://www.npmjs.com/package/dotenv"><em><strong>dotenv</strong></em></a>：允许我们使用 .env 管理环境变量文件</p></li></ul><p>创建 .env文件，并添加以下几行：</p>ELASTICSEARCH_ENDPOINT="Your Elasticsearch endpoint"
ELASTICSEARCH_API_KEY="Your Elasticssearch API"<p>这样，我们就可以使用<code>dotenv</code> 软件包导入这些变量。</p><p>创建<code>server.js</code> 文件：</p>const express = require("express");
const bodyParser = require("body-parser");
const { Client } = require("@elastic/elasticsearch");
 
require("dotenv").config(); //environment variables setup

const ELASTICSEARCH_ENDPOINT = process.env.ELASTICSEARCH_ENDPOINT;
const ELASTICSEARCH_API_KEY = process.env.ELASTICSEARCH_API_KEY;
const PORT = 3000;


const app = express();

app.listen(PORT, () =&gt; {
  console.log("Server running on port", PORT);
});
app.use(bodyParser.json());


let esClient = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: { apiKey: ELASTICSEARCH_API_KEY },  
});

app.get("/ping", async (req, res) =&gt; {
  try {
    const result = await esClient.info();

    res.status(200).json({
      success: true,
      clusterInfo: result,
    });
  } catch (error) {
    console.error("Error getting Elasticsearch info:", error);

    res.status(500).json({
      success: false,
      clusterInfo: null,
      error: error.message,
    });
  }
});<p>这段代码设置了一个基本的 Express.js 服务器，该服务器监听端口 3000，并使用 API 密钥进行身份验证，连接到 Elasticsearch 集群。它包括一个 /ping 端点，通过 GET 请求访问时，可使用 Elasticsearch 客户端的<code>.info()</code> 方法查询 Elasticsearch 集群的基本信息。 </p><p>如果查询成功，会以 JSON 格式返回群集信息；否则会返回错误信息。服务器还使用 body-parser 中间件来处理 JSON 请求体。</p><p>运行文件，启动服务器：</p><p><code>node server.js</code></p><p>答案应该是这样的</p>Server running on port 3000<p>现在，让我们查阅端点<code>/ping</code> ，检查 Elasticsearch 集群的状态。</p>curl http://localhost:3000/ping
{
    "success": true,
    "clusterInfo": {
        "name": "instance-0000000000",
        "cluster_name": "61b7e19eec204d59855f5e019acd2689",
        "cluster_uuid": "BIfvfLM0RJWRK_bDCY5ldg",
        "version": {
            "number": "9.0.0",
            "build_flavor": "default",
            "build_type": "docker",
            "build_hash": "112859b85d50de2a7e63f73c8fc70b99eea24291",
            "build_date": "2025-04-08T15:13:46.049795831Z",
            "build_snapshot": false,
            "lucene_version": "10.1.0",
            "minimum_wire_compatibility_version": "8.18.0",
            "minimum_index_compatibility_version": "8.0.0"
        },
        "tagline": "You Know, for Search"
    }
}<h2>编制文件索引</h2><p>一旦连接起来，我们就可以使用语义<a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">_文本（</a>用于语义搜索）和文本（用于全文查询）等映射对文档进行索引。有了这两种字段类型，我们还可以进行<a href="https://www.elastic.co/what-is/hybrid-search">混合搜索</a>。</p><p>我们将创建一个新的<code>load.js</code> 文件来生成映射并上传文件。</p><h3>Elasticsearch 客户端</h3><p>我们首先需要对客户端进行实例化和身份验证：</p>const { Client } = require("@elastic/elasticsearch");

const ELASTICSEARCH_ENDPOINT = "cluster/project_endpoint";
const ELASTICSEARCH_API_KEY = "apiKey";

const esClient = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: { apiKey: ELASTICSEARCH_API_KEY },
});<h3>语义映射</h3><p>我们将创建一个包含兽医院数据的索引。我们将保存主人、宠物和访问详情的信息。</p><p>我们要进行全文搜索的数据，如名称和描述，将以文本形式存储。类别中的数据，如动物的种类或品种，将以关键字的形式存储。</p><p>此外，我们还将把所有字段的值复制到一个 semantic_text 字段中，以便也能针对这些信息运行语义搜索。</p>const INDEX_NAME = "vet-visits";

const createMappings = async (indexName, mapping) =&gt; {
  try {
    const body = await esClient.indices.create({
      index: indexName,
      body: {
        mappings: mapping,
      },
    });

    console.log("Index created successfully:", body);
  } catch (error) {
    console.error("Error creating mapping:", error);
  }
};

await createMappings(INDEX_NAME, {
  properties: {
    owner_name: {
      type: "text",
      copy_to: "semantic_field",
    },
    pet_name: {
      type: "text",
      copy_to: "semantic_field",
    },
    species: {
      type: "keyword",
      copy_to: "semantic_field",
    },
    breed: {
      type: "keyword",
      copy_to: "semantic_field",
    },
    vaccination_history: {
      type: "keyword",
      copy_to: "semantic_field",
    },
    visit_details: {
      type: "text",
      copy_to: "semantic_field",
    },
    semantic_field: {
      type: "semantic_text",
    },
  },
});<h3>批量助手</h3><p>客户端的另一个优势是，我们可以使用<a href="https://www.elastic.co/guide/en/elasticsearch/client/javascript-api/current/client-helpers.html#bulk-helper">批量助手</a>来分批建立索引。通过批量辅助器，我们可以轻松处理并发、重试等问题，以及如何处理通过函数成功或失败的每个文档。</p><p>该助手的一个吸引人的特点是可以使用数据流。该功能允许您逐行发送文件，而不是将整个文件存储在内存中并一次性发送到 Elasticsearch。</p><p>要将数据上传到 Elasticsearch，请在项目根目录下创建名为 data.ndjson 的文件，并添加以下信息（也可以从<a href="https://github.com/Delacrobix/JS-client-best-practices_article/blob/main/data.ndjson">此处</a>下载包含数据集的文件）：</p>{"owner_name":"Alice Johnson","pet_name":"Buddy","species":"Dog","breed":"Golden Retriever","vaccination_history":["Rabies","Parvovirus","Distemper"],"visit_details":"Annual check-up and nail trimming. Healthy and active."}
{"owner_name":"Marco Rivera","pet_name":"Milo","species":"Cat","breed":"Siamese","vaccination_history":["Rabies","Feline Leukemia"],"visit_details":"Slight eye irritation, prescribed eye drops."}
{"owner_name":"Sandra Lee","pet_name":"Pickles","species":"Guinea Pig","breed":"Mixed","vaccination_history":[],"visit_details":"Loss of appetite, recommended dietary changes."}
{"owner_name":"Jake Thompson","pet_name":"Luna","species":"Dog","breed":"Labrador Mix","vaccination_history":["Rabies","Bordetella"],"visit_details":"Mild ear infection, cleaning and antibiotics given."}
{"owner_name":"Emily Chen","pet_name":"Ziggy","species":"Cat","breed":"Mixed","vaccination_history":["Rabies","Feline Calicivirus"],"visit_details":"Vaccination update and routine physical."}
{"owner_name":"Tomás Herrera","pet_name":"Rex","species":"Dog","breed":"German Shepherd","vaccination_history":["Rabies","Parvovirus","Leptospirosis"],"visit_details":"Follow-up for previous leg strain, improving well."}
{"owner_name":"Nina Park","pet_name":"Coco","species":"Ferret","breed":"Mixed","vaccination_history":["Rabies"],"visit_details":"Slight weight loss; advised new diet."}
{"owner_name":"Leo Martínez","pet_name":"Simba","species":"Cat","breed":"Maine Coon","vaccination_history":["Rabies","Feline Panleukopenia"],"visit_details":"Dental cleaning. Minor tartar buildup removed."}
{"owner_name":"Rachel Green","pet_name":"Rocky","species":"Dog","breed":"Bulldog Mix","vaccination_history":["Rabies","Parvovirus"],"visit_details":"Skin rash, antihistamines prescribed."}
{"owner_name":"Daniel Kim","pet_name":"Mochi","species":"Rabbit","breed":"Mixed","vaccination_history":[],"visit_details":"Nail trimming and general health check. No issues."}<p>我们使用 split2 对文件行进行流式处理，而批量助手则将它们发送到 Elasticsearch。</p>const { createReadStream } = require("fs");
const split = require("split2");
 
const indexData = async (filePath, indexName) =&gt; {
  try {
    console.log(`Indexing data from ${filePath} into ${indexName}...`);

    const result = await esClient.helpers.bulk({
      datasource: createReadStream(filePath).pipe(split()),

      onDocument: () =&gt; {
        return {
          index: { _index: indexName },
        };
      },
      onDrop(doc) {
        console.error("Error processing document:", doc);
      },
    });

    console.log("Bulk indexing successful elements:", result.items.length);
  } catch (error) {
    console.error("Error indexing data:", error);
    throw error;
  }
};

await indexData("./data.ndjson", INDEX_NAME);<p>上面的代码读取 .ndjson文件，并使用<code>helpers.bulk</code> 方法将每个 JSON 对象批量索引到指定的 Elasticsearch 索引中。它使用<code>createReadStream</code> 和<code>split2</code> 对文件进行流式处理，为每个文件设置索引元数据，并记录处理失败的文件。完成后，它会记录成功索引的项目数。</p><p>除<code>indexData</code> 功能外，您还可以使用 Kibana 直接通过用户界面上传文件，并使用<a href="https://www.elastic.co/docs/manage-data/ingest/upload-data-files">上传数据文件用户界面。</a></p><p>我们运行文件，将文件上传到 Elasticsearch 集群。</p><p><code>node load.js</code></p>Creating mappings for index vet-visits...
Index created successfully: { acknowledged: true, shards_acknowledged: true, index: 'vet-visits' }
Indexing data from ./data.ndjson into vet-visits...
Bulk indexing completed. Total documents: 10, Failed: 0<h2>在 Elasticsearch 中搜索数据</h2><p>回到<code>server.js</code> 文件，我们将创建不同的端点来执行词法、语义或混合搜索。</p><p>简而言之，这些类型的搜索并不相互排斥，而是取决于您需要回答的问题类型。</p><p>查询类型</p><p>用例</p><p>问题示例</p><p>词法查询</p><p>问题中的单词或词根很可能出现在索引文件中。问题与文件之间的标记相似性。</p><p>我在找一件蓝色运动 T 恤。</p><p>语义查询</p><p>问题中的词语不可能出现在文件中。问题与文件之间的概念相似性。</p><p>我在寻找适合寒冷天气穿的衣服。</p><p>混合搜索</p><p>问题包含词汇和/或语义成分。问题与文档之间的标记和语义相似性。</p><p>我想为海滩婚礼找一件 S 码的礼服。</p><p>问题的<em><strong>词汇 </strong></em>部分很可能是标题和说明的一部分，或者是类别名称，而<em><strong>语义 </strong></em>部分则是与这些领域相关的概念。<em><strong>蓝色</strong></em>可能是一个类别名称或描述的一部分，<em><strong>海滩婚礼</strong></em>不太可能是，但可以与亚麻服装在语义上相关。</p><h3>词法查询 (/search/lexic?q=&lt;query_term&gt;)</h3><p>词法搜索也称全文搜索，是指基于标记的相似性进行搜索；也就是说，经过分析后，将返回包含搜索标记的文档。</p><p>您可以<a href="https://www.elastic.co/demo-gallery/lexical-search">点击此处</a>查看我们的词法搜索实践教程。</p>app.get("/search/lexic", async (req, res) =&gt; {
  const { q } = req.query;

  const INDEX_NAME = "vet-visits";

  try {
    const result = await esClient.search({
      index: INDEX_NAME,
      size: 5,
      body: {
        query: {
          multi_match: {
            query: q,
            fields: ["owner_name", "pet_name", "visit_details"],
          },
        },
      },
    });

    res.status(200).json({
      success: true,
      results: result.hits.hits
    });
  } catch (error) {
    console.error("Error performing search:", error);

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p>我们测试：<em><strong>修剪指甲</strong></em></p>curl http://localhost:3000/search/lexic?q=nail%20trimming<p>请回答：</p>{
    "success": true,
    "results": [
        {
            "_index": "vet-visits",
            "_id": "-RY6RJYBLe2GoFQ6-9n9",
            "_score": 2.7075968,
            "_source": {
                "pet_name": "Mochi",
                "owner_name": "Daniel Kim",
                "species": "Rabbit",
                "visit_details": "Nail trimming and general health check. No issues.",
                "breed": "Mixed",
                "vaccination_history": []
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8BY6RJYBLe2GoFQ6-9n9",
            "_score": 2.560356,
            "_source": {
                "pet_name": "Buddy",
                "owner_name": "Alice Johnson",
                "species": "Dog",
                "visit_details": "Annual check-up and nail trimming. Healthy and active.",
                "breed": "Golden Retriever",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Distemper"
                ]
            }
        }
    ]
}<h3>语义查询 (/search/semantic?q=&lt;query_term&gt;)</h3><p>语义搜索与词汇搜索不同，它通过矢量搜索找到与搜索词含义相似的结果。</p><p>您可以<a href="https://www.elastic.co/demo-gallery/semantic-search">点击这里</a>查看我们的语义搜索实践教程。</p>app.get("/search/semantic", async (req, res) =&gt; {
  const { q } = req.query;

  const INDEX_NAME = "vet-visits";

  try {
    const result = await esClient.search({
      index: INDEX_NAME,
      size: 5,
      body: {
        query: {
          semantic: {
            field: "semantic_field",
            query: q
          },
        },
      },
    });

    res.status(200).json({
      success: true,
      results: result.hits.hits,
    });
  } catch (error) {
    console.error("Error performing search:", error);

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p>我们进行测试：<em><strong>谁做了修脚？</strong></em></p>curl http://localhost:3000/search/semantic?q=Who%20got%20a%20pedicure?<p>请回答：</p>{
    "success": true,
    "results": [
        {
            "_index": "vet-visits",
            "_id": "-RY6RJYBLe2GoFQ6-9n9",
            "_score": 4.861466,
            "_source": {
                "owner_name": "Daniel Kim",
                "pet_name": "Mochi",
                "species": "Rabbit",
                "breed": "Mixed",
                "vaccination_history": [],
                "visit_details": "Nail trimming and general health check. No issues."
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8BY6RJYBLe2GoFQ6-9n9",
            "_score": 4.7152824,
            "_source": {
                "pet_name": "Buddy",
                "owner_name": "Alice Johnson",
                "species": "Dog",
                "visit_details": "Annual check-up and nail trimming. Healthy and active.",
                "breed": "Golden Retriever",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Distemper"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "9RY6RJYBLe2GoFQ6-9n9",
            "_score": 1.6717153,
            "_source": {
                "pet_name": "Rex",
                "owner_name": "Tomás Herrera",
                "species": "Dog",
                "visit_details": "Follow-up for previous leg strain, improving well.",
                "breed": "German Shepherd",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Leptospirosis"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "9xY6RJYBLe2GoFQ6-9n9",
            "_score": 1.5600781,
            "_source": {
                "pet_name": "Simba",
                "owner_name": "Leo Martínez",
                "species": "Cat",
                "visit_details": "Dental cleaning. Minor tartar buildup removed.",
                "breed": "Maine Coon",
                "vaccination_history": [
                    "Rabies",
                    "Feline Panleukopenia"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "-BY6RJYBLe2GoFQ6-9n9",
            "_score": 1.2696637,
            "_source": {
                "pet_name": "Rocky",
                "owner_name": "Rachel Green",
                "species": "Dog",
                "visit_details": "Skin rash, antihistamines prescribed.",
                "breed": "Bulldog Mix",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus"
                ]
            }
        }
    ]
}<h3>混合查询 (/search/hybrid?q=&lt;query_term&gt;)</h3><p>混合搜索允许我们将语义搜索和词法搜索结合起来，从而获得两全其美的效果：既能获得标记搜索的精确性，又能获得语义搜索的意义接近性。</p>app.get("/search/hybrid", async (req, res) =&gt; {
  const { q } = req.query;

  const INDEX_NAME = "vet-visits";

  try {
    const result = await esClient.search({
      index: INDEX_NAME,
      body: {
        retriever: {
          rrf: {
            retrievers: [
              {
                standard: {
                  query: {
                    bool: {
                      must: {
                         multi_match: {
             query: q,
            fields: ["owner_name", "pet_name", "visit_details"],
          },
                      },
                    },
                  },
                },
              },
              {
                standard: {
                  query: {
                    bool: {
                      must: {
                        semantic: {
                          field: "semantic_field",
                          query: q,
                        },
                      },
                    },
                  },
                },
              },
            ],
          },
        },
        size: 5,
      },
    });

    res.status(200).json({
      success: true,
      results: result.hits.hits,
    });
  } catch (error) {
    console.error("Error performing search:", error);

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p>我们以 "<em><strong>谁做了修脚或牙科治疗？"</strong></em></p>curl http://localhost:3000/search/hybrid?q=who%20got%20a%20pedicure%20or%20dental%20treatment<p>响应：</p>{
    "success": true,
    "results": [
        {
            "_index": "vet-visits",
            "_id": "9xY6RJYBLe2GoFQ6-9n9",
            "_score": 0.032522473,
            "_source": {
                "pet_name": "Simba",
                "owner_name": "Leo Martínez",
                "species": "Cat",
                "visit_details": "Dental cleaning. Minor tartar buildup removed.",
                "breed": "Maine Coon",
                "vaccination_history": [
                    "Rabies",
                    "Feline Panleukopenia"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "-RY6RJYBLe2GoFQ6-9n9",
            "_score": 0.016393442,
            "_source": {
                "pet_name": "Mochi",
                "owner_name": "Daniel Kim",
                "species": "Rabbit",
                "visit_details": "Nail trimming and general health check. No issues.",
                "breed": "Mixed",
                "vaccination_history": []
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8BY6RJYBLe2GoFQ6-9n9",
            "_score": 0.015873017,
            "_source": {
                "pet_name": "Buddy",
                "owner_name": "Alice Johnson",
                "species": "Dog",
                "visit_details": "Annual check-up and nail trimming. Healthy and active.",
                "breed": "Golden Retriever",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Distemper"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "9RY6RJYBLe2GoFQ6-9n9",
            "_score": 0.015625,
            "_source": {
                "pet_name": "Rex",
                "owner_name": "Tomás Herrera",
                "species": "Dog",
                "visit_details": "Follow-up for previous leg strain, improving well.",
                "breed": "German Shepherd",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Leptospirosis"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8xY6RJYBLe2GoFQ6-9n9",
            "_score": 0.015384615,
            "_source": {
                "pet_name": "Luna",
                "owner_name": "Jake Thompson",
                "species": "Dog",
                "visit_details": "Mild ear infection, cleaning and antibiotics given.",
                "breed": "Labrador Mix",
                "vaccination_history": [
                    "Rabies",
                    "Bordetella"
                ]
            }
        }
    ]
}<h2>结论</h2><p>在本系列的第一部分中，我们介绍了如何按照客户端/服务器最佳实践设置环境并创建带有不同搜索端点的服务器，以查询 Elasticsearch 文档。查看我们系列的<a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i">第二部分</a>，您将了解生产最佳实践以及如何在无服务器环境中运行 Elasticsearch Node.js 客户端。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i</guid>
    <category><![CDATA[Javascript]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt16d00c8a548b32e8/6a17e3d8fbc5f8c740491a19/72200540ed258779d87e53a72ea189f8a138540c-1600x901.png" length="0" type="image/png"/>
    <pubDate>Thu, 15 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[将 Ollama 与推理应用程序接口结合使用]]></title>
    <description><![CDATA[了解如何使用 Inference API 将 Ollama 与 Elasticsearch 集成。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将学习如何使用 Ollama 将本地模型连接到 Elasticsearch 推理模型，然后使用 Playground 提出文档问题。</p><p>Elasticsearch 允许用户使用开放<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/inference-apis.html">推理 API</a> 连接到 LLM，并支持 Amazon Bedrock、Cohere、Google AI、Azure AI Studio、HuggingFace - as a service 等提供商。</p><p><a href="https://ollama.com">Ollama</a>是一款允许您使用自己的基础设施（本地机器/服务器）下载和执行 LLM 模型的工具。<a href="https://ollama.com/library">在这里</a>，您可以找到与 Ollama 兼容的可用型号列表。</p><p>如果你想托管和测试不同的开源模型，Ollama 是一个不错的选择，因为 Ollama 会处理好一切，让你不必担心每个模型的不同设置方式，也不必担心如何创建 API 来访问模型功能。</p><p>由于 Ollama API 与 OpenAI API 兼容，我们可以轻松集成推理模型，并使用 Playground 创建 RAG 应用程序。</p><h2>准备工作</h2><ol><li><p>Elasticsearch 8.17</p></li><li><p>Kibana 8.17</p></li><li><p>Python</p></li></ol><h2>步长</h2><ol><li><p><a href="https://www.elastic.co/cn/search-labs/blog/ollama-with-inference-api#setting-up-ollama-llm-server">设置 Ollama LLM 服务器</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/ollama-with-inference-api#creating-mappings">创建映射</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/ollama-with-inference-api#indexing-data">索引数据</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/ollama-with-inference-api#asking-questions-using-playground">使用 Playground 提问</a></p></li></ol><h2>设置 Ollama LLM 服务器</h2><p>我们将设置一个 LLM 服务器，使用 Ollama 将其连接到 Playground 实例。我们需要</p><ul><li><p>下载并运行 Ollama。</p></li><li><p>使用 ngrok 通过互联网访问托管 Ollama 的本地网络服务器</p></li></ul><h3>下载并运行 Ollama</h3><p>要使用 Ollama，我们首先需要<a href="https://ollama.com/download">下载它</a>。Ollama 支持 Linux、Windows 和 macOS，因此只需<a href="https://ollama.com/download"> 在这里 下载与你的操作系统兼容的 Ollama 版本即可 。</a>安装好 Ollama 后，我们可以从支持的 LLM<a href="https://ollama.com/library">列表</a>中选择一个模型。在本例中，我们将使用<a href="https://ollama.com/library/llama3.2">llama3.2</a> 模型，这是一个通用的多语言模型。在设置过程中，您将启用 Ollama 的命令行工具。下载完成后，您就可以运行下面一行：</p>ollama pull llama3.2<p>将输出</p>pulling manifest
pulling dde5aa3fc5ff... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 2.0 GB
pulling 966de95ca8a6... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 1.4 KB
pulling fcc5a6bec9da... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 7.7 KB
pulling a70ff7e570d9... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 6.0 KB
pulling 56bb8bd477a5... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏   96 B
pulling 34bb5ab01051... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏  561 B
verifying sha256 digest
writing manifest
success<p>安装完成后，可以使用此命令进行测试：</p>ollama run llama3.2<p>我们来提个问题：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc12f240920e897f5/6a17f39425daab32a508a367/ad1eff81c1b04d2a747c3afd0ecbc215e5bd96fd-800x501.gif" alt="运行 Ollama 并向它提问" /><p>模型运行后，Ollama 会启用一个默认在"11434" 端口运行的 API。让我们按照<a href="https://github.com/ollama/ollama/blob/main/docs/api.md">官方文档</a>，向该应用程序接口提出请求：</p>curl http://localhost:11434/api/generate -d '{                                          
  "model": "llama3.2",               
  "prompt": "What is the capital of France?"
}' <p>这就是我们得到的答复：</p>{"model":"llama3.2","created_at":"2024-11-28T21:48:42.152817532Z","response":"The","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.251884485Z","response":" capital","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.347365913Z","response":" of","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.446837322Z","response":" France","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.542367394Z","response":" is","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.644580384Z","response":" Paris","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.739865362Z","response":".","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.834347518Z","response":"","done":true,"done_reason":"stop","context":[128006,9125,128007,271,38766,1303,33025,2696,25,6790,220,2366,18,271,128009,128006,882,128007,271,3923,374,279,6864,315,9822,30,128009,128006,78191,128007,271,791,6864,315,9822,374,12366,13],"total_duration":6948567145,"load_duration":4386106503,"prompt_eval_count":32,"prompt_eval_duration":1872000000,"eval_count":8,"eval_duration":684000000}<p><em>请注意，该端点的特定响应是流式响应。</em></p><h3>使用 ngrok 将终端接入互联网</h3><p>由于我们的端点在本地环境中运行，因此无法通过互联网从另一个点（如我们的弹性云实例）进行访问。<a href="https://ngrok.com">ngrok</a>允许我们公开提供公共 IP 的端口。在 ngrok 中创建账户，并按照官方<a href="https://dashboard.ngrok.com/get-started/setup">设置指南</a>进行操作。</p><p>安装并配置好 ngrok 代理后，我们就可以公开 Ollama 正在使用的端口：</p>ngrok http 11434 --host-header="localhost:11434"<p><em>注意： </em><em><code>--host-header="localhost:11434"</code></em>头 <em> 保证请求中的"Host" 头与"localhost:11434 匹配。"</em></p><p>执行该命令将返回一个公共链接，只要 ngrok 和 Ollama 服务器在本地运行，该链接就能正常工作。</p>Session Status                online                                                                                                                                                                              
Account                       xxxx@yourEmailProvider.com (Plan: Free)                                                                                                                                             
Version                       3.18.4                                                                                                                                                                              
Region                        United States (us)                                                                                                                                                                  
Latency                       561ms                                                                                                                                                                               
Web Interface                 http://127.0.0.1:4040                                                                                                                                                               
Forwarding                    https://your-ngrok-url.ngrok-free.app -&gt; http://localhost:11434                                                                                                                   


Connections                   ttl     opn     rt1     rt5     p50     p90                                                                                                                                         
                              0       0       0.00    0.00    0.00    0.00                                                ```<p>在"Forwarding" 中，我们可以看到 ngrok 生成了一个 URL。留着以后用吧。</p><p>让我们再次尝试使用 ngrok 生成的 URL 向端点发出 HTTP 请求：</p>curl https://your-ngrok-endpoint.ngrok-free.app/api/generate -d '{                                          
  "model": "llama3.2",               
  "prompt": "What is the capital of France?"
}'<p>答复应与前一个答复类似。</p><h2>创建映射</h2><h3>ELSER 端点</h3><p>在本示例中，我们将<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/put-inference-api.html">使用 Elasticsearch 推理 API 创建一个推理端点</a>。此外，我们还将使用<a href="https://www.elastic.co/cn/guide/en/machine-learning/current/ml-nlp-elser.html">ELSER</a>生成嵌入。</p>PUT _inference/sparse_embedding/medicines-inference
{
  "service": "elasticsearch",
  "service_settings": {
    "num_allocations": 1,
    "num_threads": 1,
    "model_id": ".elser_model_2_linux-x86_64"
  }
}<p>在这个例子中，我们假设有一家药店出售两种药物：</p><ul><li><p>需要处方的药物。</p></li><li><p>无需处方的药物。</p></li></ul><p>这些信息将包含在每种药物的描述字段中。</p><p>LLM 必须对该字段进行解释，因此这就是我们要使用的数据映射：</p>PUT medicines
{
  "mappings": {
    "properties": {
      "name": {
        "type": "text",
        "copy_to": "semantic_field"
      },
      "semantic_field": {
        "type": "semantic_text",
        "inference_id": "medicines-inference"
      },
      "text_description": {
        "type": "text",
        "copy_to": "semantic_field"
      }
    }
  }
}<p>字段<code>text_description</code> 将存储描述的纯文本，而作为<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/semantic-text.html">语义_文本字段</a>类型的<code>semantic_field</code> 将存储由 ELSER 生成的嵌入。</p><p>属性<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/copy-to.html">copy_to</a>将把字段名和<code>text_description</code> 中的内容复制到语义字段中，以便为这些字段生成嵌入内容。</p><h2>索引数据</h2><p>现在，让我们使用<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/docs-bulk.html">_bulk API</a> 为数据建立索引。</p>POST _bulk
{"index":{"_index":"medicines"}}
{"id":1,"name":"Paracetamol","text_description":"An analgesic and antipyretic that does NOT require a prescription."}
{"index":{"_index":"medicines"}}
{"id":2,"name":"Ibuprofen","text_description":"A nonsteroidal anti-inflammatory drug (NSAID) available WITHOUT a prescription."}
{"index":{"_index":"medicines"}}
{"id":3,"name":"Amoxicillin","text_description":"An antibiotic that requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":4,"name":"Lorazepam","text_description":"An anxiolytic medication that strictly requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":5,"name":"Omeprazole","text_description":"A medication for stomach acidity that does NOT require a prescription."}
{"index":{"_index":"medicines"}}
{"id":6,"name":"Insulin","text_description":"A hormone used in diabetes treatment that requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":7,"name":"Cold Medicine","text_description":"A compound formula to relieve flu symptoms available WITHOUT a prescription."}
{"index":{"_index":"medicines"}}
{"id":8,"name":"Clonazepam","text_description":"An antiepileptic medication that requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":9,"name":"Vitamin C","text_description":"A dietary supplement that does NOT require a prescription."}
{"index":{"_index":"medicines"}}
{"id":10,"name":"Metformin","text_description":"A medication used for type 2 diabetes that requires a prescription."}<p>响应：</p>{
   "errors": false,
   "took": 34732020848,
   "items": [
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "mYoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 0,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "mooeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 1,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "m4oeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 2,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "nIoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 3,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "nYoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 4,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "nooeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 5,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "n4oeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 6,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "oIoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 7,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "oYoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 8,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "oooeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 9,
     	"_primary_term": 1,
     	"status": 201
   	}
 	}
   ]
 }<h2>使用 Playground 提问</h2><p><a href="https://www.elastic.co/cn/guide/en/kibana/current/playground.html">Playground</a>是一款 Kibana 工具，可让您使用 Elasticsearch 索引和 LLM 提供商快速创建 RAG 系统。您可以阅读<a href="https://www.elastic.co/cn/search-labs/blog/playground-connectors-data-chat">本文</a>了解更多信息。</p><h3>将当地的法律硕士与游乐场连接起来</h3><p>我们首先需要创建一个连接器，使用我们刚刚创建的公共 URL。在 Kibana 中，转到<strong>搜索&gt;Playground</strong>，然后点击"连接到 LLM" 。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt22148eabfabf6d3f/6a17f3963e9e459f97ba15c6/1854f0808f8150e359fe62ba5d901d32a88d477c-1600x867.png" alt="将当地的法律硕士与奥拉玛游乐场联系起来" /><p>此操作将显示 Kibana 界面左侧的菜单。在那里，点击"OpenAI" 。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6e28194d9012f141/6a17f39725daab500a08a36b/c83d3c4d7035a518124ad7d22b38764db57b6800-933x1007.png" alt="选择一个连接器：打开 AI Ollama" /><p>现在我们可以开始配置 OpenAI 连接器了。</p><p>访问"Connector settings" ，并为 OpenAI 提供商选择"Other (OpenAI Compatible Service)" ：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9984dce6f78a7c08/6a17f3990b0bed0b7add36c4/ecfcdc4b575c309bd55b4e61ca0ddb348aa84f64-917x268.png" alt="为使用推理应用程序接口的 Ollama 设置连接器设置" /><p>现在，让我们配置其他字段。在本例中，我们将模型命名为"medicines-llm" 。在 URL 字段中，使用 ngrok 生成的 URL (/v1/chat/completions)。在"Default model" 字段中，选择"llama3.2" 。我们不会使用 API 密钥，因此只需输入任意文本即可：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8a24b93a39d380fb/6a17f39b96142a15c7eb1c3c/5d3b5027c8096cbe49fb740d70aa24e849611a9d-916x688.png" alt="添加设置" /><p>点击"保存" 并点击"添加数据源" 添加索引药物：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4f107a54d5be25f9/6a17f39d4b055deb9d432338/525113da59e902c8235f62bde8fb62371a63e11b-1579x753.png" alt="使用 Playground 添加数据源，以便向文档提问" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt03fb36fe05dcdbf5/6a17f39ebe608602f40048be/96138de0bbe2c2ac619f64889d3487df62739ca4-466x805.png" alt="添加查询数据" /><p>好极了现在，我们可以使用本地运行的 LLM 作为 RAG 引擎访问 Playground。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb1b27580107259b3/6a17f3a096142abefceb1c40/cfb48b33c70f4534ab77eb01f58008237f65e6f4-1600x851.png" alt="在 Playground 中选择模型设置" /><p>在测试之前，让我们给代理添加更具体的指令，并将发送给模型的文件数量增加到 10 份，以便答案有尽可能多的可用文件。由于使用了 copy_to 属性，上下文字段将是<code>semantic_field</code> ，其中包括药品的名称和描述。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4fbc97dc87c6fc62/6a17f3a1e8fbce052c3a1aa0/0c57c9c0e1a0e7b58fffdd3ef81d67d41e2990c4-580x806.png" alt="Elastic Playground 中的 Moel 设置" /><p>现在我们来问一个问题：<em><strong>没有处方可以购买氯硝西泮吗？</strong></em>看看会发生什么：</p><p>不出所料，我们得到了正确答案。</p><h3>后续步骤</h3><p>下一步是创建自己的应用程序！Playground 提供了一个 Python 代码脚本，你可以在自己的机器上运行，并根据自己的需要进行定制。例如，将其置于<a href="https://fastapi.tiangolo.com/">FastAPI</a>服务器之后，创建一个由用户界面使用的 QA 药品聊天机器人。</p><p>点击 Playground 右上方的 "<em><strong>查看代码</strong></em>"按钮即可找到该代码：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt42fe193b8aa08830/6a17f3a33e9e4569e8ba15ca/816bfd0e5f936ad65dbe719d5df10714e550a40b-380x121.png" alt="查看代码按钮" /><p>然后使用<em><strong>Endpoints&amp; API 密钥</strong></em>生成代码中所需的<code>ES_API_KEY</code> 环境变量。</p><p>本例的代码如下：</p>## Install the required packages
## pip install -qU elasticsearch openai
import os
from elasticsearch import Elasticsearch
from openai import OpenAI
es_client = Elasticsearch(
    "https://your-deployment.us-central1.gcp.cloud.es.io:443",
    api_key=os.environ["ES_API_KEY"]
)
openai_client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
)
index_source_fields = {
    "medicines": [
        "semantic_field"
    ]
}
def get_elasticsearch_results():
    es_query = {
        "retriever": {
            "standard": {
                "query": {
                    "nested": {
                        "path": "semantic_field.inference.chunks",
                        "query": {
                            "sparse_vector": {
                                "inference_id": "medicines-inference",
                                "field": "semantic_field.inference.chunks.embeddings",
                                "query": query
                            }
                        },
                        "inner_hits": {
                            "size": 2,
                            "name": "medicines.semantic_field",
                            "_source": [
                                "semantic_field.inference.chunks.text"
                            ]
                        }
                    }
                }
            }
        },
        "size": 3
    }
    result = es_client.search(index="medicines", body=es_query)
    return result["hits"]["hits"]
def create_openai_prompt(results):
    context = ""
    for hit in results:
        inner_hit_path = f"{hit['_index']}.{index_source_fields.get(hit['_index'])[0]}"
        ## For semantic_text matches, we need to extract the text from the inner_hits
        if 'inner_hits' in hit and inner_hit_path in hit['inner_hits']:
            context += '\n --- \n'.join(inner_hit['_source']['text'] for inner_hit in hit['inner_hits'][inner_hit_path]['hits']['hits'])
        else:
            source_field = index_source_fields.get(hit["_index"])[0]
            hit_context = hit["_source"][source_field]
            context += f"{hit_context}\n"
    prompt = f"""
  Instructions:
  - You are an assistant specializing in answering questions about the sale of medicines.
  - Answer questions truthfully and factually using only the context presented.
  - If you don't know the answer, just say that you don't know, don't make up an answer.
  - You must always cite the document where the answer was extracted using inline academic citation style [], using the position.
  - Use markdown format for code examples.
  - You are correct, factual, precise, and reliable.
  Context:
  {context}
  """
    return prompt
def generate_openai_completion(user_prompt, question):
    response = openai_client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "system", "content": user_prompt},
            {"role": "user", "content": question},
        ]
    )
    return response.choices[0].message.content
if __name__ == "__main__":
    question = "my question"
    elasticsearch_results = get_elasticsearch_results()
    context_prompt = create_openai_prompt(elasticsearch_results)
    openai_completion = generate_openai_completion(context_prompt, question)
    print(openai_completion)<p>要使其与 Ollama 兼容，必须更改 OpenAI 客户端，使其连接到 Ollama 服务器，而不是 OpenAI 服务器。您可以在这里找到 OpenAI 示例和兼容端点的完整列表。</p>openai_client = OpenAI(
    # you can use http://localhost:11434/v1/ if running this code locally.
    base_url='https://your-ngrok-url.ngrok-free.app/v1/',
    # required but ignored
    api_key='ollama',
)<p>在调用完成方法时，将模型更改为 llama3.2：</p>def generate_openai_completion(user_prompt, question):
    response = openai_client.chat.completions.create(
        model="llama3.2",
        messages=[
            {"role": "system", "content": user_prompt},
            {"role": "user", "content": question},
        ]
    )
    return response.choices[0].message.content<p>让我们补充一个问题：<em><strong>我可以在没有处方的情况下购买氯硝西泮吗？ </strong></em>至 Elasticsearch 查询：</p>def get_elasticsearch_results():
    es_query = {
        "retriever": {
            "standard": {
                "query": {
                    "nested": {
                        "path": "semantic_field.inference.chunks",
                        "query": {
                            "sparse_vector": {
                                "inference_id": "medicines-inference",
                                "field": "semantic_field.inference.chunks.embeddings",
                                "query": "Can I buy Clonazepam without a prescription?"
                            }
                        },
                        "inner_hits": {
                            "size": 2,
                            "name": "medicines.semantic_field",
                            "_source": [
                                "semantic_field.inference.chunks.text"
                            ]
                        }
                    }
                }
            }
        },
        "size": 3
    }
    result = es_client.search(index="medicines", body=es_query)
    return result["hits"]["hits"]<p>此外，我们还在完成调用中打印了一些内容，以便确认我们将 Elasticsearch 结果作为问题上下文的一部分发送：</p>if __name__ == "__main__":
    question = "Can I buy Clonazepam without a prescription?"
    elasticsearch_results = get_elasticsearch_results()
    context_prompt = create_openai_prompt(elasticsearch_results)
    print("========== Context Prompt START ==========")
    print(context_prompt)
    print("========== Context Prompt END ==========")
    print("========== Ollama Completion START ==========")
    openai_completion = generate_openai_completion(context_prompt, question)
    print(openai_completion)
    print("========== Ollama Completion END ==========")<p>现在运行命令</p><p><code>pip install -qU elasticsearch openai</code></p><p><code>python main.py</code></p><p>你应该看到这样的内容：</p>========== Context Prompt START ==========
  Instructions:
  - You are an assistant specializing in answering questions about the sale of medicines.
  - Answer questions truthfully and factually using only the context presented.
  - If you don't know the answer, just say that you don't know, don't make up an answer.
  - You must always cite the document where the answer was extracted using inline academic citation style [], using the position.
  - Use markdown format for code examples.
  - You are correct, factual, precise, and reliable.
  Context:
  Clonazepam
 ---
An antiepileptic medication that requires a prescription.A nonsteroidal anti-inflammatory drug (NSAID) available WITHOUT a prescription.
 ---
IbuprofenAn anxiolytic medication that strictly requires a prescription.
 ---
Lorazepam


========== Context Prompt END ==========
========== Ollama Completion START ==========
No, you cannot buy Clonazepam over-the-counter (OTC) without a prescription [1]. It is classified as a controlled substance in the United States due to its potential for dependence and abuse. Therefore, it can only be obtained from a licensed healthcare provider who will issue a prescription for this medication.
========== Ollama Completion END ==========<h2>结论</h2><p>在本文中，当我们将 Ollama 等工具与 Elasticsearch 推论 API 和 Playground 结合使用时，我们可以看到它们的强大功能和多功能性。</p><p>经过几个简单的步骤后，我们就拥有了一个可运行的 RAG 应用程序，它可以聊天，使用 LLM 在我们自己的基础设施中运行，成本为零。这也使我们能够对资源和敏感信息有更多的控制权，此外，我们还可以使用各种模型来完成不同的任务。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ollama-with-inference-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ollama-with-inference-api</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd9c8eb0fc946920e/6a17f3a46864a4b2fbb688f0/399b9ef527be633845fb6505b68132cc03bc9e09-1150x628.png" length="0" type="image/png"/>
    <pubDate>Fri, 14 Feb 2025 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>