<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Andre Luiz - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Andre Luiz - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/cn/search-labs/author/andre-luiz</link>
    </image>
    <link>https://www.elastic.co/cn/search-labs/author/andre-luiz</link>
    <atom:link href="https://www.elastic.co/cn/search-labs/rss/author/andre-luiz.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[cn]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 03:58:15 GMT</lastBuildDate>
  <item>
    <title><![CDATA[将嵌入映射到 Elasticsearch 字段类型：semantic_text、dense_vector、sparse_vector]]></title>
    <description><![CDATA[讨论如何以及何时使用 semantic_text、dense_vector 或 sparse_vector，以及它们与嵌入生成的关系。]]></description>
    <content:encoded><![CDATA[<p>多年来，利用嵌入式技术提高信息检索相关性和准确性的做法有了长足的发展。Elasticsearch 等工具已经通过密集向量、稀疏向量和语义文本等专门字段类型支持这类数据。不过，要取得良好的效果，必须了解如何将嵌入正确映射到可用的 Elasticsearch 字段类型：<code>semantic_text</code>、<code>dense_vector</code> 和<code>sparse_vector</code> 。</p><p>在本文中，我们将讨论这些字段类型、每种类型的使用时间，以及它们与嵌入生成和使用策略的关系，包括在索引和查询过程中的关系。</p><h2>密集矢量类型</h2><p>Elasticsearch 中的<code>dense_vector</code> 字段类型用于存储密集向量，密集向量是文本、图像和音频等数据的数字表示，其中几乎所有维度都是相关的。这些向量是使用 OpenAI、Cohere 或 Hugging Face 等平台提供的嵌入模型生成的，旨在捕捉数据的整体语义，即使数据与其他文档不共享确切术语。</p><p>在 Elasticsearch 中，稠密向量的维度最多可达 4096，具体取决于所使用的模型。例如，all-MiniLM-L6-v2 模型生成的向量有 384 个维度，而 OpenAI 的 text-embedding-ada-002 模型生成的向量有 1536 个维度。</p><p><code>dense_vector</code> 字段通常被用作存储这类嵌入的默认类型，当需要更多控制时，例如使用预生成向量、应用自定义相似性函数或与外部模型集成。</p><h3>何时以及为何使用 dense_vector 类型？</h3><p>密集向量非常适合捕捉句子、段落或整个文档之间的语义相似性。当目标是比较文本的整体含义时，即使它们不共享相同的术语，它们也能很好地发挥作用。</p><p>密集矢量字段非常适合已经拥有外部嵌入生成管道，使用 OpenAI、Cohere 或 Hugging Face 等平台提供的模型，并且只想手动存储和查询这些矢量的情况。这种类型的字段与嵌入模型具有很高的兼容性，在生成和查询方面具有充分的灵活性，允许您控制矢量的生成、索引和搜索使用方式。</p><p>此外，它还支持不同形式的语义搜索，在需要调整排名逻辑的情况下，可使用 k-NN 或 script_score 等查询。这些可能性使密集矢量成为 RAG（检索增强生成）、推荐系统和基于相似性的个性化搜索等应用的理想选择。</p><p>最后，该字段允许您自定义相关性逻辑，使用<code>cosineSimilarity</code> 、<code>dotProduct</code> 或<code>l2norm</code> 等函数，根据使用情况的需要调整排名。 </p><p>对于需要灵活性、定制化和与上述高级用例兼容的用户来说，密集矢量仍然是最佳选择。</p><h3>如何使用密集矢量类型查询？</h3><p>对定义为<strong><code>dense_vector</code></strong> 的字段的搜索使用 k 近邻查询。该查询负责查找密集向量与查询向量最接近的文档。下面举例说明如何将 k-NN 查询应用于密集向量场：</p>{
  "knn": {
    "field": "my_dense_vector",
    "k": 10,
    "num_candidates": 50,
    "query_vector": [/* vector generated by model */]
  }
}<p>除 k-NN 查询外，如果需要自定义文档评分，也可以使用 script_score 查询，将其与<strong>余弦相似度、点积或 l2norm 等</strong>向量比较函数相结合，以更可控的方式计算相关性。请看示例：</p>{
"script_score": {
    "query": { "match_all": {} },
    "script": {
      "source": "cosineSimilarity(params.query_vector,
'my_dense_vector') + 1.0",
      "params": {
        "query_vector": [/* vector */]
      }
    }
  }
}<p>如果您想深入了解，我建议您阅读《<a href="https://www.elastic.co/search-labs/blog/vector-search-set-up-elasticsearch">如何在 Elasticsearch 中设置向量搜索</a>》一文。</p><p></p><h2>稀疏矢量类型</h2><p><strong><code>sparse_vector</code></strong> 字段类型用于存储稀疏矢量，稀疏矢量是一种数值表示，其中大部分值为零，只有少数项具有重要权重。这种类型的向量在基于术语的模型中很常见，如 SPLADE 或 ELSER（弹性学习稀疏 EncodeR）。</p><h3>何时以及为何使用稀疏向量类型？</h3><p>当你需要更精确的词汇搜索而又不牺牲语义智能时，稀疏向量是理想的选择。它们将文本表示为标记/值对，只突出显示最相关的术语和相关权重，从而提供清晰度、控制和效率。</p><p>这类字段在根据术语生成向量时特别有用，例如在 ELSER 或 SPLADE 模型中，这些模型会根据每个标记在文本中的相对重要性为其分配不同的权重。</p><p>如果您想控制查询中特定词语的影响，稀疏向量类型允许您手动调整词语的权重，以优化结果的排名。</p><p>它的主要优点包括：搜索透明，因为可以清楚地了解为什么某个文件被认为是相关的；存储高效，因为只保存非零值的标记，而不像密集向量那样存储所有维度。</p><p>此外，稀疏向量是混合搜索策略的理想补充，甚至可以与密集向量相结合，将词汇精确性与语义理解相结合。</p><h3>如何使用稀疏向量类型查询？</h3><p><strong><code>sparse_vector</code></strong> 查询可让您根据标记/值格式的查询向量搜索文档。请看下面的查询示例：</p>{
  "query": {
    "sparse_vector": {
      "field": "field_sparse",
      "query_vector": {
        "token1": 0.6,
        "token2": 0.2,
        "token3": 0.9
      }
    }
  }
}<p>如果希望使用训练有素的模型，可以使用推理端点自动将查询文本转换为稀疏向量：</p>{
  "query": {
    "sparse_vector": {
      "field": "field_sparse",
      "inference_id": "the inference ID to produce the token/weights",
      "query": "search text"
    }
  }
}<p>要进一步探讨这一主题，我建议阅读《<a href="https://www.elastic.co/search-labs/blog/sparse-vector-embedding">用训练有素的 ML 模型理解稀疏向量嵌入</a>》。</p><h2>语义文本类型</h2><p><strong><code>semantic_text</code></strong> 字段类型是在 Elasticsearch 中使用语义搜索的最简单、最直接的方法。它通过一个推理端点，在索引和查询时自动处理嵌入生成。这意味着你不必担心手动生成或存储矢量的问题。</p><h3>何时以及为何使用语义文本？</h3><p><code>semantic_text</code> 字段非常适合那些希望以最少的技术投入、无需手动处理矢量即可开始工作的用户。该字段可自动执行嵌入生成和矢量搜索映射等步骤，使设置更快更方便。</p><p>如果您重视<strong>简单性和抽象性</strong>，就应该考虑使用<code>semantic_text</code> ，因为它<strong>消除了手动配置映射、嵌入生成和摄取管道的复杂性</strong>。只需选择推理模型，其余的就交给 Elasticsearch 处理。</p><p>其主要优势包括在索引和查询过程中<strong>自动生成嵌入</strong>，以及<strong>可随时使用的映射</strong>，该映射经过预先配置，可支持选定的推理模型。</p><p>此外，该领域还提供<strong>对自动分割长文本（文本分块）的本地支持</strong>，可将大文本分割成较小的段落，每个段落都有自己的嵌入，从而提高搜索精度。这极大地提高了工作效率，尤其是对于那些希望在不处理语义搜索底层工程的情况下快速实现价值的团队而言。</p><p>不过，虽然<code>semantic_text</code> 提供了速度和简便性，但这种方法也有一些局限性。它允许使用市场标准模型，只要这些模型可以作为 Elasticsearch 中的推理端点。但<strong>它不支持外部生成的嵌入</strong>，而<code>dense_vector</code> 字段则可以做到这一点。</p><p>如果您需要对向量的生成方式进行更多控制，希望使用自己的嵌入，或需要将多个字段结合起来以实现高级策略，<code>dense_vector</code> 和<code>sparse_vector</code> 字段可提供更多自定义或特定领域方案所需的灵活性。</p><h3>如何使用语义文本类型查询</h3><p>在<strong><code>semantic_text</code></strong> 之前，必须根据嵌入类型（密集或稀疏）使用不同的查询。<code>sparse_vector</code> 查询用于稀疏字段，而<code>dense_vector</code> 字段则需要 KNN 查询。</p><p>使用语义文本类型时，搜索是通过<a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-semantic-query">语义查询</a>进行的，查询会自动生成查询向量，并与索引文档的嵌入进行比较。<strong><code>semantic_text</code></strong> 类型允许您定义用于嵌入查询的推理端点，但如果未指定任何推理端点，则将对查询应用索引过程中使用的相同端点。</p>{
  "query": {
    "semantic": {
      "field": "semantic_text_field",
      "query": "search text"
    }
  }
}<p>要了解更多信息，我建议您阅读<a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text"> Elasticsearch 新语义_文本映射这 篇文章 ：简化语义搜索</a> 。</p><h2>结论</h2><p>在选择如何在 Elasticsearch 中映射嵌入时，必须了解要如何生成向量以及需要对向量进行何种程度的控制。如果您追求简单，语义文本字段可实现自动和可扩展的语义搜索，使其成为许多初始用例的理想选择。当需要更多控制、微调性能或与自定义模型集成时，密集矢量和稀疏矢量场可提供必要的灵活性。</p><p>理想的字段类型取决于您的使用案例、可用基础设施以及机器学习堆栈的成熟度。最重要的是，Elastic 提供了用于构建现代化和高度适应性搜索系统的工具。</p><h2>参考资料</h2><ul><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-text.html">语义文本字段类型</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/sparse-vector.html">稀疏矢量场类型</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html">密集矢量场类型</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-semantic-query.html">语义查询</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-sparse-vector-query.html">稀疏向量查询</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html">kNN 搜索</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">Elasticsearch 新语义文本映射：简化语义搜索</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/sparse-vector-embedding">用训练有素的 ML 模型理解稀疏向量嵌入</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/mapping-embeddings-to-elasticsearch-field-types</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/mapping-embeddings-to-elasticsearch-field-types</guid>
    <category><![CDATA[向量数据库]]></category>
    <category><![CDATA[映射]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt72cd3c2601b22886/6a17083b0c4857259901a9dc/f98fdff837db55b466780c0bae672aa6f6c3a966-1200x628.png" length="0" type="image/png"/>
    <pubDate>Tue, 13 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 ML 生成过滤器和切面]]></title>
    <description><![CDATA[探索在搜索体验中使用 ML 模型自动创建过滤器和切面与传统硬编码方法的利弊。]]></description>
    <content:encoded><![CDATA[<p>过滤器和切面是用于完善搜索结果的机制，可帮助用户更快地找到相关内容或产品。在传统方法中，规则是人工定义的。例如，在电影目录中，流派等属性是预定义的，可用于筛选器和切面。另一方面，通过人工智能模型，可以自动从电影特征中提取新的属性，使整个过程更加动态和个性化。在本博客中，我们将探讨每种方法的优缺点，重点介绍它们的应用和挑战。</p><h2>筛选器与分面</h2><p>在开始之前，我们先来定义一下什么是过滤器和切面。<strong>过滤器</strong>是用于限制结果集的预定义属性。例如，在市场中，甚至在进行搜索之前就可以使用筛选器。用户可以先选择一个类别，如<strong>"Video games"</strong> ，然后再搜索<strong>"PS5"</strong> ，将搜索范围缩小到更具体的子集，而不是整个数据库。这大大增加了获得更多相关结果的机会。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0a77b38aae238938/6a170b821949f72a52e7aa51/5ed8868fa5017d034e1273e35c884a5430afdf3c-1600x937.png" alt="筛选" /><p><strong>面板的</strong>工作原理与筛选器类似，但只有在执行搜索后才可用。换句话说，搜索会返回结果，并根据这些结果生成新的细化选项列表。例如，在搜索 PS5 游戏机时，可以显示<strong>存储容量</strong>、<strong>运输成本</strong>和<strong>颜色</strong>等信息，帮助用户选择理想的产品。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt166e356b80d423ef/6a170b840e2e494ca341a10f/c5633fcc5b6fbb916110faf32144d8d43572e33a-1600x937.png" alt="面面观 " /><p>既然我们已经定义了过滤器和切面，下面我们就来讨论经典方法和基于机器学习 (ML) 的方法对其实施和使用的影响。每种方法都有影响搜索效率的优势和挑战。</p><h2>过滤器和分面的经典方法</h2><p>在这种方法中，过滤器和切面是根据预定义规则手动定义的。这意味着，考虑到目录结构和用户需求，可用于细化搜索的属性是固定的，并事先进行了规划。</p><p>例如，在市场中，"Electronics" 或"Fashion" 等类别可能有特定的筛选条件，如品牌、格式和价格范围。这些规则是静态创建的，可确保搜索体验的一致性，但每当出现新的产品或类别时，就需要进行手动调整。</p><p>虽然这种方法提供了对所显示过滤器和面的可预测性和控制，但当出现需要动态改进的新趋势时，这种方法就会受到限制。</p><p><strong>优点</strong></p><ul><li><p><strong>可预测性和控制：</strong>由于筛选器和切面是手动定义的，因此管理变得更加容易。</p></li><li><p><strong>低复杂性：</strong>无需训练模型。</p></li><li><p><strong>易于维护：</strong>由于规则是预定义的，因此可以快速进行调整和修正。</p></li></ul><p><strong>缺点</strong></p><ul><li><p><strong>新过滤器需要重新索引：</strong>每当需要使用新属性作为筛选器时，就必须对整个数据集重新索引，以确保文档包含该信息。</p></li><li><p><strong>缺乏动态适应性：</strong>过滤器是静态的，不能根据用户行为的变化自动调整。</p></li></ul><h3>滤波器/滤面的实现 - 经典方法</h3><p>在<strong>开发工具 Kibana</strong> 中，我们将使用<strong>经典方法</strong>创建过滤器/面板演示。</p><p>首先，我们定义映射来构建索引：</p>PUT videogames
{
  "mappings": {
    "properties": {
      "name": { "type": "text" },
      "brand": { "type": "keyword" },
      "storage": { "type": "keyword" },
      "price": { "type": "float" },
      "description": { "type": "text" }
    }
  }
}<p><strong>品牌</strong>和<strong>存储</strong> <strong>字段</strong>被设置为关键字，可直接用于聚合<strong>（面</strong>）。<strong>价格</strong>字段为<strong>浮动</strong>类型，可以创建<strong>价格范围</strong>。</p><p>下一步，将对产品数据编制索引：</p>POST videogames/_bulk
{ "index": { "_id": 1 } }
{ "name": "Play Station 5", "brand": "Sony", "storage": "1TB", "price": 499.99, "description": "Stunning Gaming: Marvel at stunning graphics and experience the features of the new PS5. Breathtaking Immersion: Discover a deeper gaming experience with support for haptic feedback, adaptive triggers, and 3D Audio technology. Slim Design: With the PS5 Digital Edition, gamers get powerful gaming technology in a sleek, compact design. 1TB of Storage: Have your favorite games ready and waiting for you to play with 1TB of built-in SSD storage. Backward Compatibility and Game Boost: The PS5 console can play over 4,000 PS4 games. With Game Boost, you can even enjoy faster, smoother frame rates in some of the best PS4 console games." }
{ "index": { "_id": 2 } }
{ "name": "Xbox Series X", "brand": "Microsoft", "storage": "1TB", "price": 499.99, "description": "Fastest, most powerful Xbox console ever. Play thousands of titles: Every game looks and plays better on Xbox Series X. At the heart of Series X is the Xbox Velocity. Architecture, which combines a custom SSD and built-in software to significantly reduce load times in and out of game. Switch between multiple games in an instant with Quick Resume. Explore new worlds and experience the action like never before with an unparalleled 12 teraflops of graphics processing power. Enjoy 4K gaming at up to 120 frames per second, premium advanced 3D sound, and more. 4K at 120 FPS: requires compatible content and display X version - with disc drive" }
{ "index": { "_id": 3 } }
{ "name": "Nintendo Switch", "brand": "Nintendo", "storage": "512GB", "price": 299.99, "description": "SHARPER, VIBRANT VISUALS. The new 7-inch screen on the Nintendo Switch OLED takes your gaming to the next level: vibrant colors with sharp contrasts for every moment. INTEGRATED GAMEPLAY. Enjoy the console's many multiplayer modes and connect with other players. Online or locally, the fun on the Nintendo Switch is guaranteed. ENJOY IMMERSION FOR LONGER. In addition to delivering an unparalleled experience, thanks to its improved audio, the Nintendo Switch has a rechargeable battery while you play. From 4.5 hours to 9 hours of battery life. INCLUDES SUPER MARIO BROS. WONDER. Transform your world with the phenomenal flowers in this new Mario game, full of amazing adventures, power-ups and new abilities. NINTENDO SWITCH ONLINE SUBSCRIPTION. Access online games, play with friends and enjoy the exclusive benefits of the Nintendo Switch Online subscription." }
{ "index": { "_id": 4 } }
{ "name": "Steam Deck", "brand": "Valve", "storage": "512GB", "price": 399.99, "description": "You can save games, apps, photos and videos without worrying about space. High-Level Performance: The 4-core processor and graphics ensure a dynamic experience and fast responses. High-Definition Images: Smooth transitions and sharp images provide complete immersion in the game. Wireless Connectivity: Wi-Fi technology allows you to play wherever you want, without wires or cables limiting your fun" }
{ "index": { "_id": 5 } }
{ "name": "Nintendo Switch Lite", "brand": "Nintendo", "storage": "512GB", "price": 299.99, "description": "MADE TO BE PORTABLE. Nintendo Switch Lite is designed specifically for portable gaming. The console lets you jump into your favorite games wherever you are. COMPACT AND LIGHTWEIGHT. With its sleek, lightweight design, this console is ready to hit the road wherever you are. COMPATIBLE GAMES. The Nintendo Switch Lite system plays the library of Nintendo Switch games that work in handheld mode. A WORLD OF COLOR TO CHOOSE FROM. Available in a variety of vibrant and unique colors, Nintendo Switch Lite lets you bring even more personality wherever you go." }<p>现在，让我们按照品牌、存储空间和价格范围对结果进行分组，从而检索出经典的面孔。在查询中，定义了 size:0。在这种情况下，目标是只检索聚合结果，而不包括与查询相对应的文档。</p>POST videogames/_search
{
  "size": 0,
  "aggs": {
    "brands": {
      "terms": { "field": "brand" }
    },
    "storage_sizes": {
      "terms": { "field": "storage" }
    },
    "price_ranges": {
      "range": {
        "field": "price",
        "ranges": [
          { "to": 300 },   
          { "from": 300, "to": 500 },  
          { "from": 500 }  
        ]
      }
    }
  }
}<p>回复将包括<strong>品牌</strong>、<strong>存储</strong>和<strong>价格的</strong>计数，有助于创建筛选器和面。</p>"aggregations": {
   "brands": {
     "doc_count_error_upper_bound": 0,
     "sum_other_doc_count": 0,
     "buckets": [
       {
         "key": "Microsoft",
         "doc_count": 1
       },
       {
         "key": "Nintendo",
         "doc_count": 1
       },
       {
         "key": "Sony",
         "doc_count": 1
       },
       {
         "key": "Valve",
         "doc_count": 1
       }
     ]
   },
   "storage_sizes": {
     "doc_count_error_upper_bound": 0,
     "sum_other_doc_count": 0,
     "buckets": [
       {
         "key": "1TB",
         "doc_count": 2
       },
       {
         "key": "512GB",
         "doc_count": 2
       }
     ]
   },
   "price_ranges": {
     "buckets": [
       {
         "key": "*-300.0",
         "to": 300,
         "doc_count": 1
       },
       {
         "key": "300.0-500.0",
         "from": 300,
         "to": 500,
         "doc_count": 3
       },
       {
         "key": "500.0-*",
         "from": 500,
         "doc_count": 0
       }
     ]
   }
 }<h2>基于机器学习/人工智能的筛选器和分面方法</h2><p>在这种方法中，机器学习（ML）模型（包括人工智能（AI）技术）分析数据属性，生成相关的过滤器和面。ML/AI 不依赖预定义规则，而是利用索引数据特征。这样就能动态发现新的切面和过滤器。</p><p><strong>优点</strong></p><ul><li><p><strong>自动更新：</strong>自动生成新的过滤器和切面，无需手动调整。</p></li><li><p><strong>发现新属性：</strong>它可以将<strong>以前未考虑过的 </strong>数据特征识别为过滤器，从而丰富搜索体验。</p></li><li><p><strong>减少人工操作：</strong>当人工智能从可用数据中学习时，团队无需不断定义和更新过滤规则。</p></li></ul><p><strong>缺点</strong></p><ul><li><p><strong>维护复杂性：</strong>使用模型可能需要预先验证，以确保生成的过滤器的一致性。</p></li><li><p><strong>需要 ML 和 AI 专业知识：</strong>该解决方案需要合格的专业人员来微调和监控模型性能。</p></li><li><p><strong>无关过滤器的风险：</strong>如果模型没有得到很好的校准，可能会生成对用户无用的切面。</p></li><li><p><strong>成本：</strong>使用 ML 和 AI 可能需要第三方服务，从而增加运营成本。</p></li></ul><p>值得注意的是，即使有了校准良好的模型和精心制作的提示，生成的切面仍应经过审查步骤。这种验证可以是手动的，也可以基于审核规则，以确保内容的适当性和安全性。虽然这不一定是一个缺点，但这是一个重要的考虑因素，以确保在提供给用户之前，面的质量和适用性。</p><h3>实施过滤器/面板--人工智能方法</h3><p>在本演示中，我们将使用一个人工智能模型来自动分析产品特性并提出相关属性建议。有了结构良好的提示，我们就能从目录中提取信息，并将其转化为过滤器和切面。下面，我们将介绍这一过程的每个步骤。</p><p>最初，我们将使用<strong>推理 API</strong>注册一个端点，以便与 ML 服务集成。以下是与<strong>OpenAI 服务</strong>集成的示例。</p>PUT _inference/completion/generate_filter_ia
{
   "service": "openai",
   "service_settings": {
       "api_key": "your-key",
       "model_id": "gpt-4o-mini"
   }
}<p>现在，我们定义一个管道来执行提示并获取模型生成的新过滤器。</p>PUT /_ingest/pipeline/generate_filter_ai
{
   "processors": [
     {
       "script": {
         "source": """ctx.prompt = "You are an expert in data organization for search and product categorization. Your task is to analyze the following product and identify the best dynamic facets that can be used in an e-commerce search experience. Product: " + ctx.name + "description: " + ctx.description + "Instructions: - Analyze the product name and description. - Extract only the dynamic facets (technological features or product characteristics that can be inferred from the description, try to create max 3 facets by characteristics found). Put the values into an array. Using key and value, e.g. dynamic_facets: [{ \"name\": \"Gaming Experience\", \"value\": \"Haptic Feedback\" },{ \"name\": \"Gaming Experience\", \"value\": \"Adaptive Triggers\" } - Return only a JSON."
         """
       }
     },
     {
       "inference": {
         "model_id": "generate_filter_ia",
         "input_output": {
           "input_field": "prompt",
           "output_field": "result"
         }
       }
     },
     {
       "gsub": {
         "field": "result",
         "pattern": "```json",
         "replacement": ""
       }
     },
     {
       "json" : {
         "field" : "result",
         "strict_json_parsing": false,
         "add_to_root" : true
       }
     },
     {
       "remove": {
         "field": "result"
       }
     },
     {
       "remove": {
         "field": "prompt"
       }
     }
   ]
}<p>为"PlayStation 5" 产品运行该流水线的模拟，说明如下：</p><p><em>令人惊叹的游戏：惊叹于令人惊叹的画面，体验全新 PS5 的功能。</em></p><p><em>令人惊叹的沉浸感：支持触觉反馈、自适应触发器和 3D 音频技术，探索更深层次的游戏体验。</em></p><p><em>超薄设计：通过 PS5 数字版，玩家可以在时尚、紧凑的设计中获得强大的游戏技术。</em></p><p><em>1TB 存储空间：内置 1TB SSD 存储空间，让您随时随地畅玩最喜爱的游戏。</em></p><p><em>向后兼容和游戏提升：PS5 游戏机可播放 4,000 多款 PS4 游戏。有了 Game Boost，您甚至可以在一些最好的 PS4 游戏机游戏中享受更快、更流畅的帧率。</em></p><p>让我们观察一下这次模拟产生的提示输出。</p>{
 "docs": [
   {
     "doc": {
       "_index": "index",
       "_version": "-3",
       "_id": "1",
       "_source": {
         "name": "Play Station 5",
         "result": """```json
{
 "dynamic_facets": [
   { "name": "Storage Capacity", "value": "1TB SSD" },
   { "name": "Graphics Technology", "value": "Stunning Graphics" },
   { "name": "Audio Technology", "value": "3D Audio" }
 ]
}
```""",
         "description": "Stunning Gaming: Marvel at stunning graphics and experience the features of the new PS5. Breathtaking Immersion: Discover a deeper gaming experience with support for haptic feedback, adaptive triggers, and 3D Audio technology. Slim Design: With the PS5 Digital Edition, gamers get powerful gaming technology in a sleek, compact design. 1TB of Storage: Have your favorite games ready and waiting for you to play with 1TB of built-in SSD storage. Backward Compatibility and Game Boost: The PS5 console can play over 4,000 PS4 games. With Game Boost, you can even enjoy faster, smoother frame rates in some of the best PS4 console games.",
         "model_id": "generate_filter_ia",
         "prompt": """You are an expert in data organization for search and product categorization. Your task is to analyze the following product and identify the best dynamic facets that can be used in an e-commerce search experience. Product: Play Station 5description: Stunning Gaming: Marvel at stunning graphics and experience the features of the new PS5. Breathtaking Immersion: Discover a deeper gaming experience with support for haptic feedback, adaptive triggers, and 3D Audio technology. Slim Design: With the PS5 Digital Edition, gamers get powerful gaming technology in a sleek, compact design. 1TB of Storage: Have your favorite games ready and waiting for you to play with 1TB of built-in SSD storage. Backward Compatibility and Game Boost: The PS5 console can play over 4,000 PS4 games. With Game Boost, you can even enjoy faster, smoother frame rates in some of the best PS4 console games.Instructions: - Analyze the product name and description. - Extract only the dynamic facets (technological features or product characteristics that can be inferred from the description, try create max 3 facets by characteristics found). Put the values like arrays. Using key and value, e.g. dynamic_facets: [{ "name": "Gaming Experience", "value": "Haptic Feedback" },{ "name": "Gaming Experience", "value": "Adaptive Triggers" } - Return only a JSON."""
       },
       "_ingest": {
         "timestamp": "2025-03-19T22:14:32.0161803Z"
       }
     }
   }
 ]
}<p>现在，新索引中将添加一个新字段<strong>dynamic_facets</strong>，用于存储人工智能生成的面。</p>PUT videogames_1
{
 "mappings": {
   "properties": {
     "name": { "type": "text" },
     "brand": { "type": "keyword" },
     "storage": { "type": "keyword" },
     "price": { "type": "float" },
     "description": { "type": "text" },
     "dynamic_facets": { "type": "nested",
     "properties": { "name": { "type": "keyword" },
                     "value": { "type": "keyword" } } }
   }
 }
}<p>我们将使用<strong>Reindex API</strong> 将<strong>videogames</strong>索引重新编入<strong>videogames_1</strong>，并在此过程中应用<strong>generate_filter_ai</strong>管道。该管道将在索引编制过程中自动生成动态切面。</p>POST _reindex?wait_for_completion=false
{
 "source": {
   "index": "videogames"
 },
 "dest": {
   "index": "videogames_1",
   "pipeline": "generate_filter_ai"
 }
}<p>现在，我们将运行搜索并获得新的筛选器：</p>GET videogames_1/_search
{
 "size": 0,
 "query": {
   "match": {
     "name": "nintendo"
   }
 },
 "aggs": {
   "dynamic_facets": {
     "nested": {
       "path": "dynamic_facets"
     },
     "aggs": {
       "facets": {
         "terms": {
           "field": "dynamic_facets.name"
         },
         "aggs": {
           "facets": {
             "terms": {
               "field": "dynamic_facets.value"
             }
           }
         }
       }
     }
   }
 }
}<p>结果</p>"aggregations": {
   "dynamic_facets": {
     "doc_count": 3,
     "facets": {
       "doc_count_error_upper_bound": 0,
       "sum_other_doc_count": 0,
       "buckets": [
         {
           "key": "Frame Rate",
           "doc_count": 1,
           "facets": {
             "doc_count_error_upper_bound": 0,
             "sum_other_doc_count": 0,
             "buckets": [
               {
                 "key": "120 FPS",
                 "doc_count": 1
               }
             ]
           }
         },
         {
           "key": "Gaming Resolution",
           "doc_count": 1,
           "facets": {
             "doc_count_error_upper_bound": 0,
             "sum_other_doc_count": 0,
             "buckets": [
               {
                 "key": "4K",
                 "doc_count": 1
               }
             ]
           }
         },
         {
           "key": "Graphics Processing Power",
           "doc_count": 1,
           "facets": {
             "doc_count_error_upper_bound": 0,
             "sum_other_doc_count": 0,
             "buckets": [
               {
                 "key": "12 Teraflops",
                 "doc_count": 1
               }
             ]
           }
         }
       ]
     }
   }
 }<p>下面是一个简单的前端，以表示面的实现：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb0d6aa40caf7a91a/6a170b86ab7f0839afdb9eb6/12b6d9d4f4d0985848d92841545fd22b7253ae6d-1600x1288.png" alt="执行方面" /><p><a href="https://gist.github.com/andreluiz1987/06d9ec1b381e942e9def0e969bd811a0">这里</a>提供了用户界面代码。</p><h2>结论</h2><p>这两种创建过滤器和切面的方法各有利弊。基于手动规则的传统方法可提供控制并降低成本，但需要不断更新，且无法动态适应新产品或新功能。</p><p>另一方面，基于人工智能和机器学习的方法可以自动提取切面，使搜索更加灵活，并且无需人工干预即可发现新的属性。不过，这种方法的实施和维护可能更为复杂，需要进行校准以确保结果的一致性。</p><p>在传统方法和基于人工智能的方法之间做出选择，取决于企业的需求和复杂程度。对于数据属性稳定且可预测的简单场景，传统方法可以更高效、更易于维护，从而避免基础设施和人工智能模型的不必要成本。另一方面，使用 ML/AI 提取切面可以大大增加价值，改善搜索体验，使过滤更加智能。</p><p>重要的是要评估自动化是否值得投资，或者更传统的解决方案是否已能有效满足业务需求。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/filters-facets-using-ml</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/filters-facets-using-ml</guid>
    <category><![CDATA[相关性]]></category>
    <category><![CDATA[ML 研究]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4084864dcdaa25d3/6a170b880c485781f901aaa9/6f196643d573614fe5124705c7e4db9bfce004b0-1200x628.png" length="0" type="image/png"/>
    <pubDate>Thu, 03 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[在 Elasticsearch 中使用 Amazon Nova 模型]]></title>
    <description><![CDATA[了解如何在 Elasticsearch 中使用 Amazon Nova 模型，自动从产品评论中提取情感倾向、真实性判断、内容摘要及关键词。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将讨论亚马逊的人工智能模型系列 Amazon Nova，并学习如何将其与 Elasticsearch 结合使用。</p><h2>关于亚马逊新星</h2><p>Amazon Nova 是亚马逊人工智能模型系列，可在亚马逊 Bedrock 上使用，旨在提供高性能和高成本效益。这些模型可使用文本、图像和视频输入，生成文本输出，并针对不同的精度、速度和成本需求进行了优化。</p><h3>亚马逊 Nova 主要型号</h3><ul><li><p>亚马逊 Nova Micro：该机型专门针对文本，速度快、性价比高，是翻译、推理、代码补全和解决数学问题的理想之选。其生成速度超过每秒 200 个令牌，非常适合需要即时响应的应用。</p></li><li><p>Amazon Nova Lite：一款低成本多模态模型，能够快速处理图像、视频及文本数据。该模型以其速度和准确性脱颖而出，适用于成本因素显著的大流量交互式应用场景。</p></li><li><p>亚马逊新星专业版：最先进的选择，集高精度、高速度和高性价比于一身。是视频摘要、问答、软件开发和人工智能代理等复杂任务的理想选择。专家评论证明了它在文本和视觉理解方面的卓越表现，以及它遵循指令和执行自动工作流程的能力。</p></li></ul><p>亚马逊 Nova 模型适用于各种应用，从内容创建和数据分析到软件开发和人工智能驱动的流程自动化。</p><p>下面，我们将演示如何结合 Elasticsearch 使用 Amazon Nova 模型进行自动产品评论分析。</p><p>我们将做什么</p><ol><li><p>通过 Inference API 创建一个端点，将 Amazon Bedrock 与 Elasticsearch 集成在一起。</p></li><li><p>使用推理处理器创建一个管道，该管道将调用推理 API 端点。</p></li><li><p>索引产品评论，并使用管道自动生成评论分析。</p></li><li><p>分析整合结果。</p></li></ol><h2>使用 Amazon Nova Lite 在推理 API 中创建终端</h2><p>首先，我们配置 Inference API，将 Amazon Bedrock 与 Elasticsearch 集成。我们定义了亚马逊 No<strong>va Lite</strong>，id 为amazon.nova-lite-v1:0、因为它在速度、准确性和成本之间取得了平衡。</p><p><strong>注意：</strong>使用 Amazon Bedrock 需要有效凭证。您可以<a href="https://docs.aws.amazon.com/keyspaces/latest/devguide/create.keypair.html">在此处</a>查看获取访问密钥的文档：</p>PUT _inference/completion/bedrock_completion_amazon_nova-lite
{
   "service": "amazonbedrock",
   "service_settings": {
       "access_key": "#access_key#",
       "secret_key": "#secret_key#",
       "region": "us-east-1",
       "provider": "amazontitan",
       "model": "amazon.nova-lite-v1:0"
   }
}<h2>创建审查分析管道</h2><p>现在，我们创建一个处理管道，使用推理处理器来执行审查分析提示。此提示将把评论数据发送到 Amazon Nova Lite，由其执行：</p><ul><li><p>情绪分类（积极、消极或中性）。</p></li><li><p>审查总结。</p></li><li><p>关键词生成。</p></li><li><p>真实性测量（真实 | 可疑 | 一般）。</p></li></ul>PUT /_ingest/pipeline/review_analyzer_ai
{
      "processors": [
      {
        "script": 
            {
            "source": """ctx.prompt = "Analyze the following product review and return a structured JSON. Task: - Summarize the review concisely. - Detect and classify the sentiment as positive, neutral, or negative.- Generate relevant tags (keywords) based on the review content and detected sentiment. - Evaluate the authenticity of the review (authentic, suspicious, or generic). Review: " + ctx.review + " Respond in JSON format with the following fields: \"review_analyze\": {\"sentiment\": \"&lt;positive | neutral | negative&gt;\", \"authenticity\": \"&lt;authentic | suspicious | generic&gt;\",\"summary\": \"&lt;short review summary&gt;\", \"keywords\": [\"&lt;keyword 1&gt;\", \"&lt;keyword 2&gt;\", \"...\"]}}}"
            """
            }
      },
      {
        "inference": {
          "model_id": "bedrock_completion_amazon_nova-lite",
          "input_output": {
            "input_field": "prompt",
            "output_field": "result"
          }
        }
      },
      {
        "gsub": {
          "field": "result",
          "pattern": "```json",
          "replacement": ""
        } 
      },
      {
        "json" : {
          "field" : "result",
          "strict_json_parsing": false,
          "add_to_root" : true
        }
      },
      {
        "remove": {
          "field": "result"
        }
      },
      {
        "remove": {
          "field": "prompt"
        }
      }
    ]
}<h2>索引审查</h2><p>现在，我们使用批量 API 对产品评论进行索引。先前创建的管道将自动应用，将 Nova 模型生成的分析添加到索引文档中。</p>POST bulk/
{ "index": { "_index" : "products", "_id": 1, "pipeline":"review_analyzer_ai" } }
{ "product": "Pampers Pants Premium Care Fralda", "review": "Best diaper ever! Great material, lots of cotton, without all that plastic. Doesn't leak! My baby is a boy and every diaper leaked around the waist, this model solved the problem. Even on a small baby it's worth the effort of putting on the short diaper. I put it on my baby at 9 pm and only take it off in the morning, without any leaks." }
{ "index": { "_index" : "products", "_id": 2, "pipeline":"review_analyzer_ai" } }
{ "product": "Portable Electric Body Massager", "review": "It broke in three months for no apparent reason, thank goodness I didn't review it before. I don't recommend buying it because it has a short lifespan." }
{ "index": { "_index" : "products", "_id": 3, "pipeline":"review_analyzer_ai" } }
{ "product": "Havit Fuxi-H3 Black Quad-Mode Wired and Wireless Gaming Headset", "review": "The sound is good for the price, but the connectivity is horrible. You always need to be playing audio, otherwise it loses connection (I work from home, and this is very annoying). Sometimes it loses connection and you have to turn it off and on again to get it back on. The microphone is very sensitive, so it loses connection frequently and you have to turn the headset off and on for the microphone to work again. The flexibility of the stem is useless, because if you move it, the microphone can turn off. Sometimes I need to use Linux and the headset simply doesn't work. It's light and comfortable, the sound is adequate, but the connectivity is terrible." }
{ "index": { "_index" : "products", "_id": 4, "pipeline":"review_analyzer_ai" } }
{ "product": "Air Fryer 4L Oil Free Fryer Mondial", "review": "For those looking for value for money, it's a good option, but the tray (which is underneath the perforated basket) is already peeling a lot. My mother has one just like it and said that hers is even rusting, in other words, the material is MUCH inferior. There's also something that bothers me, because it looks like a microwave, it doesn't fry evenly, it's weaker in the middle and stronger on the sides. Buy at your own risk." }<h2>查询和分析结果</h2><p>最后，我们运行一个查询，看看亚马逊 Nova Lite 模型是如何对评论进行分析和分类的。通过运行 GET products/_search，我们可以获得已经用评论内容生成的字段充实过的文档。</p><p>该模型可识别主要情绪（正面、中性或负面），生成简明摘要，提取相关关键词，并估计每条评论的真实性。这些字段有助于了解客户的意见，而无需阅读全文。</p><p>为了解释结果，我们研究了</p><ul><li><p>情感，表示消费者对产品的总体看法。</p></li><li><p>摘要，突出了所述要点。</p></li><li><p>关键词，可用于对类似评论进行分组或识别反馈模式。</p></li><li><p>真实性，表示评论是否可信。这对策划或管理非常有用。</p></li></ul>   "hits": [
      {
        "_index": "products",
        "_id": "1",
        "_score": 1,
        "_ignored": [
          "review.keyword"
        ],
        "_source": {
          "product": "Pampers Pants Premium Care Fralda",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The reviewer praises the diaper for its great material, high cotton content, and leak-proof design, especially highlighting its effectiveness for their baby.",
            "sentiment": "positive",
            "keywords": [
              "best diaper",
              "great material",
              "cotton",
              "no plastic",
              "leak-proof",
              "baby",
              "effective"
            ],
            "authenticity": "authentic"
          },
          "review": "Best diaper ever! Great material, lots of cotton, without all that plastic. Doesn't leak! My baby is a boy and every diaper leaked around the waist, this model solved the problem. Even on a small baby it's worth the effort of putting on the short diaper. I put it on my baby at 9 pm and only take it off in the morning, without any leaks."
        }
      },
      {
        "_index": "products",
        "_id": "2",
        "_score": 1,
        "_source": {
          "product": "Portable Electric Body Massager",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The product broke in three months for no apparent reason and the reviewer does not recommend it due to its short lifespan.",
            "sentiment": "negative",
            "keywords": [
              "broke",
              "short lifespan",
              "not recommend"
            ],
            "authenticity": "authentic"
          },
          "review": "It broke in three months for no apparent reason, thank goodness I didn't review it before. I don't recommend buying it because it has a short lifespan."
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 1,
        "_ignored": [
          "review.keyword"
        ],
        "_source": {
          "product": "Havit Fuxi-H3 Black Quad-Mode Wired and Wireless Gaming Headset",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The headset has good sound quality for the price but suffers from poor connectivity, especially when using the microphone or moving the headset. It also has compatibility issues with Linux.",
            "sentiment": "negative",
            "keywords": [
              "sound",
              "connectivity",
              "microphone",
              "compatibility",
              "annoying",
              "turn off and on",
              "Linux",
              "flexible stem",
              "work from home"
            ],
            "authenticity": "authentic"
          },
          "review": "The sound is good for the price, but the connectivity is horrible. You always need to be playing audio, otherwise it loses connection (I work from home, and this is very annoying). Sometimes it loses connection and you have to turn it off and on again to get it back on. The microphone is very sensitive, so it loses connection frequently and you have to turn the headset off and on for the microphone to work again. The flexibility of the stem is useless, because if you move it, the microphone can turn off. Sometimes I need to use Linux and the headset simply doesn't work. It's light and comfortable, the sound is adequate, but the connectivity is terrible."
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 1,
        "_ignored": [
          "review.keyword"
        ],
        "_source": {
          "product": "Air Fryer 4L Oil Free Fryer Mondial",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The product offers value for money but has issues with peeling, rusting, and uneven frying.",
            "sentiment": "negative",
            "keywords": [
              "value for money",
              "peeling",
              "rusting",
              "uneven frying",
              "weaker in the middle"
            ],
            "authenticity": "authentic"
          },
          "review": "For those looking for value for money, it's a good option, but the tray (which is underneath the perforated basket) is already peeling a lot. My mother has one just like it and said that hers is even rusting, in other words, the material is MUCH inferior. There's also something that bothers me, because it looks like a microwave, it doesn't fry evenly, it's weaker in the middle and stronger on the sides. Buy at your own risk."
        }
      }
    ]<h2>总结</h2><p>Amazon Nova Lite 与 Elasticsearch 的整合展示了语言模型如何将原始评论转化为结构化的有价值信息。通过管道处理评论，我们能够自动、一致地提取情感、真实性、摘要和关键词。</p><p>结果表明，该模型可以理解评论的上下文，对用户意见进行分类，并突出每个体验中最相关的要点。这将创建一个更丰富的数据集，可用于提高搜索能力。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/amazon-nova-models-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/amazon-nova-models-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdbbf13eb690294f1/6a17fddd6df73195190a115a/304713c48b568e17d0bb56b19edb28769f7801b3-721x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 02 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何使用我们的同义词 API 自动生成和上传同义词]]></title>
    <description><![CDATA[了解如何使用 LLM 自动识别和生成同义词，从而以编程方式将术语加载到 Elasticsearch 同义词 API 中。]]></description>
    <content:encoded><![CDATA[<p>提高搜索结果的质量对于提供高效的用户体验至关重要。优化搜索的方法之一是通过同义词自动扩展查询词。这样就可以更广泛地解释查询，涵盖各种语言，从而改进结果匹配。</p><p>本博客将探讨如何使用大型语言模型 (LLM) 自动识别和生成同义词，并允许以编程方式将这些术语加载到 Elasticsearch 的同义词 API 中。</p><h2>何时使用同义词？</h2><p>与矢量搜索相比，使用同义词是一种更快、更具成本效益的解决方案。它的实现较为简单，因为它不需要深厚的嵌入知识，也不需要复杂的矢量摄取过程。</p><p>此外，由于矢量搜索需要更大的存储容量和内存来嵌入索引和检索，因此资源消耗较低。</p><p>另一个重要方面是搜索区域化。有了同义词，就可以根据当地语言和习俗调整术语。这在嵌入式可能无法匹配区域表达或特定国家术语的情况下非常有用。例如，有些单词或缩略语在不同地区可能有不同的含义，但当地用户自然会将其视为同义词。在巴西，这种情况非常普遍。"Abacaxi" 和"ananás" 是同一种水果（菠萝），但在东北部的一些地区，第二个术语更常用。同样，东南部著名的"pão francês" 在东北部可能被称为"pão careca" 。</p><h2>如何使用 LLM 生成同义词？</h2><p>为了自动获取同义词，我们可以使用 LLM，它可以分析术语的上下文，并建议适当的变体。这种方法可以动态扩展同义词，确保搜索范围更广、更准确，而无需依赖固定词典。</p><p>在本演示中，我们将使用 LLM 生成电子商务产品的同义词。由于查询词的变化，许多搜索结果很少或没有结果。有了同义词，我们就可以解决这个问题。例如，搜索"智能手机" 可以涵盖不同型号的手机，确保用户找到所需的产品。</p><h3>准备工作</h3><p>在开始之前，我们需要设置环境并定义所需的依赖关系。我们将使用 Elastic 提供的解决方案，<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">在 Docker 中本地运行 Elasticsearch 和 Kibana</a>。代码将使用 Python 3.9.6 版本编写，依赖关系如下：</p>pip install openai==1.59.8 elasticsearch==8.15.1<h3>创建产品索引</h3><p>最初，我们将创建一个不支持同义词的产品索引。这样我们就可以验证查询，然后将其与包含同义词的索引进行比较。</p><p>为了创建索引，我们在 Kibana DevTools 中使用以下命令批量加载产品数据集：</p>POST _bulk
{"index": {"_index": "products", "_id": 10001}}
{"category": "Electronics", "name": "iPhone 14 Pro"}
{"index": {"_index": "products", "_id": 10007}}
{"category": "Electronics", "name": "MacBook Pro 16-inch"}
{"index": {"_index": "products", "_id": 10013}}
{"category": "Electronics", "name": "Samsung Galaxy Tab S8"}
{"index": {"_index": "products", "_id": 10037}}
{"category": "Electronics", "name": "Apple Watch Series 8"}
{"index": {"_index": "products", "_id": 10049}}
{"category": "Electronics", "name": "Kindle Paperwhite"}
{"index": {"_index": "products", "_id": 10067}}
{"category": "Electronics", "name": "Samsung QLED 4K TV"}
{"index": {"_index": "products", "_id": 10073}}
{"category": "Electronics", "name": "HP Spectre x360 Laptop"}
{"index": {"_index": "products", "_id": 10079}}
{"category": "Electronics", "name": "Apple AirPods Pro"}
{"index": {"_index": "products", "_id": 10115}}
{"category": "Electronics", "name": "Amazon Echo Show 10"}
{"index": {"_index": "products", "_id": 10121}}
{"category": "Electronics", "name": "Apple iPad Air"}
{"index": {"_index": "products", "_id": 10127}}
{"category": "Electronics", "name": "Apple AirPods Max"}
{"index": {"_index": "products", "_id": 10151}}
{"category": "Electronics", "name": "Sony WH-1000XM4 Headphones"}
{"index": {"_index": "products", "_id": 10157}}
{"category": "Electronics", "name": "Google Pixel 6 Pro"}
{"index": {"_index": "products", "_id": 10163}}
{"category": "Electronics", "name": "Apple MacBook Air"}
{"index": {"_index": "products", "_id": 10181}}
{"category": "Electronics", "name": "Google Pixelbook Go"}
{"index": {"_index": "products", "_id": 10187}}
{"category": "Electronics", "name": "Sonos Beam Soundbar"}
{"index": {"_index": "products", "_id": 10199}}
{"category": "Electronics", "name": "Apple TV 4K"}
{"index": {"_index": "products", "_id": 10205}}
{"category": "Electronics", "name": "Samsung Galaxy Watch 4"}
{"index": {"_index": "products", "_id": 10211}}
{"category": "Electronics", "name": "Apple MacBook Pro 16-inch"}
{"index": {"_index": "products", "_id": 10223}}
{"category": "Electronics", "name": "Amazon Echo Dot (4th Gen)"}<h3>用 LLM 生成同义词</h3><p>在这一步中，我们将使用 LLM 来动态生成同义词。为此，我们将整合 OpenAI 应用程序接口，定义适当的模型和提示。LLM 将接收产品类别和名称，确保同义词与上下文相关。</p>import json
import logging

from openai import OpenAI

def call_gpt(prompt, model):
    try:
        logging.info("generate synonyms by llm...")
        response = client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            temperature=0.7,
            max_tokens=1000
        )
        content = response.choices[0].message.content.strip()
        return content
    except Exception as e:
        logging.error(f"Failed to use model: {e}")
        return None

def generate_synonyms(category, products):
   synonyms = {}

   for product in products:
       prompt = f"You are an expert in generating synonyms for products. Based on the category and product name provided, generate synonyms or related terms. Follow these rules:\n"
       prompt += "1. **Format**: The first word should be the main item (part of the product name, excluding the brand), followed by up to 3 synonyms separated by commas.\n"
       prompt += "2. **Exclude the brand**: Do not include the brand name in the synonyms.\n"
       prompt += "3. **Maximum synonyms**: Generate a maximum of 3 synonyms per product.\n\n"
       prompt += f"The category is: **{category}**, and the product is: **{product}**. Return only the synonyms in the requested format, without additional explanations."

       response = call_gpt(prompt, "gpt-4o")
       synonyms[product] = response

   return synonyms<p>从创建的产品索引中，我们将检索"Electronics" 类别中的所有项目，并将其名称发送到 LLM。预期输出结果如下</p>{
  "iPhone 14 Pro": ["iPhone", "smartphone", "mobile", "handset"],
  "MacBook Pro 16-inch": ["MacBook", "Laptop", "Notebook", "Ultrabook"],
  "Samsung Galaxy Tab S8": ["Tab", "Tablet", "Slate", "Pad"],
  "Bose QuietComfort 35 Headphones": ["Headphones", "earphones", "earbuds", "headset"]
}<p>有了生成的同义词，我们就可以使用同义词 API 将其注册到 Elasticsearch 中。</p><h3>使用同义词 API 管理同义词</h3><p>同义词 API 提供了在系统内直接管理同义词集的有效方法。每个同义词集都由同义词规则组成，其中一组词在搜索中被视为等同词。</p><p><strong>创建同义词集示例</strong></p>PUT _synonyms/my-synonyms-set
{
  "synonyms_set": [
    {
      "id": "rule-1",
      "synonyms": "hello, hi"
    },
    {
      "synonyms": "bye, goodbye"
    }
  ]
}<p>
这样就创建了一个名为"my-synonyms-set," 的集合，其中"hello" 和"hi" 被视为等同词，"bye" 和"goodbye 也被视为等同词。"</p><h2>为产品目录创建同义词</h2><p>下面是建立同义词集并将其插入 Elasticsearch 的方法。同义词规则是根据 LLM 建议的同义词映射生成的。每条规则都有一个 ID（与 slug 格式的产品名称相对应）和 LLM 计算出的同义词列表。</p>import json
import logging

from elasticsearch import Elasticsearch
from slugify import slugify

es = Elasticsearch(
    "http://localhost:9200",
    api_key="your_api_key"
)

def mount_synonyms(results):
   synonyms_set = [{"id": slugify(product), "synonyms": synonyms} for product, synonyms in
                   results.items()]

   try:
       response = es.synonyms.put_synonym(id="products-synonyms-set",
                                                 synonyms_set=synonyms_set)

       logging.info(json.dumps(response.body, indent=4))
       return response.body
   except Exception as e:
       logging.error(f"Error create synonyms: {str(e)}")
       return None<p>下面是创建同义词集的请求有效载荷：</p>{
   "synonyms_set":[
      {
         "id": "iphone-14-pro",
         "synonyms": "iPhone, smartphone, mobile, handset"
      },
      {
         "id": "macbook-pro-16-inch",
         "synonyms": "MacBook, Laptop, Notebook, Computer"
      },
      {
         "id": "samsung-galaxy-tab-s8",
         "synonyms": "Tablet, Slate, Pad, Device"
      },
      {
         "id": "garmin-forerunner-945",
         "synonyms": "Forerunner, smartwatch, fitness watch, GPS watch"
      },
      {
         "id": "bose-quietcomfort-35-headphones",
         "synonyms": "Headphones, Earphones, Headset, Cans"
      }
   ]
}<p>在集群中创建同义词集后，我们就可以进行下一步，即使用定义的同义词集创建支持同义词的新索引。</p><p>下面是完整的 Python 代码，其中包含 LLM 生成的同义词和同义词 API 定义的同义词集创建：</p>import json
import logging

from elasticsearch import Elasticsearch
from openai import OpenAI
from slugify import slugify

logging.basicConfig(level=logging.INFO)

client = OpenAI(
   api_key="your-key",
)

es = Elasticsearch(
    "http://localhost:9200",
    api_key="your_api_key"
)


def call_gpt(prompt, model):
   try:
       logging.info("generate synonyms by llm...")
       response = client.chat.completions.create(
           model=model,
           messages=[{"role": "user", "content": prompt}],
           temperature=0.7,
           max_tokens=1000
       )
       content = response.choices[0].message.content.strip()
       return content
   except Exception as e:
       logging.error(f"Failed to use model: {e}")
       return None


def generate_synonyms(category, products):
   synonyms = {}

   for product in products:
       prompt = f"You are an expert in generating synonyms for products. Based on the category and product name provided, generate synonyms or related terms. Follow these rules:\n"
       prompt += "1. **Format**: The first word should be the main item (part of the product name, excluding the brand), followed by up to 3 synonyms separated by commas.\n"
       prompt += "2. **Exclude the brand**: Do not include the brand name in the synonyms.\n"
       prompt += "3. **Maximum synonyms**: Generate a maximum of 3 synonyms per product.\n\n"
       prompt += f"The category is: **{category}**, and the product is: **{product}**. Return only the synonyms in the requested format, without additional explanations."

       response = call_gpt(prompt, "gpt-4o")
       synonyms[product] = response

   return synonyms


def get_products(category):
   query = {
       "size": 50,
       "_source": ["name"],
       "query": {
           "bool": {
               "filter": [
                   {
                       "term": {
                           "category.keyword": category
                       }
                   }
               ]
           }
       }
   }
   response = es.search(index="products", body=query)

   if response["hits"]["total"]["value"] &gt; 0:
       product_names = [hit["_source"]["name"] for hit in response["hits"]["hits"]]
       return product_names
   else:
       return []


def mount_synonyms(results):
   synonyms_set = [{"id": slugify(product), "synonyms": synonyms} for product, synonyms in
                   results.items()]

   try:
       es_client = get_client_es()
       response = es_client.synonyms.put_synonym(id="products-synonyms-set",
                                                 synonyms_set=synonyms_set)

       logging.info(json.dumps(response.body, indent=4))
       return response.body
   except Exception as e:
       logging.error(f"Erro update synonyms: {str(e)}")
       return None


if __name__ == '__main__':
   category = "Electronics"
   products = get_products("Electronics")
   llm_synonyms = generate_synonyms(category, products)
   mount_synonyms(llm_synonyms)<h3>创建支持同义词的索引</h3><p>将创建一个新索引，对<code>products</code> 索引中的所有数据进行重新索引。该索引将使用<code>synonyms_filter</code> ，它应用了之前创建的<code>products-synonyms-set</code> 。</p><p>以下是配置为使用同义词的索引映射：</p>PUT products_02
{
  "settings": {
    "analysis": {
      "filter": {
        "synonyms_filter": {
          "type": "synonym",
          "synonyms_set": "products-synonyms-set",
          "updateable": true
        }
      },
      "analyzer": {
        "synonyms_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "synonyms_filter"
          ]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "ID": {
        "type": "long"
      },
      "category": {
        "type": "keyword"
      },
      "name": {
        "type": "text",
        "analyzer": "standard",
        "search_analyzer": "synonyms_analyzer"
      }
    }
  }
}<h3>重新索引<code>products</code> 索引</h3><p>现在，我们将使用<strong>Reindex API</strong>将<code>products</code> 索引中的数据迁移到包含同义词支持的新<code>products_02</code> 索引中。在 Kibana DevTools 中执行了以下代码：
</p>POST _reindex
{
  "source": {
    "index": "products"
  },
  "dest": {
    "index": "products_02"
  }
}<p>迁移后，<code>products_02</code> 索引将被填充，并可使用配置的同义词集验证搜索。</p><h3>使用同义词验证搜索</h3><p>让我们比较一下两个索引的搜索结果。我们将在两个索引上执行相同的查询，并验证是否使用同义词来检索结果。</p><h4>在<code>products</code> 索引中搜索（不含同义词）</h4><p>我们将使用 Kibana 执行搜索并分析结果。在分析&gt; 发现菜单中，我们将创建一个数据视图，以可视化我们创建的索引中的数据。</p><p>在 Discovery 中，单击数据视图并定义名称和索引模式。对于"<strong>产品</strong>" 索引，我们将使用"<strong>产品</strong>"模式。然后，我们将重复该过程，使用"<strong>products_02</strong><strong>"</strong>模式为"products_02" 索引创建一个新的数据视图。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte826fd932cfeb9df/6a17fdffec0f8912aa5a6841/3ad4a6891a3905e96532a312932fdf3a8216aec2-1600x599.png" alt="" /><p>配置好数据视图后，我们就可以返回 Analytics&gt; Discovery 并开始验证。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltba3729c60068e8a0/6a17fe01e9ea87ba2aa9c82a/422c4b2b51abae6580cad25085d1b8a365fc6b9e-1294x850.png" alt="" /><p>在这里，选择 DataView 产品并对"tablet" 一词进行搜索后，我们没有得到任何结果，尽管我们知道有"Kindle Paperwhite" 和"Apple iPad Air" 这样的产品。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2c6a383dd4cb0157/6a17fe02577262671d1bce0c/e4ae3a785fdd93f48d7c7d204185ded149126f2c-1600x862.png" alt="" /><h4>在<code>products_02</code> 索引中搜索（支持同义词）</h4><p>在支持同义词的"<strong>products_synonyms</strong>" 数据视图上执行相同查询时，产品被成功检索。这表明配置的同义词集工作正常，确保搜索词的不同变体都能返回预期结果。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt986b4706e2f70014/6a17fe043e9e454edbba16d3/e609749c39e90d5c82fa846af6124679dd62bcb8-1600x526.png" alt="" /><p>我们可以直接在 Kibana DevTools 中运行相同的查询来获得相同的结果。只需使用 Elasticsearch Search API 搜索 products_02 索引即可：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbe829a3c4d7aac60/6a17fe05e8fbce03d73a1bd7/504d0d1f96dcfbceb309063dc0716bcee64ad2f8-1600x870.png" alt="" /><h2>结论</h2><p>在 Elasticsearch 中使用同义词提高了产品目录搜索的准确性和覆盖范围。与众不同的关键在于使用了<strong>LLM</strong>，它可以根据上下文自动生成同义词，无需预定义清单。该模型分析了产品名称和类别，确保与电子商务相关的同义词。</p><p>此外，<strong>同义词 API</strong>简化了词典管理，允许动态修改同义词集。有了这种方法，搜索变得更加灵活，更能适应不同的用户查询模式。</p><p>这一过程可以通过新数据和模型调整不断改进，确保提供越来越高效的研究体验。</p><h2>参考资料</h2><p><strong>在本地运行 Elasticsearch</strong></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html</a></p><p><strong>同义词应用程序接口</strong></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/synonyms-apis.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/synonyms-apis.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-synonyms-automate</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-synonyms-automate</guid>
    <category><![CDATA[相关性]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0f0247b9bc1d1ccd/6a17fe07ec0f891c745a6845/05a3cfeaa387561d5334ca3f1609035ddfff7481-1200x628.png" length="0" type="image/png"/>
    <pubDate>Thu, 27 Mar 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何通过 Airbyte 将数据导入 Elasticsearch]]></title>
    <description><![CDATA[使用 Airbyte 将数据导入 Elasticsearch。我们将介绍前提条件、Airbyte 配置和逐步集成。]]></description>
    <content:encoded><![CDATA[<p>Airbyte 是一款数据集成工具，可让您以自动化和可扩展的方式将信息从不同来源转移到不同目的地。它使您能够从应用程序接口、数据库和其他系统中提取数据，并将其加载到 Elasticsearch 等平台中，从而提供高级搜索和高效分析。</p><p>在本文中，我们将介绍如何配置 Airbyte 以将数据摄取到 Elasticsearch，包括关键概念、前提条件和逐步集成。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt99eabff95c587c00/6a17e300e8fbce2d303a18a9/ea7af907dfd4c0b7e8673164467ee236623282d2-1360x802.png" alt="配置 Airbyte 以将数据输入 Elasticsearch" /><h2>Airbyte 基本概念</h2><p>Airbyte 的使用有几个基本概念。下面，我们将重点介绍其中的主要内容：</p><ul><li><p>来源：定义提取数据的来源。</p></li><li><p>目的地：定义数据的发送和存储位置。</p></li><li><p>连接：配置源和目标之间的关系，包括同步频率。</p></li></ul><h2>Airbyte 与 Elasticsearch 的集成</h2><p>在本演示中，我们将执行一个集成，将存储在 S3 存储桶中的数据迁移到 Elasticsearch 索引中。我们将展示如何在 Airbyte 中配置源（S3）和目标（Elasticsearch）。</p><h3>准备工作</h3><p>要观看这一演示，必须满足以下前提条件：</p><ol><li><p>在 AWS 中创建一个存储桶，用于存储包含数据的 JSON 文件。</p></li><li><p>使用 Docker<a href="https://docs.airbyte.com/using-airbyte/getting-started/oss-quickstart">在本地安装 Airbyte</a>。</p></li><li><p>在 Elastic Cloud 中创建一个 Elasticsearch 集群来存储输入的数据。</p></li></ol><p>下面，我们将详细介绍每个步骤。</p><h4>安装 Airbyte</h4><p>Airbyte 可以使用 Docker 在本地运行，也可以在云中运行，但使用时需要付费。在本演示中，我们将使用 Docker 本地版本。</p><p>安装可能需要几分钟时间。按照安装说明进行安装后，Airbyte 可在以下网址使用： http://localhost:8000。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0f9149be887b5500/6a17e302abe0f2b1c6dfe956/66147b5121413ad9baecb10c6886288e917f7c09-1600x1102.png" alt="安装 Airbyte" /><p></p><p>登录后，我们就可以开始配置集成了。</p><h4>创建水桶</h4><p>在此步骤中，您需要一个 AWS 账户来创建一个 S3 存储桶。此外，还必须通过创建策略和 IAM 用户来设置正确的权限，以允许访问存储桶。</p><p>在该桶中，我们将上传包含不同日志记录的 JSON 文件，这些记录稍后将迁移到 Elasticsearch。文件日志的内容是这样的</p>{
   "timestamp": "2025-02-15T14:00:12Z",
   "level": "INFO",
   "service": "data_pipeline",
   "message": "Pipeline execution started",
   "details": {
       "pipeline_id": "abc123",
       "source": "MySQL",
       "destination": "Elasticsearch"
   }
}<p>以下是装入邮筒的文件：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbf769c5e9bdb5dfe/6a17e3043e9e45a360ba13f7/f3e7f5889002e3a804121a880d97f1d93044e2f7-1600x680.png" alt="装入 Airbyte 文件桶的文件" /><h4>弹性云配置</h4><p>为了便于演示，我们将使用弹性云。如果您还没有账户，可在此处创建一个免费试用账户：<a href="https://cloud.elastic.co/registration">弹性云注册</a>。</p><p>在弹性云中配置部署后，您需要获得</p><ul><li><p>Elasticsearch 服务器的 URL。</p></li><li><p>访问 Elasticsearch 的用户。</p></li></ul><p>要获取 URL，请访问部署&gt; 我的部署，在应用程序中找到 Elasticsearch 并点击 "复制端点"。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltddfaaf863168e2bb/6a17e305faa913e29793c7e4/2e38c0bfb2cea83d9ef90dba0e673559fe358199-1368x1056.png" alt="弹性云配置" /><p>要创建用户，请按照以下步骤操作：</p><ol><li><p>访问 Kibana&gt; 堆栈管理&gt; 用户。</p></li><li><p>创建一个具有超级用户角色的新用户。</p></li><li><p>填写字段以创建用户。</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltca15b0d3084f4c67/6a17e3073e9e4507e2ba13fb/54d098805087a2d60772475cbb31784083a38c25-1600x1026.png" alt="在弹性云中创建用户" /><p>现在，我们已经设置好了一切，可以开始在 Airbyte 中配置连接器了。</p><h3>配置信号源连接器</h3><p>在这一步中，我们将为 S3 创建源连接器。为此，我们将进入 Airbyte 界面，在菜单中选择 "源 "选项。然后，我们将搜索 S3 连接器。下面，我们将详细介绍配置连接器所需的步骤：</p><ol><li><p>访问 Airbyte 并进入 "来源 "菜单。</p></li><li><p>搜索并选择 S3 连接器。</p></li><li><p>配置以下参数：</p><ol><li><p>源名称：定义数据源名称。</p></li><li><p>交付方法：选择复制记录（建议用于结构化数据）。</p></li><li><p>数据格式：选择 JSON 格式。</p></li><li><p>流名称：定义 Elasticsearch 中索引的名称。</p></li><li><p>存储桶名称：输入 AWS 中桶的名称。</p></li><li><p>AWS 访问密钥和 AWS 密钥：输入访问凭证。</p></li></ol></li></ol><p>点击 "设置来源"，等待验证。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdca1fb8d8f5d034f/6a17e30963baff75bc741bb2/f83566ab5ebad07ad9fd61dc20353ca5a95c8c92-1600x1099.png" alt="等待 Airbyte 和 Elasticsearch 数据摄取验证" /><h3>配置目的地连接器</h3><p>在这一步中，我们将配置目标连接器，即 Elasticsearch。为此，我们将进入菜单并选择 "目的地 "选项。然后，我们将搜索 Elasticsearch 并点击返回的结果。现在，我们将继续配置该连接：</p><ol><li><p>进入 Airbyte 并转到 "目的地 "菜单。</p></li><li><p>搜索并选择 Elasticsearch 连接器。</p></li><li><p>配置以下参数：</p><ol><li><p>验证方法：选择用户名/密码。</p></li><li><p>用户名和密码：使用在 Kibana 中创建的凭证。</p></li><li><p>服务器端点：粘贴从弹性云复制的 URL。</p></li></ol></li></ol><p>点击 "<strong>设置目的地</strong>"，等待验证。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt72056179d048e027/6a17e30b033c8d5d9d6bb115/f9246b54cc77e0589c0bf658ffc274fd28f766d5-1600x941.png" alt="在 Elastic Cloud 中为 Airbyte 数据摄取创建目的地" /><h3>创建源连接和目的地连接</h3><p>一旦创建了 "源 "和 "目标"，就会创建它们之间的连接，从而完成集成的创建。 </p><p>下面是创建连接的说明：</p><p>1.在菜单中，转到 "连接"，点击 "创建第一个连接"。</p><p>2.在下一个屏幕中，您可以选择一个现有的源或创建一个新源。由于我们已经创建了一个源，因此将选择源 S3。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc7f1e01215a7a328/6a17e30c63baff53d7741bb6/14937ccb2686bbde7cb99ffe7e13f7f4e35d7a31-1600x393.png" alt="在 Airbyte 中选择现有源或创建新源" /><p>3.下一步是选择目的地。由于我们已经创建了 Elasticsearch 连接器，因此将选择它来完成配置。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltece5c720f28eee00/6a17e30d505ac3dc68ad8aa3/d10962bc175f9ce0b2745ea913d739d3c40ba3b6-1600x431.png" alt="在 Airbyte 中选择目的地" /><p>下一步需要定义同步模式和使用的模式。由于只创建了日志模式，因此它将是唯一可供选择的选项。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7cbcccad9cfd5a8d/6a17e30f414c64d32d9450e9/d5d3a2e20038d17ef701822fe7ccc38d6e175c50-1600x805.png" alt="在 Airbyte 中定义同步模式" /><p>4.我们将进入配置连接步骤。在这里，我们可以定义连接名称和集成执行频率。频率有三种配置方式：</p><ul><li><p><strong>Cron</strong>：根据用户定义的 cron 表达式（例如 0 0 15 * * ?，每天 15:00）运行同步；</p></li><li><p><strong>预定</strong>：在指定的时间间隔内运行同步（例如每 24 小时、每 2 小时）；</p></li><li><p><strong>手动</strong>：手动： 手动运行同步。</p></li></ul><p>在本演示中，我们将选择手动选项。</p><p>最后，点击 "<strong>设置连接</strong>"，将建立源和目标之间的连接。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4f5fc0e51035b733/6a17e311af47b67ac8cddeff/1a0470f6dff341cb90e6ecfd8ae87a7d2243cbf4-1600x626.png" alt="单击 Airbyte 中的设置连接" /><h3>将数据从 S3 同步到 Elasticsearch</h3><p>返回 "连接 "屏幕后，就可以看到已创建的连接。要执行程序，只需单击 "同步"。从那时起，数据将开始从 S3 迁移到 Elasticsearch。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9b0ad7a7446f9d58/6a17e3121d1b83b4c593e3c3/c1ce54b129eafb2533b638e4df2b96f3a266ad62-1600x347.png" alt="在 Airbyte 上将数据从 S3 同步到 Elasticsearch" /><p>如果一切顺利，您将获得同步状态。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt857cdad5bd7fe03f/6a17e313be608646e00046b9/6fe9d5826787066bd8bf0f5e17905df9f8d698e6-1600x361.png" alt="将状态从 S3 同步到 Airbyte 中的 Elasticsearch" /><h3>在 Kibana 中可视化数据</h3><p>现在，我们将进入 Kibana 分析数据，检查索引是否正确。在 Kibana 发现部分，我们将创建一个名为日志的数据视图。这样，我们就能查看同步后创建的日志索引中的数据。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1bc137f7c04e50ee/6a17e3151480094f2eb486cf/6b6909647ac3187d0fee477cfd732741314f82cf-1600x566.png" alt="在 Kibana 中可视化数据：Airbyte 和 Elastic" /><p>现在，我们可以将索引数据可视化，并对其进行分析。这样，我们使用 Airbyte 验证了整个迁移流程，在 Airbyte 中加载了数据桶中的数据，并在 Elasticsearch 中编制了索引。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd8c0a36a83ee4bb1/6a17e317b1e1135ea679f20e/ca8e9b3d7f7112291af58acf514f4573b44037e8-1600x806.png" alt="Airbyte 和 Elastic：在 Kibana 中可视化索引数据并对其执行分析" /><h2>结论：Airbyte&amp; Elasticsearch 集成</h2><p>事实证明，Airbyte 是一种高效的数据集成工具，可以自动连接多个数据源和目的地。在本教程中，我们演示了如何将数据从 S3 存储桶摄取到 Elasticsearch 索引，并重点介绍了该过程的主要步骤。</p><p>这种方法有利于大量数据的摄取，并允许在 Elasticsearch 中进行分析，如复杂的搜索、聚合和数据可视化。</p><h2>参考资料</h2><p><strong>快速启动 Airbyte：</strong></p><p><a href="https://docs.airbyte.com/using-airbyte/getting-started/oss-quickstart#part-1-install-abctl">https://docs.airbyte.com/using-airbyte/getting-started/oss-quickstart#part-1-install-abctl</a></p><p><strong>核心概念：</strong></p><p><a href="https://docs.airbyte.com/using-airbyte/core-concepts/">https://docs.airbyte.com/using-airbyte/core-concepts/</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/airbyte-elasticsearch-ingest-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/airbyte-elasticsearch-ingest-data</guid>
    <category><![CDATA[索引数据]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5defe5e12935b233/6a17e318505ac3ed7fad8aa7/dce2bad9949006163af95ed05b5a1eacf5393dc7-1200x628.png" length="0" type="image/png"/>
    <pubDate>Fri, 14 Mar 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何通过 LlamaIndex 向 Elasticsearch 采集数据]]></title>
    <description><![CDATA[逐步介绍如何使用 RAG 和 LlamaIndex 采集数据并进行搜索。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将使用 LlamaIndex 来索引数据，为常见问题实现一个搜索引擎。Elasticsearch 将作为我们的矢量数据库，实现矢量搜索，而 RAG（Retrieval-Augmented Generation，检索增强生成）将丰富上下文，提供更准确的响应。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte5895fbc057ffc1b/6a17f4cf3e9e45288bba15ef/7ac65a686bdd76c145e903f5c3110c62875a525f-972x501.png" alt="LlamaIndex&amp; Elasticsearch：输入文档并构建常见问题搜索" /><h2>什么是 LlamaIndex？</h2><p>LlamaIndex 是一个框架，可帮助创建由大型语言模型（LLM）驱动的代理和工作流，以便与特定或私人数据进行交互。它可以将各种来源（应用程序接口、PDF、数据库）的数据与 LLM 集成，从而完成研究、信息提取和生成上下文化回复等任务。</p><p><strong>关键概念：</strong></p><ul><li><p>代理：使用 LLM 执行任务的智能助手，任务范围从简单的响应到复杂的操作。</p></li><li><p>工作流程：将代理、数据连接器和高级任务工具结合起来的多步骤流程。</p></li><li><p>语境增强：利用外部数据丰富 LLM 的技术，克服其训练局限性。</p></li></ul><p><strong>LlamaIndex</strong> <strong>与 Elasticsearch 集成：</strong></p><p>Elasticsearch 可以通过各种方式与 LlamaIndex 配合使用：</p><ul><li><p>数据源：使用 Elasticsearch 阅读器提取文档。</p></li><li><p>嵌入模型：将数据编码为向量，用于语义搜索。</p></li><li><p>矢量存储：将 Elasticsearch 用作搜索矢量化文档的存储库。</p></li><li><p>高级存储：配置文档摘要或知识图谱等结构。</p></li></ul><h2>使用 LlamaIndex 和 Elasticsearch 创建常见问题搜索 </h2><h3>数据准备</h3><p>我们将以<a href="https://www.elastic.co/guide/en/cloud/current/ec-faq-getting-started.html">Elasticsearch 服务常见问题解答</a>为例进行说明。每个问题都是从网站上提取的，并保存在一个单独的文本文件中。您可以使用任何方法来组织数据；在本例中，我们选择将文件保存在本地。</p><p>文件示例：</p>File Name: what-is-elasticsearch-service.txt
Content: Elasticsearch Service is hosted and managed Elasticsearch and Kibana brought to you by the creators of Elasticsearch. Elasticsearch Service is part of Elastic Cloud and ships with features that you can only get from the company behind Elasticsearch, Kibana, Beats, and Logstash. Elasticsearch is a full text search engine that suits a range of uses, from search on websites to big data analytics and more.<p>保存所有问题后，目录将如下所示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt43467eb9c5579103/6a17f4d02f4a5cc60bfa8a62/1f367d57f2e650334671a2c03156ef4412c0c615-962x704.png" alt="" /><h3>安装依赖项</h3><p>我们将使用 Python 语言实现摄取和搜索，我使用的版本是 3.9。作为前提条件，有必要安装以下依赖项：</p>llama-index-vector-stores-elasticsearch
llama-index
openai<p>Elasticsearch 和 Kibana 将使用 Docker 创建，并通过 docker-compose.yml 配置为运行 8.16.2 版本。这样就能更容易地创建本地环境。</p>version: '3.8'
services:

 elasticsearch:
   image: docker.elastic.co/elasticsearch/elasticsearch:8.16.2
   container_name: elasticsearch-8.16.2
   environment:
     - node.name=elasticsearch
     - xpack.security.enabled=false
     - discovery.type=single-node
     - "ES_JAVA_OPTS=-Xms1024m -Xmx1024m"
   ports:
     - 9200:9200
   networks:
     - shared_network

 kibana:
   image: docker.elastic.co/kibana/kibana:8.16.2
   container_name: kibana-8.16.2
   restart: always
   environment:
     - ELASTICSEARCH_URL=http://elasticsearch:9200
   ports:
     - 5601:5601
   depends_on:
     - elasticsearch
   networks:
     - shared_network

networks:
 shared_network:<h3>使用 LlamaIndex 进行文件摄取</h3><p>文件将使用 LlamaIndex 索引到 Elasticsearch 中。首先，我们使用<strong>SimpleDirectoryReader</strong> 加载文件，它允许从本地目录加载文件。加载文档后，我们将使用<strong>VectorStoreIndex 对</strong>其进行索引。</p>documents = SimpleDirectoryReader("./faq").load_data()

storage_context = StorageContext.from_defaults(vector_store=es)
index = VectorStoreIndex(documents, storage_context=storage_context, embed_model=embed_model)<p>LlamaIndex 中的矢量存储负责存储和管理文档嵌入。LlamaIndex 支持不同类型的向量存储，在本例中，我们将使用 Elasticsearch。在 StorageContext 中，我们配置 Elasticsearch 实例。由于上下文是本地的，因此不需要额外的参数。有关其他环境中的配置，请参阅文档检查必要的参数：<a href="https://docs.llamaindex.ai/en/stable/examples/vector_stores/ElasticsearchIndexDemo/#configuring-elasticsearchstore">ElasticsearchStore 配置</a>。</p><p>默认情况下，LlamaIndex 使用 OpenAI<strong>text-embedding-ada-002</strong>模型生成嵌入。不过，在本例中，我们将使用<strong>文本嵌入-3-小</strong>模型。需要注意的是，使用该模型需要一个 OpenAI API 密钥。</p><p>以下是完整的文件摄取代码。</p>import openai
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, StorageContext
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore

openai.api_key = os.environ["OPENAI_API_KEY"]

es = ElasticsearchStore(
   index_name="faq",
   es_url="http://localhost:9200"
)

def format_title(filename):
   filename_without_ext = filename.replace('.txt', '')
   text_with_spaces = filename_without_ext.replace('-', ' ')
   formatted_text = text_with_spaces.title()

   return formatted_text


embed_model = OpenAIEmbedding(model="text-embedding-3-small")

documents = SimpleDirectoryReader("./faq").load_data()

for doc in documents:
   doc.metadata['title'] = format_title(doc.metadata['file_name'])

storage_context = StorageContext.from_defaults(vector_store=es)
index = VectorStoreIndex(documents, storage_context=storage_context, embed_model=embed_model)<p>执行后，文件将被编入<strong>faq</strong>索引，如下图所示：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt391ae1bfc1daefbf/6a17f4d22f4a5c9887fa8a66/d59b85ed1f57bf80e84d6cb0d6d722a2d12ae4c0-1600x745.png" alt="" /><h3>使用 RAG 搜索</h3><p>要执行搜索，我们需要配置<strong>ElasticsearchStore</strong>客户端，用 Elasticsearch URL 设置<strong>index_name</strong>和<strong>es_url</strong>字段。在<strong>retrieval_strategy</strong> 中，我们为向量搜索定义了<strong>AsyncDenseVectorStrategy</strong>。还提供其他策略，如<strong>AsyncBM25Strategy</strong>（关键词搜索）和<strong>AsyncSparseVectorStrategy</strong>（稀疏向量）。更多详情，请参阅<a href="https://docs.llamaindex.ai/en/stable/api_reference/storage/vector_store/elasticsearch/">官方文档</a>。</p>es = ElasticsearchStore(
   index_name="faq",
   es_url="http://localhost:9200",
   retrieval_strategy=AsyncDenseVectorStrategy(
   )
)<p>接下来，我们将创建一个<strong>VectorStoreIndex</strong>对象，并使用 ElasticsearchStore 对象配置<strong>vector_store</strong>。通过<strong>as_retriever</strong>方法，我们可以搜索与查询最相关的文档，并通过<strong>similarity_top_k</strong>参数设置返回结果的数量为 5。</p>   index = VectorStoreIndex.from_vector_store(vector_store=es)
   retriever = index.as_retriever(similarity_top_k=5)
   results = retriever.retrieve(query)<p>下一步是 RAG。矢量搜索的结果被纳入本地语言管理器的格式化提示中，从而能够根据检索到的信息做出符合具体情况的响应。</p><p>在 PromptTemplate 中，我们定义了提示格式，其中包括</p><ul><li><p>Context ({context_str})：检索器检索到的文件。</p></li><li><p>Query ({query_str}): 用户的问题。</p></li><li><p>说明：指导模型根据环境做出反应，而不依赖外部知识。</p></li></ul>qa_prompt = PromptTemplate(
   "You are a helpful and knowledgeable assistant."
   "Your task is to answer the user's query based solely on the context provided below."
   "Do not use any prior knowledge or external information.\n"
   "---------------------\n"
   "Context:\n"
   "{context_str}\n"
   "---------------------\n"
   "Query: {query_str}\n"
   "Instructions:\n"
   "1. Carefully read and understand the context provided.\n"
   "2. If the context contains enough information to answer the query, provide a clear and concise answer.\n"
   "3. Do not make up or guess any information.\n"
   "Answer: "
)<p>最后，LLM 会处理提示，并根据上下文返回准确的回复。</p>llm = OpenAI(model="gpt-4o")
context_str = "\n\n".join([n.node.get_content() for n in results])
response = llm.complete(
   qa_prompt.format(context_str=context_str, query_str=query)
)

print("Answer:")
print(response)<p>完整代码如下：</p>es = ElasticsearchStore(
   index_name="faq",
   es_url="http://localhost:9200",
   retrieval_strategy=AsyncDenseVectorStrategy(
   )
)


def print_results(results):
   for rank, result in enumerate(results, start=1):
       title = result.metadata.get("title")
       score = result.get_score()
       text = result.get_text()
       print(f"{rank}. title={title} \nscore={score} \ncontent={text}")


def search(query: str):
   index = VectorStoreIndex.from_vector_store(vector_store=es)

   retriever = index.as_retriever(similarity_top_k=10)
   results = retriever.retrieve(QueryBundle(query_str=query))
   print_results(results)

   qa_prompt = PromptTemplate(
       "You are a helpful and knowledgeable assistant."
       "Your task is to answer the user's query based solely on the context provided below."
       "Do not use any prior knowledge or external information.\n"
       "---------------------\n"
       "Context:\n"
       "{context_str}\n"
       "---------------------\n"
       "Query: {query_str}\n"
       "Instructions:\n"
       "1. Carefully read and understand the context provided.\n"
       "2. If the context contains enough information to answer the query, provide a clear and concise answer.\n"
       "3. Do not make up or guess any information.\n"
       "Answer: "
   )

   llm = OpenAI(model="gpt-4o")
   context_str = "\n\n".join([n.node.get_content() for n in results])
   response = llm.complete(
       qa_prompt.format(context_str=context_str, query_str=query)
   )

   print("Answer:")
   print(response)


question = "Elastic services are free?"
print(f"Question: {question}")
search(question)<p>现在，我们可以执行搜索，例如"Elastic services are free?" ，并根据常见问题数据本身获得符合上下文的回复。</p>Question: Elastic services are free?
Answer:
Elastic services are not entirely free. However, there is a 14-day free trial available for exploring Elastic solutions. After the trial, access to features and services depends on the subscription level.<p>为了做出这一回应，我们使用了以下文件：</p>1. title=Can I Try Elasticsearch Service For Free 
score=1.0 
content=Yes, sign up for a 14-day free trial. The trial starts the moment a cluster is created.
During the free trial period get access to a deployment to explore Elastic solutions for Enterprise Search, Observability, Security, or the latest version of the Elastic Stack.

2. title=Do You Offer Elastic S Commercial Products 
score=0.9941274512218439 
content=Yes, all Elasticsearch Service customers have access to basic authentication, role-based access control, and monitoring.
Elasticsearch Service Gold, Platinum and Enterprise customers get complete access to all the capabilities in X-Pack: Security, Alerting, Monitoring, Reporting, Graph Analysis &amp; Visualization. Contact us to learn more.

3. title=What Is Elasticsearch Service 
score=0.9896776845746571 
content=Elasticsearch Service is hosted and managed Elasticsearch and Kibana brought to you by the creators of Elasticsearch. Elasticsearch Service is part of Elastic Cloud and ships with features that you can only get from the company behind Elasticsearch, Kibana, Beats, and Logstash. Elasticsearch is a full text search engine that suits a range of uses, from search on websites to big data analytics and more.

4. title=Can I Run The Full Elastic Stack In Elasticsearch Service 
score=0.9880631561979476 
content=Many of the products that are part of the Elastic Stack are readily available in Elasticsearch Service, including Elasticsearch, Kibana, plugins, and features such as monitoring and security. Use other Elastic Stack products directly with Elasticsearch Service. For example, both Logstash and Beats can send their data to Elasticsearch Service. What is run is determined by the subscription level.

5. title=What Is The Difference Between Elasticsearch Service And The Amazon Elasticsearch Service 
score=0.9835054890793161 
content=Elasticsearch Service is the only hosted and managed Elasticsearch service built, managed, and supported by the company behind Elasticsearch, Kibana, Beats, and Logstash. With Elasticsearch Service, you always get the latest versions of the software. Our service is built on best practices and years of experience hosting and managing thousands of Elasticsearch clusters in the Cloud and on premise. For more information, check the following Amazon and Elastic Elasticsearch Service comparison page.
Please note that there is no formal partnership between Elastic and Amazon Web Services (AWS), and Elastic does not provide any support on the AWS Elasticsearch Service.<h2>结论</h2><p>我们使用 LlamaIndex 演示了如何创建一个高效的常见问题搜索系统，并将 Elasticsearch 作为向量数据库提供支持。使用嵌入法对文件进行摄取和索引，从而实现矢量搜索。通过 PromptTemplate，搜索结果被纳入上下文并发送给 LLM，LLM 会根据检索到的文件生成精确且符合上下文的回复。</p><p>该工作流程将信息检索与根据上下文生成回复整合在一起，以提供准确、相关的结果。</p><h2>参考资料</h2><p><a href="https://www.elastic.co/guide/en/cloud/current/ec-faq-getting-started.html">https://www.elastic.co/guide/en/cloud/current/ec-faq-getting-started.html</a></p><p><a href="https://docs.llamaindex.ai/en/stable/api_reference/readers/elasticsearch/">https://docs.llamaindex.ai/en/stable/api_reference/readers/elasticsearch/</a></p><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_index/">https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_index/</a></p><p><a href="https://docs.llamaindex.ai/en/stable/examples/query_engine/custom_query_engine/">https://docs.llamaindex.ai/en/stable/examples/query_engine/custom_query_engine/</a></p><p><a href="https://www.elastic.co/search-labs/integrations/llama-index">https://www.elastic.co/search-labs/integrations/llama-index</a></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-llamaindex-ingest-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-llamaindex-ingest-data</guid>
    <category><![CDATA[索引数据]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf9f88afa92e90390/6a17f4d33e03d7987d4f2dd0/b8b760bfd8694df43fd74ba90ae5fc1edbe4ce76-1150x628.png" length="0" type="image/png"/>
    <pubDate>Fri, 28 Feb 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[分面搜索：利用人工智能改进搜索范围和结果]]></title>
    <description><![CDATA[探索如何在 Elasticsearch 中使用面搜索来快速缩小类别内的选项范围。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将探讨人工智能（AI），特别是使用 GPT-4 等高级语言模型，如何帮助创建更多的上下文切面，使它们对用户更加相关和有用。</p><p>切面搜索是电子商务平台的一个强大工具。它有助于根据显示项目的特征来组织和完善搜索结果。虽然滤镜经常被混淆，但切面的工作原理是不同的。筛选器是固定属性，由索引中始终存在的信息定义，如产品类别或格式。而面则是动态的，由执行搜索后返回的结果生成。</p><p>试想一下服装目录："类别" （如 T 恤、裤子）或"性别" （如男、女）等字段是帮助缩小结果范围的过滤器。而面则反映了结果中出现的产品的具体特征，如常见颜色、可用尺寸或材料。这使得搜索体验更具适应性和情境性。</p><p>下面是一张图片，我们与一个切面进行交互，可以看到由切面筛选出的搜索结果。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc4ce8ca7e5b62ad7/6a17f8ba6864a4534bb68979/74d2159706ab7248ebe5efddc74c882f0693db71-600x420.gif" alt="分面搜索示例" /><h2>人工智能如何改善面的生成</h2><p>人工智能通常与语义搜索和嵌入相关联，那么面呢？如何利用人工智能，使面面俱到对每次搜索更有用，并针对具体情况？</p><p>一个引人入胜的可能性是利用人工智能创建新的分类，超越索引中的传统分类。通过分析内容的具体特征，这些新类别可以提供更丰富、更精确的上下文，使面更相关、更符合用户需求。与原始文件类别相比，这使得结果的细化更有意义。</p><h2>人工智能如何完善电影分类以提高搜索效果</h2><p>让我们来分析一下以下影片，它们目前都被归类为剧情片类型：</p><ul><li><p>梦之安魂曲
简历：四个科尼岛人吸毒成瘾后，他们的乌托邦被打破了。</p></li><li><p>美国丽人
简历：一位在性方面受挫的郊区父亲在迷恋上女儿最好的朋友后，陷入了中年危机。</p></li><li><p>Good Will Hunting
简历：威尔-亨廷是麻省理工学院的看门人，他有数学天赋，但需要心理学家的帮助才能找到人生方向。</p></li></ul><p>这种类型划分无法捕捉到每部影片的细微差别或独特背景。通过利用人工智能分析故事梗概和中心主题，我们可以创建新的类别，更好地反映每部电影的真实背景。例如</p><ul><li><p>梦之安魂曲 - 新类别："成瘾与依赖"</p></li><li><p>美国丽人》 - 新类别："中年危机"</p></li><li><p>Good Will Hunting - 新类别："智力斗争"</p></li></ul><p>这些新类别使搜索更加精确，同时为用户提供了更有意义的筛选器来完善搜索结果。当原始分类过于笼统时，这种方法尤为有效，能让用户更轻松地准确找到他们要找的内容。</p><h2>使用 GPT-4 创建新类别：面搜索示例</h2><p>在这个例子中，我们将展示如何利用人工智能模型来创建新的电影类别，使其更加精确，并与每部作品的背景相一致。为了演示这一过程，我们将使用 Elastic 仿真管道和 OpenAI 推断服务。我们将创建一个由多个处理器组成的流水线，其中包括脚本处理器，它将负责创建提示语，并在推理处理器中执行，从而确定新的类别。其他处理器将用于处理流水线执行过程中产生的数据和辅助字段。值得一提的是，这一逻辑也可应用于其他类似的工具或模型。</p><p>首先，我们需要创建推理端点，将服务定义为 OpenAI、访问服务所需的令牌以及模型。在本例中，我使用的是 gpt-4-mini。有关 OpenAI 推断服务的更多详情，请点击<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-openai.html">此处</a>。</p>PUT _inference/completion/generate_topics_ia
{
    "service": "openai",
    "service_settings": {
        "api_key": "your-token",
        "model_id": "gpt-4o-mini"
    }
}<p>端点创建完成后，我们就可以用它来创建新类别了。下面是一个处理文档数据操作和提示生成整个过程的管道。我将详细解释每个处理器的功能。</p><p>第一个处理器将负责建立提示。明确详细的说明非常重要，这样人工智能才能正确分析和识别主题。在本提示中，我要求根据对电影标题、描述和类型的分析来确定 2 个主题。</p>{
        "script": {
          "source": """
            ctx.prompt = "You are an expert in semantic analysis and audiovisual content categorization. Your task is to generate only subcategories (max 2 topics) that describe specific aspects of movies based on their genres and descriptions. The output should be like: 'n1, n2, ...n'. Here is a movie info to analyze: Title: " + ctx.title  + "Genres: " + ctx.genres  + "Description: " + ctx.description;
          """
        }<p>下一个管道是推理管道，它将接收提示并将其发送到我们的<strong>generate_topics_ia</strong>端点。模型生成的响应将存储在结果字段中。</p>{
        "inference": {
          "model_id": "generate_topics_ia",
          "input_output": {
            "input_field": "prompt",
            "output_field": "result"
          }
        }
      }<p>接下来，除了删除我创建的临时字段外，我们还有 3 个处理器用于处理响应并将其设置到主题字段中。</p><p>执行该流水线后，我们将得到以下结果：</p>{
  "docs": [
    {
      "doc": {
        "_index": "index",
        "_version": "-3",
        "_id": "1",
        "_source": {
          "description": "While Frodo and Sam edge closer to Mordor with the help of the shifty Gollum, the divided fellowship makes a stand against Sauron's new ally, Saruman, and his hordes of Isengard.",
          "model_id": "generate_topics_ia",
          "title": "The Lord of the Rings: The Fellowship of the Ring",
          "genres": [
            "Action",
            "Adventure",
            "Drama"
          ],
          "topics": [
            "Fantasy",
            "Quest"
          ]
        },
        "_ingest": {
          "timestamp": "2024-11-22T17:51:51.340010257Z"
        }
      }
    },
    {
      "doc": {
        "_index": "index",
        "_version": "-3",
        "_id": "2",
        "_source": {
          "description": "A team of explorers travel through a wormhole in space in an attempt to ensure humanity's survival.",
          "model_id": "generate_topics_ia",
          "title": "Interstellar",
          "genres": [
            "Adventure",
            "Drama",
            "Sci-Fi"
          ],
          "topics": [
            "space exploration",
            "human survival"
          ]
        },
        "_ingest": {
          "timestamp": "2024-11-22T17:51:51.340413173Z"
        }
      }
    },
    {
      "doc": {
        "_index": "index",
        "_version": "-3",
        "_id": "3",
        "_source": {
          "description": "An astronaut becomes stranded on Mars after his team assume him dead, and must rely on his ingenuity to find a way to signal to Earth that he is alive.",
          "model_id": "generate_topics_ia",
          "title": "The Martian",
          "genres": [
            "Adventure",
            "Drama",
            "Sci-Fi"
          ],
          "topics": [
            "survival",
            "ingenuity"
          ]
        },
        "_ingest": {
          "timestamp": "2024-11-22T17:51:51.340427965Z"
        }
      }
    }
  ]
}<p>请注意，尽管有些影片最初属于同一类型，但我们有了与影片背景更加相关的新类别。</p><p>现在，我们可以使用这些新类别，并将它们与文档一起编入索引。这样，在生成切面时，除了主类别外，我们还可以根据影片的背景情况，生成更具体的子类别。</p><p>此外，还可以将这些新类别矢量化，并将其用于矢量搜索。这意味着，新的类别不仅可以用作过滤器，还可以用来计算与搜索词的语义相似性，从而进一步提高搜索结果的相关性。</p><p>完整的管道：</p>POST /_ingest/pipeline/_simulate
{
  "pipeline": {
    "processors": [
      {
        "script": {
          "source": """
            ctx.prompt = "You are an expert in semantic analysis and audiovisual content categorization. Your task is to generate only subcategories (max 2 topics) that describe specific aspects of movies based on their genres and descriptions. The output should be like string: 'n1, n2m ...n'. Here is a movies info to analyze: Title: " + ctx.title  + "Genres: " + ctx.genres  + "Description: " + ctx.description;
          """
        }
      },
      {
        "inference": {
          "model_id": "generate_topics_ia",
          "input_output": {
            "input_field": "prompt",
            "output_field": "result"
          }
        }
      },
      {
        "split": {
          "field": "result",
          "target_field": "topics",
          "separator": ", "
        }
      },
      {
        "remove": {
          "field": "result"
        }
      },
      {
        "remove": {
          "field": "prompt"
        }
      }
    ]
  },
  "docs": [
    {
      "_index": "index",
      "_id": "1",
      "_source": {
        "title": "The Lord of the Rings: The Fellowship of the Ring",
        "description": "While Frodo and Sam edge closer to Mordor with the help of the shifty Gollum, the divided fellowship makes a stand against Sauron's new ally, Saruman, and his hordes of Isengard.",
        "genres": [
          "Action",
          "Adventure",
          "Drama"
        ]
      }
    },
    {
      "_index": "index",
      "_id": "2",
      "_source": {
        "title": "Interstellar",
        "description": "A team of explorers travel through a wormhole in space in an attempt to ensure humanity's survival.",
        "genres": [
          "Adventure", "Drama", "Sci-Fi"
        ]
      }
    },
    {
      "_index": "index",
      "_id": "3",
      "_source": {
        "title": "The Martian",
        "description": "An astronaut becomes stranded on Mars after his team assume him dead, and must rely on his ingenuity to find a way to signal to Earth that he is alive.",
        "genres": [
          "Adventure", "Drama", "Sci-Fi"
        ]
      }
    }
  ]
}<h2>结论</h2><p>利用人工智能改进面，可以使搜索结果更具体、更符合上下文，从而改变搜索体验。固定类别通常比较宽泛，而人工智能生成的类别则不同，它能更好地反映背景情况。例如，在对电影进行重新分类时，我们可以捕捉到主要类别所忽略的背景，从而提供更相关的分组。</p><p>将这些新的类别添加到索引中，不仅可以改进分面，还可以实现矢量搜索。其结果是搜索体验更加高效，筛选器更加符合上下文。</p><h2>参考资料</h2><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-openai.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-openai.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/simulate-pipeline-api.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/simulate-pipeline-api.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/script-processor.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/script-processor.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/faceted-search-examples-ai</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/faceted-search-examples-ai</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd2838185214162c8/6a17f8bedbb4ff04affb58a5/25c9f9baa2326b5189ce0b1cc6240475781c755d-721x421.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 28 Jan 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何通过 Apache Airflow 将数据采集到 Elasticsearch]]></title>
    <description><![CDATA[了解如何通过 Apache Airflow 将数据摄取到 Elasticsearch。]]></description>
    <content:encoded><![CDATA[<h2>什么是阿帕奇气流？</h2><p>Apache Airflow 是一个用于创建、调度和监控工作流的平台。它用于协调 ETL 流程、数据管道和其他复杂的工作流程，具有灵活性和可扩展性。它的可视化界面和实时监控功能使管道管理更方便、更高效，让您可以跟踪执行的进度和结果。以下是其四大支柱：</p><ul><li><p><strong>动态： </strong>管道是用 Python 定义的，可以动态、灵活地生成工作流程。</p></li><li><p><strong>可扩展性：</strong>Airflow 可与各种环境集成，可创建自定义操作符，并可根据需要执行特定代码。</p></li><li><p><strong>优雅：</strong>管道的编写方式简洁明了。</p></li><li><p><strong>可扩展：</strong>它的模块化架构使用消息队列来协调任意数量的工作者。</p></li></ul><p>在实际应用中，气流可用于以下情况：</p><ul><li><p><strong>数据导入： </strong>协调将数据导入 Elasticsearch 等数据库的日常工作。</p></li><li><p><strong>日志监控：</strong>管理日志文件的收集和处理，然后在 Elasticsearch 中进行分析，以识别错误或异常。</p></li><li><p><strong>整合多个数据源：</strong>将来自不同系统（应用程序接口、数据库、文件）的信息整合到 Elasticsearch 的单层中，简化搜索和报告。</p></li></ul><h2>了解 Airflow 中的 DAG（有向无环图</h2><p>在 Airflow 中，工作流由 DAG（有向无环图）表示。DAG 是一种定义任务执行顺序的结构。DAG 的主要特点是</p><ul><li><p><strong>由独立任务组成：</strong>每个任务代表一个工作单元，可独立执行。</p></li><li><p><strong>排序： </strong>任务的执行顺序在 DAG 中明确定义。</p></li><li><p><strong>可重用性：</strong>DAG 设计为可重复执行，有利于流程自动化。</p></li></ul><h2>气流组件</h2><p>Airflow 生态系统由多个组件组成，这些组件共同协调任务：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt25a23489b8725b3e/6a17dfcfe8fbceba103a1846/bacd83aff625026d62f023e0434baa5782a2761a-1046x628.png" alt="气流主要部件" /><ul><li><p><strong>调度器：</strong>负责调度 DAG 和发送任务供工作者执行。</p></li><li><p><strong>执行者：</strong>管理任务的执行，将任务分配给员工。</p></li><li><p><strong>网络服务器：</strong>提供与 DAG 和任务交互的图形界面。</p></li><li><p><strong>Dags 文件夹：</strong>存放用 Python 编写的 DAG 的文件夹。</p></li><li><p><strong>元数据：</strong>作为工具存储库的数据库，由调度器和执行器用来存储执行状态。</p></li></ul><h2>Apache Airflow 和 Elasticsearch</h2><p>我们将演示如何使用 Apache Airflow 和 Elasticsearch 来协调任务并在 Elasticsearch 中对结果进行索引。本演示的目的是创建一个任务流水线，以更新 Elasticsearch 索引中的记录。该索引包含一个电影数据库，用户可在其中对电影进行评分和分配等级。假设每天有数百个收视率，那么就有必要不断更新收视率记录。为此，将开发一个每日执行的 DAG，负责检索新的综合评级并更新索引中的记录。</p><p>在 DAG 流程中，我们将有一个获取评级的任务，然后是一个验证结果的任务。如果数据不存在，DAG 将被定向到故障任务。否则，数据将在 Elasticsearch 中编入索引。我们的目标是通过一种方法，利用负责计算分数的机制来检索评分，从而更新索引中电影的评分字段。</p><h2>将 Apache Airflow 和 Elasticsearch 与 Docker 结合使用</h2><p>为了创建容器化环境，我们将使用带有 Docker 的 Apache Airflow。请按照<a href="https://airflow.apache.org/docs/apache-airflow/stable/howto/docker-compose/index.html"> "Running Airflow in Docker"</a> 指南中的说明实际设置 Airflow。</p><p>至于 Elasticsearch，我将使用 Elastic Cloud 上的集群，但如果你愿意，也可以使用 Docker 配置 Elasticsearch。已经创建了一个包含电影目录的索引，索引中包含电影数据。这些电影的 "评分 "字段将被更新。</p><h2>创建 DAG</h2><p>通过 Docker 安装后，将创建一个文件夹结构，其中包括 dags 文件夹，我们必须将 DAG 文件放在这里，Airflow 才能识别它们。</p><p>在此之前，我们需要确保安装了必要的依赖项。以下是此项目的依赖项</p>pip install apache-airflow apache-airflow-providers-elasticsearch<p>我们将创建文件<code>update_ratings_movies.py</code> 并开始任务编码。</p><p>现在，让我们导入必要的库：</p>from airflow import DAG
from airflow.operators.python import PythonOperator, BranchPythonOperator
from airflow.providers.elasticsearch.hooks.elasticsearch import ElasticsearchPythonHook<p>我们将使用<a href="https://airflow.apache.org/docs/apache-airflow-providers-elasticsearch/stable/hooks/elasticsearch_python_hook.html"><strong>ElasticsearchPythonHook</strong></a>，这是一个通过抽象连接和使用外部 API 来简化 Airflow 与 Elasticsearch 集群之间集成的组件。</p><p>接下来，我们定义 DAG，说明其主要参数：</p><ul><li><p><strong><code>dag_id</code></strong>：DAG 的名称。</p></li><li><p><strong><code>start_date</code></strong>：DAG 启动时间。</p></li><li><p><strong><code>schedule</code></strong>：定义周期（本例中为每日）。</p></li><li><p><strong><code>doc_md</code></strong>文件：将被导入并显示在 Airflow 界面中的文件。</p></li></ul><h2>确定任务</h2><p>现在，让我们来定义 DAG 的任务。第一个任务将负责检索电影分级数据。我们将使用<strong>PythonOperator</strong>，并将<code>task_id</code> 设置为<code>'get_movie_ratings'</code> 。<code>python_callable</code> 参数将调用负责获取评级的函数。</p>get_ratings_operator = PythonOperator(
   task_id='get_movie_ratings',
   python_callable=get_movie_ratings_task
)<p>接下来，我们需要验证结果是否有效。为此，我们将使用一个带有<strong>BranchPythonOperator</strong> 的条件。<code>task_id</code> 将是<code>'validate_result'</code> ，而<code>python_callable</code> 将调用验证函数。<code>op_args</code> 参数将用于把上一个任务<code>'get_movie_ratings'</code> 的结果传递给验证函数。</p>validate_result = BranchPythonOperator(
   task_id='validate_result',
   python_callable=validate_result,
   op_args=["{{ task_instance.xcom_pull(task_ids='get_movie_ratings') }}"]
)<p>如果验证成功，我们将从<code>'get_movie_ratings'</code> 任务中获取数据，并将其索引到 Elasticsearch 中。为此，我们将创建一个新任务<code>'index_movie_ratings'</code> ，它将使用<strong>PythonOperator</strong>。<code>op_args</code> 参数将把<code>'get_movie_ratings'</code> 任务的结果传递给索引函数。</p>index_ratings_operator = PythonOperator(
   task_id='index_movie_ratings',
   python_callable=index_movie_ratings_task,
   op_args=["{{ task_instance.xcom_pull(task_ids='get_movie_ratings') }}"]
)<p>如果验证显示失败，DAG 将继续执行失败通知任务。在本例中，我们只打印了一条信息，但在实际应用中，我们可以配置警报来通知故障。</p>failed_get_rating_operator = PythonOperator(
   task_id='failed_get_rating_operator',
   python_callable=lambda: print('Ratings were False, skipping indexing.')
)<p>最后，我们定义任务依赖关系，确保它们以正确的顺序执行：</p>get_ratings_operator &gt;&gt; validate_result &gt;&gt; [index_ratings_operator, failed_get_rating_operator]<p>下面是我们 DAG 的完整代码：</p>"""
DAG update Rating Movies
"""
import ast
import random

from airflow import DAG
from datetime import datetime

from airflow.operators.python import PythonOperator, BranchPythonOperator
from airflow.providers.elasticsearch.hooks.elasticsearch import ElasticsearchPythonHook


def index_movie_ratings_task(movies):
   es_hook = ElasticsearchPythonHook(hosts=None,
                                     es_conn_args={
                                         "cloud_id": "cloud_id"
                                         "api_key": "api-key"
                                     })
   es_client = es_hook.get_conn
   actions = []
   for movie in ast.literal_eval(movies):
       actions.append(
           {
               "update": {
                   "_id": movie["id"],
                   "_index": "movies"
               }
           }
       )
       actions.append(
           {
               "doc": {
                   "rating": movie["rating"]
               },
               "doc_as_upsert": True
           }
       )
   result = es_client.bulk(operations=actions)
   print(f"Ingestion completed.")
   print(result)
   return True


def get_movie_ratings_task():
   movies = [
       {"id": i, "rating": round(random.uniform(1, 10), 1)}
       for i in range(1, 100)
   ]
   return movies

def validate_result(result):
   if not result:
       return 'failed_get_rating_operator'
   else:
       return 'index_movie_ratings'


with DAG(
       dag_id="update_ratings_movies_2024",
       start_date=datetime(2024, 12, 29),
       schedule="@daily",
       doc_md=__doc__,
):
   get_ratings_operator = PythonOperator(
       task_id='get_movie_ratings',
       python_callable=get_movie_ratings_task
   )

   validate_result = BranchPythonOperator(
       task_id='validate_result',
       python_callable=validate_result,
       op_args=["{{ task_instance.xcom_pull(task_ids='get_movie_ratings') }}"],
       provide_context=True
   )

   index_ratings_operator = PythonOperator(
       task_id='index_movie_ratings',
       python_callable=index_movie_ratings_task,
       op_args=["{{ task_instance.xcom_pull(task_ids='get_movie_ratings') }}"]
   )

   failed_get_rating_operator = PythonOperator(
       task_id='failed_get_rating_operator',
       python_callable=lambda: print('Ratings were False, skipping indexing.')
   )

get_ratings_operator &gt;&gt; validate_result &gt;&gt; [index_ratings_operator, failed_get_rating_operator]<h2>可视化 DAG 执行</h2><p>在 Apache Airflow 界面中，我们可以直观地看到 DAG 的执行情况。只需转到"DAGs" 选项卡，找到您创建的 DAG。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9c5de210a5a264ad/6a17dfd00b0bed0290dd34cf/905b9191c4e3191e8b5608174d4c555370bf25eb-1600x760.png" alt="在 Apache Airflow 与 Elasticsearch 的接口中可视化 DAG 的执行情况" /><p>下面，我们可以直观地看到任务的执行情况及其各自的状态。通过选择特定日期的执行情况，我们可以访问每个任务的日志。请注意，在<strong><code>index_movie_ratings</code></strong> 任务中，我们可以看到索引中的索引结果，而且索引已成功完成。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0beddf311a808d24/6a17dfd2033c8d76de6bb0ba/73c3f738d27500cf153377bedf1aa2b67a94a8c8-1600x648.png" alt="通过 Elasticsearch 在 Apache Airflow 中可视化任务的执行情况及其状态" /><p>在其他选项卡中，可以获取有关任务和 DAG 的更多信息，帮助分析和解决潜在问题。</p><h2>结论</h2><p>在本文中，我们演示了如何将 Apache Airflow 与 Elasticsearch 集成以创建数据摄取解决方案。我们展示了如何配置 DAG，定义负责检索、验证和索引电影数据的任务，以及如何在 Airflow 界面监控和可视化这些任务的执行。</p><p>这种方法可以很容易地适应不同类型的数据和工作流，使 Airflow 成为在各种情况下协调数据管道的有用工具。</p><h2>参考资料</h2><p>阿帕奇气流</p><p><a href="https://airflow.apache.org/">https://airflow.apache.org/</a></p><p>使用 Docker 安装 Apache Airflow</p><p><a href="https://airflow.apache.org/docs/apache-airflow/stable/howto/docker-compose/index.html">https://airflow.apache.org/docs/apache-airflow/stable/howto/docker-compose/index.html</a></p><p>Elasticsearch Python 挂钩</p><p><a href="https://airflow.apache.org/docs/apache-airflow-providers-elasticsearch/stable/hooks/elasticsearch_python_hook.html">https://airflow.apache.org/docs/apache-airflow-providers-elasticsearch/stable/hooks/elasticsearch_python_hook.html</a></p><p>Python 操作员</p><p><a href="https://airflow.apache.org/docs/apache-airflow/stable/howto/operator/python.html">https://airflow.apache.org/docs/apache-airflow/stable/howto/operator/python.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/apache-airflow-elasticsearch-ingest-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/apache-airflow-elasticsearch-ingest-data</guid>
    <category><![CDATA[索引数据]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt28475ace97d1d989/6a1704c9286714d6dd93e1ec/5d4b47ac5d2ba453fc19dcc15efa2aed5f55d88b-1440x1355.png" length="0" type="image/png"/>
    <pubDate>Fri, 17 Jan 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何通过 Kafka 向 Elasticsearch 摄取数据]]></title>
    <description><![CDATA[逐步指导如何使用 Python、Docker Compose 和 Kafka Connect 将 Apache Kafka 与 Elasticsearch 集成，以实现高效的数据摄取、索引和可视化。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将展示如何将 Apache Kafka 与 Elasticsearch 集成，以实现数据摄取和索引。我们将概述 Kafka 及其生产者和消费者的概念，并创建一个日志索引，通过 Apache Kafka 接收消息并编制索引。该项目使用 Python 实现，代码可在<a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/elasticsearch-through-apache-kafka">GitHub</a> 上获取。</p><h3><strong>准备工作</strong></h3><ul><li><p>Docker 和 Docker Compose：确保计算机上安装了 Docker 和 Docker Compose。</p></li><li><p>Python 3.x：运行生产者和消费者脚本。</p></li></ul><h3><strong>Apache Kafka 简介</strong></h3><p>Apache Kafka 是一个分布式流平台，可实现高扩展性和可用性以及容错。在 Kafka 中，数据管理通过主要组件进行：</p><ul><li><p><strong>经纪人</strong>：负责在生产者和消费者之间存储和分发信息。</p></li><li><p><strong>Zookeeper</strong>：管理和协调 Kafka 代理，控制集群状态、分区领导者和消费者信息。</p></li><li><p><strong>主题</strong>：发布和存储数据以供消费的渠道。</p></li><li><p><strong>消费者和生产者</strong>：生产者向主题发送数据，消费者检索数据。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4aae32304d7417f6/6a17f7be577262b47c1bcdac/89a37243baec48bbdfa85e3298fc91082322ed4e-1600x868.png" alt="图解 Apache Kafka" /><p>这些组件共同构成了 Kafka 生态系统，为数据流提供了一个强大的框架。</p><h3><strong>项目结构</strong></h3><p>为了解数据摄取过程，我们将其分为几个阶段：</p><ul><li><p><strong>基础架构调配</strong>：设置 Docker 环境以支持 Kafka、Elasticsearch 和 Kibana。</p></li><li><p><strong>创建生产者</strong>：实现向日志主题发送数据的 Kafka 生产者。</p></li><li><p><strong>创建消费者</strong>：开发 Kafka 消费者，以便在 Elasticsearch 中读取信息并编制索引。</p></li><li><p><strong>输入验证</strong>：验证和确认发送和消耗的数据。</p></li></ul><h3><strong>使用 Docker Compose 配置基础设施</strong></h3><p>我们利用 Docker Compose 配置和管理必要的服务。下面是 Docker Compose 代码，用于设置集成 Apache Kafka、Elasticsearch 和 Kibana 所需的各项服务，确保数据摄取过程。</p>version: "3"

services:

  zookeeper:
    image: confluentinc/cp-zookeeper:latest
    container_name: zookeeper
    environment:
      ZOOKEEPER_CLIENT_PORT: 2181

  kafka:
    image: confluentinc/cp-kafka:latest
    container_name: kafka
    depends_on:
      - zookeeper
    ports:
      - "9092:9092"
      - "9094:9094"
    environment:
      KAFKA_BROKER_ID: 1
      KAFKA_ZOOKEEPER_CONNECT: zookeeper:2181
      KAFKA_ADVERTISED_LISTENERS: PLAINTEXT://kafka:29092,PLAINTEXT_HOST:${HOST_IP}:9092
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: PLAINTEXT:PLAINTEXT,PLAINTEXT_HOST:PLAINTEXT
      KAFKA_INTER_BROKER_LISTENER_NAME: PLAINTEXT
      KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1

  elasticsearch:
    image: docker.elastic.co/elasticsearch/elasticsearch:8.15.1
    container_name: elasticsearch-8.15.1
    environment:
      - node.name=elasticsearch
      - xpack.security.enabled=false
      - discovery.type=single-node
      - "ES_JAVA_OPTS=-Xms512m -Xmx512m"
    volumes:
      - ./elasticsearch:/usr/share/elasticsearch/data
    ports:
      - 9200:9200

  kibana:
    image: docker.elastic.co/kibana/kibana:8.15.1
    container_name: kibana-8.15.1
    ports:
      - 5601:5601
    environment:
      ELASTICSEARCH_URL: http://elasticsearch:9200
      ELASTICSEARCH_HOSTS: '["http://elasticsearch:9200"]'<p>您可以直接从 Elasticsearch Labs<a href="https://github.com/andreluiz1987/elasticsearch-labs/tree/supporting-blog/elasticsearch-apache-kafka/supporting-blog-content/elasticsearch-through-apache-kafka">GitHub</a>repo 访问该文件。</p><h3><strong>使用 Kafka 生产者发送数据</strong></h3><p>生产者负责向日志主题发送消息。通过分批发送信息，它提高了网络使用效率，允许使用<code>batch_size</code> 和<code>linger_ms</code> 设置进行优化，这两个设置分别控制批次的数量和延迟。配置<code>acks='all'</code> 可确保信息的持久存储，这对重要的日志数据至关重要。</p>producer = KafkaProducer(
   bootstrap_servers=['localhost:9092'],  # Specifies the Kafka server to connect
   value_serializer=lambda x: json.dumps(x).encode('utf-8'),  # Serializes data as JSON and encodes it to UTF-8 before sending
   batch_size=16384,     # Sets the maximum batch size in bytes (here, 16 KB) for buffered messages before sending
   linger_ms=10,         # Sets the maximum delay (in milliseconds) before sending the batch
   acks='all'            # Specifies acknowledgment level; 'all' ensures message durability by waiting for all replicas to acknowledge
)


def generate_log_message():
   levels = ["INFO", "WARNING", "ERROR", "DEBUG"]
   messages = [
       "User login successful",
       "User login failed",
       "Database connection established",
       "Database connection failed",
       "Service started",
       "Service stopped",
       "Payment processed",
       "Payment failed"
   ]
   log_entry = {
       "level": random.choice(levels),
       "message": random.choice(messages),
       "timestamp": time.time()
   }
   return log_entry

def send_log_batches(topic, num_batches=5, batch_size=10):
   for i in range(num_batches):
       logger.info(f"Sending batch {i + 1}/{num_batches}")
       for  in range(batch_size):
           log_message = generate_log_message()
           producer.send(topic, value=log_message)
       producer.flush()


if __name__ == "__main__":
   topic = "logs"
   send_log_batches(topic)
   producer.close()<p>启动生产者时，消息会分批发送到主题，如下图所示：</p>INFO:kafka.conn:Set configuration …
INFO:log_producer:Sending batch 1/5 
INFO:log_producer:Sending batch 2/5
INFO:log_producer:Sending batch 3/5
INFO:log_producer:Sending batch 4/5<h3><strong>使用 Kafka 消费者消费数据并编制索引</strong></h3><p>消费者旨在高效处理消息，从日志主题中批量消费，并将其索引到 Elasticsearch 中。通过<code>auto_offset_reset='latest'</code> ，它可以确保消费者开始处理最新的邮件，而忽略较早的邮件，并且<code>max_poll_records=10</code> 将批量限制为 10 封邮件。使用<code>fetch_max_wait_ms=2000</code> 时，消费者最多等待 2 秒钟，积累足够的报文后再处理批处理。</p><p>在其主循环中，消费者消耗日志信息，处理每个批次并将其索引到 Elasticsearch 中，从而确保持续的数据摄取。</p>consumer = KafkaConsumer(
   'logs',                               
   bootstrap_servers=['localhost:9092'],
   auto_offset_reset='latest',            # Ensures reading from the latest offset if the group has no offset stored
   enable_auto_commit=True,               # Automatically commits the offset after processing
   group_id='log_consumer_group',         # Specifies the consumer group to manage offset tracking
   max_poll_records=10,                   # Maximum number of messages per batch
   fetch_max_wait_ms=2000                 # Maximum wait time to form a batch (in ms)
)

def create_bulk_actions(logs):
   for log in logs:
       yield {
           "_index": "logs",
           "_source": {
               'level': log['level'],
               'message': log['message'],
               'timestamp': log['timestamp']
           }
       }

if __name__ == "__main__":
   try:
       print("Starting message processing…")
       while True:

           messages = consumer.poll(timeout_ms=1000)  # Poll receive messages

           # process each batch messages
           for _, records in messages.items():
               logs = [json.loads(record.value) for record in records]
               bulk_actions = create_bulk_actions(logs)
               response = helpers.bulk(es, bulk_actions)
               print(f"Indexed {response[0]} logs.")
   except Exception as e:
       print(f"Erro: {e}")
   finally:
       consumer.close()
       print(f"Finish")<h3><strong>在 Kibana 中可视化数据</strong></h3><p>有了 Kibana，我们就能探索和验证从 Kafka 采集并在 Elasticsearch 中编入索引的数据。通过访问 Kibana 中的 "<strong>开发工具</strong>"，您可以查看已编入索引的信息，并确认数据符合预期。例如，如果我们的 Kafka 生产者发送了 5 个批次，每个批次 10 条消息，那么我们应该在索引中看到总共 50 条记录。</p><p>要验证数据，可以使用<strong>开发工具</strong>部分的以下查询：</p>GET /logs/_search
{
  "query": {
    "match_all": {}
  }
}<p>响应：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc15fb278fe6f984e/6a17f7c0b1e1131ac279f404/f44f95fc27bba50991412d5c7e7728519b9bdec4-688x1024.png" alt="响应验证数据 - Kafka&amp; Elasticsearch" /><p>此外，Kibana 还提供创建可视化和仪表盘的功能，有助于使分析更加直观和互动。下面，您可以看到我们创建的仪表盘和可视化的一些示例，这些仪表盘和可视化以各种格式展示了数据，增强了我们对所处理信息的理解。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltabc2856c9feedc47/6a17f7c13e9e4522bfba1651/a18e0ebb543e929136d4786651bc6cee32fa69bc-1600x470.png" alt="Kibana 可视化 - Kafka&amp; Elasticsearch" /><h3><strong>使用 Kafka Connect 进行数据摄取</strong></h3><p>Kafka Connect 是一项服务，旨在促进数据源和目的地（汇）（如数据库或文件系统）之间的集成。它通过预定义的连接器自动处理数据移动。在我们的案例中，Elasticsearch 发挥着数据汇的作用。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt53e3acaf62e7dbed/6a17f7c3577262c5151bcdb0/52a6982c864fdc04cb7a8a5fb02e67dca0ba8226-1600x819.png" alt="使用 Kafka Connect 进行数据输入" /><p>使用 Kafka Connect，我们可以简化数据摄取流程，无需在 Elasticsearch 中手动实施数据摄取工作流。有了适当的连接器，Kafka Connect 就能将发送到 Kafka 主题的数据直接编入 Elasticsearch 索引，只需极少的设置，也无需额外编码。</p><h4><strong>使用 Kafka Connect</strong></h4><p>为了实现 Kafka Connect，我们将在 Docker Compose 设置中添加<a href="https://github.com/andreluiz1987/es-apache-kafka/blob/main/docker-compose.yml#L31"> kafka-connect 服务 </a>。此配置的关键部分是安装 Elasticsearch 连接器，它将处理数据索引。</p><p>配置服务并创建 Kafka Connect 容器后，需要为 Elasticsearch 连接器创建配置文件。该文件定义了基本参数，如</p><ul><li><p><code>connection.url</code>:Elasticsearch 的连接 URL。</p></li><li><p><code>topics</code>:连接器将监控的 Kafka 主题（本例中为"logs" ）。</p></li><li><p><code>type.name</code>:Elasticsearch 中的文档类型（通常为 _doc）。</p></li><li><p><code>value.converter</code>:将 Kafka 消息转换为 JSON 格式。</p></li><li><p><code>value.converter.schemas.enable</code>:指定是否包含模式。</p></li><li><p><code>schema.ignore</code> 和<code>key.ignore</code> ：在编制索引时忽略 Kafka 模式和键的设置。</p></li></ul><p>下面是在 Kafka Connect 中创建 Elasticsearch 连接器的<code>curl</code> 命令：</p>curl --location '{{url}}/connectors' \
--header 'Content-Type: application/json' \
--data '{
    "name": "elasticsearch-sink-connector",
    "config": {
        "connector.class": "io.confluent.connect.elasticsearch.ElasticsearchSinkConnector",
        "topics": "logs",
        "connection.url": "http://elasticsearch:9200",
        "type.name": "_doc",
        "value.converter": "org.apache.kafka.connect.json.JsonConverter",
        "value.converter.schemas.enable": "false",
        "schema.ignore": "true",
        "key.ignore": "true"
    }
}'<p>使用此配置后，Kafka Connect 将自动开始摄取发送到"logs" 主题的数据，并在 Elasticsearch 中编制索引。这种方法可实现全自动数据摄取和索引，无需额外编码，从而简化了整个集成过程。</p><h3><strong>结论</strong></h3><p>集成 Kafka 和 Elasticsearch 可为实时数据摄取和分析创建一个强大的管道。本指南为构建强大的数据摄取架构提供了基本方法，可在 Kibana 中实现无缝可视化和分析，随时适应未来更复杂的要求。</p><p>此外，使用 Kafka Connect 使 Kafka 和 Elasticsearch 之间的集成更加简化，无需额外的代码来处理数据和编制索引。Kafka Connect 使发送到特定主题的数据只需最少的配置就能在 Elasticsearch 中自动编入索引。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-apache-kafka-ingest-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-apache-kafka-ingest-data</guid>
    <category><![CDATA[索引数据]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt53e3acaf62e7dbed/6a17f7c3577262c5151bcdb0/52a6982c864fdc04cb7a8a5fb02e67dca0ba8226-1600x819.png" length="0" type="image/png"/>
    <pubDate>Tue, 24 Dec 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何在电子商务产品目录中使用混合搜索]]></title>
    <description><![CDATA[了解如何使用混合搜索建立电子商务产品目录，使用分面、促销、个性化和行为分析。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将演示如何实现混合搜索，将全文搜索和矢量搜索的结果结合起来。混合搜索将这两种方法统一起来，充分利用了两种搜索策略的优点，从而提高了搜索结果的广度。</p><p>除了集成混合搜索，我们还将演示如何添加功能，使您的搜索解决方案更加强大。其中包括切面和个性化产品促销。此外，我们还将向您展示如何使用 Elastic 的行为分析工具捕捉用户互动并生成有价值的见解。</p><p>在本实现中，您将看到如何构建允许用户查看搜索结果并与之交互的界面，以及负责返回信息的应用程序接口。要访问包含源代码的资源库，请点击下面的链接：</p><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/product-store-search">https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/product-store-search</a></p></li><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/app-product-store">https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/app-product-store</a> </p></li></ul><p>我们将本指南分为几个步骤，从创建索引到实施分面和结果个性化等高级功能。最后，您将拥有一个强大的搜索解决方案，可以在电子商务场景中使用。</p><h2>电子商务混合搜索的环境设置</h2><p>在开始实施之前，我们需要设置环境。您可以选择使用 Elastic Cloud 上的服务或容器化解决方案来管理 Elasticsearch。如果选择容器化，可在此版本库中找到通过 Docker Compose 进行的配置：<a href="https://github.com/andreluiz1987/product-store-search/blob/main/docker/docker-compose.yml">docker-compose.yml</a>。</p><h2>创建索引和录入产品目录</h2><p>索引将根据化妆品目录创建，其中包括名称、描述、照片、类别和标签等字段。用于全文搜索的字段，如"名称" 和"描述，" 将被映射为<code>text</code> ，而用于聚合的字段，如"类别" 和"品牌，" 将被映射为<code>keyword</code> ，以便进行分面搜索。</p><p>"description" 字段将用于矢量搜索，因为它提供了有关产品的更多背景信息。这个字段将被定义为<code>dense_vector,</code> ，存储描述的矢量表示。</p><p>索引映射如下</p>{
   "mappings":{
      "properties":{
         "id":{
            "type":"keyword"
         },
         "brand":{
            "type":"text",
            "fields":{
               "keyword":{
                  "type":"keyword"
               }
            }
         },
         "name":{
            "type":"text"
         },
         "price":{
            "type":"float"
         },
         "price_sign":{
            "type":"keyword"
         },
         "currency":{
            "type":"keyword"
         },
         "image_link":{
            "type":"keyword"
         },
         "description":{
            "type":"text"
         },
         "description_embeddings":{
            "type":"dense_vector",
            "dims":384
         },
         "rating":{
            "type":"keyword"
         },
         "category":{
            "type":"keyword"
         },
         "product_type":{
            "type":"keyword"
         },
         "tag_list":{
            "type":"keyword"
         }
      }
   }
}<p>创建索引的脚本可在<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/product-store-search/infra/create_index.py">此处</a>找到。</p><h2>嵌入式生成</h2><p>为了将产品描述矢量化，我们使用了 all-MiniLM-L6-v2 模型。在这种情况下，应用程序负责在编制索引前生成嵌入。另一种方法是将模型导入 Elasticsearch 集群，但在本地环境中，我们选择直接在应用程序中执行矢量化。</p><p>我们使用<a href="https://www.kaggle.com/datasets/shivd24coder/cosmetic-brand-products-dataset">Kaggle</a>上的化妆品数据集来填充索引，为了提高数据摄取的效率，我们使用了批处理方法。在同一摄取阶段，我们将生成"description" 字段的嵌入，并将其索引到新字段"description_embeddings" 中。</p><p>整个数据摄取过程可通过存储库中的<strong>Jupyter Notebook</strong>直接跟踪和执行。该笔记本提供了关于如何读取、处理数据并将其编入 Elasticsearch 索引的分步指南，便于复制和实验。</p><p>您可以通过以下链接访问该笔记本：<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/product-store-search/ingestion/ingestion.ipynb">摄入笔记本。</a></p><h2>混合搜索实施</h2><p>现在，让我们来实现混合搜索。对于基于关键字的搜索，我们使用<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-multi-match-query.html">multi_match</a>查询，目标字段为"name、" " category、" 和"description。"这样就能确保检索到这些字段中包含搜索词的文档。</p><p>对于向量搜索，我们使用<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html">KNN 查询</a>。在执行查询之前，需要对搜索词进行矢量化，这需要使用对输入词进行矢量化的方法来完成。请注意，摄取时使用的同一模型也用于搜索词。</p><p>这两种搜索的组合是通过<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/rrf.html">互易等级融合（RRF） </a>算法完成的，该算法可合并两种查询的结果，并通过减少噪音来提高搜索精度。RRF 允许基于关键字的搜索和矢量搜索共同发挥作用，从而增强对用户查询的理解。</p>query = {
   "retriever": {
       "rrf": {
           "retrievers": [
               {
                   "standard": {
                       "query": organic_query['query']
                   }
               },
               {
                   "knn": {
                       "field": "description_embeddings",
                       "query_vector": vector,
                       "k": 5,
                       "num_candidates": 20
                   }
               }
           ],
           "rank_window_size": 20,
           "rank_constant": 5
       }
   },
   "_source": organic_query['_source']
}<h3>结果比较：关键词搜索与混合搜索</h3><p>现在，让我们比较一下传统关键词搜索和混合搜索的结果。当使用关键字搜索"干性皮肤粉底" 时，我们会得到以下结果：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcb5405df787bc8f5/6a17023e961e697ca1c4cdb4/52e717aa3c9c1fadfadb639f5fb77cf8e47e3b34-1600x1021.png" alt="比较结果：关键词搜索与混合搜索" /><ol><li><p><strong>Revlon ColorStay Makeup for Normal / Dry SkinDescription</strong>：露华浓持久彩妆（Revlon ColorStay Makeup）质地轻盈，遮瑕持久，不会结块、褪色或脱妆。这款无油、保湿平衡配方采用了 "定时释放技术"，特别适合中性或干性肌肤使用，能够持续为肌肤补充水分：妆感舒适，可持续使用长达 24 小时；中等至完全遮盖；有多种美丽色调可供选择。
</p></li><li><p><strong>美宝莲梦幻柔滑慕斯粉底液描述</strong>：你会爱上它的原因独特的乳霜状粉底提供 100% 婴儿般柔滑的完美肤质。不含油、不含香料，通过皮肤科医生测试，通过过敏测试，不致粉刺，不会堵塞毛孔。适合敏感性皮肤</p></li></ol><p><strong>分析</strong>：在搜索"干性皮肤粉底时，" 搜索结果是通过搜索关键词与产品标题和描述之间的精确匹配获得的。然而，这种匹配并不总是最佳选择。例如，<strong>Revlon ColorStay Makeup for Normal / Dry Skin</strong>就是一个不错的选择，因为它是专为干性皮肤配制的。尽管它不含油分，但其配方设计却能提供持续的保湿效果。相比之下，我们还收到了<strong>美宝莲梦幻柔滑慕斯粉底液</strong>，虽然这款粉底液不含油分，也能补充水分，但一般更推荐油性或混合性皮肤使用，因为无油产品往往侧重于控油，而不是提供干性皮肤所需的额外水分。这凸显了基于关键词搜索的局限性，因为这种搜索可能会返回无法完全满足干性皮肤患者特殊需求的产品。</p><p>现在，使用混合方法进行相同的搜索：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt412e38b327cf6a4f/6a17024066c4f94c52f8beb9/13b562a5120619355ba0861976102e208964a49f-1600x1021.png" alt="使用混合方法进行搜索" /><ol><li><p><strong>CoverGirl Outlast Stay Luminous Foundation Creamy Natural (820):产品简介</strong>：CoverGirl Outlast Stay Luminous 粉底液是打造露光妆效和微妙光泽的完美之选。它不含油分，配方不油腻，能为肌肤带来全天候的自然亮泽！这款全天候粉底液可为肌肤补充水分，同时提供无瑕遮瑕。<strong>分析： </strong>这款产品非常适合干性皮肤的用户，因为它强调补水。"Hydrates skin" 和"dewy finish" 这两个词符合用户寻找干性皮肤粉底的意图。矢量搜索很可能理解了补水的概念，并将其与解决皮肤干燥问题的粉底需求联系起来。
</p></li><li><p><strong>Revlon ColorStay Makeup for Normal / Dry Skin:说明：</strong>Revlon ColorStay Makeup 采用轻盈配方，具有持久遮瑕效果，不会结块、褪色或脱落。这款不含油分的保湿平衡配方采用了 "定时释放技术"，特别适合中性或干性肌肤使用，能持续为肌肤补充水分。<strong>分析： </strong>这款产品直接针对干性皮肤用户的需求，明确指出其配方适用于中性或干性皮肤。"水分平衡配方" 和持续的保湿效果非常适合寻找适合干性皮肤的粉底的人。矢量搜索成功检索到这一结果，不仅是因为关键词匹配，还因为重点关注补水，并特别提到干性皮肤是目标人群。
</p></li><li><p><strong>精华粉底液说明： </strong>精华粉底液质地轻盈，遮瑕度适中，共有 21 种色调可供选择。这些粉底液的遮瑕度适中，看起来很自然，精华液质地非常轻盈。它们的粘度很低，可使用随附的泵或单独购买的玻璃滴管进行分配。<strong>分析：</strong>在这款产品中，说明强调的是一款质地轻盈、妆感自然的精华粉底液，这与干性皮肤人群的需求不谋而合，因为干性皮肤人群通常需要的是温和、保湿、妆感不结块的产品。尽管"干性皮肤" 这个词没有被明确提及，但矢量搜索可能从更广泛的背景中捕捉到了轻盈、自然的遮盖力和类似精华液的质地，这与保湿度和涂抹舒适度有关，使其与干性皮肤相关。</p></li></ol><h2>面的实施</h2><p>面孔对于有效提炼和过滤搜索结果至关重要，可为用户提供更有针对性的导航，尤其是在电子商务等产品种类繁多的情况下。它们允许用户根据类别、品牌或价格等属性调整搜索结果，使搜索更加准确。为了实现这一功能，我们在<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html"> </a><code>category</code>和 字段上使用了 术语聚合<code>brand</code> ，这两个 字段<code>keyword</code> 在索引创建阶段被定义为 。</p>    query = build_query(term, categories, product_types, brands)
    query["aggs"] = {
        "product_types": {"terms": {"field": "product_type"}},
        "categories": {"terms": {"field": "category"}},
        "brands": {"terms": {"field": "brand.keyword"}}
    }<p>实施的完整代码可在<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/product-store-search/api/api.py#L144">此处</a>找到。</p><p>下面是搜索"适用于干性皮肤的粉底" 的面结果：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt31080652a60d8600/6a1702416234e0af7ddb18e7/e8907af359c021eaa38ec7a6a53f10c325ee8ade-1146x1248.png" alt="搜索&quot;适用于干性皮肤的粉底液的面部结果&quot;" /><h2>自定义结果：固定查询</h2><p>在某些情况下，在搜索结果中推广某些产品可能是有益的。为此，我们使用了 "<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-pinned-query.html"><strong>固定查询"</strong></a>，它允许特定产品出现在搜索结果的顶部。下面，我们将在不推销任何产品的情况下搜索"Foundation" ：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt31e75c815262993e/6a170243509168ae42e1b95d/9396c4ed358e7a68eed47f17dd6915d0bad492aa-1600x1157.png" alt="在不推销任何产品的情况下搜索&quot;Foundation&quot; " /><p>在我们的例子中，我们可以推广带有"无麸质标签的产品。"通过使用产品 ID，我们可以确保产品在搜索结果中的优先级。具体而言，我们将推广以下产品：<strong>精华粉底液</strong>（编号：1043）、<strong>遮瑕粉底液</strong>（编号：1042）和<strong>Realist 隐形定妆粉</strong>（编号：1039）。</p>{
   "query":{
      "pinned":{
         "ids":[
            "1043",
            "1042",
            "1039"
         ],
         "organic":{
            "bool":{
               "must":[
                  {
                     "multi_match":{
                        "query":"foundation",
                        "fields":[
                           "name",
                           "category",
                           "description"
                        ]
                     }
                  }
               ]
            }
         }
      }
   }
}<p>我们使用特定的产品 ID 来确保它们在查询结果中的优先级。查询结构包括一个产品 ID 列表，该列表应将"钉在" 的顶部（在本例中，ID 为 1043、1042 和 1039），而其余结果则按照搜索的有机流程，使用"name" 、"category" 和"description" 等字段中的文本查询条件组合。这样，就有可能以可控的方式推广项目，确保其可见性，同时保持搜索的其他部分基于通常的相关性。</p><p>下面是查询执行的结果和促销产品：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt53a6d5cbef227f16/6a1702458b73cb68f0189ef3/8c1a547bfe31360c405eae0891c666635051a51b-1600x1039.png" alt="带促销产品的查询执行结果" /><p>完整的查询代码可在<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/product-store-search/api/api.py#L112">此处</a>找到。</p><h2>利用行为分析技术分析搜索行为</h2><p>到目前为止，我们已经增加了一些功能，以提高搜索结果的相关性和产品的可发现性。现在，我们将通过加入一项功能来最终确定搜索解决方案，该功能将帮助我们分析用户的搜索行为，识别有结果或无结果的查询以及搜索结果的点击等模式。为此，我们将使用 Elastic 提供的<strong>行为分析</strong>功能。有了它，只需几个步骤，我们就能监控和分析用户的搜索行为，获得宝贵的见解，从而优化搜索体验。</p><h3>创建行为分析集合</h3><p>我们的第一项操作是创建一个集合，负责接收所有行为分析事件。要创建集合，请访问<strong>Search&gt; Behavioral Analytics</strong> 中的 Kibana 界面。在下面的示例中，我们创建了名为<code>tracking-search</code> 的集合。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9b4cb5480ab0bfef/6a1702474a531b189936a7e7/c05e533212a3690c5ea9f2d5226cbcdc901c482a-1600x1009.png" alt="行为分析--为您的收藏命名" /><h3>将行为分析整合到界面中</h3><p>我们的<a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/hybrid-search-for-an-e-commerce-product-catalogue/app-product-store">前端</a>应用程序是用 JavaScript 开发的，为了集成行为分析，我们将按照 Elastic 官方文档中描述的步骤安装<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/behavioral-analytics-start.html#behavioral-analytics-start-ui-integration-js-client"><strong>行为分析 JavaScript 跟踪器</strong></a>。</p><h3>实施 JavaScript 跟踪器</h3><p>现在，我们将把跟踪器客户端导入应用程序，并使用<code>trackPageView</code> 、<code>trackSearch</code> 和<code>trackSearchClick</code> 方法来捕捉用户交互。</p><p><strong>免责声明</strong>：虽然我们正在使用一种工具来收集用户交互数据，但这对于确保遵守<strong>GDPR</strong> 至关重要。这意味着要明确告知用户正在收集哪些数据、将如何使用这些数据，并提供选择退出跟踪的选项。此外，我们必须采取强有力的安全措施来保护收集到的信息，并尊重用户的权利，如数据访问和删除，确保所有步骤都符合 GDPR 原则。
</p><p><strong>步骤 1：创建跟踪器实例</strong></p><p>首先，我们将创建用于监控交互的跟踪器实例。在此配置中，我们定义了目标端点、集合名称和 API 密钥：</p>createTracker({
  endpoint: "https://endpoint:443",
  collectionName: "tracking-search",
  apiKey: "api-key"
});<p><strong>步骤 2：获取页面浏览量</strong></p><p>要跟踪页面浏览量，我们可以配置<code>trackPageView</code> 事件：</p>    trackPageView({
      page: {
        title: "home-page"
      },
    });<p>有关<code>trackPageView</code> 事件的详细信息，请参阅本<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/behavioral-analytics-event-reference.html#behavioral-analytics-event-reference-pageview-fields">文档</a>。</p><p><strong>步骤 3：捕捉搜索查询</strong></p><p>为了监控用户的搜索行为，我们将使用<code>trackSearch</code> 方法：</p>      trackSearch({
        search: {
          query: searchTerm,
          results: {
            items: documents,
            total_results: response.data.length,
          },
        },
      });<p>在这里，我们收集搜索词和搜索结果。</p><p><strong>步骤 4：跟踪搜索结果的点击率</strong></p><p>最后，为了捕捉搜索结果的点击，我们将使用<code>trackSearchClick</code> 方法：</p>trackSearchClick({
      document: { id: product.id, index: "products-catalog"},
      search: {
        query: searchTerm,
        page: {
          current: 1,
          size: products.length,
        },
        results: {
          items: documents,
          total_results: products.length,
        },
        search_application: "app-product-store"
      },
    });<p>我们收集被点击文档的 ID 信息以及搜索词和搜索结果。</p><h3>在 Kibana 中分析数据</h3><p>既然已经捕捉到了用户交互事件，我们就可以获得有关搜索操作的宝贵数据。Kibana 使用行为分析工具来可视化和分析这些行为数据。要查看结果，只需导航到<strong>搜索&gt; 行为分析&gt; 我的收藏</strong>，就会显示捕获事件的概览。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdab4f4507c5a4732/6a17024860084b34ca3c4411/e219c4af2e01b0ccc1b9079459a07832835abc75-1600x1155.png" alt="在 Kibana 中分析数据" /><p>在此概览中，我们可以大致了解集成到界面中的每个操作所捕获的事件。从这些信息中，我们可以获得有关用户搜索行为的宝贵见解。不过，如果您想创建个性化的仪表盘，其中包含与您的特定场景更相关的指标，Kibana 提供了用于构建仪表盘的强大工具，允许您创建各种指标可视化。</p><p>下面，我创建了一些可视化和图表来监测，例如，一段时间内搜索次数最多的词语、没有结果的查询、突出显示搜索次数最多的词语的词云，以及最后的地理可视化，以确定搜索访问来自哪里。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt04ced4406436f942/6a17024a5091685116e1b961/84a2c39dab77a3f9366c258a3a45d2cbc7df125e-1600x689.png" alt="可视化和图表监控" /><h2>结论</h2><p>在本文中，我们实施了一种混合搜索解决方案，将关键词和矢量搜索相结合，为用户提供更准确、更相关的搜索结果。我们还探索了如何使用其他功能，如面和钉住查询的个性化结果，以创建更完整、更高效的搜索体验。</p><p>此外，我们还集成了 Elastic 的<strong>行为分析</strong>功能，以捕捉和分析用户与搜索引擎交互过程中的行为。通过使用<code>trackPageView</code> 、<code>trackSearch</code> 和<code>trackSearchClick</code> 等方法，我们能够监控搜索查询、搜索结果点击量和页面浏览量，从而对搜索行为产生有价值的见解。</p><h2>参考资料</h2><p>数据集</p><p><a href="https://www.kaggle.com/datasets/shivd24coder/cosmetic-brand-products-dataset">https://www.kaggle.com/datasets/shivd24coder/cosmetic-brand-products-dataset</a></p><p>变压器</p><p><a href="https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2">https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2</a></p><p>互惠等级融合</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/rrf.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/rrf.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/retriever.html#rrf-retriever">https://www.elastic.co/guide/en/elasticsearch/reference/current/retriever.html#rrf-retriever</a></p><p>Knn 查询</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html</a></p><p>固定查询</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-pinned-query.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-pinned-query.html</a></p><p>行为分析应用程序接口</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/behavioral-analytics-apis.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/behavioral-analytics-apis.html</a></p><p>https://www.elastic.co/guide/en/elasticsearch/reference/current/behavioral-analytics-overview.html</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/hybrid-search-ecommerce</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/hybrid-search-ecommerce</guid>
    <category><![CDATA[向量数据库]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt63c711d1bf9b2501/6a17024c66c4f9c2b7f8bebd/05578fc595a12f6b1ebf88a10a2a31e9971b545e-1200x628.png" length="0" type="image/png"/>
    <pubDate>Tue, 12 Nov 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[语义搜索的实现：使用 Elasticsearch 构建食谱搜索]]></title>
    <description><![CDATA[在电子商务网站中实施语义搜索。]]></description>
    <content:encoded><![CDATA[<h2>引言</h2><p>许多电子商务网站都希望提升食谱搜索体验。语义搜索如果应用得当，可以让客户根据更自然的查询快速找到所需的配料，例如"情人节的配料" 或"感恩节大餐。"</p><p>本文将演示如何使用 Elasticsearch 实现支持此类查询的语义搜索。我们将配置一个索引来存储超市的配料和产品的目录，并演示如何使用该索引来改进食谱搜索。在本文中，我们将介绍如何创建这种数据结构，并应用自然语言处理技术提供符合客户意图的相关结果。</p><p>本文中介绍的所有代码都是用 Python 开发的，可在<a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/building-a-recipe-search-with-elasticsearch">GitHub</a> 上获取。您可以访问资源库，查看源代码，根据需要进行调整，并直接在您的开发环境中实施解决方案。</p><h2>开始实施语义搜索</h2><p>要开始实施语义搜索，我们首先需要定义自然语言模型。Elastic 提供了自己的模型<a href="https://www.elastic.co/guide/en/machine-learning/8.15/ml-nlp-elser.html"><strong>ELSER</strong></a>，但也支持整合来自不同供应商的 NLP 模型，如 Hugging Face。这种灵活性使您可以选择最适合您需求的方案。</p><p>在本文中，我们将使用<strong>ELSER</strong>，它可以降低部署和管理 NLP 模型的复杂性。此外，Elastic 还提供<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search-semantic-text.html"><strong>语义文本</strong></a>功能，大大简化了流程。有了<strong>semantic_text</strong>，整个嵌入生成过程就变得简单而自动化。您只需定义一个推理点，并在索引映射中指定接收嵌入的字段。在编制文档索引时，将生成嵌入并自动与指定字段关联。</p><h3>设置步骤</h3><p>以下是创建支持语义搜索的索引的步骤。按照这些说明，您就可以配置好索引，并为语义搜索做好准备：</p><ol><li><p><strong>创建 </strong><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-elser.html"><strong>推理点</strong></a>。</p></li><li><p><a href="https://github.com/andreluiz1987/semantic-search-market/blob/main/infra.py"><strong>创建索引</strong></a>，将描述字段设置为 semantic_text，以便接收嵌入信息。</p></li><li><p><a href="https://github.com/andreluiz1987/semantic-search-market/blob/main/ingestion.py"><strong>将数据索引</strong></a>到杂货目录索引中，该索引将存储产品目录。该目录是从<a href="https://www.kaggle.com/datasets/bhavikjikadara/grocery-store-dataset?select=GroceryDataset.csv">此处</a>提供的数据集中获取的。</p></li></ol><h2>语义搜索在超市中的应用</h2><p>现在，我们已经用杂货店产品数据填充了索引，我们正在测试和验证查询，以便使用语义搜索改进搜索结果。我们的目标是提供更智能的搜索体验，了解上下文和用户意图，提供更相关、更准确的搜索结果。</p><h3>语义搜索解决的挑战</h3><p>基于产品目录，让我们来探讨一下语义搜索如何通过解决词汇和上下文问题来改变杂货店的搜索体验，而传统的词汇搜索往往难以解决这些问题。</p><h4><strong>1.解读烹饪意图</strong></h4><p><strong>问题 01</strong>：客户可能会搜索"烧烤海鲜" ，但词法搜索系统可能无法完全理解查询背后的意图。它可能无法识别所有适合烧烤的海鲜产品，只返回产品标题中包含"seafood" 或"grill" 的确切术语的产品。</p><p>首先，我们将进行词性搜索并分析结果。然后，我们将进行同样的语义搜索，比较同一搜索词的搜索结果。</p><p><strong>查询词法搜索</strong></p> response = client.search(
        index="grocery-catalog",
        size=5,
        source_excludes="description_embedding",
        query={
            "multi_match": {
                "query": "seafood for grilling",
                "fields": [
                    "name",
                    "description"]
            }
        }
    )<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>词法</p><p>西北鱼阿拉斯加贝尔迪雪蟹</p><p>10.453125</p><p>词法</p><p>吉田先生，酱汁原味美食</p><p>7.2289705</p><p>词法</p><p>优质海鲜品种包 - 20 件</p><p>7.1924105</p><p>词法</p><p>美国红鲷鱼 - 整只、头朝上、洗净</p><p>6.998647</p><p>词法</p><p>龙虾爪&amp; 胳膊，可持续野生捕捞</p><p>6.438654</p><p>词条搜索返回了一些适合烧烤的海鲜产品，如美国红鲷鱼和西北鱼阿拉斯加 Bairdi 雪蟹。然而，词法搜索返回的相关性较低的产品排在了列表的前列，例如吉田先生酱，它不是一种海鲜产品，而是一种肉酱，这表明词法算法很难完全理解"用于烧烤的语境。"</p><p><strong>语义搜索解决方案</strong></p><p>我们使用的查询方式是将"seafood" 与"grilling" 等烹饪上下文结合起来，返回一个全面的选项列表，如鱼片、虾和扇贝，这些都是烧烤的理想选择--即使"grill" 或"seafood" 这些词没有直接出现在产品名称中。这可确保搜索结果更贴近客户的意图。</p><p><strong>查询语义搜索：</strong></p>es_client.search(
   index="grocery-catalog-elser",
   size=size,
   source_excludes="description_embedding",
   query={
       "semantic": {
           "field": "description_embedding",
           "query": "seafood for grilling"

       }
   })<p>搜索类型</p><p>名称</p><p>得分</p><p>语义学</p><p>去头洗净的整条鲈鱼</p><p>16.175909</p><p>语义学</p><p>阿拉斯加黑鳕鱼（黑貂鱼）</p><p>15.855331</p><p>语义学</p><p>美国红鲷鱼 - 整只，头朝下</p><p>15.454779</p><p>语义学</p><p>西北鱼阿拉斯加贝尔迪雪蟹</p><p>15.855331</p><p>语义学</p><p>美国红鲷鱼 - 整只，头朝下</p><p>15.3892355</p><p>语义搜索不仅返回了与"seafood," 这一术语直接相关的产品，而且还理解了"grilling," 这一上下文，并带来了适合烧烤的整鱼和鱼片。关键在于结果的精确性，其中包括烤制常用的全鱼，如白鲷鱼和阿拉斯加黑鳕鱼。</p><p><strong>问题 02 </strong> ：许多客户在工作一天后会搜索快速简便的晚餐解决方案，使用的术语包括"Easy Weeknight meals。"传统的词法搜索可能无法完全捕捉到快餐的概念，通常只关注名称中包含"easy" 一词的产品。</p><p>与上一个问题一样，我们将首先进行词法搜索。之后，我们将采用语义搜索来解决问题。</p><p><strong>查询词法搜索</strong></p> response = client.search(
        index="grocery-catalog",
        size=5,   
        source_excludes="description_embedding",
        query={
            "multi_match": {
                "query": "easy weeknight meals",
                "fields": [
                    "name",
                    "description"]
            }
        }
    )<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>词法</p><p>艾利易撕地址标签，4200 个装</p><p>8.017723</p><p>词法</p><p>自热应急/便携餐 32</p><p>6.592727</p><p>词法</p><p>海岸海鲜黄鳍金枪鱼块 Poke</p><p>5.836883</p><p>词法</p><p>Hefty 超重 12 盎司泡沫塑料</p><p>5.8116536</p><p>词法</p><p>Vanity Fair Everyday餐巾纸，2层，110片装</p><p>5.752989</p><p>词法搜索返回的相关结果要少得多，包括与餐饮完全无关的物品，如 Avery Easy Peel Address Labels 和 Vanity Fair Everyday Napkins。这些产品无法满足用户对快餐的需求。虽然词法搜索确实返回了一个有用的产品（Omeals Self Heating Emergency Meals），但其他结果，如餐巾纸和标签，仅与描述中的"easy" 或"weeknight" 匹配，没有真正满足用户对快速用餐解决方案的需求。</p><p><strong>语义搜索解决方案</strong></p><p>我们实施了一项查询，了解快速简便餐饮背后的意图。它将可快速烹制的产品，如预煮肉类、冷冻意大利面或套餐联系起来，即使这些产品的名称中没有明确包含"easy" 这个词。这种方法可确保顾客找到最合适的选择，在周末快速享用晚餐，满足对便利性的需求。</p><p><strong>查询语义搜索</strong></p>es_client.search(
   index="grocery-catalog-elser",
   size=size,
   source_excludes="description_embedding",
   query={
       "semantic": {
           "field": "description_embedding",
           "query": "easy weeknight meals"

       }
   })<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>语义学</p><p>自热应急/便携餐 32</p><p>14.610006</p><p>语义学</p><p>Nissin，杯面，虾，2.5 盎司</p><p>13.751424</p><p>语义学</p><p>Namaste 无谷蛋白华夫饼&amp; 煎饼预拌粉</p><p>13.73376</p><p>语义学</p><p>爱达荷土豆、黄金烤土豆饼</p><p>12.549422</p><p>语义学</p><p>Nissin 杯面，鸡肉，24 支装</p><p>12.034527</p><p>语义搜索返回的产品明显与方便快捷的膳食有关，如方便面（杯面）、预煮土豆和煎饼粉，这些都是简单的周末晚餐的典型选择。这表明，语义搜索可以抓住"简易隔夜饭这一短语背后的概念，" ，从而捕捉到用户寻找快捷方便饭菜的意图。有趣的是，其他类别的产品，如"苏打水、" ，在相关情况下（如佐餐饮料）也可能包括在内。</p><h4><strong>2.地区术语和词汇变化</strong></h4><p><strong>问题</strong>：一位客户可能会搜索"soda," ，而另一位客户可能会使用"pop" 来搜索相同的产品。传统的词库检索无法识别这两个词指的是同一个项目。</p><p><strong>查询词法搜索</strong></p> response = client.search(
        index="grocery-catalog",
        size=5,
        source_excludes="description_embedding",
        query={
            "multi_match": {
                "query": "refreshing pop drink low sugar",
                "fields": [
                    "name",
                    "description"]
            }
        }
    )<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>词法</p><p>Prime Hydration+ Sticks 电解质混合饮料</p><p>14.492869</p><p>词法</p><p>Capri Sun，100% 果汁，多种包装</p><p>12.340851</p><p>词法</p><p>Joyburst 能量饮料, Frose Rose, 12瓶装</p><p>11.839179</p><p>词法</p><p>Kellogg's Pop-Tarts, Frosted Brown Sugar Cinnamon</p><p>9.97788</p><p>词法</p><p>Kind 迷你巧克力棒, 多样包装, 0.7</p><p>9.336912</p><p>词法搜索侧重于精确的词语匹配。虽然它返回了 Prime Hydration 和 Capri Sun 等产品，但与"pop" 一词直接匹配也导致了不相关的结果，如 Kellogg's Pop-Tarts，它是一种零食而不是饮料。这凸显了当一个术语有多种含义或可能含糊不清时，词汇搜索的效果会如何降低。</p><p><strong>语义搜索解决方案</strong></p><p>在语义查询中，我们可以克服词汇搜索无法解决的词汇变化问题。通过扩展搜索条件，我们能够根据上下文的含义获得结果，提供更相关、更全面的回复。</p><p><strong>查询：</strong></p>es_client.search(
   index="grocery-catalog-elser",
   size=size,
   source_excludes="description_embedding",
   query={
       "semantic": {
           "field": "description_embedding",
           "query": "refreshing pop drink low sugar"

       }
   })<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>语义学</p><p>奥利葆 12 盎司益生元苏打水品种</p><p>14.776867</p><p>语义学</p><p>佰草集抗氧化可可粉，多种包装，18 磅</p><p>14.663253</p><p>语义学</p><p>怪物能量饮料，零度超能，24 磅</p><p>14.486348</p><p>语义学</p><p>Joyburst 能量饮料，12 盎司</p><p>14.007214</p><p>语义学</p><p>Joyburst 能量饮料, Frose Rose, 12瓶装</p><p>13.641038</p><p>语义搜索会返回与"pop" 这一概念直接匹配的产品，作为"soda" 的同义词（如 Olipop Prebiotics Soda），即使产品名称中可能没有"pop" 这一确切术语。搜索理解了用户的意图--清爽的低糖饮料--并能够返回相关产品，包括益生元苏打水（Olipop）和无糖能量饮料（Monster Energy Drink）等选项。</p><h2>结论</h2><p>事实证明，在食品杂货店中实施语义搜索对于理解复杂的查询非常有效，如"烧烤海鲜" 和"简易周末餐。"这种方法使我们能够更准确地解读用户意图，返回高度相关的产品。</p><p>通过使用 Elasticsearch 和 ELSER 简化流程，我们能够快速高效地应用语义搜索，显著改善搜索结果，提供更灵活、更有针对性的购物体验。这不仅优化了搜索过程，还提高了为客户提供的搜索结果的相关性。</p><h2>参考资料</h2><p>ELSER 型：</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/put-inference-api.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/put-inference-api.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-elser.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-elser.html</a></p><p></p><p>语义文本：</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-text.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-text.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html</a></p><p></p><p>数据集：</p><p><a href="https://www.kaggle.com/datasets/bhavikjikadara/grocery-store-dataset?select=GroceryDataset.csv">https://www.kaggle.com/datasets/bhavikjikadara/grocery-store-dataset?select=GroceryDataset.csv</a></p><p></p><p>语义搜索：</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search-semantic-text.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search-semantic-text.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/semantic-search-elasticsearch-ecommerce</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/semantic-search-elasticsearch-ecommerce</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c545fc80b6d79d6/6a170214839dfad776dcfd6f/d968e646240cd3ef7c79b5124d562a5f951d812b-1440x840.png" length="0" type="image/png"/>
    <pubDate>Thu, 07 Nov 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何通过 Apache Camel 向 Elasticsearch 采集数据]]></title>
    <description><![CDATA[通过实际示例了解如何通过 Apache Camel 将数据摄入 Elasticsearch。]]></description>
    <content:encoded><![CDATA[<p>使用 Apache Camel 将数据导入 Elasticsearch 的过程结合了搜索引擎的鲁棒性和集成框架的灵活性。在本文中，我们将探讨 Apache Camel 如何简化和优化 Elasticsearch 的数据摄取。为了说明这一功能，我们将实施一个入门应用程序，逐步演示如何配置和使用 Apache Camel 将数据发送到 Elasticsearch。</p><h2>什么是 Apache Camel？</h2><p>Apache Camel 是一个开源集成框架，可简化不同系统之间的连接，让开发人员专注于业务逻辑，而不必担心系统通信的复杂性。Camel 的核心概念是"routes，即" ，它定义了信息从原点到目的地的路径，可能包括转换、验证和过滤等中间步骤。</p><h3>Apache Camel 架构</h3><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt05652327efd4d5d1/6a17e6dafbc5f8a588491a8b/bef8145623a8fa80f929f9faa57ce0c460be2d0b-884x458.png" alt="Apache Camel 架构" /><p>Camel 使用"组件" 来连接不同的系统和协议，如数据库和消息服务，并使用"端点" 来表示消息的入口和出口。这些概念提供了模块化和灵活的设计，使其更易于高效、可扩展地配置和管理复杂的集成。</p><h2>使用 Elasticsearch 和 Apache Camel</h2><p>我们将演示如何配置一个简单的 Java 应用程序，使用 Apache Camel 将数据摄取到 Elasticsearch 集群中。还将介绍使用 Apache Camel 中定义的路由在 Elasticsearch 中创建、更新和删除数据的过程。</p><h3>1.添加依赖项</h3><p>配置此集成的第一步是在项目的<code>pom.xml</code> 文件中添加必要的依赖项。这将包括 Apache Camel 和 Elasticsearch 库。我们将使用新的 Java API 客户端库，因此必须导入<code>camel-elasticsearch</code> 组件，且版本必须与<code>camel-core</code> 库相同。</p><p>如果要使用 Java 低级 Rest Client，则必须使用 Elasticsearch 低级 Rest Client 组件。</p>&lt;dependency&gt;
   &lt;groupId&gt;org.apache.camel&lt;/groupId&gt;
   &lt;artifactId&gt;camel-core&lt;/artifactId&gt;
   &lt;version&gt;4.7.0&lt;/version&gt;
&lt;/dependency&gt;

&lt;dependency&gt;
   &lt;groupId&gt;org.apache.camel&lt;/groupId&gt;
   &lt;artifactId&gt;camel-elasticsearch&lt;/artifactId&gt;
   &lt;version&gt;4.7.0&lt;/version&gt;
&lt;/dependency&gt;

&lt;dependency&gt;
   &lt;groupId&gt;org.apache.camel&lt;/groupId&gt;
   &lt;artifactId&gt;camel-jackson&lt;/artifactId&gt;
   &lt;version&gt;4.7.0&lt;/version&gt;
&lt;/dependency&gt;

&lt;dependency&gt;
   &lt;groupId&gt;co.elastic.clients&lt;/groupId&gt;
   &lt;artifactId&gt;elasticsearch-java&lt;/artifactId&gt;
   &lt;version&gt;8.14.3&lt;/version&gt;
&lt;/dependency&gt;
<h3>2.配置和运行 Camel 内核</h3><p>配置的第一步是使用<code>DefaultCamelContext</code> 类创建一个新的 Camel 上下文，作为定义和执行路由的基础。接下来，我们配置 Elasticsearch 组件，它将允许 Apache Camel 与 Elasticsearch 集群交互。<code>ESlasticsearchComponent</code> 实例被配置为连接到<code>localhost:9200</code> 地址，这是本地 Elasticsearch 集群的默认地址。对于需要身份验证的环境设置，应阅读有关如何配置组件和启用基本身份验证的文档，即<strong>"配置组件和启用基本身份验证"</strong> 。</p>public class ESComponent {

    public static ElasticsearchComponent getInstance() {
        var elasticsearch = new ElasticsearchComponent();
        elasticsearch.setHostAddresses("localhost:9200");
        return elasticsearch;
    }

    public static String getName() {
        return "elasticsearch";
    }
}
<p>然后，该组件会被添加到 Camel 上下文中，使已定义的路由能够使用该组件在 Elasticsearch 中执行操作。</p>try (var context = new DefaultCamelContext()) {
   context.addComponent(ESComponent.getName(), ESComponent.getInstance());
   context.addRoutes(new OperationBulkRoute());
   context.start();
}
<p>之后，路由会被添加到上下文中。我们将创建用于批量索引、更新和删除文档的路由。</p><h3>3.配置 Camel 路由</h3><h4>数据索引</h4><p>我们要配置的第一个路由是用于数据索引。我们将使用一个包含电影目录的 JSON 文件。路由将被配置为读取位于<a href="https://gist.github.com/andreluiz1987/40756874b5fbea0a29586f9376d7f1f4"><code>src/main/resources/movies.json</code></a> 的文件，将 JSON 内容反序列化为 Java 对象，然后应用聚合策略将多条信息合并为一条，以便在 Elasticsearch 中进行批量操作。每条信息的大小配置为 500 条，也就是说，批量索引每次将索引 500 部影片。</p><p>批量路由 Elasticsearch 操作</p>String URI_BULK_OPERATION = String
       .format("elasticsearch://elasticsearch?operation=%s&amp;indexName=%s",
               IndexOperationConfig.BULK_OPERATION,
               INDEX_NAME);
public class OperationBulkRoute extends RouteBuilder {
   private static final Log log = LogFactory.getLog(OperationBulkRoute.class);
   private static final int BULK_SIZE = 500;

   @Override
   public void configure() {
       from("file:src/main/resources?fileName=movies.json&amp;noop=true")
               .routeId("route-bulk-ingest")
               .unmarshal().json()
               .split(body())
               .aggregate(constant(true), new BulkAggregationStrategy())
               .completionSize(BULK_SIZE)
               .to(URI_BULK_OPERATION)
               .process(exchange -&gt; {
                   var body = exchange.getIn().getBody(String.class);
                   log.info(String.format("Response: %s", body));
               })
               .end();
   }
}
<p>批量文件将被发送到 Elasticsearch 的批量操作端点。这种方法确保了处理大量数据时的效率和速度。</p><h4>数据更新</h4><p>下一步将是更新文件。在上一步中，我们为一些电影编制了索引，现在我们将创建新的路径，通过参考代码搜索文档，然后更新评级字段。</p><p>我们建立了一个 Camel 上下文<code>(DefaultCamelContext)</code> ，在其中注册了一个 Elasticsearch 组件，并添加了一个自定义路由 IngestionRoute。操作开始时，先通过 ProducerTemplate 发送文档代码，然后从 direct:update-ingestion 端点启动路由。</p>try (var context = new DefaultCamelContext()) {
    context.addComponent(ESComponent.getName(), ESComponent.getInstance());
    context.addRoutes(new IngestionRoute());
    context.start();
    ProducerTemplate producerTemplate = context.createProducerTemplate();
    producerTemplate.sendBody("direct:update-ingestion", documentCode);
    Thread.sleep(5000);
}
<p>接下来是 IngestionRoute，它是该流程的输入端点。路由执行多个流水线操作。首先，在 Elasticsearch 中进行搜索，按代码查找文件<code>(direct:search-by-id)</code> ，其中 SearchByCodeProcessor 根据代码组合查询。然后，UpdateRatingProcessor 对检索到的文档进行处理，将结果转换为电影对象，将电影分级更新为特定值，并准备将更新后的文档发回 Elasticsearch 进行更新。</p>public class IngestionRoute extends RouteBuilder {
    private static final Log log = LogFactory.getLog(IngestionRoute.class);

    @Override
    public void configure() throws Exception {

        from("direct:update-ingestion")
                .pipeline()
                .to("direct:search-by-id")
                .to(URI_SEARCH_OPERATION)
                .to("direct:update-rating")
                .to(URI_UPDATE_OPERATION)
                .process(exchange -&gt; {
                    var body = exchange.getIn().getBody(String.class);
                    log.info(String.format("Response: %s", body));
                })
                .end();

        from("direct:search-by-id")
                .process(new SearchByCodeProcessor());

        from("direct:update-rating")
                .process(new UpdateRatingProcessor());
    }
}
<p><code>SearchByCodeProcessor</code> 处理器的配置仅用于执行搜索查询：</p>public class SearchByCodeProcessor implements Processor {
    @Override
    public void process(Exchange exchange) throws Exception {
        var code = exchange.getIn().getBody();

        String query = "{\n" +
                "  \"query\": {\n" +
                "   \"term\": {\n" +
                "     \"code\": {\n" +
                "       \"value\":" + code + "\n" +
                "     }\n" +
                "   }\n" +
                "  }\n" +
                "}";
        exchange.setProperty("document_code", code);
        exchange.getIn().setBody(query);
    }
}
<p><code>UpdateRatingProcessor</code> 处理器负责更新评级字段。</p>public class UpdateRatingProcessor implements Processor {

    private final ObjectMapper objectMapper;

    public UpdateRatingProcessor() {
        this.objectMapper = new ObjectMapper();
        this.objectMapper.configure(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES, false);
    }

    @Override
    public void process(Exchange exchange) throws Exception {

        HitsMetadata response = exchange.getIn().getBody(HitsMetadata.class);
        var code = Long.parseLong(exchange.getProperty("document_code").toString());

        if (response != null &amp;&amp; response.hits() != null) {

            var documents = parseToMovies(response);

            var optionalMovie = documents.stream()
                    .filter(document -&gt; code == (document.getSource().getCode())).findAny();

            optionalMovie.ifPresent(document -&gt; {
                document.getSource().setRating(13.0);
                Map&lt;String, Object&gt; updateMap = new HashMap&lt;&gt;();
                updateMap.put("doc", document.getSource());
                exchange.getIn().setHeader("indexId", document.getId());
                exchange.getIn().setBody(updateMap);
            });
        }
    }
<h4>数据删除</h4><p>最后，配置删除文件的路径。在这里，我们将使用文档 ID 删除文档。在 Elasticsearch 中，要删除文档，我们需要知道文档标识符、存储该文档的索引并执行删除请求。在 Apache Camel 中，我们将通过创建一个新路由来执行此操作，如下图所示。</p><p>路由从 direct:op-delete 端点开始，它是入口点。需要删除文件时，会在邮件正文中收到其标识符<code>(_id)</code> 。然后，路由使用简单的<code>("${body}")</code> ，从报文正文中提取_id，用该标识符的值设置 indexId 头。</p>public class OperationDeleteRoute extends RouteBuilder {
   private static final Log log = LogFactory.getLog(OperationDeleteRoute.class);

   @Override
   public void configure() {
       from("direct:op-delete")
               .routeId("route-delete")
               .setHeader("indexId", simple("${body}"))
               .to(URI_DELETE_OPERATION)
               .process(exchange -&gt; {
                   var body = exchange.getIn().getBody(String.class);
                   log.info(String.format("Response: %s", body));
               })
               .end();
       ;
   }
}
String URI_DELETE_OPERATION = String
       .format("elasticsearch://elasticsearch?operation=%s&amp;indexName=%s",
               IndexOperationConfig.DELETE_OPERATION,
               INDEX_NAME);
<p>最后，消息会被定向到 URI_DELETE_OPERATION 指定的端点，该端点会连接到 Elasticsearch，以便在相应索引中执行文档删除操作。现在我们已经创建了路由，可以创建一个 Camel 上下文<code>(DefaultCamelContext)</code> ，该上下文已配置为包含 Elasticsearch 组件。</p>try (var context = new DefaultCamelContext()) {
   context.addComponent(ESComponent.getName(), ESComponent.getInstance());
   context.addRoutes(new OperationDeleteRoute());
   context.start();
   ProducerTemplate producerTemplate = context.createProducerTemplate();
   producerTemplate.sendBody("direct:op-delete", documentId);
}
<p>接下来，由<code>OperationDeleteRoute</code> 类定义的删除路由被添加到上下文中。初始化上下文后，<code>ProducerTemplate</code> ，将应删除文档的标识符传递给<code>direct:op-delete</code> 端点，从而触发删除路由。</p><h2>结论</h2><p>Apache Camel 和 Elasticsearch 之间的集成可实现稳健高效的数据摄取，利用 Camel 的灵活性定义路由，从而处理不同的数据操作场景，如索引、更新和删除。通过这种设置，您可以以可扩展的方式协调和自动化复杂的流程，确保您的数据在 Elasticsearch 中得到有效管理。该示例演示了如何将这些工具结合使用，以创建高效、适应性强的数据摄取解决方案。</p><h2>参考资料</h2><ul><li><p><a href="https://camel.apache.org/manual/">Apache Camel</a></p></li><li><p><a href="https://camel.apache.org/manual/architecture.html">Apache Camel 架构</a></p></li><li><p><a href="https://camel.apache.org/components/4.4.x/eips/aggregate-eip.html">聚合 Apache Camel</a></p></li><li><p><a href="https://camel.apache.org/components/4.4.x/file-component.html">文件组件</a></p></li><li><p><a href="https://camel.apache.org/components/4.4.x/elasticsearch-component.html">Elasticsearch 组件</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-apache-camel-ingest-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-apache-camel-ingest-data</guid>
    <category><![CDATA[索引数据]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt05652327efd4d5d1/6a17e6dafbc5f8a588491a8b/bef8145623a8fa80f929f9faa57ce0c460be2d0b-884x458.png" length="0" type="image/png"/>
    <pubDate>Mon, 09 Sep 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>