<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[基础功能 - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[基础功能 - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/cn/search-labs/blog/category/basics</link>
    </image>
    <link>https://www.elastic.co/cn/search-labs/blog/category/basics</link>
    <atom:link href="https://www.elastic.co/cn/search-labs/rss/category/basics.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[cn]]></language>
    <lastBuildDate>Tue, 29 Sep 2026 01:24:03 GMT</lastBuildDate>
  <item>
    <title><![CDATA[如何在 Azure AKS 上自动部署 Elasticsearch]]></title>
    <description><![CDATA[了解如何使用 AKS Automatic 和 ECK 在 Azure 上部署带有 Kibana 的 Elasticsearch，以实现部分托管的 Elasticsearch 设置配置。]]></description>
    <content:encoded><![CDATA[<p>本文是系列文章的一部分，我们将学习如何使用不同的基础架构安装 Elasticsearch。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt45071aec499098c0/6a17fe770b0beda3d1dd37f4/0a65ca8b62fd8a42d7751b8f4bf057e33d877304-940x458.png" alt="Elasticsearch 部署工作" /><p>与基于 Marketplace 的弹性云解决方案相比，ECK 需要付出更多努力，但它比自己部署虚拟机更加自动化，因为 Kubernetes 操作员将负责系统协调和节点扩展。</p><p>这一次，我们将使用自动功能与 Azure Kubernetes 服务 (AKS) 配合工作。在其他文章中，您将学习如何使用<a href="https://www.elastic.co/search-labs/blog/azure-elasticsearch-vm-deployment">Azure VM</a>和<a href="https://www.elastic.co/search-labs/blog/deploy-elasticsearch-azure-marketplace">Azure Marketplace</a>。</p><h2>什么是 AKS 自动系统？</h2><p><a href="https://learn.microsoft.com/en-us/azure/aks/intro-aks-automatic">Azure Kubernetes 服务（AKS）可自动 </a>管理集群设置、动态分配资源并集成安全最佳实践，同时保持 Kubernetes 的灵活性，使开发人员能够在几分钟内从容器镜像转为部署应用程序。</p><p>AKS Automatic 消除了大部分集群管理开销，在简单性和灵活性之间取得了良好的平衡。正确的选择取决于您的使用情况，但如果您计划这样做，决定就会容易得多：</p><ul><li><p><strong>部署测试环境： </strong>部署快速而简单，是快速实验或短期集群的理想选择。</p></li><li><p><strong>无需严格的虚拟机、存储或网络要求即可工作： </strong>AKS Automatic 提供预定义的默认设置，因此，如果这些设置符合您的需求，就可以省去额外的配置。</p></li><li><p><strong>首次使用 Kubernetes： </strong>通过处理集群的大部分设置工作，AKS Automatic 可降低学习曲线，让团队专注于自己的应用。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9e74556cda9b56bd/6a17fe791d1b830cdc93e681/2e4c09b8c5e0ce5e8ea9c369626a373b7030a5ba-854x489.png" alt=" 如何创建 Azure Kubernetes 服务 (AKS) 自动集群" /><p>对于 Elasticsearch，我们将使用<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s">Elastic Cloud on Kubernetes </a>(ECK)，它是官方的 Elastic Kubernetes 操作员，可以简化 Elastic Stack 的 Kubernetes 部署协调。</p><h2>如何设置 AKS 自动系统</h2><p>1.登录<a href="https://azure.microsoft.com/">Microsoft Azure 门户</a>。</p><p>2.在<strong>右上角， </strong>单击上的<strong> Cloud Shell</strong>按钮访问控制台，并从那里部署 AKS 群集。或者，您也可以使用<a href="https://learn.microsoft.com/en-us/azure/cloud-shell/overview">Azure 云外壳</a>。</p><p><em><strong>请记住，在教程中将项目 ID 更新为您的项目 ID。</strong></em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt06acd165140f9ab5/6a17fe7ae9ea876604a9c849/0aa60605777c0a6e3aef8faa4e54388c2cb582c8-624x495.png" alt="" /><p><em>打开 AKS 时的样子应该如上截图所示。</em></p><p>3.安装 aks-preview Azure CLI 扩展。该预览版允许我们在创建群集时选择<code>--sku automatic</code> ，从而启用 AKS 自动功能。</p>az extension add --name aks-preview<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt703add301bb94b89/6a17fe7c57726213da1bce20/2e05ab67fc554c5fb5208683c179fdeaeadd95db-624x56.png" alt="" /><p><em>如果看到此信息，说明 AKS 扩展已正确安装。</em></p><p>4.使用<code>az feature register</code> 命令注册<a href="https://learn.microsoft.com/en-us/azure/azure-app-configuration/concept-feature-management"> 功能标志</a></p>az feature register --namespace Microsoft.ContainerService --name AutomaticSKUPreview<p><em>您将看到我们刚刚创建的功能订阅的详细信息：</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt024ca2930c3a02e3/6a17fe7ee9ea87e003a9c84d/3aca710c1f312ba91de461638e518386919ec722-801x138.png" alt="" /><p>确认注册状态，直到从 "<em><strong>正在注册</strong></em>"变为 "<em><strong>已注册</strong></em>"。完成注册可能需要几分钟时间。</p>az feature show --namespace Microsoft.ContainerService --name AutomaticSKUPreview<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0b8d717b484bd461/6a17fe7f445de951aa4d0382/186486b08ab8e1c372efaff50f10cbddeaf4e0cd-844x177.png" alt="" /><p>运行<code>az provider register</code> 以传播更改。</p>az provider register --namespace Microsoft.ContainerService<p>5.创建资源组</p><p>资源组是要管理和部署的 Azure 资源的逻辑组。</p>az group create --name elastic-resource --location eastus<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbbf4b6e75c2dd560/6a17fe8163baff6114741e72/d1952269e97d94f914020754bd02702f9eafd037-770x212.png" alt="" /><p>6.创建自动驾驶仪群集。我们将把它命名为<em><strong>myAKSAutomaticCluster </strong></em>，并使用刚刚创建的资源组。确保以下任何一种虚拟机大小都有<em><strong>16 个</strong></em>可用vCPU：<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/dpsv5-series">Standard_D4pds_v5</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/dldsv5-series">Standard_D4lds_v5</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/dadsv5-series">Standard_D4ads_v5</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/ddsv5-series">Standard_D4ds_v5</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/ddv5-series">Standard_D4d_v5</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/ddv4-series">Standard_D4d_v4</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/general-purpose/dsv3-series">Standard_DS3_v2</a>、<a href="https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/memory-optimized/dv2-dsv2-series-memory">Standard_DS12_v2</a>，以便 AKS 分配资源。</p>az aks create \
    --resource-group elastic-resource \
    --name myAKSAutomaticCluster \
    --sku automatic \
    --generate-ssh-keys<p><em>* 如果出现 </em><em><code>MissingSubscriptionRegistration</code></em><em> 错误，请带着缺失的订阅返回第 4 步。例如， </em><em><code>The subscription is not registered to use namespace '</code></em><em><strong><code>microsoft.insights</code></strong></em> <em><code>'</code></em> 需要运行 <em><code>az provider register --namespace Microsoft.Insights.</code></em></p><p>按照交互式登录：</p><p><em>此时会出现一条要求运行 "az login "的信息。您必须运行该命令，然后等待。</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc711a241388cf786/6a17fe83dbb4fff454fb5929/14c0238f755fe6347519e69d3cb28c0fa52ec044-775x203.png" alt="" /><p>7.等待准备就绪。制作大约需要 10 分钟。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc389eecd12e3b2c8/6a17fe8425daab0b2408a442/eb00c3ad18f884f47db6645b196808ebec07c1fc-797x177.png" alt="" /><p>8.配置 kubectl 命令行访问权限。</p>az aks get-credentials --resource-group elastic-resource --name myAKSAutomaticCluster<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfdb13d3950950b06/6a17fe85b1e11319e179f4bd/5136d72a5d455345b0b6205bb232c4bdf7762998-793x52.png" alt="" /><p><em>请注意，我们安装的扩展正在启用 AKS Automatic。</em></p><p>9.确认节点已部署。</p>kubectl get nodes<p>您将看到一条禁止的错误信息；请复制错误信息中的用户 ID。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4b2bb8558e1a39f4/6a17fe87a292991ce3d02ea8/d6c021fa54f4db00d2d795f5ba9b5a93376d03cd-818x47.png" alt="" /><p>10.将用户添加到 AKS 访问控制中。</p><p>获取 AKS ID。复制命令输出。</p>az aks show --resource-group elastic-resource  --name myAKSAutomaticCluster --query id --output tsv<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0ebc4fa29348d561/6a17fe896df7317a3b0a1166/22a1cdc538bd379812a752c6a368a0651000abb8-810x36.png" alt="" /><p>使用 AKS ID 和用户的主要 ID 创建角色分配。</p>az role assignment create --role "Azure Kubernetes Service RBAC Cluster Admin" --assignee &lt;PRINCIPAL_ID&gt; --scope &lt;AKS_ID&gt;<p>11.尝试再次确认节点已部署。</p>kubectl get nodes<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltda416366f0e53d49/6a17fe8ab1e113f87c79f4c1/c9b3a5c1cc540ef732c3e7f60b0a973bdbd0b6fd-617x99.png" alt="" /><p>12.安装 Kubernetes 上的弹性云（ECK）操作员。</p># Install ECK Custom Resource Definitions
kubectl create -f https://download.elastic.co/downloads/eck/2.16.1/crds.yaml

# Install the ECK operator
kubectl apply -f https://download.elastic.co/downloads/eck/2.16.1/operator.yaml<p>13.让我们使用默认值创建一个单节点 Elasticsearch 实例。</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: elasticsearch.k8s.elastic.co/v1
kind: Elasticsearch
metadata:
  name: quickstart
spec:
  version: 9.0.0
  nodeSets:
  - name: default
    count: 1
    config:
      node.store.allow_mmap: false
EOF<p>我们禁用<code>nmap</code> 是因为默认 AKS 机器的<code>vm.max_map_count</code> 值过低。不建议在生产中禁用它，但可以增加<code>vm.max_map_count</code> 的值。您可以<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s/virtual-memory">在这里</a>阅读更多关于如何做到这一点的信息。</p><p>14.我们也来部署一个 Kibana 单节点集群。对于 Kibana，我们将添加一个负载平衡器，它将为我们提供一个外部 IP，我们可以用它从我们的设备访问 Kibana。</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: kibana.k8s.elastic.co/v1
kind: Kibana
metadata:
  name: quickstart
spec:
  version: 9.0.0
  http:
    service:
      spec:
        type: LoadBalancer
  count: 1
  elasticsearchRef:
    name: quickstart
EOF<p>默认情况下，AKS Automatic 会将负载平衡器配置为公共负载平衡器；您可以通过设置元数据注释来更改行为：</p><p><code>service.beta.kubernetes.io/azure-load-balancer-internal: "true"</code></p><p>15.检查 pod 是否正在运行。</p>kubectl get pods<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltea98380109a7a0b1/6a17fe8c4b055dda6c43241e/213a897176c0af6cea19c7c777cfaf8734e3ee6e-616x84.png" alt="" /><p>16.您还可以运行<code>kubectl get elasticsearch</code> 和<code>kubectl get kibana</code> 获取更具体的统计信息，如 Elasticsearch 版本、节点和健康状况。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt89119d59582673d9/6a17fe8d4b055d4257432422/c84988e725ef892eddd8fb7e5a03d58c35a8f9d6-470x62.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte714550df3504cd4/6a17fe8f7b54f982a28b3b19/452dd03d314cd00c8a3c19e19862b968592a0435-415x62.png" alt="" /><p>17.获取您的服务。</p>kubectl get svc<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc9c24dc9b78a354/6a17fe900b0beda9f8dd37f8/b2d3e8f368be22b89aa2ed4d4d514f97dd6cbabd-624x115.png" alt="" /><p>这将在 EXTERNAL-IP 下显示 Kibana 的<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s/accessing-services">外部 URL</a>。负载平衡器的调配可能需要几分钟时间。<em><strong>复制 EXTERNAL-IP 的值。</strong></em></p><p>18.获取 "elastic "用户的 Elasticsearch 密码：</p>kubectl get secret quickstart-es-elastic-user -o=jsonpath='{.data.elastic}' | base64 --decode<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltec8ef3a4df60e7a8/6a17fe9263baff3381741e76/bd74537f8c35c4e027c518913fdb0a0524621d56-624x31.png" alt="" /><p>19.通过浏览器<strong>访问 Kibana</strong>：</p><p>a.url: https://&lt;EXTERNAL_IP&gt;:5601</p><p>b.用户名：elastic</p><p>c.密码：c44A295CaEt44D6xIzN6Zs5m（来自上一步）</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7e8fbd1593bae329/6a17fe937b54f91d7c8b3b1d/a601112527d80721b292328ed8da58386d2837eb-463x503.png" alt="" /><p>20.从浏览器访问 Elastic Cloud 时，您将看到欢迎屏幕。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9ad7cd3196eb7dbe/6a17fe95b1e11336a579f4c5/f91e71fa961d215a8d886601d1a9fc5c452ce329-1999x1256.png" alt="" /><p>如果要更改 Elasticsearch 集群规格，如更改或调整节点大小，可以使用新设置再次应用 YML 清单：</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: elasticsearch.k8s.elastic.co/v1
kind: Elasticsearch
metadata:
  name: quickstart
spec:
  version: 9.0.0
  nodeSets:
    - name: default
      count: 2
      config:
        node.store.allow_mmap: false
      podTemplate:
        spec:
          containers:
            - name: elasticsearch
              resources:
                requests:
                  memory: 1.5Gi
                  cpu: 2
                limits:
                  memory: 1.5Gi
                  cpu: 2
EOF<p>在本例中，我们将增加一个节点，并修改 RAM 和 CPU。如您所见，现在<code>kubectl get elasticsearch</code> 显示了 2 个节点：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdfe6ff383e03814e/6a17fe974b055dd514432426/4b139a476b50933d45d99e09479112817964f76a-624x60.png" alt="" /><p>Kibana 也是如此：</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: kibana.k8s.elastic.co/v1
kind: Kibana
metadata:
  name: quickstart
spec:
  version: 9.0.0
  http:
    service:
      spec:
        type: LoadBalancer
  count: 1
  elasticsearchRef:
    name: quickstart
  podTemplate:
    spec:
      containers:
        - name: kibana
          env:
            - name: NODE_OPTIONS
              value: "--max-old-space-size=1024"
          resources:
            requests:
              memory: 0.5Gi
              cpu: 0.5
            limits:
              memory: 1Gi
              cpu: 1
EOF<p>我们可以调整容器的 CPU/RAM，也可以调整<a href="https://nodejs.org/">Node.js </a>的内存使用量<a href="https://nodejs.org/api/cli.html#--max-old-space-sizesize-in-mib">（max-old-space-size</a>）。</p><p>请记住，<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s/volume-claim-templates">现有的批量索赔不能缩减</a>。应用更新后，操作员将在最短的时间内完成更改。</p><p>测试完成后，请记住删除群集，以避免不必要的成本。</p>az aks delete --name myAKSAutomaticCluster --resource-group elastic-resource<h2>结论</h2><p>使用 Azure AKS Automatic 和 ECK 可为部署 Elasticsearch 和 Kibana 提供一个平衡的解决方案：它降低了操作复杂性，确保了自动扩展和更新，并充分利用了 Kubernetes 的灵活性。这种方法非常适合需要可靠、可重复和可维护的部署流程，而无需手动管理每个基础架构细节的团队，使其成为测试和生产环境的实用选择。</p><h2>后续步骤</h2><p>如果您想了解有关 Kubernetes 的更多信息，可点击此处查看官方文档：</p><ul><li><p><a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s">Kubernetes 上的弹性云 | 弹性文档</a></p></li><li><p><a href="https://learn.microsoft.com/en-us/azure/aks/intro-aks-automatic">Azure Kubernetes 服务 (AKS) 自动化介绍（预览版）</a></p></li><li><p><a href="https://azure.github.io/AKS/2024/05/22/aks-automatic">AKS Automatic - AKS 工程博客</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-azure-aks-automatic-deployment</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-azure-aks-automatic-deployment</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Eduard Martin]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt58d08dac9d3586db/6a17fe983e9e45002eba16e7/4d821659a606e04390b09215e9a0d32eb01f0d1b-854x489.png" length="0" type="image/png"/>
    <pubDate>Fri, 14 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[在 Elasticsearch 中为结构化文档配置递归分块]]></title>
    <description><![CDATA[了解如何在 Elasticsearch 中使用分块大小、分隔符组和自定义分隔符列表配置递归分块，以优化结构文档索引。]]></description>
    <content:encoded><![CDATA[<p>自 8.16 版起，用户可以配置将长文档导入语义文本字段时使用的分块策略。从 9.1 / 8.19 版开始，我们引入了一种新的可配置递归分块策略，使用正则表达式列表对文档进行分块。分块的目的是将长文档分割成囊括相关内容的部分。我们现有的策略会按单词/句子的粒度分割文本，但以结构化格式编写的文档（例如："......"）则不会这样做。Markdown）通常会在由一些分隔字符串定义的部分内包含相关内容（例如："......"）。标题）。对于这些类型的文档，我们正在引入递归分块策略，以利用结构化文档的格式来创建更好的分块！</p><h2>什么是递归分块？</h2><p>递归分块法会遍历所提供的分块模式列表，逐步将文档分成更小的分块，直到达到所需的最大分块大小。</p><h3>如何配置递归分块？</h3><p>以下是用户为递归分块提供的可配置值：</p><ul><li><p>(必填）<code>max_chunk_size</code> ：字块中的最大字数。</p></li><li><p>任选其一：</p><ul><li><p><code>separators</code>:用于将文档分割成块的 regex 字符串模式列表。</p></li><li><p><code>separator_group</code>:一个字符串，它将映射到 Elastic 定义的默认分隔符列表，用于特定类型的文档。目前，<code>markdown</code> 和<code>plaintext</code> 。</p></li></ul></li></ul><h3>递归分块是如何工作的？</h3><p>递归分块的过程如下：给定输入文档、<code>max_chunk_size</code> （以字数为单位）和分隔符字符串列表：</p><ol><li><p>如果输入文档已经在最大分块大小范围内，则返回一个涵盖整个输入文档的分块。</p></li><li><p>根据分隔符的出现次数，将文本分割成潜在的文本块。对于每个潜在的数据块</p><ol><li><p>如果潜在数据块在最大数据块大小范围内，则将其添加到要返回给用户的数据块列表中。</p></li><li><p>否则，从第 2 步开始重复，只使用潜在文本块中的文本，并使用列表中的下一个分隔符进行分割。如果没有其他分隔符可以尝试，就退回到基于句子的分块。</p></li></ol></li></ol><h2>配置递归分块的示例</h2><p>除了分块大小，递归分块的主要配置是选择应使用哪些分隔符来分割文档。如果您不确定从哪里开始，Elasticsearch 提供了一些默认的分离器组，可用于常见的使用情况。</p><h3>利用分离器组</h3><p>要使用分隔组，只需在配置分块设置时提供要使用的组名即可。例如</p>"chunking_settings": {
    "strategy": "recursive",
    "max_chunk_size": 25,
    "separator_group": "plaintext"
}<p>这样就可以利用分隔符列表<code>["(?&lt;!\\n)\\n\\n(?!\\n)", "(?&lt;!\\n)\\n(?!\\n)")]</code> 来实现递归分块策略。对于一般的纯文本应用程序，这种方法效果很好，可以在 2 个换行符后再分隔出 1 个换行符。</p><p>我们还提供一个分隔符组<code>markdown</code> ，它将利用分隔符列表：</p>[
"\n# ",
       "\n## ",
       "\n### ",
       "\n#### ",
       "\n##### ",
       "\n###### ",
       "\n^(?!\\s*$).*\\n-{1,}\\n",
       "\n^(?!\\s*$).*\\n={1,}\\n"
]<p>这个分隔符列表可以很好地适用于一般的标记符使用情况，在 6 个标题层次和分节符上分别进行分隔。</p><p>创建资源（推理端点/语义文本字段）时，与当时分隔符组相对应的分隔符列表将存储在您的配置中。如果以后更新了分隔符组，也不会改变已创建资源的行为。</p><h3>使用自定义分隔符列表</h3><p>如果预定义的分隔符组不适合您的使用情况，您可以定义一个符合您需求的自定义分隔符列表。请注意，可以在分隔符列表中提供正则表达式。以下是使用自定义分隔符配置分块设置的示例：</p>"chunking_settings": {
    "strategy": "recursive",
    "max_chunk_size": 25,
    "separators": ["\n\n", "\n", "&lt;my-custom-separator&gt;"]
}<p>上述分块策略将在 2 个换行符、1 个换行符和一个字符串<code>“&lt;my-custom-separator&gt;”</code> 上进行分割。</p><h2>递归分块的实际应用示例</h2><p>让我们来看一个递归分块的实例。在本示例中，我们将使用以下分块设置和自定义分隔符列表，使用顶部两层标题分割标记符文档：</p>"chunking_settings": {
    "strategy": "recursive",
    "max_chunk_size": 25,
    "separators": ["\n# ", "\n## "]
}<p>让我们来看看一个简单的未分块 Markdown 文档：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdb5f41d1bd43ba50/6a17e831e9ea87c1d8a9c5f3/3a5507f4a1288065097231548e5b18e240508785-1302x1446.png" alt="未分块的 Markdown 文档" /><p>现在，让我们使用上面定义的分块设置对文档进行分块：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfffda162c7b9c87a/6a17e83296142aefa8eb1b0b/a3313c4c40ff39b8dbcdd7c4878c723f088e6c1a-1600x1187.png" alt="在 Elasticsearch 中将文档分块" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt96f65346a8e09e3a/6a17e834445de9157b4d015e/79a2921943191ea631df94c9d465818ec8d3e738-1600x1206.png" alt="在第二个分隔符上拆分--在 Elasticsearch 中将文档分块" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt28381c8f85aedf07/6a17e836ec0f89801e5a6640/459e695cce7540267422396b9a62ff4ad35f61db-1600x1260.png" alt="Elasticsearch 中基于句子的分块处理后文档中的最终分块" /><p>注意：每个分块（分块 3 除外）末尾的换行符不会突出显示，而是包含在实际分块边界内。</p><h3>今天就开始使用递归分块技术！</h3><p>有关使用该功能的更多信息，请查看有关配置分块设置的文档。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/recursive-chunking-structured-documents-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/recursive-chunking-structured-documents-elasticsearch</guid>
    <category><![CDATA[基础功能]]></category>
    <category><![CDATA[在 Elastic 内部]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Daniel Rubinstein]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf442dc4941f37be7/6a17e838505ac3eaf8ad8b3d/591872e31880768ca927507654a621addc0d124d-1600x960.png" length="0" type="image/png"/>
    <pubDate>Tue, 11 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[介绍 Kibana 中的 Elasticsearch 查询规则用户界面]]></title>
    <description><![CDATA[了解如何使用 Elasticsearch 查询规则用户界面，在 Kibana 中使用可定制的规则集从搜索查询中添加或排除文档，而不影响有机排名。]]></description>
    <content:encoded><![CDATA[<p>搜索引擎的工作就是返回相关结果。然而，有些业务需求并不限于此，比如突出销售、优先考虑季节性产品或展示赞助项目，而开发人员不可能总是在搜索查询中做到这一点。</p><p>此外，这些用例通常具有时间敏感性，而经历典型的开发阶段（创建代码分支，然后等待新版本发布）是一个耗时的过程。</p><p>那么，如果我们只需调用 API，或者在 Kibana 中点击几下就能完成整个过程，那会怎样呢？</p><h2>查询规则用户界面</h2><p>Elasticsearch 8.10 引入了<a href="https://www.elastic.co/blog/introducing-query-rules-elasticsearch-8-10"><strong>查询规则</strong></a>和<a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/retrievers/rule-retriever"><strong>规则检索器</strong></a>。这些工具旨在根据规则在不影响有机结果排名的情况下将<a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-pinned-query"><em>钉入结果</em></a>注入查询。它们只是以声明式的简单方式在结果之上添加业务逻辑。</p><p>查询规则的一些常见用例包括</p><ul><li><p><strong>突出显示促销列表或销售</strong>：在顶部显示促销或赞助商品。</p></li><li><p><strong>根据上下文或地理位置排除</strong>：当当地法规不允许显示某些项目时，隐藏这些项目。</p></li><li><p><strong>优先处理关键结果</strong>：确保热门搜索或固定搜索始终排在前面，无论有机搜索排名如何。</p></li></ul><p>要访问界面并与这些工具互动，需要点击 Kibana 侧边菜单，然后转到相关性下的<strong>查询规则</strong> <strong>：</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltac12541cddd58e36/6a170853a29299941cd00fc2/242e33e89d1a07ffa0e76009c46b3a9236722741-458x1010.png" alt="在相关性下访问 Elasticsearch 中的查询规则" /><p>查询规则菜单弹出后，点击<strong>创建第一个规则集：</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcc28329c0f3c3aa9/6a17085547d49c67e22d893b/30b3a91bbbf243d314cf38298e01ca5cff784430-1600x945.png" alt="在 Elasticsearch 中创建第一个查询规则集" /><p>接下来，您需要为规则集命名。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb37d271297a4f148/6a170856a29299782cd00fc6/26c5462f88678867776f933b5655ca0df0d72a16-708x446.png" alt="在 Elasticsearch 中命名查询规则集" /><p>定义每条规则的表格有三个关键部分：</p><ul><li><p><strong>标准</strong>：适用规则必须满足的条件。例如，"当 query_string 字段包含<em>Christmas</em>值时 "或 "当 country 字段为<em>CO 时"。</em></p></li><li><p><strong>行动</strong>：这是您希望在条件满足时发生的事情。它可以被固定（将文档固定到顶部结果）或排除（隐藏文档）。</p></li><li><p><strong>元数据</strong>：这些字段在查询运行时会随查询一起出现。它们可以包括用户信息（如位置或语言）以及搜索数据（query_string）。这些值是标准用于决定是否应用规则的值。</p></li></ul><h2>例如：热门项目</h2><p>假设我们有一个电子商务网站，上面有不同的商品。在查看这些指标时，我们注意到在游戏机类别中，"DualShock 4 无线控制器 "是销售量最大的商品之一，尤其是当用户搜索关键词 "PS4 "或 "PlayStation 4 "时。因此，我们决定在用户搜索这些关键词时，将该产品放在搜索结果的顶部。</p><p>首先，让我们使用批量 API 请求为每个项目的文档建立索引：</p>POST _bulk
{ "index": { "_index": "products", "_id": "1" } }
{ "id": "1", "name": "PlayStation 4 Slim 1TB", "category": "console", "brand": "Sony", "price": 1200 }
{ "index": { "_index": "products", "_id": "2" } }
{ "id": "2", "name": "DualShock 4 Wireless Controller", "category": "accessory", "brand": "Sony", "price": 250 }
{ "index": { "_index": "products", "_id": "3" } }
{ "id": "3", "name": "PlayStation 4 Camera", "category": "accessory", "brand": "Sony", "price": 200 }
{ "index": { "_index": "products", "_id": "4" } }
{ "id": "4", "name": "PlayStation 4 VR Headset", "category": "accessory", "brand": "Sony", "price": 900 }
{ "index": { "_index": "products", "_id": "5" } }
{ "id": "5", "name": "Charging Station for DualShock 4", "category": "accessory", "brand": "Sony", "price": 80 }<p>如果我们不干预查询，该项目通常会出现在第四位。问题是这样的</p>GET products/_search
{
 "query": {
   "match": {
     "name": "PlayStation 4"
   }
 }
}<p>结果如下</p>{
 "took": 1,
 "timed_out": false,
 "_shards": {
   "total": 1,
   "successful": 1,
   "skipped": 0,
   "failed": 0
 },
 "hits": {
   "total": {
     "value": 5,
     "relation": "eq"
   },
   "max_score": 0.6973252,
   "hits": [
     {
       "_index": "products",
       "_id": "3",
       "_score": 0.6973252,
       "_source": {
         "id": "3",
         "name": "PlayStation 4 Camera",
         "category": "accessory",
         "brand": "Sony",
         "price": 200
       }
     },
     {
       "_index": "products",
       "_id": "1",
       "_score": 0.6260078,
       "_source": {
         "id": "1",
         "name": "PlayStation 4 Slim 1TB",
         "category": "console",
         "brand": "Sony",
         "price": 1200
       }
     },
     {
       "_index": "products",
       "_id": "4",
       "_score": 0.6260078,
       "_source": {
         "id": "4",
         "name": "PlayStation 4 VR Headset",
         "category": "accessory",
         "brand": "Sony",
         "price": 900
       }
     },
     {
       "_index": "products",
       "_id": "2",
       "_score": 0.08701137,
       "_source": {
         "id": "2",
         "name": "DualShock 4 Wireless Controller",
         "category": "accessory",
         "brand": "Sony",
         "price": 250
       }
     },
     {
       "_index": "products",
       "_id": "5",
       "_score": 0.07893815,
       "_source": {
         "id": "5",
         "name": "Charging Station for DualShock 4",
         "category": "accessory",
         "brand": "Sony",
         "price": 80
       }
     }
   ]
 }
}<p>让我们创建一个查询规则来改变这种情况。首先，让我们像这样把它添加到规则集中：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1576d4f4a2e60548/6a170858cdacbfccb07d298d/fdc42646fb3e76a09bca7d19047a76efe343f7a2-1600x650.png" alt="如何在 Elasticsearch 中编辑查询规则集" /><p>或相应的<a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-query-rules-put-ruleset">API 请求</a>：</p>PUT _query_rules/my-rules
{
  "rules": [
    {
      "rule_id": "rule-1232",
      "type": "pinned",
      "criteria": [
        {
          "type": "exact",
          "metadata": "query_string",
          "values": [
            "PS4",
            "PlayStation 4"
          ]
        }
      ],
      "actions": {
        "docs": [
          {
            "_index": "products",
            "_id": "2"
          }
        ]
      }
    }
  ]
}<p>要在查询中使用<strong>规则集 </strong>，我们必须使用查询规则类型。这种查询主要由两部分组成：</p>GET /products/_search
{
 "retriever": {
   "rule": {
     "retriever": {
       "standard": {
         "query": {
           "match": { "name": "PlayStation 4" }
         }
       }
     },
     "match_criteria": {
       "query_string": "PlayStation 4"
     },
     "ruleset_ids": ["my-rules"]
   }
 }
}<ul><li><p><strong>匹配标准</strong>：这些是用于与用户查询进行比较的元数据。在本例中，当 query_string 字段的值为 "PlayStation 4 "时，规则集被激活。</p></li><li><p><strong>query</strong>：实际查询，用于搜索和获取有机结果。</p></li></ul><p>这样，首先运行有机查询，然后 Elasticsearch 应用规则集中的规则：</p>{
 "took": 17,
 "timed_out": false,
 "_shards": {
   "total": 1,
   "successful": 1,
   "skipped": 0,
   "failed": 0
 },
 "hits": {
   "total": {
     "value": 5,
     "relation": "eq"
   },
   "max_score": 1.7014122e+38,
   "hits": [
     {
       "_index": "products",
       "_id": "2",
       "_score": 1.7014122e+38,
       "_source": {
         "id": "2",
         "name": "DualShock 4 Wireless Controller",
         "category": "accessory",
         "brand": "Sony",
         "price": 250
       }
     },
     {
       "_index": "products",
       "_id": "3",
       "_score": 0.6973252,
       "_source": {
         "id": "3",
         "name": "PlayStation 4 Camera",
         "category": "accessory",
         "brand": "Sony",
         "price": 200
       }
     },
     {
       "_index": "products",
       "_id": "1",
       "_score": 0.6260078,
       "_source": {
         "id": "1",
         "name": "PlayStation 4 Slim 1TB",
         "category": "console",
         "brand": "Sony",
         "price": 1200
       }
     },
     {
       "_index": "products",
       "_id": "4",
       "_score": 0.6260078,
       "_source": {
         "id": "4",
         "name": "PlayStation 4 VR Headset",
         "category": "accessory",
         "brand": "Sony",
         "price": 900
       }
     },
     {
       "_index": "products",
       "_id": "5",
       "_score": 0.07893815,
       "_source": {
         "id": "5",
         "name": "Charging Station for DualShock 4",
         "category": "accessory",
         "brand": "Sony",
         "price": 80
       }
     }
   ]
 }
}<h2>示例：基于用户的元数据</h2><p>查询规则的另一个有趣应用是使用元数据，根据用户或网页的上下文信息显示特定文档。</p><p>例如，假设我们想根据用户的忠诚度（用数值表示）来突出显示商品或定制销售。</p><p>我们可以直接将这些元数据导入查询，这样当所述值满足特定条件时，规则就会激活。</p><p>首先，我们将为一份只有忠诚度高的用户才能看到的文档建立索引：</p>POST _bulk
{ "index": { "_index": "products", "_id": "6" } }
{ "id": "6", "name": "PlayStation Plus Deluxe Card - 12 months", "category": "membership", "brand": "Sony", "price": 300 }<p>现在，让我们在同一规则集内创建一条新规则，这样当忠诚度_级别等于或高于 80 时，项目就会出现在结果的顶部。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt158578005df8c76d/6a17085aab7f086dc0db9de3/58de12dff93305440608f51465462fcc68653a08-1421x496.png" alt="如何在 Elasticsearch 中编辑查询规则集" /><p>保存规则和规则集。</p><p>以下是相应的 REST 请求：</p>PUT _query_rules/my-rules
{
  "rules": [
    {
      "rule_id": "pin-premiun-user",
      "type": "pinned",
      "criteria": [
        {
          "type": "gte",
          "metadata": "loyalty_level",
          "values": [
            80
          ]
        }
      ],
      "actions": {
        "docs": [
          {
            "_index": "products",
            "_id": "6"
          }
        ]
      }
    }
  ]
}<p>现在，在运行查询时，我们需要在元数据中包含新参数<strong>loyalty_level </strong>。如果满足规则中的条件，新文档将出现在结果的顶部。</p><p>例如，在发送忠诚度级别为 80 的查询时：</p>POST /products/_search
{
  "retriever": {
    "rule": {
      "retriever": {
        "standard": {
          "query": {
            "match": {
              "name": "PlayStation"
            }
          }
        }
      },
      "match_criteria": {
        "query_string": "PlayStation",
        "loyalty_level": 80
      },
      "ruleset_ids": ["my-rules"]
    }
  }
}<p>我们将在结果上方看到忠诚度文件：</p>{
  "took": 31,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4,
      "relation": "eq"
    },
    "max_score": 1.7014122e+38,
    "hits": [
      {
        "_index": "products",
        "_id": "6",
        "_score": 1.7014122e+38,
        "_source": {
          "id": "6",
          "name": "PlayStation Plus Deluxe Card - 12 months",
          "category": "membership",
          "brand": "Sony",
          "price": 300
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 0.5054567,
        "_source": {
          "id": "3",
          "name": "PlayStation 4 Camera",
          "category": "accessory",
          "brand": "Sony",
          "price": 200
        }
      },
      {
        "_index": "products",
        "_id": "1",
        "_score": 0.45618832,
        "_source": {
          "id": "1",
          "name": "PlayStation 4 Slim 1TB",
          "category": "console",
          "brand": "Sony",
          "price": 1200
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 0.45618832,
        "_source": {
          "id": "4",
          "name": "PlayStation 4 VR Headset",
          "category": "accessory",
          "brand": "Sony",
          "price": 900
        }
      }
    ]
  }
}<p>在下面的例子中，由于忠诚度等级为 70，因此不符合规则，物品不应出现在顶部：</p>POST /products/_search
{
  "retriever": {
    "rule": {
      "retriever": {
        "standard": {
          "query": {
            "match": {
              "name": "PlayStation"
            }
          }
        }
      },
      "match_criteria": {
        "query_string": "PlayStation",
        "loyalty_level": 70
      },
      "ruleset_ids": ["my-rules"]
    }
  }
}<p>结果如下：</p>{
  "took": 7,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4,
      "relation": "eq"
    },
    "max_score": 0.5054567,
    "hits": [
      {
        "_index": "products",
        "_id": "3",
        "_score": 0.5054567,
        "_source": {
          "id": "3",
          "name": "PlayStation 4 Camera",
          "category": "accessory",
          "brand": "Sony",
          "price": 200
        }
      },
      {
        "_index": "products",
        "_id": "1",
        "_score": 0.45618832,
        "_source": {
          "id": "1",
          "name": "PlayStation 4 Slim 1TB",
          "category": "console",
          "brand": "Sony",
          "price": 1200
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 0.45618832,
        "_source": {
          "id": "4",
          "name": "PlayStation 4 VR Headset",
          "category": "accessory",
          "brand": "Sony",
          "price": 900
        }
      },
      {
        "_index": "products",
        "_id": "6",
        "_score": 0.3817649,
        "_source": {
          "id": "6",
          "name": "PlayStation Plus Deluxe Card - 12 months",
          "category": "membership",
          "brand": "Sony",
          "price": 300
        }
      }
    ]
  }
}<h2>例如：立即排除</h2><p>假设我们的<strong>DualShock 4 无线控制器（ID 2）</strong>暂时缺货，无法出售。因此，业务团队决定在此期间将其从搜索结果中删除，而不是手动删除文档或等待某些数据流程启动。</p><p>我们将使用与刚才应用于热门项目类似的过程，但这次我们不选择 "<em>已固定"</em>，而是选择 "<em>排除</em>"。这条规则就像一个黑名单。将条件改为 "<strong>始终"</strong>，这样每次运行查询时，排除都会起作用。</p><p>规则应该是这样的</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt38564c0b7f4a6ee2/6a17085c1949f78692e7a989/f10971e4f1bc9520105111adfa3a476581a27130-1600x623.png" alt="Elasticsearch 中立即排除规则集的示例" /><p>保存规则和规则集以应用更改。以下是相应的 REST 请求：</p>PUT _query_rules/my-rules
{
  "rules": [
    {
      "rule_id": "rule-6358",
      "type": "pinned",
      "criteria": [
        {
          "type": "always"
        }
      ],
      "actions": {
        "docs": [
          {
            "_index": "products",
            "_id": "2"
          }
        ]
      }
    }
  ]
}<p>现在，当我们再次运行查询时，你会发现结果中不再有该项目，尽管之前的规则是将其固定。这是因为<strong>排除结果的优先级高于钉牢结果</strong>。</p>{
 "took": 6,
 "timed_out": false,
 "_shards": {
   "total": 1,
   "successful": 1,
   "skipped": 0,
   "failed": 0
 },
 "hits": {
   "total": {
     "value": 4,
     "relation": "eq"
   },
   "max_score": 2.205655,
   "hits": [
     {
       "_index": "products",
       "_id": "3",
       "_score": 2.205655,
       "_source": {
         "id": "3",
         "name": "PlayStation 4 Camera",
         "category": "accessory",
         "brand": "Sony",
         "price": 200
       }
     },
     {
       "_index": "products",
       "_id": "1",
       "_score": 1.9738505,
       "_source": {
         "id": "1",
         "name": "PlayStation 4 Slim 1TB",
         "category": "console",
         "brand": "Sony",
         "price": 1200
       }
     },
     {
       "_index": "products",
       "_id": "4",
       "_score": 1.9738505,
       "_source": {
         "id": "4",
         "name": "PlayStation 4 VR Headset",
         "category": "accessory",
         "brand": "Sony",
         "price": 900
       }
     },
     {
       "_index": "products",
       "_id": "5",
       "_score": 0.69247496,
       "_source": {
         "id": "5",
         "name": "Charging Station for DualShock 4",
         "category": "accessory",
         "brand": "Sony",
         "price": 80
       }
     }
   ]
 }
}<h2>结论</h2><p><strong>查询规则</strong>使调整相关性变得非常容易，无需修改任何代码。新的<strong>Kibana</strong> <strong>UI </strong>允许在几秒钟内做出这些更改，让您和您的业务团队对搜索结果拥有更多控制权。</p><p>除电子商务外，查询规则还能支持许多其他应用场景：在支持门户中突出显示故障排除指南，在知识库中显示关键的内部文档，在新闻网站中宣传突发事件，或过滤掉过期的职位或内容列表。它们甚至可以执行合规规则，如根据用户角色或地区隐藏受限资料。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-query-rules-ui-introduction</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-query-rules-ui-introduction</guid>
    <category><![CDATA[基础功能]]></category>
    <category><![CDATA[开发者体验]]></category>
    <dc:creator><![CDATA[Jhon Guzmán]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt565ed0eb407e098d/6a17085d8b73cb363d189fb1/1fb10bd31c509cc9b9bb4f71f49970f140e6c36f-1600x945.png" length="0" type="image/png"/>
    <pubDate>Fri, 07 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何在 AWS Marketplace 上部署 Elasticsearch]]></title>
    <description><![CDATA[通过这份分步指南，您将了解如何在 AWS Marketplace 上使用 Elastic Cloud Service 来设置和运行 Elasticsearch。]]></description>
    <content:encoded><![CDATA[<p>本文将介绍如何使用 Marketplace 产品在 AWS 上部署 Elasticsearch。</p><p>我们将在AWS上使用Elastic Cloud Service，这是Elastic提供的正式托管型Elasticsearch服务，通过AWS的原生基础架构简化了Elastic Stack所有组件的部署和编排。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6c1107b202d48147/6a16f67d92262a04071cc077/f15814051b53b50bec38f9a9f515a1e6dc08a56c-884x440.png" alt="" /><p>如想了解如何在 AWS EC2 上安装并配置 Elasticsearch，请查看<a href="https://www.elastic.co/search-labs/blog/elasticsearch-on-aws-ec2-deployment-guide">这篇博客</a>。
</p><h2>什么是 AWS Marketplace？</h2><p><a href="https://aws.amazon.com/marketplace"><strong>AWS Marketplace 上的 Elastic</strong></a> 提供全托管的搜索和分析体验，AWS 负责基础架构配置、安全和扩展，而开发者则专注于开发搜索应用。这使得团队能够在几分钟内部署具有内置 AWS 集成的企业级 Elasticsearch 集群。</p><h2>何时在 AWS Marketplace 上使用 Elastic？</h2><p>AWS Marketplace 上的 Elastic 最适合拥有现有 AWS 基础架构，希望在无运营开销的情况下部署具有托管服务、内置安全性和无缝 AWS 集成的 Elasticsearch 的组织。</p><h2>如何在AWS Marketplace上设置Elastic Cloud</h2><h3>步骤 1：访问 AWS Marketplace</h3><p>1. 登录 <a href="https://console.aws.com/">AWS</a></p><ul><li><p>在搜索栏中，搜索 AWS Marketplace</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd7b4a64a97929ca0/6a16f67e60084b1ee43c4333/fc9928f79482c2c01e33978c88d390a2bfa2a3bf-1600x340.png" alt="" /><p>2. 在左侧导航面板中，点击<strong>探索产品</strong> (Discover products)，然后搜索 Elasticsearch</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta3efc233beb103ab/6a16f680b0367d681572babd/ca4232271cb13ebfe33de406ecaec085033ec8a0-1454x760.png" alt="" /><p>3. 点击 <strong>Elastic Cloud (Elasticsearch Service)</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd7fe524f5d0cc84e/6a16f682cdacbfabdf7d27c4/e59aa276e55532f2ac3461d0ca983af4d41ad7a6-1600x611.png" alt="" /><h3>第 2 步：订阅服务</h3><p>1. 选择<strong>购买选项</strong>或点击<strong>免费试用</strong> (Try for free)</p><p>2. 查看<strong>定价详细信息</strong>、<strong>条款和条件</strong>以及<strong>购买详情</strong></p><p>3. 点击<strong>订阅</strong> (Subscribe) 按钮。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt027c8adb6bfa2b73/6a16f683a292995855d00de4/1c30d12b6b1061e76771d518011e522285f939f1-1600x290.png" alt="" /><p>4. 现在要设置 Elastic 帐户。按照 AWS 的步骤操作</p><p>a. 点击“启用集成”(Enable integration) 按钮</p><p>b. 点击“登录或创建供应商帐户“(Sign in or create a vendor account) 按钮</p><p>c. 点击“启动模板”(Launch template) 按钮</p><p>d。点击“启动软件”(Launch software) 按钮</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d290ceb4c00c3ca/6a16f685839dfa35f6dcfc9b/879d9f0f01406e1955e1b38a2f6f2192ef040344-852x722.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt20832a47a3a5ce7b/6a16f68767045b0c8b45bf92/fc59be78cf776aa12867f40810598419576cbd39-1600x1143.png" alt="" /><h3>步骤 3。在 Elastic 中配置您的新帐户。</h3><p>1. 创建您的 Elastic 账户。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2a47e1e23a0be123/6a16f68875879e400ffe15c0/5efeaf0737062a55470b17b67651f220e12183f2-986x905.png" alt="" /><p>2. 验证您的电子邮件地址。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltaad961a0644018b5/6a16f68a0811ae0734e9feb6/e0cfaac278614e317ce278935040bfa5a58edd13-853x894.png" alt="" /><p>3. 输入您的姓名和公司信息</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt52f4a502993fa5d4/6a16f68b2b835fb4b9f4afc4/d5658fe66c3b1bcced73e822eae006846f0ddd9e-997x903.png" alt="" /><p>4. 完成简短的 Elastic 调查</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt375dca1904b2f0c1/6a16f68dcdacbf4f8d7d27c8/a3f53c00dadfd22f7d739a920c87d5f387182833-892x805.png" alt="" /><p>5. 选择要托管 Elastic Cloud 的地区。默认情况下，系统会选择您的实际 AWS 地区</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0c23f8b8b63af517/6a16f68fb0367d6a6072bac1/c1dcdf3bf91c305821daaa25a60aa03be6454c1c-1207x1032.png" alt="" /><p>6. 等待 Elastic 部署</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd88a26701fb26db5/6a16f69075879ef0d8fe15cc/50903e57ebea7cc47bdfabf4750b4ba2a7a91148-1370x1266.png" alt="" /><p>7. 您的部署已连接到您的 AWS Marketplace 订阅。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt822e8acb553ffac4/6a16f69275879ed57ffe15d0/3bb731e2d5de5053ccecb77e45dbdbcdaf294dba-1600x1288.png" alt="" /><h2>取消您的订阅</h2><p>取消您的订阅</p><p>1. 进入<a href="https://console.aws.com/">AWS 控制台</a></p><p>在搜索栏中搜索 AWS Marketplace。点击 <strong>AWS Marketplace</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7a19e162d4cb818a/6a16f694839dfa21a6dcfc9f/aeed3d1e67b4cef91934de257a6fd6daa9737a12-1600x554.png" alt="" /><p>2. 点击<strong>Elastic Cloud 订阅</strong> (Elastic Cloud subscription)</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta189d9ff9b518243/6a16f695a292991b60d00de8/04e6cc41850226df223dbe2d1b0e4b45265f6c39-1600x564.png" alt="" /><p>3. 点击<strong>操作</strong> (Actions) 按钮，然后点击<strong>取消订阅</strong> (Cancel subscription)</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt40fc41f3911b8b66/6a16f6971949f70896e7a7a9/e33d334ea6541c637a223de3ebd6209def75a6d3-1600x1039.png" alt="" /><p>4. 确认取消，然后点击<strong>是</strong> (Yes) 和<strong>取消订阅</strong> (Cancel subscription) 按钮。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3eb1839197854313/6a16f699ab7f08469fdb9c31/b73b3187168adc7aefdd46f95be33c1bce3da1e4-1103x698.png" alt="" /><p>5. 页面顶部将会显示确认信息。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt513bdb916a3bff8e/6a16f69a1949f74c5de7a7ad/c5ba66a23d535e866a8b458e5aca82c5f0b93037-1600x639.png" alt="" /><h2>后续步骤</h2><p>通过 7 天免费试用开始您的 Elastic Cloud 之旅，其中包括一次部署和三个项目<a href="https://aws.amazon.com/marketplace/pp/prodview-voru33wi6xs7k"> Elastic Cloud (Elasticsearch Service)</a>。只需登录您的 AWS 帐户，然后点击“查看购买选项”(View Purchase Options)，即可立即在 Elastic<a href="https://aws.amazon.com/marketplace/pp/prodview-voru33wi6xs7k"> Cloud (Elasticsearch Service)</a> 上开始使用 Elastic 的 Search AI Platform。试用版提供对搜索、安全性和可观测性解决方案的全面访问权限，无需支付任何基础架构管理开销费用。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/aws-elasticsearch-service-set-up</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/aws-elasticsearch-service-set-up</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Eduard Martin]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcbf8a044704b64d5/6a16f69ccf4f25d524b2cf3f/a80776d2ef85db26f850d932339fac2d26b90278-1086x620.png" length="0" type="image/png"/>
    <pubDate>Fri, 03 Oct 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch 分片和副本：实用指南]]></title>
    <description><![CDATA[掌握 Elasticsearch 分片和副本的概念，并学习如何优化它们。]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch 在 Lucene 的基础上建立了一个分布式系统，解决了可扩展性和容错问题，从而增强了 Lucene 的功能。它还提供基于 JSON 的 REST 应用程序接口，使与其他系统的互操作性变得非常简单。</p><p>Elasticsearch 等分布式系统可能非常复杂，影响其性能和稳定性的因素很多。<strong>分片</strong>是 Elasticsearch 中最基本的概念之一，了解分片的工作原理将使您能够有效地管理 Elasticsearch 集群。</p><p>本文将解释什么是主分片和副本分片，它们对 Elasticsearch 集群的影响，以及有哪些工具可以调整它们以适应不同的需求。</p><h2>了解碎片</h2><p>Elasticsearch 索引中的数据可能会大量增长。为了便于管理，每条数据都保存在一个索引中，而索引是将一个索引分割成若干<strong>碎片</strong>。每个 Elasticsearch 分区都是一个 Apache Lucene 索引，每个单独的 Lucene 索引都包含 Elasticsearch 索引中文档的一个子集。以这种方式拆分索引可以控制资源使用量。Apache Lucene 索引的上限为 2,147,483,519 (2³¹ - 129) 个文档。</p><p>有时，出于重新平衡的目的，需要在节点间移动指数。由于这一过程需要大量时间和资源，因此索引不应过大，这有助于保持可控的恢复时间。此外，由于索引是由需要不断合并在一起的 Lucene 段组成的，因此段不能太大，这一点很重要。由于这些原因，Elasticsearch 将索引数据分割成更易于管理的小块（称为<strong>主分片</strong>），这些分片可以更方便地分布在多台计算机上。<strong>复制</strong>分区只是相应主分区的一个精确副本，我们将在本文稍后部分介绍它们的功能。</p><p>拥有适当数量的分片对性能非常重要。因此，提前制定计划是明智之举。当查询在不同分片上并行运行时，其执行速度要快于由单个分片组成的索引，但前提是每个分片位于不同的节点上，且集群中有足够多的节点。但与此同时，分片也会消耗内存和磁盘空间，包括索引数据和集群元数据。分片过多（也称为过度分片）会降低查询、索引请求和管理操作的速度，因此保持适当的平衡至关重要。</p><p>主分区的数量是在<strong>为特定索引实例</strong>创建索引时定义的。如果以后需要不同数量的主分片，可以使用<strong> 调整大小</strong><a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-split">API</a> --拆分（更多的主分片）、<a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-shrink">收缩</a>（更少的主分片）或<a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-clone">克隆</a>（相同数量的主分片，并对副本进行新的设置）。创建索引时，可以<strong>将主</strong>分片和副本分片的数量设置为索引的设置：</p>PUT /sensor
{
   "settings" : {
       "index" : {
           "number_of_shards" : 6,
           "number_of_replicas" : 2
       }
   }
}<p>(如果没有指定分片或副本的数量，从 Elasticsearch 7.0 开始，两者的默认值都是 1）。理想的分片数量应根据索引中的数据量来确定。一般来说，<a href="https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/size-shards">一个最佳分区应容纳 10-50GB 的数据</a>，每个分区的文件数少于 2 亿。例如，如果您预计一天内会积累约 300GB 的应用程序日志，那么在该索引中设置约 10 个分片是合理的，前提是您有足够多的节点来托管这些分片。</p><p>碎片在其生命周期中会经历多种状态，包括</p><ul><li><p><strong>初始化：</strong>使用分片前的初始状态。</p></li><li><p><strong>已启动：</strong>分片处于激活状态，可以接收请求。</p></li><li><p><strong>搬迁：</strong>当分片正在被移动到不同节点时出现的一种状态。这在某些情况下可能是必要的，例如，当它们所在的节点快用完磁盘空间时。</p></li><li><p><strong>未分配：</strong>未能分配的分区的状态。发生这种情况时会给出原因，例如，托管分片的节点已不在集群中<em>（NODE_LEFT）</em>或由于恢复到一个已关闭的索引<em>中（EXISTING_INDEX_RESTORED）。</em></p></li></ul><p>要查看所有分片、它们的状态和其他元数据，可以使用以下请求：</p>GET _cat/shards<p>要查看特定索引的分片，可以在 URL 中添加索引名称，例如传感器：</p>GET _cat/shards/sensor<p>该命令会产生输出结果，如下面的示例。默认情况下，显示的列包括索引名称、名称（即编号）、是主分片还是副本、状态、文件数量、磁盘大小以及分片所在节点的 IP 地址和节点 ID。</p>sensor 5 p STARTED    0  283b 127.0.0.1 ziap
sensor 5 r UNASSIGNED                  
sensor 2 p STARTED    1 3.7kb 127.0.0.1 ziap
sensor 2 r UNASSIGNED                  
sensor 3 p STARTED    3 7.2kb 127.0.0.1 ziap
sensor 3 r UNASSIGNED                  
sensor 1 p STARTED    1 3.7kb 127.0.0.1 ziap
sensor 1 r UNASSIGNED                  
sensor 4 p STARTED    2 3.8kb 127.0.0.1 ziap
sensor 4 r UNASSIGNED                  
sensor 0 p STARTED    0  283b 127.0.0.1 ziap
sensor 0 r UNASSIGNED<h2>了解副本</h2><p>每个分区只包含一份数据副本，而索引则可以包含多个分区副本。因此有两种分片，<strong>即主分片</strong>和副本或<strong>复制</strong> 分片。主分片的每个副本总是位于不同的节点上，这就确保了在节点发生故障时数据的高可用性。除了冗余及其在防止数据丢失和宕机方面的作用外，副本还可以帮助提高搜索性能，因为它允许查询与主分片并行处理，因此速度更快。</p><p>主分片和副本分片的行为方式存在一些重要差异。虽然两者都能处理查询、索引请求（即向索引添加数据）必须先经过主分片，然后才能复制到副本分片。如上所述，如果主分片不可用--例如，由于节点断开或硬件故障--副本就会被提升以接替其角色。</p><p>虽然复制可以在节点发生故障时提供帮助，但重要的是不要有太多的复制，因为它们会在编制索引时消耗内存、磁盘空间和计算能力。主分片和副本之间的另一个区别是，虽然主分片的数量在索引创建后无法更改，但副本的数量可以通过更新索引设置随时动态更改。</p><p>复制的另一个考虑因素是可用节点的数量。副本总是放在与主分片不同的节点上，因为如果节点发生故障，同一节点上的两个相同数据副本将无法提供保护。因此，一个系统要支持<em>n 个</em>副本，集群中至少需要有<em>n + 1 个</em>节点。例如，如果集群中有两个节点，而索引配置了六个副本，则只会分配一个副本。另一方面，拥有七个节点的系统完全可以处理一个主分片和六个副本。</p><h2>优化分片和副本</h2><p>即使在创建了主分片和副本分片平衡得当的索引后，也需要对这些分片进行监控，因为索引的动态会随着时间的推移而发生变化。例如，在处理时间序列数据时，最新数据的指数通常比旧数据的指数更活跃。如果不对这些指数进行调整，它们将消耗相同数量的资源，尽管它们的需求非常不同。</p><p>翻转索引 API 可用于区分新旧索引。可以对其进行设置，一旦达到某个阈值（磁盘上索引的大小、文档数量或年限），它就会自动创建新索引。该 API 对于控制分片大小也很有用。由于索引创建后无法轻易更改分片数量，因此如果不满足翻转条件，分片将继续积累数据。对于只需不经常访问的旧索引，缩小和强制合并索引是减少其内存和磁盘占用的两种不同方法。前者减少了索引中分片的数量，后者则减少了 Lucene 片段的数量，并释放了已删除文档的空间。</p><h2>作为 Elasticsearch 基础的主分片和副本分片</h2><p>Elasticsearch 作为适用于海量数据的分布式存储、搜索和分析平台，已经建立了良好的声誉。然而，在如此大规模的运作中，挑战将不可避免地出现。这就是为什么了解主分片和副本分片如何工作对 Elasticsearch 如此重要和基础的原因，因为这有助于优化平台的可靠性和性能。</p><p>了解它们如何工作以及如何优化它们，对于实现更强大、更高性能的 Elasticsearch 集群至关重要。如果您经常遇到查询响应迟缓或中断的情况，这些知识可能是克服这些障碍的关键。</p><p>请关注 Elasticsearch 的官方文档，了解有关<a href="https://www.elastic.co/docs/deploy-manage/distributed-architecture/clusters-nodes-shards">群集、节点和分片</a>、<a href="https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/size-shards">如何确定分片大小</a>、<a href="https://www.elastic.co/docs/deploy-manage/distributed-architecture/shard-allocation-relocation-recovery">分片分配和恢复的</a>更多信息。</p><p>本主题还可作为入门课程在<a href="https://youtu.be/sAySPSyL2qE">Elastic Community YouTube 频道</a>上观看。</p><p>最后但并非最不重要的一点：如果你不想担心节点、分片或副本，可以试试<a href="https://www.elastic.co/docs/deploy-manage/deploy/elastic-cloud/serverless">Elastic Cloud Serverless</a>。该 Elastic 云产品由 Elastic 全面管理，并可根据您的工作负载自动扩展。免费试用可以帮助您熟悉无服务器方法的其他优势。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-shards-and-replicas-guide</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-shards-and-replicas-guide</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Piotr Przybyl]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt71dd92d939d383a2/6a17e9d53e03d769c44f2cc7/7775c44f01f2516c4ff4cce6d6bbe9e7b2c38908-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 14 Aug 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[揭开独特模式的面纱：Elasticsearch 中重要术语聚合指南]]></title>
    <description><![CDATA[了解如何使用重要术语聚合来发现数据中的洞察力。]]></description>
    <content:encoded><![CDATA[<p>在 Elasticsearch 中，<a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-significantterms-aggregation">重要术语聚合</a>超出了<a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-terms-aggregation">最常见术语的</a>范围，可在数据集中找到统计上不寻常的值。这使我们能够发现有价值的见解和非显而易见的模式。一个重要的术语集合提供了两个有用参数的响应：</p><ul><li><p><strong>bg_count（背景计数）： </strong>在父数据集中找到的文件数</p></li><li><p><strong>doc_count：</strong>结果数据集中找到的文件数</p></li></ul><p>例如，在手机销售数据集中，我们可以像这样查找 iPhone 16 销售的重要术语：</p>GET phone_sales_analysis/_search
{
 "size": 0,
 "query": {
   "term": {
     "phone_model": {
       "value": "iPhone 16"
     }
   }
 },
 "aggs": {
   "significant_cities": {
     "significant_terms": {
       "field": "city_region",
       "size": 1
     }
   }
 }
}<p>然后，答复给了我们：</p>{
 "aggregations": {
   "significant_cities": {
     "doc_count": 122,
     "bg_count": 424,
     "buckets": [
       {
         "key": "Houston",
         "doc_count": 12,
         "score": 0.1946481360617346,
         "bg_count": 14
       }

     ]
   }
 }
}<p>在整个数据集中，休斯顿既不是排名前十的城市，也不是 iPhone 16 的热门城市。不过，重要术语汇总显示，与其他数据相比，<em><strong> 该城市购买 iPhone 16 的比例过高</strong></em>。让我们深入了解这些数字：</p><ul><li><p><strong>在最高层：</strong></p><ul><li><p><strong>doc_count：122 - </strong>查询总共匹配了 122 份文件</p></li><li><p><strong>bg_count：424 - </strong>背景集（所有销售文件）包含 424 份文件</p></li></ul></li><li><p><strong>在休斯顿的水桶里：</strong></p><ul><li><p><strong>doc_count：12 - </strong>休斯顿出现在 122 条查询结果中的 12 条中</p></li><li><p><strong>bg_count：14 - </strong>在背景数据集的 424 份文件中，休斯顿出现在 14 份文件中</p></li></ul></li></ul><p>这告诉我们，在 424 次总购物中，只有 14 次发生在休斯顿，占总购物次数的 3.3% 。然而，如果我们只看 iPhone 16 的销售情况，就会发现 122 件中有 12 件发生在休斯顿，比整个数据集多 3 倍，即 9.8% ；这是非常重要的！</p><p>以下是可视化效果图：每个城市/地区的销售总额。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blted1af6606708267e/6a17f5d614800960e1b488d5/f31335b0b7793650025f941820f238dd35bfb09f-1486x1066.png" alt="" /><p>我们可以看到，休斯顿有 14 笔销售，是数据集中销售额第 14 高的城市。</p><p>现在，如果我们只对 iPhone 16 的销售情况进行筛选，休斯顿就有 12 台，成为该机型销售量第二大的城市：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd2882d5e87d02406/6a17f5d71d1b83851e93e5b2/6516040db77e6c62af5541a74c723b18008ad3c6-1472x1038.png" alt="" /><h2>了解重要术语汇总</h2><p>根据 Elastic 文档，<a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-significantterms-aggregation">重要的术语是聚合</a>：</p><p><em>"（查找）在前景集和背景集之间流行度发生显著变化的术语"。</em></p><p>这意味着它使用统计指标，将数据子集（前景集）中某个术语的频率与父数据集（背景集）中同一术语的频率进行比较。这样，评分反映的是统计意义，而不是术语在数据中出现的频率。</p><p>重要术语聚合与普通术语聚合的主要区别在于</p><ul><li><p>重要术语对数据的子集进行比较，而术语聚合只对查询产生的数据集起作用。</p></li><li><p>术语聚合的结果是数据集中最常见的术语，而重要术语的结果则忽略了常见术语，以找出数据集的独特之处。</p></li><li><p>重要术语对性能的影响更大，因为它需要从磁盘而不是内存中获取数据，就像术语聚合所做的那样。</p></li></ul><h2>实际应用（消费者行为分析）</h2><h3>为分析准备数据</h3><p>为了进行分析，我们生成了一个合成的手机销售数据集，其中包括价格、手机规格、购买者的人口统计数据和反馈信息。我们还根据用户的反馈生成了嵌入信息，以便日后进行语义查询。我们使用了 Elasticsearch 上开箱即用的<a href="https://huggingface.co/intfloat/multilingual-e5-small">多语言 e5 小型模型</a>。</p><p></p><p>要在 Elasticsearch 上使用此数据集：</p><ol><li><p>使用 Kibana<a href="https://www.elastic.co/docs/manage-data/ingest/upload-data-files"> 上传数据文件</a> 功能上传 CSV 文件（可从<a href="https://github.com/Alex1795/significant_terms_blog_dataset/blob/main/phone_sales_analysis_dataset.csv"> 此处 下载）。</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/chat-with-pdf-elastic-playground#upload-pdfs-to-kibana">如本博客</a>所示，设置一个名为 "嵌入 "的语义字段，使用 <code>multilingual-e5-small model</code></p></li><li><p>使用字段类型默认值完成导入（除<code>purchase_date</code> 和<code>user_feedback)</code> 外，每个字段都使用关键字。请确保添加索引名称<code>phone_sales_analysis</code> ，以便能够按原样运行此处提供的查询。</p></li></ol><p>这项分析的主要重点是发现<em><strong>"iPhone 16 购买者与其他人群的不同之处</strong></em>"，并为营销目的对购买者进行细分。 </p><p>这是数据集中的一份样本文件：</p>{
         "customer_type": "Returning",
         "user_feedback": "I have to say, quality is great for the price. The battery life is really good.",
         "upgrade_frequency": "2 years",
         "storage_capacity": "256GB",
         "occupation": "Technology &amp; Data",
         "color": "Phantom Black",
         "gender": "Male",
         "price_paid": 899,
         "previous_brand_loyalty": "Mixed",
         "location_type": "Urban",
         "phone_model": "Samsung Galaxy S24",
         "city_region": "San Francisco Bay Area",
         "@timestamp": "2024-03-15T00:00:00.000-05:00",
         "income_bracket": "75000-100000",
         "purchase_channel": "Online",
         "feedback_sentiment": "positive",
         "education_level": "Bachelor",
         "embedding": "I have to say, quality is great for the price. The battery life is really good.",
         "customer_id": "C001",
         "purchase_date": "2024-03-15",
         "age": 34,
         "trade_in_model": "iPhone 13"
}<h3>了解人口模式</h3><p>在此，我们将对一般人群进行分析，并将其与 iPhone 16 用户重要术语汇总的有趣发现进行比较。</p><h4>正常模式</h4><p>为了了解正常的购买模式，我们可以汇总不同领域所有文档的数据。为简单起见，我们将重点探讨购买手机的人的职业。我们可以通过向 Elasticsearch 提出请求来实现这一点。</p>GET phone_sales_analysis/_search
{
 "aggs": {
   "occupation_distribution": {
     "terms": {
       "size": 5,
       "field": "occupation"
     }
   }
 },
 "size": 0
}<p>这告诉我们，数据集中的主要职业（按记录数计）是</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltec11e3ae5b9cb3c0/6a17f5d9505ac3f28aad8cba/99136ddddd7abad5d74481158a04501b6915441b-1518x480.png" alt="" /><h4>iPhone 16 用户的使用模式</h4><p>为了了解购买了 iPhone 16 的人有什么不同，让我们在同一字段上运行术语聚合，并在查询中使用过滤器找到这些人，就像这样：</p>GET phone_sales_analysis/_search
{
  "query": {
    "term": {
      "phone_model": "iPhone 16"
    }
  },
  "aggs": {
    "occupation_distribution": {
      "terms": {
        "size": 5,
        "field": "occupation"
      }
    }
  },
  "size": 0
}<p>因此，iPhone 16 用户的主要职业是</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt26a5415e2f30d164/6a17f5da445de9637a4d028a/36ce86475beb03810c6ad81d7c776d1eec736654-1500x484.png" alt="" /><p>我们可以看到，iPhone 16 用户的职业模式与其他型号手机的用户不同。让我们使用 Kibana 来轻松实现结果的可视化：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4e0ddb09fe58e454/6a17f5dce317912cfa2d596b/b70ab05bc962a274e1617b6caf20575c489a62d8-1448x1128.png" alt="" /><p></p><p>在这张图表中，我们可以看到 iPhone 16 的趋势与整个人群的趋势不同。</p><p>我们可以跳过整个分析，通过一个重要项的汇总，来看看 iPhone 16 用户与普通用户的不同之处：</p>GET phone_sales_analysis/_search
{
  "query": {
    "term": {
      "phone_model": "iPhone 16"
    }
  },
  "aggs": {
    "occupation_distribution": {
      "significant_terms": {
        "size": 5,
        "field": "occupation"
      }
    }
  },
  "size": 0
}<p>简而言之，我们得到了这样的答复：</p><p>iPhone 16 的职业值</p><p>文件数量</p><p>bg_count</p><p>职业分布（最高级别）</p><p>122</p><p>424</p><p>医疗&amp; 保健桶</p><p>45</p><p>57</p><p>这些回复清楚地表明，iPhone 16 用户有一个不常见的（读作 "重要！"）问题。与普通人相比，医疗&amp; 保健领域的人数更多。让我们看看回复中的数字意味着什么：</p><ul><li><p><strong>在最高层：</strong></p><ul><li><p><strong>doc_count：122 - </strong>查询总共匹配了 122 份文件</p></li><li><p><strong>bg_count：424 - </strong>背景集（所有销售文件）包含 424 份文件</p></li></ul></li><li><p><strong>在医疗&amp; 保健桶中：</strong></p><ul><li><p><strong>doc_count：45 - </strong>"医疗&amp; 保健" 在 122 条查询结果中出现了 45 条</p></li><li><p><strong>bg_count：57 - </strong>"医疗&amp; 保健" 在背景数据集中的全部 424 份文件中出现 57 份</p></li></ul></li></ul><p>在 424 位买家中，有 57 位在医疗&amp; 保健领域工作，即 13.44% 。但是，当我们查看 iPhone 16 的购买者时，122 位购买者中有 45 位从事医疗&amp; ，即 36.88% 。这意味着在 iPhone 16 用户中，从事医疗&amp; 保健工作的可能性要高出一倍！</p><p>我们可以将同样的分析应用于其他领域（年龄、地点、收入阶层等），从而发现更多有关 iPhone 16 用户独特之处的信息。 </p><h3>消费者细分</h3><p>我们可以利用重要术语聚合来提取产品、类别和客户群之间的关系洞察。为此，我们为感兴趣的类别建立一个父聚合。我们还使用了重要术语和普通术语子分类，以发现对该类别的有趣见解，并将其与该职业中大多数人使用的术语进行比较。</p><p>例如，让我们看看某些工作领域的人喜欢什么：</p><ol><li><p>为了更清楚地进行分析，我们将搜索范围限制在 3 个工作领域：["行政&amp; 支持","技术&amp; 数据","医疗&amp; 保健"]</p></li><li><p>在汇总方面，我们首先按职业进行术语汇总</p></li><li><p>增加一个子分类：按手机型号分类--查找在各个领域工作的用户正在购买哪些手机型号</p></li><li><p>添加第二个子分类：按手机型号分类的重要术语，以找出每个工作领域中的特殊型号</p></li></ol>GET phone_sales_analysis/_search
{
 "query": {
   "terms": {
     "occupation": [
       "Administrative &amp; Support",
       "Technology &amp; Data",
       "Medical &amp; Healthcare"
     ]
   }
 },
 "aggs": {
   "occupations": {
     "terms": {
       "size": 15,
       "field": "occupation"
     },
     "aggs": {
       "general_models": {
         "terms": {
           "field": "phone_model"
         }
       },
       "significant_models": {
         "significant_terms": {
           "field": "phone_model"
         }
       }
     }
   }
 },
 "size": 0
}<p>让我们来分析一下汇总结果：</p><p><strong>职业</strong>行政&amp; 支持</p><p><strong>术语汇总</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3a325cde37b40504/6a17f5dd4b055dbb68432372/a4ad519c9013867a3f4cee032160eadd8a47804a-1506x398.png" alt="" /><p><strong>重要术语汇总</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltef514b8bbb7f429c/6a17f5df3e9e456488ba1612/e5604fa8036667bdfe733576a5e7c6153760dd3a-306x220.png" alt="" /><p>从该表中我们可以推断出，该职业的趋势与整个人口的趋势之间没有显著差异</p><p><strong>职业</strong>：技术&amp; 数据</p><p><strong>术语汇总</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt94189446d8f3100e/6a17f5e142022983b029f76e/13b09039bb7d183276451007d2d69dc190b1d3c0-1508x836.png" alt="" /><p></p><p><strong>重要术语汇总</strong></p><p>文件总数424</p><p>该职业的文件：71</p><p>手机型号</p><p>doc_count （本职业中的本模型）</p><p>bg_count （所有文件中都有此模型）</p><p>% 在所有文件中</p><p>% 从事这一职业</p><p>谷歌 Pixel 8</p><p>12</p><p>220</p><p>5.19%</p><p>16.90%</p><p>OnePlus 11</p><p>9</p><p>14</p><p>3.30%</p><p>12.68%</p><p>OnePlus 12 Pro</p><p>3</p><p>3</p><p>0.71%</p><p>4.23%</p><p>谷歌 Pixel 8 Pro</p><p>9</p><p>21</p><p>4.95%</p><p>12.68%</p><p>无手机 2</p><p>5</p><p>8</p><p>1.89%</p><p>7.04%</p><p>三星 Galaxy Z Fold5</p><p>4</p><p>6</p><p>1.42%</p><p>5.63%</p><p>OnePlus 12</p><p>8</p><p>20</p><p>4.72%</p><p>11.27%</p><p><strong>职业</strong>：医疗&amp; 保健</p><p><strong>术语汇总</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt270b0a861a16488a/6a17f5e23e03d76a934f2df0/b008e996742fc0bb48dc6bacff17cfbc56cf0d73-1492x398.png" alt="" /><p><strong>重要术语汇总</strong></p><p>文件总数424</p><p>该职业的文件：57</p><p>手机型号</p><p>doc_count （本职业中的本模型）</p><p>bg_count （所有文件中都有此模型）</p><p>% 在所有文件中</p><p>% 从事这一职业</p><p>iPhone 16</p><p>45</p><p>122</p><p>28.77%</p><p>78.95%</p><p>iPhone 15 Pro Max</p><p>3</p><p>13</p><p>3.07%</p><p>5.26%</p><p>iPhone 15</p><p>7</p><p>40</p><p>9.43%</p><p>12.28%</p><p>让我们看看这些数据告诉了我们什么故事：</p><ul><li><p>医疗&amp; 医疗保健专业人士更喜欢 iPhone 16，而且普遍倾向于使用苹果手机。</p></li><li><p>技术&amp; 数据专业人士更喜欢高端安卓手机，但不一定使用三星品牌。在这一类别中，iPhone 也有相当大的发展趋势。</p></li><li><p>行政管理&amp; 支持专业人员更喜欢三星和谷歌手机，但没有形成强烈而独特的趋势。</p></li></ul><h3>重要术语汇总和混合搜索</h3><p>混合搜索结合了文本搜索和语义结果，可提供更好的搜索体验。在这种情况下，一个重要的术语聚合可以通过回答问题来深入了解上下文感知搜索的结果：<strong>与所有文档相比，这个数据集有什么特别之处？</strong>为了展示这一特点，让我们看看当用户谈论良好性能时，哪些模型的代表性过高： </p><ul><li><p>让我们建立一个语义查询，通过字段嵌入找到最接近输入 "性能良好 "的用户反馈</p></li><li><p>我们还将在文本字段 user_feedback 中使用相同的术语进行文本搜索</p></li><li><p>我们还将添加一个重要术语查询，以找到在这些结果中出现频率高于完整数据集的手机型号
</p></li></ul>GET phone_sales_analysis/_search
{
 "retriever": {
   "rrf": {
     "retrievers": [
       {
         "standard": {
           "query": {
             "bool": {
               "must": [
                 {
                   "match": {
                     "user_feedback": {
                       "query": "good performance",
                       "operator": "and"
                     }
                   }
                 }
               ]
             }
           }
         }
       },
       {
         "standard": {
           "query": {
             "semantic": {
               "field": "embedding",
               "query": "good performance"
             }
           }
         }
       }
     ],
    "rank_window_size": 20
   }
 },
 "aggs": {
   "Models": {
     "significant_terms": {
       "field": "phone_model"
     }
   }
 }
}<p>让我们来看一个匹配文件的例子：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1c3dc221896c8832/6a17f5e4445de91eac4d028e/4cb488097a382f0c28c21540db4f593d23633473-1600x162.png" alt="" /><p>这就是我们得到的答复：</p>{
  "took": 388,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 20,
      "relation": "eq"
    },
    "max_score": 0.016393442,
    "hits": [...]
  },
  "aggregations": {
    "Models": {
      "doc_count": 20,
      "bg_count": 424,
      "buckets": [
        {
          "key": "iPhone 15",
          "doc_count": 5,
          "score": 0.4125,
          "bg_count": 40
        }
      ]
    }
  }
}<p></p><p>这告诉我们，虽然 iPhone 15 在总共 424 篇文档中出现了 40 次（占文档总数的 9.4% ），但在符合语义搜索 "良好表现 "的 20 篇文档（占文档总数的 25% ）中却能找到 5 次。因此，我们可以得出这样的结论：在谈论良好性能时，发现 iPhone 15 的可能性是偶然发现的 2.7 倍。</p><h2>结论</h2><p>重要术语聚合可以通过将数据集与全局文档进行比较，发现数据集的独特细节。这可以揭示数据中意想不到的关系，而不仅仅是出现次数的计算。例如，我们可以在各种使用案例中应用重要术语，从而实现非常有趣的功能：</p><ul><li><p>在<a href="https://www.elastic.co/blog/significant-terms-aggregation#credit"> 侦查 欺诈行为时找出模式 --识别被盗信用卡的常见交易。</a></p></li><li><p>从用户评论中洞察品牌质量--发现差评过多的品牌。</p></li><li><p><a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-significantterms-aggregation#_use_on_free_text_fields">发现 </a>分类错误的文档--发现属于某个类别（术语过滤器）但在描述中使用了该类别不常用词的文档（重要术语汇总）。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/significant-terms-aggregation-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/significant-terms-aggregation-elasticsearch</guid>
    <category><![CDATA[基础功能]]></category>
    <category><![CDATA[查询 DSL]]></category>
    <dc:creator><![CDATA[Alexander Dávila]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta4d95b6134c2a110/6a17f5e6505ac3fd3cad8cbf/13adbc901837835bb56abf15e377127b017cfac8-1536x1024.png" length="0" type="image/png"/>
    <pubDate>Mon, 07 Jul 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何在 GCP GKE Autopilot 上部署 Elasticsearch。]]></title>
    <description><![CDATA[了解如何通过 GKE Autopilot 与 ECK 在 GCP 上部署 Elasticsearch 集群，实现部分托管的 Elasticsearch 设置配置。]]></description>
    <content:encoded><![CDATA[<p>在本文中，我们将学习如何使用 Autopilot 在 Google Cloud Kubernetes (GKE) 上部署 Elasticsearch。</p><p>对于 Elasticsearch，我们将使用 <a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s">Elastic Cloud on Kubernetes</a> (ECK)，这是正式的 Elasticsearch Kubernetes 运维工具，简化了所有 Elastic Stack 组件的 Kubernetes 部署协调。</p><p>要了解更多关于如何在不同 GCP 基础架构上部署 Elasticsearch 集群的信息，您可以阅读我们关于 <a href="https://www.elastic.co/search-labs/blog/elasticsearch-gpc-google-compute-engine">Google Cloud Compute</a> 和 <a href="https://www.elastic.co/search-labs/blog/deploy-elastic-gcp-marketplace">Google Cloud Marketplace</a> 的入门文章。</p><h2>Elasticsearch 部署步骤</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt61969357430c94f8/6a17f6d125daab024708a3ae/56b54d718dcff9af9050873c41fdf738074851da-1428x582.png" alt="Elasticsearch ECK 部署工作。" /><h3>什么是 GKE Autopilot？</h3><p><a href="https://cloud.google.com/kubernetes-engine/docs/concepts/autopilot-overview?hl=es-419"><strong>Google Kubernetes Engine (GKE) Autopilot</strong></a> 提供完全托管的 Kubernetes 体验，由 Google 负责集群配置、节点管理、安全与扩展，而开发人员只需专注于应用程序部署，使团队能够凭借内置的最佳实践在几分钟内完成从代码到生产环境的转化。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt91e10f44290aeac4/6a17f6d33e9e451f67ba1638/bbf6de63fa0a199326352f521cb22654818799f6-1600x958.png" alt="GKE Autopilot 容器。" /><h2>何时在 Google Cloud 中使用 ECK？</h2><p>Elastic Cloud on Kubernetes (ECK) 最适合那些拥有现有 Kubernetes 基础架构并希望部署具有高级功能（如专用节点角色、高可用性和自动化）的 Elasticsearch 的组织。</p><h2>如何在 Google Cloud 中设置 ECK？</h2><p>1. 登录 <a href="https://console.cloud.google.com">Google Cloud Console </a>。</p><p>2. 在<strong>右上角</strong>点击“<strong>Cloud Shell</strong>”按钮进入控制台，并从那里部署 GKE 集群。或者，您也可以使用 <a href="https://cloud.google.com/cli">gcloud CLI</a>。</p><p><em><strong>操作过程中记得将项目 ID 替换为您自己的项目 ID。</strong></em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc69e47bf97e4be31/6a17f6d5505ac3bbe6ad8cdb/999b03861d4fe44f360ab4c7e2616e1dc10cf182-1558x1248.png" alt="如何在 GCP GKE Autopilot 上设置 Elasticsearch。" /><p>3. 启用 <a href="https://console.cloud.google.com/flows/enableapi?apiid=container.googleapis.com">Google Kubernetes Engine API</a>。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1a4d95465f563ece/6a17f6d7e8fbced6393a1af3/03827d3dc0e987c019e7747d33e7c01920047beb-911x246.png" alt="启用 Google Kubernetes Engine API。" /><p>点击<em><strong>下一步</strong></em>。</p><p>现在，搜索 Kubernetes Engine API 时，应该显示 Kubernetes Engine API 已启用。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6ba41b82d995347f/6a17f6d8505ac3036bad8cdf/d5cd46f0333086bcb31b80cf9c08a469b449ec0f-640x250.png" alt="Kubernetes Engine API。" /><p>4. 在 Cloud Shell 中创建一个 Autopilot 集群。我们将其命名为 autopilot-cluster-1，并请将 autopilot-test 替换为您的项目 ID。</p>gcloud beta container --project "autopilot-test-457216" clusters create-auto "autopilot-cluster-1" --region "us-central1" --release-channel "regular" --tier "standard" --enable-ip-access --no-enable-google-cloud-access --network "projects/autopilot-test-457216/global/networks/default" --subnetwork "projects/autopilot-test-457216/regions/us-central1/subnetworks/default" --cluster-ipv4-cidr "/17" --binauthz-evaluation-mode=DISABLED<p>5. 等待集群就绪。创建过程大约需要 10 分钟。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4e821e77d5db5d71/6a17f6da148009381cb48900/81fbc45ba56d0f16ba42724cb8ae45e60b327dbc-1581x258.png" alt="Autopilot 集群设置的图像。" /><p>正确设置集群后，将显示一条确认消息。</p><p>6. 配置 kubectl 命令行访问。</p>gcloud container clusters get-credentials autopilot-cluster-1 --region us-central1 --project autopilot-test-457216<p>您应该看到：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf1c828e5059c8ec8/6a17f6dcabe0f215e5dfebb7/b0beba1ee00ce9029f586ee32693fc2aa58c7f65-3442x142.png" alt="" /><p><em>已为 autopilot-cluster-1 生成 kubeconfig 条目。</em></p><p>7. 安装 <a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s">Elastic Cloud on Kubernetes</a>(ECK) 运维工具。</p># Install ECK Custom Resource Definitions
kubectl create -f https://download.elastic.co/downloads/eck/2.16.1/crds.yaml

# Install the ECK operator
kubectl apply -f https://download.elastic.co/downloads/eck/2.16.1/operator.yaml<p>8. 让我们创建一个具有默认值的单节点 Elasticsearch 实例。</p><p>如果您想查看不同设置的配方，可以访问<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s/recipes">此链接</a>。</p><p>请注意，如果未指定 <code>storageClass</code>，ECK 将使用默认设置，GKE 的默认设置为 <code>standard-rwo</code>，该配置使用 <a href="https://cloud.google.com/kubernetes-engine/docs/how-to/persistent-volumes/gce-pd-csi-driver?cloudshell=true">Compute Engine 持久化磁盘 CSI 驱动</a>）并创建 1GB 的卷。</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: elasticsearch.k8s.elastic.co/v1
kind: Elasticsearch
metadata:
  name: quickstart
spec:
  version: 9.0.0
  nodeSets:
  - name: default
    count: 1
    config:
      node.store.allow_mmap: false
EOF<p>我们禁用<code>nmap</code>，是因为默认 GKE 机器的 <code>vm.max_map_count</code> 值过低。不建议在生产环境中禁用它，但建议增加 <code>vm.max_map_count</code> 的值。您可以<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s/virtual-memory">在这里</a>阅读更多关于如何做到这一点的信息。</p><p>9. 我们还要部署一个 Kibana 单节点集群。对于 Kibana，我们将添加一个 LoadBalancer，它将为我们提供一个外部 IP，我们可以使用该 IP 从我们的设备访问 Kibana。</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: kibana.k8s.elastic.co/v1
kind: Kibana
metadata:
  name: quickstart
spec:
  version: 9.0.0
  http:
    service:
      metadata:
        annotations:
          cloud.google.com/l4-rbs: "enabled"
      spec:
        type: LoadBalancer
  count: 1
  elasticsearchRef:
    name: quickstart
EOF<p>请注意注释： </p><p><code>cloud.google.com/l4-rbs: "enabled"</code></p><p><em><strong>该注释非常重要，因为它指示 Autopilot 提供一个面向公众的 LoadBalancer。如果未设置，LoadBalancer 将为内部类型。</strong></em></p><p>10. 检查您的 pod 是否正在运行</p>kubectl get pods<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt627dd8f482340bc2/6a17f6de3e03d779a44f2e16/99da1270581a137683770efdb9c6e1577ec9fc01-3150x442.png" alt="" /><p>11. 您还可以使用 <code>run kubectl get elasticsearch</code> 和 <code>kubectl get kibana</code> 来获取更具体的统计信息，例如 Elasticsearch 版本、节点和健康状况。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt32d179d1c4f785b3/6a17f6e04b055db05a432392/86234f307970fd5f78b8acd41496e8cc89ff82d3-3414x326.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt75b60dc0e4ad9636/6a17f6e296142a27eaeb1ca6/29160286ccc88928734c8ea11b1923db8e85d49d-3142x318.png" alt="" /><p>12. 获取您的服务。</p>kubectl get svc<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3806ad91791373a2/6a17f6e46df731bc800a10ab/ed1a07314b84a99b4aa1fec3db4b9badeb9587ee-3446x610.png" alt="" /><p>这将显示 Kibana 在 EXTERNAL-IP 下的外部 URL。可能需要几分钟时间来配置 LoadBalancer。<em><strong>复制 EXTERNAL-IP 的值。</strong></em></p><p>13. 获取“elastic”用户的 Elasticsearch 密码：</p>kubectl get secret quickstart-es-elastic-user -o=jsonpath='{.data.elastic}' | base64 --decode<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ab53ce5a685491b/6a17f6e63e03d74fcc4f2e1a/ab5054219216ebc15fc0d96e27605aaf13b720c6-3448x210.png" alt="" /><p>14. 通过浏览器<strong>访问 Kibana</strong>：</p><ul><li><p>URL: https://&lt;EXTERNAL_IP&gt;:5601</p></li><li><p>用户名：elastic</p></li><li><p>密码：28Pao50lr2GpyguX470L2uj5（来自上一步）</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd86d22f797132c21/6a17f6e7e8fbce62b43a1afb/47cbe88dc14db64db3a256f3f7504cc86a843475-463x503.png" alt="Elastic 欢迎界面。" /><p>15. 通过浏览器访问时，您将看到欢迎界面。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6f44dc23e6f1f625/6a17f6e9be6086ea71004948/a75c151c0144b7efe2b730698c0ed0156fa9b16a-1600x1005.png" alt="Elasticsearch 主页。" /><p>如果您想更改 Elasticsearch 集群规格，例如更改或调整节点大小，可重新应用包含新设置的 yml 配置文件：</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: elasticsearch.k8s.elastic.co/v1
kind: Elasticsearch
metadata:
  name: quickstart
spec:
  version: 9.0.0
  nodeSets:
    - name: default
      count: 2
      config:
        node.store.allow_mmap: false
      podTemplate:
        spec:
          containers:
            - name: elasticsearch
              resources:
                requests:
                  memory: 1.5Gi
                  cpu: 2
                limits:
                  memory: 1.5Gi
                  cpu: 2
EOF<p>在此示例中，我们将再添加一个节点，并修改 RAM 和 CPU。如您所见，现在 <code>kubectl get elasticsearch</code> 显示 2 个节点：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt33a7ba0be483195a/6a17f6ebe8fbcea7da3a1aff/48b475622cc48890bff8105d151f2cbde28d7021-3418x298.png" alt="" /><p>这同样适用于 Kibana：</p>cat &lt;&lt;EOF | kubectl apply -f -
apiVersion: kibana.k8s.elastic.co/v1
kind: Kibana
metadata:
  name: quickstart
spec:
  version: 9.0.0
  http:
    service:
      metadata:
        annotations:
          cloud.google.com/l4-rbs: "enabled"
      spec:
        type: LoadBalancer
  count: 1
  elasticsearchRef:
    name: quickstart
  podTemplate:
    spec:
      containers:
        - name: kibana
          env:
            - name: NODE_OPTIONS
              value: "--max-old-space-size=1024"
          resources:
            requests:
              memory: 0.5Gi
              cpu: 0.5
            limits:
              memory: 1Gi
              cpu: 1
EOF<p>我们可以调整容器 CPU/RAM 以及 <a href="https://nodejs.org/">Node.js</a> 的内存使用量（<a href="https://nodejs.org/api/cli.html#--max-old-space-sizesize-in-mib">max-old-space-size</a>）。</p><p>请注意，<a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s/volume-claim-templates">现有的卷声明无法缩小容量</a>。应用更新后，运维工具将在最短中断时间内完成更改。</p><p>请记得在测试结束后删除集群，以避免产生不必要的成本。</p>gcloud container clusters delete autopilot-cluster-1<h2>后续步骤</h2><p>如果您想了解更多关于 Kubernetes 和 Google Kubernetes Engine 的信息，请查看以下文章：</p><ul><li><p><a href="https://www.elastic.co/docs/deploy-manage/deploy/cloud-on-k8s">Kubernetes 上的 Elastic Cloud | Elastic 文档</a></p></li><li><p><a href="https://cloud.google.com/blog/products/containers-kubernetes/introducing-gke-autopilot">推出 GKE Autopilot ｜ Google Cloud 博客</a></p></li><li><p><a href="https://cloud.google.com/kubernetes-engine/docs/concepts/autopilot-overview">Autopilot 概述 | Google Kubernetes Engine (GKE)</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/eck-gke-autopilot</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/eck-gke-autopilot</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Eduard Martin]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltea045d2d12606d11/6a17f6edb1e113c86179f3e7/d9c462fe63011356671479ccfedd435eec1ede52-1200x628.png" length="0" type="image/png"/>
    <pubDate>Thu, 19 Jun 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[正确使用 JavaScript 的 Elasticsearch，第二部分]]></title>
    <description><![CDATA[了解生产环境最佳实践，并学习如何在 Serverless 环境中运行 Elasticsearch Node.js 客户端，以减少代码错误。 ]]></description>
    <content:encoded><![CDATA[<p>这是 Elasticsearch in JavaScript 系列的第二部分。在<a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i"> 第一部分 中 ，</a> 我们学习了如何正确设置环境、配置 Node.js 客户端、索引数据和搜索。在第二部分中，我们将学习如何实施生产最佳实践，并在无服务器环境中运行 Elasticsearch<a href="http://node.js">Node.js</a>客户端。</p><p>我们将审查</p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#production-best-practices">生产最佳实践</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#error-handling">错误处理能力</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#testing">测试</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#serverless-environments">无服务器环境</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#running-the-client-on-elastic-serverless">在 Elastic Serverless 上运行客户端</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii#running-the-client-on-function-as-a-service-environment">在功能即服务环境中运行客户端</a></p></li></ul></li></ul><p><em>您可以 </em><a href="https://github.com/Delacrobix/JS-client-best-practices_article"><em><strong>在这里</strong></em></a>查看示例的源代码 <em><strong>。</strong></em></p><h2>生产最佳实践</h2><h3>Elasticsearch 中的错误处理</h3><p>Node.js 中 Elasticsearch 客户端的一个有用功能是，它为 Elasticsearch 中可能出现的错误提供了对象，因此您可以用不同的方式验证和处理这些错误。</p><p>要<a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript/connecting#client-error-handling">查看全部内容</a>，请执行此操作： </p>const { errors } = require('@elastic/elasticsearch')
console.log(errors)<p>让我们回到搜索示例，处理一些可能出现的错误：</p>app.get("/search/lexic", async (req, res) =&gt; {
 ....
  } catch (error) {
    if (error instanceof errors.ResponseError) {
      let errorMessage =
        "Response error!, query malformed or server down, contact the administrator!";

      if (error.body.error.type === "parsing_exception") {
        errorMessage = "Query malformed, make sure mappings are set correctly";
      }

      res.status(error.meta.statusCode).json({
        erroStatus: error.meta.statusCode,
        success: false,
        results: null,
        error: errorMessage,
      });
    }

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p><code>ResponseError</code> 尤其是当答案为<code>4xx</code> 或<code>5xx</code> 时，即表示请求不正确或服务器不可用。</p><p>我们可以通过生成错误查询来测试这类错误，比如尝试<strong>在文本类型字段上进行术语查询：</strong></p><p>默认错误：</p> {
    "success": false,
    "results": null,
    "error": "parsing_exception\n\tRoot causes:\n\t\tparsing_exception: [terms] query does not support [visit_details]"
}<p>定制错误： </p>{
    "erroStatus": 400,
    "success": false,
    "results": null,
    "error": "Response error!, query malformed or server down; contact the administrator!"
}<p>我们还可以以某种方式捕捉和处理每种类型的错误。例如，我们可以在<code>TimeoutError</code> 中添加重试逻辑。</p>app.get("/search/semantic", async (req, res) =&gt; {
    try {
  ...
  } catch (error) {
    if (error instanceof errors.TimeoutError) {


     // Retry logic...

      res.status(error.meta.statusCode).json({
        erroStatus: error.meta.statusCode,
        success: false,
        results: null,
        error:
          "The request took more than 10s after 3 retries. Try again later.",
      });
    }
  }
});<h3>测试</h3><p>测试是保证应用程序稳定性的关键。为了以一种与 Elasticsearch 隔离的方式测试代码，我们可以在创建集群时使用<a href="https://github.com/elastic/elasticsearch-js-mock">elasticsearch-js-mock</a>库。</p><p>通过该库，我们可以实例化一个与真实客户端非常相似的客户端，但只需将客户端的 HTTP 层替换为模拟层，其他部分与原始客户端保持一致，就能满足我们的配置要求。</p><p>我们将安装 mocks 库和用于自动测试的<a href="https://github.com/avajs/ava">AVA</a>。</p><p><code>npm install @elastic/elasticsearch-mock</code></p><p><code>npm install --save-dev ava</code></p><p>我们将配置<code>package.json</code> 文件以运行测试。确保它看起来是这样的：</p>"type": "module",
	"scripts": {
		"test": "ava"
	},
	"devDependencies": {
		"ava": "^5.0.0"
	}<p>现在，让我们创建<code>test.js</code> 文件并安装我们的模拟客户端：</p>const { Client } = require('@elastic/elasticsearch')
const Mock = require('@elastic/elasticsearch-mock')

const mock = new Mock()
const client = new Client({
  node: 'http://localhost:9200',
  Connection: mock.getConnection()
})<p>现在，为语义搜索添加一个模拟：</p>function createSemanticSearchMock(query, indexName) {
  mock.add(
    {
      method: "POST",
      path: `/${indexName}/_search`,
      body: {
        query: {
          semantic: {
            field: "semantic_field",
            query: query,
          },
        },
      },
    },
    () =&gt; {
      return {
        hits: {
          total: { value: 2, relation: "eq" },
          hits: [
            {
              _id: "1",
              _score: 0.9,
              _source: {
                owner_name: "Alice Johnson",
                pet_name: "Buddy",
                species: "Dog",
                breed: "Golden Retriever",
                vaccination_history: ["Rabies", "Parvovirus", "Distemper"],
                visit_details:
                  "Annual check-up and nail trimming. Healthy and active.",
              },
            },
            {
              _id: "2",
              _score: 0.7,
              _source: {
                owner_name: "Daniel Kim",
                pet_name: "Mochi",
                species: "Rabbit",
                breed: "Mixed",
                vaccination_history: [],
                visit_details:
                  "Nail trimming and general health check. No issues.",
              },
            },
          ],
        },
      };
    }
  );
}<p>现在我们可以为代码创建一个测试，确保 Elasticsearch 部分始终返回相同的结果：</p>import test from 'ava';

test("performSemanticSearch must return formatted results correctly", async (t) =&gt; {
  const indexName = "vet-visits";
  const query = "Which pets had nail trimming?";

  createSemanticSearchMock(query, indexName);

  async function performSemanticSearch(esClient, q, indexName = "vet-visits") {
    try {
      const result = await esClient.search({
        index: indexName,
        body: {
          query: {
            semantic: {
              field: "semantic_field",
              query: q,
            },
          },
        },
      });

      return {
        success: true,
        results: result.hits.hits,
      };
    } catch (error) {
      if (error instanceof errors.TimeoutError) {
        return {
          success: false,
          results: null,
          error: error.body.error.reason,
        };
      }

      return {
        success: false,
        results: null,
        error: error.message,
      };
    }
  }

  const result = await performSemanticSearch(esClient, query, indexName);

  t.true(result.success, "The search must be successful");
  t.true(Array.isArray(result.results), "The results must be an array");

  if (result.results.length &gt; 0) {
    t.true(
      "_source" in result.results[0],
      "Each result must have a _source property"
    );
    t.true(
      "pet_name" in result.results[0]._source,
      "Results must include the pet_name field"
    );
    t.true(
      "visit_details" in result.results[0]._source,
      "Results must include the visit_details field"
    );
  }
});<p>让我们进行测试。</p><p><code>npm run test</code></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt36304e286146f362/6a170559d7c02237b2de638f/42feae845ae8eae03c37ad7ad114e8db35984812-1186x302.png" alt="" /><p>完成！从现在起，我们就可以测试我们的应用程序，100% 专注于代码而不是外部因素。</p><h2>无服务器环境</h2><h3>如何在 Elastic Serverless 上运行客户端</h3><p>我们介绍了在云端或内部运行 Elasticsearch 的情况；不过，Node.js 客户端也支持与<a href="https://www.elastic.co/guide/en/serverless/current/intro.html">Elastic Cloud Serverless</a> 的连接。</p><p>Elastic Cloud Serverless 允许您创建一个项目，在这个项目中，您无需担心基础设施问题，因为 Elastic 会在内部处理这些问题，您只需担心您想索引的数据以及您想在多长时间内访问这些数据。</p><p>从使用角度来看，Serverless 将计算与存储分离，为<a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-tier-autoscaling">搜索</a>和<a href="https://www.elastic.co/search-labs/blog/elasticsearch-ingest-autoscaling">索引</a>提供了自动扩展功能。这样，您就可以只增长实际需要的资源。</p><p>客户端会进行以下调整，以连接到无服务器：</p><ul><li><p>关闭嗅探，忽略任何与嗅探相关的选项</p></li><li><p>忽略配置中传递的除第一个节点外的所有节点，并忽略任何节点过滤和选择选项</p></li><li><p>启用压缩和 "TLSv1_2_method"（与为弹性云配置时相同）</p></li><li><p>为所有请求添加 "elastic-api-version "HTTP 头信息</p></li><li><p>默认使用 "云连接池"，而不是 "加权连接池</p></li><li><p>关闭卖方 "内容类型 "和 "接受 "标头，转而使用标准 MIME 类型</p></li></ul><p>要连接无服务器项目，需要使用参数 serverMode：serverless。</p>const { Client } = require('@elastic/elasticsearch')
const client = new Client({
  node: 'ELASTICSEARCH_ENDPOINT',
  auth: { apiKey: 'ELASTICSEARCH_API_KEY' },
  serverMode: "serverless",
});<h3>如何在函数即服务环境中运行客户端</h3><p>在示例中，我们使用了 Node.js 服务器，但您也可以使用功能即服务环境连接 AWS lambda、GCP Run 等功能。</p>'use strict'

const { Client } = require('@elastic/elasticsearch')

const client = new Client({
  // client initialisation
})

exports.handler = async function (event, context) {
  // use the client
}<p>另一个例子是连接像 Vercel 这样的服务，它也是无服务器的。您可以查看这个<a href="https://github.com/elastic/elasticsearch-js/blob/main/docs/examples/proxy/README.md">完整的示例</a>，了解如何做到这一点，但<a href="https://github.com/elastic/elasticsearch-js/blob/main/docs/examples/proxy/api/search.js">搜索端点</a>最相关的部分如下所示：</p>const response = await client.search(
  {
    index: INDEX,
    // You could directly send from the browser
    // the Elasticsearch's query DSL, but it will
    // expose you to the risk that a malicious user
    // could overload your cluster by crafting
    // expensive queries.
    query: {
      match: { field: req.body.text },
    },
  },
  {
    headers: {
      Authorization: `ApiKey ${token}`,
    },
  }
);<p>该端点位于 /api 文件夹中，从服务器端运行，因此客户端只能控制与搜索词相对应的 "文本 "参数。</p><p>使用 "功能即服务 "的意义在于，与全天候运行的服务器不同，功能只启动运行该功能的机器，一旦完成，机器就会进入休息模式，以减少资源消耗。</p><p>如果应用程序没有收到太多请求，这种配置会很方便；否则，成本会很高。您还需要考虑<a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtime-environment.html">函数的生命周期</a>和运行时间（在某些情况下可能只有几秒钟）。</p><h2>结论</h2><p>在本文中，我们学习了如何处理错误，这在生产环境中至关重要。我们还介绍了在模拟 Elasticsearch 服务的过程中测试应用程序的方法，无论集群的状态如何，这种方法都能提供可靠的测试，让我们专注于我们的代码。</p><p>最后，我们演示了如何通过配置 Elastic Cloud Serverless 和 Vercel 应用程序来启动完全无服务器堆栈。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-ii</guid>
    <category><![CDATA[Javascript]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc58be329ffebcd60/6a17043e47d49c0bc62d88ab/70fb0ff949f6db9ac9b8a28ecb4329ab915ebf46-720x420.png" length="0" type="image/png"/>
    <pubDate>Mon, 19 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何优化 Elasticsearch 磁盘空间和使用率]]></title>
    <description><![CDATA[了解如何预防并处理 Elasticsearch 磁盘使用率过高（超载）以及容量利用率不足的情况，从而优化集群成本。]]></description>
    <content:encoded><![CDATA[<p>磁盘管理对任何数据库都很重要，Elasticsearch 也不例外。如果没有足够的可用磁盘空间，Elasticsearch 将停止向节点分配分片。这将最终导致您无法向群集写入数据，并有可能导致应用程序中的数据丢失。另一方面，如果磁盘空间过大，则需要为超出需要的资源付费。</p><h2>水印背景</h2><p>Elasticsearch 集群上有各种 "水印 "阈值，可帮助您跟踪可用磁盘空间。当节点上的磁盘填满时，第一个越过的阈值就是 "低磁盘水印"。 第二个阈值就是 "高磁盘水印阈值"。 最后，将达到 "磁盘淹没阶段"。一旦过了这个阈值，群集就会阻止写入已通过水印的节点上有一个分片（主分片或副本）的所有索引。 仍可进行读取（搜索）。</p><h2>如何预防和处理磁盘过满（利用率过高）的情况</h2><p>有多种方法可以处理 Elasticsearch 磁盘过满的情况：</p><ol><li><p><strong>删除</strong> <strong>旧数据：</strong>通常情况下，数据不应无限期保存。防止和解决磁盘过满的方法之一是确保当数据达到一定年限时，对其进行可靠的归档和删除。一种方法是使用<a href="https://www.elastic.co/docs/manage-data/lifecycle/index-lifecycle-management">ILM</a>。</p></li><li><p><strong>增加存储容量：</strong>如果无法删除数据，可能需要添加更多数据节点或增加磁盘大小，以便在不影响性能的情况下保留所有数据。如果需要为群集增加存储容量，则应考虑是否只需增加存储容量，还是同时按比例增加存储容量以及 RAM 和 CPU 资源（请参阅下文有关<a href="https://www.elastic.co/search-labs/blog/optimize-elasticsearch-disk-space-and-usage#the-relationship-between-disk-size,-ram-and-cpu">磁盘大小、RAM 和 CPU 比例的</a>部分）。</p></li></ol><h2>如何为 Elasticsearch 集群增加存储容量</h2><ol><li><p><strong>增加数据节点的数量： </strong>请记住，新节点的大小应与现有节点相同，并使用相同的 Elasticsearch 版本。</p></li><li><p><strong>扩大现有节点的规模： </strong>在基于云的环境中，增加现有节点的磁盘大小和内存/CPU 通常很容易。</p></li><li><p><strong>只增加磁盘大小： </strong>在基于云的环境中，增加磁盘大小通常相对容易。</p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore"><strong>快照</strong></a><a href="https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore"> </a><a href="https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore"><strong>和</strong></a><a href="https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore"> </a><a href="https://www.elastic.co/docs/deploy-manage/tools/snapshot-and-restore"><strong>恢复</strong></a><strong>：</strong>如果您愿意让旧数据根据要求通过自动流程从备份中检索出来，您可以对旧索引进行快照、删除，并根据要求从快照中临时恢复数据。 </p></li><li><p><strong>减少每个分片的副本数量：</strong>减少数据的另一个方法是减少每个分片的副本数量。为了实现高可用性，您希望每个分片有一个副本，但当数据变旧时，您可能不需要副本也能工作。如果数据是持久性的，或者您有备份可以在需要时恢复，那么这种方法通常是可行的。</p></li><li><p><strong>创建警报：</strong>为了防止磁盘将来被填满并采取主动行动，应根据磁盘使用情况创建警报，以便在磁盘开始填满时发出通知。 </p></li></ol><h2>如何预防和处理磁盘容量利用不足的情况</h2><p>如果磁盘容量未得到充分利用，有多种选择可以减少群集的存储容量。</p><h3>如何减少 Elasticsearch 集群的存储容量</h3><p>减少群集存储容量的方法有很多种。</p><p><strong>1.减少数据节点数量</strong></p><p>如果你想减少数据存储，同时按相同比例减少 RAM 和 CPU 资源，那么这是最简单的策略。停用不必要的节点可能会节省最大的成本。</p><p>在停止节点运行之前，您应该</p><ul><li><p>确保要停用的节点不需要作为 MASTER 节点。应始终至少有三个节点具有 MASTER 节点角色。</p></li><li><p>将数据碎片从要退役的节点上移走。</p></li></ul><p><strong>2.用较小的节点取代现有节点</strong></p><p>如果无法进一步减少节点数量（通常最低配置为 3 个），则可能需要缩小现有节点的规模。请记住，最好确保所有数据节点的 RAM 内存和磁盘大小相同，因为分片是根据每个节点的分片数量进行平衡的。</p><p>过程如下</p><ul><li><p>向群集添加新的、较小的节点</p></li><li><p>将碎片迁移到远离将要退役的节点的地方</p></li><li><p>关闭旧节点</p></li></ul><p><strong>3.缩小节点上的磁盘大小</strong></p><p>如果只想减少节点上的磁盘大小，而不改变群集的整体 RAM 或 CPU，那么可以减少每个节点的磁盘大小。减少 Elasticsearch 节点上的磁盘大小并非易事。</p><p>最简单的方法通常是</p><ul><li><p>从节点迁移碎片</p></li><li><p>停止节点</p></li><li><p>在节点上挂载新数据卷，并设置适当大小</p></li><li><p>将旧磁盘卷中的所有数据复制到新卷中</p></li><li><p>分离旧卷 A</p></li><li><p>启动节点并将碎片迁移回节点</p></li></ul><p>这就要求其他节点上有足够的容量，以便在此过程中临时存储节点上的额外碎片。在许多情况下，管理这一流程的成本可能会超过潜在的磁盘使用节余。因此，用具有所需磁盘大小的新节点完全替换该节点可能更简单（请参阅上文 "用较小节点替换现有节点"）。</p><p>在为不必要的资源付费时，显然可以通过优化资源利用率来降低成本。</p><h2>磁盘大小、内存和 CPU 之间的关系</h2><p>集群中磁盘容量与内存的理想比例取决于您的具体使用情况。因此，在考虑更改存储容量时，还应考虑当前的磁盘/内存/CPU 比例是否适当平衡，以及是否需要按相同比例增加/减少内存/CPU。</p><p>内存和 CPU 需求取决于<a href="https://opster.com/guides/elasticsearch/glossary/elasticsearch-indexing/">索引</a>活动量、查询次数和类型，以及搜索和汇总的数据量。这通常与群集上存储的数据量成正比，因此也应与磁盘大小相关。</p><p>磁盘容量和内存之间的比例可根据使用情况进行调整。请看这里的几个例子：</p><p></p><p>指数活动</p><p>保留</p><p>搜索活动</p><p>磁盘容量</p><p>内存</p><p>企业搜索应用程序</p><p>适度摄入原木</p><p>长</p><p>灯光</p><p>2TB</p><p>32GB</p><p>应用程序监控</p><p>大量摄入原木</p><p>短</p><p>灯光</p><p>1TB</p><p>32GB</p><p>电子商务</p><p>轻型数据索引</p><p>无限期</p><p>重型</p><p>500GB</p><p>32GB</p><p><em>请记住，修改节点机器配置时必须小心谨慎，因为这可能会导致节点宕机，而且需要确保分片不会开始迁移到其他已经过度紧张的节点上。</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/optimize-elasticsearch-disk-space-and-usage</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/optimize-elasticsearch-disk-space-and-usage</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt087c3d95b6cb59c5/6a17dbda445de986f54cffd9/5d41a078dd03e4480a0ff4e9591c8618b9bab4d0-720x420.png" length="0" type="image/png"/>
    <pubDate>Fri, 16 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[正确使用 JavaScript 的 Elasticsearch，第一部分]]></title>
    <description><![CDATA[讲解如何用 JavaScript 创建可投入生产的 Elasticsearch 后端。  

探索如何使用 JavaScript 与 Elasticsearch，遵循客户端/服务器最佳实践，搭建包含多个搜索端点的服务器，用于查询 Elasticsearch 文档。]]></description>
    <content:encoded><![CDATA[<p>本文是系列文章的第一篇，介绍如何使用 JavaScript 使用 Elasticsearch。在本系列中，您将学习如何在 JavaScript 环境中使用 Elasticsearch 的基础知识，并回顾创建搜索应用程序的最相关功能和最佳实践。最后，您将了解使用 JavaScript 运行 Elasticsearch 所需的一切。</p><p>在第一部分中，我们将回顾</p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#environment">环境</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#frontend,-backend,-or-serverless?">前端、后端还是无服务器？</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#connecting-the-client">连接客户端</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#indexing-documents">编制文件索引</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#elasticsearch-client">Elasticsearch 客户端</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#semantic-mappings">语义映射</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#bulk-helper">批量助手</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#searching-data">搜索数据</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#lexical-query-(/search/lexic?q=%3Cquery-term%3E)">词法查询</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#semantic-query-(/search/semantic?q=%3Cquery-term%3E)">语义查询</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i#hybrid-query-(/search/hybrid?q=%3Cquery-term%3E)">混合查询</a></p></li></ul></li></ul><p><em>您可以 </em><a href="https://github.com/Delacrobix/JS-client-best-practices_article"><em><strong>在这里</strong></em></a>查看示例的源代码 <em><strong>。</strong></em></p><h3>什么是 Elasticsearch Node.js 客户端？</h3><p><a href="https://www.elastic.co/guide/en/elasticsearch/client/javascript-api/current/index.html">Elasticsearch Node.js 客户端</a>是一个 JavaScript 库，它将 Elasticsearch API 的 HTTP REST 调用放到了 JavaScript 中。这样就能更轻松地处理和使用帮助程序，简化批量编制文档索引等任务。</p><h2>环境</h2><h3>前端、后端还是无服务器？</h3><p>要使用 JavaScript 客户端创建搜索应用程序，我们至少需要两个组件：Elasticsearch 集群和运行客户端的 JavaScript 运行时。</p><p>JavaScript 客户端支持所有 Elasticsearch 解决方案（云、on-prem 和 Serverless），它们之间没有重大区别，因为客户端内部会处理所有变化，所以你不必担心使用哪一种。</p><p>不过，JavaScript 运行时必须从<strong>服务器</strong>运行，而<strong>不能直接从浏览器</strong>运行。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd3ec469c83e3a71a/6a17e3d5445de91da44d00b6/92ce6cfd923c8008fa44f617a58193642d9d5879-661x410.png" alt="在 JavaScript 环境中使用 Elasticsearch。" /><p>这是因为从浏览器调用 Elasticsearch 时，用户可能会获得敏感信息，如集群 API 密钥、主机或查询本身。Elasticsearch 建议<strong>永远不要将集群直接暴露在互联网上 </strong>，而是使用一个中间层来抽象所有这些信息，这样用户只能看到参数。您可以<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/es-security-principles.html#security-protect-cluster-traffic">在这里</a>了解更多相关信息。</p><p>我们建议使用这样的模式：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d7f215f2e70230a/6a17e3d6fbc5f83de6491a13/a08769f08ec73fe57bf2e961cfdfbb1cdd57919d-972x429.png" alt="设置 Elasticsearch Node.js 客户端。" /><p>在这种情况下，客户端只向服务器发送搜索条件和验证密钥，而服务器则完全控制查询和与 Elasticsearch 的通信。</p><h3>连接客户端</h3><p>首先，按照<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">以下步骤</a>创建一个 API 密钥。</p><p>按照前面的示例，我们将创建一个简单的 Express 服务器，并使用 Node.JS 服务器的客户端连接到该服务器。</p><p>我们将使用 NPM 初始化项目，并安装 Elasticsearch 客户端和<a href="https://expressjs.com/">Express。</a>后者是一个在 Node.js 中调用服务器的库。使用 Express，我们可以通过 HTTP 与后端交互。</p><p>让我们初始化项目：</p><p><code>npm init -y</code></p><p>安装依赖项：</p><p><code>npm install @elastic/elasticsearch express split2 dotenv</code></p><p>让我来为你分析一下：</p><ul><li><p><a href="https://www.npmjs.com/package/@elastic/elasticsearch"><em><strong>@elastic/elasticsearch</strong></em></a>：它是 Node.js 的官方客户端</p></li><li><p><a href="https://www.npmjs.com/package/express"><em><strong>快递</strong></em></a>：它将使我们能够运行一个轻量级的 nodejs 服务器，以暴露 Elasticsearch</p></li><li><p><a href="https://www.npmjs.com/package/split2"><em><strong>split2</strong></em></a>： 将文本行分割成数据流。每次处理一行 ndjson 文件时非常有用</p></li><li><p><a href="https://www.npmjs.com/package/dotenv"><em><strong>dotenv</strong></em></a>：允许我们使用 .env 管理环境变量文件</p></li></ul><p>创建 .env文件，并添加以下几行：</p>ELASTICSEARCH_ENDPOINT="Your Elasticsearch endpoint"
ELASTICSEARCH_API_KEY="Your Elasticssearch API"<p>这样，我们就可以使用<code>dotenv</code> 软件包导入这些变量。</p><p>创建<code>server.js</code> 文件：</p>const express = require("express");
const bodyParser = require("body-parser");
const { Client } = require("@elastic/elasticsearch");
 
require("dotenv").config(); //environment variables setup

const ELASTICSEARCH_ENDPOINT = process.env.ELASTICSEARCH_ENDPOINT;
const ELASTICSEARCH_API_KEY = process.env.ELASTICSEARCH_API_KEY;
const PORT = 3000;


const app = express();

app.listen(PORT, () =&gt; {
  console.log("Server running on port", PORT);
});
app.use(bodyParser.json());


let esClient = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: { apiKey: ELASTICSEARCH_API_KEY },  
});

app.get("/ping", async (req, res) =&gt; {
  try {
    const result = await esClient.info();

    res.status(200).json({
      success: true,
      clusterInfo: result,
    });
  } catch (error) {
    console.error("Error getting Elasticsearch info:", error);

    res.status(500).json({
      success: false,
      clusterInfo: null,
      error: error.message,
    });
  }
});<p>这段代码设置了一个基本的 Express.js 服务器，该服务器监听端口 3000，并使用 API 密钥进行身份验证，连接到 Elasticsearch 集群。它包括一个 /ping 端点，通过 GET 请求访问时，可使用 Elasticsearch 客户端的<code>.info()</code> 方法查询 Elasticsearch 集群的基本信息。 </p><p>如果查询成功，会以 JSON 格式返回群集信息；否则会返回错误信息。服务器还使用 body-parser 中间件来处理 JSON 请求体。</p><p>运行文件，启动服务器：</p><p><code>node server.js</code></p><p>答案应该是这样的</p>Server running on port 3000<p>现在，让我们查阅端点<code>/ping</code> ，检查 Elasticsearch 集群的状态。</p>curl http://localhost:3000/ping
{
    "success": true,
    "clusterInfo": {
        "name": "instance-0000000000",
        "cluster_name": "61b7e19eec204d59855f5e019acd2689",
        "cluster_uuid": "BIfvfLM0RJWRK_bDCY5ldg",
        "version": {
            "number": "9.0.0",
            "build_flavor": "default",
            "build_type": "docker",
            "build_hash": "112859b85d50de2a7e63f73c8fc70b99eea24291",
            "build_date": "2025-04-08T15:13:46.049795831Z",
            "build_snapshot": false,
            "lucene_version": "10.1.0",
            "minimum_wire_compatibility_version": "8.18.0",
            "minimum_index_compatibility_version": "8.0.0"
        },
        "tagline": "You Know, for Search"
    }
}<h2>编制文件索引</h2><p>一旦连接起来，我们就可以使用语义<a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">_文本（</a>用于语义搜索）和文本（用于全文查询）等映射对文档进行索引。有了这两种字段类型，我们还可以进行<a href="https://www.elastic.co/what-is/hybrid-search">混合搜索</a>。</p><p>我们将创建一个新的<code>load.js</code> 文件来生成映射并上传文件。</p><h3>Elasticsearch 客户端</h3><p>我们首先需要对客户端进行实例化和身份验证：</p>const { Client } = require("@elastic/elasticsearch");

const ELASTICSEARCH_ENDPOINT = "cluster/project_endpoint";
const ELASTICSEARCH_API_KEY = "apiKey";

const esClient = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: { apiKey: ELASTICSEARCH_API_KEY },
});<h3>语义映射</h3><p>我们将创建一个包含兽医院数据的索引。我们将保存主人、宠物和访问详情的信息。</p><p>我们要进行全文搜索的数据，如名称和描述，将以文本形式存储。类别中的数据，如动物的种类或品种，将以关键字的形式存储。</p><p>此外，我们还将把所有字段的值复制到一个 semantic_text 字段中，以便也能针对这些信息运行语义搜索。</p>const INDEX_NAME = "vet-visits";

const createMappings = async (indexName, mapping) =&gt; {
  try {
    const body = await esClient.indices.create({
      index: indexName,
      body: {
        mappings: mapping,
      },
    });

    console.log("Index created successfully:", body);
  } catch (error) {
    console.error("Error creating mapping:", error);
  }
};

await createMappings(INDEX_NAME, {
  properties: {
    owner_name: {
      type: "text",
      copy_to: "semantic_field",
    },
    pet_name: {
      type: "text",
      copy_to: "semantic_field",
    },
    species: {
      type: "keyword",
      copy_to: "semantic_field",
    },
    breed: {
      type: "keyword",
      copy_to: "semantic_field",
    },
    vaccination_history: {
      type: "keyword",
      copy_to: "semantic_field",
    },
    visit_details: {
      type: "text",
      copy_to: "semantic_field",
    },
    semantic_field: {
      type: "semantic_text",
    },
  },
});<h3>批量助手</h3><p>客户端的另一个优势是，我们可以使用<a href="https://www.elastic.co/guide/en/elasticsearch/client/javascript-api/current/client-helpers.html#bulk-helper">批量助手</a>来分批建立索引。通过批量辅助器，我们可以轻松处理并发、重试等问题，以及如何处理通过函数成功或失败的每个文档。</p><p>该助手的一个吸引人的特点是可以使用数据流。该功能允许您逐行发送文件，而不是将整个文件存储在内存中并一次性发送到 Elasticsearch。</p><p>要将数据上传到 Elasticsearch，请在项目根目录下创建名为 data.ndjson 的文件，并添加以下信息（也可以从<a href="https://github.com/Delacrobix/JS-client-best-practices_article/blob/main/data.ndjson">此处</a>下载包含数据集的文件）：</p>{"owner_name":"Alice Johnson","pet_name":"Buddy","species":"Dog","breed":"Golden Retriever","vaccination_history":["Rabies","Parvovirus","Distemper"],"visit_details":"Annual check-up and nail trimming. Healthy and active."}
{"owner_name":"Marco Rivera","pet_name":"Milo","species":"Cat","breed":"Siamese","vaccination_history":["Rabies","Feline Leukemia"],"visit_details":"Slight eye irritation, prescribed eye drops."}
{"owner_name":"Sandra Lee","pet_name":"Pickles","species":"Guinea Pig","breed":"Mixed","vaccination_history":[],"visit_details":"Loss of appetite, recommended dietary changes."}
{"owner_name":"Jake Thompson","pet_name":"Luna","species":"Dog","breed":"Labrador Mix","vaccination_history":["Rabies","Bordetella"],"visit_details":"Mild ear infection, cleaning and antibiotics given."}
{"owner_name":"Emily Chen","pet_name":"Ziggy","species":"Cat","breed":"Mixed","vaccination_history":["Rabies","Feline Calicivirus"],"visit_details":"Vaccination update and routine physical."}
{"owner_name":"Tomás Herrera","pet_name":"Rex","species":"Dog","breed":"German Shepherd","vaccination_history":["Rabies","Parvovirus","Leptospirosis"],"visit_details":"Follow-up for previous leg strain, improving well."}
{"owner_name":"Nina Park","pet_name":"Coco","species":"Ferret","breed":"Mixed","vaccination_history":["Rabies"],"visit_details":"Slight weight loss; advised new diet."}
{"owner_name":"Leo Martínez","pet_name":"Simba","species":"Cat","breed":"Maine Coon","vaccination_history":["Rabies","Feline Panleukopenia"],"visit_details":"Dental cleaning. Minor tartar buildup removed."}
{"owner_name":"Rachel Green","pet_name":"Rocky","species":"Dog","breed":"Bulldog Mix","vaccination_history":["Rabies","Parvovirus"],"visit_details":"Skin rash, antihistamines prescribed."}
{"owner_name":"Daniel Kim","pet_name":"Mochi","species":"Rabbit","breed":"Mixed","vaccination_history":[],"visit_details":"Nail trimming and general health check. No issues."}<p>我们使用 split2 对文件行进行流式处理，而批量助手则将它们发送到 Elasticsearch。</p>const { createReadStream } = require("fs");
const split = require("split2");
 
const indexData = async (filePath, indexName) =&gt; {
  try {
    console.log(`Indexing data from ${filePath} into ${indexName}...`);

    const result = await esClient.helpers.bulk({
      datasource: createReadStream(filePath).pipe(split()),

      onDocument: () =&gt; {
        return {
          index: { _index: indexName },
        };
      },
      onDrop(doc) {
        console.error("Error processing document:", doc);
      },
    });

    console.log("Bulk indexing successful elements:", result.items.length);
  } catch (error) {
    console.error("Error indexing data:", error);
    throw error;
  }
};

await indexData("./data.ndjson", INDEX_NAME);<p>上面的代码读取 .ndjson文件，并使用<code>helpers.bulk</code> 方法将每个 JSON 对象批量索引到指定的 Elasticsearch 索引中。它使用<code>createReadStream</code> 和<code>split2</code> 对文件进行流式处理，为每个文件设置索引元数据，并记录处理失败的文件。完成后，它会记录成功索引的项目数。</p><p>除<code>indexData</code> 功能外，您还可以使用 Kibana 直接通过用户界面上传文件，并使用<a href="https://www.elastic.co/docs/manage-data/ingest/upload-data-files">上传数据文件用户界面。</a></p><p>我们运行文件，将文件上传到 Elasticsearch 集群。</p><p><code>node load.js</code></p>Creating mappings for index vet-visits...
Index created successfully: { acknowledged: true, shards_acknowledged: true, index: 'vet-visits' }
Indexing data from ./data.ndjson into vet-visits...
Bulk indexing completed. Total documents: 10, Failed: 0<h2>在 Elasticsearch 中搜索数据</h2><p>回到<code>server.js</code> 文件，我们将创建不同的端点来执行词法、语义或混合搜索。</p><p>简而言之，这些类型的搜索并不相互排斥，而是取决于您需要回答的问题类型。</p><p>查询类型</p><p>用例</p><p>问题示例</p><p>词法查询</p><p>问题中的单词或词根很可能出现在索引文件中。问题与文件之间的标记相似性。</p><p>我在找一件蓝色运动 T 恤。</p><p>语义查询</p><p>问题中的词语不可能出现在文件中。问题与文件之间的概念相似性。</p><p>我在寻找适合寒冷天气穿的衣服。</p><p>混合搜索</p><p>问题包含词汇和/或语义成分。问题与文档之间的标记和语义相似性。</p><p>我想为海滩婚礼找一件 S 码的礼服。</p><p>问题的<em><strong>词汇 </strong></em>部分很可能是标题和说明的一部分，或者是类别名称，而<em><strong>语义 </strong></em>部分则是与这些领域相关的概念。<em><strong>蓝色</strong></em>可能是一个类别名称或描述的一部分，<em><strong>海滩婚礼</strong></em>不太可能是，但可以与亚麻服装在语义上相关。</p><h3>词法查询 (/search/lexic?q=&lt;query_term&gt;)</h3><p>词法搜索也称全文搜索，是指基于标记的相似性进行搜索；也就是说，经过分析后，将返回包含搜索标记的文档。</p><p>您可以<a href="https://www.elastic.co/demo-gallery/lexical-search">点击此处</a>查看我们的词法搜索实践教程。</p>app.get("/search/lexic", async (req, res) =&gt; {
  const { q } = req.query;

  const INDEX_NAME = "vet-visits";

  try {
    const result = await esClient.search({
      index: INDEX_NAME,
      size: 5,
      body: {
        query: {
          multi_match: {
            query: q,
            fields: ["owner_name", "pet_name", "visit_details"],
          },
        },
      },
    });

    res.status(200).json({
      success: true,
      results: result.hits.hits
    });
  } catch (error) {
    console.error("Error performing search:", error);

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p>我们测试：<em><strong>修剪指甲</strong></em></p>curl http://localhost:3000/search/lexic?q=nail%20trimming<p>请回答：</p>{
    "success": true,
    "results": [
        {
            "_index": "vet-visits",
            "_id": "-RY6RJYBLe2GoFQ6-9n9",
            "_score": 2.7075968,
            "_source": {
                "pet_name": "Mochi",
                "owner_name": "Daniel Kim",
                "species": "Rabbit",
                "visit_details": "Nail trimming and general health check. No issues.",
                "breed": "Mixed",
                "vaccination_history": []
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8BY6RJYBLe2GoFQ6-9n9",
            "_score": 2.560356,
            "_source": {
                "pet_name": "Buddy",
                "owner_name": "Alice Johnson",
                "species": "Dog",
                "visit_details": "Annual check-up and nail trimming. Healthy and active.",
                "breed": "Golden Retriever",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Distemper"
                ]
            }
        }
    ]
}<h3>语义查询 (/search/semantic?q=&lt;query_term&gt;)</h3><p>语义搜索与词汇搜索不同，它通过矢量搜索找到与搜索词含义相似的结果。</p><p>您可以<a href="https://www.elastic.co/demo-gallery/semantic-search">点击这里</a>查看我们的语义搜索实践教程。</p>app.get("/search/semantic", async (req, res) =&gt; {
  const { q } = req.query;

  const INDEX_NAME = "vet-visits";

  try {
    const result = await esClient.search({
      index: INDEX_NAME,
      size: 5,
      body: {
        query: {
          semantic: {
            field: "semantic_field",
            query: q
          },
        },
      },
    });

    res.status(200).json({
      success: true,
      results: result.hits.hits,
    });
  } catch (error) {
    console.error("Error performing search:", error);

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p>我们进行测试：<em><strong>谁做了修脚？</strong></em></p>curl http://localhost:3000/search/semantic?q=Who%20got%20a%20pedicure?<p>请回答：</p>{
    "success": true,
    "results": [
        {
            "_index": "vet-visits",
            "_id": "-RY6RJYBLe2GoFQ6-9n9",
            "_score": 4.861466,
            "_source": {
                "owner_name": "Daniel Kim",
                "pet_name": "Mochi",
                "species": "Rabbit",
                "breed": "Mixed",
                "vaccination_history": [],
                "visit_details": "Nail trimming and general health check. No issues."
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8BY6RJYBLe2GoFQ6-9n9",
            "_score": 4.7152824,
            "_source": {
                "pet_name": "Buddy",
                "owner_name": "Alice Johnson",
                "species": "Dog",
                "visit_details": "Annual check-up and nail trimming. Healthy and active.",
                "breed": "Golden Retriever",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Distemper"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "9RY6RJYBLe2GoFQ6-9n9",
            "_score": 1.6717153,
            "_source": {
                "pet_name": "Rex",
                "owner_name": "Tomás Herrera",
                "species": "Dog",
                "visit_details": "Follow-up for previous leg strain, improving well.",
                "breed": "German Shepherd",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Leptospirosis"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "9xY6RJYBLe2GoFQ6-9n9",
            "_score": 1.5600781,
            "_source": {
                "pet_name": "Simba",
                "owner_name": "Leo Martínez",
                "species": "Cat",
                "visit_details": "Dental cleaning. Minor tartar buildup removed.",
                "breed": "Maine Coon",
                "vaccination_history": [
                    "Rabies",
                    "Feline Panleukopenia"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "-BY6RJYBLe2GoFQ6-9n9",
            "_score": 1.2696637,
            "_source": {
                "pet_name": "Rocky",
                "owner_name": "Rachel Green",
                "species": "Dog",
                "visit_details": "Skin rash, antihistamines prescribed.",
                "breed": "Bulldog Mix",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus"
                ]
            }
        }
    ]
}<h3>混合查询 (/search/hybrid?q=&lt;query_term&gt;)</h3><p>混合搜索允许我们将语义搜索和词法搜索结合起来，从而获得两全其美的效果：既能获得标记搜索的精确性，又能获得语义搜索的意义接近性。</p>app.get("/search/hybrid", async (req, res) =&gt; {
  const { q } = req.query;

  const INDEX_NAME = "vet-visits";

  try {
    const result = await esClient.search({
      index: INDEX_NAME,
      body: {
        retriever: {
          rrf: {
            retrievers: [
              {
                standard: {
                  query: {
                    bool: {
                      must: {
                         multi_match: {
             query: q,
            fields: ["owner_name", "pet_name", "visit_details"],
          },
                      },
                    },
                  },
                },
              },
              {
                standard: {
                  query: {
                    bool: {
                      must: {
                        semantic: {
                          field: "semantic_field",
                          query: q,
                        },
                      },
                    },
                  },
                },
              },
            ],
          },
        },
        size: 5,
      },
    });

    res.status(200).json({
      success: true,
      results: result.hits.hits,
    });
  } catch (error) {
    console.error("Error performing search:", error);

    res.status(500).json({
      success: false,
      results: null,
      error: error.message,
    });
  }
});<p>我们以 "<em><strong>谁做了修脚或牙科治疗？"</strong></em></p>curl http://localhost:3000/search/hybrid?q=who%20got%20a%20pedicure%20or%20dental%20treatment<p>响应：</p>{
    "success": true,
    "results": [
        {
            "_index": "vet-visits",
            "_id": "9xY6RJYBLe2GoFQ6-9n9",
            "_score": 0.032522473,
            "_source": {
                "pet_name": "Simba",
                "owner_name": "Leo Martínez",
                "species": "Cat",
                "visit_details": "Dental cleaning. Minor tartar buildup removed.",
                "breed": "Maine Coon",
                "vaccination_history": [
                    "Rabies",
                    "Feline Panleukopenia"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "-RY6RJYBLe2GoFQ6-9n9",
            "_score": 0.016393442,
            "_source": {
                "pet_name": "Mochi",
                "owner_name": "Daniel Kim",
                "species": "Rabbit",
                "visit_details": "Nail trimming and general health check. No issues.",
                "breed": "Mixed",
                "vaccination_history": []
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8BY6RJYBLe2GoFQ6-9n9",
            "_score": 0.015873017,
            "_source": {
                "pet_name": "Buddy",
                "owner_name": "Alice Johnson",
                "species": "Dog",
                "visit_details": "Annual check-up and nail trimming. Healthy and active.",
                "breed": "Golden Retriever",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Distemper"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "9RY6RJYBLe2GoFQ6-9n9",
            "_score": 0.015625,
            "_source": {
                "pet_name": "Rex",
                "owner_name": "Tomás Herrera",
                "species": "Dog",
                "visit_details": "Follow-up for previous leg strain, improving well.",
                "breed": "German Shepherd",
                "vaccination_history": [
                    "Rabies",
                    "Parvovirus",
                    "Leptospirosis"
                ]
            }
        },
        {
            "_index": "vet-visits",
            "_id": "8xY6RJYBLe2GoFQ6-9n9",
            "_score": 0.015384615,
            "_source": {
                "pet_name": "Luna",
                "owner_name": "Jake Thompson",
                "species": "Dog",
                "visit_details": "Mild ear infection, cleaning and antibiotics given.",
                "breed": "Labrador Mix",
                "vaccination_history": [
                    "Rabies",
                    "Bordetella"
                ]
            }
        }
    ]
}<h2>结论</h2><p>在本系列的第一部分中，我们介绍了如何按照客户端/服务器最佳实践设置环境并创建带有不同搜索端点的服务器，以查询 Elasticsearch 文档。查看我们系列的<a href="https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i">第二部分</a>，您将了解生产最佳实践以及如何在无服务器环境中运行 Elasticsearch Node.js 客户端。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/how-to-use-elasticsearch-in-javascript-part-i</guid>
    <category><![CDATA[Javascript]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt16d00c8a548b32e8/6a17e3d8fbc5f8c740491a19/72200540ed258779d87e53a72ea189f8a138540c-1600x901.png" length="0" type="image/png"/>
    <pubDate>Thu, 15 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何配置 Elasticsearch 索引中的副本数量]]></title>
    <description><![CDATA[了解如何在 Elasticsearch 索引中配置 number_of_replicas 以提升搜索性能并提供节点故障恢复能力。 
]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch 设计为分布式系统，可处理大量数据并提供高可用性。实现这一点的关键功能之一是索引复制概念，该概念由<code>number_of_replicas</code> 设置控制。本文将深入探讨这一设置的细节、影响以及如何正确配置。</p><h2>副本在 Elasticsearch 中的作用</h2><p>在 Elasticsearch 中，索引是在多个主分片上分区的文档集合。每个主分片都是一个独立的 Apache Lucene 索引，索引中的文档分布在所有主分片中。为确保高可用性和数据冗余，Elasticsearch 允许每个分片拥有一个或多个副本（称为副本）。<code>number_of_replicas</code> 设置可控制 Elasticsearch 为索引中每个主分区创建的副本分区（拷贝）数量。默认情况下，Elasticsearch 会为每个主分片创建一个副本，但可以根据系统要求进行更改。</p><h2>配置复制数</h2><p><code>number_of_replicas</code> 设置可在创建索引时配置，也可稍后更新。下面是创建索引时的设置方法：</p>PUT /my_index
{
  "settings": {
    "number_of_replicas": 2
  }
}<p>在此示例中，Elasticsearch 将为<code>my_index</code> 索引中的每个主分区创建两个副本。</p><p>要更新现有索引的<code>number_of_replicas</code> 设置，可以使用<code>_settings</code> API：</p>PUT /my_index/_settings
{
  "number_of_replicas": 3
}<p>该命令将更新<code>my_index</code> 索引，使每个主分区都有三个副本。</p><h2>复制数设置的影响</h2><p><code>number_of_replicas</code> 设置对 Elasticsearch<a href="https://opster.com/guides/elasticsearch/glossary/elasticsearch-cluster/">集群的</a>性能和弹性有重大影响。以下是一些需要考虑的要点：</p><ol><li><p><strong>数据冗余和可用性：</strong>通过为每个分片创建更多副本，增加<code>number_of_replicas</code> 可提高数据的可用性。如果某个节点发生故障，Elasticsearch 仍可从其余<a href="https://opster.com/guides/elasticsearch/glossary/elasticsearch-node/">节点</a>上的副本分片提供数据。</p></li><li><p><strong>搜索性能：</strong>副本分片可以为读取请求提供服务，因此拥有更多的副本可以通过在更多分片上分配负载来提高搜索性能。</p></li><li><p><strong>写性能：</strong>不过，每次写操作都必须在分片的每个副本上执行。因此，较高的<code>number_of_replicas</code> 会降低<a href="https://opster.com/guides/elasticsearch/glossary/elasticsearch-indexing/">索引</a>性能，因为它会增加每次写入必须执行的操作次数。</p></li><li><p><strong>存储要求：</strong>更多的副本意味着更多的存储空间。应确保群集有足够的容量来存储额外的副本。</p></li><li><p><strong>节点故障恢复能力：</strong> <code>number_of_replicas</code> 的设置应考虑群集中的节点数量。如果<code>number_of_replicas</code> 等于或大于节点数，则群集可以承受多个节点的故障而不会丢失数据。</p></li></ol><h2>设置复制数的最佳做法</h2><p><code>number_of_replicas</code> 的最佳设置取决于系统的具体要求。不过，这里有一些通用的最佳做法：</p><ul><li><p>对于单节点集群，<code>number_of_replicas</code> 应设置为 0，因为没有其他节点可以容纳副本。</p></li><li><p>对于多节点集群，<code>number_of_replicas</code> 至少应设置为 1，以确保数据冗余和高可用性。</p></li><li><p>如果搜索性能是一个优先事项，请考虑增加<code>number_of_replicas</code> 。不过，请注意写入性能和存储要求之间的权衡。</p></li><li><p>始终确保集群有足够的容量来存储额外的副本。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-index-number-of_replicas</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-index-number-of_replicas</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd041e871a8935448/6a17de320b0bedf404dd34ab/23b96aaa1a38b1f4747b4a87695d816f24c0cf70-720x421.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 14 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[从 Elasticsearch 文档中删除字段]]></title>
    <description><![CDATA[了解如何使用 Update API、脚本或重新索引，从 Elasticsearch 文档中删除字段，支持单次及批量操作。]]></description>
    <content:encoded><![CDATA[<p>在 Elasticsearch 中，从文档中删除字段是一项常见需求。当您想从索引中删除不必要或过时的信息时，这将非常有用。在本文中，我们将讨论从 Elasticsearch 文档中删除字段的不同方法以及示例和分步说明。 </p><h2>方法 1：使用更新 API</h2><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/update-document">Update API</a> 允许您通过提供一段修改文档源数据的脚本来更新文档。您可以通过将字段值设为 null 来使用此 API 删除文档中的字段。以下是详细的操作步骤：</p><p>1.确定要更新的文档的索引、文档类型（如果使用 Elasticsearch 6.x 或更早版本）和文档 ID。</p><p>2.使用更新 API 并编写脚本，将字段设置为空，或者更好的做法是从源文档中删除该字段。下面的示例演示了如何从 "my_index "索引中 ID 为 "1 "的文档中删除 "field_to_delete "字段：</p>POST /my_index/_update/1
{
  "script": "ctx._source.remove('field_to_delete')"
}<p>3.执行请求。如果成功，Elasticsearch 将返回一个响应，表明文档已被更新。</p><p>注意：此方法只能从指定文档中删除字段。该字段仍将存在于映射和索引中的其他文档中。</p><h2>方法二：使用修改后的源数据进行重新索引</h2><p>若要从某个索引的所有文档中删除一个字段，您可以使用 <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-reindex">Reindex API</a> 来创建一个包含修改后源数据的新索引。具体操作如下：</p><p>1.创建一个新索引，其设置和映射与原始索引相同。您可以使用获取索引 API 来检索原始索引的设置和映射。</p><p>2.使用重新索引 API 将原始索引中的文档复制到新索引中，同时从源中删除字段。下面的示例演示了如何从 "my_index "索引中的所有文档中删除 "field_to_delete "字段：</p>POST /_reindex
{
  "source": {
    "index": "my_index"
  },
  "dest": {
    "index": "new_index"
  },
  "script": {
    "source": "ctx._source.remove('field_to_delete')"
  }
}<p>
3.验证新索引是否包含已删除字段的正确文件。</p><p>4.如果一切正常，就可以删除原始索引，如有必要，还可以为新索引添加一个别名，其名称与原始索引名称相同。</p><h2>方法三：更新映射并重新索引</h2><p>如果要从映射和索引中的所有文档中删除某个字段，可以更新映射，然后重新索引文档。具体方法如下</p><p>1.创建一个新索引，设置与原始索引相同。</p><p>2.使用获取映射 API 检索原始索引的映射。</p><p>3.修改映射，删除要删除的字段。</p><p>4.使用 Put Mapping API 将修改后的映射应用到新索引。</p><p>5.如方法 2 所述，使用重新索引 API 将原始索引中的文档复制到新索引中。</p><p>6.验证新索引是否包含已删除字段的正确文件，以及映射中是否不存在该字段。</p><p>7. 如果一切正常，您可以删除原始索引，并在必要时为新的索引添加一个与原始索引同名的别名。</p><h2>结论</h2><p>在本文中，我们讨论了从 Elasticsearch 文档中删除字段的三种方法：使用 Update API、使用修改后的源重新索引，以及更新映射并重新索引。每种方法都有自己的用例和权衡，因此请选择最适合您要求的方法。在将更改应用到生产环境之前，请务必记住测试更改并验证结果。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-delete-field-from-document</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-delete-field-from-document</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8deb617c89943b69/6a17e26c4b055d209e43212f/89278eb7309b7f3018c61be2b514d1fd25b9564d-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 09 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何在 Elasticsearch 中连接两个索引]]></title>
    <description><![CDATA[解释如何使用术语查询、Logstash elasticsearch 过滤器、浓缩处理器和 ES|QL 来连接 Elasticsearch 中的两个索引。]]></description>
    <content:encoded><![CDATA[<p>在 Elasticsearch 中，连接两个索引不像在传统 SQL 关系数据库中那么简单。不过，使用 Elasticsearch 提供的某些技术和功能也可以实现类似的结果。</p><p>历史上，许多人使用<a href="https://www.elastic.co/cn/docs/reference/elasticsearch/mapping-reference/nested"><code>nested</code></a><a href="https://www.elastic.co/cn/docs/reference/elasticsearch/mapping-reference/nested"> 字段类型</a>作为将不同索引连接在一起的机制。然而，由于 Kibana 的查询成本高昂且支持不完整，特别是镜头可视化功能，该功能受到了限制。</p><p>本文将深入探讨在 Elasticsearch 中连接两个索引的过程，重点介绍以下方法： </p><ol><li><p>使用<code>terms</code> 查询</p></li><li><p>在摄取流水线中使用<code>enrich</code> 处理器</p></li><li><p>Logstash<code>elasticsearch</code> 过滤器插件</p></li><li><p>ES|QL <code>ENRICH</code></p></li><li><p>ES|QL <code>LOOKUP JOIN</code></p></li></ol><h2>使用术语查询</h2><p><a href="https://www.elastic.co/cn/docs/reference/query-languages/query-dsl/query-dsl-terms-query">术语查询</a>是 Elasticsearch 中连接两个索引的最有效方法之一。该查询用于检索在特定字段中包含一个或多个精确术语的文档。下面我们讨论如何使用它来连接两个索引。</p><p>首先，您需要从第一个索引中获取所需的数据。这可以通过简单的 GET 请求和从<code>_source</code> 属性中提取值来实现。</p># Simple GET request
GET first_index/_search<p>获得第一个索引的数据后，就可以用它来查询第二个索引。这是通过<code>terms</code> 查询完成的，您可以在查询中指定要匹配的字段和值。</p><p>下面就是一个例子：</p>GET second_index/_search
{
  "query": {
    "terms": {
      "field_in_second_index": ["value1_from_first_index", "value2_from_first_index"]
    }
  }
}<p>
在本例中，<code>field_in_second_index</code> 是第二个索引中要与第一个索引中的值匹配的字段。<code>value1_from_first_index</code> 和<code>value2_from_first_index</code> 是第一个索引中要在第二个索引中匹配的值。</p><p>术语查询还支持使用<a href="https://www.elastic.co/cn/docs/reference/query-languages/query-dsl/query-dsl-terms-query#query-dsl-terms-lookup">术语查找</a>技术一次性完成上述两个步骤。Elasticsearch 会以透明方式从另一个索引中检索要匹配的值。例如，如果您有一个包含球员列表的球队索引：</p>PUT teams/_doc/team1
{
  "players":   ["john", "bill", "michael"]
}
PUT teams/_doc/team2
{
  "players":   ["aaron", "joe", "donald"]
}<p>如下图所示，可以通过人员索引查询在 team1 队中比赛的所有人员：</p>GET people/_search?pretty
{
  "query": {
    "terms": {
        "name" : {
            "index" : "teams",
            "id" : "team1",
            "path" : "players"
        }
    }
  }
}<p>在上面的示例中，Elasticsearch 会以透明方式从球队索引中 id 为 team1 的文档中检索球员姓名（即"john"、"bill "和 "michael"），并查找人物索引中所有在姓名字段中包含这些值的文档。</p><p>对于那些好奇的人来说，等价的 SQL 查询应该是这样的：</p><h2>使用浓缩处理器</h2><p><a href="https://www.elastic.co/cn/docs/reference/enrich-processor/enrich-processor"><code>enrich</code></a><a href="https://www.elastic.co/cn/docs/reference/enrich-processor/enrich-processor"> 处理器</a>是另一个可用于连接 Elasticsearch 中两个索引的强大工具。该处理器通过添加来自预定义丰富索引的数据来丰富输入文件的数据。</p><p>下面介绍如何使用浓缩处理器连接两个索引：</p><p>1.首先，您需要创建一个浓缩策略。该策略定义了使用哪个索引来丰富输入文档、匹配哪个字段以及使用哪个字段来丰富输入文档。</p><p>下面就是一个例子：</p>PUT _enrich/policy/my_enrich_policy
{
  "match": {
    "indices": "first_index",
    "match_field": "field_in_first_index",
    "enrich_fields": ["field_to_enrich"]
  }
}<p>2.创建策略后，需要执行该策略，以便根据新创建的策略创建 enrich 索引：</p>PUT _enrich/policy/my_enrich_policy/_execute<p>这将建立一个新的隐藏浓缩索引，在浓缩过程中使用。根据源索引的大小，这一操作可能需要一些时间。在进行下一步之前，请确保充实政策已完全制定。</p><p>3.建立丰富策略后，就可以在摄取管道中使用丰富处理器来丰富传入文档的数据：</p>PUT _ingest/pipeline/my_pipeline
{
  "processors": [
    {
      "enrich": {
        "policy_name": "my_enrich_policy",
        "field": "field_in_second_index",
        "target_field": "enriched_field"
      }
    }
  ]
}<p>在本例中，<code>field_in_second_index</code> 是第二个索引中需要与第一个索引中的<code>match_field</code> 匹配的字段。<code>enriched_field</code> 是第二个索引中的新字段，将包含第一个索引<code>enrich_fields</code> 中的丰富数据。</p><p>这种方法的一个缺点是，如果<code>first_index</code> 中的数据发生变化，则需要重新执行浓缩策略。丰富索引不会自动更新或同步源索引。但是，如果<code>first_index</code> 相对稳定，那么这种方法就很有效。</p><h2>Logstash elasticsearch 过滤器插件</h2><p>如果使用 Logstash，另一个与上述<code>enrich</code> 处理器类似的选项是使用<code>elasticsearch</code> 过滤器插件，根据指定的查询将相关字段添加到事件中。Logstash 管道的配置位于<code>.conf</code> 文件中，如<code>my-pipeline.conf</code> 。</p><p>假设我们的管道使用<a href="https://www.elastic.co/cn/docs/reference/logstash/plugins/plugins-inputs-elasticsearch"><code>elasticsearch</code></a><a href="https://www.elastic.co/cn/docs/reference/logstash/plugins/plugins-inputs-elasticsearch"> 输入插件</a>从 Elasticsearch 中提取日志，并通过查询缩小选择范围：</p>input {
  # Read all documents from Elasticsearch matching the given query
  elasticsearch {
    hosts =&gt; "localhost"
    query =&gt; '{ "query": { "match": { "statuscode": 200 } }, "sort": [ "_doc" ] }'
  }
}<p>如果我们想用给定索引的信息来丰富这些信息，可以使用<code>filter</code> 部分的<a href="https://www.elastic.co/cn/docs/reference/logstash/plugins/plugins-filters-elasticsearch"><code>elasticsearch</code></a><a href="https://www.elastic.co/cn/docs/reference/logstash/plugins/plugins-filters-elasticsearch"> 过滤器插件</a>来丰富我们的日志：</p>filter {
   elasticsearch {
      hosts =&gt; ["localhost"]
      index =&gt; "index_name"
      query =&gt; "type:start AND operation:%{[opid]}"
      fields =&gt; { "@timestamp" =&gt; "started" }
   }
}<p>上述代码将从索引<code>index_name</code> 中查找文件，其中<code>type</code> 为起始值，操作字段与指定的<code>opid</code> 匹配，然后将<code>@timestamp</code> 字段的值复制到名为<code>started</code> 的新字段中。</p><p>然后，丰富的文档将被发送到适当的输出源，在本例中是使用<a href="https://www.elastic.co/cn/docs/reference/logstash/plugins/plugins-outputs-elasticsearch"><code>elasticsearch</code></a><a href="https://www.elastic.co/cn/docs/reference/logstash/plugins/plugins-outputs-elasticsearch"> 输出插件</a>发送到 Elasticsearch：</p>output {
    elasticsearch {
        hosts =&gt; "localhost"
        data_stream =&gt; "true"
    }
}<p>如果您已经在使用 Logstash，该选项可能有助于将丰富逻辑整合到一个地方，并在新事件发生时进行处理。但是，如果您不这样做，就会增加解决方案的复杂性，而且您还需要运行和维护另一个组件。</p><h2>ES|QL ENRICH</h2><p>在 8.14 版本中引入的<a href="https://www.elastic.co/cn/docs/explore-analyze/query-filter/languages/esql">ES|QL</a> 是 Elasticsearch 支持的管道式查询语言，可用于过滤、转换和分析数据。使用 ENRICH 处理命令，我们就可以使用丰富策略从现有索引中添加数据。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbeb8992bde773461/6a17f6b663baff00c9741dd4/03aadddc08afffff3f6526c9c052999c97fa09dd-1600x989.png" alt="丰富 esql" /><p>以原浓缩处理器示例中的相同策略<code>my_enrich_policy</code> 为例，ES|QL 示例如下：</p><p>也可以覆盖匹配字段和丰富字段，在我们的例子中分别是<code>field_in_first_index</code> 和<code>field_to_enrich</code> ：</p><p>虽然 ES|QL 的明显限制是需要先指定丰富策略，但 ES|QL 确实提供了根据需要调整字段的灵活性。</p><h2>es|ql 查找连接</h2><p>Elasticsearch 8.18 引入了一种在 Elasticsearch 中连接索引的新方法，即<code>LOOKUP JOIN</code> 命令。该命令在连接的右侧使用新的<a href="https://www.elastic.co/cn/docs/reference/elasticsearch/index-settings/index-modules#index-mode-setting">查找索引模式</a>，以 SQL 风格的 LEFT OUTER JOIN 方式运行。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt783ffb3f9802f92d/6a17f6b8e9ea870608a9c788/1d73495979c4d6bb675c4c966ea86d9a72dc1c48-510x605.png" alt="es|ql 查找连接" /><p>再看我们之前的例子，新的查询如下，其中<code>match_field</code> 需要同时出现在<code>first_index</code> 和<code>second_index</code> 中：</p><p>与其他方法相比，LOOKUP JOIN 的优势在于它不需要任何<code>enrich</code> 策略，因此也不需要与设置策略相关的额外处理。与本文讨论的其他方法不同，它在处理经常变化的丰富数据时非常有用。</p><h2>结论</h2><p>总之，虽然 Elasticsearch 不支持传统的连接操作，但它提供了各种功能，可用于实现类似的结果。具体来说，我们介绍了如何使用连接操作：</p><ol><li><p><code>terms</code> 查询</p></li><li><p><code>enrich</code> 摄录流水线中的处理器</p></li><li><p>Logstash<code>elasticsearch</code> 过滤器插件</p></li><li><p>ES|QL <code>ENRICH</code></p></li><li><p>ES|QL <code>LOOKUP JOIN</code></p></li></ol><p>需要注意的是，这些方法都有其局限性，应根据具体要求和数据性质谨慎使用。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-join-two-indexes</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-join-two-indexes</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Carly Richmond]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt74822b3b7cb2a41a/6a17f6b97f6f156288c09cc7/0d4736d10fa3e12e6233cd59993299c7bd48911b-680x450.png" length="0" type="image/png"/>
    <pubDate>Wed, 07 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[了解 Elasticsearch 评分和解释 API]]></title>
    <description><![CDATA[了解 Elasticsearch 的评分机制与实用评分函数，借助 Explain API 检查搜索相关性并提升文档排名效果。]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch 是一个功能强大的搜索引擎，通过计算索引中每个文档的得分，提供快速、相关的搜索结果。这个分数是决定搜索结果排序的关键因素。在本文中，我们将深入探讨 Elasticsearch 的评分机制，并探索 Explain API，这有助于理解评分过程。</p><h2>Elasticsearch 中的评分机制</h2><p>Elasticsearch 默认使用一种名为实用评分函数 (BM25) 的评分模型。该模型以概率信息检索理论为基础，考虑了术语频率、反向文档频率和字段长度规范化等因素。让我们简要讨论一下这些因素：</p><ol><li><p><strong>术语频率 (TF)：</strong>它表示术语在文档中出现的次数。术语频率越高，说明术语与文档之间的关系越密切。</p></li><li><p><strong>反向文档频率 (IDF)：</strong>该因子用于衡量术语在整个文档集中的重要性。出现在许多文件中的术语被认为不太重要，而出现在较少文件中的术语则被认为更重要。</p></li><li><p><strong>字段长度归一化</strong>：该因子考虑了术语所在字段的长度。较短字段的权重更大，因为在较短字段中，术语被认为更重要。</p></li></ol><h2>使用解释 API</h2><p>Elasticsearch 中的解释 API 是了解评分过程的重要工具。它详细解释了如何计算特定文件的得分。要使用解释 API，您需要向以下端点发送 GET 请求：</p>GET /&lt;index&gt;/_explain/&lt;document_id&gt;<p>在请求正文中，您需要提供想要了解评分的查询。这里有一个例子：</p>{
  "query": {
    "match": {
      "title": "elasticsearch"
    }
  }
}<p>解释 API 的回复将包括评分过程的详细分类，包括各个因素（TF、IDF 和字段长度正常化）及其对最终得分的贡献。下面是一个答复样本：</p>{
  "_index": "example_index",
  "_type": "_doc",
  "_id": "1",
  "matched": true,
  "explanation": {
    "value": 1.2,
    "description": "weight(title:elasticsearch in 0) [PerFieldSimilarity], result of:",
    "details": [
      {
        "value": 1.2,
        "description": "score(doc=0,freq=1.0 = termFreq=1.0\n), product of:",
        "details": [
          {
            "value": 2.2,
            "description": "idf, computed as log(1 + (docCount - docFreq + 0.5) / (docFreq + 0.5)) from:",
            "details": [
              {
                "value": 1,
                "description": "docFreq",
                "details": []
              },
              {
                "value": 1,
                "description": "docCount",
                "details": []
              }
            ]
          },
          {
            "value": 0.5,
            "description": "tfNorm, computed as (freq * (k1 + 1)) / (freq + k1 * (1 - b + b * fieldLength / avgFieldLength)) from:",
            "details": [
              {
                "value": 1,
                "description": "termFreq=1.0",
                "details": []
              },
              {
                "value": 1.2,
                "description": "parameter k1",
                "details": []
              },
              {
                "value": 0.75,
                "description": "parameter b",
                "details": []
              },
              {
                "value": 1,
                "description": "avgFieldLength",
                "details": []
              },
              {
                "value": 1,
                "description": "fieldLength",
                "details": []
              }
            ]
          }
        ]
      }
    ]
  }
}<p>在本例中，回复显示 1.2 分是 IDF 值（2.2）和 tfNorm 值（0.5）的乘积。详细的解释有助于了解评分因素，并有助于微调搜索相关性。</p><h2>结论</h2><p>Elasticsearch 评分是提供相关搜索结果的一个重要方面。通过了解评分机制和使用解释 API，您可以深入了解影响搜索结果的因素，并优化搜索查询以提高相关性和性能。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-scoring-and-explain-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-scoring-and-explain-api</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbe7de1872f1527e3/6a17de303e9e452974ba1374/a70c5403064d5bbceff66a17373332362227f13c-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Mon, 05 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[通过两个字段进行 Elasticsearch 搜索]]></title>
    <description><![CDATA[探索在两个字段中搜索的技巧，包括 multi_match 查询、bool 查询，以及查询时字段加权。]]></description>
    <content:encoded><![CDATA[<p>在 Elasticsearch 中跨多个字段搜索是许多应用程序的共同要求。本文将探讨通过两个字段执行搜索的高级技术，包括多重匹配查询、布尔查询和查询时字段提升。这些技术将帮助您为用户创建更准确、更相关的搜索结果。</p><h2>通过两个字段执行搜索的高级技术</h2><h3>1.多匹配查询</h3><p>多匹配查询允许您在多个字段中搜索单个查询字符串。当您要查找两个字段中任何一个包含给定查询字符串的文档时，这种方法非常有用。下面是一个多重匹配查询的示例，在 "标题 "或 "描述 "字段中搜索术语 "示例"：</p>{
  "query": {
    "multi_match": {
      "query": "example",
      "fields": ["title", "description"]
    }
  }
}<h3>2.布尔查询</h3><p>bool 查询允许您使用布尔逻辑组合多个查询。您可以使用 "should "子句搜索两个字段中任何一个与查询匹配的文档。下面是一个在 "标题 "和 "描述 "字段中搜索术语 "示例 "的 bool 查询示例：</p>{
  "query": {
    "bool": {
      "should": [
        {"match": {"title": "example"}},
        {"match": {"description": "example"}}
      ]
    }
  }
}<h3>3.查询时间字段增强</h3><p>有时，您可能希望在搜索过程中更重视一个字段而不是另一个字段。您可以在查询时对字段应用提升因子来实现这一点。提升值越高，该字段的权重就越大，从而更有可能影响最终搜索得分。下面是一个多匹配查询的示例，在 "标题 "字段中应用了提升因子：</p>{
  "query": {
    "multi_match": {
      "query": "example",
      "fields": ["title^3", "description"]
    }
  }
}<p>在这个例子中，"标题 "字段的提升系数为 3，因此在决定搜索得分时，它比 "描述 "字段重要三倍。</p><h3>4.组合使用不同提升因子的查询</h3><p>您还可以使用 bool 查询将具有不同提升因子的多个查询组合起来。这样，您就可以微调搜索结果中每个字段的重要性。下面是一个 bool 查询的示例，"标题 "和 "描述 "字段应用了不同的提升因子：</p>{
  "query": {
    "bool": {
      "should": [
        {"match": {"title": {"query": "example", "boost": 3}}},
        {"match": {"description": {"query": "example", "boost": 1}}}
      ]
    }
  }
}<p>在这个例子中，"标题 "字段的提升因子为 3，而 "描述 "字段的提升因子为 1。</p><h2>结论</h2><p>在 Elasticsearch 中，可以使用多匹配查询、布尔查询和查询时字段增强等高级技术实现两个字段的搜索。通过结合这些技术，您可以为用户创建更准确、更相关的搜索结果。尝试使用不同的查询组合和提升因素，为您的特定使用情况找到最佳搜索配置。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-search-by-two-fields</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-search-by-two-fields</guid>
    <category><![CDATA[基础功能]]></category>
    <category><![CDATA[查询 DSL]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltda47d75430c4fa7c/6a17f5cae3179149242d5963/d5d04bbcfc3925f48f3487ea4c7e0dd2205316d0-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 30 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何在使用案例中实施更好的二进制量化 (BBQ)]]></title>
    <description><![CDATA[探讨为什么要在用例中实施更好的二进制量化 (BBQ) 以及如何实施。]]></description>
    <content:encoded><![CDATA[<p>矢量搜索为实现文本的语义搜索或图像、视频或音频的相似性搜索提供了基础。在矢量搜索中，矢量是数据的数学表示，可能非常庞大，有时也会比较迟钝。更好的二进制量化（以下简称 BBQ）是一种矢量压缩方法。它可以让你找到正确的匹配，同时缩小矢量，使搜索和处理速度更快。本文将介绍 BBQ 和 rescore_vector，这是一个仅适用于量化索引的字段，可自动对向量重新评分。</p><p>本文中提到的所有完整查询和输出都可以在我们的<a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/how-and-why-bbq">Elasticsearch Labs 代码库中</a>找到。</p><h2>为什么要在使用案例中实施更好的二进制量化 (BBQ)？</h2>注：要深入了解 BBQ 背后的数学原理，请查看下面的<a href="https://www.elastic.co/cn/search-labs/blog/bbq-implementation-into-use-case#further-learning">"进一步学习 "部分</a>。就本博客而言，重点是实施。<p>数学知识固然耐人寻味，但要想完全掌握矢量搜索保持精确的原因，这一点至关重要。归根结底，这一切都与压缩有关，因为事实证明，目前的矢量搜索算法受到数据读取速度的限制。因此，如果能将所有数据都存储到内存中，那么与从存储设备中读取数据相比，速度将得到显著提升 （内存的 读取<a href="https://sre.google/static/pdf/rule-of-thumb-latency-numbers-letter.pdf"> 速度约为固态硬盘的 200 倍</a> ）。</p><p>有几点需要注意：</p><ul><li><p>基于图形的索引，如<a href="https://arxiv.org/pdf/1603.09320">HNSW</a>（层次导航小世界），对于向量检索来说是最快的。</p><ul><li><p>HNSW：一种近似近邻搜索算法，可构建多层图结构，从而实现高效的高维相似性搜索。</p></li></ul></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt760bd95c206bfa8f/6a17e2ad505ac393f7ad8a95/590f3b3c72a76023a38a0436cd9ff90a9f80e936-1964x1262.png" alt="HNSW：一种近似近邻搜索算法，可构建多层图结构，从而实现高效的高维相似性搜索。" /><ul><li><p>从根本上说，HNSW 的速度受限于从内存读取数据的速度，或者在最糟糕的情况下，受限于从存储器读取数据的速度。</p><ul><li><p>理想情况下，您希望能够将所有存储的向量加载到内存中。</p></li></ul></li><li><p>嵌入模型通常以 float32 的精度生成向量，每个浮点数 4 个字节。</p></li><li><p>最后，根据向量和/或维数的多少，内存很快就会不够存放所有向量。</p></li></ul><p>如果把这看作是理所当然的，那么一旦你开始摄入数百万甚至数十亿的向量，每个向量都可能有数百甚至数千个维度，你就会发现问题很快就出现了。题为 "<a href="https://www.elastic.co/cn/search-labs/blog/bbq-implementation-into-use-case#approximate-numbers-on-the-compression-ratios">压缩比近似值</a>"的部分提供了一些粗略的数字。</p><h2>开始需要什么？</h2><p>要开始使用，您需要具备以下条件：</p><ul><li><p>如果使用 Elastic Cloud 或内部部署，则需要高于 8.18 的 Elasticsearch 版本。虽然 BBQ 是在 8.16 中引入的，但在本文中，您将使用<code>vector_rescore</code> ，它是在 8.18 中引入的。</p></li><li><p>此外，您还需要确保集群中有一个<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/ml-settings.html">机器学习（ML）节点</a>。(注意：加载模型需要至少 4GB 的 ML 节点，但如果要完成生产工作负载，可能需要更大的节点）。</p></li><li><p>如果使用的是无服务器，则需要选择针对向量进行了优化的实例。</p></li><li><p>您还需要具备矢量数据库方面的基础知识。如果您还不熟悉 Elastic 中的矢量搜索概念，可能需要先查看以下资源：</p><ul><li><p><a href="https://www.elastic.co/cn/search-labs/blog/elastic-vector-database-practical-example">导航弹性矢量数据库</a></p></li><li><p><a href="https://www.elastic.co/cn/blog/retrieval-augmented-generation-explained">检索增强生成背后的重大理念</a></p></li></ul></li></ul><h2>更好的二进制量化 (BBQ) 实现</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt18df00df95ff2ca7/6a17e2af414c6411989450df/4d388078495566f0527e931e0c2e38facdce83c6-1503x748.png" alt="Elasticsearch bbq 实现。" /><p>为了使本博客简单明了，您将在可用时使用内置函数。在这种情况下，<a href="https://www.elastic.co/cn/guide/en/machine-learning/8.17/ml-nlp-e5.html"><code>.multilingual-e5-small</code></a> 向量嵌入模型将直接在 Elasticsearch 内部的机器学习节点上运行。请注意，您可以用自己选择的嵌入器<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/infer-service-openai.html">（OpenAI</a>、<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/infer-service-google-ai-studio.html">Google AI Studio</a>、<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/infer-service-cohere.html">Cohere</a>等）替换<code>text_embedding</code> 模型。如果您喜欢的模型尚未集成，您也可以<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/bring-your-own-vectors.html">自带密集向量嵌入</a>模型）。</p><p>首先，您需要创建一个推理端点，为给定文本生成向量。您将从 Kibana<a href="https://www.elastic.co/cn/guide/en/kibana/8.18/console-kibana.html">Dev Tools 控制台</a>运行所有这些命令。该命令将下载<code>.multilingual-e5-small</code>.如果端点还不存在，它将为您设置端点；这可能需要一分钟的时间。你可以在 Outputs 文件夹中的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/01-create-an-inference-endpoint-output.json">01-create-an-inference-endpoint-output.json</a>文件中看到预期输出。 </p>PUT _inference/text_embedding/my_e5_model
{
  "service": "elasticsearch",
  "service_settings": {
    "num_threads": 1,
    "model_id": ".multilingual-e5-small",
    "adaptive_allocations": {
      "enabled": true,
      "min_number_of_allocations": 1
    }
  }
}<p>返回后，模型就设置好了，您可以使用以下命令测试模型是否按预期运行。你可以在 Outputs 文件夹中的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/02-embed-text-output.json">02-embed-text-output.json</a>文件中看到预期输出。</p>POST _inference/text_embedding/my_e5_model
{
  "input": "my awesome piece of text"
}<p>如果遇到训练好的模型没有分配到任何节点的问题，可能需要手动启动模型。</p>POST _ml/trained_models/.multilingual-e5-small/deployment/_start<p>现在，让我们创建一个带有 2 个属性的新映射，一个标准文本字段 (<code>my_field</code>) 和一个 384 维的密集矢量字段 (<code>my_vector</code>) ，以匹配嵌入模型的输出。您还可以覆盖<code>index_options.type to bbq_hnsw</code> 。你可以在 Outputs 文件夹中的文件<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/03-create-byte-qauntized-index-output.json">03-create-byte-qauntized-index-output.json</a>中看到预期输出。</p>PUT bbq-my-byte-quantized-index
{
  "mappings": {
    "properties": {
      "my_field": {
        "type": "text"
      },
      "my_vector": {
        "type": "dense_vector",
        "dims": 384,
        "index_options": {
          "type": "bbq_hnsw"
        }
      }
    }
  }
}<p>要确保 Elasticsearch 生成向量，可以使用<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/ingest.html">Ingest Pipeline</a>。该管道需要三样东西：端点 (<code>model_id</code>)、要为其创建向量的<code>input_field</code> 以及用于存储这些向量的<code>output_field</code> 。下面的第一条命令将创建推理摄取管道，该管道使用<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/inference-apis.html">推理服务 </a>，第二条命令将测试管道是否正常工作。你可以在 Outputs 文件夹中的文件<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/04-create-and-simulate-ingest-pipeline-output.json">04-create and-simulate-ingest-pipeline-output.json</a>中看到预期输出。 </p>PUT _ingest/pipeline/my_inference_pipeline
{
  "processors": [
    {
      "inference": {
        "model_id": "my_e5_model",
        "input_output": [
          {
            "input_field": "my_field",
            "output_field": "my_vector"
          }
        ]
      }
    }
  ]
}

POST _ingest/pipeline/my_inference_pipeline/_simulate
{
  "docs": [
    {
      "_source": {
        "my_field": "my awesome text field"
      }
    }
  ]
}<p>现在，您可以使用下面的前 2 个命令添加一些文档，并使用第 3 个命令测试搜索是否有效。你可以在 Outputs 文件夹中的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/05-bbq-index-output.json">05-bbq-index-output.json</a>文件中查看预期输出。 </p>PUT bbq-my-byte-quantized-index/_doc/1?pipeline=my_inference_pipeline
{
    "my_field": "my awesome text field"
}

PUT bbq-my-byte-quantized-index/_doc/2?pipeline=my_inference_pipeline
{
    "my_field": "some other sentence"
}

GET bbq-my-byte-quantized-index/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "knn": {
            "field": "my_vector",
            "query_vector_builder": {
              "text_embedding": {
                "model_id": "my_e5_model",
                "model_text": "my awesome search field"
              }
            },
            "k": 10,
            "num_candidates": 100
          }
        }
      ]
    }
  },
  "_source": [
    "my_field"
  ]
}<p>正如<a href="https://www.elastic.co/cn/search-labs/blog/better-binary-quantization-lucene-elasticsearch#lucene-benchmarking">本文章</a>所建议的，当您扩展到非数量级的数据时，建议使用重采样和超采样，因为它们有助于在受益于压缩优势的同时保持较高的召回准确率。从 Elasticsearch 8.18 版开始，您可以使用<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.18/knn-search.html#dense-vector-knn-search-rescoring">rescore_vector</a> 这样做。预期输出在 Outputs 文件夹中的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/06-bbq-search-8-18-output.json">06-bbq-search-8-18-output.json</a>文件中。</p>GET bbq-my-byte-quantized-index/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "knn": {
            "field": "my_vector",
            "query_vector_builder": {
              "text_embedding": {
                "model_id": "my_e5_model",
                "model_text": "my awesome search field"
              }
            },
            "rescore_vector": {
              "oversample": 3
            },
            "k": 10,
            "num_candidates": 100
          }
        }
      ]
    }
  },
  "_source": [
    "my_field"
  ]
}<p>这些分数与原始数据的分数相比如何？如果您再次进行上述操作，但使用<code>index_options.type: hnsw</code> ，您会发现得分非常接近。你可以在 Outputs 文件夹中的<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/how-and-why-bbq/Outputs/07-raw-vector-output.json">07-raw-vector-output.json</a>文件中看到预期输出。</p>PUT my-raw-vector-index
{
  "mappings": {
    "properties": {
      "my_field": {
        "type": "text"
      },
      "my_vector": {
        "type": "dense_vector",
        "dims": 384,
        "index_options": {
          "type": "hnsw"
        }
      }
    }
  }
}

PUT my-raw-vector-index/_doc/1?pipeline=my_inference_pipeline
{
    "my_field": "my awesome text field"
}

PUT my-raw-vector-index/_doc/2?pipeline=my_inference_pipeline
{
    "my_field": "some other sentence"
}

GET my-raw-vector-index/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "knn": {
            "field": "my_vector",
            "query_vector_builder": {
              "text_embedding": {
                "model_id": "my_e5_model",
                "model_text": "my awesome search field"
              }
            },
            "k": 10,
            "num_candidates": 100
          }
        }
      ]
    }
  },
  "_source": [
    "my_field"
  ]
}<h2>压缩比的近似值</h2><p>在使用矢量搜索时，存储和内存需求很快就会成为一项重大挑战。下面的细目说明了不同的量化技术如何显著减少矢量数据的内存占用。</p><p>向量 (V)</p><p>尺寸（D）</p><p>未加工（V x D x 4）</p><p>int8 (V x (D x 1 + 4))</p><p>int4 (V x (D x 0.5 + 4))</p><p>bbq (V x (D x 0.125 + 4))</p><p>10,000,000</p><p>384</p><p>14.31GB</p><p>3.61GB</p><p>1.83GB</p><p>0.58GB</p><p>50,000,000</p><p>384</p><p>71.53GB</p><p>18.07GB</p><p>9.13GB</p><p>2.89GB</p><p>100,000,000</p><p>384</p><p>143.05GB</p><p>36.14GB</p><p>18.25GB</p><p>5.77GB</p><h2>结论</h2><p>BBQ 是一种优化方法，可用于压缩矢量数据而不影响精度。它的工作原理是将向量转换为比特，让您能够有效地搜索数据，并使您能够扩展人工智能工作流程，加快搜索速度并优化数据存储。</p><h2>进一步学习</h2><p>如果您想了解有关烧烤的更多信息，请务必查看以下资源：</p><ul><li><p><a href="https://www.elastic.co/cn/search-labs/blog/better-binary-quantization-lucene-elasticsearch">Lucene 和 Elasticsearch 中的二进制量化 (BBQ)</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/bit-vectors-elasticsearch-bbq-vs-pq">更好的二进制量化（BBQ）与乘积量化比较</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/optimized-scalar-quantization-elasticsearch">优化的标量量化更好的二进制量化</a></p></li><li><p><a href="https://www.youtube.com/watch?v=04NzMt2Nigc">更好的二进制量化 (BBQ)：从字节到烧烤，更好的矢量搜索的秘密》，本-特伦特著</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/bbq-implementation-into-use-case</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/bbq-implementation-into-use-case</guid>
    <category><![CDATA[向量数据库]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Sachin Frayne,Jessica Garson]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3dd0495b536b2615/6a17e2b0414c6488459450e3/66842055367cdd795532b01c167f2a4b03dc65e3-1200x628.png" length="0" type="image/png"/>
    <pubDate>Wed, 23 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch 堆大小使用情况和 JVM 垃圾收集]]></title>
    <description><![CDATA[探索 Elasticsearch 堆大小使用和 JVM 垃圾收集，包括最佳实践以及如何在堆内存使用率过高或 JVM 性能不佳时解决问题。]]></description>
    <content:encoded><![CDATA[<p>堆大小是指分配给 Elasticsearch 节点 Java 虚拟机的 RAM 容量。</p><p>从 7.11 版开始，Elasticsearch 默认会根据节点的角色和总内存自动设置 JVM 堆大小。建议大多数生产环境使用默认大小。不过，如果要手动设置 JVM 堆大小，一般来说，应将 -Xms 和 -Xmx 设置为相同的值，即总可用内存的 50% ，最大（约）为 31GB。</p><p>堆大小越大，节点用于索引和搜索操作的内存就越大。不过，节点也需要内存进行缓存，因此使用 50% 可以在两者之间保持健康的平衡。出于同样的原因，在生产中应避免在 Elasticsearch 的同一节点上使用其他内存密集型进程。</p><p>通常情况下，堆使用量会呈现锯齿状，在最大堆使用量的 30 到 70% 之间摇摆。这是因为 JVM 会稳步增加堆使用百分比，直到垃圾回收过程再次释放内存。当垃圾回收进程跟不上时，就会出现堆使用率高的情况。堆使用率高的一个指标是垃圾回收无法将堆使用率降低到 30% 左右。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt03908d8eea824755/6a17dbe63e03d71e314f2b3e/0a17a67cc589a3c1fbf9e918eadc119df7bd7619-858x278.png" alt="" /><p>在上图中，您可以看到 JVM 堆的正常锯齿形。</p><p>你还会看到有两种类型的垃圾回收，即年轻的 GC 和年老的 GC。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0d527a7905c78a45/6a17dbe84b055d09484320c2/8df5c24c4894404de4617be7a13683c9027d607d-875x281.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt681db7f60d9dbe40/6a17dbe97f6f152b9bc099f0/e01eb2537310b052580411153b8eddc187d97687-890x264.png" alt="" /><p>在健康的 JVM 中，垃圾回收最好能满足以下条件：</p><ul><li><p>年轻 GC 的处理速度很快（50 毫秒内）。</p></li><li><p>年轻的 GC 执行频率不高（约 10 秒）。</p></li><li><p>旧 GC 处理速度很快（1 秒内）。</p></li><li><p>旧 GC 执行频率不高（每 10 分钟或更长时间执行一次）。</p></li></ul><h3><strong>如何解决堆内存使用率过高或 JVM 性能不佳的问题</strong></h3><p>堆内存使用量增加有多种原因：</p><h4><strong>过度包庇</strong></h4><p>请<a href="https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/size-shards#sizing-shard-guidelines">点击此处</a>查看有关过度储藏的文件。</p><h4><strong>聚合规模大</strong></h4><p>为了避免聚合大小过大，请尽量减少查询中聚合桶的数量（大小）。</p>GET /_search
{
   "aggs" : {
       "products" : {
           "terms" : {
               "field" : "product",
               "size" : 5
                          }
       }
   }
}<p>您可以使用慢查询日志（慢日志），并通过以下方式在特定索引上实施。</p>PUT /my_index/_settings
{
   "index.search.slowlog.threshold.query.warn": "10s",
   "index.search.slowlog.threshold.query.info": "5s",
   "index.search.slowlog.threshold.query.debug": "2s",
   "index.search.slowlog.threshold.query.trace": "500ms",
   "index.search.slowlog.threshold.fetch.warn": "1s",
   "index.search.slowlog.threshold.fetch.info": "800ms",
   "index.search.slowlog.threshold.fetch.debug": "500ms",
   "index.search.slowlog.threshold.fetch.trace": "200ms",
   "index.search.slowlog.level": "info"
}<p>需要很长时间才能返回结果的查询可能是资源密集型查询。</p><h4><strong>批量索引大小过大</strong></h4><p>如果您发送的请求很大，那么这可能是堆消耗大的原因。尝试减少批量索引请求的大小。</p><h4><strong>制图问题</strong></h4><p>特别是，如果您使用 "fielddata: true"，那么这将成为 JVM 堆的主要用户。</p><h4><strong>堆大小设置错误</strong></h4><p>堆大小可以通过以下方式手动定义：</p><p>设置环境变量</p>ES_JAVA_OPTS="-Xms2g -Xmx2g"<p>编辑 Elasticsearch 配置目录中的 jvm.options 文件：</p>-Xms2g
-Xmx2g<p>环境变量设置优先于文件设置。</p><p>必须重新启动节点才能将设置考虑在内。</p><h4><strong>JVM 新比率设置错误</strong></h4><p>通常无需设置，因为 Elasticsearch 默认设置此值。该参数定义 JVM 中 "新一代 "和 "老一代 "对象的可用空间比例。</p><p>如果发现旧 GC 变得非常频繁，可以尝试在 Elasticsearch 配置目录下的 jvm.options 文件中专门设置该值。</p>-XX:NewRatio=3<h3><strong>在大型 Elasticsearch 集群中，管理堆大小使用和 JVM 垃圾收集的最佳实践是什么？</strong></h3><p>在大型 Elasticsearch 集群中，管理堆大小使用和 JVM 垃圾收集的最佳做法是确保堆大小最多设置为可用 RAM 的 50% ，并根据具体使用情况优化 JVM 垃圾收集设置。必须监控堆大小和垃圾回收指标，以确保群集以最佳状态运行。具体来说，监控 JVM 堆大小、垃圾收集时间和垃圾收集暂停非常重要。此外，监控垃圾回收周期的次数和垃圾回收所花费的时间也很重要。通过监控这些指标，可以发现堆大小或垃圾回收设置方面的任何潜在问题，并在必要时采取纠正措施。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-heap-size-jvm-garbage-collection</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-heap-size-jvm-garbage-collection</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc58290fbc9f4efb9/6a1705f97d8d67cae970e632/b162c28623b9070fd1980bcd891b9dd1e868f2f0-720x421.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 22 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何增加 Elasticsearch 中的主分区数量]]></title>
    <description><![CDATA[了解如何使用 split 和 reindex API 在 Elasticsearch 中增加主分片数量，以实现最佳分片扩展。]]></description>
    <content:encoded><![CDATA[<p>无法增加现有索引的主分区数，这意味着如果要增加主分区数，必须重新创建索引。在这种情况下，通常使用两种方法：_reindex API 和 _split API。</p><p>与 _reindex API 相比，_split API 通常是更快的方法。在进行这两项操作前，<strong> 必须停止</strong><strong> 索引 ，否则源索引和目标索引的文档计数将不同。</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt46dd6abe0e6fe1eb/6a17e368148009d6a7b486d3/aa0ae010c2f5691ca00440fb453ed6b47bacd24f-1200x628.png" alt="通过重新创建索引来增加 Elasticsearch 中的分片数量" /><h2>方法 1 - 使用拆分 API</h2><p>拆分 API 用于通过复制设置和映射现有索引来创建一个新索引，其中包含所需的主分片数量。可在创建过程中设置所需的主分区数量。在实施拆分 API 之前，应检查以下设置：</p><ol><li><p>源索引必须是只读的。这意味着需要停止索引进程。</p></li><li><p>目标索引中主分区的数量必须是源索引中主分区数量的倍数。例如，如果源索引有 5 个主分区，那么目标索引的主分区可以设置为 10、15、20，以此类推。</p></li></ol><p>注意：如果只需要更改主分区编号，则首选拆分 API，因为它比 Reindex API 快得多。</p><h3>实施拆分应用程序接口</h3><p>创建测试索引：</p>POST test_split_source/_doc
{
  "test": "test"
}<p>源索引必须是只读的，才能进行拆分：</p>PUT test_split_source/_settings
{
  "index.blocks.write": true
}<p>设置和映射将自动从源索引中复制：</p>POST /test_split_source/_split/test_split_target
{
  "settings": {
    "index.number_of_shards": 3
  }
}<p>您可以通过以下方式检查进度：</p>GET _cat/recovery/test_split_target?v&amp;h=index,shard,time,stage,files_percent,files_total<p>由于设置和映射是从源索引中复制的，因此目标索引是只读的。让我们启用目标索引的写操作：</p>PUT test_split_target/_settings
{
    "index.blocks.write": null
}<p>删除原始索引前，检查源索引和目标索引的 docs.count：</p>GET _cat/indices/test_split*?v&amp;h=index,pri,rep,docs.count<p>索引名称和别名不能相同。您需要删除源索引，并将源索引名称作为别名添加到目标索引中：</p>DELETE test_split_source
PUT /test_split_target/_alias/test_split_source<p>将<strong>test_split_source</strong>别名添加到<strong>test_split_target</strong>索引后，应使用</p>GET test_split_source
POST test_split_source/_doc
{
  "test": "test"
}<h2>方法 2 - 使用重新索引 API</h2><p>通过使用 Reindex API 创建新索引，可以给出任意数量的主分区计数。在创建具有预定主分片数量的新索引后，源索引中的所有数据都可以重新索引到这个新索引。</p><p>除了拆分 API 功能外，还可以使用 reindex AP 中的 ingest_pipeline 对数据进行操作。通过摄取管道，只有符合过滤器的指定字段才会被索引到使用查询的目标索引中。数据内容可通过简单的脚本进行更改，多个索引可合并为一个索引。</p><h3>实施重新索引 API</h3><p>创建测试重新索引：</p>POST test_reindex_source/_doc
{
    "test": "test"
}<p>从源索引中复制设置和映射：</p>GET test_reindex_source<p>创建包含设置、映射和所需分片数的目标索引：</p>PUT test_reindex_target
{
  "mappings" : {},
  "settings": {
    "number_of_shards": 10,
    "number_of_replicas": 0,
    "refresh_interval": -1
  }
}<p>*注意：设置 number_of_replicas：0 和 refresh_interval: -1 会提高重新索引的速度。</p><p>启动重新索引程序。设置 requests_per_second=-1 和 slices=auto 可以调整重新索引的速度。</p>POST _reindex?requests_per_second=-1&amp;slices=auto&amp;wait_for_completion=false
{
  "source": {
    "index": "test_reindex_source"
  },
  "dest": {
    "index": "test_reindex_target"
  }
}<p>运行 reindex API 时，您将看到 task_id。复制并使用 _tasks API 进行检查：</p>GET _tasks/&lt;task_id&gt;<p>重新索引完成后更新设置：</p>PUT test_reindex_target/_settings
{
  "number_of_replicas": 1,
  "refresh_interval": "1s"
}<p>删除原始索引前，请检查源索引和目标索引的 docs.count，它们应该是一样的：</p>GET _cat/indices/test_reindex_*?v&amp;h=index,pri,rep,docs.count<p>索引名称和别名不能相同。删除源索引，并将源索引名称作为别名添加到目标索引中：</p>DELETE test_reindex_source
PUT /test_reindex_target/_alias/test_reindex_source<p>将 test_split_source 别名添加到 test_split_target 索引后，使用</p>GET test_reindex_source<h2>总结</h2><p>如果要增加现有索引的主分区数，则需要重新创建新索引的设置和映射。主要有两种方法：重新索引 API 和拆分 API。在使用这两种方法之前，必须先停止主动索引。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-increase-primary-shard-count</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-increase-primary-shard-count</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta8aa774fc00d7233/6a17e223dbb4ff68b3fb5611/7034b76019a0cba52c25eda29fceb18afc96ed0b-720x420.png" length="0" type="image/png"/>
    <pubDate>Thu, 17 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何在不同版本的 Elasticsearch& 集群之间迁移数据]]></title>
    <description><![CDATA[探索在 Elasticsearch 版本和集群之间传输数据的方法。]]></description>
    <content:encoded><![CDATA[<p>当你想升级 Elasticsearch 集群时，有时创建一个新的、独立的集群并将数据从旧集群转移到新集群会更容易。这样做的好处是，用户可以在新的集群上测试其所有数据和配置以及所有应用程序，而不会有任何停机或数据丢失的风险。</p><p>这种方法的缺点是需要一些重复的硬件，在试图顺利传输和同步所有数据时可能会造成困难。</p><p>如果需要将应用程序从一个数据中心迁移到另一个数据中心，也可能需要执行类似的程序。</p><p>本文将讨论并详细介绍在 Elasticsearch 集群之间传输数据的三种方法。</p><p><strong>如何在 Elasticsearch 集群之间迁移数据？</strong></p><p>在 Elasticsearch 集群之间传输数据有 3 种方法：</p><ol><li><p><a href="https://www.elastic.co/cn/search-labs/blog/elasticsearch-migrate-data-versions-clusters#1.-reindexing-data-from-a-remote-cluster">从远程群集重新编排索引</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/elasticsearch-migrate-data-versions-clusters#2.-transferring-data-using-snapshots">使用快照传输数据</a></p></li><li><p><a href="https://www.elastic.co/cn/search-labs/blog/elasticsearch-migrate-data-versions-clusters#3.-transferring-data-using-logstash">使用 Logstash 传输数据</a></p></li></ol><p>使用快照通常是最快速、最可靠的数据传输方法。不过，请记住，您只能将快照还原到相同或更高版本的群集上，而不能还原到相差一个主版本以上的群集上。这意味着您可以将 6.x 快照还原到 7.x 群集上，但不能还原到 8.x 群集上。</p><p>如果需要增加一个以上的主要版本，则需要重新索引或使用 Logstash。</p><p>现在，让我们分别详细了解在 Elasticsearch 集群之间传输数据的三个选项。</p><h2>1.从远程群集重新索引数据</h2><p>在开始重新索引之前，请记住您需要为新群集上的所有索引设置适当的映射。为此，您必须使用适当的映射直接创建索引，或者使用索引模板。</p><h3>从远程重新索引 - 需要配置</h3><p>要从远程重新索引，应将以下配置添加到接收数据的群集的 elasticseearch.yml 文件中，在 Linux 系统中，该文件通常位于此处：/etc/elasticsearch/elasticsearch.yml。要添加的配置如下：</p>reindex.remote.whitelist: "192.168.1.11:9200"<p>如果使用 SSL，则应将 CA 证书添加到每个节点，并在 elasticsearch.yml 中每个节点的命令中包含以下内容：</p>reindex.ssl.certificate_authorities: “/path/to/ca.pem”<p>或者，也可以在所有 Elasticsearch 节点上添加下面一行，以禁用 SSL 验证。不过，由于这种方法不如前一种方法安全，因此不太推荐使用：</p>reindex.remote.whitelist: "192.168.1.11:9200"
reindex.ssl.verification_mode: none
systemctl restart elasticsearch service <p>您需要在每个节点上进行这些修改，并进行滚动重启。有关如何操作的详细信息，请参阅<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/8.17/restart-cluster.html#restart-cluster-rolling">我们的指南</a>。</p><h3>重编索引命令</h3><p>在 elasticsearch.yml 文件中定义远程主机，并在必要时添加 SSL 证书后，就可以使用下面的命令开始重新索引数据了：</p>POST _reindex
{
  "source": {
    "remote": {
      "host": "http://192.168.1.11:9200",
      "username": "elastic",
      "password": "123456",
     "socket_timeout": "1m",
      "connect_timeout": "1m"

    },
    "index": "companydatabase"
  },
  "dest": {
    "index": "my-new-index-000001"
  }
}<p>在此过程中，您可能会遇到超时错误，因此最好为超时设置宽松的值，而不是依赖默认值。</p><p>现在，让我们来看看从远程重新索引时可能会遇到的其他一些常见错误。</p><h3>从远程重新索引时的常见错误</h3><h4>1.重新索引未列入白名单</h4>{
  "error": {
    "root_cause": [
      {
        "type": "illegal_argument_exception",
        "reason": "[192.168.1.11:9200] not whitelisted in reindex.remote.whitelist"
      }
    ],
    "type": "illegal_argument_exception",
    "reason": "[192.168.1.11:9200] not whitelisted in reindex.remote.whitelist"
  },
  "status": 400
}<p>如果遇到此错误，则表明您没有按照上文所述在 Elasticsearch 中定义远程主机 IP 地址或节点名称 DNS，或者忘记重启 Elasticsearch 服务。</p><p>要解决 Elasticsearch 集群的问题，需要将远程主机添加到所有 Elasticsearch 节点，并重新启动 Elasticsearch 服务。</p><h4>2.SSL 握手异常</h4>{
  "error": {
    "root_cause": [
      {
        "type": "s_s_l_handshake_exception",
        "reason": "PKIX path building failed: sun.security.provider.certpath.SunCertPathBuilderException: unable to find valid certification path to requested target"
      }
    ],
    "type": "s_s_l_handshake_exception",
    "reason": "PKIX path building failed: sun.security.provider.certpath.SunCertPathBuilderException: unable to find valid certification path to requested target",
    "caused_by": {
      "type": "validator_exception",
      "reason": "PKIX path building failed: sun.security.provider.certpath.SunCertPathBuilderException: unable to find valid certification path to requested target",
      "caused_by": {
        "type": "sun_cert_path_builder_exception",
        "reason": "unable to find valid certification path to requested target"
      }
    }
  },
  "status": 500
}<p>这个错误说明，你忘记按上文所述在 elasticsearch.yml 中添加 reindex.ssl.certificate_authorities。加进去：</p>#elasticsearch.yml
reindex.ssl.certificate_authorities: "/path/to/ca.pem"<h2>2.使用快照传输数据</h2><p>请记住，如上所述，您只能将快照还原到相同或更高版本的群集上，而不能还原到相差一个主要版本以上的群集上。</p><p>如果需要增加一个以上的主要版本，则需要重新索引或使用 Logstash。</p><p>通过快照传输数据需要以下步骤：</p><p>步骤 1.将版本库插件添加到第一个 Elasticsearch 群集--为了通过快照在群集间传输数据，需要确保新旧群集都能访问版本库。AWS、Google 和 Azure 等云存储库通常是理想的选择。要拍摄快照，请参阅<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/snapshot-restore.html">我们的指南</a>，并按照其中描述的步骤操作。</p><p>步骤 2.重启 Elasticsearch 服务（滚动重启）。</p><p>第 3 步为第一个 Elasticsearch 集群创建一个存储库。</p><p>步骤 4- 将版本库插件添加到第二个 Elasticsearch 集群。</p><p>第 5 步--向第二个 Elasticsearch 集群添加只读版本库--您需要重复创建第一个 Elasticsearch 集群的相同步骤来添加版本库。</p><p>重要提示：将第二个 Elasticsearch 集群连接到同一个 AWS S3 存储库时，应将存储库定义为只读存储库：</p>PUT _snapshot/my_s3_repository
{
  "type": "s3",
  "settings": {
    "bucket": "my-analytic-data",
    "endpoint": "s3.eu-de.cloud-object-storage.appdomain.cloud",
    "readonly": "true"
  }
}<p>这一点很重要，因为你要防止在同一个快照存储库中出现混合 Elasticsearch 版本的风险。</p><p>第 6 步 - 将数据还原到第二个 Elasticsearch 群集 - 采取上述步骤后，您可以还原数据并将其传输到新的群集。请按照<a href="https://www.elastic.co/cn/guide/en/elasticsearch/reference/current/snapshot-restore.html">本文</a>所述步骤将数据恢复到新群集。 </p><h2>3.使用 Logstash 传输数据</h2><p>在开始使用 logstash 传输数据之前，请记住您需要为新群集上的所有索引设置适当的映射。为此，您需要直接创建索引或使用索引模板。</p><p>要在两个 Elasticsearch 集群之间传输数据，可以设置一个临时 Logstash 服务器，并用它在两个集群之间传输数据。对于小型集群，2GB 内存的实例就足够了。对于较大的集群，可以使用配备 8GB 内存的四核 CPU。</p><p>有关安装 Logstash 的指导，<a href="https://www.elastic.co/cn/guide/en/logstash/current/installing-logstash.html">请参阅此处</a>。</p><h3>从一个群集向另一个群集传输数据的 Logstash 配置</h3><p>将单个索引从群集 A 复制到群集 B 的基本配置是</p>iinput
{
elasticsearch
      {
        hosts =&gt; ["192.168.1.11:9200"]
        index =&gt; "index_name"
       docinfo =&gt; true      
      }
}

output 
{
  elasticsearch {
        hosts =&gt; "https://192.168.1.12:9200"
        index =&gt; "index_name"
        
  }
}<p>对于安全的 elasticsearch，可以使用下面的配置：</p>input
{
  elasticsearch
      {
        hosts =&gt; ["192.168.1.11:9200"]
        index =&gt; "index_name"
        docinfo =&gt; true 
        user =&gt; "elastic"
        password =&gt; "elastic_password"
        ssl =&gt; true
        ssl_certificate_verification =&gt; false
            
      }
}

output 
{
  elasticsearch {
        hosts =&gt; "https://192.168.1.12:9200"
        index =&gt; "index_name"
        user =&gt; "elastic"
        password =&gt; "elastic_password"
        ssl =&gt; true
        ssl_certificate_verification =&gt; false
  }
}<h3>索引元数据</h3><p>上述命令将写入一个命名索引。如果要传输多个索引并保留索引名称，则需要在 Logstash 输出中添加以下一行：</p>index =&gt; "%{[@metadata][_index]}"<p>此外，如果您想保留文件的原始 ID，则需要添加：</p>document_id =&gt; "%{[@metadata][_id]}"<p>请注意，设置文档 ID 会大大降低数据传输的速度，因此只有在必要时才保留原始 ID。</p><h2>同步更新</h2><p>上述所有方法都需要相对较长的时间，您可能会发现在等待过程完成时，原始群集中的数据已被更新。</p><p>有多种策略可以同步数据传输过程中可能出现的任何更新，在开始这一过程之前，你应该考虑一下这些问题。您尤其需要考虑</p><ul><li><p>您有什么方法来识别自数据传输过程开始以来已更新/添加的任何数据（例如，数据中的 "最后更新时间 "字段）？</p></li><li><p>您可以使用什么方法来传输最后一个数据？</p></li><li><p>是否存在记录重复的风险？通常会有，除非您使用的方法在重新索引时将文档 ID 设置为已知值）。</p></li></ul><p>下文介绍了实现同步更新的不同方法。</p><h3>1.使用排队系统</h3><p>有些摄取/更新系统使用队列，可以 "重放 "过去 x 天内收到的数据修改。这可以为同步任何更改提供一种手段。 </p><h3>2.从远程重新索引</h3><p>对 "last_update_time"&gt; x 天前的所有项目重复重新索引过程。您可以在重新索引请求中添加一个 "查询 "参数。</p><h3>3.Logstash</h3><p>在 Logstash 输入中，可以添加一个查询来过滤所有 "last_update_time"&gt; x 天前的项目。不过，除非设置了 document_id，否则这一过程会导致非时间序列数据出现重复。</p><h3>4.快照</h3><p>不可能只恢复索引的一部分，因此必须使用上述其他数据传输方法之一（或脚本）来更新数据传输过程后发生的任何更改。</p><p>不过，快照恢复比重新索引/Logstash 快得多，因此可以在传输快照时短暂暂停更新，以完全避免问题。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-migrate-data-versions-clusters</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-migrate-data-versions-clusters</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Kofi Bartlett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc041ebca11476c6/6a16f70560084b31b93c4344/01fde3b1d714f12bf8673140c9f2f940d443de31-1440x823.jpg" length="0" type="image/jpeg"/>
    <pubDate>Mon, 14 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[如何使用我们的同义词 API 自动生成和上传同义词]]></title>
    <description><![CDATA[了解如何使用 LLM 自动识别和生成同义词，从而以编程方式将术语加载到 Elasticsearch 同义词 API 中。]]></description>
    <content:encoded><![CDATA[<p>提高搜索结果的质量对于提供高效的用户体验至关重要。优化搜索的方法之一是通过同义词自动扩展查询词。这样就可以更广泛地解释查询，涵盖各种语言，从而改进结果匹配。</p><p>本博客将探讨如何使用大型语言模型 (LLM) 自动识别和生成同义词，并允许以编程方式将这些术语加载到 Elasticsearch 的同义词 API 中。</p><h2>何时使用同义词？</h2><p>与矢量搜索相比，使用同义词是一种更快、更具成本效益的解决方案。它的实现较为简单，因为它不需要深厚的嵌入知识，也不需要复杂的矢量摄取过程。</p><p>此外，由于矢量搜索需要更大的存储容量和内存来嵌入索引和检索，因此资源消耗较低。</p><p>另一个重要方面是搜索区域化。有了同义词，就可以根据当地语言和习俗调整术语。这在嵌入式可能无法匹配区域表达或特定国家术语的情况下非常有用。例如，有些单词或缩略语在不同地区可能有不同的含义，但当地用户自然会将其视为同义词。在巴西，这种情况非常普遍。"Abacaxi" 和"ananás" 是同一种水果（菠萝），但在东北部的一些地区，第二个术语更常用。同样，东南部著名的"pão francês" 在东北部可能被称为"pão careca" 。</p><h2>如何使用 LLM 生成同义词？</h2><p>为了自动获取同义词，我们可以使用 LLM，它可以分析术语的上下文，并建议适当的变体。这种方法可以动态扩展同义词，确保搜索范围更广、更准确，而无需依赖固定词典。</p><p>在本演示中，我们将使用 LLM 生成电子商务产品的同义词。由于查询词的变化，许多搜索结果很少或没有结果。有了同义词，我们就可以解决这个问题。例如，搜索"智能手机" 可以涵盖不同型号的手机，确保用户找到所需的产品。</p><h3>准备工作</h3><p>在开始之前，我们需要设置环境并定义所需的依赖关系。我们将使用 Elastic 提供的解决方案，<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">在 Docker 中本地运行 Elasticsearch 和 Kibana</a>。代码将使用 Python 3.9.6 版本编写，依赖关系如下：</p>pip install openai==1.59.8 elasticsearch==8.15.1<h3>创建产品索引</h3><p>最初，我们将创建一个不支持同义词的产品索引。这样我们就可以验证查询，然后将其与包含同义词的索引进行比较。</p><p>为了创建索引，我们在 Kibana DevTools 中使用以下命令批量加载产品数据集：</p>POST _bulk
{"index": {"_index": "products", "_id": 10001}}
{"category": "Electronics", "name": "iPhone 14 Pro"}
{"index": {"_index": "products", "_id": 10007}}
{"category": "Electronics", "name": "MacBook Pro 16-inch"}
{"index": {"_index": "products", "_id": 10013}}
{"category": "Electronics", "name": "Samsung Galaxy Tab S8"}
{"index": {"_index": "products", "_id": 10037}}
{"category": "Electronics", "name": "Apple Watch Series 8"}
{"index": {"_index": "products", "_id": 10049}}
{"category": "Electronics", "name": "Kindle Paperwhite"}
{"index": {"_index": "products", "_id": 10067}}
{"category": "Electronics", "name": "Samsung QLED 4K TV"}
{"index": {"_index": "products", "_id": 10073}}
{"category": "Electronics", "name": "HP Spectre x360 Laptop"}
{"index": {"_index": "products", "_id": 10079}}
{"category": "Electronics", "name": "Apple AirPods Pro"}
{"index": {"_index": "products", "_id": 10115}}
{"category": "Electronics", "name": "Amazon Echo Show 10"}
{"index": {"_index": "products", "_id": 10121}}
{"category": "Electronics", "name": "Apple iPad Air"}
{"index": {"_index": "products", "_id": 10127}}
{"category": "Electronics", "name": "Apple AirPods Max"}
{"index": {"_index": "products", "_id": 10151}}
{"category": "Electronics", "name": "Sony WH-1000XM4 Headphones"}
{"index": {"_index": "products", "_id": 10157}}
{"category": "Electronics", "name": "Google Pixel 6 Pro"}
{"index": {"_index": "products", "_id": 10163}}
{"category": "Electronics", "name": "Apple MacBook Air"}
{"index": {"_index": "products", "_id": 10181}}
{"category": "Electronics", "name": "Google Pixelbook Go"}
{"index": {"_index": "products", "_id": 10187}}
{"category": "Electronics", "name": "Sonos Beam Soundbar"}
{"index": {"_index": "products", "_id": 10199}}
{"category": "Electronics", "name": "Apple TV 4K"}
{"index": {"_index": "products", "_id": 10205}}
{"category": "Electronics", "name": "Samsung Galaxy Watch 4"}
{"index": {"_index": "products", "_id": 10211}}
{"category": "Electronics", "name": "Apple MacBook Pro 16-inch"}
{"index": {"_index": "products", "_id": 10223}}
{"category": "Electronics", "name": "Amazon Echo Dot (4th Gen)"}<h3>用 LLM 生成同义词</h3><p>在这一步中，我们将使用 LLM 来动态生成同义词。为此，我们将整合 OpenAI 应用程序接口，定义适当的模型和提示。LLM 将接收产品类别和名称，确保同义词与上下文相关。</p>import json
import logging

from openai import OpenAI

def call_gpt(prompt, model):
    try:
        logging.info("generate synonyms by llm...")
        response = client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            temperature=0.7,
            max_tokens=1000
        )
        content = response.choices[0].message.content.strip()
        return content
    except Exception as e:
        logging.error(f"Failed to use model: {e}")
        return None

def generate_synonyms(category, products):
   synonyms = {}

   for product in products:
       prompt = f"You are an expert in generating synonyms for products. Based on the category and product name provided, generate synonyms or related terms. Follow these rules:\n"
       prompt += "1. **Format**: The first word should be the main item (part of the product name, excluding the brand), followed by up to 3 synonyms separated by commas.\n"
       prompt += "2. **Exclude the brand**: Do not include the brand name in the synonyms.\n"
       prompt += "3. **Maximum synonyms**: Generate a maximum of 3 synonyms per product.\n\n"
       prompt += f"The category is: **{category}**, and the product is: **{product}**. Return only the synonyms in the requested format, without additional explanations."

       response = call_gpt(prompt, "gpt-4o")
       synonyms[product] = response

   return synonyms<p>从创建的产品索引中，我们将检索"Electronics" 类别中的所有项目，并将其名称发送到 LLM。预期输出结果如下</p>{
  "iPhone 14 Pro": ["iPhone", "smartphone", "mobile", "handset"],
  "MacBook Pro 16-inch": ["MacBook", "Laptop", "Notebook", "Ultrabook"],
  "Samsung Galaxy Tab S8": ["Tab", "Tablet", "Slate", "Pad"],
  "Bose QuietComfort 35 Headphones": ["Headphones", "earphones", "earbuds", "headset"]
}<p>有了生成的同义词，我们就可以使用同义词 API 将其注册到 Elasticsearch 中。</p><h3>使用同义词 API 管理同义词</h3><p>同义词 API 提供了在系统内直接管理同义词集的有效方法。每个同义词集都由同义词规则组成，其中一组词在搜索中被视为等同词。</p><p><strong>创建同义词集示例</strong></p>PUT _synonyms/my-synonyms-set
{
  "synonyms_set": [
    {
      "id": "rule-1",
      "synonyms": "hello, hi"
    },
    {
      "synonyms": "bye, goodbye"
    }
  ]
}<p>
这样就创建了一个名为"my-synonyms-set," 的集合，其中"hello" 和"hi" 被视为等同词，"bye" 和"goodbye 也被视为等同词。"</p><h2>为产品目录创建同义词</h2><p>下面是建立同义词集并将其插入 Elasticsearch 的方法。同义词规则是根据 LLM 建议的同义词映射生成的。每条规则都有一个 ID（与 slug 格式的产品名称相对应）和 LLM 计算出的同义词列表。</p>import json
import logging

from elasticsearch import Elasticsearch
from slugify import slugify

es = Elasticsearch(
    "http://localhost:9200",
    api_key="your_api_key"
)

def mount_synonyms(results):
   synonyms_set = [{"id": slugify(product), "synonyms": synonyms} for product, synonyms in
                   results.items()]

   try:
       response = es.synonyms.put_synonym(id="products-synonyms-set",
                                                 synonyms_set=synonyms_set)

       logging.info(json.dumps(response.body, indent=4))
       return response.body
   except Exception as e:
       logging.error(f"Error create synonyms: {str(e)}")
       return None<p>下面是创建同义词集的请求有效载荷：</p>{
   "synonyms_set":[
      {
         "id": "iphone-14-pro",
         "synonyms": "iPhone, smartphone, mobile, handset"
      },
      {
         "id": "macbook-pro-16-inch",
         "synonyms": "MacBook, Laptop, Notebook, Computer"
      },
      {
         "id": "samsung-galaxy-tab-s8",
         "synonyms": "Tablet, Slate, Pad, Device"
      },
      {
         "id": "garmin-forerunner-945",
         "synonyms": "Forerunner, smartwatch, fitness watch, GPS watch"
      },
      {
         "id": "bose-quietcomfort-35-headphones",
         "synonyms": "Headphones, Earphones, Headset, Cans"
      }
   ]
}<p>在集群中创建同义词集后，我们就可以进行下一步，即使用定义的同义词集创建支持同义词的新索引。</p><p>下面是完整的 Python 代码，其中包含 LLM 生成的同义词和同义词 API 定义的同义词集创建：</p>import json
import logging

from elasticsearch import Elasticsearch
from openai import OpenAI
from slugify import slugify

logging.basicConfig(level=logging.INFO)

client = OpenAI(
   api_key="your-key",
)

es = Elasticsearch(
    "http://localhost:9200",
    api_key="your_api_key"
)


def call_gpt(prompt, model):
   try:
       logging.info("generate synonyms by llm...")
       response = client.chat.completions.create(
           model=model,
           messages=[{"role": "user", "content": prompt}],
           temperature=0.7,
           max_tokens=1000
       )
       content = response.choices[0].message.content.strip()
       return content
   except Exception as e:
       logging.error(f"Failed to use model: {e}")
       return None


def generate_synonyms(category, products):
   synonyms = {}

   for product in products:
       prompt = f"You are an expert in generating synonyms for products. Based on the category and product name provided, generate synonyms or related terms. Follow these rules:\n"
       prompt += "1. **Format**: The first word should be the main item (part of the product name, excluding the brand), followed by up to 3 synonyms separated by commas.\n"
       prompt += "2. **Exclude the brand**: Do not include the brand name in the synonyms.\n"
       prompt += "3. **Maximum synonyms**: Generate a maximum of 3 synonyms per product.\n\n"
       prompt += f"The category is: **{category}**, and the product is: **{product}**. Return only the synonyms in the requested format, without additional explanations."

       response = call_gpt(prompt, "gpt-4o")
       synonyms[product] = response

   return synonyms


def get_products(category):
   query = {
       "size": 50,
       "_source": ["name"],
       "query": {
           "bool": {
               "filter": [
                   {
                       "term": {
                           "category.keyword": category
                       }
                   }
               ]
           }
       }
   }
   response = es.search(index="products", body=query)

   if response["hits"]["total"]["value"] &gt; 0:
       product_names = [hit["_source"]["name"] for hit in response["hits"]["hits"]]
       return product_names
   else:
       return []


def mount_synonyms(results):
   synonyms_set = [{"id": slugify(product), "synonyms": synonyms} for product, synonyms in
                   results.items()]

   try:
       es_client = get_client_es()
       response = es_client.synonyms.put_synonym(id="products-synonyms-set",
                                                 synonyms_set=synonyms_set)

       logging.info(json.dumps(response.body, indent=4))
       return response.body
   except Exception as e:
       logging.error(f"Erro update synonyms: {str(e)}")
       return None


if __name__ == '__main__':
   category = "Electronics"
   products = get_products("Electronics")
   llm_synonyms = generate_synonyms(category, products)
   mount_synonyms(llm_synonyms)<h3>创建支持同义词的索引</h3><p>将创建一个新索引，对<code>products</code> 索引中的所有数据进行重新索引。该索引将使用<code>synonyms_filter</code> ，它应用了之前创建的<code>products-synonyms-set</code> 。</p><p>以下是配置为使用同义词的索引映射：</p>PUT products_02
{
  "settings": {
    "analysis": {
      "filter": {
        "synonyms_filter": {
          "type": "synonym",
          "synonyms_set": "products-synonyms-set",
          "updateable": true
        }
      },
      "analyzer": {
        "synonyms_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "synonyms_filter"
          ]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "ID": {
        "type": "long"
      },
      "category": {
        "type": "keyword"
      },
      "name": {
        "type": "text",
        "analyzer": "standard",
        "search_analyzer": "synonyms_analyzer"
      }
    }
  }
}<h3>重新索引<code>products</code> 索引</h3><p>现在，我们将使用<strong>Reindex API</strong>将<code>products</code> 索引中的数据迁移到包含同义词支持的新<code>products_02</code> 索引中。在 Kibana DevTools 中执行了以下代码：
</p>POST _reindex
{
  "source": {
    "index": "products"
  },
  "dest": {
    "index": "products_02"
  }
}<p>迁移后，<code>products_02</code> 索引将被填充，并可使用配置的同义词集验证搜索。</p><h3>使用同义词验证搜索</h3><p>让我们比较一下两个索引的搜索结果。我们将在两个索引上执行相同的查询，并验证是否使用同义词来检索结果。</p><h4>在<code>products</code> 索引中搜索（不含同义词）</h4><p>我们将使用 Kibana 执行搜索并分析结果。在分析&gt; 发现菜单中，我们将创建一个数据视图，以可视化我们创建的索引中的数据。</p><p>在 Discovery 中，单击数据视图并定义名称和索引模式。对于"<strong>产品</strong>" 索引，我们将使用"<strong>产品</strong>"模式。然后，我们将重复该过程，使用"<strong>products_02</strong><strong>"</strong>模式为"products_02" 索引创建一个新的数据视图。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte826fd932cfeb9df/6a17fdffec0f8912aa5a6841/3ad4a6891a3905e96532a312932fdf3a8216aec2-1600x599.png" alt="" /><p>配置好数据视图后，我们就可以返回 Analytics&gt; Discovery 并开始验证。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltba3729c60068e8a0/6a17fe01e9ea87ba2aa9c82a/422c4b2b51abae6580cad25085d1b8a365fc6b9e-1294x850.png" alt="" /><p>在这里，选择 DataView 产品并对"tablet" 一词进行搜索后，我们没有得到任何结果，尽管我们知道有"Kindle Paperwhite" 和"Apple iPad Air" 这样的产品。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2c6a383dd4cb0157/6a17fe02577262671d1bce0c/e4ae3a785fdd93f48d7c7d204185ded149126f2c-1600x862.png" alt="" /><h4>在<code>products_02</code> 索引中搜索（支持同义词）</h4><p>在支持同义词的"<strong>products_synonyms</strong>" 数据视图上执行相同查询时，产品被成功检索。这表明配置的同义词集工作正常，确保搜索词的不同变体都能返回预期结果。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt986b4706e2f70014/6a17fe043e9e454edbba16d3/e609749c39e90d5c82fa846af6124679dd62bcb8-1600x526.png" alt="" /><p>我们可以直接在 Kibana DevTools 中运行相同的查询来获得相同的结果。只需使用 Elasticsearch Search API 搜索 products_02 索引即可：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbe829a3c4d7aac60/6a17fe05e8fbce03d73a1bd7/504d0d1f96dcfbceb309063dc0716bcee64ad2f8-1600x870.png" alt="" /><h2>结论</h2><p>在 Elasticsearch 中使用同义词提高了产品目录搜索的准确性和覆盖范围。与众不同的关键在于使用了<strong>LLM</strong>，它可以根据上下文自动生成同义词，无需预定义清单。该模型分析了产品名称和类别，确保与电子商务相关的同义词。</p><p>此外，<strong>同义词 API</strong>简化了词典管理，允许动态修改同义词集。有了这种方法，搜索变得更加灵活，更能适应不同的用户查询模式。</p><p>这一过程可以通过新数据和模型调整不断改进，确保提供越来越高效的研究体验。</p><h2>参考资料</h2><p><strong>在本地运行 Elasticsearch</strong></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html</a></p><p><strong>同义词应用程序接口</strong></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/synonyms-apis.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/synonyms-apis.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-synonyms-automate</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-synonyms-automate</guid>
    <category><![CDATA[相关性]]></category>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0f0247b9bc1d1ccd/6a17fe07ec0f891c745a6845/05a3cfeaa387561d5334ca3f1609035ddfff7481-1200x628.png" length="0" type="image/png"/>
    <pubDate>Thu, 27 Mar 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[实时日志，顺利使用：Elasticsearch 新推出的专用 logsdb 索引模式]]></title>
    <description><![CDATA[Elasticsearch 在日志管理方面的最新创新 logsdb 最多可将日志数据的存储空间减少 65% ，使可观察性和安全团队能够在不超出预算的情况下扩大可见性，同时保持所有数据的可访问性和可搜索性。]]></description>
    <content:encoded><![CDATA[<h2>Elasticsearch 的新索引模式 logsdb 可将日志存储需求减少最多 65%</h2><p>今天，我们宣布 Elasticsearch 的新索引模式 logsdb 全面上市，与没有 logsdb 的近期版本 Elasticsearch 相比，它可<strong>将日志数据的存储空间占用减少多达 65%</strong>。这一显著改进使可观察性和安全团队能够在不超出预算的情况下扩大可观察性，同时保持所有数据可立即访问以进行分析。</p><p>Logsdb 索引模式优化了数据排序，通过合成 <code>synthetic _source</code> 实时重建非存储字段值来消除重复，并通过先进的算法和编解码器提高压缩效率，同时在 Elasticsearch 中使用列式存储以实现高效的日志存储和检索。</p><h2>通过 logsdb 索引模式提高存储效率，增强分析能力并降低成本</h2><p>日志为检测和解决可观测性和安全问题提供了关键信号，随着 AI 的进步简化了文本数据的分析，其作用日益增加，因此高效的存储和高性能的访问比以往任何时候都更为重要。</p><p>遗憾的是，基础架构和应用程序产生的日志量不断增加，导致成本上升，不得不做出一些对分析产生负面影响的妥协：限制收集、减少保留或将新数据归入孤立的存档层级。</p><p>Logsdb 直接应对这些挑战。有了更高的存储效率，您就可以收集更多数据，并避免复杂的数据筛选带来的麻烦。您可以将日志保留更长时间，以满足威胁搜寻、事件响应和合规要求。而且，由于所有数据均可搜索，因此无论数据集有多大，都能快速获得见解。</p><h2>logsdb 索引模式所依托的技术创新</h2><p>Logsdb 索引模式通过智能索引排序、合成 _source 和高级压缩显著减少了日志数据的磁盘占用。实施它可以将日志存储需求减少高达 65%，与没有 logsdb 的最新版本的 Elasticsearch 相比。尽管 logsdb 目前在索引过程中使用更多的 CPU，但其高效的存储降低了大多数客户的总体成本。对于需要长期保留的客户，我们预计总拥有成本（TCO）可降低最多 50%。</p><p><strong>智能索引排序</strong>可提高存储效率达 30% ，并通过将类似数据放在一起来减少某些日志数据集的查询延迟。默认情况下，它会按 host.name 和 @timestamp 对索引进行排序。如果您的数据有更合适的字段，您可以指定这些字段。</p><p>通过 Zstandard 压缩 (Zstd)、delta 编码、run-length 编码和其他自动选择的智能编解码器，<strong>高级压缩</strong>大大降低了日志等文本数据的存储要求。Doc-values 以专为压缩和性能优化的列格式存储，可高效存储和检索字段值，用于排序、聚合和脚本编写。</p><p>通过舍弃<strong>_source</strong>字段并按需全部或部分重建 _source 字段，合成_source 使企业能够将存储需求再减少 20-40% 。虽然该功能有时需要更多的计算来进行索引和检索，但测试表明，它带来了可衡量的净效率改进。Synthetic _source 基于近两年的生产使用指标，对日志进行了大量改进，包括支持几乎所有字段类型。</p><p>由此产生的存储节省效益会贯穿于索引的整个生命周期阶段。热层存储减少 65%，温层、冷层和冻结层的存储也会相应减少，同时还能减少在桶存储中存储快照所占用的空间。</p><h2>不影响可见性：保留所有日志，确保可观测和安全性</h2><p>日志是基础架构和应用程序可见性的基础，提供了监测和故障排除最简单和最基本的信号。然而，随着日志记录量的增长，成本也在不断上升。这一挑战迫使客户实施复杂的筛选和管理策略，过早地删除数据，并将相关日志滞留在需要一天或更长时间才能恢复成可用状态的存储区中，然后再进行分析。如果没有完整、易于搜索和可访问的数据集，查找并解决问题的难度就会大大增加。</p><p>Logsdb 索引模式以<a href="https://www.elastic.co/cn/elasticsearch/elasticsearch-searchable-snapshots">可搜索快照</a>和<a href="https://www.elastic.co/cn/blog/automatic-import-ai-data-integration-builder">自动导入</a>等突破性 Elasticsearch 功能为基础，解决了运营和安全团队的这些痛点：</p><p><strong>降低成本： </strong>Logsdb 可减少日志的存储空间占用达 65% ，使企业能够在保留更多数据的同时降低存储费用。这意味着所有存储层（从热存储到冷冻存储）都能节约成本，使用这些数据的可观察性和安全团队也能提高工作效率。</p><p><strong>保存宝贵数据： </strong>Logsdb 可保存所有日志数据并提高运行效率，而无需依赖额外的工具或复杂的过滤器。利用合成 _source 等功能，无需存储整个源文件即可保留数据的价值。</p><p><strong>扩大可见性：</strong>Logsdb 可在一个平台上高效访问所有数据，而无需为可观察性、安全性和历史数据建立单独的孤岛。对于站点可靠性工程师（SRE）来说，它可以在分析日志的同时分析指标、跟踪和业务数据，从而加快问题的解决。同样，对于安全运营中心 (SOC) 团队来说，它可以通过消除盲点来加快调查和修复工作。</p><p><strong>简化数据访问：</strong>Logsdb 可让 SRE 团队有效保留可操作的数据，用于故障排除、趋势分析和分析。同样，SOC 团队可以迅速搜索所有数据，进行调查和威胁猎取，而无需承担高昂的成本。</p><h2>Logsdb 已为您的环境做好准备</h2><p>Elasticsearch logsdb 索引模式从 8.17 版开始普遍适用于 Elastic Cloud 托管和自管理客户，并默认为<a href="https://www.elastic.co/cn/elasticsearch/serverless"> Elastic Cloud Serverless</a> 中的日志启用。</p><p>拥有标准、黄金级或白金级许可证的组织可以使用基本 logsdb 功能（包括智能索引排序和高级压缩）。无服务器客户和拥有企业许可证的组织可使用能进一步降低存储要求的完整 logsdb 功能（包括合成 _source）。</p><h2>运行中的 Elasticsearch logsdb</h2><p>Logsdb 使您能够保留所有日志数据并提高运营效率，而不会缩小收集范围或丢弃或孤立数据。通过智能索引排序、高级压缩和合成 _source 等功能，您可以在预算允许的范围内保存并分析所需的数据。</p><p>想亲自体验一下吗？<a href="https://cloud.elastic.co/registration">免费试用 Elastic</a>。</p><p><em>本文中描述的任何功能或功能性的发布和时间均由 Elastic 自行决定。当前尚未发布的任何功能或功能性可能无法按时提供或根本无法提供。</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-logsdb-index-mode</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-logsdb-index-mode</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Mark Settle,George Kobar,Amena Siddiqi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt982a3c761d66169f/6a17de7c7b54f9788d8b37ff/d3daacafea7a1d78c825a18f8281460c7106d3a5-721x420.png" length="0" type="image/png"/>
    <pubDate>Thu, 12 Dec 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[语义搜索的实现：使用 Elasticsearch 构建食谱搜索]]></title>
    <description><![CDATA[在电子商务网站中实施语义搜索。]]></description>
    <content:encoded><![CDATA[<h2>引言</h2><p>许多电子商务网站都希望提升食谱搜索体验。语义搜索如果应用得当，可以让客户根据更自然的查询快速找到所需的配料，例如"情人节的配料" 或"感恩节大餐。"</p><p>本文将演示如何使用 Elasticsearch 实现支持此类查询的语义搜索。我们将配置一个索引来存储超市的配料和产品的目录，并演示如何使用该索引来改进食谱搜索。在本文中，我们将介绍如何创建这种数据结构，并应用自然语言处理技术提供符合客户意图的相关结果。</p><p>本文中介绍的所有代码都是用 Python 开发的，可在<a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/building-a-recipe-search-with-elasticsearch">GitHub</a> 上获取。您可以访问资源库，查看源代码，根据需要进行调整，并直接在您的开发环境中实施解决方案。</p><h2>开始实施语义搜索</h2><p>要开始实施语义搜索，我们首先需要定义自然语言模型。Elastic 提供了自己的模型<a href="https://www.elastic.co/guide/en/machine-learning/8.15/ml-nlp-elser.html"><strong>ELSER</strong></a>，但也支持整合来自不同供应商的 NLP 模型，如 Hugging Face。这种灵活性使您可以选择最适合您需求的方案。</p><p>在本文中，我们将使用<strong>ELSER</strong>，它可以降低部署和管理 NLP 模型的复杂性。此外，Elastic 还提供<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search-semantic-text.html"><strong>语义文本</strong></a>功能，大大简化了流程。有了<strong>semantic_text</strong>，整个嵌入生成过程就变得简单而自动化。您只需定义一个推理点，并在索引映射中指定接收嵌入的字段。在编制文档索引时，将生成嵌入并自动与指定字段关联。</p><h3>设置步骤</h3><p>以下是创建支持语义搜索的索引的步骤。按照这些说明，您就可以配置好索引，并为语义搜索做好准备：</p><ol><li><p><strong>创建 </strong><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-elser.html"><strong>推理点</strong></a>。</p></li><li><p><a href="https://github.com/andreluiz1987/semantic-search-market/blob/main/infra.py"><strong>创建索引</strong></a>，将描述字段设置为 semantic_text，以便接收嵌入信息。</p></li><li><p><a href="https://github.com/andreluiz1987/semantic-search-market/blob/main/ingestion.py"><strong>将数据索引</strong></a>到杂货目录索引中，该索引将存储产品目录。该目录是从<a href="https://www.kaggle.com/datasets/bhavikjikadara/grocery-store-dataset?select=GroceryDataset.csv">此处</a>提供的数据集中获取的。</p></li></ol><h2>语义搜索在超市中的应用</h2><p>现在，我们已经用杂货店产品数据填充了索引，我们正在测试和验证查询，以便使用语义搜索改进搜索结果。我们的目标是提供更智能的搜索体验，了解上下文和用户意图，提供更相关、更准确的搜索结果。</p><h3>语义搜索解决的挑战</h3><p>基于产品目录，让我们来探讨一下语义搜索如何通过解决词汇和上下文问题来改变杂货店的搜索体验，而传统的词汇搜索往往难以解决这些问题。</p><h4><strong>1.解读烹饪意图</strong></h4><p><strong>问题 01</strong>：客户可能会搜索"烧烤海鲜" ，但词法搜索系统可能无法完全理解查询背后的意图。它可能无法识别所有适合烧烤的海鲜产品，只返回产品标题中包含"seafood" 或"grill" 的确切术语的产品。</p><p>首先，我们将进行词性搜索并分析结果。然后，我们将进行同样的语义搜索，比较同一搜索词的搜索结果。</p><p><strong>查询词法搜索</strong></p> response = client.search(
        index="grocery-catalog",
        size=5,
        source_excludes="description_embedding",
        query={
            "multi_match": {
                "query": "seafood for grilling",
                "fields": [
                    "name",
                    "description"]
            }
        }
    )<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>词法</p><p>西北鱼阿拉斯加贝尔迪雪蟹</p><p>10.453125</p><p>词法</p><p>吉田先生，酱汁原味美食</p><p>7.2289705</p><p>词法</p><p>优质海鲜品种包 - 20 件</p><p>7.1924105</p><p>词法</p><p>美国红鲷鱼 - 整只、头朝上、洗净</p><p>6.998647</p><p>词法</p><p>龙虾爪&amp; 胳膊，可持续野生捕捞</p><p>6.438654</p><p>词条搜索返回了一些适合烧烤的海鲜产品，如美国红鲷鱼和西北鱼阿拉斯加 Bairdi 雪蟹。然而，词法搜索返回的相关性较低的产品排在了列表的前列，例如吉田先生酱，它不是一种海鲜产品，而是一种肉酱，这表明词法算法很难完全理解"用于烧烤的语境。"</p><p><strong>语义搜索解决方案</strong></p><p>我们使用的查询方式是将"seafood" 与"grilling" 等烹饪上下文结合起来，返回一个全面的选项列表，如鱼片、虾和扇贝，这些都是烧烤的理想选择--即使"grill" 或"seafood" 这些词没有直接出现在产品名称中。这可确保搜索结果更贴近客户的意图。</p><p><strong>查询语义搜索：</strong></p>es_client.search(
   index="grocery-catalog-elser",
   size=size,
   source_excludes="description_embedding",
   query={
       "semantic": {
           "field": "description_embedding",
           "query": "seafood for grilling"

       }
   })<p>搜索类型</p><p>名称</p><p>得分</p><p>语义学</p><p>去头洗净的整条鲈鱼</p><p>16.175909</p><p>语义学</p><p>阿拉斯加黑鳕鱼（黑貂鱼）</p><p>15.855331</p><p>语义学</p><p>美国红鲷鱼 - 整只，头朝下</p><p>15.454779</p><p>语义学</p><p>西北鱼阿拉斯加贝尔迪雪蟹</p><p>15.855331</p><p>语义学</p><p>美国红鲷鱼 - 整只，头朝下</p><p>15.3892355</p><p>语义搜索不仅返回了与"seafood," 这一术语直接相关的产品，而且还理解了"grilling," 这一上下文，并带来了适合烧烤的整鱼和鱼片。关键在于结果的精确性，其中包括烤制常用的全鱼，如白鲷鱼和阿拉斯加黑鳕鱼。</p><p><strong>问题 02 </strong> ：许多客户在工作一天后会搜索快速简便的晚餐解决方案，使用的术语包括"Easy Weeknight meals。"传统的词法搜索可能无法完全捕捉到快餐的概念，通常只关注名称中包含"easy" 一词的产品。</p><p>与上一个问题一样，我们将首先进行词法搜索。之后，我们将采用语义搜索来解决问题。</p><p><strong>查询词法搜索</strong></p> response = client.search(
        index="grocery-catalog",
        size=5,   
        source_excludes="description_embedding",
        query={
            "multi_match": {
                "query": "easy weeknight meals",
                "fields": [
                    "name",
                    "description"]
            }
        }
    )<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>词法</p><p>艾利易撕地址标签，4200 个装</p><p>8.017723</p><p>词法</p><p>自热应急/便携餐 32</p><p>6.592727</p><p>词法</p><p>海岸海鲜黄鳍金枪鱼块 Poke</p><p>5.836883</p><p>词法</p><p>Hefty 超重 12 盎司泡沫塑料</p><p>5.8116536</p><p>词法</p><p>Vanity Fair Everyday餐巾纸，2层，110片装</p><p>5.752989</p><p>词法搜索返回的相关结果要少得多，包括与餐饮完全无关的物品，如 Avery Easy Peel Address Labels 和 Vanity Fair Everyday Napkins。这些产品无法满足用户对快餐的需求。虽然词法搜索确实返回了一个有用的产品（Omeals Self Heating Emergency Meals），但其他结果，如餐巾纸和标签，仅与描述中的"easy" 或"weeknight" 匹配，没有真正满足用户对快速用餐解决方案的需求。</p><p><strong>语义搜索解决方案</strong></p><p>我们实施了一项查询，了解快速简便餐饮背后的意图。它将可快速烹制的产品，如预煮肉类、冷冻意大利面或套餐联系起来，即使这些产品的名称中没有明确包含"easy" 这个词。这种方法可确保顾客找到最合适的选择，在周末快速享用晚餐，满足对便利性的需求。</p><p><strong>查询语义搜索</strong></p>es_client.search(
   index="grocery-catalog-elser",
   size=size,
   source_excludes="description_embedding",
   query={
       "semantic": {
           "field": "description_embedding",
           "query": "easy weeknight meals"

       }
   })<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>语义学</p><p>自热应急/便携餐 32</p><p>14.610006</p><p>语义学</p><p>Nissin，杯面，虾，2.5 盎司</p><p>13.751424</p><p>语义学</p><p>Namaste 无谷蛋白华夫饼&amp; 煎饼预拌粉</p><p>13.73376</p><p>语义学</p><p>爱达荷土豆、黄金烤土豆饼</p><p>12.549422</p><p>语义学</p><p>Nissin 杯面，鸡肉，24 支装</p><p>12.034527</p><p>语义搜索返回的产品明显与方便快捷的膳食有关，如方便面（杯面）、预煮土豆和煎饼粉，这些都是简单的周末晚餐的典型选择。这表明，语义搜索可以抓住"简易隔夜饭这一短语背后的概念，" ，从而捕捉到用户寻找快捷方便饭菜的意图。有趣的是，其他类别的产品，如"苏打水、" ，在相关情况下（如佐餐饮料）也可能包括在内。</p><h4><strong>2.地区术语和词汇变化</strong></h4><p><strong>问题</strong>：一位客户可能会搜索"soda," ，而另一位客户可能会使用"pop" 来搜索相同的产品。传统的词库检索无法识别这两个词指的是同一个项目。</p><p><strong>查询词法搜索</strong></p> response = client.search(
        index="grocery-catalog",
        size=5,
        source_excludes="description_embedding",
        query={
            "multi_match": {
                "query": "refreshing pop drink low sugar",
                "fields": [
                    "name",
                    "description"]
            }
        }
    )<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>词法</p><p>Prime Hydration+ Sticks 电解质混合饮料</p><p>14.492869</p><p>词法</p><p>Capri Sun，100% 果汁，多种包装</p><p>12.340851</p><p>词法</p><p>Joyburst 能量饮料, Frose Rose, 12瓶装</p><p>11.839179</p><p>词法</p><p>Kellogg's Pop-Tarts, Frosted Brown Sugar Cinnamon</p><p>9.97788</p><p>词法</p><p>Kind 迷你巧克力棒, 多样包装, 0.7</p><p>9.336912</p><p>词法搜索侧重于精确的词语匹配。虽然它返回了 Prime Hydration 和 Capri Sun 等产品，但与"pop" 一词直接匹配也导致了不相关的结果，如 Kellogg's Pop-Tarts，它是一种零食而不是饮料。这凸显了当一个术语有多种含义或可能含糊不清时，词汇搜索的效果会如何降低。</p><p><strong>语义搜索解决方案</strong></p><p>在语义查询中，我们可以克服词汇搜索无法解决的词汇变化问题。通过扩展搜索条件，我们能够根据上下文的含义获得结果，提供更相关、更全面的回复。</p><p><strong>查询：</strong></p>es_client.search(
   index="grocery-catalog-elser",
   size=size,
   source_excludes="description_embedding",
   query={
       "semantic": {
           "field": "description_embedding",
           "query": "refreshing pop drink low sugar"

       }
   })<p><strong>结果</strong></p><p>搜索类型</p><p>名称</p><p>得分</p><p>语义学</p><p>奥利葆 12 盎司益生元苏打水品种</p><p>14.776867</p><p>语义学</p><p>佰草集抗氧化可可粉，多种包装，18 磅</p><p>14.663253</p><p>语义学</p><p>怪物能量饮料，零度超能，24 磅</p><p>14.486348</p><p>语义学</p><p>Joyburst 能量饮料，12 盎司</p><p>14.007214</p><p>语义学</p><p>Joyburst 能量饮料, Frose Rose, 12瓶装</p><p>13.641038</p><p>语义搜索会返回与"pop" 这一概念直接匹配的产品，作为"soda" 的同义词（如 Olipop Prebiotics Soda），即使产品名称中可能没有"pop" 这一确切术语。搜索理解了用户的意图--清爽的低糖饮料--并能够返回相关产品，包括益生元苏打水（Olipop）和无糖能量饮料（Monster Energy Drink）等选项。</p><h2>结论</h2><p>事实证明，在食品杂货店中实施语义搜索对于理解复杂的查询非常有效，如"烧烤海鲜" 和"简易周末餐。"这种方法使我们能够更准确地解读用户意图，返回高度相关的产品。</p><p>通过使用 Elasticsearch 和 ELSER 简化流程，我们能够快速高效地应用语义搜索，显著改善搜索结果，提供更灵活、更有针对性的购物体验。这不仅优化了搜索过程，还提高了为客户提供的搜索结果的相关性。</p><h2>参考资料</h2><p>ELSER 型：</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/put-inference-api.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/put-inference-api.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-elser.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-elser.html</a></p><p></p><p>语义文本：</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-text.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-text.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html</a></p><p></p><p>数据集：</p><p><a href="https://www.kaggle.com/datasets/bhavikjikadara/grocery-store-dataset?select=GroceryDataset.csv">https://www.kaggle.com/datasets/bhavikjikadara/grocery-store-dataset?select=GroceryDataset.csv</a></p><p></p><p>语义搜索：</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search-semantic-text.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search-semantic-text.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/semantic-search-elasticsearch-ecommerce</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/semantic-search-elasticsearch-ecommerce</guid>
    <category><![CDATA[基础功能]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c545fc80b6d79d6/6a170214839dfad776dcfd6f/d968e646240cd3ef7c79b5124d562a5f951d812b-1440x840.png" length="0" type="image/png"/>
    <pubDate>Thu, 07 Nov 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>