<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Srikanth Manvi - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Srikanth Manvi - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/cn/search-labs/author/srikanth-manvi</link>
    </image>
    <link>https://www.elastic.co/cn/search-labs/author/srikanth-manvi</link>
    <atom:link href="https://www.elastic.co/cn/search-labs/rss/author/srikanth-manvi.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[cn]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 12:38:41 GMT</lastBuildDate>
  <item>
    <title><![CDATA[如何使用用于微软语义内核（Microsoft Semantic Kernel）的Elasticsearch矢量存储连接器进行人工智能代理开发]]></title>
    <description><![CDATA[微软语义内核（Microsoft Semantic Kernel）是一款轻量级开源开发工具包，可让您轻松构建人工智能代理，并将最新的人工智能模型集成到您的 C#、Python 或 Java 代码库中。随着Semantic Kernel Elasticsearch向量存储连接器（Elasticsearch Vector Store Connector）的发布，使用Semantic Kernel构建人工智能代理的开发人员现在可以将Elasticsearch作为可扩展的企业级向量存储插件，同时继续使用Semantic Kernel抽象。]]></description>
    <content:encoded><![CDATA[<p>我们与<a href="https://learn.microsoft.com/en-us/semantic-kernel/overview/"> 微软语义内核</a> （ Microsoft<a href="https://learn.microsoft.com/en-us/semantic-kernel/overview/"> Semantic Kernel ）团队合作，宣布面向 微软语义内核</a> （.NET）用户推出<a href="https://github.com/elastic/semantic-kernel-net/"> Semantic Kernel Elasticsearch矢量存储连接器（Vector Store Connector ）。</a>语义内核（Semantic Kernel）简化了企业级人工智能代理的构建过程，包括利用来自矢量存储库（Vector Store）的更多相关数据驱动响应来增强大型语言模型（LLM）的能力。语义内核（Semantic Kernel）为与Elasticsearch等矢量存储进行交互提供了一个无缝的抽象层，可提供创建、列出和删除记录集合以及上传、检索和删除单条记录等基本功能。</p><p><a href="https://learn.microsoft.com/en-us/semantic-kernel/concepts/vector-store-connectors/out-of-the-box-connectors/elasticsearch-connector?pivots=programming-language-csharp">开箱即用的Semantic Kernel Elasticsearch向量存储连接器（Vector Store Connector</a>）支持Semantic Kernel<a href="https://learn.microsoft.com/en-us/semantic-kernel/concepts/vector-store-connectors/?pivots=programming-language-csharp#the-vector-store-abstraction">向量存储抽象</a>，这使得开发人员在构建人工智能代理时能够非常容易地将Elasticsearch作为向量存储插件。</p><p>Elasticsearch 在开源社区拥有坚实的基础，最近采用了<a href="https://www.elastic.co/blog/elasticsearch-is-open-source-again">AGPL 许可证</a>。这些工具与开源的微软语义内核（Microsoft Semantic Kernel）相结合，可提供强大的企业级解决方案。您可以通过运行此命令<code>curl -fsSL https://elastic.co/start-local | sh </code> ，在几分钟内启动 Elasticsearch（详情请参考<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">start-local</a>），然后在生产人工智能代理的同时，迁移到<a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;utm_source=semantickernel&amp;utm_content=documentation">云托管</a>或<a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.16/install-elasticsearch.html">自托管</a>版本。</p><p>在本篇博客中，我们将探讨在使用Semantic Kernel（语义内核）时，如何使用<a href="https://github.com/elastic/semantic-kernel-net/">Semantic Kernel Elasticsearch向量存储连接器</a>。该连接器的 Python 版本将在未来推出。</p><h2>高级应用场景：利用 Semantic Kernel&amp; Elasticsearch 构建 RAG 应用程序</h2><p>下面我们将举例说明。在高层次上，我们正在构建一个 RAG（检索增强生成）应用程序，它将用户的问题作为输入，并返回一个答案。我们将使用 Azure OpenAI （ 也可使用<a href="https://devblogs.microsoft.com/semantic-kernel/introducing-new-ollama-connector-for-local-models/"> 本地 LLM</a> ）作为 LLM，Elasticsearch 作为向量存储，Semantic Kernel (.net) 作为将所有组件连接在一起的框架。</p><p>如果您不熟悉 RAG 架构，可以通过以下文章快速了解<a href="https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag">： https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag。</a></p><p>答案由 LLM 生成，LLM 从 Elasticsearch 向量存储中获取与问题相关的上下文。答复还包括法律硕士用作背景的资料来源。</p><h3>RAG 示例</h3><p>在这个具体例子中，我们创建了一个应用程序，允许用户就内部酒店数据库中存储的酒店提出问题。例如，用户可以根据不同标准搜索特定酒店，或要求提供酒店列表。</p><p>在示例数据库中，我们生成了一个包含 100 个条目的<a href="https://github.com/elastic/semantic-kernel-net/blob/main/Elastic.SemanticKernel.Playground/hotels.csv">酒店列表</a>。为了让您尽可能轻松地试用连接器演示，我们特意设置了较小的样本量。在实际应用中，Elasticsearch 连接器将显示出其优于其他选项（如 "InMemory "向量存储实现）的优势，尤其是在处理超大数据量时。</p><p>完整的演示应用程序可在 Elasticsearch 向量存储连接器存储<a href="https://github.com/elastic/semantic-kernel-net/tree/main/Elastic.SemanticKernel.Playground">库中</a>找到。</p><p>让我们先将所需的 NuGet 软件包和指令添加到项目中：</p>dotnet add package "Elastic.Clients.Elasticsearch" -v 8.16.2
dotnet add package "Elastic.SemanticKernel.Connectors.Elasticsearch" -v 0.1.2
dotnet add package "Microsoft.Extensions.Hosting" -v 9.0.0
dotnet add package "Microsoft.SemanticKernel.Connectors.AzureOpenAI" -v 1.30.0
dotnet add package "Microsoft.SemanticKernel.PromptTemplates.Handlebars" -v 1.30.0using System;
using System.IO;
using System.Linq;
using System.Threading.Tasks;

using Elastic.Clients.Elasticsearch;
using Elastic.Transport;

using Microsoft.Extensions.DependencyInjection;
using Microsoft.Extensions.Hosting;
using Microsoft.Extensions.VectorData;
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.Data;
using Microsoft.SemanticKernel.Embeddings;
using Microsoft.SemanticKernel.PromptTemplates.Handlebars;<p>现在，我们可以创建我们的数据模型，并为其提供语义内核（Semantic Kernel）的特定属性，以定义存储模型模式和文本搜索的一些提示：</p>/// &lt;summary&gt;
/// Data model for storing a "hotel" with a name, a description, a  description embedding and an optional reference link.
/// &lt;/summary&gt;
public sealed record Hotel
{
	[VectorStoreRecordKey]
	public required string HotelId { get; set; }

	[TextSearchResultName]
	[VectorStoreRecordData(IsFilterable = true)]
	public required string HotelName { get; set; }

	[TextSearchResultValue]
	[VectorStoreRecordData(IsFullTextSearchable = true)]
	public required string Description { get; set; }

	[VectorStoreRecordVector(Dimensions: 1536, DistanceFunction.CosineSimilarity, IndexKind.Hnsw)]
	public ReadOnlyMemory&lt;float&gt;? DescriptionEmbedding { get; set; }

	[TextSearchResultLink]
	[VectorStoreRecordData]
	public string? ReferenceLink { get; set; }
}<p>存储模型模式属性（`VectorStore*`）与 Elasticsearch 向量存储连接器的实际使用最为相关，即</p><p></p><ul><li><p><code>VectorStoreRecordKey</code> 来标记记录类上的一个属性，作为记录存储在向量存储中的键。</p></li><li><p><code>VectorStoreRecordData</code> 将记录类的一个属性标记为 "数据"。</p></li><li><p><code>VectorStoreRecordVector</code> 将记录类的一个属性标记为矢量。</p></li></ul><p>所有这些属性都接受各种可选参数，可用于进一步定制存储模型。以<code>VectorStoreRecordKey </code> 为例，可以指定不同的距离函数或不同的索引类型。</p><p>文本搜索属性 (<code>TextSearch*</code>) 在本示例的最后一步中非常重要。我们稍后再谈。</p><p>下一步，我们将初始化语义内核引擎，并获取核心服务的引用。在实际应用中，应使用<a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/dependency-injection">依赖注入</a>而不是直接访问服务集合。同样的道理也适用于硬编码的配置和秘密，它们应该使用<a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/configuration">配置提供程序</a>来读取：</p>var builder = Host.CreateApplicationBuilder(args);

// Register AI services.
var kernelBuilder = builder.Services.AddKernel();

kernelBuilder.AddAzureOpenAIChatCompletion("gpt-4o", "https://my-service.openai.azure.com", "my_token");

kernelBuilder.AddAzureOpenAITextEmbeddingGeneration("ada-002", "https://my-service.openai.azure.com", "my_token");

// Register text search service.
kernelBuilder.AddVectorStoreTextSearch&lt;Hotel&gt;();

// Register Elasticsearch vector store.
var elasticsearchClientSettings = new ElasticsearchClientSettings(new Uri("https://my-elasticsearch-instance.cloud"))
    .Authentication(new BasicAuthentication("elastic", "my_password"));

kernelBuilder.AddElasticsearchVectorStoreRecordCollection&lt;string, Hotel&gt;("skhotels", elasticsearchClientSettings);

// Build the host.
using var host = builder.Build();

// For demo purposes, we access the services directly without using a DI context.

var kernel = host.Services.GetService&lt;Kernel&gt;()!;
var embeddings = host.Services.GetService&lt;ITextEmbeddingGenerationService&gt;()!;
var vectorStoreCollection = host.Services.GetService&lt;IVectorStoreRecordCollection&lt;string, Hotel&gt;&gt;()!;

// Register search plugin.
var textSearch = host.Services.GetService&lt;VectorStoreTextSearch&lt;Hotel&gt;&gt;()!;
kernel.Plugins.Add(textSearch.CreateWithGetTextSearchResults("SearchPlugin"));<p>现在可以使用<code>vectorStoreCollection</code> 服务创建数据集，并摄取一些<a href="https://github.com/elastic/semantic-kernel-net/blob/main/Elastic.SemanticKernel.Playground/hotels.csv">演示记录</a>：</p>await vectorStoreCollection.CreateCollectionIfNotExistsAsync();

// CSV format: ID;Hotel Name;Description;Reference Link
var hotels = (await File.ReadAllLinesAsync("hotels.csv"))
    .Select(x =&gt; x.Split(';'));

foreach (var chunk in hotels.Chunk(25))
{
    var descriptionEmbeddings = await embeddings.GenerateEmbeddingsAsync(chunk.Select(x =&gt; x[2]).ToArray());
    
    for (var i = 0; i &lt; chunk.Length; ++i)
    {
        var hotel = chunk[i];
        await vectorStoreCollection.UpsertAsync(new Hotel
        {
            HotelId = hotel[0],
            HotelName = hotel[1],
            Description = hotel[2],
            DescriptionEmbedding = descriptionEmbeddings[i],
            ReferenceLink = hotel[3]
        });
    }
}<p>由此可见，语义内核（Semantic Kernel）是如何将向量存储的使用及其复杂性简化为几个简单的方法调用的。</p><p>在 Elasticsearch 中创建一个新索引，并创建所有必要的属性映射。然后，我们的数据集会完全透明地映射到存储模型中，并最终存储到索引中。下面是映射在 Elasticsearch 中的显示方式。</p>{
  "mappings": {
    "properties": {
      "descriptionEmbedding": {
        "dims": 1536,
        "index": true,
        "index_options": {
          "type": "hnsw"
        },
        "similarity": "cosine",
        "type": "dense_vector"
      },
      "hotelName": {
        "type": "keyword"
      },
      "description": {
        "type": "text"
      }
    }
  }
}<p><code>embeddings.GenerateEmbeddingsAsync()</code> 会透明地调用已配置的 Azure AI 嵌入生成服务。</p><p>在这个演示的最后一个步骤中，我们还可以看到更多的神奇之处。</p><p>当用户就数据提问时，只需调用<code>InvokePromptAsync</code> ，就能执行以下所有操作：</p><p>1.为用户的问题生成嵌入代码</p><p>2.在矢量存储器中搜索相关条目</p><p>3.将查询结果插入提示模板</p><p>4.最终提示形式的实际查询将发送到人工智能聊天完成服务</p>// Invoke the LLM with a template that uses the search plugin to
// 1. get related information to the user query from the vector store
// 2. add the information to the LLM prompt.
var response = await kernel.InvokePromptAsync(
    promptTemplate: """
                    Please use this information to answer the question:
                    {{#with (SearchPlugin-GetTextSearchResults question)}}
                      {{#each this}}
                        Name: {{Name}}
                        Value: {{Value}}
                        Source: {{Link}}
                        -----------------
                      {{/each}}
                    {{/with}}
                    
                    Include the source of relevant information in the response.

                    Question: {{question}}
                    """,
    arguments: new KernelArguments
    {
        { "question", "Please show me all hotels that have a rooftop bar." },
    },
    templateFormat: "handlebars",
    promptTemplateFactory: new HandlebarsPromptTemplateFactory());<p>还记得我们之前在数据模型上定义的<code>TextSearch*</code> 属性吗？有了这些属性，我们就能在提示模板中使用相应的占位符，这些占位符会根据向量存储中的条目信息自动填充。</p><p>对于我们的问题"，请告诉我所有拥有屋顶酒吧的酒店。" ，最终答复如下：</p>Console.WriteLine(response.ToString());

// &gt; The hotel that has a rooftop bar is Skyline Suites. You can find more information about this hotel [here](https://example.com/yz567).<p>正确答案是指 hotels.csv 中的以下条目</p>9;
Skyline Suites;
Offering panoramic city views from every suite, this hotel is perfect for those who love the urban landscape. Enjoy luxurious amenities, a rooftop bar, and close proximity to attractions. Luxurious and contemporary.;
https://example.com/yz567<p>这个例子很好地说明了微软语义内核的使用是如何通过其深思熟虑的抽象功能大大降低复杂性，并实现高度灵活性的。例如，只需修改一行代码，就可以更换向量存储或所使用的人工智能服务，而无需重构代码的任何其他部分。</p><p>同时，该框架还提供了大量高级功能，如 "InvokePrompt "函数或模板或搜索插件系统。</p><p>完整的演示应用程序可在 Elasticsearch 向量存储连接器存储库中找到。</p><h2>Elasticsearch 还能做什么</h2><ul><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">Elasticsearch 新语义文本映射：简化语义搜索</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-reranking-with-retrievers">利用检索器在 Elasticsearch 中进行语义重排</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">高级 RAG 技术第 1 部分：数据处理</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">高级 RAG 技术第 2 部分：查询和测试</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-rag-with-llama3-opensource-and-elastic">使用 Llama 3 开放源代码和 Elastic 构建 RAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/local-rag-agent-elasticsearch-langgraph-llama3">使用 LangGraph、LLaMA3 和 Elasticsearch 向量存储从零开始构建本地代理的教程</a></p></li></ul><h2>Elasticsearch&amp; Semantic Kernel（语义内核）：下一步是什么？</h2><ul><li><p>我们展示了在.NET中构建GenAI应用时，如何将Elasticsearch向量存储轻松插入Semantic Kernel。敬请期待下一步的 Python 集成。</p></li><li><p>由于Semantic Kernel（语义内核）为<a href="https://www.elastic.co/search-labs/tutorials/search-tutorial/vector-search/hybrid-search">混合</a>搜索等高级搜索功能建立了抽象，Elasticsearch连接将使.NET开发人员能够在使用Semantic Kernel（语义内核）的同时轻松实现这些功能。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-connector-microsoft-semantic-kernel</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-connector-microsoft-semantic-kernel</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[.NET]]></category>
    <category><![CDATA[向量数据库]]></category>
    <dc:creator><![CDATA[Florian Bernd,Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2d8725035e86f8a8/6a17fe447f6f1564f8c09d74/0564fe794e4c66d0507317822d7aa71826183d20-1311x762.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 06 Dec 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[利用 Elasticsearch 和 LlamaIndex 保护 RAG 中的敏感信息和 PII 信息]]></title>
    <description><![CDATA[如何使用 Elasticsearch 和 LlamaIndex 保护 RAG 应用程序中的敏感数据和 PII 数据。]]></description>
    <content:encoded><![CDATA[<p></p><p></p><p>在本篇文章中，我们将探讨在 RAG（检索增强生成）流程中使用公共 LLM 时保护个人身份信息 (PII) 和敏感数据的方法。我们将探索使用开源库和正则表达式屏蔽 PII 和敏感数据，以及在调用公共 LLM 之前使用本地 LLM 屏蔽数据。</p><p>在开始之前，让我们回顾一下我们在本篇文章中使用的一些术语。</p><h2>术语</h2><p><a href="https://www.llamaindex.ai/">LlamaIndex</a>是用于构建 LLM（大型语言模型）应用程序的领先数据框架。LlamaIndex 为构建 RAG（检索增强生成）应用程序的各个阶段提供了抽象概念。像 LlamaIndex 和 LangChain 这样的框架提供了抽象，因此应用程序不会与任何特定 LLM 的应用程序接口紧密耦合。</p><p><a href="https://www.elastic.co/enterprise-search">Elasticsearch</a>由<a href="https://elastic.co/">Elastic</a> 提供。Elastic 是 Elasticsearch 背后的行业领导者，Elasticsearch 是一个可扩展的数据存储和矢量数据库，支持精确的全文搜索、语义理解的矢量搜索以及两全其美的混合搜索。Elasticsearch 是一个分布式 RESTful 搜索和分析引擎、可扩展数据存储和矢量数据库。我们在本博客中使用的 Elasticsearch 功能在 Elasticsearch 的免费开放版本中提供。</p><p><a href="https://www.promptingguide.ai/techniques/rag">检索增强生成（RAG）</a>是一种人工智能技术/模式，在这种模式下，LLM 可以利用外部知识生成对用户查询的回复。这样，法律硕士的答复就可以根据具体情况量身定做，而不是泛泛而谈。</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.13/semantic-search.html">嵌入</a>是文本/媒体意义的数字表示。它们是高维信息的低维表示。</p><h2>RAG 和数据保护</h2><p>一般来说，大型语言模型（LLM）善于根据模型中的可用信息生成响应，而这些信息可能是在互联网数据上训练出来的。然而，对于那些模型中没有信息的查询，则需要向法律硕士提供模型中没有的外部知识或具体细节。这些信息可能存在于您的数据库或内部知识系统中。检索增强生成（RAG）是一种技术，对于给定的用户查询，首先从外部（LLM）系统（如数据库）检索相关的上下文/信息，然后将上下文与用户查询一起发送给 LLM，以生成更具体、更相关的响应。</p><p>这使得 RAG 技术在问题解答、内容创建以及任何有利于深入理解上下文和细节的应用中都非常有效。</p><p>因此，在 RAG 管道中，您有可能将 PII（个人身份信息）等内部信息和敏感信息（如姓名、出生日期、账号等）暴露给公共 LLM。</p><p>虽然在使用 Elasticsearch 等矢量数据库时，数据是安全的（通过各种杠杆，如<a href="https://www.elastic.co/guide/en/cloud-enterprise/current/ece-configure-rbac.html">基于角色的访问控制</a>、<a href="https://www.elastic.co/search-labs/blog/dls-internal-knowledge-search">文档级安全</a>等），但在向外部公共 LLM 发送数据时必须小心谨慎。</p><p>在使用大型语言模型 (LLM) 时，出于多种原因，保护个人身份信息 (PII) 和敏感数据至关重要：</p><ul><li><p><strong>隐私合规</strong>：许多地区都有严格的法规，如欧洲的《通用数据保护条例》（GDPR）或美国的《加利福尼亚消费者隐私法》（CCPA），这些法规都要求保护个人数据。要避免法律后果和罚款，就必须遵守这些法律。</p></li><li><p><strong>用户信任</strong>：确保敏感信息的保密性和完整性可建立用户信任。用户更愿意使用他们认为能保护其隐私的系统，并与之互动。</p></li><li><p><strong>数据安全</strong>：防止数据泄露至关重要。如果没有足够的保障措施，暴露在法律硕士面前的敏感数据很容易被窃取或滥用，从而导致身份被盗或金融欺诈等潜在危害。</p></li><li><p><strong>道德方面的考虑</strong>：从道德角度讲，尊重用户隐私并负责任地处理他们的数据非常重要。对 PII 处理不当会导致歧视、侮辱或其他负面社会影响。</p></li><li><p><strong>企业声誉</strong>：未能保护敏感数据的公司可能会声誉受损，这可能会对其业务造成长期负面影响，包括失去客户和收入。</p></li><li><p><strong>减少滥用风险</strong>：安全处理敏感数据有助于防止对数据或模型的恶意使用，例如在有偏见的数据上训练模型，或利用数据操纵或伤害个人。</p></li></ul><p>总之，为了确保法律合规、维护用户信任、确保数据安全、坚持道德标准、保护企业声誉和降低滥用风险，必须对 PII 和敏感数据进行强有力的保护。</p><h2>快速回顾</h2><p>在<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch">上一篇文章</a>中，我们讨论了如何使用 RAG 技术，将 Elasticsearch 作为向量数据库，同时使用 LlamaIndex 和本地运行的 Mistral LLM 来实现 Q&amp;A 体验。在此基础上，我们将继续努力。</p><p>阅读上一篇文章是可有可无的，因为我们现在将快速讨论/复述上一篇文章中的内容。</p><p>我们有一个样本数据集，内容是一家虚构的家庭保险公司的座席人员与客户之间的呼叫中心对话。我们开发了一个简单的 RAG 应用程序，可以回答 "客户因哪些与水有关的问题而提出索赔 "等问题。</p><p>从高度上看，流程是这样的。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" alt="RAG 流程" /><p>在索引阶段，我们使用 LlamaIndex 管道加载并索引文档。文档被分块并连同其嵌入一起存储在 Elasticsearch 向量数据库中。</p><p>在查询阶段，当用户提出问题时，LlamaIndex 会检索与查询相关的前 K 个相似文档。这些排名前 K 位的相关文档连同查询结果被发送到本地运行的 Mistral LLM，然后由其生成响应并发送回用户。请随时查看上一篇文章或<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/tree/main">探索代码</a>。</p><p>在上一篇文章中，我们在本地运行了 LLM。不过，在生产过程中，您可能希望使用<a href="https://openai.com/">OpenAI</a>、<a href="https://mistral.ai/">Mistral</a>、<a href="https://www.anthropic.com/claude">Anthropic</a>等公司提供的外部 LLM。这可能是因为您的使用案例需要一个更大的基础模型，或者由于企业生产的需要（如可扩展性、可用性、性能等），在本地运行不是一种选择。</p><p>在 RAG 管道中引入外部 LLM 会使您面临不慎将敏感和 PII 泄露给 LLM 的风险。在本篇文章中，我们将探讨如何在将文档发送给外部法律硕士之前，将 PII 信息屏蔽作为 RAG 管道的一部分。</p><h2>持有公共法学硕士学位的 RAG</h2><p>在讨论如何在 RAG 管道中保护 PII 和敏感信息之前，我们将首先使用 LlamaIndex、Elasticsearch 向量数据库和 OpenAI LLM 构建一个简单的 RAG 应用程序。</p><h3>准备工作</h3><p>我们需要以下材料</p><ul><li><p>运行<strong>Elasticsearch</strong>作为向量数据库来存储嵌入。按照上一篇文章中关于<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#install-elasticsearch">安装 Elasticsearch</a> 的说明进行操作。</p></li><li><p>开放人工智能应用程序接口密钥。</p></li></ul><h3>简单的 RAG 应用</h3><p>整个代码可在<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/tree/protecting-pii">Github Repository</a>（branch:protection-pii）中找到，以供参考。克隆该 repo 是可选的，因为我们将在下文中详细介绍代码。</p><p>在您最喜欢的集成开发环境中，用以下 3 个文件创建一个新的 Python 应用程序。</p><ul><li><p><code>index.py</code> 与索引数据有关的代码的位置。</p></li><li><p><code>query.py</code> 与查询和 LLM 交互相关的代码都放在这里。</p></li><li><p><code>.env</code> 配置属性（如 API 密钥）的位置。</p></li></ul><p>我们需要安装一些软件包。首先，我们要在应用程序的根文件夹中创建一个新的 python<a href="https://docs.python.org/3/library/venv.html">虚拟环境</a>。</p>python3 -m venv .venv
<p>激活虚拟环境并安装以下所需软件包。</p>source .venv/bin/activate
pip install llama-index 
pip install llama-index-embeddings-openai
pip install llama-index-vector-stores-elasticsearch
pip install sentence-transformers
pip install python-dotenv
pip install openai
<p>在 .env 中配置 OpenAI 和 Elasticsearch 连接属性锉刀</p>OPENAI_API_KEY="REPLACEME"
ELASTIC_CLOUD_ID="REPLACEME"
ELASTIC_API_KEY="REPLACEME"
<h4>索引数据</h4><p>下载<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/blob/main/conversations.json">conversations.json</a>文件，其中包含客户与我们虚构的房屋保险公司呼叫中心座席之间的<em>对话</em>。将该文件与 2 个 python 文件和 .env 文件一起放在应用程序的根目录中。文件。下面是该文件内容的示例。</p>{
"conversation_id": 103,
"customer_name": "Sophia Jones",
"agent_name": "Emily Wilson",
"policy_number": "JKL0123",
"conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
"summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
<p>在<code>index.py</code> 中粘贴下面的代码，该代码负责索引数据。</p># index.py
# pip install sentence-transformers
# pip install llama-index-embeddings-openai
# pip install llama-index-embeddings-huggingface

import json
import os
from dotenv import load_dotenv
from llama_index.core import Document
from llama_index.core import Settings
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore


def get_documents_from_file(file):
   """Reads a json file and returns list of Documents"""

   with open(file=file, mode='rt') as f:
       conversations_dict = json.loads(f.read())

   # Build Document objects using fields of interest.
   documents = [Document(text=item['conversation'],
                         metadata={"conversation_id": item['conversation_id']})
                for
                item in conversations_dict]
   return documents

# Load .env file contents into env
load_dotenv('.env')
Settings.embed_model = HuggingFaceEmbedding(
   model_name="BAAI/bge-small-en-v1.5"
)

def main():
   # ElasticsearchStore is a VectorStore that
   # takes care of Elasticsearch Index and Data management.
   es_vector_store = ElasticsearchStore(index_name="convo_index",
                                        vector_field='conversation_vector',
                                        text_field='conversation',
                                        es_cloud_id=os.getenv("ELASTIC_CLOUD_ID"),
                                        es_api_key=os.getenv("ELASTIC_API_KEY"))

   # LlamaIndex Pipeline configured to take care of chunking, embedding
   # and storing the embeddings in the vector store.
   llamaindex_pipeline = IngestionPipeline(
       transformations=[
           SentenceSplitter(chunk_size=350, chunk_overlap=50),
           Settings.embed_model
       ],
       vector_store=es_vector_store
   )

   # Load data from a json file into a list of LlamaIndex Documents
   documents = get_documents_from_file(file="conversations.json")
   llamaindex_pipeline.run(documents=documents)
   print(".....Indexing Data Completed.....\n")

if __name__ == "__main__":
   main()
<p>运行上述代码可看到在 Elasticsearch 中创建了一个索引，将嵌入信息存储在名为<code>convo_index</code> 的 Elasticsearch 索引中。</p><p>如果您需要有关 LlamaIndex IngestionPipeline 的解释，请参阅上一篇文章中的<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#indexing-data">创建 IngestionPipeline</a> 部分。</p><h4>查询</h4><p>在上一篇文章中，我们使用了本地 LLM 进行<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#querying">查询</a>。</p><p>在本篇文章中，我们将使用公共 LLM OpenAI，如下所示。</p># query.py
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.llms.openai import OpenAI
from index import es_vector_store

# Public LLM where we send user query and Related Documents
llm = OpenAI()

index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents are sent as-is. So any PII/Sensitive data is sent to the LLM.
query_engine = index.as_query_engine(llm, similarity_top_k=10)

query="Give me summary of water related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>上述代码将打印 OpenAI 的响应如下。</p><p>客户提出了各种与水有关的索赔，包括地下室水渍、水管爆裂、冰雹对屋顶造成的损坏等问题，以及由于未及时通知、维护问题、逐渐磨损和原有损坏等原因造成的拒赔。在每个案例中，客户都对索赔被拒表示沮丧，并寻求对其索赔进行公平的评估和决定。</p><h2>在 RAG 中屏蔽 PII</h2><p>到目前为止，我们所做的工作是将文档原样连同用户查询一起发送给 OpenAI。</p><p>在 RAG 管道中，从矢量存储中检索到相关上下文后，我们有机会在将查询和上下文发送到 LLM 之前屏蔽 PII 和敏感信息。</p><p>在向外部法律硕士发送 PII 信息之前，有多种方法可以掩盖 PII 信息，每种方法都有自己的优点。下面我们来看看其中的一些选择</p><ol><li><p>使用 spacy.io 或<a href="https://microsoft.github.io/presidio/">Presidio</a>（微软维护的开源库）等 NLP 库。</p></li><li><p>使用开箱即用的 LlamaIndex <code>NERPIINodePostprocessor.</code></p></li><li><p>通过 <code>PIINodePostprocessor</code></p></li></ol><p>使用上述任何一种方法实现屏蔽逻辑后，您就可以使用后处理器（您自己定制的后处理器或 LlamaIndex 开箱即用的后处理器）配置 LlamaIndex 的 IngestionPipeline。</p><h3>使用 NLP 库</h3><p>作为 RAG 管道的一部分，我们可以使用 NLP 库屏蔽敏感数据。我们将在本演示中使用 spacy.io 软件包。</p><p>创建一个新文件<code>query_masking_nlp.py</code> 并添加以下代码。</p># query_masking_nlp.py

# pip install spacy
# python3 - m spacy download en_core_web_sm
import re
from typing import List, Optional

import spacy
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor.types import BaseNodePostprocessor
from llama_index.core.schema import NodeWithScore
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai import OpenAI
from index import es_vector_store

# Load the spaCy model
nlp = spacy.load("en_core_web_sm")

# Compile regex patterns for performance
phone_pattern = re.compile(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b')
email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b')
date_pattern = re.compile(r'\b(\d{1,2}[-/]\d{1,2}[-/]\d{2,4}|\d{2,4}[-/]\d{1,2}[-/]\d{1,2})\b')
dob_pattern = re.compile(
r"(January|February|March|April|May|June|July|August|September|October|November|December)\s(\d{1,2})(st|nd|rd|th),\s(\d{4})")
address_pattern = re.compile(r'\d+\s+[\w\s]+\,\s+[A-Za-z]+\,\s+[A-Z]{2}\s+\d{5}(-\d{4})?')
zip_code_pattern =  re.compile(r'\b\d{5}(?:-\d{4})?\b')
policy_number_pattern = re.compile(r"[A-Z]{3}\d{4}\.$")  # 3 characters followed by 4 digits, in our case e.g XYZ9876

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# match = re.match(policy_number_pattern, "XYZ9876")
# print(match)


def mask_pii(text):
   """
   Masks Personally Identifiable Information (PII) in the given
   text using pre-defined regex patterns and spaCy's named entity recognition.
   Args:
       text (str): The input text containing potential PII.
   Returns:
       str: The text with PII masked.
   """

   # Process the text with spaCy for NER
   doc = nlp(text)

   # Mask entities identified by spaCy NER (e.g First/Last Names etc)
   for ent in doc.ents:
       if ent.label_ in ["PERSON", "ORG", "GPE"]:
           text = text.replace(ent.text, '[MASKED]')

   # Apply regex patterns after NER to avoid overlapping issues
   text = phone_pattern.sub('[PHONE MASKED]', text)
   text = email_pattern.sub('[EMAIL MASKED]', text)
   text = date_pattern.sub('[DATE MASKED]', text)
   text = address_pattern.sub('[ADDRESS MASKED]', text)
   text = dob_pattern.sub('[DOB MASKED]', text)
   text = zip_code_pattern.sub('[ZIP MASKED]', text)
   text = policy_number_pattern.sub('[POLICY MASKED]', text)

   return text


class CustomPostProcessor(BaseNodePostprocessor):
   """
   Custom Postprocessor which masks Personally Identifiable Information (PII).
   PostProcessor is called on the Documents before they are sent to the LLM.
   """
   def _postprocess_nodes(
           self, nodes: List[NodeWithScore], query_bundle: Optional[QueryBundle]
   ) -&gt; List[NodeWithScore]:
       # Masks PII
       for n in nodes:
          n.node.set_content(mask_pii(n.text))
       return nodes

   
# Use Public LLM to send user query and Related Documents
llm = OpenAI()
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents are masked based on custom logic defined in CustomPostProcessor._postprocess_nodes.
query_engine = index.as_query_engine(llm, similarity_top_k=10, node_postprocessors=[CustomPostProcessor()])



query = "Give me summary of water related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
response = query_engine.query(bundle)
print(response)

<p>法律硕士的答复如下。</p>客户提出了各种与水有关的索赔，包括地下室水渍、水管爆裂、冰雹损坏屋顶以及暴雨期间的洪水等问题。这些索赔导致了基于缺乏及时通知、维护问题、逐渐磨损和预先存在的损坏等原因的拒赔而产生的挫折感。客户对这些拒赔表示失望、压力和经济负担，要求对其索赔进行公平评估和彻底审查。一些客户还面临索赔处理延迟的问题，这进一步引起了对保险公司服务的不满。<p>在上述代码中，当创建 Llama 索引查询引擎时，我们提供了一个 CustomPostProcessor。</p><p>QueryEngine 调用的逻辑在<code>CustomPostProcessor</code> 的<code>_postprocess_nodes</code> 方法中定义。我们正在使用 SpaCy.io 库来检测文件中的命名实体，然后使用一些正则表达式来替换这些名称以及敏感信息，然后再将文件发送到 LLM。</p><p>以下是自定义 PostProcessor 创建的原始会话和屏蔽会话的部分示例。</p><p>原文如此：</p>客户：你好，我是马修-洛佩兹（Matthew Lopez），出生日期是 1984 年 10 月 12 日，住在纽约州斯莫尔敦市雪松街 456 号，邮编 34567。我的保单号码是 TUV8901。探员下午好 马修我今天能为您提供什么帮助？客户：你好，我对贵公司拒绝我索赔的决定感到非常失望。<p>由 CustomPostProcessor 生成的屏蔽文本。</p>顾客：你好，我是 [蒙面]，[蒙面] 是 [生日蒙面]，我住在 34567 [蒙面] [蒙面] 的西达街 456 号。我的保单号码是 [屏蔽]。探员下午好，[蒙面]我今天能为您提供什么帮助？客户：你好，我对贵公司拒绝我索赔的决定感到非常失望。<p>请注意：</p><p><em>识别和屏蔽 PII 和敏感信息并不是一项简单的任务。要涵盖敏感信息的各种格式和语义，就必须充分了解自己的领域和数据。虽然上述代码可能适用于某些使用情况，但您可能需要根据自己的需求和测试情况进行修改。</em></p><h3>使用开箱即用的 LlamaIndex <code>NERPIINodePostprocessor</code></h3><p>LlamaIndex 通过引入以下功能，使保护 RAG 管道中的 PII 信息变得更加容易 <code>NERPIINodePostprocessor.</code></p>from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor import NERPIINodePostprocessor
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai import OpenAI
from index import es_vector_store

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# Use Public LLM to send user query and Related Documents
llm = OpenAI()

ner_processor = NERPIINodePostprocessor()
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents masked using the NERPIINodePostprocessor so that PII/Sensitive data is not sent to the LLM.
query_engine = index.as_query_engine(llm, similarity_top_k=10, node_postprocessors=[ner_processor])

query = "Give me summary of fire related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
response = query_engine.query(bundle)
print(response)
<p>答复如下</p>客户提出了与火灾有关的财产损失索赔。在一个案例中，由于纵火被排除在承保范围之外，车库火灾损失索赔被拒绝。另一位客户就其住宅遭受的火灾损失提出索赔，该损失属于其保单的承保范围。此外，一位客户报告了厨房火灾，并得到了火灾损失赔偿的保证。<h3>通过 <code>PIINodePostprocessor</code></h3><p>我们还可以利用本地或专用网络中运行的 LLM，在将数据发送到公共 LLM 之前完成屏蔽工作。</p><p>我们将使用运行在本地机器 Ollama 上的 Mistral 来进行屏蔽。</p><h4>本地运行 Mistral</h4><p>下载并安装<a href="https://ollama.com/">Ollama</a>。安装 Ollama 后，运行此命令下载并运行<a href="https://ollama.com/library/mistral">mistral</a></p>ollama run mistral
<p>首次下载并在本地运行模型可能需要几分钟时间。通过提出类似下面 "写一首关于云的诗 "的问题来验证 mistral 是否在运行，并验证诗歌是否符合您的要求。保持 ollama 运行，因为我们稍后需要通过代码与 mistral 模型交互。</p><p>新建一个名为<code>query_masking_local_LLM.py</code> 的文件，并添加以下代码。</p># pip install llama-index-llms-ollama
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor import PIINodePostprocessor
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.ollama import Ollama
from llama_index.llms.openai import OpenAI
from index import es_vector_store

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# Use Public LLM to send user query and Related Documents and Local LLM to mask
public_llm = OpenAI()
local_llm = Ollama(model="mistral")

pii_processor = PIINodePostprocessor(llm=local_llm)
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the public LLM.
# Note that documents are masked using the local llm via PIINodePostprocessor
# so that PII/Sensitive data is not sent to the public LLM.
query_engine = index.as_query_engine(public_llm, similarity_top_k=10, node_postprocessors=[pii_processor])


query = "Give me summary of fire related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>响应如下所示</p>客户提出了与火灾有关的财产损失索赔。在一个案例中，由于纵火被排除在承保范围之外，车库火灾损失索赔被拒绝。另一位客户就其住宅遭受的火灾损失提出索赔，该损失属于其保单的承保范围。此外，一位客户报告了厨房火灾，并得到了火灾损失赔偿的保证。<h3>结论</h3><p>在这篇文章中，我们介绍了在 RAG 流程中使用公共 LLM 时如何保护 PII 和敏感数据。我们展示了实现这一目标的多种方法。强烈建议在采用这些方法之前，根据您的用例和需求对其进行测试。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/rag-security-masking-pii</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/rag-security-masking-pii</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" length="0" type="image/png"/>
    <pubDate>Thu, 25 Jul 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[使用 LlamaIndex、Elasticsearch 和 Mistral 的 RAG（检索增强生成]]></title>
    <description><![CDATA[了解如何使用 LlamaIndex、Elasticsearch 和本地运行的 Mistral 实现 RAG（检索增强生成）系统。]]></description>
    <content:encoded><![CDATA[<p>在本博客中，我们将讨论如何使用 RAG 技术（检索增强生成）和作为向量数据库的 Elasticsearch 来实现 Q&amp;A 体验。我们将使用 LlamaIndex 和本地运行的 Mistral LLM。</p><p>在开始之前，我们先来了解一些术语。</p><h3>术语</h3><p><a href="https://www.llamaindex.ai/">LlamaIndex</a>是用于构建 LLM（大型语言模型）应用程序的领先数据框架。LlamaIndex 为构建 RAG（检索增强生成）应用程序的各个阶段提供了抽象概念。像 LlamaIndex 和 LangChain 这样的框架提供了抽象，因此应用程序不会与任何特定 LLM 的应用程序接口紧密耦合。</p><p><a href="https://www.elastic.co/enterprise-search">Elasticsearch</a>由<a href="https://elastic.co/">Elastic</a> 提供。Elastic 是 Elasticsearch 背后的行业领导者，Elasticsearch 是一个搜索和分析引擎，支持精确的全文搜索、语义理解的矢量搜索以及两全其美的混合搜索。Elasticsearch 是一种可扩展的数据存储和矢量数据库。我们在本博客中使用的 Elasticsearch 功能在 Elasticsearch 的免费开放版本中提供。</p><p><a href="https://www.promptingguide.ai/techniques/rag">检索增强生成（RAG）</a>是一种人工智能技术/模式，它为 LLM 提供外部知识，以生成对用户查询的回复。这样，法律硕士的答复就可以根据具体情况量身定做，答复也更加具体。</p><p><a href="https://docs.mistral.ai/">Mistral</a>提供开源和优化的企业级 LLM 模型。在本教程中，我们将使用可在笔记本电脑上运行的开源模型<a href="https://docs.mistral.ai/models/#mistral-7b">mistral-7b。</a>如果你不想在笔记本电脑上运行模型，也可以使用他们的云版本，在这种情况下，你必须修改本博客中的代码，以使用正确的 API 密钥和软件包。</p><p><a href="https://ollama.com/">Ollama</a>可帮助您在笔记本电脑上本地运行 LLM。我们将使用 Ollama 在本地运行开源的 Mistral-7b 模型。</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.13/semantic-search.html">嵌入</a>是文本/媒体意义的数字表示。它们是高维信息的低维表示。</p><h3>使用 LlamaIndex、Elasticsearch&amp; Mistral 构建 RAG 应用程序：场景概述</h3><p><strong>场景</strong></p><p>我们有一个样本数据集（JSON 文件），内容是一家虚构的家庭保险公司的座席人员与客户之间的呼叫中心对话。我们将建立一个简单的 RAG 应用程序，它可以回答以下问题</p><p><code>Give me summary of water related issues.</code></p><h3>高位流量</h3><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" alt="RAG 流程" /><p>我们使用 Ollama 在本地运行 Mistral LLM。</p><p>接下来，我们以<code>Documents</code> 的形式将 JSON 文件中的<em>对话</em>加载到<a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a>（由 Elasticsearch 支持的 VectorStore）中。在加载文档时，我们使用本地运行的 Mistral 模型创建嵌入。我们将这些嵌入和<em>对话</em>一起存储在 LlamaIndex Elasticsearch 向量<a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">存储（ElasticsearchStore</a>）中。</p><p>我们配置一个 LlamaIndex IngestionPipeline，并向其提供我们使用的本地 LLM，在本例中是通过 Ollama 运行的 Mistral。</p><p>当我们提出 "请简要介绍与水有关的问题 "这样的问题时、Elasticsearch 可进行语义搜索，并返回与水问题有关的<em>对话</em>。这些<em>对话</em>与原始问题一起发送给本地运行的 LLM，以生成答案。</p><h3>建立 RAG 应用程序的步骤</h3><h4>本地运行 Mistral</h4><p>下载并安装<a href="https://ollama.com/">Ollama</a>。安装 Ollama 后，运行此命令下载并运行<a href="https://ollama.com/library/mistral">mistral</a></p>ollama run mistral
<p>首次下载并在本地运行模型可能需要几分钟时间。通过提出类似下面 "写一首关于云的诗 "的问题来验证 mistral 是否在运行，并验证诗歌是否符合您的要求。保持 ollama 运行，因为我们稍后需要通过代码与 mistral 模型交互。</p><h4>安装 Elasticsearch</h4><p>通过创建云部署<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">（在此说明</a>）或在 docker 中运行<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/docker">（在此说明</a>），启动并运行 Elasticsearch。您也可以从<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/docker#self-hosted-production-deployments">这里</a>开始创建生产级的 Elasticsearch 自托管部署。</p><p>假设您使用的是云部署，请按照说明中的要求获取部署的 API 密钥和云 ID。我们稍后将使用它们。</p><h4>RAG 申请</h4><p>整个代码可在此<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany">Github 代码库中</a>找到，以供参考。克隆 repo 是可选项，因为我们将在下文中介绍代码。</p><p>在您最喜欢的集成开发环境中，用以下 3 个文件创建一个新的 Python 应用程序。</p><ul><li><p><code>index.py</code> 与索引数据有关的代码的位置。</p></li><li><p><code>query.py</code> 与查询和 LLM 交互相关的代码都放在这里。</p></li><li><p><code>.env</code> 配置属性（如 API 密钥）的位置。</p></li></ul><p>我们需要安装一些软件包。首先，我们要在应用程序的根文件夹中创建一个新的 python<a href="https://docs.python.org/3/library/venv.html">虚拟环境</a>。</p>python3 -m venv .venv
<p>激活虚拟环境并安装以下所需软件包。</p>source .venv/bin/activate
pip install llama-index 
pip install llama-index-embeddings-ollama
pip install llama-index-llms-ollama
pip install llama-index-vector-stores-elasticsearch
pip install sentence-transformers
pip install python-dotenv
<h4>索引数据</h4><p>下载<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/blob/main/conversations.json">conversations.json</a>文件，其中包含客户与我们 fictionaly 房屋保险公司呼叫中心座席之间的<em>对话</em>。将该文件与 2 个 python 文件和 .env 文件一起放在应用程序的根目录中。文件。下面是该文件内容的示例。</p>{
    "conversation_id": 103,
    "customer_name": "Sophia Jones",
    "agent_name": "Emily Wilson",
    "policy_number": "JKL0123",
    "conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
    "summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
<p>我们在<code>index.py</code> 中定义了一个名为<code>get_documents_from_file</code> 的函数，用于读取 json 文件并创建文档列表。<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/documents_and_nodes/">文档</a>对象是 LlamaIndex 处理信息的基本单位。</p># index.py
import json, os
from llama_index.core import Document, Settings
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.ingestion import IngestionPipeline
from llama_index.embeddings.ollama import OllamaEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore
from dotenv import load_dotenv

def get_documents_from_file(file):
   """Reads a json file and returns list of Documents"""

   with open(file=file, mode='rt') as f:
       conversations_dict = json.loads(f.read())
      
   # Build Document objects using fields of interest.
   documents = [Document(text=item['conversation'],
                         metadata={"conversation_id": item['conversation_id']})
                for
                item in conversations_dict]
   return documents
<p>创建摄取管道</p><p>首先，将在<code>Install Elasticsearch</code> 部分获得的 Elasticsearch CloudID 和 API 密钥添加到<code>.env</code> 文件中。<code>.env</code> 文件应如下所示（使用真实值）。</p>ELASTIC_CLOUD_ID=&lt;REPLACE WITH YOUR CLOUD ID&gt;
ELASTIC_API_KEY=&lt;REPLACE WITH YOUR API_KEY&gt;
<p>通过 LlamaIndex<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/ingestion_pipeline/">IngestionPipeline</a>，您可以使用多个组件组成一个管道。在<code>index.py</code> 文件中添加以下代码。</p># index.py

# Load .env file contents into env
# ELASTIC_CLOUD_ID and ELASTIC_API_KEY are expected to be in the .env file.
load_dotenv('.env')

# ElasticsearchStore is a VectorStore that
# takes care of ES Index and Data management.
es_vector_store = ElasticsearchStore(index_name="calls",
                                     vector_field='conversation_vector',
                                     text_field='conversation',
                                     es_cloud_id=os.getenv("ELASTIC_CLOUD_ID"),
                                     es_api_key=os.getenv("ELASTIC_API_KEY"))


def main():
    # Embedding Model to do local embedding using Ollama.
    ollama_embedding = OllamaEmbedding("mistral")

    # LlamaIndex Pipeline configured to take care of chunking, embedding
    # and storing the embeddings in the vector store.
    pipeline = IngestionPipeline(
        transformations=[
            SentenceSplitter(chunk_size=350, chunk_overlap=50),
            ollama_embedding,
        ],
        vector_store=es_vector_store
    )

    # Load data from a json file into a list of LlamaIndex Documents
    documents = get_documents_from_file(file="conversations.json")

    pipeline.run(documents=documents)
    print(".....Done running pipeline.....\n")


if __name__ == "__main__":
    main()

<p>如前所述，LlamaIndex IngestPipeline 可由多个组件组成。我们将在<code>pipeline = IngestionPipeline(...</code> 行的管道中添加 3 个组件。</p><ul><li><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/?h=sentencesp#sentencesplitter">SentenceSplitter</a>：从<code>get_documents_from_file()</code> 的定义中可以看出，每个文档都有一个文本字段，用来保存 json 文件中的对话内容。该文本字段是一段较长的文本。为了使语义搜索工作顺利进行，需要将其分解成小块文本。<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/?h=sentencesp#sentencesplitter">SentenceSplitter</a>类可以帮我们做到这一点。这些块在 LlamaIndex 术语中称为节点。节点中的元数据指向它们所属的文档。或者，也可以使用 Elasticsearch Ingestpipeline 进行分块，如本<a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">博客</a>所示。</p></li><li><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/models/embeddings/">OllamaEmbedding</a>：嵌入模型可将一段文字转换成数字（也称为向量）。有了数字表示法，我们就可以进行<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">语义搜索</a>，搜索结果与词义相匹配，而不仅仅是进行文本搜索。我们为 IngestionPipeline 提供<code>OllamaEmbedding("mistral")</code> 。使用 SentenceSplitter 分割的语块会通过 Ollama 发送到本地机器上运行的 Mistral 模型，然后 mistral 会为这些语块创建嵌入。</p></li><li><p><a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a>LlamaIndex ElasticsearchStore 向量存储会将正在创建的嵌入式内容备份到 Elasticsearch 索引中。ElasticsearchStore 负责创建和填充指定 Elasticsearch 索引的内容。在创建 ElasticsearchStore（由<code>es_vector_store</code> 引用）时，我们提供要创建的 Elasticsearch 索引的名称（本例中为<code>calls</code> ）、索引中要存储嵌入的字段（本例中为<code>conversation_vector</code> ）以及要存储文本的字段（本例中为<code>conversation</code> ）。总之，根据我们的配置<code>ElasticsearchStore</code> 在 Elasticsearch 中创建了一个新索引，并将<code>conversation_vector</code> 和<code>conversation</code> 作为字段（以及其他自动创建的字段）。</p></li></ul><p>将这一切联系起来，我们通过调用<code>pipeline.run(documents=documents)</code> 运行管道。</p><p>运行 index.py 脚本执行摄取管道：</p>python index.py
<p>管道运行完成后，我们应该会在 Elasticsearch 中看到一个名为<code>calls</code> 的新索引。使用开发控制台运行一个简单的 elasticsearch 查询，就能看到数据和嵌入式内容一起加载。</p>GET calls/_search?size=1
<p>概括地说，我们从 JSON 文件创建了文档，将文档分割成块，为这些块创建了嵌入，并将嵌入（和文本对话）存储在向量存储（ElasticsearchStore）中。</p><h4>查询</h4><p>llamaIndex<a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_guide/">VectorStoreIndex</a>可以让你检索相关文档和查询数据。默认情况下，VectorStoreIndex 会将嵌入存储在<a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_guide/">SimpleVectorStore</a> 的内存中。不过，也可以使用外部向量存储（如<a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a>）来使嵌入持久化。</p><p>打开<code>query.py</code> 并粘贴以下代码</p># query.py
from llama_index.core import VectorStoreIndex, QueryBundle, Response, Settings
from llama_index.embeddings.ollama import OllamaEmbedding
from llama_index.llms.ollama import Ollama
from index import es_vector_store

# Local LLM to send user query to
local_llm = Ollama(model="mistral")
Settings.embed_model= OllamaEmbedding("mistral")

index = VectorStoreIndex.from_vector_store(es_vector_store)
query_engine = index.as_query_engine(local_llm, similarity_top_k=10)

query="Give me summary of water related issues"
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>我们定义了一个本地 LLM (<code>local_llm</code>) 来指向在 Ollama 上运行的 Mistral 模型。接下来，我们从之前创建的 ElasticssearchStore 向量存储中创建一个 VectorStoreIndex (<code>index</code>)，然后从索引中获取一个查询引擎。在创建查询引擎时，我们会引用本地 LLM 来进行响应，我们还会提供 (<code>similarity_top_k=10</code>) 来配置应从向量存储中检索并发送给 LLM 以获得响应的文档数量。</p><p>运行<code>query.py</code> 脚本执行 RAG 流程：</p>python query.py
<p>我们发送查询<code>Give me summary of water related issues</code> （可随意定制<code>query</code> ），法律硕士的回复应与提供的相关文件类似。</p>在所提供的背景下，我们看到客户询问与水有关的损害的承保范围。在两个案例中，洪水对地下室造成了破坏，在另一个案例中，屋顶漏水也是问题所在。代理商确认，这两种水渍都在他们各自的保险范围内。因此，与水有关的问题，包括洪水和屋顶漏水，通常都在房屋保险的承保范围之内。<h4>注意事项</h4><p>这篇博文是关于使用 Elasticsearch 的 RAG 技术的初学者介绍，因此省略了一些功能配置，而这些功能配置将使您能够将这一起点应用到生产中。在构建生产用例时，您需要考虑更复杂的方面，如使用<a href="https://www.elastic.co/search-labs/blog/dls-internal-knowledge-search">文档级安全</a>保护数据，将数据分块作为 Elasticsearch<a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">Ingest 管道</a>的一部分，甚至在用于 GenAI/Chat/Q&amp;A 用例的相同数据上运行其他<a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-overview.html">ML 作业</a>。</p><p>您还可以考虑使用<a href="https://www.elastic.co/guide/en/enterprise-search/current/connectors.html">Elastic Connectors</a> 从各种外部来源（如 Azure Blob Storage、Dropbox、Gmail 等）获取数据并创建嵌入。</p><p>Elastic 可实现上述所有功能，并为 GenAI 用例及其他应用提供全面的企业级解决方案。</p><h4>接下来呢？</h4><ul><li><p>您可能已经注意到，我们正在将 10 个相关会话与用户问题一起发送给 LLM，以便制定回复。这些对话可能包含 PII（个人身份信息），如姓名、出生日期、地址等。在我们的案例中，LLM 是本地的，因此数据泄漏不是问题。但是，如果您想使用在云中运行的 LLM（例如 OpenAI），则不宜发送包含 PII 信息的文本。在后续博客中，我们将介绍如何在向 RAG 流程中的外部 LLM 发送 PII 信息之前完成屏蔽。</p></li><li><p>在本篇文章中，我们使用了本地 LLM，在接下来的文章 "在 RAG 中屏蔽 PII 数据 "中，我们将介绍如何轻松地从本地 LLM 切换到公共 LLM。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" length="0" type="image/png"/>
    <pubDate>Fri, 12 Apr 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>