<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Srikanth Manvi - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Srikanth Manvi - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/jp/search-labs/author/srikanth-manvi</link>
    </image>
    <link>https://www.elastic.co/jp/search-labs/author/srikanth-manvi</link>
    <atom:link href="https://www.elastic.co/jp/search-labs/rss/author/srikanth-manvi.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[jp]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 10:21:25 GMT</lastBuildDate>
  <item>
    <title><![CDATA[AIエージェント開発におけるMicrosoftセマンティックカーネル向けElasticsearch Vector Store Connectorの使い方]]></title>
    <description><![CDATA[Microsoft Semantic Kernel は、AI エージェントを簡単に構築し、最新の AI モデルを C#、Python、または Java コードベースに統合できる軽量のオープンソース開発キットです。Semantic Kernel Elasticsearch Vector Store Connector のリリースにより、AI エージェントの構築に Semantic Kernel を使用する開発者は、Semantic Kernel の抽象化を引き続き使用しながら、Elasticsearch をスケーラブルなエンタープライズ グレードのベクター ストアとしてプラグインできるようになりました。]]></description>
    <content:encoded><![CDATA[<p><a href="https://learn.microsoft.com/en-us/semantic-kernel/overview/">Microsoft Semantic Kernel</a> チームと連携して、<a href="https://learn.microsoft.com/en-us/semantic-kernel/overview/"> Microsoft Semantic</a> Kernel (.NET) ユーザー向けに<a href="https://github.com/elastic/semantic-kernel-net/"> Semantic Kernel Elasticsearch Vector Store Connector が利用可能になったことを発表します。</a>セマンティック カーネルは、ベクター ストアからのより関連性の高いデータ駆動型の応答を使用して大規模言語モデル (LLM) を強化する機能など、エンタープライズ グレードの AI エージェントの構築を簡素化します。Semantic Kernel は、Elasticsearch などの Vector Stores と対話するためのシームレスな抽象化レイヤーを提供し、レコードのコレクションの作成、一覧表示、削除や、個々のレコードのアップロード、取得、削除などの重要な機能を提供します。</p><p><a href="https://learn.microsoft.com/en-us/semantic-kernel/concepts/vector-store-connectors/out-of-the-box-connectors/elasticsearch-connector?pivots=programming-language-csharp">すぐに使用できるセマンティック カーネル Elasticsearch ベクター ストア コネクタは、</a>セマンティック カーネル<a href="https://learn.microsoft.com/en-us/semantic-kernel/concepts/vector-store-connectors/?pivots=programming-language-csharp#the-vector-store-abstraction">ベクター ストアの抽象化</a>をサポートしており、開発者は AI エージェントの構築時に Elasticsearch をベクター ストアとしてプラグインすることが非常に簡単になります。</p><p>Elasticsearch はオープンソース コミュニティに強固な基盤を持ち、最近<a href="https://www.elastic.co/blog/elasticsearch-is-open-source-again">AGPL ライセンスを</a>採用しました。これらのツールは、オープンソースの Microsoft Semantic Kernel と組み合わせることで、強力なエンタープライズ対応ソリューションを提供します。このコマンド<code>curl -fsSL https://elastic.co/start-local | sh </code> (詳細については<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">start-local</a>を参照) を実行して数分で Elasticsearch を起動し、ローカルで開始できます。その後、AI エージェントを本番稼働させながら、<a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;utm_source=semantickernel&amp;utm_content=documentation">クラウドホスト バージョン</a>または<a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.16/install-elasticsearch.html">セルフホスト</a>バージョンに移行できます。</p><p>このブログでは、Semantic Kernel を使用する際に<a href="https://github.com/elastic/semantic-kernel-net/">Semantic Kernel Elasticsearch Vector Store Connector を</a>使用する方法について説明します。コネクタの Python バージョンは将来提供される予定です。</p><h2>高レベルのシナリオ: Semantic Kernel と Elasticsearch を使用した RAG アプリの構築</h2><p>次のセクションでは例を見ていきます。大まかに言うと、ユーザーの質問を入力として受け取り、回答を返す RAG (Retrieval Augmented Generation) アプリケーションを構築しています。LLM として Azure OpenAI (<a href="https://devblogs.microsoft.com/semantic-kernel/introducing-new-ollama-connector-for-local-models/">ローカル LLM</a>も使用可能)、ベクター ストアとして Elasticsearch、すべてのコンポーネントを結び付けるフレームワークとして Semantic Kernel (.net) を使用します。</p><p>RAG アーキテクチャに精通していない場合は、次の記事で簡単に概要を把握できます: <a href="https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag">https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag</a> 。</p><p>回答は、Elasticsearch vectorstore から取得され、質問に関連するコンテキストが入力する LLM によって生成されます。応答には、LLM によってコンテキストとして使用されたソースも含まれます。</p><h3>RAGの例</h3><p>この具体的な例では、社内のホテル データベースに保存されているホテルについてユーザーが質問できるアプリケーションを構築します。ユーザーは例えばさまざまな基準に基づいて特定のホテルを検索したり、ホテルのリストを要求したりできます。</p><p>サンプル データベースでは、100 件のエントリを含む<a href="https://github.com/elastic/semantic-kernel-net/blob/main/Elastic.SemanticKernel.Playground/hotels.csv">ホテルのリスト</a>を生成しました。コネクタのデモをできるだけ簡単に試せるように、サンプル サイズは意図的に小さくなっています。実際のアプリケーションでは、特に非常に大量のデータを扱う場合、Elasticsearch コネクタは `InMemory` ベクトル ストア実装などの他のオプションよりも優位性を発揮します。</p><p>完全なデモ アプリケーションは、Elasticsearch ベクター ストア コネクタ<a href="https://github.com/elastic/semantic-kernel-net/tree/main/Elastic.SemanticKernel.Playground">リポジトリ</a>にあります。</p><p>まず、必要な NuGet パッケージと using ディレクティブをプロジェクトに追加することから始めましょう。</p>dotnet add package "Elastic.Clients.Elasticsearch" -v 8.16.2
dotnet add package "Elastic.SemanticKernel.Connectors.Elasticsearch" -v 0.1.2
dotnet add package "Microsoft.Extensions.Hosting" -v 9.0.0
dotnet add package "Microsoft.SemanticKernel.Connectors.AzureOpenAI" -v 1.30.0
dotnet add package "Microsoft.SemanticKernel.PromptTemplates.Handlebars" -v 1.30.0using System;
using System.IO;
using System.Linq;
using System.Threading.Tasks;

using Elastic.Clients.Elasticsearch;
using Elastic.Transport;

using Microsoft.Extensions.DependencyInjection;
using Microsoft.Extensions.Hosting;
using Microsoft.Extensions.VectorData;
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.Data;
using Microsoft.SemanticKernel.Embeddings;
using Microsoft.SemanticKernel.PromptTemplates.Handlebars;<p>これで、データ モデルを作成し、セマンティック カーネル固有の属性を指定して、ストレージ モデル スキーマとテキスト検索のヒントを定義できるようになりました。</p>/// &lt;summary&gt;
/// Data model for storing a "hotel" with a name, a description, a  description embedding and an optional reference link.
/// &lt;/summary&gt;
public sealed record Hotel
{
	[VectorStoreRecordKey]
	public required string HotelId { get; set; }

	[TextSearchResultName]
	[VectorStoreRecordData(IsFilterable = true)]
	public required string HotelName { get; set; }

	[TextSearchResultValue]
	[VectorStoreRecordData(IsFullTextSearchable = true)]
	public required string Description { get; set; }

	[VectorStoreRecordVector(Dimensions: 1536, DistanceFunction.CosineSimilarity, IndexKind.Hnsw)]
	public ReadOnlyMemory&lt;float&gt;? DescriptionEmbedding { get; set; }

	[TextSearchResultLink]
	[VectorStoreRecordData]
	public string? ReferenceLink { get; set; }
}<p>ストレージ モデル スキーマ属性 (`VectorStore*`) は、Elasticsearch Vector Store Connector の実際の使用に最も関連しています。具体的には次のようになります。</p><p></p><ul><li><p><code>VectorStoreRecordKey</code> レコード クラスのプロパティを、ベクトル ストアにレコードが格納されるキーとしてマークします。</p></li><li><p><code>VectorStoreRecordData</code> レコード クラスのプロパティを 'data' としてマークします。</p></li><li><p><code>VectorStoreRecordVector</code> レコード クラスのプロパティをベクトルとしてマークします。</p></li></ul><p>これらの属性はすべて、ストレージ モデルをさらにカスタマイズするために使用できるさまざまなオプション パラメーターを受け入れます。たとえば、 <code>VectorStoreRecordKey </code>の場合、異なる距離関数や異なるインデックス タイプを指定することが可能です。</p><p>テキスト検索属性 ( <code>TextSearch*</code> ) は、この例の最後のステップで重要になります。これらについては後ほど説明します。</p><p>次のステップでは、セマンティック カーネル エンジンを初期化し、コア サービスへの参照を取得します。実際のアプリケーションでは、サービス コレクションに直接アクセスするのではなく、<a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/dependency-injection">依存性注入を</a>使用する必要があります。同じことがハードコードされた構成とシークレットにも当てはまります。これらは、代わりに<a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/configuration">構成プロバイダー</a>を使用して読み取る必要があります。</p>var builder = Host.CreateApplicationBuilder(args);

// Register AI services.
var kernelBuilder = builder.Services.AddKernel();

kernelBuilder.AddAzureOpenAIChatCompletion("gpt-4o", "https://my-service.openai.azure.com", "my_token");

kernelBuilder.AddAzureOpenAITextEmbeddingGeneration("ada-002", "https://my-service.openai.azure.com", "my_token");

// Register text search service.
kernelBuilder.AddVectorStoreTextSearch&lt;Hotel&gt;();

// Register Elasticsearch vector store.
var elasticsearchClientSettings = new ElasticsearchClientSettings(new Uri("https://my-elasticsearch-instance.cloud"))
    .Authentication(new BasicAuthentication("elastic", "my_password"));

kernelBuilder.AddElasticsearchVectorStoreRecordCollection&lt;string, Hotel&gt;("skhotels", elasticsearchClientSettings);

// Build the host.
using var host = builder.Build();

// For demo purposes, we access the services directly without using a DI context.

var kernel = host.Services.GetService&lt;Kernel&gt;()!;
var embeddings = host.Services.GetService&lt;ITextEmbeddingGenerationService&gt;()!;
var vectorStoreCollection = host.Services.GetService&lt;IVectorStoreRecordCollection&lt;string, Hotel&gt;&gt;()!;

// Register search plugin.
var textSearch = host.Services.GetService&lt;VectorStoreTextSearch&lt;Hotel&gt;&gt;()!;
kernel.Plugins.Add(textSearch.CreateWithGetTextSearchResults("SearchPlugin"));<p><code>vectorStoreCollection</code>サービスを使用してコレクションを作成し、いくつかの<a href="https://github.com/elastic/semantic-kernel-net/blob/main/Elastic.SemanticKernel.Playground/hotels.csv">デモ レコード</a>を取り込むことができるようになりました。</p>await vectorStoreCollection.CreateCollectionIfNotExistsAsync();

// CSV format: ID;Hotel Name;Description;Reference Link
var hotels = (await File.ReadAllLinesAsync("hotels.csv"))
    .Select(x =&gt; x.Split(';'));

foreach (var chunk in hotels.Chunk(25))
{
    var descriptionEmbeddings = await embeddings.GenerateEmbeddingsAsync(chunk.Select(x =&gt; x[2]).ToArray());
    
    for (var i = 0; i &lt; chunk.Length; ++i)
    {
        var hotel = chunk[i];
        await vectorStoreCollection.UpsertAsync(new Hotel
        {
            HotelId = hotel[0],
            HotelName = hotel[1],
            Description = hotel[2],
            DescriptionEmbedding = descriptionEmbeddings[i],
            ReferenceLink = hotel[3]
        });
    }
}<p>これは、セマンティック カーネルが、複雑なベクトル ストアの使用を、いくつかの単純なメソッド呼び出しにまで削減する方法を示しています。</p><p>内部的には、Elasticsearch に新しいインデックスが作成され、必要なすべてのプロパティ マッピングが作成されます。その後、データ セットは完全に透過的にストレージ モデルにマッピングされ、最終的にインデックスに保存されます。以下は Elasticsearch でのマッピングの様子です。</p>{
  "mappings": {
    "properties": {
      "descriptionEmbedding": {
        "dims": 1536,
        "index": true,
        "index_options": {
          "type": "hnsw"
        },
        "similarity": "cosine",
        "type": "dense_vector"
      },
      "hotelName": {
        "type": "keyword"
      },
      "description": {
        "type": "text"
      }
    }
  }
}<p><code>embeddings.GenerateEmbeddingsAsync()</code>は、構成された Azure AI Embeddings Generation サービスを透過的に呼び出しました。</p><p>このデモの最後のステップでは、さらに多くの魔法が観察できます。</p><p><code>InvokePromptAsync</code>を 1 回呼び出すだけで、ユーザーがデータについて質問したときに、次のすべての操作が実行されます。</p><p>1.ユーザーの質問の埋め込みが生成される</p><p>2. ベクトルストアで関連するエントリを検索する</p><p>3. クエリの結果はプロンプトテンプレートに挿入されます</p><p>4. 最終プロンプトの形式で実際のクエリがAIチャット補完サービスに送信されます。</p>// Invoke the LLM with a template that uses the search plugin to
// 1. get related information to the user query from the vector store
// 2. add the information to the LLM prompt.
var response = await kernel.InvokePromptAsync(
    promptTemplate: """
                    Please use this information to answer the question:
                    {{#with (SearchPlugin-GetTextSearchResults question)}}
                      {{#each this}}
                        Name: {{Name}}
                        Value: {{Value}}
                        Source: {{Link}}
                        -----------------
                      {{/each}}
                    {{/with}}
                    
                    Include the source of relevant information in the response.

                    Question: {{question}}
                    """,
    arguments: new KernelArguments
    {
        { "question", "Please show me all hotels that have a rooftop bar." },
    },
    templateFormat: "handlebars",
    promptTemplateFactory: new HandlebarsPromptTemplateFactory());<p>以前データ モデルで定義した<code>TextSearch*</code>属性を覚えていますか?これらの属性により、プロンプト テンプレート内の対応するプレースホルダーを使用できるようになります。これらのプレースホルダーには、ベクター ストア内のエントリからの情報が自動的に入力されます。</p><p>「屋上バーがあるホテルをすべて教えてください。」という質問に対する最終的な回答は次のとおりです。</p>Console.WriteLine(response.ToString());

// &gt; The hotel that has a rooftop bar is Skyline Suites. You can find more information about this hotel [here](https://example.com/yz567).<p>答えは、hotels.csvの次のエントリを正しく参照しています。</p>9;
Skyline Suites;
Offering panoramic city views from every suite, this hotel is perfect for those who love the urban landscape. Enjoy luxurious amenities, a rooftop bar, and close proximity to attractions. Luxurious and contemporary.;
https://example.com/yz567<p>この例は、Microsoft Semantic Kernel を使用すると、よく考えられた抽象化によって複雑さが大幅に軽減され、非常に高いレベルの柔軟性が実現されることを示しています。たとえば、コードの 1 行を変更するだけで、コードの他の部分をリファクタリングすることなく、使用されているベクトル ストアまたは AI サービスを置き換えることができます。</p><p>同時に、このフレームワークは、`InvokePrompt` 関数やテンプレート、検索プラグイン システムなどの膨大な高レベル機能を提供します。</p><p>完全なデモ アプリケーションは、Elasticsearch ベクター ストア コネクタ リポジトリにあります。</p><h2>Elasticsearchで他に何ができるのか</h2><ul><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">Elasticsearchの新しいsemantic_textマッピング：セマンティック検索の簡素化</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-reranking-with-retrievers">Elasticsearch におけるセマンティックリランキング（リトリーバー使用）</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">高度なRAGテクニックパート1：データ処理</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">高度なRAGテクニックパート2：クエリとテスト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-rag-with-llama3-opensource-and-elastic">Llama 3オープンソースとElasticでRAGを構築する</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/local-rag-agent-elasticsearch-langgraph-llama3">LangGraph、LLaMA3、Elasticsearchベクターストアを使用してローカルエージェントをゼロから構築するチュートリアル</a></p></li></ul><h2>Elasticsearch とセマンティックカーネル: 次は何?</h2><ul><li><p>.NET で GenAI アプリケーションを構築する際に、Elasticsearch ベクター ストアを Semantic Kernel に簡単にプラグインする方法を示しました。次回の Python 統合にご期待ください。</p></li><li><p>Semantic Kernel は<a href="https://www.elastic.co/search-labs/tutorials/search-tutorial/vector-search/hybrid-search">ハイブリッド検索</a>などの高度な検索機能の抽象化を構築するため、Elasticsearch Connect を使用すると、.NET 開発者は Semantic Kernel を使用しながらそれらを簡単に実装できるようになります。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-connector-microsoft-semantic-kernel</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-connector-microsoft-semantic-kernel</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[.NET]]></category>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Florian Bernd,Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2d8725035e86f8a8/6a17fe447f6f1564f8c09d74/0564fe794e4c66d0507317822d7aa71826183d20-1311x762.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 06 Dec 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch と LlamaIndex を使用して RAG の機密情報と PII 情報を保護する]]></title>
    <description><![CDATA[Elasticsearch と LlamaIndex を使用して RAG アプリケーション内の機密データと PII データを保護する方法。]]></description>
    <content:encoded><![CDATA[<p></p><p></p><p>この記事では、RAG (Retrieval Augmented Generation) フローでパブリック LLM を使用する際に、個人識別情報 (PII) と機密データを保護する方法について説明します。オープンソース ライブラリと正規表現を使用して PII と機密データをマスキングする方法と、パブリック LLM を呼び出す前にローカル LLM を使用してデータをマスキングする方法を検討します。</p><p>始める前に、この投稿で使用するいくつかの用語を確認しましょう。</p><h2>用語について</h2><p><a href="https://www.llamaindex.ai/">LlamaIndex は</a>、LLM (大規模言語モデル) アプリケーションを構築するための主要なデータ フレームワークです。LlamaIndex は、RAG (Retrieval Augmented Generation) アプリケーションの構築のさまざまな段階に抽象化を提供します。LlamaIndex や LangChain などのフレームワークは、アプリケーションが特定の LLM の API に密結合されないように抽象化を提供します。</p><p><a href="https://www.elastic.co/enterprise-search">Elasticsearch</a>は<a href="https://elastic.co/">Elastic</a>によって提供されています。Elastic は、精度の高い全文検索、意味理解のためのベクトル検索、両方の長所を生かしたハイブリッド検索をサポートするスケーラブルなデータ ストアおよびベクトル データベースである Elasticsearch を提供する業界リーダーです。Elasticsearch は、分散型の RESTful 検索および分析エンジン、スケーラブルなデータ ストア、およびベクター データベースです。このブログで使用している Elasticsearch 機能は、Elasticsearch の無料およびオープン バージョンで利用できます。</p><p><a href="https://www.promptingguide.ai/techniques/rag">検索拡張生成 (RAG)</a>は、LLM に外部知識を提供してユーザーのクエリに対する応答を生成する AI テクニック/パターンです。これにより、LLM 応答を特定のコンテキストに合わせてカスタマイズし、汎用性を抑えることができます。</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.13/semantic-search.html">埋め込みは</a>、テキスト/メディアの意味を数値的に表現したものです。それらは高次元情報の低次元表現です。</p><h2>RAGとデータ保護</h2><p>一般的に、大規模言語モデル (LLM) は、インターネット データでトレーニングできるモデル内で利用可能な情報に基づいて応答を生成するのに適しています。ただし、モデル内で情報が得られないクエリの場合、LLM にはモデル内に含まれていない外部知識または特定の詳細が提供される必要がありま す。このような情報は、データベースまたは社内のナレッジ システム内にある可能性があります。検索拡張生成 (RAG) は、特定のユーザー クエリに対して、まず外部 (LLM に対して) システム (データベースなど) から関連するコンテキスト/情報を取得し、そのコンテキストをユーザー クエリとともに LLM に送信して、より具体的で関連性の高い応答を生成する手法です。</p><p>これにより、RAG テクニックは、質問への回答、コンテンツの作成、コンテキストと詳細の深い理解が役立つあらゆるアプリケーションに非常に効果的になります。</p><p>その結果、RAG パイプラインでは、PII (個人識別情報) などの内部情報や機密情報 (名前、生年月日、口座番号など) がパブリック LLM に公開されるリスクがあります。</p><p>Elasticsearch のようなベクター データベースを使用する場合、データは安全です (<a href="https://www.elastic.co/guide/en/cloud-enterprise/current/ece-configure-rbac.html">ロール ベースのアクセス制御</a>、<a href="https://www.elastic.co/search-labs/blog/dls-internal-knowledge-search">ドキュメント レベルのセキュリティ</a>などのさまざまな手段を通じて)。ただし、データを外部のパブリック LLM に送信する場合は注意が必要です。</p><p>大規模言語モデル (LLM) を使用する場合、個人を特定できる情報 (PII) と機密データを保護することは、いくつかの理由から重要です。</p><ul><li><p><strong>プライバシーコンプライアンス</strong>: 多くの地域では、ヨーロッパの一般データ保護規則 (GDPR) や米国のカリフォルニア州消費者プライバシー法 (CCPA) など、個人データの保護を義務付ける厳格な規制があります。法的責任や罰金を回避するには、これらの法律を遵守する必要があります。</p></li><li><p><strong>ユーザーの信頼</strong>: 機密情報の機密性と整合性を確保することで、ユーザーの信頼が構築されます。ユーザーは、自分のプライバシーが保護されると信じるシステムを使用したり、やり取りしたりする可能性が高くなります。</p></li><li><p><strong>データ セキュリティ</strong>: データ侵害に対する保護が不可欠です。適切な保護措置を講じずに LLM に公開された機密データは盗難や悪用される可能性があり、個人情報の盗難や金融詐欺などの潜在的な危害につながる可能性があります。</p></li><li><p><strong>倫理的な考慮事項</strong>: 倫理的には、ユーザーのプライバシーを尊重し、ユーザーのデータを責任を持って扱うことが重要です。個人情報を不適切に取り扱うと、差別、汚名、その他の社会的悪影響が生じる可能性があります。</p></li><li><p><strong>企業の評判</strong>: 機密データを保護できない企業は評判が損なわれる可能性があり、顧客や収益の喪失など、ビジネスに長期的な悪影響を及ぼす可能性があります。</p></li><li><p><strong>悪用リスクの軽減</strong>: 機密データを安全に扱うことで、偏ったデータでモデルをトレーニングしたり、データを使用して個人を操作したり害を与えたりするなど、データやモデルの悪意のある使用を防ぐことができます。</p></li></ul><p>全体として、PII と機密データの堅牢な保護は、法令遵守の確保、ユーザーの信頼の維持、データ セキュリティの確保、倫理基準の遵守、ビジネスの評判の保護、不正使用のリスクの軽減に必要です。</p><h2>簡単な要約</h2><p><a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch">前回の投稿</a>では、LlamaIndex とローカルで実行される Mistral LLM を使用しながら、Elasticsearch をベクター データベースとして RAG テクニックを使用して Q&amp;A エクスペリエンスを実装する方法について説明しました。ここではそれを基に構築します。</p><p>前回の投稿を読むことはオプションです。今回は前回の投稿で行った内容を簡単に説明/要約します。</p><p>架空の住宅保険会社のエージェントと顧客間のコールセンターの会話のサンプル データセットがありました。私たちは、「顧客はどのような水関連の問題について請求をしているのか？」といった質問に答えるシンプルな RAG アプリケーションを構築しました。</p><p>大まかに言うと、フローは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" alt="RAGフロー" /><p>インデックス作成フェーズでは、LlamaIndex パイプラインを使用してドキュメントを読み込み、インデックスを作成しました。ドキュメントはチャンク化され、埋め込みとともに Elasticsearch ベクター データベースに保存されました。</p><p>ユーザーが質問したクエリフェーズで、LlamaIndex はクエリに関連する上位 K 件の類似ドキュメントを取得しました。これらの上位 K 件の関連ドキュメントはクエリとともに、ローカルで実行されている Mistral LLM に送信され、ユーザーに返される応答が生成されました。ぜひ前回の投稿をご覧になったり、<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/tree/main">コードを調べたりしてください</a>。</p><p>前回の投稿では、LLM をローカルで実行しました。ただし、本番環境では、 <a href="https://openai.com/">OpenAI</a> 、 <a href="https://mistral.ai/">Mistral</a> 、 <a href="https://www.anthropic.com/claude">Anthropic</a>などのさまざまな企業が提供する外部 LLM を使用する必要がある場合があります。ユースケースでより大きな基礎モデルが必要であるか、スケーラビリティ、可用性、パフォーマンスなどのエンタープライズ生産のニーズによりローカルでの実行がオプションではないことが原因である可能性があります。</p><p>RAG パイプラインに外部 LLM を導入すると、機密情報や PII が LLM に誤って漏洩するリスクが生じます。この投稿では、ドキュメントを外部 LLM に送信する前に、RAG パイプラインの一部として PII 情報をマスクする方法について説明します。</p><h2>公立LLM取得のRAG</h2><p>RAG パイプラインで PII と機密情報を保護する方法について説明する前に、まず LlamaIndex、Elasticsearch Vector データベース、OpenAI LLM を使用してシンプルな RAG アプリケーションを構築します。</p><h3>要件</h3><p>以下のものが必要となります。</p><ul><li><p>埋め込みを保存するためのベクター データベースとして<strong>Elasticsearch を</strong>起動して実行します。<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#install-elasticsearch">Elasticsearch のインストール</a>に関する前回の投稿の手順に従ってください。</p></li><li><p>AI API キーを開きます。</p></li></ul><h3>シンプルなRAGアプリケーション</h3><p>参考までに、コード全体はこの<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/tree/protecting-pii">Githubリポジトリ</a>（branch:protecting-pii）にあります。以下のコードを見ていくので、リポジトリのクローンは任意です。</p><p>お気に入りの IDE で、以下の 3 つのファイルを使用して新しい Python アプリケーションを作成します。</p><ul><li><p><code>index.py</code> データのインデックス作成に関連するコードが配置される場所。</p></li><li><p><code>query.py</code> クエリと LLM の相互作用に関連するコードが配置される場所。</p></li><li><p><code>.env</code> API キーなどの構成プロパティが配置される場所。</p></li></ul><p>いくつかのパッケージをインストールする必要があります。まず、アプリケーションのルート フォルダーに新しい Python<a href="https://docs.python.org/3/library/venv.html">仮想環境</a>を作成します。</p>python3 -m venv .venv
<p>仮想環境をアクティブ化し、以下の必要なパッケージをインストールします。</p>source .venv/bin/activate
pip install llama-index 
pip install llama-index-embeddings-openai
pip install llama-index-vector-stores-elasticsearch
pip install sentence-transformers
pip install python-dotenv
pip install openai
<p>.envでOpenAIとElasticsearchの接続プロパティを設定するファイル。</p>OPENAI_API_KEY="REPLACEME"
ELASTIC_CLOUD_ID="REPLACEME"
ELASTIC_API_KEY="REPLACEME"
<h4>データのインデックス作成</h4><p>架空の住宅保険会社の顧客とコールセンターエージェント間の<em> 会話</em> が含まれる<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/blob/main/conversations.json"> conversations.json ファイルをダウンロードします。</a>アプリケーションのルートディレクトリに、2つのPythonファイルと.envファイルと一緒にファイルを配置します。先ほど作成したファイル。以下はファイルの内容の例です。</p>{
"conversation_id": 103,
"customer_name": "Sophia Jones",
"agent_name": "Emily Wilson",
"policy_number": "JKL0123",
"conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
"summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
<p>データのインデックス作成を処理する以下のコードを<code>index.py</code>に貼り付けます。</p># index.py
# pip install sentence-transformers
# pip install llama-index-embeddings-openai
# pip install llama-index-embeddings-huggingface

import json
import os
from dotenv import load_dotenv
from llama_index.core import Document
from llama_index.core import Settings
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore


def get_documents_from_file(file):
   """Reads a json file and returns list of Documents"""

   with open(file=file, mode='rt') as f:
       conversations_dict = json.loads(f.read())

   # Build Document objects using fields of interest.
   documents = [Document(text=item['conversation'],
                         metadata={"conversation_id": item['conversation_id']})
                for
                item in conversations_dict]
   return documents

# Load .env file contents into env
load_dotenv('.env')
Settings.embed_model = HuggingFaceEmbedding(
   model_name="BAAI/bge-small-en-v1.5"
)

def main():
   # ElasticsearchStore is a VectorStore that
   # takes care of Elasticsearch Index and Data management.
   es_vector_store = ElasticsearchStore(index_name="convo_index",
                                        vector_field='conversation_vector',
                                        text_field='conversation',
                                        es_cloud_id=os.getenv("ELASTIC_CLOUD_ID"),
                                        es_api_key=os.getenv("ELASTIC_API_KEY"))

   # LlamaIndex Pipeline configured to take care of chunking, embedding
   # and storing the embeddings in the vector store.
   llamaindex_pipeline = IngestionPipeline(
       transformations=[
           SentenceSplitter(chunk_size=350, chunk_overlap=50),
           Settings.embed_model
       ],
       vector_store=es_vector_store
   )

   # Load data from a json file into a list of LlamaIndex Documents
   documents = get_documents_from_file(file="conversations.json")
   llamaindex_pipeline.run(documents=documents)
   print(".....Indexing Data Completed.....\n")

if __name__ == "__main__":
   main()
<p>上記のコードを実行すると、Elasticsearch にインデックスが作成され、 <code>convo_index</code>という名前の Elasticsearch インデックスに埋め込みが保存されます。</p><p>LlamaIndex IngestionPipeline に関する説明が必要な場合は、 <a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#indexing-data">IngestionPipeline の作成</a>セクションの前の投稿を参照してください。</p><h4>クエリ</h4><p>前回の投稿では、<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#querying">クエリ</a>にローカル LLM を使用しました。</p><p>この投稿では、以下に示すように、公開 LLM である OpenAI を使用します。</p># query.py
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.llms.openai import OpenAI
from index import es_vector_store

# Public LLM where we send user query and Related Documents
llm = OpenAI()

index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents are sent as-is. So any PII/Sensitive data is sent to the LLM.
query_engine = index.as_query_engine(llm, similarity_top_k=10)

query="Give me summary of water related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>上記のコードは、OpenAI からの応答を以下のように出力します。</p><p>顧客は、地下室の水害、水道管の破裂、屋根への雹害、適時通知の欠如、メンテナンスの問題、徐々に進行する消耗、既存の損傷などの理由による請求の拒否など、さまざまな水関連の請求を提起しています。いずれの場合も、顧客は請求の却下に対する不満を表明し、請求に関する公正な評価と決定を求めました。</p><h2>RAG での PII のマスキング</h2><p>これまで説明してきたのは、ユーザークエリとともにドキュメントをそのまま OpenAI に送信することです。</p><p>RAG パイプラインでは、関連するコンテキストが Vector ストアから取得された後、クエリとコンテキストを LLM に送信する前に、PII と機密情報をマスクする機会があります。</p><p>外部 LLM に送信する前に PII 情報をマスクする方法はいくつかあり、それぞれにメリットがあります。以下のオプションをいくつか見てみましょう</p><ol><li><p>spacy.io や<a href="https://microsoft.github.io/presidio/">Presidio</a> (Microsoft が管理するオープン ソース ライブラリ) などの NLP ライブラリを使用します。</p></li><li><p>LlamaIndexをそのまま使用する <code>NERPIINodePostprocessor.</code></p></li><li><p>ローカルLLMの使用 <code>PIINodePostprocessor</code></p></li></ol><p>上記のいずれかの方法を使用してマスキング ロジックを実装したら、PostProcessor (独自のカスタム PostProcessor または LlamaIndex が提供するすぐに使用できる PostProcessor) を使用して LlamaIndex IngestionPipeline を構成できます。</p><h3>NLPライブラリの使用</h3><p>RAG パイプラインの一部として、NLP ライブラリを使用して機密データをマスクできます。このデモでは、spacy.io パッケージを使用します。</p><p>新しいファイル<code>query_masking_nlp.py</code>を作成し、以下のコードを追加します。</p># query_masking_nlp.py

# pip install spacy
# python3 - m spacy download en_core_web_sm
import re
from typing import List, Optional

import spacy
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor.types import BaseNodePostprocessor
from llama_index.core.schema import NodeWithScore
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai import OpenAI
from index import es_vector_store

# Load the spaCy model
nlp = spacy.load("en_core_web_sm")

# Compile regex patterns for performance
phone_pattern = re.compile(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b')
email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b')
date_pattern = re.compile(r'\b(\d{1,2}[-/]\d{1,2}[-/]\d{2,4}|\d{2,4}[-/]\d{1,2}[-/]\d{1,2})\b')
dob_pattern = re.compile(
r"(January|February|March|April|May|June|July|August|September|October|November|December)\s(\d{1,2})(st|nd|rd|th),\s(\d{4})")
address_pattern = re.compile(r'\d+\s+[\w\s]+\,\s+[A-Za-z]+\,\s+[A-Z]{2}\s+\d{5}(-\d{4})?')
zip_code_pattern =  re.compile(r'\b\d{5}(?:-\d{4})?\b')
policy_number_pattern = re.compile(r"[A-Z]{3}\d{4}\.$")  # 3 characters followed by 4 digits, in our case e.g XYZ9876

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# match = re.match(policy_number_pattern, "XYZ9876")
# print(match)


def mask_pii(text):
   """
   Masks Personally Identifiable Information (PII) in the given
   text using pre-defined regex patterns and spaCy's named entity recognition.
   Args:
       text (str): The input text containing potential PII.
   Returns:
       str: The text with PII masked.
   """

   # Process the text with spaCy for NER
   doc = nlp(text)

   # Mask entities identified by spaCy NER (e.g First/Last Names etc)
   for ent in doc.ents:
       if ent.label_ in ["PERSON", "ORG", "GPE"]:
           text = text.replace(ent.text, '[MASKED]')

   # Apply regex patterns after NER to avoid overlapping issues
   text = phone_pattern.sub('[PHONE MASKED]', text)
   text = email_pattern.sub('[EMAIL MASKED]', text)
   text = date_pattern.sub('[DATE MASKED]', text)
   text = address_pattern.sub('[ADDRESS MASKED]', text)
   text = dob_pattern.sub('[DOB MASKED]', text)
   text = zip_code_pattern.sub('[ZIP MASKED]', text)
   text = policy_number_pattern.sub('[POLICY MASKED]', text)

   return text


class CustomPostProcessor(BaseNodePostprocessor):
   """
   Custom Postprocessor which masks Personally Identifiable Information (PII).
   PostProcessor is called on the Documents before they are sent to the LLM.
   """
   def _postprocess_nodes(
           self, nodes: List[NodeWithScore], query_bundle: Optional[QueryBundle]
   ) -&gt; List[NodeWithScore]:
       # Masks PII
       for n in nodes:
          n.node.set_content(mask_pii(n.text))
       return nodes

   
# Use Public LLM to send user query and Related Documents
llm = OpenAI()
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents are masked based on custom logic defined in CustomPostProcessor._postprocess_nodes.
query_engine = index.as_query_engine(llm, similarity_top_k=10, node_postprocessors=[CustomPostProcessor()])



query = "Give me summary of water related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
response = query_engine.query(bundle)
print(response)

<p>LLM による応答を以下に示します。</p>顧客からは、地下室の水害、水道管の破裂、屋根への雹害、大雨による浸水など、さまざまな水関連のクレームが出ています。これらの請求は、適時の通知の欠如、メンテナンスの問題、徐々に進行する消耗、既存の損傷などの理由に基づいて請求が拒否されたために、フラストレーションを招いています。顧客は、こうした請求拒否の結果、失望、ストレス、経済的負担を感じており、請求の公正な評価と徹底的な見直しを求めています。一部の顧客は保険金請求処理の遅延にも直面しており、保険会社が提供するサービスに対するさらなる不満が生じています。<p>上記のコードでは、Llama Index QueryEngine を作成するときに CustomPostProcessor を指定します。</p><p>QueryEngine によって呼び出されるロジックは、 <code>CustomPostProcessor</code>の<code>_postprocess_nodes</code>メソッドで定義されています。私たちは SpaCy.io ライブラリを使用して名前付きエンティティを検出し、ドキュメントを LLM に送信する前に、いくつかの正規表現を使用してそれらの名前と機密情報を置き換えます。</p><p>以下に、元の会話の一部と、CustomPostProcessor によって作成されたマスクされた会話の例を示します。</p><p>原文:</p>顧客: こんにちは。私はマシュー ロペスです。生年月日は 1984 年 10 月 12 日、住所は 456 Cedar St, Smalltown, NY 34567 です。私の保険証券番号はTUV8901です。エージェント: こんにちは、マシュー。本日はどのようなご用件でしょうか？顧客: こんにちは。私の請求を却下するという貴社の決定に、大変失望しております。<p>CustomPostProcessor によってマスクされたテキスト。</p>顧客: こんにちは。私は [MASKED] です。[MASKED] は [DOB MASKED] で、456 Cedar St, [MASKED], [MASKED] 34567 に住んでいます。私の保険証券番号は[MASKED]です。エージェント: こんにちは、[MASKED]。本日はどのようなご用件でしょうか？顧客: こんにちは。私の請求を却下するという貴社の決定に、大変失望しております。<p>注記：</p><p><em>個人情報や機密情報を識別してマスキングすることは簡単な作業ではありません。機密情報のさまざまな形式とセマンティクスをカバーするには、ドメインとデータに関する十分な理解が必要です。上記のコードは一部のユースケースでは機能する可能性がありますが、ニーズとテストに基づいて変更する必要がある場合もあります。</em></p><h3>LlamaIndexをそのまま使用する <code>NERPIINodePostprocessor</code></h3><p>LlamaIndexは、RAGパイプラインでPII情報を保護しやすくするために、 <code>NERPIINodePostprocessor.</code></p>from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor import NERPIINodePostprocessor
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai import OpenAI
from index import es_vector_store

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# Use Public LLM to send user query and Related Documents
llm = OpenAI()

ner_processor = NERPIINodePostprocessor()
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents masked using the NERPIINodePostprocessor so that PII/Sensitive data is not sent to the LLM.
query_engine = index.as_query_engine(llm, similarity_top_k=10, node_postprocessors=[ner_processor])

query = "Give me summary of fire related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
response = query_engine.query(bundle)
print(response)
<p>応答は以下のとおりです</p>顧客は、火災により所有物件に損害が生じたとして損害賠償請求を起こした。あるケースでは、放火が補償範囲外であったため、ガレージの火災による損害に対する請求が却下されました。別の顧客は、保険でカバーされていた自宅の火災による損害について請求をしました。さらに、顧客がキッチンの火災を報告し、火災による損害が補償されることが保証されました。<h3>ローカルLLMの使用 <code>PIINodePostprocessor</code></h3><p>また、データをパブリック LLM に送信する前に、ローカルまたはプライベート ネットワークで実行されている LLM を活用してマスキング作業を行うこともできます。</p><p>マスキングを行うには、ローカル マシン上の Ollama で実行されている Mistral を使用します。</p><h4>Mistralをローカルで実行する</h4><p><a href="https://ollama.com/">Ollama</a>をダウンロードしてインストールします。Ollamaをインストールした後、このコマンドを実行して<a href="https://ollama.com/library/mistral">mistralを</a>ダウンロードして実行します。</p>ollama run mistral
<p>モデルを初めてローカルにダウンロードして実行するには数分かかる場合があります。以下のような「雲についての詩を書いてください」という質問をして、ミストラルが実行されているかどうかを確認し、その詩が気に入ったものかどうかを確認します。後でコードを通じてミストラル モデルとやり取りする必要があるため、ollama を実行したままにしておきます。</p><p><code>query_masking_local_LLM.py</code>という新しいファイルを作成し、以下のコードを追加します。</p># pip install llama-index-llms-ollama
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor import PIINodePostprocessor
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.ollama import Ollama
from llama_index.llms.openai import OpenAI
from index import es_vector_store

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# Use Public LLM to send user query and Related Documents and Local LLM to mask
public_llm = OpenAI()
local_llm = Ollama(model="mistral")

pii_processor = PIINodePostprocessor(llm=local_llm)
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the public LLM.
# Note that documents are masked using the local llm via PIINodePostprocessor
# so that PII/Sensitive data is not sent to the public LLM.
query_engine = index.as_query_engine(public_llm, similarity_top_k=10, node_postprocessors=[pii_processor])


query = "Give me summary of fire related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>応答は以下のようなものです</p>顧客は、火災により所有物件に損害が生じたとして損害賠償請求を起こした。あるケースでは、放火が補償範囲外であったため、ガレージの火災による損害に対する請求が却下されました。別の顧客は、保険でカバーされていた自宅の火災による損害について請求をしました。さらに、顧客がキッチンの火災を報告し、火災による損害が補償されることが保証されました。<h3>まとめ</h3><p>この投稿では、RAG フロー内でパブリック LLM を使用する際に PII と機密データを保護する方法を説明しました。私たちはそれを実現する複数の方法を実証しました。採用する前に、ユースケースとニーズに基づいてこれらのアプローチをテストすることを強くお勧めします。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/rag-security-masking-pii</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/rag-security-masking-pii</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" length="0" type="image/png"/>
    <pubDate>Thu, 25 Jul 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[LlamaIndex、Elasticsearch、Mistral を使用した RAG (検索拡張生成)]]></title>
    <description><![CDATA[LlamaIndex、Elasticsearch、ローカルで実行される Mistral を使用して RAG (Retrieval Augmented Generation) システムを実装する方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>このブログでは、Elasticsearch をベクター データベースとして使用し、RAG テクニック (Retrieval Augmented Generation) を使用して Q&amp;A エクスペリエンスを実装する方法について説明します。LlamaIndex とローカルで実行されている Mistral LLM を使用します。</p><p>始める前に、いくつかの用語を見てみましょう。</p><h3>用語について</h3><p><a href="https://www.llamaindex.ai/">LlamaIndex は</a>、LLM (大規模言語モデル) アプリケーションを構築するための主要なデータ フレームワークです。LlamaIndex は、RAG (Retrieval Augmented Generation) アプリケーションの構築のさまざまな段階に抽象化を提供します。LlamaIndex や LangChain などのフレームワークは、アプリケーションが特定の LLM の API に密結合されないように抽象化を提供します。</p><p><a href="https://www.elastic.co/enterprise-search">Elasticsearch</a>は<a href="https://elastic.co/">Elastic</a>によって提供されています。Elastic は、精度の高い全文検索、意味理解のためのベクトル検索、両方の長所を生かしたハイブリッド検索をサポートする検索および分析エンジンである Elasticsearch を提供する業界リーダーです。Elasticsearch はスケーラブルなデータ ストアおよびベクター データベースです。このブログで使用している Elasticsearch の機能は、Elasticsearch の無料およびオープン バージョンで利用できます。</p><p><a href="https://www.promptingguide.ai/techniques/rag">検索拡張生成 (RAG)</a>は、LLM に外部知識を提供してユーザークエリへの応答を生成する AI テクニック/パターンです。これにより、LLM 応答を特定のコンテキストに合わせて調整できるようになり、応答がより具体的になります。</p><p><a href="https://docs.mistral.ai/">Mistral は</a>、オープンソースと最適化されたエンタープライズ グレードの LLM モデルの両方を提供します。このチュートリアルでは、ラップトップで実行されるオープン ソース モデル<a href="https://docs.mistral.ai/models/#mistral-7b">mistral-7b</a>を使用します。ラップトップでモデルを実行したくない場合は、代わりにクラウド バージョンを使用することもできます。その場合、適切な API キーとパッケージを使用するようにこのブログのコードを変更する必要があります。</p><p><a href="https://ollama.com/">Ollama は、</a>ラップトップ上で LLM をローカルに実行するのに役立ちます。Ollama を使用して、オープンソースの Mistral-7b モデルをローカルで実行します。</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.13/semantic-search.html">埋め込みは</a>、テキスト/メディアの意味を数値的に表現したものです。それらは高次元情報の低次元表現です。</p><h3>LlamaIndex、Elasticsearch、Mistral を使用した RAG アプリケーションの構築: シナリオの概要</h3><p><strong>シナリオ：</strong></p><p>架空の住宅保険会社のエージェントと顧客間のコールセンター会話のサンプル データセット (JSON ファイル) があります。次のような質問に答えられるシンプルなRAGアプリケーションを構築します。</p><p><code>Give me summary of water related issues.</code></p><h3>高レベルフロー</h3><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" alt="RAGフロー" /><p>Ollama を使用して Mistral LLM をローカルで実行しています。</p><p>次に、JSON ファイルからの<em>会話を</em><code>Documents</code>として<a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a> (Elasticsearch を基盤とする VectorStore) に読み込みます。ドキュメントをロードする際に、ローカルで実行されている Mistral モデルを使用して埋め込みを作成します。これらの埋め込みは<em>会話</em>とともに LlamaIndex Elasticsearch ベクター ストア ( <a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a> ) に保存されます。</p><p>LlamaIndex IngestionPipeline を設定し、使用するローカル LLM (この場合は Ollama 経由で実行される Mistral) に提供します。</p><p>「水に関する問題の概要を教えてください。」のような質問をすると、Elasticsearch はセマンティック検索を実行し、水問題に関連する<em>会話</em>を返します。これらの<em>会話は</em>元の質問とともに、ローカルで実行されている LLM に送信され、回答が生成されます。</p><h3>RAGアプリケーションの構築手順</h3><h4>Mistralをローカルで実行する</h4><p><a href="https://ollama.com/">Ollama</a>をダウンロードしてインストールします。Ollamaをインストールした後、このコマンドを実行して<a href="https://ollama.com/library/mistral">mistralを</a>ダウンロードして実行します。</p>ollama run mistral
<p>モデルを初めてローカルにダウンロードして実行するには数分かかる場合があります。以下のような「雲についての詩を書いてください」という質問をして、ミストラルが実行されているかどうかを確認し、その詩が気に入ったものかどうかを確認します。後でコードを通じてミストラル モデルとやり取りする必要があるため、ollama を実行したままにしておきます。</p><h4>Elasticsearchをインストール</h4><p>クラウド デプロイメントを作成する (<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">手順はこちら</a>) か、Docker で実行する (<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/docker">手順はこちら</a>) ことで、Elasticsearch を起動して実行します。<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/docker#self-hosted-production-deployments">ここから</a>開始して、Elasticsearch の実稼働グレードのセルフホスト型デプロイメントを作成することもできます。</p><p>クラウド デプロイメントを使用している場合は、手順に記載されているように、デプロイメント用の API キーとクラウド ID を取得します。後で使用します。</p><h4>RAGアプリケーション</h4><p>参考までに、コード全体はこの<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany">Github リポジトリ</a>にあります。以下のコードを実行するため、リポジトリのクローン作成はオプションです。</p><p>お気に入りの IDE で、以下の 3 つのファイルを使用して新しい Python アプリケーションを作成します。</p><ul><li><p><code>index.py</code> データのインデックス作成に関連するコードが配置される場所。</p></li><li><p><code>query.py</code> クエリと LLM の相互作用に関連するコードが配置される場所。</p></li><li><p><code>.env</code> API キーなどの構成プロパティが配置される場所。</p></li></ul><p>いくつかのパッケージをインストールする必要があります。まず、アプリケーションのルート フォルダーに新しい Python<a href="https://docs.python.org/3/library/venv.html">仮想環境</a>を作成します。</p>python3 -m venv .venv
<p>仮想環境をアクティブ化し、以下の必要なパッケージをインストールします。</p>source .venv/bin/activate
pip install llama-index 
pip install llama-index-embeddings-ollama
pip install llama-index-llms-ollama
pip install llama-index-vector-stores-elasticsearch
pip install sentence-transformers
pip install python-dotenv
<h4>データのインデックス作成</h4><p>架空の住宅保険会社の顧客とコール センター エージェント間の<em> 会話</em> が含まれる<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/blob/main/conversations.json"> conversations.json ファイルをダウンロードします。</a>アプリケーションのルートディレクトリに、2つのPythonファイルと.envファイルと一緒にファイルを配置します。先ほど作成したファイル。以下はファイルの内容の例です。</p>{
    "conversation_id": 103,
    "customer_name": "Sophia Jones",
    "agent_name": "Emily Wilson",
    "policy_number": "JKL0123",
    "conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
    "summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
<p><code>index.py</code>に、json ファイルを読み取ってドキュメントのリストを作成する<code>get_documents_from_file</code>という関数を定義します。<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/documents_and_nodes/">ドキュメント</a>オブジェクトは、LlamaIndex が扱う情報の基本単位です。</p># index.py
import json, os
from llama_index.core import Document, Settings
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.ingestion import IngestionPipeline
from llama_index.embeddings.ollama import OllamaEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore
from dotenv import load_dotenv

def get_documents_from_file(file):
   """Reads a json file and returns list of Documents"""

   with open(file=file, mode='rt') as f:
       conversations_dict = json.loads(f.read())
      
   # Build Document objects using fields of interest.
   documents = [Document(text=item['conversation'],
                         metadata={"conversation_id": item['conversation_id']})
                for
                item in conversations_dict]
   return documents
<p>IngestionPipelineを作成する</p><p>まず、 <code>Install Elasticsearch</code>セクションで取得した Elasticsearch CloudID と API キーを<code>.env</code>ファイルに追加します。<code>.env</code>ファイルは以下のようになります (実際の値を使用)。</p>ELASTIC_CLOUD_ID=&lt;REPLACE WITH YOUR CLOUD ID&gt;
ELASTIC_API_KEY=&lt;REPLACE WITH YOUR API_KEY&gt;
<p>LlamaIndex <a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/ingestion_pipeline/">IngestionPipeline を</a>使用すると、複数のコンポーネントを使用してパイプラインを構成できます。以下のコードを<code>index.py</code>ファイルに追加します。</p># index.py

# Load .env file contents into env
# ELASTIC_CLOUD_ID and ELASTIC_API_KEY are expected to be in the .env file.
load_dotenv('.env')

# ElasticsearchStore is a VectorStore that
# takes care of ES Index and Data management.
es_vector_store = ElasticsearchStore(index_name="calls",
                                     vector_field='conversation_vector',
                                     text_field='conversation',
                                     es_cloud_id=os.getenv("ELASTIC_CLOUD_ID"),
                                     es_api_key=os.getenv("ELASTIC_API_KEY"))


def main():
    # Embedding Model to do local embedding using Ollama.
    ollama_embedding = OllamaEmbedding("mistral")

    # LlamaIndex Pipeline configured to take care of chunking, embedding
    # and storing the embeddings in the vector store.
    pipeline = IngestionPipeline(
        transformations=[
            SentenceSplitter(chunk_size=350, chunk_overlap=50),
            ollama_embedding,
        ],
        vector_store=es_vector_store
    )

    # Load data from a json file into a list of LlamaIndex Documents
    documents = get_documents_from_file(file="conversations.json")

    pipeline.run(documents=documents)
    print(".....Done running pipeline.....\n")


if __name__ == "__main__":
    main()

<p>前述のように、LlamaIndex IngestPipeline は複数のコンポーネントで構成できます。パイプラインの行<code>pipeline = IngestionPipeline(...</code>に 3 つのコンポーネントを追加しています。</p><ul><li><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/?h=sentencesp#sentencesplitter">SentenceSplitter</a> : <code>get_documents_from_file()</code>の定義からわかるように、各ドキュメントには、json ファイルにある会話を保持するテキスト フィールドがあります。このテキスト フィールドは長いテキストです。セマンティック検索がうまく機能するには、小さなテキストのチャンクに分割する必要があります。<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/?h=sentencesp#sentencesplitter">SentenceSplitter</a>クラスがこれを実行します。これらのチャンクは、LlamaIndex 用語ではノードと呼ばれます。ノードには、それが属するドキュメントを指すメタデータが存在します。あるいは、この<a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">ブログ</a>で示されているように、Elasticsearch Ingestpipeline を使用してチャンク化することもできます。</p></li><li><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/models/embeddings/">OllamaEmbedding</a> : 埋め込みモデルは、テキストの一部を数値 (ベクトルとも呼ばれます) に変換します。数値表現を使用すると、単なるテキスト検索ではなく、単語の意味と一致する検索結果を表示する<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">セマンティック検索</a>を実行できます。IngestionPipeline に<code>OllamaEmbedding("mistral")</code>を提供します。SentenceSplitter を使用して分割したチャンクは、Ollama を介してローカル マシンで実行されている Mistral モデルに送信され、Mistral によってチャンクの埋め込みが作成されます。</p></li><li><p><a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a> : LlamaIndex ElasticsearchStore ベクター ストアは、作成される埋め込みを Elasticsearch インデックスにバックアップします。ElasticsearchStore は、指定された Elasticsearch インデックスの内容の作成と入力を担当します。ElasticsearchStore ( <code>es_vector_store</code>で参照) を作成する際に、作成する Elasticsearch インデックスの名前 (この場合は<code>calls</code> )、埋め込みを保存するインデックスのフィールド (この場合は<code>conversation_vector</code> )、およびテキストを保存するフィールド (この場合は<code>conversation</code> ) を指定します。要約すると、設定に基づいて、 <code>ElasticsearchStore</code> Elasticsearch に新しいインデックスを作成し、 <code>conversation_vector</code>と<code>conversation</code>フィールドとして（他の自動作成されたフィールドとともに）使用します。</p></li></ul><p>これらすべてを結び付けて、 <code>pipeline.run(documents=documents)</code>を呼び出してパイプラインを実行します。</p><p>index.py スクリプトを実行して、取り込みパイプラインを実行します。</p>python index.py
<p>パイプラインの実行が完了すると、Elasticsearch に<code>calls</code>という新しいインデックスが表示されます。開発コンソールを使用して単純な elasticsearch クエリを実行すると、埋め込みとともに読み込まれたデータが表示されるはずです。</p>GET calls/_search?size=1
<p>これまでに行ったことをまとめると、JSON ファイルからドキュメントを作成し、それをチャンクに分割し、それらのチャンクの埋め込みを作成し、埋め込み (およびテキスト会話) をベクター ストア (ElasticsearchStore) に保存しました。</p><h4>クエリ</h4><p>llamaIndex <a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_guide/">VectorStoreIndex を</a>使用すると、関連するドキュメントを取得したり、データをクエリしたりできます。デフォルトでは、 VectorStoreIndex は埋め込みを<a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_guide/">SimpleVectorStore</a>内のメモリ内に格納します。ただし、埋め込みを永続化するために、代わりに外部のベクトル ストア ( <a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a>など) を使用することもできます。</p><p><code>query.py</code>を開いて以下のコードを貼り付けます</p># query.py
from llama_index.core import VectorStoreIndex, QueryBundle, Response, Settings
from llama_index.embeddings.ollama import OllamaEmbedding
from llama_index.llms.ollama import Ollama
from index import es_vector_store

# Local LLM to send user query to
local_llm = Ollama(model="mistral")
Settings.embed_model= OllamaEmbedding("mistral")

index = VectorStoreIndex.from_vector_store(es_vector_store)
query_engine = index.as_query_engine(local_llm, similarity_top_k=10)

query="Give me summary of water related issues"
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>Ollama 上で実行されている Mistral モデルを指すようにローカル LLM ( <code>local_llm</code> ) を定義します。次に、先ほど作成した ElasticssearchStore ベクトル ストアから VectorStoreIndex ( <code>index</code> ) を作成し、インデックスからクエリ エンジンを取得します。クエリ エンジンを作成するときに、応答に使用するローカル LLM を参照し、ベクター ストアから取得して LLM に送信して応答を取得するドキュメントの数を構成する ( <code>similarity_top_k=10</code> ) も提供します。</p><p>RAG フローを実行するには、 <code>query.py</code>スクリプトを実行します。</p>python query.py
<p>クエリ<code>Give me summary of water related issues</code>を送信します ( <code>query</code>は自由にカスタマイズできます)。関連するドキュメントとともに提供される LLM からの応答は次のようになります。</p>提供されたコンテキストでは、水に関連する損害の補償について顧客が問い合わせた例がいくつかあります。2件のケースでは洪水により地下室が損傷し、別のケースでは屋根の漏水が問題となった。代理店は、両方のタイプの水害がそれぞれの保険でカバーされていることを確認しました。したがって、浸水や屋根の漏水などの水関連の問題は、通常、住宅保険でカバーされます。<h4>注意点:</h4><p>このブログ投稿は、Elasticsearch を使用した RAG テクニックの初心者向け紹介であるため、この開始点を本番環境に移行できるようにする機能の構成については省略しています。実稼働ユースケース向けに構築する場合は、<a href="https://www.elastic.co/search-labs/blog/dls-internal-knowledge-search">ドキュメント レベルのセキュリティ</a>でデータを保護したり、Elasticsearch<a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">取り込みパイプ</a>ラインの一部としてデータをチャンク化したり、GenAI/チャット/Q&amp;A ユースケースで使用されているのと同じデータで他の<a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-overview.html">ML ジョブ</a>を実行したりするなど、より高度な側面を考慮する必要があります。</p><p><a href="https://www.elastic.co/guide/en/enterprise-search/current/connectors.html">Elastic Connectors</a>を使用して、さまざまな外部ソース (Azure Blob Storage、Dropbox、Gmail など) からデータを取得し、埋め込みを作成することも検討してください。</p><p>Elastic は、上記すべてとそれ以上のことを可能にし、GenAI ユースケースなどに対応する包括的なエンタープライズ グレードのソリューションを提供します。</p><h4>What’s next?</h4><ul><li><p>お気づきかもしれませんが、応答を作成するために、ユーザーの質問とともに 10 件の関連する会話が LLM に送信されています。これらの会話には、名前、生年月日、住所などの PII (個人を特定できる情報) が含まれる場合があります。私たちの場合、LLM はローカルなので、データ漏洩は問題になりません。ただし、クラウドで実行される LLM (OpenAI など) を使用する場合は、PII 情報を含むテキストを送信することは望ましくありません。次回のブログでは、RAG フローで外部 LLM に送信する前に PII 情報をマスキングする方法について説明します。</p></li><li><p>この投稿ではローカル LLM を使用しました。RAG での PII データのマスキングに関する次の投稿では、ローカル LLM からパブリック LLM に簡単に切り替える方法について説明します。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" length="0" type="image/png"/>
    <pubDate>Fri, 12 Apr 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>