<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[AI - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[AI - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/jp/search-labs/blog/category/ai</link>
    </image>
    <link>https://www.elastic.co/jp/search-labs/blog/category/ai</link>
    <atom:link href="https://www.elastic.co/jp/search-labs/rss/category/ai.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[jp]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 09:35:55 GMT</lastBuildDate>
  <item>
    <title><![CDATA[描くのではなく、説明する：MCPとES|QLによるAIネイティブのKibanaダッシュボード]]></title>
    <description><![CDATA[プロンプトからダッシュボードへ。example-mcp-dashbuilderを使って、自然言語でKibanaダッシュボードを構築する方法を学びましょう。ES|QLクエリを書き、インタラクティブなグラフを作成し、全面的に機能するダッシュボードをKibanaに直接エクスポートするオープンソースのMCPアプリケーションです。]]></description>
    <content:encoded><![CDATA[<p>example-mcp-dashbuilderはオープンソースのMCPアプリケーションで、平易な英語のプロンプトをライブでインタラクティブなKibanaのダッシュボードに変換します。これらはすべて、エディタのチャット画面内で行われます。ダッシュボードの要望を記述すると、AIがインデックス構造を検出し、各可視化に適切なES|QLアグリゲーションを記述し、作業中にインラインでプレビューをレンダリングします。完了後、1つのコマンドで完全に機能するKibanaのダッシュボードがエクスポートされます。実際のLens可視化、正確なグリッドレイアウト、カスタムカラーもそのまま保持されます。現在、6種類のグラフがサポートされており、Kibana Lensの全機能がロードマップに設定されています。</p><h2>Kibanaダッシュボードビルダーとは？</h2><p>必要なダッシュボードをわかりやすい日本語で説明すると、インタラクティブなグラフ、ドラッグアンドドロップのレイアウト、Kibanaへのワンクリックエクスポートが表示・実行されるとしたらどうでしょうか。</p><p>それがまさに <a href="https://github.com/elastic/example-mcp-dashbuilder.git"><strong>example-mcp-dashbuilder</strong></a> の役割です。これはオープンソースのモデルコンテキストプロトコル（MCP）アプリケーションで、AIアシスタントをElasticsearchに接続し、会話を通じて総合的なKibanaのダッシュボードを作成できます。メニューをクリックしたり、手動で可視化設定を書く必要はありません。必要なものを説明するだけで、AIがデータを調査し、Elasticsearch Query Language（ES|QL）でクエリを書き、グラフを作成し、ライブかつインタラクティブなダッシュボードを提供します。これらはすべて、エディタのチャット画面内で行われます。</p><h2><strong>プロンプトからダッシュボードまでを数秒で</strong></h2><p>実際の様子を以下で紹介します。次のように入力します：</p><p>「「logstash-*」から、リクエストの合計数、時間の経過に伴う転送バイト数、上位の地理的ソース、対応コードの内訳を含むWebトラフィックダッシュボードを作成してください。」</p><p>AIは次のように動作します：</p><ol><li><p><strong>データを発見：</strong>インデックスを一覧表示し、フィールドマッピングを検査します。</p></li><li><p><strong>ES|QLのクエリを作成：</strong>スキーマに合わせ、適切なアグリゲーションを使用します。</p></li><li><p><strong>可視化を作成：</strong>棒グラフ、折れ線グラフ、スパークライン付きメトリクス、ヒートマップ、円グラフを作成できます。</p></li><li><p><strong>すべてを整理整頓：</strong>折りたたみ可能なセクション、わかりやすいタイトル、適切なレイアウト。</p></li><li><p><strong>インタラクティブなプレビューを表示：</strong>チャット内でツールチップ、時間選択機能、ドラッグ＆ドロップ機能を利用できます。</p></li></ol><p>各グラフは作成されると同時にインラインで表示されるため、リアルタイムで進捗状況を確認できます。次に<code>view_dashboard</code>は、Kibanaの48列グリッドにすべてのパネルが配置された完全なダッシュボードを表示します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt75af5d9042d141b5/6a17e99dbe608675a4004792/dcbf47c4f17bf1a184fb0167408ebeb861ef6c9d-1404x1568.png" alt="example-mcp-dashbuilderインターフェースに2つのチャートが表示されます。1つ目は、「上位の地理的ソース」と題された縦棒グラフで、国コード別のリクエスト数を示しています。2つ目は、「HTTP応答コードの内訳」というタイトルの円グラフで、200、404、503の応答のセグメントを示しています。" /><p><em>単一のグラフを本文中にプレビュー表示します。</em></p><h2><strong>ES|QLで構築</strong></h2><p>すべてのデータ検索には、Elasticsearchのパイプ型クエリ言語である<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL</a>を使用しています。AIは単に未加工のクエリをそのまま渡すだけでなく、ES|QL構文に関して標準搭載された知識とお客様のデータ構造に関する情報を使用して、各可視化タイプに対して正確かつ効率的なクエリを作成します。</p><p>サーバーには包括的なES|QLリファレンスがMCPリソースとして含まれています。クエリを書く前に、AIはこのリファレンスを読み取り、利用可能なコマンド、関数、およびパターンを理解します。データ可視化のベストプラクティス・ガイド（リソースとしても機能）と組み合わせることで、AIはクエリの<em>方法</em>だけでなく、<em>何が</em>可視化を優れたものにするのかも知ることができます：</p><ul><li><p>時系列には <code>BUCKET(@timestamp, 1 day)</code> を使い、常に時間フィールドで <code>SORT</code> します。</p></li><li><p>円グラフは<code>| SORT value DESC | LIMIT 6</code>で6切れに制限します。</p></li><li><p>カテゴリー比較のための棒グラフ、トレンドを示す折れ線グラフ、主要業績評価指標（KPI）のためのメトリクスを選択します。</p></li></ul><h2><strong>オープンエンド分析によるAI主導のデータ探索</strong></h2><p>頭の中で既に設計したダッシュボードを、実際に作成することはまた別の作業です。「このインデックスの何が興味深いのか？」と尋ねること、そして有用な答えを得ることはより難しく、AIが単に描くのではなく、<em>探索</em>する方法を知る必要があります。</p><p>example-mcp-dashbuilder は、構造化された探索フローを定義する <code>analysis://guidelines</code> リソースを提供します。そのリソースとは、データのプロファイリング、ターゲットを絞ったアグリゲーションの実行、調査に値するパターンの抽出、最も興味深い発見のためのチャートの作成、ユーザーが次に望むかもしれないドリルダウンクエリの提案などです。トリガーフレーズ（例えば「ログを分析」や「このインデックス内のパターンを発見」）は、AIが何かを行う前にプレイブックを読み込むようにするため、オープンエンドなプロンプトはランダムなチャートの集合ではなく、一貫性のある調査を生成します。</p><p>結果：AIに馴染みのないインデックスを渡すと、開始点が返されます。開始点は、ダッシュボードと「以下に気づきました。これらの中に詳しく調べたいものはありますか？」というプロンプトと短いリストです。</p><h2><strong>Kibanaダッシュボードのエクスポートとインポート：完全な往復処理</strong></h2><p>エクスポート/インポートの往復処理は、既に Kibana を使用しているチームにとって example-mcp-dashbuilder が真に役立つ部分です。example-mcp-dashbuilder は独自の機能を持ち、エディタ内に存在する対話型のダッシュボード画面ですが、作業内容をエディタ内に閉じ込めることはありません。ここで構築されたダッシュボードは、必要に応じてKibanaに移動できます。既存のKibanaダッシュボードは、AI支援による編集のために逆方向に移動させることが可能です。</p><h3><strong>Kibanaにエクスポート</strong></h3><p>ダッシュボードにご満足いただけましたら、次のコマンド1つでエクスポートできます：</p><p>「このダッシュボードをKibanaにエクスポートしてください」</p><p>すべてのパネルは実際のKibana Lensの可視化に変換されます。変換後も以下は保持されます：</p><ul><li><p><strong>ES|QLクエリ：</strong>LensにおけるES|QLのデータソースとして直接転送されます。</p></li><li><p><strong>グリッド位置：</strong>Kibanaと同じ48列システムを使用しているため、レイアウトはKibanaと全く同じに見えます。</p></li><li><p><strong>カスタムカラー：</strong>シリーズパレット、メトリックの背景、ヒートマップのカラーランプ。</p></li></ul><p>その結果として、全面的に機能するKibanaのダッシュボードができます。スクリーンショットでも埋め込みでもありません。共有してKibanaで編集を続けることができるダッシュボードです。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1921c74c2833cabe/6a17e99f6864a4a712b687da/5e27777bc0a82cafb373943f65298bdb21d66176-1999x902.png" alt="2つのダッシュボードが並んで表示されています。「Webトラフィックの編集 — Logstash（Dashbuilder）」と題された左側のダッシュボードには、最近のトラフィック指標が、トラフィック量パネル、地理情報に関する棒グラフ、レスポンスコードの円グラフとともに表示されています。右側のダッシュボードには、より高い合計値を示す同様のレイアウトがあり、トラフィック量パネル、地理情報に関する棒グラフ、レスポンスコードの円グラフが含まれています。" /><p><em>Kibanaダッシュボードとカーソルチャットのダッシュボードを並べて表示します。</em></p><h3><strong>Kibanaからインポート</strong></h3><p>往復処理は逆方向でも機能します。</p><p>「ID abc-123でKibanaのダッシュボードをインポート」</p><p>これは既存のKibanaのダッシュボードを取得し、そのLensの可視化を編集可能なチャート構成に変換し、グリッドのレイアウトとセクションを保持して、すべてを example-mcp-dashbuilder に読み込みます。そこから自然言語で修正し、再エクスポートできます。</p><p>このように、AIは既存のKibanaワークフローにおける共同作業者となり、それを置き換えるものではありません。</p><h2><strong>カスタムテーマと色</strong></h2><p>ブランド化されたダッシュボードをご希望ですか？お問い合わせください：</p><p>「カスタムカラーを使用したピンクを基調としたダッシュボードを作成」</p><p>すべての可視化タイプはカスタムカラー設定をサポートしています：</p><ul><li><p><strong>チャート：</strong><code>palette</code> はシリーズとスライスに対して16進数の色の配列を指定できます。</p></li><li><p><strong>指標：</strong><code>color</code> が背景色を設定します。</p></li><li><p><strong>ヒートマップ：</strong> <code>colorRamp</code> は、低い値から高い値への勾配を定義します。</p></li></ul><p>AIはテーマのリクエストを自然に受け取ります。「海のテーマ」と伝えると、青やティールの色合いが選択されます。「自社のブランドカラーと一致させてください」と伝えて16進数の値を指定すると、エクスポート時にKibanaに引き継がれます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2bc7cddbdef81354/6a17e9a1ec0f89ee155a665e/4aceba013ac9cbb4a541109efd6acddf8a6ec47d-1562x1568.png" alt="カスタムピンクカラーの、テーマを使用したeコマースダッシュボード。レイアウトでは、上部に収益と注文のKPIが表示され、その下に折りたたみ式のトレンドセクション、さらにその下にカテゴリ別の2つのグラフ（カテゴリ別の収益を示す棒グラフと、カテゴリ別の注文数を示す円グラフ）が表示されています。" /><p><em>カスタムカラーの、テーマを使用したダッシュボード。</em></p><p><strong>example-mcp-dashbuilder の仕組み：MCPアーキテクチャ</strong></p><p>example-mcp-dashbuilderは、AIアシスタントを外部ツールやデータに接続するためのオープン標準である <a href="https://modelcontextprotocol.io/">MCP</a>に基づいて構築されています。アーキテクチャの概要は以下の通りです：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6c0cd879646e9947/6a17e9a36864a4c408b687df/cbfeabe151ec1ee2b0655f4d17468c9bb358df7e-1024x559.png" alt="MCPサーバーに接続されたMCPホストを示すアーキテクチャ図で、ツール、リソース、手順の説明が含まれています。その下にあるMCPアプリボックスには、Elastic ChartsとKibanaのグリッドレイアウトが含まれています。ElasticsearchとKibanaが下部に表示され、それらをMCPアプリに接続する矢印があります。" /><p><strong>MCPサーバー</strong>は、AIが直接呼び出せる25のツールを公開しています。これには、ES|QLクエリの実行からダッシュボードのエクスポートまで、網羅的な内容が含まれています。さらに、インラインプレビューがデータの取得、レイアウト変更の永続化、時間フィールドの検出に使用する内部専用の「アプリのみ」ツールもいくつかあります。また、3つのリソースも提供しており、データ可視化のベストプラクティスガイド、ES|QLリファレンス、そしてオープンエンドのプロンプト（「ログを分析」、「このインデックスで何が興味深いのか」）に対応する深度分析プレイブックがあります。そして、stdioまたはHTTPのいずれかで実行されます。HTTPトランスポートはストリーム可能な対応とセッション管理をサポートしているため、複数のクライアントが1つのサーバーに接続できます。</p><p><strong>MCPアプリ</strong>は、インタラクティブなプレビューを表示します。React、<a href="https://elastic.github.io/elastic-charts">Elastic Charts</a>、<a href="https://eui.elastic.co/">Elastic UI</a>を組み合わせて構築されており、1つの独立したHTMLファイルにまとめられています。AIが <code>view_dashboard</code> を呼び出したり、チャートを作成したりすると、ホストはこのHTMLをサンドボックス化されたiframe内にレンダリングします。アプリ全体は<a href="https://modelcontextprotocol.io/extensions/apps/overview">MCP Appsのプロトコル</a>を通じてサーバーと通信し、postMessage上の <code>callServerTool()</code> を使ってデータの取得、レイアウトの保存、時間フィールドの検出を行います。localhostサーバーもなく、ポート設定も、外部ネットワーク依存もありません。</p><p>これは、あらゆるMCP互換のクライアント（Cursor、Claude Desktop、Claude.ai、VS CodeとCopilotの併用など）と動作することを意味します。</p><h2><strong>example-mcp-dashbuilder はどのようなチャートタイプをサポートしていますか？</strong></h2><p>執筆時点では、最も一般的なダッシュボードのシナリオをカバーする以下の6種類のチャートタイプがサポートされています。</p><p>タイプ</p><p>最適な用途</p><p>例</p><p>棒グラフ</p><p>カテゴリ比較</p><p>地理的ソース別のリクエスト</p><p>折れ線グラフ</p><p>一定期間におけるトレンドの変化</p><p>1時間あたりの転送バイト数</p><p>エリア</p><p>時間経過に伴うボリューム</p><p>時間経過に伴うリクエスト量</p><p>円グラフ</p><p>全体に占める割合（最大6切れ）</p><p>対応コードの分布</p><p>メトリック</p><p>スパークライン付きの単一KPI</p><p>時間別トレンド付きのリクエスト総数</p><p>ヒートマップ</p><p>二次元領域全体でのパターン</p><p>曜日・時間別リクエスト</p><p>ダッシュボードは、整理のための折りたたみ可能なセクション、自動時間フィールド検出を備えたタイムピッカー、および複数のダッシュボードを保存して切り替える機能をサポートしています。並行チャットセッションは、すべてのツールコールで<code>dashboardId</code>がスレッド化されているため、互いに分離された状態を維持します。</p><h2><strong>example-mcp-dashbuilder のインストールと実行方法</strong></h2><p>example-mcp-dashbuilder はオープンソースであり、すぐに利用可能です。Node.js 22+、Elasticsearchインスタンス（ローカルまたはElastic Cloud）、およびMCP互換のクライアントが必要です。</p><p><strong>Claude Desktop：</strong> <a href="https://github.com/elastic/example-mcp-dashbuilder/releases">GitHub Releases</a>から最新版<code>.mcpb</code>をダウンロードし、ダブルクリックします。Claude Desktopから、Elasticsearchの認証情報を入力するよう促されます。</p><p><strong>Cursor / Claude Code / VS Code Copilot：</strong>MCP設定をリリースされた tarball に指定します。クローンや <code>npm install</code> は不要です。</p>{
  "mcpServers": {
    "example-mcp-dashbuilder": {
      "type": "stdio",
      "command": "npx",
      "args": ["https://github.com/elastic/example-mcp-dashbuilder/releases/latest/download/example-mcp-dashbuilder.tgz"]
    }
  }
}<p>環境変数として <code>ES_NODE, ES_API_KEY</code>（または <code>ES_USERNAME / ES_PASSWORD</code>）と <code>KIBANA_URL</code> を設定します。ソースから作業したい場合は、リポジトリをクローンし、<code>npm run setup</code> を実行して、ローカルのElasticsearchとElastic Cloud（Cloud ID + APIキー）の両方を処理するインタラクティブウィザードを使用します。</p><p>次のように、構築を開始できます：</p><p>「ログのインデックスを探索し、可能な限り洞察に富むダッシュボードを作成してください」</p><p>AIがその後を引き継ぎます。😉</p><h2><strong>ロードマップ：example-mcp-dashbuilder の今後</strong></h2><p>これは初期リリースであり、現在も鋭意開発を進めています。以下の分野などに注力しています。</p><ul><li><p><strong>より多くのチャートタイプ：</strong>ゲージグラフ、ドーナツグラフ、ツリーマップ、データテーブル、タグクラウドなど、Lensの全機能に対応。</p></li><li><p><strong>ダッシュボードをGitにプッシュ：</strong>ダッシュボードの設定をリポジトリに書き込み、バージョン管理やコードレビューのワークフローを行います。</p></li><li><p><strong>優れたエラーUX：</strong>ES|QLクエリが失敗した場合、一般的な修正案を含むより詳細なフィードバックを提供します。</p></li><li><p><strong>より高度な分析フロー：</strong>詳細分析のプレイブックを拡張し、より多くのデータ形式（ログ、メトリクス、トレース）に対応します。</p></li></ul><p>お客様が構築されたものを、ぜひご紹介ください。お試しの後で問題があればご報告いただき、お客様のチームにとって最も役立つ可視化やワークフローはどのようなものかお知らせください。</p><p><a href="https://github.com/elastic/example-mcp-dashbuilder">GitHub: elastic/example-mcp-dashbuilder</a></p><h3>謝辞</h3><p><a href="mailto:walter.rafelsberger@elastic.co">ウォルター・ラフェルズバーガー</a>と<a href="mailto:tim.schnell@elastic.co">ティム・シュネル</a>の実装への貢献に感謝します。</p><h3>FAQ</h3><p><strong>example-mcp-dashbuilder とは？</strong>example-mcp-dashbuilder は、AIアシスタントをElasticsearchに接続するオープンソースのMCP（Model Context Protocol）アプリケーションです。Kibanaのダッシュボードを平易な日本語で説明し、ES|QLクエリを自動生成し、可視化を作成し、エディタのチャット画面内にライブかつインタラクティブなダッシュボードを表示できます。</p><p><strong>example-mcp-dashbuilder はデータ取得にどのようなクエリ言語を使っていますか？</strong>すべてのデータ取得には、Elasticsearchのパイプクエリ言語であるES|QLを使用しています。MCPサーバーには、クエリを書く前にAIが読み取るES|QLリファレンスが標準搭載されているため、各可視化タイプの正しい構文と効率的なアグリゲーションが確保されます。</p><p><strong>example-mcp-dashbuilder で作成したダッシュボードをKibanaにエクスポートできますか？</strong>はい。「このダッシュボードをKibanaにエクスポート」を実行すると、すべてのパネルが実際の Kibana Lens の可視化に変換され、ES|QLクエリ、48列のグリッドレイアウト、カスタムカラー、シリーズパレットが保持されます。結果は、スクリーンショットや埋め込みではなく、全面的に機能するKibanaのダッシュボードです。</p><p><strong>既存のKibanaのダッシュボードを example-mcp-dashbuilder にインポートして、AI支援型編集はできますか？</strong>はい。KibanaのダッシュボードIDを指定すると、既存のダッシュボードが取得され、Lensの可視化が編集可能なグラフ構成に変換され、example-mcp-dashbuilder に読み込まれます。その後、自然言語を使用してダッシュボードを変更し、Kibanaに再エクスポートできます。</p><p><strong>example-mcp-dashbuilder と互換性のあるMCPクライアントはどれですか？</strong>example-mcp-dashbuilder は、Cursor、Claude Desktop、Claude.ai、VS Code with Copilotなど、あらゆるMCP互換クライアントで動作します。stdioとHTTPトランスポートの両方をサポートしており、localhostサーバーやポートの設定は不要です。</p><p><strong>example-mcp-dashbuilder はどのチャートタイプをサポートしていますか？</strong>現在のリリースでは、棒グラフ、折れ線グラフ、面グラフ、円グラフ、メトリクス（スパークライン付き）、ヒートマップの6種類のグラフがサポートされています。Kibana Lensの全機能に合わせて、ゲージグラフ、ドーナツ、ツリーマップ、データテーブル、タグクラウドなどを追加する予定です。</p><p><strong>example-mcp-dashbuilder を実行するには何が必要ですか？</strong>Node.js 22以上、Elasticsearchインスタンス（ローカルまたはElastic Cloud）、およびMCP互換クライアントが必要です。環境変数 ES_NODE、ES_API_KEY（またはES_USERNAME/ES_PASSWORD）、KIBANA_URLを設定します。Claude Desktopの場合は、GitHub Releasesから.mcpbファイルをダウンロードし、ダブルクリックしてインストールします。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/kibana-dashboard-builder-mcp-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/kibana-dashboard-builder-mcp-esql</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Stratoula Kalafateli]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2a69a35d6d51ff47/6a17e9a5b1e11339cd79f2b3/0d38385fd64c1445b2e955ba20532570f7f38679-1280x720.png" length="0" type="image/png"/>
    <pubDate>Fri, 22 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearchによるエンティティ解決、パート4：究極のチャレンジ]]></title>
    <description><![CDATA[ショートカットを防ぐために設計された、非常に多様な「究極のチャレンジ」データセットにおけるエンティティ解決の課題の解決と評価。]]></description>
    <content:encoded><![CDATA[<p>これまでのインテリジェントなエンティティ解決は2つの方法で実装されてきました。いずれのアプローチも、エンティティの準備と抽出、そしてElasticsearchによる候補の取得という同じ方法で始まります。そこから、プロンプトベースのJSON生成または関数呼び出しのいずれかを通じて、大規模言語モデル（LLM）を使用して候補を評価し、モデルにその判断について透明性のある説明を提供することを要求します。</p><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-entity-resolution-llm-function-calling">前回の記事</a>で見たように、関数呼び出しによってもたらされる一貫性は、単に便利な最適化ではなく、不可欠なものです。構造的なエラーを評価ループから除去したところ、標準的なシナリオ（ティア4データセットなど）の結果が劇的に向上しました。</p><p>しかし、答えるべき明白な疑問はまだ残っています。</p><p><em>状況が本当に複雑になってきた場合でも、このアプローチは有効でしょうか？</em></p><p>現実世界におけるエンティティ解決が単純なケースで失敗することはめったにありませんが、名前が言語、文化、文字体系、時代、組織の境界を越える場合に失敗します。人が名前ではなく肩書きで言及されている場合、会社名が変更された場合、音訳が一貫していない場合、そして（スペルではなく）文脈だけが言及と現実世界の実体を結びつける唯一の要素である場合、この方法は失敗します。</p><p>そこで、このシリーズの最後の記事として、このシステムにいわば<strong>究極のチャレンジ</strong>を課すこととしました。</p><h2>なぜこれが究極の挑戦なのでしょうか？</h2><p>以前の評価では、ますます複雑になるデータセットを用いてシステムをテストしました。前回の記事で触れた第4段階に到達する頃には、すでにニックネーム、称号、多言語名、意味的な参照などが混在する状況になっていました。これらのテストにより、アーキテクチャ自体は健全であることが示されましたが、信頼性の問題、特に不正な形式のJSONが原因で、リコールが抑制されていることがわかりました。</p><p>関数呼び出しの仕組みが整ったことで、ようやく安定した基盤ができました。そのおかげで、さらに興味深い質問をする機会が得られました。</p><p><em>1つの統一されたパイプラインで </em><em><strong>多くの異なる種類の</strong></em><em>エンティティ解決問題を一度に処理することは可能でしょうか？</em></p><p>究極のチャレンジデータセットは、まさにその側面を徹底的に追求するために設計されました。</p><p>このデータセットは、（ニックネームや音訳といった）単一の困難に焦点を当てるのではなく、 <strong>50種類以上の異なる課題タイプ</strong>を組み合わせています。</p><ul><li><p>文化的な命名規則。</p></li><li><p>タイトルに基づく参照。</p></li><li><p>事業上の関係性と過去の社名変更。</p></li><li><p>多言語および異文字表記での言及。</p></li><li><p>上記のうち複数を組み合わせた複合的な課題。</p></li></ul><p>重要なのは、この試みが特定の狭い用途向けに最適化することではなく、ルールがエンティティごとに変化した場合でも<em>設計パターン</em>が通用するかどうかをテストすることです。</p><h2>データセットの概要</h2><p>究極のチャレンジデータセットは以下で構成されます。</p><ul><li><p>個人、組織、機関などの<strong>50のエンティティ</strong>。</p></li><li><p>構造と言語の複雑さが異なる<strong>約60本の記事</strong>。</p></li><li><p>大きく以下に分類される<strong>51種類の異なるチャレンジカテゴリー</strong>。</p><ul><li><p>文化的な命名規則。</p></li><li><p>肩書きと職務上の背景。</p></li><li><p>事業と組織間の関係。</p></li><li><p>多言語および音訳の課題。</p></li><li><p>複合シナリオとエッジケースのシナリオ。</p></li></ul></li></ul><p>本シリーズの前半で、生成AIを用いてデータセットを作成することは諸刃の剣であることを確認しました。生成AIがなければ十分な規模と多様性を備えたテストデータを収集することは極めて困難になりますが、このモデルは放置すると、物事をあまりにも単純化しすぎる傾向があります。</p><p>例えば、初期世代の検証段階で、モデルに「ロシアの大統領」といったフレーズがウラジーミル・プーチンの明示的な別名として含まれていることが判明しました。それは今日では妥当に思えるかもしれませんが、文脈解決能力をテストするという目的を損なうことになります。記事が1990年代のロシアについて論じている場合はどうなるでしょうか？システムは、ハードコードされたエイリアスに頼るのではなく、文脈から正しいエンティティを推論するべきです。</p><p>そのため、このデータセットは<strong>ショートカットが効かない</strong>ように意図的に設計されています。システムが意味を推測することが想定されている場合、別名は明示的にリスト化されません。記述的なフレーズはエンティティにあらかじめリンクされていません。正確な一致は、単なるローカルテキストだけでなく、記事レベルの文脈によって決まることが多いです。</p><p><strong>重要な注意点：</strong>本システムは多様なシナリオにおける機能を実証していますが、これはあくまで教育用プロトタイプです。実際の制裁対象組織の監視を扱う本番システムでは、追加の検証、コンプライアンスチェック、監査証跡、および機密性の高いユースケースに対する特別な処理が必要となります。</p><h2>これらのシナリオが難しい理由</h2><p>このシリーズの最初の投稿で、単純であいまいな例「新しいSwiftアップデートが登場しました！」を紹介しました。課題は、「Swift」という単語が、文脈によって複数の現実世界の実体として解釈される可能性があることです。この例はより広範な真実、つまり、自然言語は本質的に曖昧であるということを捉えています。</p><p>したがって、エンティティ解決は単なる文字列照合の問題ではありません。人間は日常的に、共通の知識、文化的規範、状況的文脈に頼って参照関係を解決していますが、私たちは自分がそうしていることにほとんど気づきません。</p><p>よくあるケースをいくつか考えてみましょう。</p><ul><li><p>「大統領」という称号は地政学的・時間的な文脈なしには意味がありません。</p></li><li><p>会社名は、記事がいつ書かれたかによって、親会社、子会社、または以前のブランドを指す場合があります。</p></li><li><p>人名は、言語や文化によって、異なる順序、書体、または音訳で表記されることがあります。</p></li><li><p>同じフレーズでも、文脈によって異なる対象を指す場合があり、システムは一致を受け入れるのと同じくらい確信を持って一致を<em>拒否</em>できなければなりません。</p></li></ul><p>これらすべてを適切に処理する単一のルールセットは存在しないため、このプロトタイプは懸念事項を非常に積極的に分離しています。</p><ul><li><p>Elasticsearchは候補の範囲を効率的かつ分かりやすく絞り込みます。</p></li><li><p>LLMは、判断が必要で、それ自体を説明しなければならない場合にのみ使用されます。</p></li><li><p>検索と推論は別個のステップのままです。</p></li></ul><p>課題の種類が多様化するにつれて、この区分けはさらに重要になります。</p><h2>システムが特別なケースなしに多様性を処理する仕組み</h2><p>この評価で最も興味深い結果の一つは、<em>変更しなかった</em>点にあります。</p><ul><li><p>日本語名に関する特別なロジックは追加して<strong>いません</strong>。</p></li><li><p>アラビア語の父称に関するカスタムルールは追加して<strong>いません</strong>。</p></li><li><p>ハードコーディングされたマッピングを過去の会社名に追加して<strong>いません</strong>。</p></li></ul><p>その代わりに、このシステムはシリーズ前半で紹介したものと同じ主要要素に依存していました。</p><ul><li><p>セマンティック検索のためにインデックス化されたコンテキスト強化エンティティ。</p></li><li><p>Elasticsearchでのハイブリッド検索（完全検索、エイリアス、セマンティック）。</p></li><li><p>少数の、明確に定義された一致候補セット。</p></li><li><p>関数呼び出しと最小スキーマによって制約されたLLM判断。</p></li></ul><p>これは、システムの柔軟性が、増え続けるルールのコレクションからではなく、<strong>表現とアーキテクチャ</strong>から生まれることを示唆しています。</p><p>システムが成功するのは、適切な候補が取得され、LLMが参照が特定のエンティティにマッピングされる（またはされない）理由を説明できる十分なコンテキストがある場合です。</p><h2>結果：パフォーマンスの概要</h2><p>究極のチャレンジデータセットにおいて、システムは以下のような全体的な結果を生み出しました。</p><ul><li><p><strong>精度：</strong>約91％</p></li><li><p><strong>再現率：</strong>約86％</p></li><li><p><strong>F1スコア：</strong>約89%</p></li><li><p><strong>LLM合格率：</strong>約72％</p></li></ul><h3>チャレンジの種類ごとのパフォーマンス</h3><p>チャレンジの種類ごとに結果を分解すると、強みと限界が明らかになります。</p><p><strong>最も優れたパフォーマンス（F1スコア100%）</strong>が見られた分野は以下のとおりです。</p><ul><li><p>文字体系間の照合（キリル文字、韓国語、中国語の企業名）。</p></li><li><p>ヘブライ語のシナリオ（父称、専門職称、宗教称号、音写）。</p></li><li><p>事業階層構造（航空宇宙、多角化製造業、多部門企業）。</p></li><li><p>職業上の肩書き（学術、軍事、政治、宗教）。</p></li><li><p>複数の文字体系を含む日本語シナリオの組み合わせ。</p></li></ul><p><strong>優れたパフォーマンス（F1スコア80～99％）</strong>には以下が含まれます。</p><ul><li><p>国際的な政治家（98％）。</p></li><li><p>歴史的な名称変更（90%）。</p></li><li><p>複雑なビジネス階層（89％）。</p></li><li><p>日本の企業名（93％）。</p></li><li><p>異言語間の音訳（86％）。</p></li><li><p>アラビア語の父称（86％）。</p></li></ul><p><strong>より困難な分野</strong>には以下が含まれます。</p><ul><li><p>高度な音訳（中国語、韓国語）：0% F1。</p></li><li><p>特定の日本語シナリオ（敬称、名前の順序、表記体系のバリエーション）：約67% F1。</p></li><li><p>一部のアラビア語のシナリオ（会社名、機関の参考文献）：約40％ F1。</p></li></ul><p>ここで重要なのは、<em>なぜ</em>システムがこれらのケースで機能不全に陥ったのかという点です。失敗の原因は、全体的なアプローチが破綻したことではなく、特定のコンポーネントの限界、特に特定の多言語シナリオにおけるセマンティック検索に使用される高密度ベクトルモデルの限界にありました。</p><p>検索と判断が明確に分離されているため、パフォーマンスを向上させるためにシステムを書き換える必要はありません。より高性能な多言語埋め込みモデルの採用、エンティティコンテキストの強化、または検索戦略の洗練により、コアアーキテクチャを変更することなく、これらのカテゴリー全体で結果が向上します。</p><p>アーキテクチャーの観点から見ると、それが真の成功指標です。</p><h2>この結果が設計について教えてくれること</h2><p>シリーズを振り返ると、いくつかのパターンが際立っています。</p><ul><li><p><strong>準備は巧みなマッチングよりも重要です。 </strong>エンティティに事前にコンテキストを付加することで、後々の曖昧さを劇的に減らすことができます。</p></li><li><p><strong>LLMは、レトリバーではなく、判断者として最も価値があります。</strong>したがって、検索を求めるよりも、<em>なぜ</em>一致が意味をなすかを説明するよう求めることの方がはるかに強力です。</p></li><li><p><strong>信頼性が精度を実現します。</strong>関数呼び出しは、JSONを整理しただけでなく、取得ステップにすでに潜在していた想起を解放しました。</p></li><li><p><strong>一般化は専門化に勝ります。</strong>厳選された少数の抽象化によって、独自のロジックを必要とせずに数十種類の課題に対応できました。</p></li></ul><p>これが、プロトタイプが意図的にElasticsearchネイティブであり、LLMの使用方法が意図的に保守的である理由です。目標は検索を置き換えることではなく、意味が重要な状況において、検索を説明可能なものにすることです。</p><h2>結びに</h2><p>究極のチャレンジとは、完璧な指標を追い求めることではなく、より根本的な問いに答えることでした。</p><p><em>透明性が高く、検索優先で、LLMを活用したアーキテクチャは、ルールやブラックボックスに陥ることなく、現実世界のエンティティの曖昧さを処理できるでしょうか？</em></p><p>その回答は、この教育用プロトタイプに関しては「はい」ですが、本番環境での強化、コンプライアンス、監視、データの品質に関する明確な注意事項があります。エンティティの一致が行われた<em>理由</em>を正当化する必要のあるシステムを構築している場合、このパターンは真剣に検討する価値があります。このシリーズを通して、エンティティ解決は必ずしも難解なものではないということが伝われば幸いです。適切に関心事を分離することで、それは論理的に考え、測定し、改善できるものになります。</p><p>この研究はまた、より広範なアーキテクチャパターンを示唆しています。浮かび上がってくるのは、古典的な検索拡張生成（RAG）の、わずかではあるが重要な進化です。検索結果を直接生成に供給するのではなく、明示的な評価ステップを導入します。LLMはまず、取得された候補を評価し、妥当性を確認するために使用され、承認された結果のみが生成の強化に使用されます。これは、Generation-Augmented Retrieval-Augmented Generation with Evaluation、つまりGARAGEと名付けられるでしょう。うまい頭字語が嫌いな人なんていませんから。</p><p>このパターンは、他にどのような用途で活用できるでしょうか？信頼性、透明性、そして論理的な説明を必要とするシステムは、まさにうってつけの候補と言えます。この分野における今後の研究は、今回得られた成果と同様に説得力のあるものとなるはずであり、コミュニティが今後どのような展開を見せるのか、非常に楽しみです。</p><h2>次のステップ：試してみましょう</h2><p>究極のチャレンジが実際に動作する様子をご覧になりたいですか？実際の実装、詳細な説明、実践的な例を含む完全なウォークスルーについては、<a href="https://github.com/jesslm/entity-resolution-lab-public/tree/main/notebooks#:~:text=5%20minutes%20ago-,05_ultimate_challenge_v3.ipynb,-Initial%20public%20lab"><strong>Ultimate Challenge notebook</strong></a>を参照してください。</p><p>完全なエンティティ解決パイプラインにより、本番での使用に必要なコアコンセプトとアーキテクチャが示されています。これを基盤に、透明性と説明可能性を維持しながら、ニュース記事を監視し、エンティティの言及を追跡し、どのエンティティがどの記事に登場するのかについての質問に回答するシステムを構築できます。
</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/entity-resolution-elasticsearch-llm-challenges</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/entity-resolution-elasticsearch-llm-challenges</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[ハイブリッド検索]]></category>
    <dc:creator><![CDATA[Jessica Moszkowicz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc58be329ffebcd60/6a17043e47d49c0bc62d88ab/70fb0ff949f6db9ac9b8a28ecb4329ab915ebf46-720x420.png" length="0" type="image/png"/>
    <pubDate>Fri, 13 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[ElasticsearchとLLMによるエンティティ解決（第2部）：LLM判定とセマンティック検索によるエンティティのマッチング]]></title>
    <description><![CDATA[Elasticsearch でのエンティティ解決にセマンティック検索と透過的なLLM判断を使用します。]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/search-labs/blog/entity-resolution-llm-elasticsearch">第1部</a>では、ウォッチリストを作成し、エンティティの言及を抽出しました。これで、「言及が実際にどのエンティティを指しているのか」という難しい質問に答える準備ができました。このシリーズの最初のブログの例に戻りましょう。ここでは、エンティティ解決が必要な理由を説明しています。「新しいSwiftアップデートが登場しました！」この見出しにもう少し文脈が添えられていると想像してください。</p><ol><li><p>新しいSwiftアップデートが登場しました！開発者たちは新しい機能を試したがっています。</p></li><li><p>新しいSwiftアップデートが登場しました！新しいアルバムは来月リリースされます。</p></li></ol><p>この追加されたコンテキストにより、「Swift」という名前を正しいエンティティに解決できるはずです。</p><p><a href="https://www.elastic.co/search-labs/blog/entity-resolution-llm-elasticsearch">前回の投稿</a>では、ウォッチリストを設定し、追加のコンテキストでエンティティを充実させました。上記の例を見ると、リストには少なくとも「Taylor Swift」と「Swift Programming Language」の2つのエンティティが必要です。また、テキストからエンティティの言及を抽出する方法も説明しました。これらの例はどちらも「Swift」を抽出します。これらの材料、強化された監視リスト、抽出されたエンティティが揃ったところで、いよいよショーの主役であるエンティティマッチングを紹介する準備が整いました。</p><p><strong>注意：</strong>これは、エンティティマッチングの概念を教えるために設計された教育用プロトタイプです。本番システムは、異なる大規模言語モデル（LLM）、カスタムマッチングルール、特殊な判断パイプライン、または複数のマッチング戦略を組み合わせたアンサンブルアプローチを使用する可能性があります。</p><h2>問題：マッチングが難しい理由</h2><p>人間の言語とは驚くべきものです。その最も興味深い特性の1つは、その無限の創造性です。無限の数の新しい文を生成し、理解することができます。そうであるなら、エンティティ解決において正確な一致が稀なのも不思議ではありません。作家は可能な限り創造的であろうと努めます。エンティティが言及されるたびにフルネームを書いたり読んだりしなければならないとしたら、かなり面倒です。そのため、厳密な一致は簡単ですが、現実には、より洗練されたエンティティ解決アプローチが必要です。それは、人間の作者の無限の創造性に少なくとも部分的には対応できるほどに堅牢なアプローチであるべきです。そのため、私たちは問題を2つのステップに分けます。まずはElasticsearchを使用して大規模な候補を取得し、次にLLMを使用してそれらの候補が実際に同じ現実世界のエンティティを指しているかどうかを判断します。</p><h2>解決策：透明性の高いLLM判断による3段階のマッチング</h2><p>私たちはコンピューターの使い方におけるパラダイムシフトの真っ只中にあります。インターネットの台頭がローカルコンピューティングからグローバルに接続されたネットワークへと私たちを導いたように、生成AIはコンテンツ、コード、情報の作成方法を根本的に変えています。実際、このシリーズに付随する教育プロトタイプは、作者の慎重な指示のもと、LLMを使用してほぼ「バイブコーディング」のみで作成されました。これは、LLMが人間の言語に本来備わっている生産性を実現している、あるいは実現するだろうということと同義ではありませんが、エンティティ解決を支援する強力なリソースが手に入ったことを意味します。</p><p>生成AIでよく使うパターンは、Retrieval-Augmented Generation（RAG）です。ここにおいて、<em>取得（retrieval）</em>とは、エンティティ候補を取得すること（回答を生成することではない）を意味し、LLMは一致の評価と説明にのみ使用されます。エンドツーエンドのエンティティ解決についてLLMに支援を依頼する<em>こともできます</em>が、これは時間と費用の両面でコストのかかるアプローチです。RAGは、より効率的な方法でLLMにコンテキストを提供することでLLMの作業を支援し、それによってLLMがエンティティ解決を効率的に支援できるようにします。</p><p>RAGの取得部分については、再びElasticsearchを利用します。まず、正確な一致、エイリアスとの一致、そしてキーワード検索とセマンティック検索を組み合わせたハイブリッド検索という組み合わせを使用して、潜在的な一致を検索します。一致する可能性のある項目が見つかったら、LLMに送信して判断を仰ぎます。LLMは最終的な一致評価者として機能します。また、LLMにその理由を説明させます。これは他のエンティティ解決システムとの重要な差別化要因です。これらの説明がなければ、エンティティ解決はブラックボックスになります。説明があれば、一致にどんな意味があるのか自分で確認できます。</p><h2>主な概念：3段階マッチング、ハイブリッド検索、透過的なLLM判断</h2><p><strong>3段階マッチングとは？</strong>このプロジェクトの開始時に、セマンティック検索がシステムの重要な一部になるという仮説を立てましたが、すべての一致にこのような高度な検索が必要なわけではありません。効率的にマッチングを見つけるために、私たちは段階的なアプローチを取ります。まず、キーワード検索で正確な一致を確認します。そのような一致が見つかった場合、作業は完了し、先に進むことができます。完全一致が失敗した場合は、エイリアス一致を使用します。このプロトタイプでは、簡素化のために、キーワードとの完全一致によるエイリアスマッチングも行われています。本番環境では、正規化、翻字ルール、あいまい一致、またはキュレートされたエイリアステーブルを使用してこのステップを拡張する場合があります。それでも最初の2つのステップで一致する可能性のあるものが見つからない場合は、Elasticsearchの逆順位融合（RRF）を使用したハイブリッド検索によるセマンティック検索を導入します。</p><p><strong>ハイブリッド検索とは？</strong>Elasticsearchでは、セマンティック検索を使用して、コンテキストを考慮した意味のある一致を見つけることができます。Elasticsearchは、ベクトル検索とハイブリッド検索に広く使用されています。セマンティック類似性は意味を理解する上で強力ですが、構造化されたフィルタリング（例えば、時間範囲、場所、または識別子による）の代替にはならず、正確な一致が利用可能な場合は多くの場合不必要です。Elasticsearchは語彙検索で名声を博しており、これはセマンティック検索が適さないタスクに最適です。両方のアプローチを最大限に活用するために、単一のハイブリッドクエリで語彙検索とセマンティック検索を併用します。次に、結果をマージして、RRFを使用して最も一致する可能性が高いものを見つけます。このプロトタイプでは、上位2つの結果が、LLM判定に送信できる潜在的な一致となります。</p><p><strong>LLM判定を使用する理由とは？</strong>LLMの判断と説明により、システムは曖昧さとコンテキストを透過的に処理できます。これは「the president」のような場合において重要です。コンテキストによって複数のエンティティを指す可能性がありますが、システム内でニックネームや文化的なバリエーションをうまく機能させることもできます。最後に、制裁リストからエンティティを識別するなどのミッションクリティカルなタスクを検討する場合、システムを信頼するために、一致が受け入れられた理由を把握する必要があります。重要なのは、LLMはコーパス全体を検索せず、Elasticsearchによって返された少数の候補のみを評価するということです。</p><h2>実際の結果：LLM推論によるマッチング</h2><p>あらゆる自然言語処理タスクにおける大きな課題は、期待される結果が何であるかを示す「答えの鍵」となるゴールデンドキュメントを作成することです。これがなければ、システムがタスクをどの程度うまく実行するかを判断することはほぼ不可能ですが、そのようなドキュメントを作成するのは面倒なプロセスになる可能性があります。エンティティ解決のプロトタイプでは、テストに使用できるデータの設定に生成AIを再度利用しました。</p><p>まず、ニックネームや翻字などのいくつかのチャレンジタイプを定義し、次にLLMに、システムにとって徐々に大きく、より困難になる階層化されたデータセットコレクションを作成するように依頼しました。データセットの作成は期待していたほど簡単ではありませんでした。LLMでは、正解を得るのがあまりにも簡単すぎるため、「チート」が行われる傾向が強くなりました。例えば、あるチャレンジタイプは意味的なコンテキストに重点を置いています。このタイプには、「ロシアの作家」を「レフ・トルストイ」に解決することなどが含まれます。LLMは誤って「ロシアの作家」を「レフ・トルストイ」の別名として入力したため、一致を見つけるためのハイブリッド検索の必要性がなくなりました。</p><p>このような問題を修正するために何度かリファクタリングを行った結果、5つのデータセット層が使用できるようになりました。第1〜4層は徐々に規模が大きくなり、チャレンジの種類も増えました。第5層は「究極のチャレンジ」データセットで、すべてのチャレンジタイプから最も難しい例で構成されていました。すべてのテストデータは<a href="https://github.com/jesslm/entity-resolution-lab-public/tree/main/comprehensive_evaluation">包括的な評価ディレクトリ</a>で利用可能です。</p><p>プロンプトベースのエンティティ解決アプローチを評価するため、私たちは第4層データセットに注目しました。重要な注意点は、エンティティの一致品質に焦点を当てることができるように、評価が制御された実験として実施されたことです。ウォッチリストデータは事前にコンテキストで強化されており、エンティティは事前に記事から抽出され、評価で抽出精度ではなくマッチングに重点が置かれることが保証されました。これにより、一致品質が分離されます。エンドツーエンドのパフォーマンスは、抽出リコールとエンリッチメント品質にも依存します。</p><h3>評価データセット</h3><p>第4層の評価データセットは、システムの機能の包括的なテストを提供します。[1]</p><ul><li><p><strong>監視リストのエンティティ：</strong>さまざまなタイプ（人、組織、場所）にわたる66個のエンティティ。</p></li><li><p><strong>テスト記事：</strong>実際のエンティティ解決シナリオを網羅した69件の記事。</p></li><li><p><strong>予想される一致数：</strong>すべての記事で206件のエンティティが一致すると予想。</p></li><li><p><strong>チャレンジタイプ：</strong>エンティティ解決のさまざまな側面をテストする15種類のチャレンジタイプ。</p></li></ul><p>データセットに含まれる課題の種類は以下の通りです。</p><ul><li><p><strong>ニックネーム：</strong> 「ボブ・スミス」→「ロバート・スミス」（7つの記事）。</p></li><li><p><strong>称号と敬称：</strong>「Dr. Sarah Williams」→「Sarah Williams」（5つの記事）。</p></li><li><p><strong>意味的文脈：</strong> 「ロシアの作家」→「レフ・トルストイ」（8 つの記事）。</p></li><li><p><strong>多言語名：</strong>異なる文字での名前の取り扱い（6つの記事）。</p></li><li><p><strong>事業体：</strong>会社名のバリエーション（7つの記事）。</p></li><li><p><strong>役員紹介：</strong> 「Microsoft CEO」→「Satya Nadella」（5つの記事）。</p></li><li><p><strong>政治指導者：</strong>タイトルベースの参考文献（5つの記事）。</p></li><li><p><strong>イニシャル：</strong> 「J. Smith」→「John Smith」（3つの記事）。</p></li><li><p><strong>名前の順序のバリエーション：</strong>さまざまな名前の順序付け規則（3つの記事）。</p></li><li><p><strong>切り捨てられた名前：</strong>名前の一部一致（3つの記事）。</p></li><li><p><strong>名前の分割：</strong>名前がテキストに分割（3つの記事）。</p></li><li><p><strong>スペース/ハイフンの欠落：</strong>書式のバリエーション（2つの記事）。</p></li><li><p><strong>翻字：</strong>文字間の名前の一致（2つの記事）。</p></li><li><p><strong>複合チャレンジ：</strong>1つの記事に複数のチャレンジ（6つの記事）。</p></li><li><p><strong>複雑なビジネス：</strong>階層的なビジネス関係（5つの記事）。</p></li></ul><p>プロンプトベースのエンティティ解決がどのように機能したか見てみましょう。</p><h3>全体的なパフォーマンス</h3><p>結果は、LLMを活用したマッチ評価には大きな可能性があることを示していますが、重大な信頼性の問題も明らかにしています。各候補ペアはLLMによって評価される必要があるため、構造化された出力の失敗により、検索が適切に機能している場合でも受け入れと呼び出しが抑制される可能性があります。</p><p>メトリック</p><p>値</p><p>精度</p><p>83.8%</p><p>リコール</p><p>62.6％</p><p>F1スコア</p><p>71.7％</p><p>見つかった一致の合計</p><p>344</p><p>LLM合格率</p><p>44.8％</p><p>エラー率</p><p>30.2%</p><h3>エラー率の問題</h3><p>このプロトタイプで最初に行うステップは、Elasticsearchを使用して潜在的な一致ペアを作成することであることを思い出してください。これらの潜在的な一致はそれぞれ、LLMによって評価される必要があります。これらすべての一致を効率的に処理するために、LLM呼び出しをバッチ処理します。これにより、APIのコストと待ち時間が削減されますが、出力に不正な形式のJSONが表示されるリスクも高まります。バッチサイズが大きくなると、JSONはより長く複雑になり、LLMが無効なJSONを生成する可能性が高くなります。これがエラー率30%となる原因です。評価では、リクエストごとに5つの一致のバッチサイズを使用しました。この保守的なバッチサイズでも、JSON解析エラーが発生し、評価結果が大幅に歪んでいます。</p><h2>次のステップ：LLM統合の最適化</h2><p>セマンティック検索とLLMによる判断を用いてエンティティをマッチングしたことで、完全なエンティティ解決パイプラインが完成しました。ただし、このアプローチでは、モデルの判断は正しいものの、その出力が使用できない場合に、新たな障害モードが発生します。LLM統合を最適化することで、信頼性とコスト効率を向上させることができます。次の投稿では、エラーとコストを削減しながら構造と型の安全性を保証する構造化出力に関数呼び出しを使用する方法について説明します。</p><h2>はじめましょう</h2><p>エンティティマッチングの実際の動作を確認したいですか？実際の実装、詳細な説明、実践的な例を含む完全なウォークスルーについては、<a href="https://github.com/jesslm/entity-resolution-lab-public/tree/main/notebooks#:~:text=5%20minutes%20ago-,03_entity_matching_v3.ipynb,-Initial%20public%20lab">エンティティマッチングノートブック</a>を参照してください。このノートブックでは、3段階の検索、RRFを使用したハイブリッド検索、LLMを利用した推論による判断を使用してエンティティを一致させる方法を正確に示します。</p><p><strong>注意：</strong>これは、概念を教えるために設計された教育用プロトタイプです。本番システムを構築するときは、モデルの選択、コストの最適化、レイテンシ要件、品質検証、エラー処理、監視など、教育に重点を置いたこのプロトタイプではカバーされていない追加の要素を考慮してください。</p><h2>メモ</h2><ol><li><p>これらのデータセットは合成されたもので教育用に設計されており、実際の課題に近似していますが、単一の本番ドメインを代表するものではありません。</p></li></ol>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-entity-resolution-llm-semantic-search</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-entity-resolution-llm-semantic-search</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[ハイブリッド検索]]></category>
    <dc:creator><![CDATA[Jessica Moszkowicz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltefc59243d9990405/6a17056ab339d5778f769ebf/473ca4357c7d60f690edbd2a844acda169aca9c3-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic Agent BuilderとStrands Agents SDKの使用を開始]]></title>
    <description><![CDATA[Elastic Agent Builderでエージェントを作成する方法を学び、次にStrands Agents SDKで管理されたA2Aプロトコルを介してエージェントを使用する方法を学びましょう。]]></description>
    <content:encoded><![CDATA[<p>AIエージェントのアイデアをお持ちですか？おそらく、データを使って何かを行うことが関係しているでしょう。エージェントが有用なアクションを開始するには、決定を下す必要があり、正しい決定を下すには正しいデータが必要だからです。</p><p>Elastic Agent Builderは、データ接続型AIエージェントを簡単に構築できるようにします。このブログ記事でその方法を説明します。まず、Elasticに格納されているデータにアクセスするMCPツールを使ってエージェントを作成するのに必要なすべてのステップを見ていきましょう。次に、Strands Agents SDKとそのAgent2Agent（A2A）機能を使用してエージェントを操作します。<a href="https://strandsagents.com/">Strands Agents SDK</a>は、望む結果を得るために十分なコードでエージェント向けアプリを構築するマルチエージェントAI開発プラットフォームです。</p><p>AIエージェントを構築しましょう。このエージェントは、RPS+というゲームをプレイします。これは古典的な「じゃんけん」に追加のひねりを加えたもので、プレイヤーにいくつかの追加の選択肢を与えます。</p><h2>要件</h2><p>こちらのブログ記事の手順に従うために必要なものは次のとおりです。</p><ul><li><p>ローカルコンピューターで実行されているテキストエディター</p><ul><li><p>このブログ記事の例では<a href="https://code.visualstudio.com/download">Visual Studio Code</a>を使用します。</p></li></ul></li><li><p>ローカルコンピューターで実行されている<a href="https://www.python.org/downloads/">Python 3.10以上</a></p></li></ul><h2>Serverlessプロジェクトを作成する</h2><p>最初に必要なのは、Elastic Agentビルダーを含むElasticsearch Serverlessプロジェクトです。</p><p><a href="http://cloud.elastic.co/">cloud.elastic.co</a>に移動して新しいElasticsearch Serverlessプロジェクトを作成します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3472edce39ec0b81/6a17060e66c4f93c17f8bf57/31b6a5c1c30dacbb4d5e58d1c566071e7143a0c8-1600x879.gif" alt="" /><h2>インデックスを作成してデータを追加する</h2><p>次に、Elasticsearchプロジェクトにデータを追加します。開発者ツールを開き、コマンドを実行して新しいインデックスを作成し、そこにデータを挿入します。トップレベルのナビゲーションメニューから「開発者向けツール」を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltaedaa94068c07a17/6a17060f961e697558c4ce5f/f97d5af077504463155655a9e27c171a7f974f71-1600x879.jpg" alt="" /><p>コピーして、以下のPUTコマンドを開発者向けツールコンソールのリクエストインプットエリアに貼り付けてください。この文は「game-docs」という名前のElasticsearchインデックスを作成します。</p>PUT /game-docs
{
  "mappings": {
    "properties": {
      "title": { "type": "text" },
      "content": { 
        "type": "text"
      },
      "filename": { "type": "keyword" },
      "last_modified": { "type": "date" }
    }
  }
}<p>開発者ツールのステートメントの右側に表示される <strong>[リクエストの送信]</strong> ボタンをクリックします。開発者向けツールの対応エリアに<em>game-docs</em>インデックスが作成されたことを確認する通知が表示されるはずです。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt430c357b479d93af/6a170611a6c2b98191e79624/be0555a1930e4d4f58b7ed8b669c9b702532ed17-1600x880.jpg" alt="" /><p><em>game-docs</em>という名前のインデックスは、作成中のゲームのデータを保存するのに最適な場所です。ゲームに必要なすべてのデータを含むこのインデックスに、<em>rps+-md</em>という名前のドキュメントを配置しましょう。次のPUTコマンドをコピーして、開発者向けツールコンソールに貼り付けます。</p>PUT /game-docs/_doc/rps+-md
{
  "title": "Rock Paper Scissors +",
  "content": "
# Game Name
RPS+

# Starting Prompt
Let's play RPS+ !
---
What do you choose?

# Game Objects
1. Rock 🪨 👊
2. Paper 📜 🖐
3. Scissors ✄ ✌️
4. Light ☼ 👍
5. Dark Energy ☄ 🫱

# Judgement of Victory
* Rock beats Scissors
  * because rocks break scissors
* Paper beats Rock
  * because paper covers rock
* Scissors beat Paper
  * because scissors cut paper
* Rock beats Light
  * because you can build a rock structure to block out light
* Paper beats Light
  * because knowledge stored in files and paper books helps us understand light
* Light beats Dark Energy
  * because light enables humans to lighten up and laugh in the face of dark energy as it causes the eventual heat death of the universe
* Light beats Scissors
  * because light is needed to use scissors safely
* Dark Energy beats Rock
  * because dark energy rocks more than rocks. It rocks rocks and everything else in its expansion of the universe
* Dark Energy beats Paper
  * because humans, with their knowledge stored in files and paper books, can't explain dark energy 
* Scissors beat Dark Energy
  * because a human running with scissors is darker than dark energy

# Invalid Input
I was hoping for an worthy opponent
  - but alas it appears that time has past
  - but alas there's little time for your todo list when [todo:fix this] is so vast

# Cancel Game
The future belongs to the bold. Goodbye..
",
  "filename": "RPS+.md",
  "last_modified": "2025-11-25T12:00:00Z"
}<p>ステートメントの横にある<strong> [リクエストの送信] </strong>ボタンをクリックして実行し、<em>rps+-md</em>ドキュメントをgame-docsのインデックスに追加してください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt64d49e13754d5b25/6a17061214b270524be3c55d/3c01d8a4602de5c33337457591a388a4a4e3fad3-1600x879.jpg" alt="" /><p>クエリを実行するためのデータが用意されているはずです。Agent Builderを使用すると、クエリはこれまで以上に簡単になります。</p><p>トップレベルのナビゲーションメニューから<strong> [エージェント]</strong> を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb4d374bf2ba9135c/6a1706147d8d67468570e63e/82dbd2e9a439cabd5a5eea3d0ce005b87df0c3ea-1600x879.jpg" alt="" /><p>あとは、デフォルトのElastic AI Agentに「どんなデータがありますか？」と聞くだけです。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc0f879cf28772718/6a1706161949f7f25ee7a92d/f7a2f39c9d1486bdf02d9e88a732b540ac2e2cd1-1600x872.gif" alt="" /><p>Elastic AI Agentはデータを評価し、保有するデータの簡潔な説明を返します。</p><h2>ツールを作成する</h2><p>さて、Elasticにデータがいくつか入ったので、それを活用してみましょう。Agent Builderには、エージェントが必要なデータにアクセスし、タスクに適したコンテキストを得られるように<a href="https://modelcontextprotocol.io/">MCP</a>ツールを作成するための組み込みサポートが含まれています。ゲームデータを取得できるシンプルなツールを作りましょう。</p><p>Agent Builderのアクションメニューをクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7802a6b94e81440c/6a170618ab7f085287db9db4/0e327c202674dda33bcc0e494d2b588fa8b32e4f-1600x879.png" alt="" /><p>メニューオプションから <strong>[すべてのツールを表示]</strong> を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f52ffe114fb6ea7/6a17061a4a531b801b36a884/1ebf58650e9fb56750d3f0b1700fab50b44f9bdf-1600x879.png" alt="" /><p><strong>[+ 新しいツール] </strong>をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8090769f6c4d1899/6a17061c286714294093e219/6c03a7f28b99ac2d805f34f39948979893316a00-1600x879.png" alt="" /><p><strong>ツール作成</strong>フォームで<a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql"><strong>ES|QL</strong></a>を選択します。ツール<strong>タイプ</strong>として次の値を入力します。</p><p><strong>ツールID</strong>について：</p>example.get_game_docs<p><strong>説明</strong>について：</p>Get RPS+ doc from Elasticsearch game-docs index.<p><strong>構成</strong>については、以下のクエリを<strong>ES|QLクエリ</strong>テキスト領域にします。</p>FROM game-docs | WHERE filename == "RPS+.md"<p>完了した<strong>ツール作成</strong>フォームは次のようになります。ツールを作成するには、 <strong>[保存]</strong> をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt77034c305198217a/6a17061e66c4f9e54ef8bf5e/b6c93e344600f319b9d2c3030020cf2d171ac1c4-1600x1312.png" alt="" /><p>ツールラックに新しいツールが追加されました。ツールはラックに掛けておけばよいというものではなく、有効に活用されるべきものです。新しいカスタムツールを使用できるエージェントを作成しましょう。</p><h2>エージェントを作成し、ツールを割り当てます。</h2><p>Agent Builderを使えば、エージェントの作成は驚くほど簡単です。いくつかの詳細を記載したエージェントの指示を入力するだけで十分です。それではエージェントを作成しましょう。</p><p><strong>[エージェントを管理]</strong> をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltaa8a83fc2f3758a9/6a1706201949f71a10e7a931/53934b93db07187e251d4b321cb9ca647e2fd51b-1600x858.png" alt="" /><p><strong>[+ 新しいエージェント] </strong>をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3778403c5101a000/6a17062160084be12f3c449e/fae3ad8f31e71a6dfd044e1daa025a4e280b4e68-1600x490.png" alt="" /><p><strong>新しいエージェント</strong>フォームに次の情報を入力します。</p><p><strong>エージェントID</strong>には以下のテキストを入力します。</p>rps_plus_agent<p><strong>カスタム指示</strong>テキスト領域には次の指示を入力します。</p>When prompted, if the prompt contains an integer, then select the corresponding numbered item in the list of "Game Objects" from your documents. Otherwise select a random game object. This is your chosen game object for a single round of the game.

# General Game Rules
* 2 players
    - the user: the person playing the game
    - you: the agent playing the game and serving as the game master
* Each player chooses a game object which will be compared and cause them to tie, win or lose.

# Start the game
1. This is the way each new game always starts. You make the first line of your response only the name of your chosen game object. 

2. The remainder of your response should be the "Starting Prompt" text from your documents and generate a list of "Game Objects" for the person playing the game to choose a game object from.  

# End of Game: The game ends in one of the following three outcomes:
1. Invalid Input: If the player responds with an invalid game object choice, respond with variations of the "Invalid Input" text from your documents and then end the game.

2. Tie: The game ends in a tie if the user chooses the same game object as your game object choice.

3. Win or Lose: The game winner is decided based on the "Judgement of Victory" conditions from your documents. Compare the user's game object choice and your game object choice and determine who chose the winning game object.

# Game conclusion
Respond with a declaration of the winner of the game by outputting the corresponding text in the "Judgement of Victory" section of your documents.<p><strong>表示名</strong>には以下のテキストを入力します。</p>RPS+ Agent<p><strong>表示の説明</strong>には以下のテキストを入力します。</p>An agent that plays the game RPS+<p><strong>[ツール]</strong> タブをクリックして、以前に作成したカスタムツールをエージェントに提供します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte5b0fe00abdde07c/6a17062314b2704bc4e3c563/1778f64bc3a1b4004998dc3668ef7f666788e193-1600x1390.png" alt="" /><p>先ほど作成した<em>example.get_game_docs</em>ツールのみを選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2210212e07e06104/6a170625a929cf3277ae08d1/7d734cd80161bcc058817482eb330ffcf1cb567b-1600x1363.png" alt="" /><p><strong>[保存]</strong> をクリックして新しいエージェントを作成します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6e3afc1918e26f14/6a170627ab7f084746db9db8/c0014faf605ce50c03679ed0d073bd9f3ae7234d-1600x468.png" alt="" /><p>新しいエージェントをテストしてみましょう。エージェントのリストから任意のエージェントとチャットを開始するための便利なリンクがあります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blteb4b69dc5971d3a0/6a1706286f7f046840914743/b7d6943ad90a4f68691207caf66b81742e712145-1600x560.png" alt="" /><p>「start game」と入力すると、ゲームが始まります。うまくいきました！</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte5b621d602223dff/6a17062ab339d568a1769ef8/984d008e4cc3f08cc1f101720673b0f7347c066c-1600x874.gif" alt="" /><p>エージェントが応答の上部にゲームオブジェクトの選択を表示することがわかります。これは、エージェントの選択を確認し、ゲームが期待どおりに機能していることを確認できる点で便利です。しかし、自分が選択する前に相手の選択がわかっていると、じゃんけんゲームはあまり楽しくありません。ゲームを最終形に磨き上げるために、コードでエージェントを制御できるエージェントオーケストレーションプラットフォームを使用できます。</p><p>Strands Agents SDKがチャットに参加します。</p><h2>Strands Agents SDK</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt73901ec745a97fbf/6a17062c964cea23c808bab3/c195bba6ff2754f5d8fda174a0c1d247bc283710-456x156.png" alt="" /><p>新しいエージェント開発フレームワークを試してみたい場合は、<a href="https://strandsagents.com/latest/">Strands Agents SDK</a>がおすすめです。<a href="https://aws.amazon.com/blogs/opensource/introducing-strands-agents-an-open-source-ai-agents-sdk/">Strands Agents SDKはAWSから2025年5月に</a>オープンソースの<a href="https://github.com/strands-agents/sdk-python">Python</a>実装としてリリースされ、現在は<a href="https://dev.to/aws/strands-agents-now-speaks-typescript-a-side-by-side-guide-12b3">Typescript</a>版もあります。</p><h2>PythonでStrands Agents SDKの使用を開始</h2><p>コーディングエンジンを起動して、Strandsエージェントを使用してA2Aプロトコル経由で<em>RPS+エージェント</em>を制御するサンプルアプリのクローン作成と実行のプロセスを早速実行してみましょう。RPS+ゲームの微調整バージョンを作成し、エージェントの選択がプレイヤーの選択後に明らかになるようにしてみましょう。結局のところ、じゃんけんのようなゲームを楽しいものにするのは推測と驚きの結果だからです。</p><p>ローカルコンピューターで<a href="https://code.visualstudio.com/download">Visual Studio Code</a>を開き、新しいターミナルを開きます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3de752025d62993f/6a17062d0c4857f16501a997/2339cc37c89a3524f2b2a21684bc61dae958e1cf-915x460.jpg" alt="" /><p>新しく開いたターミナルで、以下のコマンドを実行してElasticsearch Labsリポジトリをクローンします。</p>git clone https://github.com/elastic/elasticsearch-labs<p>次の<em>cd</em>コマンドを実行して、ディレクトリをelasticsearch-labsディレクトリに変更します。</p>cd elasticsearch-labs<p>次に、次のコマンドを実行して、Visual Studio Codeでリポジトリを開きます。</p>code .<p>Visual Studio File Explorerで、<em>supporting-blog-content</em>フォルダーと<em>agent-builder-a2a-strands-agents</em>フォルダーを展開し、<em>elastic_agent_builder_a2a_rps+.py</em>ファイルを開きます。Visual Studio Codeで開いたファイルは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt65ef8036a70bcaf1/6a17062f1949f7af36e7a935/d153b19e0e016c701576edb99ccab5af7c554f34-1484x1530.jpg" alt="" /><p>テキストエディターに表示される<em>elastic_agent_builder_a2a_rps+.py</em>の内容は次のとおりです。</p>import asyncio
from dotenv import load_dotenv
from uuid import uuid4
import httpx
import os
import random
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, TextPart

DEFAULT_TIMEOUT = 60  # set request timeout to 1 minute


def create_message(*, role: Role = Role.user, text: str, context_id=None) -&gt; Message:
    return Message(
        kind="message",
        role="user",
        parts=[Part(TextPart(kind="text", text=text))],
        message_id=uuid4().hex,
        context_id=context_id,
    )


async def main():
    load_dotenv()
    a2a_agent_host = os.getenv("ES_AGENT_URL")
    a2a_agent_key = os.getenv("ES_API_KEY")
    custom_headers = {"Authorization": f"ApiKey {a2a_agent_key}"}

    async with httpx.AsyncClient(
        timeout=DEFAULT_TIMEOUT, headers=custom_headers
    ) as httpx_client:
        # Get agent card
        resolver = A2ACardResolver(httpx_client=httpx_client, base_url=a2a_agent_host)
        agent_card = await resolver.get_agent_card(
            relative_card_path="/rps_plus_agent.json"
        )
        # Create client using factory
        config = ClientConfig(
            httpx_client=httpx_client,
            streaming=True,
        )
        factory = ClientFactory(config)
        client = factory.create(agent_card)
        # Use the client to communicate with the agent
        print("\nSending 'start game' message to Elastic A2A agent...")
        random_game_object = random.randint(1, 5)
        msg = create_message(text=f"start with game object {random_game_object}")
        async for event in client.send_message(msg):
            if isinstance(event, Message):
                context_id = event.context_id
                response_complete = event.parts[0].root.text
                # Get agent choice from the first line of the response
                parsed_response = response_complete.split("\n", 1)
                agent_choice = parsed_response[0]
                print(parsed_response[1])
        # User choice sent for game results from the agent
        prompt = input("Your Choice  : ")
        msg = create_message(text=prompt, context_id=context_id)
        async for event in client.send_message(msg):
            if isinstance(event, Message):
                print(f"Agent Choice : {agent_choice}")
                print(event.parts[0].root.text)


if __name__ == "__main__":
    asyncio.run(main())<p>このコードで何が起きているのか見てみましょう。<em><code>main()</code></em>メソッドから始めて、コードはエージェントのURLとAPIキーの環境変数にアクセスすることから始まります。その値を用いてエージェントカードを取得するための<em><code>httpx</code></em><code> client</code>を作成します。次に、クライアントはエージェントカードの詳細を使用して、「start game」リクエストをエージェントに送信します。ここで注目すべき興味深い点は、 <code>"start game"</code>リクエストの一部として<code>random_game_object</code>値が含まれていることです。この値は、Python の標準ライブラリの<em>random</em>モジュールで生成された乱数です。これを行う理由は、（AIエージェントを可能にする）強力なLLMがランダム性に関してはそれほど優れていないことが判明したためです。Pythonが助けてくれますので問題ありません。</p><p>コードの続きですが、エージェントが「start game」リクエストに応答すると、コードはエージェントのゲームオブジェクトセレクションを取り除き、<em>agent_choice</em>変数に保存します。対応の残りの部分は、エンドユーザーに対してテキストとして表示されます。次に、ユーザーはゲームオブジェクトの選択を入力するように求められ、それがエージェントに送信されます。次に、コードはエージェントのゲームオブジェクトの選択と、エージェントの最終的なゲーム結果の決定を表示します。</p><h2>エージェントのURLとAPIキーを環境変数として設定する</h2><p>サンプルアプリはローカルコンピュータ上で実行されるため、Agent Builderエージェントと通信するためには、Strands Agents SDKにエージェントのA2A URLとAPI Keyを提供する必要があります。この例のアプリは<em>`.env`</em>というファイルを使用してこれらの値を格納します。</p><p><em>env.example</em>ファイルのコピーを作成し、新しいファイル名を<em>.env</em>とします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta17961cbcb42985c/6a170631b0367dc5a072bc55/25ead5f15a17dedb777132a082097cffb06cae4d-1600x843.jpg" alt="" /><p>Elastic Agent Builderに戻りましょう。ここで必要な両方の値を取得できます。</p><p>ページの右上にあるAgent Builderアクションメニューから <strong>[すべてのツールを表示]</strong> を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt140885d7ebfcb969/6a1706327d8d67b17670e646/9c4f4e4a3bd76e11e0a182fa007a2f6aec7777b4-1600x880.jpg" alt="" /><p>ツールページ上部の<strong>MCPサーバー</strong>ドロップダウンをクリックし、<strong>[MCPサーバーURLをコピー] </strong>を選択してください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc153c2caa27e949b/6a170634a292997793d00f6d/6cde0de678bb6f81bef8a59deffb110ad6c6ce26-1600x882.jpg" alt="" /><p><strong>MCPサーバーのURL</strong>を<em>.env</em>に貼り付けます。<strong>&lt;YOUR-ELASTIC-AGENT-BUILDER-URL&gt;</strong>プレースホルダー値の代わりにファイルを使用します。ここで、URLを1箇所更新する必要があります。つまり、末尾のテキスト「mcp」を「a2a」に置き換えます。これは、Agent Strands SDKがElastic Agent Builderで実行されているエージェントと通信するために使用するプロトコルが<a href="https://a2a-protocol.org/">A2Aプロトコル</a>であるためです。</p><p>編集したURLは次のようになるはずです。</p>https://rps-game-project-12345a.kb.us-east-1.aws.elastic.cloud/api/agent_builder/a2a<p>Elastic Cloudで取得する必要があるもう1つの値は、APIキーです。最上位ナビゲーションで<strong>Elasticsearch</strong>をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltada5de819f31d8ff/6a170635b339d55ae9769efc/651676b9be65178cdad50b5d24f26441c0bf3f97-1600x549.jpg" alt="" /><p><strong>[APIキーをコピー] ボタン</strong>をクリックして、APIキーをコピーします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta18f85790df00706/6a170637cf4f257145b2d0bd/17f1e2ed5c7682630c71e75b0b09ffb1d9036210-1600x879.jpg" alt="" /><p>次に、Visual Studio Codeに戻り、<em>.env</em>ファイルにAPIキーを貼り付けて、<strong>&lt;YOUR-ELASTIC-API-KEY&gt;</strong>プレースホルダーテキストを置き換えます。<em>.env</em>ファイルは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt92ab4b37cdcca85e/6a1706386f7f0472ed914747/a357947e07f29c8c03382e00c7baedf04a399297-1600x286.jpg" alt="" /><h2>サンプルアプリを実行してください</h2><p>Visual Studio Codeで新しいターミナルを開いてください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8702d826849755d0/6a17063a60084b45ca3c44a2/33e1174c68ea1ed47c7fe62ab6a6da657c606f56-1413x711.jpg" alt="" /><p>まず、ターミナルで次の<em>cd</em>コマンドを実行します。</p>cd elasticsearch-labs/supporting-blog-content/agent-builder-a2a-strands-agents<p>次のコマンドを実行して、Python仮想環境を作成します。</p>python -m venv .venv<p>お使いのローカルコンピューターのオペレーティングシステムに応じて、以下のコマンドを実行して仮想環境を有効にしてください。</p><ul><li><p>MacOS/Linux</p></li></ul>source .venv/bin/activate<ul><li><p>Windows</p></li></ul>.venv\Scripts\activate<p>サンプルアプリはStrands Agents SDKを使用するため、このチュートリアルではこれをインストールする必要があります。以下のコマンドを実行して、Strands Agents SDKとその必要なPythonライブラリの依存関係をインストールします。</p>pip install -r requirements.txt<p>発射台を片付けてカウントダウンを開始する時間です。アプリを起動する準備ができました。後ろに下がってください。次のコマンドを使用して実行しましょう：</p>python elastic_agent_builder_a2a_rps+.py<p>RPS+のゲームに挑戦してみましょう。幸運を祈ります！</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbb3715672995fcfa/6a17063c6234e07b76db195f/041df81fbf1776f09e1243af0a435c4c0af6aca1-1600x948.gif" alt="" /><h2>関連コンテキストでAIアプリを構築</h2><p>AIエージェントの構築のスキルを習得できました。また、Strands Agents SDKのようなエージェント開発フレームワークで、A2Aを介してElastic Agent Builderエージェントを使用することがいかに簡単であるかをお分かりいただけたと思います。カスタムデータの関連コンテキストに接続されたAIエージェントの構築には<a href="https://cloud.elastic.co/registration?utm_source=agentic-ai-category&amp;utm_medium=search-labs&amp;utm_campaign=agent-builder">Elasticをお試し</a>ください。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-builder-a2a-strands-agents-guide</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-builder-a2a-strands-agents-guide</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[エージェント型AI]]></category>
    <dc:creator><![CDATA[Jonathan Simon]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3472edce39ec0b81/6a17060e66c4f93c17f8bf57/31b6a5c1c30dacbb4d5e58d1c566071e7143a0c8-1600x879.gif" length="0" type="image/gif"/>
    <pubDate>Mon, 15 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[LangGraph.jsとElasticsearchを使用して金融AI検索ワークフローを構築]]></title>
    <description><![CDATA[LangGraph.jsとElasticsearchを使用して、自然言語クエリを投資や市場分析のための動的な条件付きフィルターに変換するAIを活用した金融検索ワークフローを構築する方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>AI検索アプリケーションの構築では、多くの場合、複数のタスク、データ取得、データ抽出をシームレスなワークフローに調整する必要があります。LangGraphは、開発者がnodeベースの構造を使用してAIエージェントを管理することで、このプロセスを簡素化します。この記事では、<a href="https://langchain-ai.github.io/langgraphjs/">LangGraph.js</a>を使用して金融ソリューションを構築します。</p><h2>LangGraphの概要</h2><p><a href="https://langchain-ai.github.io/langgraphjs/">LangGraph</a>は、AIエージェントを構築し、ワークフロー内で管理してAI支援アプリケーションを作成するためのフレームワークです。LangGraphには、タスクを表す関数を宣言し、それらをワークフローのノードとして割り当てることができるノードアーキテクチャがあります。複数のノードが相互作用した結果がグラフになります。LangGraphは、モジュール式かつ構成可能なAIシステムを構築するためのツールを提供する、より広範な<a href="https://js.langchain.com/docs/introduction/">LangChain</a>エコシステムの一部です。</p><p>LangGraphが有用である理由をより深く理解するために、LangGraphを使用して問題のある状況を解決してみましょう。</p><h2>ソリューションの概要</h2><p>ベンチャーキャピタル企業では、投資家は多くのフィルタリングオプションを備えた大規模なデータベースにアクセスできますが、基準を組み合わせたい場合には困難で時間がかかります。これにより、関連するスタートアップの一部が投資対象として見つからない可能性があります。その結果、最適な候補を見つけるために多くの時間を費やしたり、機会を逃したりすることになります。</p><p>LangGraphとElasticsearchを使用することで、自然言語を用いてフィルターで検索することが可能となり、ユーザーが手動で複雑なリクエストを何十ものフィルターで構築する必要がなくなります。柔軟性を高めるために、ワークフローはユーザーの入力に基づいて2つのクエリタイプを自動的に決定します。</p><ul><li><p><strong>投資に焦点を当てたクエリ</strong>：スタートアップ企業の財務および資金調達の側面を対象としており、<a href="https://www.investopedia.com/articles/personal-finance/102015/series-b-c-funding-what-it-all-means-and-how-it-works.asp">資金調達ラウンド</a>、バリュエーション、<a href="https://www.investopedia.com/terms/r/revenue.asp">収益</a>を含みます。<em>例：</em>「シリーズAまたはシリーズBの資金調達額が800万ドル～2,500万ドルで、月間収益が50万ドルを超えるスタートアップを探してください。」</p></li><li><p><strong>市場重視のクエリ</strong>：<a href="https://en.wikipedia.org/wiki/Vertical_market">業界分野</a>、<a href="https://en.wikipedia.org/wiki/Target_market">地理的市場</a>、<a href="https://www.investopedia.com/terms/b/businessmodel.asp">ビジネスモデル</a>に重点を置き、特定のセクターまたは地域での機会の特定に役立ちます。<em>例：</em>「サンフランシスコ、ニューヨーク、ボストンのフィンテックおよびヘルスケアのスタートアップ企業を探してください」</p></li></ul><p>クエリを強固に保つため、LLMに<a href="https://www.elastic.co/docs/solutions/search/search-templates">検索テンプレート</a>を構築させ、完全な<a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/querydsl">DSLクエリ</a>の代わりとします。このようにすれば、必要なクエリを常に取得でき、LLMは空白を埋めるだけで済み、毎回必要なクエリを構築する責任を負う必要がなくなります。</p><h2>始めるために必要なもの</h2><ul><li><p>Elasticsearch APIキー</p></li><li><p>OpenAPI APIキー</p></li><li><p>Node 18以降</p></li></ul><h2>ステップ別のガイド</h2><p>このセクションでは、アプリがどのように見えるかを見てみましょう。<a href="https://www.typescriptlang.org/">TypeScript</a>はJavaScriptのスーパーセットで、静的な型を追加することでコードの信頼性を高め、保守性を向上させ、エラーを早期に発見して安全性を高めます。既存のJavaScriptとの完全な互換性を保ちながら、これを実現します。</p><p>ノードのフローは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt90db8f03f372608c/6a170986dc55de6e16e00d93/b47d7f238c4964a6febc0de7fe5e68b186f539c3-363x555.png" alt="" /><p>上記の画像はLangGraphによって生成されたもので、ノード間の実行順序と条件付きロジックを定義するワークフローを表しています。</p><ul><li><p><strong>decideStrategy：</strong>LLMを用いてユーザーのクエリを分析し、投資重視か市場重視の2つの専門的な検索戦略のどちらかを判断します。</p></li><li><p><strong>prepareInvestmentSearch：</strong>クエリからフィルター値を抽出し、財務および資金調達関連のパラメータを強調した定義済みテンプレートを構築します。</p></li><li><p><strong>prepareMarketSearch</strong> : フィルター値も抽出しますが、市場、業界、地理的コンテキストを重視したパラメータを動的に構築します。</p></li><li><p><strong>executeSearch：</strong>検索テンプレートを使用して構築されたクエリをElasticsearchに送信し、一致するスタートアップドキュメントを取得します。</p></li><li><p><strong>visualizeResults：</strong>最終結果を、資金、業界、収益などの主要なスタートアップ属性を示す明確で読みやすい要約にフォーマットします。</p></li></ul><p>このフローには「if」ステートメントとして機能する<a href="https://langchain-ai.github.io/langgraphjs/how-tos/branching/?h=conditional#how-to-create-branches-for-parallel-node-execution">条件分岐が</a>含まれており、ユーザーの入力に基づいて投資検索パスを使用するか、市場検索パスを使用するかを決定します。LLMにより駆動されるこの意思決定ロジックにより、ワークフローは適応的でコンテキストに応じたものになります。このメカニズムについては次のセクションで詳しく説明します。</p><h3>LangGraphの状態</h3><p>各ノードを個別に見る前に、ノードがどのように通信し、データを共有するかを理解する必要があります。そのために、LangGraphではワークフローの状態を定義することができます。これはノード間で共有される状態を定義します。</p><p>状態は、ワークフロー全体の中間データを保存する共有コンテナとして機能します。ユーザーの自然言語クエリから始まり、選択された検索戦略、Elasticsearch用に準備されたパラメータ、取得された検索結果、最後にフォーマットされた出力が保持されます。</p><p>この構造により、すべてのノードが状態を読み取って更新できるようになり、ユーザー入力から最終的な視覚化までの一貫した情報の流れが保証されます。</p>const VCState = Annotation.Root({
  input: Annotation&lt;string&gt;(), // User's natural language query
  searchStrategy: Annotation&lt;string&gt;(), // Search strategy chosen by LLM
  searchParams: Annotation&lt;any&gt;(), // Prepared search parameters
  results: Annotation&lt;any[]&gt;(), // Search results
  final: Annotation&lt;string&gt;(), // Final formatted response
});<h3>アプリケーションをセットアップする</h3><p>このセクションのすべてのコードは<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch">elasticsearch-labsリポジトリ</a>で見つけることができます。</p><p>アプリが置かれるフォルダーでターミナルを開き、以下のコマンドで Node.js アプリケーションを初期化します。</p>npm init -y<p>これで、このプロジェクトに必要な依存関係をインストールできます。</p>npm install @elastic/elasticsearch @langchain/langgraph @langchain/openai @langchain/core dotenv zod &amp;&amp; npm install --save-dev @types/node tsx typescript<ul><li><p><strong><code>@elastic/elasticsearch</code></strong>: Elasticsearchのデータインジェストや検索などのリクエストを処理するのに役立ちます。</p></li><li><p><strong><code>@langchain/langgraph</code></strong>: すべてのLangGraphツールを提供するためのJS依存関係。</p></li><li><p><strong><code>@langchain/openai</code></strong>: LangChain用のOpenAI LLMクライアント。</p></li><li><p>@langchain/core：プロンプトテンプレートなど、LangChainアプリのコアとなる基本的な構成要素を提供します。</p></li><li><p><strong><code>dotenv</code></strong>:JavaScriptで環境変数を使用するために必要な依存関係。</p></li><li><p><strong><code>zod</code></strong>：型データへの依存関係。</p></li></ul><p><code>@types/node</code> <code>tsx</code> <code>typescript</code> により、TypeScriptコードを記述して実行できるようになります。</p><p>次に、以下のファイルを作成します。</p><ul><li><p><code>elasticsearchSetup</code><a href="http://ingest.ts/"><code>.ts</code></a>: Elasticsearchのマッピングを作成し、JSONファイルからデータを取り込み、Elasticsearchにデータを取り込みます。</p></li><li><p><a href="http://main.ts/"><code>main.ts</code></a>: LangGraphアプリケーションが含まれます。</p></li><li><p><code>.env</code>：環境変数を格納するファイル</p></li></ul><p><code>.env</code>ファイルに以下の環境変数を追加します。</p>ELASTICSEARCH_ENDPOINT="your-endpoint-here"
ELASTICSEARCH_API_KEY="your-key-here"
OPENAI_API_KEY="your-key-here"<p>OpenAPI APIKeyはコード上で直接使用されることはなく、ライブラリ<code>@langchain/openai</code>によって内部的に使用されます。</p><p>マッピングの作成、検索テンプレートの作成、データセットのインジェストに関するすべてのロジックは、<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts"><code>elasticsearchSetup.ts</code></a>ファイルにあります。次のステップでは、<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/main.ts"><code>main.ts</code></a>ファイルに焦点を当てていきます。また、データセットをチェックして、 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/dataset.json"><code>dataset.json</code></a>でデータがどのように表示されるかをよりよく理解することもできます。</p><h3>LangGraphアプリ</h3><p><code>main.ts</code>ファイルで、LangGraphアプリを統合するために必要な依存関係をいくつかインポートしましょう。このファイルには、ノード関数と状態宣言も含める必要があります。グラフの宣言は、次のステップで <code>main</code> メソッドで行われます。<code>elasticsearchSetup.ts</code>ファイルには、以降のステップでノード内で使用する Elasticsearch ヘルパーが含まれます。</p>import { writeFileSync } from "node:fs";
import { StateGraph, Annotation, START, END } from "@langchain/langgraph";
import { ChatOpenAI } from "@langchain/openai";
import { z } from "zod";
import {
  esClient,
  ingestDocuments,
  createSearchTemplates,
  INDEX_NAME,
  INVESTMENT_FOCUSED_TEMPLATE,
  MARKET_FOCUSED_TEMPLATE,
  createIndex,
} from "./elasticsearchSetup.js";

const llm = new ChatOpenAI({ model: "gpt-4o-mini" });<p>前述のように、LLMクライアントは、ユーザーの質問に基づいてElasticsearch検索テンプレートパラメーターを生成するために使用されます。</p>async function saveGraphImage(app: any): Promise&lt;void&gt; {
  try {
    const drawableGraph = app.getGraph();
    const image = await drawableGraph.drawMermaidPng();
    const arrayBuffer = await image.arrayBuffer();

    const filePath = "./workflow_graph.png";
    writeFileSync(filePath, new Uint8Array(arrayBuffer));
    console.log(`📊 Workflow graph saved as: ${filePath}`);
  } catch (error: any) {
    console.log("⚠️  Could not save graph image:", error.message);
  }
}<p>上記の方法はグラフ画像をpng形式で生成し、裏で<a href="https://mermaid.ink/">Mermaid.INK API</a>を利用しています。これは、スタイル設定された視覚化を使用してアプリノードがどのように相互作用するかを確認する場合に便利です。</p><h3>LangGraphノード</h3><p>次に、各ノードの詳細を見てみましょう。</p><h3>decideSearchStrategyノード</h3><p><code>decideSearchStrategy</code>ノードはユーザー入力を分析し、投資重視の検索を実行するか、市場重視の検索を実行するかを決定します。構造化された出力スキーマ（Zodで定義）を持つLLMを使用してクエリタイプを分類します。決定を下す前に、集計を使用してインデックスから利用可能なフィルターを取得し、モデルが業界、場所、資金調達データに関する最新のコンテキストを持っていることを確認します。</p><p>フィルタの可能な値を抽出してLLMに送信するために、<a href="https://www.elastic.co/docs/explore-analyze/query-filter/aggregations">集計</a>クエリを使ってElasticsearchインデックスから直接値を取得してみましょう。このロジックは<code>getAvailableFilters</code>というメソッドに割り当てられます。</p>async function getAvailableFilters() {
  try {
    const response = await esClient.search({
      index: INDEX_NAME,
      size: 0,
      aggs: {
        industries: {
          terms: { field: "industry", size: 100 },
        },
        locations: {
          terms: { field: "location", size: 100 },
        },
        funding_stages: {
          terms: { field: "funding_stage", size: 20 },
        },
        business_models: {
          terms: { field: "business_model", size: 10 },
        },
        lead_investors: {
          terms: { field: "lead_investor", size: 100 },
        },
        funding_amount_stats: {
          stats: { field: "funding_amount" },
        },
      },
    });

    return response.aggregations;
  } catch (error) {
    console.error("❌ Error getting available filters:", error);
    return {};
  }
}<p>上記の集約クエリを用いると、以下の結果が得られます。</p>{
  "industries": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "logistics",
        "doc_count": 5
      },
      ...
    ]
  },
  "locations": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "San Francisco, CA",
        "doc_count": 4
      },
      {
        "key": "New York, NY",
        "doc_count": 3
      },
      ...
    ]
  },
  "funding_stages": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "Series A",
        "doc_count": 8
      },
      ...
    ]
  },
  "business_models": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "B2B",
        "doc_count": 13
      },
      ...
    ]
  },
  "lead_investors": {
    "doc_count_error_upper_bound": 0,
    "sum_other_doc_count": 0,
    "buckets": [
      {
        "key": "Battery Ventures",
        "doc_count": 1
      },
      {
        "key": "Benchmark Capital",
        "doc_count": 1
      },
      ...
    ]
  },
  "funding_amount_stats": {
    "count": 20,
    "min": 4500000,
    "max": 35000000,
    "avg": 14075000,
    "sum": 281500000
  }
}<p>すべての結果は<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/responses/aggregationsResponse.json">こちらで</a>ご覧いただけます。</p><p>両方の戦略において、ハイブリッド検索を行うことにより、質問の構造化された部分（フィルター）とより主観的な部分（セマンティック）の両方を検出します。以下は<a href="https://www.elastic.co/docs/solutions/search/search-templates">検索テンプレート</a>を使用した両方のクエリの例です。</p>await esClient.putScript({
      id: INVESTMENT_FOCUSED_TEMPLATE,
      script: {
        lang: "mustache",
        source: `{
          "size": 5,
          "retriever": {
            "rrf": {
              "retrievers": [
                {
                  "standard": {
                    "query": {
                      "semantic": {
                        "field": "semantic_field",
                        "query": "{{query_text}}"
                      }
                    }
                  }
                },
                {
                  "standard": {
                    "query": {
                      "bool": {
                        "filter": [
                          {"terms": {"funding_stage": {{#join}}{{#toJson}}funding_stage{{/toJson}}{{/join}}}},
                          {"range": {"funding_amount": {"gte": {{funding_amount_gte}}{{#funding_amount_lte}},"lte": {{funding_amount_lte}}{{/funding_amount_lte}}}}},
                          {"terms": {"lead_investor": {{#join}}{{#toJson}}lead_investor{{/toJson}}{{/join}}}},
                          {"range": {"monthly_revenue": {"gte": {{monthly_revenue_gte}}{{#monthly_revenue_lte}},"lte": {{monthly_revenue_lte}}{{/monthly_revenue_lte}}}}}
                        ]
                      }
                    }
                  }
                }
              ],
              "rank_window_size": 100,
              "rank_constant": 20
            }
          }
        }`,
      },
    });<p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts#L119"><code>elasticsearchSetup.ts</code></a>ファイルに詳細が記載されているクエリを確認します。次のノードでは、2つのクエリのどちらを使用するかが決定されます。</p>// Node 1: Decide search strategy using LLM
async function decideSearchStrategy(state: typeof VCState.State) {
  // Zod schema for specialized search strategy decision
  const SearchDecisionSchema = z.object({
    search_type: z
      .enum(["investment_focused", "market_focused"])
      .describe("Type of specialized search strategy to use"),
    reasoning: z
      .string()
      .describe("Brief explanation of why this search strategy was chosen"),
  });

  const decisionLLM = llm.withStructuredOutput(SearchDecisionSchema);

  // Get dynamic filters from Elasticsearch
  const availableFilters = await getAvailableFilters();

  const prompt = `Query: "${state.input}"
    Available filters: ${JSON.stringify(availableFilters, null, 2)}

    Choose between two specialized search strategies:
    
    - investment_focused: For queries about funding stages, funding amounts, monthly revenue, lead investors, financial performance
    
    - market_focused: For queries about industries, locations, business models, market segments, geographic markets
    
    Analyze the query intent and choose the most appropriate strategy.
  `;

  try {
    const result = await decisionLLM.invoke(prompt);
    console.log(
      `🤔 Search strategy: ${result.search_type} - ${result.reasoning}`
    );

    return {
      searchStrategy: result.search_type,
    };
  } catch (error: any) {
    console.error("❌ Error in decideSearchStrategy:", error.message);
    return {
      searchStrategy: "investment_focused",
    };
  }
}<h3>prepareInvestmentSearchノードとprepareMarketSearchノード</h3><p>どちらのノードも共有ヘルパー関数<code>extractFilterValues</code>を使用します。この関数はLLMを活用して、業界、場所、資金調達段階、ビジネスモデルなど、ユーザーの入力に記載されている関連フィルターを識別します。このスキーマを使用して<a href="https://www.elastic.co/docs/solutions/search/search-templates">検索テンプレート</a>を構築します。</p>// Extract all possible filter values from user input
async function extractFilterValues(input: string) {
  const FilterValuesSchema = z.object({
    // Investment-focused filters
    funding_stage: z
      .array(z.string())
      .default([])
      .describe("Funding stage values mentioned in query"),
    funding_amount_gte: z
      .number()
      .default(0)
      .describe("Minimum funding amount in USD"),
    funding_amount_lte: z
      .number()
      .default(100000000)
      .describe("Maximum funding amount in USD"),
    lead_investor: z
      .array(z.string())
      .default([])
      .describe("Lead investor values mentioned in query"),
    monthly_revenue_gte: z
      .number()
      .default(0)
      .describe("Minimum monthly revenue in USD"),
    monthly_revenue_lte: z
      .number()
      .default(10000000)
      .describe("Maximum monthly revenue in USD"),
    industry: z
      .array(z.string())
      .default([])
      .describe("Industry values mentioned in query"),
    location: z
      .array(z.string())
      .default([])
      .describe("Location values mentioned in query"),
    business_model: z
      .array(z.string())
      .default([])
      .describe("Business model values mentioned in query"),
  });

  const extractorLLM = llm.withStructuredOutput(FilterValuesSchema);
  const availableFilters = await getAvailableFilters();

  const extractPrompt = `Extract ALL relevant filter values from: "${input}"
    Available options: ${JSON.stringify(availableFilters, null, 2)}
    Extract only values explicitly mentioned in the query. Leave fields empty if not mentioned.`;

  return await extractorLLM.invoke(extractPrompt);
}<p>検出された意図に応じて、ワークフローは2つのパスのいずれかを選択します。</p><p><strong>prepareInvestmentSearch：</strong>資金調達段階、資金調達額、投資家、更新情報などの財務指向の検索パラメータを構築します。クエリ テンプレート全体は<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts"><code>elasticsearchSetup.ts</code></a>ファイルで確認できます。</p>// Node 2A: Prepare Investment-Focused Search Parameters 
async function prepareInvestmentSearch(state: typeof VCState.State) {
  console.log(
    "💰 Preparing INVESTMENT-FOCUSED search parameters with financial emphasis..."
  );

  try {
    // Extract all filter values from input
    const values = await extractFilterValues(state.input);

    let searchParams: any = {
      template_id: INVESTMENT_FOCUSED_TEMPLATE,
      query_text: state.input,
      ...values,
    };

    return { searchParams };
  } catch (error) {
    console.error("❌ Error preparing investment-focused params:", error);
    return {
      searchParams: {},
    };
  }
}<p><strong>prepareMarketSearch：</strong>業界、地域、ビジネスモデルに重点を置いた市場主導のパラメータを作成します。クエリ全文は<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/langgraph-js-elasticsearch/elasticsearchSetup.ts"><code>elasticsearchSetup.ts</code></a>ファイルをご覧ください。</p>// Node 2B: Prepare Market-Focused Search Parameters
async function prepareMarketSearch(state: typeof VCState.State) {
  console.log(
    "🔍 Preparing MARKET-FOCUSED search parameters with market emphasis..."
  );

  try {
    // Extract all filter values from input
    const values = await extractFilterValues(state.input);

    let searchParams: any = {
      template_id: MARKET_FOCUSED_TEMPLATE,
      query_text: state.input,
      ...values,
    };

    return { searchParams };
  } catch (error) {
    console.error("❌ Error preparing market-focused params:", error);
    return {};
  }
}<h3>executeSearchノード</h3><p>このノードは、生成された検索パラメータを状態から取得し、最初にElasticsearchに送信します。次に、<a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-render-search-template">_render API</a>を使用してデバッグの目的でクエリを視覚化し、次に結果を取得するためのリクエストを送信します。</p>// Node 3: Execute Search
async function executeSearch(state: typeof VCState.State) {
  const { searchParams } = state;

  try {
    // getting formed query from template for debugging
    const renderedTemplate = await esClient.renderSearchTemplate({
      id: searchParams.template_id,
      params: searchParams,
    });

    console.log(
      "📋 Complete query:",
      JSON.stringify(renderedTemplate.template_output, null, 2)
    );

    const results = await esClient.searchTemplate({
      index: INDEX_NAME,
      id: searchParams.template_id,
      params: searchParams,
    });

    return {
      results: results.hits.hits.map((hit: any) =&gt; hit._source),
    };
  } catch (error: any) {
    console.error(`❌ ${state.searchParams.search_type} search error:`, error);
    return { results: [] };
  }
}<h3>visualizeResultsノード</h3><p>最後に、このnodeはElasticsearchの結果を表示します。</p>// Node 4: Visualize results
async function visualizeResults(state: typeof VCState.State) {
  const results = state.results || [];

  let formattedResults = `🎯 Found ${results.length} startups matching your criteria:\n\n`;

  results.forEach((startup: any, index: number) =&gt; {
    formattedResults += `${index + 1}. **${startup.company_name}**\n`;
    formattedResults += `   📍 ${startup.location} | 🏢 ${startup.industry} | 💼 ${startup.business_model}\n`;
    formattedResults += `   💰 ${startup.funding_stage} - $${(
      startup.funding_amount / 1000000
    ).toFixed(1)}M\n`;
    formattedResults += `   👥 ${startup.employee_count} employees | 📈 $${(
      startup.monthly_revenue / 1000
    ).toFixed(0)}K MRR\n`;
    formattedResults += `   🏦 Lead: ${startup.lead_investor}\n`;
    formattedResults += `   📝 ${startup.description}\n\n`;
  });

  return {
    final: formattedResults,
  };
}<p>プログラム的には、グラフ全体は次のようになります。</p>  const workflow = new StateGraph(VCState)
    // Register nodes - these are the processing functions
    .addNode("decideStrategy", decideSearchStrategy)
    .addNode("prepareInvestment", prepareInvestmentSearch)
    .addNode("prepareMarket", prepareMarketSearch)
    .addNode("executeSearch", executeSearch)
    .addNode("visualizeResults", visualizeResults)
    // Define execution flow with conditional branching
    .addEdge(START, "decideStrategy") // Start with strategy decision
    .addConditionalEdges(
      "decideStrategy",
      (state: typeof VCState.State) =&gt; state.searchStrategy, // Conditional function
      {
        investment_focused: "prepareInvestment", // If investment focused -&gt; RRF template preparation
        market_focused: "prepareMarket", // If market focused -&gt; dynamic query preparation
      }
    )
    .addEdge("prepareInvestment", "executeSearch") // Investment prep -&gt; execute
    .addEdge("prepareMarket", "executeSearch") // Market prep -&gt; execute
    .addEdge("executeSearch", "visualizeResults") // Execute -&gt; visualize
    .addEdge("visualizeResults", END); // End workflow<p>ご覧のとおり、アプリが次にどの「パス」またはノードを実行するかを決定する条件付きエッジがあります。この特徴は、ワークフローに分岐ロジックが必要な場合、例えば複数のツールから選択する場合や、人間が関与するステップを含む場合に有用です。</p><p>LangGraph のコア機能を理解したら、コードが実行されるアプリケーションをセットアップできます。</p><p>すべてを<code>main</code>メソッドで組み合わせ、ここではすべての要素をワークフロー変数下のグラフとして宣言します。</p>async function main() {
  await createIndex();
  await createSearchTemplates();
  await ingestDocuments();

  // Create the workflow graph with shared state
  const workflow = new StateGraph(VCState)
    // Register nodes - these are the processing functions
    .addNode("decideStrategy", decideSearchStrategy)
    .addNode("prepareInvestment", prepareInvestmentSearch)
    .addNode("prepareMarket", prepareMarketSearch)
    .addNode("executeSearch", executeSearch)
    .addNode("visualizeResults", visualizeResults)
    // Define execution flow with conditional branching
    .addEdge(START, "decideStrategy") // Start with strategy decision
    .addConditionalEdges(
      "decideStrategy",
      (state: typeof VCState.State) =&gt; state.searchStrategy, // Conditional function
      {
        investment_focused: "prepareInvestment", // If investment focused -&gt; RRF template preparation
        market_focused: "prepareMarket", // If market focused -&gt; dynamic query preparation
      }
    )
    .addEdge("prepareInvestment", "executeSearch") // Investment prep -&gt; execute
    .addEdge("prepareMarket", "executeSearch") // Market prep -&gt; execute
    .addEdge("executeSearch", "visualizeResults") // Execute -&gt; visualize
    .addEdge("visualizeResults", END); // End workflow


  const app = workflow.compile();

  await saveGraphImage(app);

  const query =
    "Find startups with Series A or Series B funding between $8M-$25M and monthly revenue above $500K";

  const marketResult = await app.invoke({ input: query });
  console.log(marketResult.final);
}<p>クエリ変数は、仮想の検索バーに入力されたユーザー入力をシミュレートします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltba7189d5f4e63403/6a1709880e2e49cc3041a076/e8d76909eb2bc1bb62f3ca9a8b3e4b85fcec2893-1600x164.png" alt="" /><p>「シリーズAまたはシリーズBの資金調達額が800万ドル～2,500万ドルで、月間収益が50万ドルを超えるスタートアップを探してください。」という自然言語フレーズから、すべてのフィルターが抽出されます。</p><p>最後にmainメソッドを呼び出します。</p>main().catch(console.error);<h3>成果</h3>🔍 Checking if index exists...
🏗️ Creating index...
✅ Index created successfully!
Ingesting documents...
✅ Documents ingested successfully!
✅ Investment-focused template created successfully!
✅ Market-focused template created successfully!

📊 Workflow graph saved as: ./workflow_graph.png

🔍 Query: "Find startups with Series A or Series B funding between $8M-$25M and monthly revenue above $500K"

🤔 Search strategy: investment_focused - The query specifically seeks profitable fintech startups with defined funding amounts and high monthly revenue, which aligns closely with financial performance metrics and investment-related criteria.

💰 Preparing INVESTMENT-FOCUSED search parameters with financial emphasis...

📋 Complete query: {
  "size": 5,
  "retriever": {
    "rrf": {
      "retrievers": [
        {
          "standard": {
            "query": {
              "semantic": {
                "field": "semantic_field",
                "query": "Find startups with Series A or Series B funding between $8M-$25M and monthly revenue above $500K"
              }
            }
          }
        },
        {
          "standard": {
            "query": {
              "bool": {
                "filter": [
                  {
                    "terms": {
                      "funding_stage": [
                        "Series A",
                        "Series B"
                      ]
                    }
                  },
                  {
                    "range": {
                      "funding_amount": {
                        "gte": 8000000,
                        "lte": 25000000
                      }
                    }
                  },
                  {
                    "terms": {
                      "lead_investor": []
                    }
                  },
                  {
                    "range": {
                      "monthly_revenue": {
                        "gte": 500000,
                        "lte": 0
                      }
                    }
                  }
                ]
              }
            }
          }
        }
      ],
      "rank_window_size": 100,
      "rank_constant": 20
    }
  }
}
🎯 Found 5 startups matching your criteria:

1. **TechFlow**
   📍 San Francisco, CA | 🏢 logistics | 💼 B2B
   💰 Series A - $8.0M
   👥 45 employees | 📈 $500K MRR
   🏦 Lead: Sequoia Capital
   📝 TechFlow optimizes supply chain operations using AI-powered route optimization and real-time tracking. Founded in 2023, shows remarkable growth with $500K monthly revenue.

2. **DataViz**
   📍 New York, NY | 🏢 enterprise software | 💼 B2B
   💰 Series A - $10.0M
   👥 42 employees | 📈 $450K MRR
   🏦 Lead: Battery Ventures
   📝 DataViz creates intuitive data visualization tools for enterprise customers. No-code platform allows business users to create dashboards without technical expertise.

3. **FinanceAI**
   📍 San Francisco, CA | 🏢 fintech | 💼 B2C
   💰 Series C - $25.0M
   👥 120 employees | 📈 $1200K MRR
   🏦 Lead: Tiger Global Management
   📝 FinanceAI provides AI-powered investment advisory services to retail investors. Uses machine learning to analyze market trends with over 100,000 active users.

4. **UrbanMobility**
   📍 New York, NY | 🏢 logistics | 💼 B2B2C
   💰 Series B - $15.0M
   👥 78 employees | 📈 $750K MRR
   🏦 Lead: Kleiner Perkins
   📝 UrbanMobility revolutionizes urban transportation through autonomous delivery drones and smart logistics hubs. Partners with major retailers for same-day delivery across Manhattan and Brooklyn.

5. **HealthTech Solutions**
   📍 Boston, MA | 🏢 healthcare | 💼 B2B
   💰 Series B - $18.0M
   👥 95 employees | 📈 $900K MRR
   🏦 Lead: General Catalyst
   📝 HealthTech Solutions develops medical devices and software for remote patient monitoring. Comprehensive telehealth platform reducing hospital readmissions by 30%.

✨  Done in 18.80s.<p>送信された入力に対して、アプリケーションは<strong>投資に重点を置いた</strong>パスを選択し、その結果、ユーザー入力から値と範囲を抽出するLangGraphワークフローによって生成されたElasticsearchクエリを確認できます。また、抽出された値が適用された状態でElasticsearchに送信されたクエリと、最後に<code>visualizeResults</code>ノードによって結果がフォーマットされた結果も確認できます。</p><p>次に、<strong>市場重視</strong>のノードを、クエリ「サンフランシスコ、ニューヨーク、ボストンのフィンテックおよびヘルスケアのスタートアップ企業を探してください」を使用してテストしてみましょう。</p>...

🔍 Query: Find fintech and healthcare startups in San Francisco, New York, or Boston

🤔 Search strategy: market_focused - The query is focused on finding fintech startups in San Francisco that are disrupting traditional banking and payment systems, which pertains to specific industries (fintech) and locations (San Francisco). Thus, a market-focused strategy is more appropriate.

🔍 Preparing MARKET-FOCUSED search parameters with market emphasis...

📋 Complete query: {
  "size": 5,
  "retriever": {
    "rrf": {
      "retrievers": [
        {
          "standard": {
            "query": {
              "semantic": {
                "field": "semantic_field",
                "query": "Find fintech and healthcare startups in San Francisco, New York, or Boston"
              }
            }
          }
        },
        {
          "standard": {
            "query": {
              "bool": {
                "filter": [
                  {
                    "terms": {
                      "industry": [
                        "fintech",
                        "healthcare"
                      ]
                    }
                  },
                  {
                    "terms": {
                      "location": [
                        "San Francisco, CA",
                        "New York, NY",
                        "Boston, MA"
                      ]
                    }
                  },
                  {
                    "terms": {
                      "business_model": []
                    }
                  }
                ]
              }
            }
          }
        }
      ],
      "rank_window_size": 50,
      "rank_constant": 10
    }
  }
}
🎯 Found 5 startups matching your criteria:

1. **FinanceAI**
   📍 San Francisco, CA | 🏢 fintech | 💼 B2C
   💰 Series C - $25.0M
   👥 120 employees | 📈 $1200K MRR
   🏦 Lead: Tiger Global Management
   📝 FinanceAI provides AI-powered investment advisory services to retail investors. Uses machine learning to analyze market trends with over 100,000 active users.

2. **CryptoWallet**
   📍 Miami, FL | 🏢 fintech | 💼 B2C
   💰 Series B - $16.0M
   👥 73 employees | 📈 $820K MRR
   🏦 Lead: Coinbase Ventures
   📝 CryptoWallet provides secure digital wallet solutions for cryptocurrency trading and storage. Multi-chain support with enterprise-grade security features.

...

✨  Done in 7.41s.<h2>学び</h2><p>執筆の過程で次のことを学びました。</p><ul><li><p>LLMにフィルターの正確な値を表示する必要があります。そうしないと、ユーザーが正確な値を入力することになります。カーディナリティが低い場合はこのアプローチで問題ありませんが、カーディナリティが高い場合は結果をフィルタリングする何らかのメカニズムが必要です。</p></li><li><p>検索テンプレートを使用すると、LLMにElasticsearchクエリを記述させるよりも結果の一貫性が大幅に向上し、速度も速くなります。</p></li><li><p>条件付きエッジは、複数のバリアントと分岐パスを持つアプリケーションを構築するための強力なメカニズムです。</p></li><li><p>構造化された出力は、予測可能でタイプセーフな応答を強制するため、LLMを使用して情報を生成する場合に非常に役立ちます。これにより、信頼性が向上し、プロンプトの誤解が減少します。</p></li></ul><p>ハイブリッド検索を通じてセマンティック検索と構造化検索を組み合わせることで、精度とコンテキスト理解のバランスを保ちながら、より適切で関連性の高い結果が生成されます。</p><h2>まとめ</h2><p>この例では、LangGraph.jsとElasticsearchを組み合わせて、自然言語クエリをElasticsearchで検索し、金融と市場のいずれかワークフローを焦点を当てた検索戦略をワークフローで決定できる動的なワークフローを作成します。このアプローチにより、手動クエリ作成の複雑さが軽減され、ベンチャーキャピタルアナリストの柔軟性と精度が向上します。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-agent-workflow-finance-langgraph-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-agent-workflow-finance-langgraph-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[エージェント型AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt013eba5d152f11f3/6a1709892b835f6784f4b1a6/12b6057d84c6356267cd178a3c6c1a5c61123ece-2000x1256.png" length="0" type="image/png"/>
    <pubDate>Fri, 05 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic Agent Builder と GPT-OSS を使用した HR 向け AI エージェントの構築]]></title>
    <description><![CDATA[Elastic Agent Builder と GPT-OSS を使用して、従業員の HR データに関する自然言語クエリに回答できる AI エージェントを構築する方法を学びます。]]></description>
    <content:encoded><![CDATA[<h2>はじめに</h2><p>この記事では<a href="https://openai.com/index/introducing-gpt-oss/">、GPT-OSS</a>と Elastic Agent Builder を使用して HR 向けの AI エージェントを構築する方法を説明します。エージェントは、OpenAI、Anthropic、その他の外部サービスにデータを送信せずに質問に答えることができます。</p><p>LM Studio を使用して GPT-OSS をローカルで提供し、Elastic Agent Builder に接続します。</p><p>この記事を読み終える頃には、情報とモデルを完全に制御しながら、従業員データに関する自然言語の質問に答えることができるカスタム AI エージェントが完成しているはずです。</p><h2>要件</h2><p>この記事には以下が必要です:</p><ul><li><p><a href="https://www.elastic.co/cloud">Elastic Cloud</a>ホスト 9.2、サーバーレスまたは<a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">ローカル</a>展開</p></li><li><p>32GB RAM搭載マシンを推奨（GPT-OSS 20Bの場合は最低16GB）</p></li><li><p><a href="https://lmstudio.ai/">LM Studio</a>がインストール済み</p></li><li><p><a href="https://www.docker.com/products/docker-desktop/">Dockerデスクトップ</a>がインストール済み</p></li></ul><h2>GPT-OSS を使用する理由は何ですか?</h2><p>ローカル LLM を使用すると、独自のインフラストラクチャに LLM を展開し、独自のニーズに合わせて微調整することができます。モデルと共有するデータの制御を維持しながら、これらすべてを実行できます。もちろん、外部プロバイダーにライセンス料を支払う必要はありません。</p><p>OpenAI は、オープン モデル エコシステムへの取り組みの一環として、2025 年 8 月 5 日に<a href="https://openai.com/index/introducing-gpt-oss/">GPT-OSS をリリースしました</a>。</p><p>20B パラメータ モデルは以下を提供します。</p><ul><li><p><strong>ツール使用能力</strong></p></li><li><p><strong>効率的な推論</strong></p></li><li><p><strong>OpenAI SDK対応</strong></p></li><li><p><strong>エージェントワークフローと互換性あり</strong></p></li></ul><p>ベンチマーク比較:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt58fab956edb40412/6a170cfcb0367da43a72bd80/29160e3345352088e8213297630882f252b00c47-1600x680.png" alt="" /><h2>ソリューションアーキテクチャ</h2><p>アーキテクチャは完全にローカル マシン上で実行されます。Elastic (Docker で実行) は LM Studio を介してローカル LLM と直接通信し、Elastic Agent Builder はこの接続を使用して従業員データを照会できるカスタム AI エージェントを作成します。</p><p>詳細については、 こちらの<a href="https://www.elastic.co/docs/solutions/observability/connect-to-own-local-llm">ドキュメント</a>を参照してください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt80db5bb0a797f51b/6a170cfd0e2e492f2c41a16f/a4a886750ff25fa8bb7aefc7448161e52cf73ed3-1600x896.png" alt="" /><h2>HR向けAIエージェントの構築：手順</h2><p>実装は 5 つのステップに分けられます。</p><ol><li><p>ローカルモデルでLMスタジオを構成する</p></li><li><p>DockerでローカルElasticをデプロイする</p></li><li><p>ElasticでOpenAIコネクタを作成する</p></li><li><p>従業員データをElasticsearchにアップロードする</p></li><li><p>AIエージェントを構築してテストする</p></li></ol><h2>ステップ1：LM StudioをGPT-OSS 20Bで構成する</h2><p>LM Studio は、大規模な言語モデルをコンピュータ上でローカルに実行できるユーザーフレンドリーなアプリケーションです。OpenAI 互換の API サーバーを提供するため、複雑なセットアップ プロセスなしで Elastic などのツールと簡単に統合できます。詳細については、 <a href="https://lmstudio.ai/docs/app">LM Studio ドキュメント</a>を参照してください。</p><p>まず、公式サイトからLM Studioをダウンロードしてインストールします。インストールしたら、アプリケーションを開きます。</p><h3>LM Studio インターフェースの場合:</h3><ol><li><p>検索タブに移動して「GPT-OSS」を検索します。</p></li><li><p>OpenAIから<code>openai/gpt-oss-20b</code>を選択してください</p></li><li><p>ダウンロードをクリック</p></li></ol><p>このモデルのサイズは約<strong>12.10 GB</strong>になります。インターネット接続によっては、ダウンロードに数分かかる場合があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2dc341a6625e34b7/6a170cff839dfa2eb4dcff44/5d01bc4dcb377b5259fc6b521fe2425a31b90ca4-1312x872.png" alt="" /><h4>モデルをダウンロードしたら:</h4><ol><li><p>ローカルサーバータブに移動します</p></li><li><p>openai/gpt-oss-20bを選択します</p></li><li><p>デフォルトのポート1234を使用する</p></li><li><p>右側のパネルで、 <strong>「ロード」</strong>に移動し、コンテキストの長さを<strong>40K</strong>以上に設定します。</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3704ca1b28465cc4/6a170d00d7c022ed8fde64ef/e546033f916381647b876815b2c1f1ae2a08365f-326x337.png" alt="" /><p>5. サーバーの開始をクリック</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7b9170a4945ff857/6a170d0266c4f9ffadf8c0a6/28ee78a3caa84d14e04db3d42f30acbe4d4d005a-1312x872.png" alt="" /><p>サーバーが実行中の場合はこれが表示されます。</p>[LM STUDIO SERVER] Success! HTTP server listening on port 1234
[LM STUDIO SERVER] Supported endpoints:
[LM STUDIO SERVER] -&gt;	GET  http://localhost:1234/v1/models
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/responses
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/chat/completions
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/completions
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/embeddings
Server started.<h2>ステップ2: DockerでローカルElasticをデプロイする</h2><p>ここで、Docker を使用して Elasticsearch と Kibana をローカルにセットアップします。Elastic は、セットアッププロセス全体を処理する便利なスクリプトを提供します。詳細については、<a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">公式ドキュメント</a>を参照してください。</p><h3>start-local スクリプトを実行する</h3><p>ターミナルで次のコマンドを実行します。</p>curl -fsSL https://elastic.co/start-local | sh<p>このスクリプトは次のことを行います。</p><ul><li><p>ElasticsearchとKibanaをダウンロードして設定する</p></li><li><p>Docker Composeを使用して両方のサービスを開始します</p></li><li><p>30日間のプラチナトライアルライセンスを自動的に有効化</p></li></ul><h3>期待される出力</h3><p>次のメッセージが表示されるまで待ち、表示されるパスワードと API キーを保存します。これらは Kibana にアクセスするために必要になります。</p>🎉 Congrats, Elasticsearch and Kibana are installed and running in Docker!
🌐 Open your browser at http://localhost:5601
   Username: elastic
   Password: KSUlOMNr
🔌 Elasticsearch API endpoint: http://localhost:9200
🔑 API key: cnJGX0pwb0JhOG00cmNJVklUNXg6cnNJdXZWMnM4bncwMllpQlFlUTlWdw==
Learn more at https://github.com/elastic/start-local<h3>Kibanaにアクセスする</h3><p>ブラウザを開いて次の場所に移動します:</p>http://localhost:5601<p>ターミナル出力で取得した資格情報を使用してログインします。</p><h3>エージェントビルダーを有効にする</h3><p>Kibana にログインしたら、 <strong>[Management]</strong> &gt; <strong>[AI]</strong> &gt; <strong>[Agent Builder]</strong>に移動して、Agent Builder をアクティブ化します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0a934bd99fa6a0ce/6a170d046234e019c3db1a5a/92e104cb846c20d875865ded8a3d37f5c7daae9b-1491x1528.png" alt="" /><h2>ステップ3: ElasticでOpenAIコネクタを作成する</h2><p>ここで、ローカル LLM を使用するように Elastic を構成します。</p><h3>アクセスコネクタ</h3><ol><li><p>キバナで</p></li><li><p><strong>プロジェクト設定</strong>&gt;<strong>管理</strong>に移動します</p></li><li><p><strong>アラートとインサイトの</strong>下で、<strong>コネクタ</strong>を選択します。</p></li><li><p>コネクタの作成をクリック</p></li></ol><h3>コネクタを構成する</h3><p>コネクタのリストから<strong>OpenAI を</strong>選択します。LM Studio は OpenAI SDK を使用しているため、互換性があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt762023c39781eb78/6a170d06a29299a59ed01087/5ac87042e086c7a2bd47a8039e646ec831f0dcc6-923x974.png" alt="" /><p>次の値をフィールドに入力します。</p><ul><li><p><strong>コネクタ名:</strong> LM Studio - GPT-OSS 20B</p></li><li><p><strong>OpenAIプロバイダーを選択:</strong>その他 (OpenAI互換サービス)</p></li><li><p><strong>URL: </strong><code>http://host.docker.internal:1234/v1/chat/completions</code></p></li><li><p><strong>デフォルトモデル:</strong> openai/gpt-oss-20b</p></li><li><p><strong>API キー:</strong> testkey-123 (LM Studio Server では認証が不要なので、任意のテキストを使用できます。)</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt980e595f80e2be2e/6a170d086f7f0468a19148cc/2084ac32fcf1fb810c8b54ecab1c85a1e3e8905b-672x1302.png" alt="" /><p>設定を完了するには、 <strong>「保存してテスト」</strong>をクリックします。</p><p><strong>重要:</strong> 「<strong>ネイティブ関数の呼び出しを有効にする</strong>」をオンにします。これは、Agent Builder が正しく動作するために必要です。これを有効にしないと、 <strong><code>No tool calls found in the response</code></strong>エラーが発生します。</p><h3>接続をテストする</h3><p>Elastic は自動的に接続をテストするはずです。すべてが正しく構成されている場合、次のような成功メッセージが表示されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d2e815dd558f881/6a170d090e2e49076541a177/f567d767f1969c4730c1daa92f651789dc3742ac-1042x812.png" alt="" /><p>対応：</p>{
  "status": "ok",
  "data": {
    "id": "chatcmpl-flj9h0hy4wcx4bfson00an",
    "object": "chat.completion",
    "created": 1761189456,
    "model": "openai/gpt-oss-20b",
    "choices": [
      {
        "index": 0,
        "message": {
          "role": "assistant",
          "content": "Hello! 👋 How can I assist you today?",
          "reasoning": "Just greet.",
          "tool_calls": []
        },
        "logprobs": null,
        "finish_reason": "stop"
      }
    ],
    "usage": {
      "prompt_tokens": 69,
      "completion_tokens": 23,
      "total_tokens": 92
    },
    "stats": {},
    "system_fingerprint": "openai/gpt-oss-20b"
  },
  "actionId": "ee1c3aaf-bad0-4ada-8149-118f52dad757"
}<h2>ステップ4: 従業員データをElasticsearchにアップロードする</h2><p>ここで、 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/gpt-oss-with-elasticsearch/hr-employees-bulk.json">HR 従業員データセット</a>をアップロードして、エージェントが機密データをどのように処理するかを説明します。私はこの構造を持つ架空のデータセットを生成しました。</p><h3>データセットの構造</h3>{
  "employee_id": "0f4dce68-2a09-4cb1-b2af-6bcb4821539b",
  "full_name": "Daffi Stiebler",
  "email": "lscutchings0@huffingtonpost.com",
  "date_of_birth": "1975-06-20T15:39:36Z",
  "hire_date": "2025-07-28T00:10:45Z",
  "job_title": "Physical Therapy Assistant",
  "department": "HR",
  "salary": "108455",
  "performance_rating": "Needs Improvement",
  "years_of_experience": 2,
  "skills": "Java",
  "education_level": "Master's Degree",
  "manager": "Carl MacGibbon",
  "emergency_contact": "Leigha Scutchings",
  "home_address": "5571 6th Park"
}<h3>マッピングを使用してインデックスを作成する</h3><p>まず、適切なマッピングを使用してインデックスを作成します。一部のキー フィールドに<a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">semantic_text</a>フィールドを使用していることに注意してください。これにより、インデックスのセマンティック検索機能が有効になります。</p>​​PUT hr-employees
{
  "mappings": {
    "properties": {
      "@timestamp": {
        "type": "date"
      },
      "employee_id": {
        "type": "keyword"
      },
      "full_name": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "email": {
        "type": "keyword"
      },
      "date_of_birth": {
        "type": "date",
        "format": "iso8601"
      },
      "hire_date": {
        "type": "date",
        "format": "iso8601"
      },
      "job_title": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "department": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "salary": {
        "type": "double"
      },
      "performance_rating": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "years_of_experience": {
        "type": "long"
      },
      "skills": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "education_level": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "manager": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "emergency_contact": {
        "type": "keyword"
      },
      "home_address": {
        "type": "keyword"
      },
      "employee_semantic": {
        "type": "semantic_text"
      }
    }
  }
}<h3>Bulk APIを使用したインデックス</h3><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/gpt-oss-with-elasticsearch/hr-employees-bulk.json">データセット</a>をコピーして Kibana の開発ツールに貼り付け、実行します。</p>POST hr-employees/_bulk
{"index": {}}
{"employee_id": "57728b91-e5d7-4fa8-954a-2384040d3886", "full_name": "Filide Gane", "email": "vhallahan1@booking.com", "job_title": "Business Systems Development Analyst", "department": "Marketing", "salary": "$52330.27", "performance_rating": "Meets Expectations", "years_of_experience": 12, "skills": "Java", "education_level": "Bachelor's Degree", "date_of_birth": "2000-02-07T16:49:32Z", "hire_date": "2023-11-07T13:03:16Z", "manager": "Freedman Kings", "emergency_contact": "Vilhelmina Hallahan", "home_address": "75 Dennis Junction"}
{"index": {}}
{"employee_id": "...", ...}<h3>データを検証する</h3><p>クエリを実行して確認します。</p>GET hr-employees/_search<h2>ステップ5: AIエージェントを構築してテストする</h2><p>すべての設定が完了したら、Elastic Agent Builder を使用してカスタム AI エージェントを構築します。詳細については、 <a href="https://www.elastic.co/docs/solutions/search/agent-builder/get-started">Elastic のドキュメント</a>を参照してください。</p><h3>コネクタを追加する</h3><p>新しいエージェントを作成する前に、デフォルトのコネクタは<a href="https://www.elastic.co/docs/reference/kibana/connectors-kibana/elastic-managed-llm">Elastic Managed LLM</a>であるため、 <code>LM Studio - GPT-OSS 20B</code>というカスタム コネクタを使用するようにエージェント ビルダーを設定する必要があります。そのためには、 <strong>「プロジェクト設定」</strong> &gt; <strong>「管理」</strong> &gt; <strong>「GenAI 設定」</strong>に移動し、作成した設定を選択して<strong>「保存」</strong>をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc42f079c5e756057/6a170d0acf4f2501d9b2d1c7/11e830c3e2fb4c298b020c928fa5422f3397ba08-1600x1152.png" alt="" /><h3>アクセスエージェントビルダー</h3><ol><li><p><strong>エージェント</strong>へ</p></li><li><p><strong>「新しいエージェントを作成」</strong>をクリックします</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb8e734817c5a7c6a/6a170d0ca929cf867cae0a34/c1e60541563650163f972ac9088dc1ed1de759a7-1600x1054.png" alt="" /><h3>エージェントを構成する</h3><p>新しいエージェントを作成するには、<strong>エージェント ID</strong> 、<strong>表示名</strong>、および<strong>表示手順</strong>が必須フィールドです。</p><p>ただし、システム プロンプトに似ていますが、カスタム エージェント用の、エージェントの動作やツールとの対話方法をガイドするカスタム インストラクションなど、さらに多くのカスタマイズ オプションがあります。ラベルは、エージェント、アバターの色、アバター シンボルを整理するのに役立ちます。</p><p>データセットに基づいてエージェント用に選択したものは次のとおりです。

<strong>エージェントID:</strong> <code>hr_assistant</code></p><p><strong>カスタム指示:</strong></p>You are an HR Analytics Assistant that helps answer questions about employee data.
When responding to queries:
- Provide clear, concise answers
- Include relevant employee details (name, department, salary, skills)
- Format monetary values with currency symbols
- Be professional and maintain data confidentiality<p>
ラベル: <code>Human Resources</code>および <code>GPT-OSS</code></p><p>表示名： <code>HR Analytics Assistant</code></p><p>表示の説明:</p>A specialized AI assistant for Human Resources that helps analyze employee data, compensation, performance metrics, and talent management. Ask questions about employees, departments, salaries, or performance analytics.<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt23fb011e5b4f4d49/6a170d0e7d8d67f47a70e77f/f94bb2bf08497e5e756ca76b30a3a51f42927756-1424x1217.png" alt="" /><p>すべてのデータが入力されたら、新しいエージェントの<strong>「保存」</strong>をクリックします。</p><h3>エージェントをテストする</h3><p>従業員データについて自然言語で質問できるようになり、GPT-OSS 20B が意図を理解して適切な応答を生成します。</p><h4>プロンプト：</h4>Which employee is the one with the highest salary in the hr-employees index?<h4>答え：</h4><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc0c52faacf63b583/6a170d0f0e2e497bfd41a17b/94ad19f80b96304028a59f60beca51dfc9aecc8a-899x631.png" alt="" /><p>エージェントのプロセスは次のとおりです。</p><p>1. GPT-OSSコネクタを使用して質問を理解する</p><p>2. 適切なElasticsearchクエリを生成する（組み込みツールまたはカスタム<a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a>を使用）</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte32a8a7e6363c7f2/6a170d115091680077e1bb44/6f2961d0d1b97475f6dda300acee84da540938e6-844x466.png" alt="" /><p>3. 一致する従業員レコードを取得する</p><p>4. 適切なフォーマットで自然言語で結果を提示する</p><p>従来の語彙検索とは異なり、GPT-OSS を搭載したエージェントは意図とコンテキストを理解するため、正確なフィールド名やクエリ構文を知らなくても情報を簡単に見つけることができます。エージェントの思考プロセスの詳細については、こちらの<a href="https://www.elastic.co/search-labs/blog/ai-agent-builder-experiments-performance">記事</a>を参照してください。</p><h2>まとめ</h2><p>この記事では、Elastic の Agent Builder を使用してカスタム AI エージェントを構築し、ローカルで実行されている OpenAI GPT-OSS モデルに接続しました。このアーキテクチャでは、Elastic と LLM の両方をローカルマシンにデプロイすることで、外部サービスに情報を送信することなく、データに対する完全な制御を維持しながら生成 AI 機能を活用できます。</p><p>実験としてはGPT-OSS 20Bを使用しましたが、Elastic Agent Builderの公式推奨モデルは<a href="https://www.elastic.co/docs/solutions/search/agent-builder/models#recommended-models">こちらを</a>参考にしています。より高度な推論機能が必要な場合は、複雑なシナリオでより優れたパフォーマンスを発揮する<a href="https://huggingface.co/openai/gpt-oss-120b">120B パラメータ バリアント</a>もありますが、ローカルで実行するにはより高性能なマシンが必要です。詳細については、 <a href="https://openai.com/open-models/">OpenAI の公式ドキュメント</a>を参照してください。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/build-an-ai-agent-hr-elastic-agent-builder-gpt-oss</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/build-an-ai-agent-hr-elastic-agent-builder-gpt-oss</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Tomás Murúa]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt664f490053e46e6b/6a170d13b0367d2d7e72bd84/05d2d0513fff67d975f9223d75108aa9f50646bc-1600x914.png" length="0" type="image/png"/>
    <pubDate>Wed, 26 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Cal Hacks 12.0 で取り上げた Elastic Agent Builder のトッププロジェクトと学習内容]]></title>
    <description><![CDATA[Cal Hacks 12.0 のトップ Elastic Agent Builder プロジェクトを探索し、サーバーレス、ES|QL、エージェント アーキテクチャに関する技術的なポイントを詳しく調べます。]]></description>
    <content:encoded><![CDATA[<p>数週間前、私たちは、世界中から 2,000 人を超える参加者が集まる最大規模の対面ハッカソンの 1 つである<a href="https://cal-hacks-12-0.devpost.com/">Cal Hacks 12.0 を</a>スポンサーするという素晴らしい機会を得ました。Elastic Agent Builder on Serverless の最も優れた活用方法に専用の賞品トラックを設けましたが、反響は驚くほど大きかったです。わずか 36 時間で、山火事インテリジェンス ツールの構築から StackOverflow バリデーターまで、Agent Builder を独創的な方法で使用した 29 件の応募を受け取りました。</p><p>Cal Hacks 12.0 での経験は、印象的なプロジェクト以外にも、同様に貴重なものをもたらしてくれました。それは、初めて当社のスタックに遭遇した開発者からの、迅速でフィルターされていないフィードバックです。ハッカソンは、厳しい期限、事前の知識ゼロ、そして予測不可能な障害（悪名高い WiFi の停止など）を伴う、ユニークなプレッシャーテストです。開発者エクスペリエンスが優れている点と、まだ改善が必要な点が正確に明らかになります。開発者が LLM 主導のワークフローを通じて新しい方法で Elastic Stack を操作することが増えているため、これは現在さらに重要になっています。このブログ投稿では、参加者が Agent Builder を使用して構築したものと、そのプロセスで学んだことについてさらに詳しく説明します。</p><h2>受賞プロジェクト</h2><h3>1位: AgentOverflow</h3><p>LLM およびエージェント時代に合わせて再構築された Stack Overflow。</p><p>AgentOverflow の詳細については、<a href="https://devpost.com/software/agentoverflow">こちらを</a>ご覧ください。</p><p>AgentOverflow は、ほとんどの AI 開発者が遭遇する問題、つまり LLM が幻覚を起こし、チャット履歴が消え、開発者が同じ問題を再度解決するのに時間を無駄にする問題に対処します。</p><p>AgentOverflow は実際の問題と解決策のペアをキャプチャ、検証、再表示するため、開発者は幻覚スパイラルを打破し、より早く製品を出荷できます。</p><h4>仕組み：</h4><p><strong>1. JSON（「ソリューション スキーマ」）を共有します。</strong></p><p>Claude の共有から 1 回クリックすると、次の内容を含む構造化形式である Share Solution JSON がスクレイピング、抽出、組み立てられます。</p><ul><li><p>問題</p></li><li><p>コンテクスト</p></li><li><p>コード</p></li><li><p>タグ</p></li><li><p>検証済みの解決手順。</p></li></ul><p>バリデーター (LAVA) が構造をチェックして強制し、ユーザーが追加のコンテキストの行を追加すると、Elasticsearch 内に保存されてインデックスが作成されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte7bc35b6d54921e8/6a17f0176df73162760a0fe6/45a3e96f4474050a855419628c2a7338bb12c706-1600x877.png" alt="「ソリューションを共有」をクリックすると、現在のセッションと関連するメタデータがスクレイピングされます。" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9967f52007fff99e/6a17f019ec0f8987c45a6701/2d65cb154d8ee32fc96ff17dfa5b0bf2636e3777-1600x1002.png" alt="ユーザーはWebフロントエンドを通じて追加のコンテキストを提供し、JSONはElasticsearchでインデックス化されます。" /><p><strong>2. 解決策を見つける</strong></p><p>行き詰まったら、 <code>Find Solution</code>をクリックすると、AgentOverflow が現在の会話をスクレイピングし、それを使用してクエリを作成し、ハイブリッド Elasticsearch 検索を実行して次の内容を表示します。</p><ul><li><p>ランク付けされたコミュニティ検証済みの修正</p></li><li><p>当初問題を解決した正確なプロンプト</p></li></ul><p>これにより、開発者は現在のセッションをすばやくコピー、貼り付け、ブロック解除できます。</p><p><strong>3. MCP - LLMのコンテキスト注入</strong></p><p>MCP (モデル コンテキスト プロトコル) を介して Elasticsearch 内に保存された構造化ソリューションに接続することにより、LLM には実行時に余分なノイズなしで高度な信号コンテキスト (コード、ログ、構成、以前の修正) が供給されます。</p><p>AgentOverflow は、関連するコンテキストを LLM に挿入する構造化メモリ レイヤーとして、Elasticsearch を備えた Agent Builder を使用します。これにより、受動的なチャットボットからコンテキストを認識した問題解決者へと変化します。</p><h3>準優勝：マーケットマインド</h3><p>6 つの Elastic Agent を活用した、市場エネルギーのリアルタイムの解釈可能なビュー。</p><p>MarketMindの詳細については、<a href="https://devpost.com/software/marketmind-b6cy2q">こちらを</a>ご覧ください。</p><p>MarketMind は、初心者トレーダーに、断片化された市場データを明確でリアルタイムなシグナルに変換するプラットフォームを提供することで、その地位を獲得しました。MarketMind は、さまざまなツール間で価格変動、ファンダメンタルズ、センチメント、ボラティリティを調整する代わりに、これらすべての情報を 1 つのプラットフォームに統合し、トレーダーが実用的な洞察を得られるよう支援します。このプロジェクトでは、エージェントの構築時に複雑な ES|QL クエリも使用しました。</p><h4>仕組み：</h4><p><strong>1. リアルタイムの市場データを収集する</strong></p><p>MarketMind は、Yahoo Finance から価格動向、ファンダメンタルズ、センチメント、ボラティリティ、リスク指標を取得します。このデータは複数の Elasticsearch インデックスに取り込まれ、整理されます。</p><p><strong>2. 6人の専門エージェントが市場を分析</strong></p><p>Agent Builder で構築された各エージェントは、市場の異なる層に焦点を当てています。これらは Elasticsearch インデックスから読み取り、独自のドメイン固有のメトリックを計算し、スコアと推論を含む標準化された JSON 出力を生成します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd4ba582f9872b65b/6a17f01b7f6f15c2d8c09c1c/7d9716cca06a047a2b3584378b5c7e592a785ba1-1284x878.png" alt="市場を分析する6つの専門的GOOGL AIエージェント" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd86ed3bfe4b8bd2b/6a17f01c5ea30f868164b6ba/5aac6a833347c0d2e596c02049ec4b4d3aae5cd7-794x764.png" alt="GOOGL特化エージェントのボリューム異常と大惨事検出分析機能" /><p><strong>3. シグナルを統合した「市場エネルギー」モデルに集約する</strong></p><p>組み合わせた出力は各株の周囲に光るパルスとして表示され、勢いが高まっているのか、リスクが高まっているのか、感情が変化しているのかを示します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5af7c7c838308275/6a17f01e42022917b629f6ca/46b3da8e3d528c5dd4e2829416c5446098acb3aa-744x718.png" alt="GOOGL特化エージェントの統一「市場エネルギー」モデル" /><p><strong>4. 洞察を視覚化する</strong></p><p>フロントエンドは、TypeScript、SVG 物理ベースのビジュアル、ライブ ローソク足チャート用の<a href="https://github.com/chartjs"> Chart.js</a> を使用して、React と<a href="https://github.com/vercel/next.js"> Next.js で構築されました。</a>これにより、生の分析がリアルタイムで実用的なフィードバックに変換されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt775e1880aa7afacc/6a17f01f1d1b83ce1f93e528/3f000c043117b77ed4127202be5a49c12e3682ba-1600x930.png" alt="GOOGL特化エージェント分析の洞察を視覚化する方法" /><h2>その他の興味深いプロジェクト:</h2><p>スタックのさまざまな部分で Elastic を使用した他の有力な候補を次に示します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltffe292009e446a70/6a17f0216df731068c0a0fea/76c49a853426844f475cd6b2a74999e60af20e8c-926x1080.png" alt="" /><p>私たちのトラックに提出されたプロジェクトの完全なリストは、<a href="https://cal-hacks-12-0.devpost.com/submissions/search?utf8=%E2%9C%93&amp;prize_filter%5Bprizes%5D%5B%5D=91882">こちらで</a>ご覧ください。</p><h2>開発者から学んだこと</h2><ul><li><p><strong>Agent Builder はユーザーフレンドリーです:</strong></p></li></ul><p>ほとんどのチームはこれまで Elastic を使用したことがありませんでしたが、それでもほとんどサポートなしでエージェントを迅速に構築できました。さらに詳しい指導が必要な人向けにワークショップを開催しましたが、ほとんどの人はデータを取り込み、そのデータに基づいてアクションを実行するエージェントを構築することができました。</p><ul><li><p>LLM は<strong><code>kNN</code></strong>クエリ<strong>に優れています</strong><strong>が、ES|QL の生成には依然としてガイダンスが必要です。</strong></p></li></ul><p>ChatGPT-5 に ES|QL クエリの生成を依頼すると、ES|QL と SQL が混在するなど、誤った情報が返されることがよくありました。LLM にマークダウン ファイルでドキュメントを供給することは、実行可能な修正であるように思われました。</p><ul><li><p><strong>スナップショット専用の ES|QL 関数がドキュメントに漏洩しました:</strong></p></li></ul><p>今後登場する<code>FIRST</code>および<code>LAST</code>集計関数が、意図せず ES|QL ドキュメントに紛れ込んでしまいました。これらのドキュメントを ChatGPT に渡したため、Serverless ではまだ利用できないにもかかわらず、モデルはこれらの関数を忠実に使用しました。グループからのフィードバックのおかげで、エンジニアリングはすぐに修正を公開し、マージして、公開されたドキュメントから関数を削除しました ( <a href="https://github.com/elastic/elasticsearch/pull/137341">PR #137341</a> )。</p><ul><li><p><strong>サーバーレス固有のガイダンスが不足しています:</strong></p></li></ul><p>チームは、ルックアップ モードで作成されなかったインデックスで<code>LOOKUP JOIN</code>有効にしようとしました。エラー メッセージにより、Serverless に存在しないコマンドが追跡されました。私たちはこれを製品チームに伝え、製品チームはすぐに Serverless 固有の実用的なメッセージの修正を開始しました。長期的には、再インデックスの複雑さを完全に隠すことがビジョンです (<a href="https://github.com/elastic/elasticsearch-serverless/issues/4838">問題 #4838</a> )。</p><ul><li><p><strong>対面イベントの価値:</strong></p></li></ul><p>オンライン ハッカソンは素晴らしいですが、ビルダーと肩を並べてデバッグしているときに得られる迅速なフィードバック ループに匹敵するものはありません。私たちは、チームがさまざまなユースケースにわたって Agent Builder を統合する様子を観察し、ES|QL を使用した開発者エクスペリエンスを改善できる部分を見つけ、非同期チャネルで解決するよりもはるかに迅速に問題を修正しました。</p><h2>まとめ</h2><p>Cal Hacks 12.0 では、素晴らしいデモを週末にわたって披露するだけでなく、新しい開発者が Elastic Stack とどのように関わっているかについても理解することができました。わずか 36 時間で、チームは Agent Builder を導入し、Elasticsearch にデータを取り込み、マルチエージェント システムを設計し、さまざまな方法で機能をテストするようになりました。このイベントは、対面イベントがなぜ重要なのかを私たちに思い出させてくれました。迅速なフィードバック ループ、実際の会話、実践的なデバッグにより、現在の開発者のニーズを理解することができました。私たちが学んだことをエンジニアリング チームに還元できることを嬉しく思います。次回のハッカソンでお会いしましょう。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-builder-projects-learnings-cal-hacks-12-0</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-builder-projects-learnings-cal-hacks-12-0</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[エージェント型AI]]></category>
    <dc:creator><![CDATA[JD Armada]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0f079179be9832d4/6a17f023631730a69c585b6d/8ba034a6f19b50521f541b8131756a8acdb52975-1280x960.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 25 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch で A2A プロトコルと MCP を使用して LLM エージェント ニュースルームを作成する: パート II]]></title>
    <description><![CDATA[エージェントのコラボレーションに A2A プロトコルを使用し、Elasticsearch でのツール アクセスに MCP を使用して、特殊なハイブリッド LLM エージェント ニュースルームを構築する方法を説明します。]]></description>
    <content:encoded><![CDATA[<h2>A2AとMCP：コードの動作</h2><p>これは、記事「Elasticsearch で A2A プロトコルと MCP を使用して LLM エージェント ニュースルームを作成する」の補足記事です。この記事では、同じエージェント内に A2A と MCP の両方のアーキテクチャを実装して、両方のフレームワークの独自のメリットを最大限に活用するメリットについて説明しました。自分でデモを実行したい場合、<a href="https://github.com/justincastilla/elastic-newsroom">リポジトリ</a>が利用可能です。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt232e466d2153c764/6a17f15f631730042d585b8d/7196f004089127f83547b2e5dc3f663205cfcdce-1162x1600.png" alt="A2A &amp; MCP プロトコルエージェントワークフロー" /><p>ニュースルームのエージェントが A2A と MCP の両方を使用して協力し、ニュース記事を作成する方法を見ていきましょう。エージェントの動作を確認するための付属リポジトリは、<a href="https://github.com/justincastilla/elastic-newsroom">ここに</a>あります。</p><h3>ステップ1：ストーリーの割り当て</h3><p><strong>ニュースチーフ</strong>（クライアントとして行動）がストーリーを割り当てます。</p>{
  "message_type": "task_request",
  "sender": "news_chief",
  "receiver": "reporter_agent",
  "payload": {
    "task_id": "story_renewable_energy_2024",
    "assignment": {
      "topic": "Renewable Energy Adoption in Europe",
      "angle": "Policy changes driving solar and wind expansion",
      "target_length": 1200,
      "deadline": "2025-09-30T18:00:00Z"
    }
  }
}<h3>ステップ2: 記者が調査を依頼する</h3><p><strong>レポーター エージェントは</strong>背景情報が必要であることを認識し、A2A を介して<strong>リサーチャー エージェント</strong>に委任します。</p>{
  "message_type": "task_request",
  "sender": "reporter_agent",
  "receiver": "researcher_agent",
  "payload": {
    "task_id": "research_eu_renewable_2024",
    "parent_task_id": "story_renewable_energy_2024",
    "capability": "fact_gathering",
    "parameters": {
      "queries": [
        "EU renewable energy capacity 2024",
        "Solar installations growth Europe",
        "Wind energy policy changes 2024"
      ],
      "depth": "comprehensive"
    }
  }
}<h3>ステップ3: 報告者がアーカイブエージェントに歴史的背景をリクエストする</h3><p><strong>レポーターエージェントは</strong>、歴史的背景が記事の内容を強めることを認識しています。A2A 経由で<strong>アーカイブエージェント</strong>( <a href="https://www.elastic.co/docs/solutions/search/elastic-agent-builder">Elastic の A2A エージェント</a>を搭載) に委任し、ニュースルームの Elasticsearch 搭載記事アーカイブを検索します。</p>{
  "message_type": "task_request",
  "sender": "reporter_agent",
  "receiver": "archive_agent",
  "payload": {
    "task_id": "archive_search_renewable_2024",
    "parent_task_id": "story_renewable_energy_2024",
    "capability": "search_archive",
    "parameters": {
      "query": "European renewable energy policy changes and adoption trends over past 5 years",
      "focus_areas": ["solar", "wind", "policy", "Germany", "France"],
      "time_range": "2019-2024",
      "result_count": 10
    }
  }
}<h3>ステップ4: アーカイブエージェントはMCPでElastic A2Aエージェントを使用する</h3><p><strong>アーカイブ エージェントは</strong>Elastic の A2A エージェントを使用し、A2A エージェントは MCP を使用して Elasticsearch ツールにアクセスします。これは、A2A がエージェントのコラボレーションを可能にし、MCP がツール アクセスを提供するハイブリッド アーキテクチャを示しています。</p># Archive Agent using Elastic A2A Agent
async def search_historical_articles(self, query_params):
    # The Archive Agent sends a request to Elastic's A2A Agent
    elastic_response = await self.a2a_client.send_request(
        agent="elastic_agent",
        capability="search_and_analyze",
        parameters={
            "natural_language_query": query_params["query"],
            "index_pattern": "newsroom-articles-*",
            "filters": {
                "topics": query_params["focus_areas"],
                "date_range": query_params["time_range"]
            },
            "analysis_type": "trend_analysis"
        }
    )
    
    # Elastic's A2A Agent internally uses MCP tools:
    # - platform.core.search (to find relevant articles)
    # - platform.core.generate_esql (to analyze trends)
    # - platform.core.index_explorer (to identify relevant indices)
    
    return elastic_response<p><strong>アーカイブエージェントは</strong>Elastic の A2A エージェントから包括的な履歴データを受信し、それをレポーターに返します。</p>{
  "message_type": "task_response",
  "sender": "archive_agent",
  "receiver": "reporter_agent",
  "payload": {
    "task_id": "archive_search_renewable_2024",
    "status": "completed",
    "archive_data": {
      "historical_articles": [
        {
          "title": "Germany's Energiewende: Five Years of Solar Growth",
          "published": "2022-06-15",
          "key_points": [
            "Germany added 7 GW annually 2020-2022",
            "Policy subsidies drove 60% of growth"
          ],
          "relevance_score": 0.94
        },
        {
          "title": "France Balances Nuclear and Renewables",
          "published": "2023-03-20",
          "key_points": [
            "France increased renewable target to 40% by 2030",
            "Solar capacity doubled 2021-2023"
          ],
          "relevance_score": 0.89
        }
      ],
      "trend_analysis": {
        "coverage_frequency": "EU renewable stories increased 150% since 2019",
        "emerging_themes": ["policy incentives", "grid modernization", "battery storage"],
        "coverage_gaps": ["Small member states", "offshore wind permitting"]
      },
      "total_articles_found": 47,
      "search_confidence": 0.91
    }
  }
}<p>このステップでは、Elastic の A2A エージェントがニュースルームのワークフローにどのように統合されるかを示します。Archive Agent（ニュースルーム固有のエージェント）は、Elastic の A2A Agent（サードパーティの専門家）と連携して、Elasticsearch の強力な検索および分析機能を活用します。Elastic のエージェントは内部的に MCP を使用して Elasticsearch ツールにアクセスし、エージェント調整 (A2A) とツール アクセス (MCP) を明確に分離します。</p><h3>ステップ5: 研究者はMCPサーバーを使用する</h3><p><strong>研究者エージェントは</strong>複数の MCP サーバーにアクセスして情報を収集します。</p># Researcher Agent using MCP to access tools
async def gather_facts(self, queries):
    results = []
    
    # Use News API MCP Server
    news_data = await self.mcp_client.invoke_tool(
        server="news_api_mcp",
        tool="search_articles",
        parameters={
            "query": queries[0],
            "date_range": "2024-01-01 to 2024-09-30",
            "sources": ["reuters", "bloomberg", "euractiv"]
        }
    )
    
    # Use Statistics MCP Server
    stats_data = await self.mcp_client.invoke_tool(
        server="eurostat_mcp",
        tool="get_energy_statistics",
        parameters={
            "metrics": ["solar_capacity_mw", "wind_capacity_mw"],
            "countries": ["EU27"],
            "year": 2024
        }
    )
    
    # Use Fact-Checking MCP Server
    verified_facts = await self.mcp_client.invoke_tool(
        server="factcheck_mcp",
        tool="verify_claims",
        parameters={
            "claims": [
                "EU solar capacity increased 40% in 2024",
                "Germany leads in wind installations"
            ]
        }
    )
    
    return self._synthesize_research(news_data, stats_data, verified_facts)<h3>ステップ6: 研究者が報告者にデータを返す</h3><p><strong>研究者エージェントは、</strong>包括的な研究を A2A 経由で返送します。</p>{
  "message_type": "task_response",
  "sender": "researcher_agent",
  "receiver": "reporter_agent",
  "payload": {
    "task_id": "research_eu_renewable_2024",
    "status": "completed",
    "research_data": {
      "key_facts": [
        {
          "claim": "EU solar capacity reached 260 GW in 2024",
          "source": "Eurostat Q3 2024 Report",
          "confidence": 0.95,
          "verified": true
        },
        {
          "claim": "Germany installed 12 GW of wind capacity in 2024",
          "source": "German Federal Network Agency",
          "confidence": 0.92,
          "verified": true
        }
      ],
      "statistics": {
        "solar_growth_rate": "35%",
        "wind_growth_rate": "28%"
      },
      "sources_count": 15
    }
  }
}<h3>ステップ7：記者が記事を書く</h3><p><strong>Reporter Agent は、</strong>調査データと独自の LLM 機能を使用して記事を作成します。書き込み中、Reporter はスタイルとテンプレートに MCP サーバーを使用します。</p># Reporter Agent writing with MCP assistance
async def write_article(self, research_data, assignment):
    # Get style guidelines via MCP
    style_guide = await self.mcp_client.get_resource(
        server="newsroom_mcp",
        resource="style://editorial/ap_style_guide"
    )
    
    # Get article template via MCP
    template = await self.mcp_client.get_resource(
        server="newsroom_mcp",
        resource="template://articles/news_story"
    )
    
    # Generate article using LLM + research + style
    draft = await self.llm.generate(
        prompt=f"""
        Write a news article following these guidelines:
        {style_guide}
        
        Using this template:
        {template}
        
        Based on this research:
        {research_data}
        
        Assignment: {assignment}
        """
    )
    
    # Self-evaluate confidence in claims
    confidence_check = await self._evaluate_confidence(draft)
    
    return draft, confidence_check<h3>ステップ8：自信が低い場合は再調査を促します</h3><p><strong>レポーター エージェントは</strong>下書きを評価し、1 つの主張の信頼性が低いことを発見しました。<strong>研究者エージェント</strong>に別のリクエストを送信します:</p>{
  "message_type": "collaboration_request",
  "sender": "reporter_agent",
  "receiver": "researcher_agent",
  "payload": {
    "request_type": "fact_verification",
    "claims": [
      {
        "text": "France's nuclear phase-down contributed to 15% increase in renewable capacity",
        "context": "Discussing policy drivers for renewable growth",
        "current_confidence": 0.45,
        "required_confidence": 0.80
      }
    ],
    "urgency": "high"
  }
}<p><strong>研究者は</strong>ファクトチェックMCPサーバーを使用して主張を検証し、更新された情報を返します。</p>{
  "message_type": "collaboration_response",
  "sender": "researcher_agent",
  "receiver": "reporter_agent",
  "payload": {
    "verified_claims": [
      {
        "original_claim": "France's nuclear phase-down contributed to 15% increase...",
        "verified_claim": "France's renewable capacity increased 18% in 2024, partially offsetting reduced nuclear output",
        "confidence": 0.88,
        "corrections": "Percentage was 18%, not 15%; nuclear phase-down is gradual, not primary driver",
        "sources": ["RTE France", "French Energy Ministry Report 2024"]
      }
    ]
  }
}<h3>ステップ9: 記者が修正して編集者に提出する</h3><p><strong>記者は</strong>検証された事実を組み込み、完成した原稿を A2A 経由で<strong>編集者エージェント</strong>に送信します。</p>{
  "message_type": "task_request",
  "sender": "reporter_agent",
  "receiver": "editor_agent",
  "payload": {
    "task_id": "edit_renewable_story",
    "parent_task_id": "story_renewable_energy_2024",
    "content": {
      "headline": "Europe's Renewable Revolution: Solar and Wind Surge 30% in 2024",
      "body": "[Full article text...]",
      "word_count": 1185,
      "sources": [/* array of sources */]
    },
    "editing_requirements": {
      "check_style": true,
      "check_facts": true,
      "check_seo": true
    }
  }
}<h3>ステップ10: MCPツールを使用した編集者のレビュー</h3><p><strong>エディター エージェントは</strong>複数の MCP サーバーを使用して記事をレビューします。</p># Editor Agent using MCP for quality checks
async def review_article(self, content):
    # Grammar and style check
    grammar_issues = await self.mcp_client.invoke_tool(
        server="grammarly_mcp",
        tool="check_document",
        parameters={"text": content["body"]}
    )
    
    # SEO optimization check
    seo_analysis = await self.mcp_client.invoke_tool(
        server="seo_mcp",
        tool="analyze_content",
        parameters={
            "headline": content["headline"],
            "body": content["body"],
            "target_keywords": ["renewable energy", "Europe", "solar", "wind"]
        }
    )
    
    # Plagiarism check
    originality = await self.mcp_client.invoke_tool(
        server="plagiarism_mcp",
        tool="check_originality",
        parameters={"text": content["body"]}
    )
    
    # Generate editorial feedback
    feedback = await self._generate_feedback(
        grammar_issues, 
        seo_analysis, 
        originality
    )
    
    return feedback<p><strong>編集者は</strong>記事を承認し、送信します。</p>{
  "message_type": "task_response",
  "sender": "editor_agent",
  "receiver": "reporter_agent",
  "payload": {
    "status": "approved",
    "quality_score": 9.2,
    "minor_edits": [
      "Changed 'surge' to 'increased' in paragraph 3 for AP style consistency",
      "Added Oxford comma in list of countries"
    ],
    "approved_content": "[Final edited article]"
  }
}<h3>ステップ11: パブリッシャーがCI/CD経由でパブリッシュする</h3><p>最後に、<strong>プリンター エージェントは</strong>、CMS および CI/CD パイプラインの MCP サーバーを使用して承認された記事を公開します。</p># Publisher Agent publishing via MCP
async def publish_article(self, content, metadata):
    # Upload to CMS via MCP
    cms_result = await self.mcp_client.invoke_tool(
        server="wordpress_mcp",
        tool="create_post",
        parameters={
            "title": content["headline"],
            "body": content["body"],
            "status": "draft",
            "categories": metadata["categories"],
            "tags": metadata["tags"],
            "featured_image_url": metadata["image_url"]
        }
    )
    
    post_id = cms_result["post_id"]
    
    # Trigger CI/CD deployment via MCP
    deploy_result = await self.mcp_client.invoke_tool(
        server="cicd_mcp",
        tool="trigger_deployment",
        parameters={
            "pipeline": "publish_article",
            "environment": "production",
            "post_id": post_id,
            "schedule": "immediate"
        }
    )
    
    # Track analytics
    await self.mcp_client.invoke_tool(
        server="analytics_mcp",
        tool="register_publication",
        parameters={
            "post_id": post_id,
            "publish_time": datetime.now().isoformat(),
            "story_id": metadata["story_id"]
        }
    )
    
    return {
        "status": "published",
        "post_id": post_id,
        "url": f"https://newsroom.example.com/articles/{post_id}",
        "deployment_id": deploy_result["deployment_id"]
    }<p><strong>出版社は</strong>A2Aを通じて出版を確認します。</p>{
  "message_type": "task_complete",
  "sender": "printer_agent",
  "receiver": "news_chief",
  "payload": {
    "task_id": "story_renewable_energy_2024",
    "status": "published",
    "publication": {
      "url": "https://newsroom.example.com/articles/renewable-europe-2024",
      "published_at": "2025-09-30T17:45:00Z",
      "post_id": "12345"
    },
    "workflow_metrics": {
      "total_time_minutes": 45,
      "agents_involved": ["reporter", "researcher", "archive", "editor", "printer"],
      "iterations": 2,
      "mcp_calls": 12
    }
  }
}<p>以下は、上記と同じエージェントを使用した付属のリポジトリ内の A2A ワークフローの完全なシーケンスです。</p><p>#</p><p>から</p><p>に</p><p>アクション</p><p>プロトコル</p><p>説明</p><p>1</p><p>ユーザー</p><p>ニュースチーフ</p><p>ストーリーの割り当て</p><p>HTTP ポスト</p><p>ユーザーがストーリーのトピックと角度を提出する</p><p>2</p><p>ニュースチーフ</p><p>内部</p><p>ストーリーを作成する</p><p>-</p><p>固有のIDを持つストーリーレコードを作成します</p><p>3</p><p>ニュースチーフ</p><p>記者</p><p>委任の割り当て</p><p>A2A</p><p>A2Aプロトコル経由でストーリー割り当てを送信します</p><p>4</p><p>記者</p><p>内部</p><p>割り当てを受け入れる</p><p>-</p><p>割り当てを内部に保存する</p><p>5</p><p>記者</p><p>MCP サーバー</p><p>アウトラインを生成</p><p>MCP/HTTP</p><p>記事のアウトラインと研究の質問を作成します</p><p>6a</p><p>記者</p><p>研究者</p><p>調査依頼</p><p>A2A</p><p>質問を送信します（6bと並行）</p><p>6b</p><p>記者</p><p>アーキビスト</p><p>アーカイブを検索</p><p>A2A JSONRPC</p><p>歴史的な記事を検索します（6aと並行）</p><p>7</p><p>研究者</p><p>MCP サーバー</p><p>研究上の質問</p><p>MCP/HTTP</p><p>MCP経由でAnthropicを使用して質問に答えます</p><p>8</p><p>研究者</p><p>記者</p><p>リターンリサーチ</p><p>A2A</p><p>調査の回答を返す</p><p>9</p><p>アーキビスト</p><p>Elasticsearch</p><p>検索インデックス</p><p>ES REST API</p><p>news_archiveインデックスをクエリ</p><p>10</p><p>アーキビスト</p><p>記者</p><p>アーカイブに戻る</p><p>A2A JSONRPC</p><p>過去の検索結果を返します</p><p>11</p><p>記者</p><p>MCP サーバー</p><p>記事を生成する</p><p>MCP/HTTP</p><p>研究/アーカイブの文脈で記事を作成する</p><p>12</p><p>記者</p><p>内部</p><p>ストアドラフト</p><p>-</p><p>下書きを内部に保存</p><p>13</p><p>記者</p><p>ニュースチーフ</p><p>下書きを送信</p><p>A2A</p><p>完成した草稿を提出する</p><p>14</p><p>ニュースチーフ</p><p>内部</p><p>ストーリーを更新</p><p>-</p><p>下書きを保存し、ステータスを「draft_submitted」に更新します</p><p>15</p><p>ニュースチーフ</p><p>エディタ</p><p>レビュー草稿</p><p>A2A</p><p>レビューのために編集者に自動ルーティング</p><p>16</p><p>エディタ</p><p>MCP サーバー</p><p>総説</p><p>MCP/HTTP</p><p>MCP経由でAnthropicを使用してコンテンツを分析します</p><p>17</p><p>エディタ</p><p>ニュースチーフ</p><p>返品レビュー</p><p>A2A</p><p>編集上のフィードバックと提案を送信します</p><p>18</p><p>ニュースチーフ</p><p>内部</p><p>ストアレビュー</p><p>-</p><p>編集者のフィードバックを保存</p><p>19</p><p>ニュースチーフ</p><p>記者</p><p>編集を適用</p><p>A2A</p><p>レビューのフィードバックをレポーターに転送する</p><p>20</p><p>記者</p><p>MCP サーバー</p><p>編集を適用</p><p>MCP/HTTP</p><p>フィードバックに基づいて記事を修正する</p><p>21</p><p>記者</p><p>内部</p><p>下書きの更新</p><p>-</p><p>修正を加えて下書きを更新する</p><p>22</p><p>記者</p><p>ニュースチーフ</p><p>返品修正</p><p>A2A</p><p>修正された記事を返す</p><p>23</p><p>ニュースチーフ</p><p>内部</p><p>ストーリーを更新</p><p>-</p><p>修正した下書きを保存し、ステータスを「修正済み」にする</p><p>24</p><p>ニュースチーフ</p><p>出版社</p><p>記事を公開する</p><p>A2A</p><p>パブリッシャーへの自動ルーティング</p><p>25</p><p>出版社</p><p>MCP サーバー</p><p>タグを生成する</p><p>MCP/HTTP</p><p>タグとカテゴリを作成する</p><p>26</p><p>出版社</p><p>Elasticsearch</p><p>インデックス記事</p><p>ES REST API</p><p>記事をnews_archiveインデックスにインデックスします</p><p>27</p><p>出版社</p><p>ファイルシステム</p><p>マークダウンを保存</p><p>ファイルI/O</p><p>記事を.mdとして保存します/articles内のファイル</p><p>28</p><p>出版社</p><p>ニュースチーフ</p><p>公開の確認</p><p>A2A</p><p>成功ステータスを返します</p><p>29</p><p>ニュースチーフ</p><p>内部</p><p>ストーリーを更新</p><p>-</p><p>ストーリーのステータスを「公開済み」に更新します</p><h2>まとめ</h2><p>A2A と MCP はどちらも、現代の拡張 LLM インフラストラクチャ パラダイムにおいて重要な役割を果たします。A2A は複雑なマルチエージェント システムに柔軟性を提供しますが、移植性が低くなり、運用が複雑になる可能性があります。MCP は、マルチエージェント オーケストレーションを処理するようには設計されていませんが、実装と保守がより簡単なツール統合のための標準化されたアプローチを提供します。</p><p>選択は二者択一ではありません。私たちのニュースルームの例で示されているように、最も洗練され効果的な LLM 対応システムは、多くの場合、両方のアプローチを組み合わせています。つまり、エージェントは A2A プロトコルを通じて調整と専門化を行いながら、MCP サーバーを通じてツールやリソースにアクセスします。このハイブリッド アーキテクチャは、MCP の標準化とエコシステムの利点に加えて、マルチエージェント システムの組織上の利点も提供します。これは、選択する必要が全くないかもしれないことを示唆している。単に両方を標準的なアプローチとして使うだけでよい。</p><p>開発者またはアーキテクトとして、両方のソリューションの最適な組み合わせをテストして決定し、特定のユースケースに適した結果を生み出すのはあなた次第です。それぞれのアプローチの長所、制限、適切な適用を理解することで、より効果的で保守性と拡張性に優れた AI システムを構築できるようになります。</p><p>デジタル ニュースルーム、顧客サービス プラットフォーム、リサーチ アシスタント、またはその他の LLM を利用したアプリケーションを構築する場合でも、調整ニーズ (A2A) とツール アクセス要件 (MCP) を慎重に検討することで、成功への道が開かれます。</p><h2>参考資料</h2><ul><li><p><strong>Elasticsearch エージェントビルダー:</strong> <a href="https://www.elastic.co/docs/solutions/search/elastic-agent-builder">https://www.elastic.co/docs/solutions/search/elastic-agent-builder</a></p></li><li><p><strong>A2A仕様</strong>: <a href="https://a2a-protocol.org/latest/specification/">https://a2a-protocol.org/latest/specification/</a></p></li><li><p><strong>A2A と MCP の統合</strong>: <a href="https://a2a-protocol.org/latest/topics/a2a-and-mcp/">https://a2a-protocol.org/latest/topics/a2a-and-mcp/</a></p></li><li><p><strong>モデルコンテキストプロトコル</strong>: <a href="https://modelcontextprotocol.io/">https://modelcontextprotocol.io</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/a2a-protocol-mcp-llm-agent-workflow-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/a2a-protocol-mcp-llm-agent-workflow-elasticsearch</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Justin Castilla]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1b1f22cdc2130333/6a17f161ec0f8917fa5a6712/f87330e5d4ca961593b3cfb861ca850a4cc34186-1519x1173.png" length="0" type="image/png"/>
    <pubDate>Mon, 24 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[コンテキストのためのYou Know - パートII：エージェントAIとコンテキストエンジニアリングの必要性]]></title>
    <description><![CDATA[LLM がエージェント AI へと進化するにつれ、RAG コンテキストの制限とメモリ管理を解決するためのコンテキスト エンジニアリングの必要性がどのように高まるのかを学びます。]]></description>
    <content:encoded><![CDATA[<p>LLM が情報検索の基本的なプロセスをどのように変えてきたかについての (かなり広範囲にわたる)<a href="https://www.elastic.co/search-labs/blog/context-engineering-hybrid-search-evolution-agentic-ai">背景</a>を踏まえて、LLM がデータのクエリ方法をどのように変えてきたかを見てみましょう。</p><h2>データと対話する新しい方法</h2><p>ジェネレーティブ (genAI) AI とエージェント AI は、従来の検索とは異なる処理を行います。かつて私たちが情報を調べ始める方法は検索（「グーグルで検索してみます…」）でしたが、gen AI とエージェントの両方にとって、開始アクションは通常、チャット インターフェースに入力された自然言語を通じて行われます。チャット インターフェースは、意味理解を使用して質問を簡潔な回答、つまりあらゆる種類の情報に関する幅広い知識を持つ予言者から出されたような要約された応答に変換する LLM とのディスカッションです。本当に売れているのは、LLM が表面化した知識の断片をつなぎ合わせて首尾一貫した思慮深い文章を生成する能力です。たとえそれが不正確であったり完全に幻覚的であったりしても、そこには<a href="https://en.wikipedia.org/wiki/Truthiness">真実味</a>があります。</p><p>私たちが使い慣れている古い検索バーは<em><strong>、私たち自身が</strong></em>推論エージェントであったときに使用した RAG エンジンと考えることができます。現在では、インターネット検索エンジンでさえ、使い古された「ハント・アンド・ペック」という語彙検索エクスペリエンスを、クエリに対する結果の要約で答える AI 主導の概要へと変えつつあり、ユーザーがクリックして個々の結果を自分で評価する必要がないようにしています。</p><h2>生成AIとRAG</h2><p>生成 AI は、世界の意味理解を活用してチャット リクエストを通じて表明された主観的な意図を解析し、推論能力を使用して専門的な回答を即座に作成します。生成 AI インタラクションにはいくつかの部分があります。ユーザーの入力/クエリから始まり、チャット セッションでの以前の会話が追加のコンテキストとして使用でき、LLM に推論方法と応答の構築手順を指示する指示プロンプトがあります。プロンプトは、「5 歳児に説明するように説明してください」という単純なタイプのガイダンスから、リクエストを処理する方法の完全な詳細へと進化しました。これらの内訳には、AI のペルソナ/役割、生成前の推論/内部思考プロセス、客観的な基準、制約、出力形式、対象者、および期待される結果を示すのに役立つ例の詳細を説明する個別のセクションが含まれることがよくあります。</p><p>ユーザーのクエリとシステム プロンプトに加えて、検索拡張生成 (RAG) は、「コンテキスト ウィンドウ」と呼ばれる追加のコンテキスト情報を提供します。RAG はアーキテクチャへの重要な追加機能であり、世界の意味理解において欠落している部分を LLM に通知するために使用します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbfa000ccfdd9d184/6a17ddb57b54f955f38b37da/5b9671d5d07d4caefde372bb3188000754a91eed-1470x746.png" alt="LLMがユーザークエリを処理してコンテキストを作成する方法" /><p>コンテキスト ウィンドウは、何を、どこに、どれだけ与えるかという点では、かなり<a href="https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html">細かい指定が</a>必要になる場合があります。もちろん、どのコンテキストが選択されるかは非常に重要ですが、提供されたコンテキストの信号対雑音比やウィンドウの長さも重要です。</p><h3>情報が少なすぎる</h3><p>クエリ、プロンプト、またはコンテキスト ウィンドウに提供される情報が少なすぎると、LLM が応答を生成するための正しいセマンティック コンテキストを正確に判断できないため、幻覚が発生する可能性があります。また、ドキュメント チャンク サイズのベクトル類似性にも問題があります。つまり、短くて単純な質問は、ベクトル化された知識ベースにある豊富で詳細なドキュメントと意味的に一致しない可能性があります。<a href="https://medium.com/data-science/how-to-use-hyde-for-better-llm-rag-retrieval-a0aa5d0e23e8">Hypothetical Document Embeddings (HyDE)</a>などのクエリ拡張手法が開発され、LLM を使用して、短いクエリよりも豊富で表現力豊かな仮説的な回答を生成します。もちろん、ここでの危険は、仮説文書自体が LLM を正しい文脈からさらに逸脱させる幻覚であるということです。</p><h3>情報が多すぎる</h3><p>私たち人間と同じように、コンテキスト ウィンドウに情報が多すぎると、LLM は重要な部分が何であるのかについて混乱し、圧倒されてしまう可能性があります。コンテキスト オーバーフロー (または「<a href="https://research.trychroma.com/context-rot">コンテキスト ロット</a>」) は、生成 AI 操作の品質とパフォーマンスに影響します。LLM の「注意予算」(作業メモリ) に大きな影響を与え、競合する多くのトークン間の関連性を薄めます。「コンテキスト腐敗」の概念には、LLM が<a href="https://alexandrabarr.beehiiv.com/p/context-windows">位置の偏り</a>を持つ傾向があるという観察も含まれます。つまり、LLM はコンテキスト ウィンドウの中央セクションのコンテンツよりも、コンテキスト ウィンドウの先頭または末尾のコンテンツを優先します。</p><h3>気が散ったり矛盾したりする情報</h3><p>コンテキスト ウィンドウが大きくなるほど、LLM が正しいコンテキストを選択して処理する妨げとなる余分な情報や矛盾した情報が含まれる可能性が高くなります。ある意味、これは「ガベージ イン/ガベージ アウト」の問題になります。つまり、ドキュメント結果セットをコンテキスト ウィンドウにダンプするだけで、LLM に処理すべき大量の情報が提供されます (多すぎる可能性があります)。ただし、コンテキストの選択方法によっては、矛盾した情報や無関係な情報が入り込む可能性が高くなります。</p><h2>エージェント型AI</h2><p>カバーすべき内容がたくさんあると言いましたが、ついにエージェント AI のトピックについて話すことができました。エージェント AI は、LLM チャット インターフェイスの非常にエキサイティングな新しい使用法であり、独自の知識とユーザーが提供するコンテキスト情報に基づいて応答を合成する生成 AI (すでに「レガシー」と呼んでもいいでしょうか?) の機能を拡張します。生成 AI が成熟するにつれて、当初は人間が簡単に確認/検証できる、面倒でリスクの低いアクティビティに限定されていた、一定レベルのタスク処理と自動化を LLM に実行させることができることに気付きました。短期間で、当初のスコープは拡大しました。LLM チャット ウィンドウは、AI エージェントが自律的に計画、実行し、指定された目標を達成するためにその計画を反復的に評価および適応させるきっかけとなることができるようになりました。エージェントは、LLM 自身の推論、チャット履歴、思考メモリ (現状のまま) にアクセスでき、その目標達成に向けて活用できる特定のツールも利用できます。また、トップレベルのエージェントが、それぞれ独自のロジック チェーン、命令セット、コンテキスト、ツールを持つ複数の<a href="https://www.philschmid.de/the-rise-of-subagents">サブエージェント</a>のオーケストレーターとして機能することを可能にするアーキテクチャも登場しています。</p><p>エージェントは、ほぼ自動化されたワークフローへのエントリ ポイントです。エージェントは自己主導型であり、ユーザーとチャットしてから「ロジック」を使用して、ユーザーの質問に答えるために使用できるツールを決定します。ツールは通常、エージェントに比べて受動的であると考えられており、1 種類のタスクを実行するために構築されています。ツールが実行できるタスクの<em>種類</em>はほぼ無限です (これは本当に素晴らしいことです!) が、ツールが実行する主なタスクは、エージェントがワークフローを実行する際に考慮するコンテキスト情報を収集することです。</p><p>技術としては、エージェント AI はまだ初期段階にあり、注意欠陥障害に相当する LLM になりがちです。つまり、指示されたことをすぐに忘れてしまい、指示にまったく含まれていない他の作業に走り出してしまうことがよくあります。一見魔法のように見えますが、LLM の「推論」機能は、シーケンス内で次に最も可能性の高いトークンを予測することに基づいています。推論（あるいは将来的には、汎用人工知能（AGI））が信頼できるものになるためには、正確で最新の情報が与えられたときに、私たちが期待する通りに推論してくれるか（そしておそらく、私たち自身では考えつかなかったようなちょっとした追加情報を提供してくれるか）を検証できなければなりません。これを実現するには、エージェント アーキテクチャに、明確に通信する機能 (プロトコル)、指定されたワークフローと制約を順守する機能 (ガードレール)、タスク内の位置を記憶する機能 (状態)、使用可能なメモリ領域を管理する機能、応答が正確でありタスクの基準を満たしていることを検証する機能が必要になります。</p><h2>私に理解できる言語で話してください</h2><p>新しい開発分野ではよくあることですが (特に LLM の世界ではそうです)、当初はエージェントとツール間の通信にはかなり多くのアプローチがありましたが、すぐに<a href="https://modelcontextprotocol.io/docs/getting-started/intro">モデル コンテキスト プロトコル (MCP) が</a>事実上の標準として採用されました。モデル コンテキスト プロトコルの定義はまさにその名前の通りで、<strong> モデルが</strong><strong> コンテキスト</strong> 情報を要求および受信するために使用する<strong> プロトコル</strong> です。MCP は、LLM エージェントが外部ツールやデータ ソースに接続するためのユニバーサル アダプタとして機能し、さまざまな LLM フレームワークやツールが簡単に相互運用できるように API を簡素化および標準化します。そのため、MCP は、エージェントが目的を達成するために自律的に実行するために与えられるオーケストレーション ロジックとシステム プロンプトと、より分離された形式 (少なくとも開始エージェントに関しては分離された形式) で実行するためにツールに送信される操作との間の、一種のピボット ポイントになります。</p><p>このエコシステムは非常に新しいため、あらゆる方向への拡大が新たなフロンティアのように感じられます。エージェント間のインタラクション（もちろん<a href="https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/">Agent2Agent (A2A)</a> ）用の類似プロトコルのほか、エージェントの推論メモリを改善するプロジェクト（ <a href="https://venturebeat.com/ai/new-memory-framework-builds-ai-agents-that-can-handle-the-real-worlds">ReasoningBank</a> ）、手元のジョブに最適な MCP サーバーを選択するプロジェクト（ <a href="https://arxiv.org/abs/2505.03275">RAG-MCP</a> ）、ゼロショット分類や入力と出力のパターン検出などのセマンティック分析を<a href="https://openai.github.io/openai-guardrails-python/">ガードレール</a>として使用してエージェントが操作できる内容を制御するプロジェクトもあります。</p><p>これらの各プロジェクトの根本的な目的は、エージェント/genAI コンテキスト ウィンドウに返される情報の品質と制御を向上させることであることにお気づきでしょうか。エージェント AI エコシステムは、コンテキスト情報をより適切に処理する (制御、管理、操作する) 能力の開発を継続していますが、エージェントが処理するための<em>最も関連性の高い</em>コンテキスト情報を取得する必要性は常に存在します。</p><h2>コンテキストエンジニアリングへようこそ!</h2><p>生成 AI の用語に詳しい方なら、おそらく「プロンプト エンジニアリング」という言葉を聞いたことがあるでしょう。現時点では、プロンプト エンジニアリングはそれ自体がほぼ疑似科学となっています。プロンプト エンジニアリングは、LLM が応答を生成する際に使用する動作を積極的に記述するための最良かつ最も効率的な方法を見つけるために使用されます。「<a href="https://www.elastic.co/search-labs/blog/context-engineering-overview">コンテキスト エンジニアリング</a>」は、「プロンプト エンジニアリング」の手法をエージェント側を超えて拡張し、MCP プロトコルのツール側で利用可能なコンテキスト ソースとシステムもカバーし、コンテキストの管理、処理、生成という幅広いトピックを扱います。</p><ul><li><p><strong>コンテキスト管理</strong>- 長時間実行される、またはより複雑なエージェント ワークフロー全体で状態とコンテキストの効率を維持することに関連します。エージェントの目標を達成するために、タスクとツールの呼び出しを繰り返し計画、追跡、オーケストレーションします。エージェントが動作しなければならない「注意予算」は限られているため、コンテキスト管理は主に、コンテキスト ウィンドウを絞り込んでコンテキストの最大限の範囲と最も重要な部分 (精度と再現率) の両方をキャプチャするのに役立つ手法に関係しています。技術には、圧縮、要約、前のステップまたはツール呼び出しからのコンテキストを永続化して、後続のステップで追加のコンテキストのために作業メモリ内にスペースを確保することが含まれます。</p></li><li><p><strong>コンテキスト処理</strong>- エージェントがすべてのコンテキストをある程度統一された方法で推論できるように、異なるソースから取得したコンテキストを統合、正規化、または調整するための論理的かつできればほとんどプログラム的な手順。基本的な作業は、すべてのソース (プロンプト、RAG、メモリなど) からのコンテキストを、エージェントが可能な限り効率的に使用できるようにすることです。 </p></li><li><p><strong>コンテキスト生成</strong>- コンテキスト処理が、取得したコンテキストをエージェントが使用できるようにすることであるならば、コンテキスト生成は、追加のコンテキスト情報を自由に、しかし制約付きで要求して受け取るための範囲をエージェントに提供します。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5e1e68c08fe050bc/6a17ddb7414c645035945073/4a8240e1eb078b2294b8d981b9caa8593589cac4-1600x900.png" alt="LLMにおけるコンテキストエンジニアリング" /><p>LLM チャット アプリケーションのさまざまな一時的な機能は、コンテキスト エンジニアリングの高レベル機能に直接 (場合によっては重複して) マッピングされます。</p><ul><li><p><strong>指示 / システム プロンプト</strong>- プロンプトは、生成的 (またはエージェント的) AI アクティビティがユーザーの目標を達成するためにどのように思考を導くかを示す足場です。プロンプトはそれ自体がコンテキストです。単なる音声による指示ではなく、回答がユーザーの要求に完全に応えているかどうかを確認するために、応答する前に「段階的に考える」や「深呼吸する」などのタスク実行ロジックやルールも含まれることがよくあります。最近のテストでは、マークアップ言語はプロンプトのさまざまな部分を組み立てるのに非常に効果的であることが示されていますが、指示を曖昧になりすぎず、具体的になりすぎないように注意して調整する必要があります。LLM が適切なコンテキストを見つけるのに十分な指示を与える必要がありますが、予期しない洞察を見逃すほど規範的であってはなりません。</p></li><li><p><strong>短期記憶</strong>(状態/履歴) - 短期記憶は、基本的にユーザーと LLM 間のチャット セッションのやり取りです。これらはライブ セッションのコンテキストを絞り込むのに役立ち、将来の取得や続行のために保存できます。 </p></li><li><p><strong>長期記憶</strong>- 長期記憶は、複数のセッションにわたって役立つ情報で構成されている必要があります。また、RAG を通じてアクセスされるのはドメイン固有の知識ベースだけではありません。最近の研究では、以前のエージェント/生成 AI 要求の結果を使用して、現在のエージェントのやり取り内で学習および参照を行っています。長期記憶領域における最も興味深い革新のいくつかは、エージェントが中断したところから再開できるように、状態がどのように<a href="https://steve-yegge.medium.com/introducing-beads-a-coding-agent-memory-system-637d7d92514a">保存され、リンクされるかを</a>調整することに関係しています。 </p></li><li><p><strong>構造化された出力</strong>- 認知には努力が必要なので、推論能力があっても、LLM が (人間と同じように) 考えるときにあまり努力を費やしたくないのは当然です。また、定義された API やプロトコルがない場合、ツール呼び出しから返されたデータを読み取る方法のマップ (スキーマ) を持つことは非常に役立ちます。<a href="https://platform.openai.com/docs/guides/structured-outputs?lang=javascript">構造化出力</a>をエージェント フレームワークの一部として組み込むと、思考主導の解析の必要性が減り、マシン間のやり取りがより高速かつ信頼性が高くなるようになります。</p></li><li><p><strong>利用可能なツール</strong>- ツールは、追加情報の収集 (エンタープライズ データ リポジトリへの RAG クエリの発行、またはオンライン API 経由の RAG クエリの発行など) から、エージェントに代わって自動アクションを実行すること (エージェントからのリクエストの基準に基づいてホテルの部屋を予約するなど) まで、さまざまな処理を実行できます。ツールは、独自のエージェント処理チェーンを持つサブエージェントになることもできます。 </p></li><li><p><strong>検索拡張生成 (RAG)</strong> - RAG の「動的な知識統合」という説明がとても気に入っています。前述のように、RAG は LLM がトレーニング時にアクセスできなかった追加情報を提供するための手法であり、主観的なクエリに最も関連性の高い正しい答えを得るために最も重要だと考えられるアイデアを繰り返し述べたものです。</p></li></ul><h2>驚異的な宇宙のパワー、小さな居住空間！</h2><p>エージェント AI には、探索すべき魅力的でエキサイティングな新しい領域が数多くあります。解決すべき従来のデータ検索および処理の問題はまだたくさんありますが、LLM の新時代に初めて日の目を見るようになったまったく新しい種類の課題もあります。私たちが現在取り組んでいる差し迫った問題の多くは、コンテキスト エンジニアリング、つまり、LLM の限られた作業メモリ空間を圧迫することなく、必要な追加のコンテキスト情報を取得することに関係しています。</p><p>さまざまなツール (および他のエージェント) にアクセスできる半自律エージェントの柔軟性により、AI を実装するための非常に多くの新しいアイデアが生まれ、さまざまな方法でそれらを組み合わせることができるのかを推測するのは困難です。現在の研究のほとんどはコンテキスト エンジニアリングの分野に属し、大量のコンテキストを処理および追跡できるメモリ管理構造の構築に重点を置いています。これは、LLM に解決してほしい深い思考の問題には、記憶することが極めて重要となる、複雑さが増し、実行時間が長く、多段階の思考ステップが含まれるためです。</p><p>この分野で現在行われている多くの実験では、エージェントの口を満たすための最適なタスク管理とツール構成を見つけようとしています。エージェントの推論チェーンにおける各ツール呼び出しは、そのツールの機能を実行するための計算と、制限されたコンテキスト ウィンドウへの影響の両方の点で累積的なコストを発生させます。LLM<a href="https://venturebeat.com/ai/ace-prevents-context-collapse-with-evolving-playbooks-for-self-improving-ai"> </a>エージェントのコンテキストを管理する最新の技術の一部は、長時間実行されるタスクの蓄積されたコンテキストを圧縮/要約すると損失が<em> 大きくなりすぎる 「</em> コンテキストの崩壊 」などの意図しない連鎖効果を引き起こしています。望ましい結果は、貴重なコンテキスト ウィンドウのメモリ領域に余分な情報が漏れることなく、簡潔で正確なコンテキストを返すツールです。</p><h3>可能性が多すぎる</h3><p>私たちはツール/コンポーネントを再利用するための柔軟性を備えた職務の分離を望んでいるため、特定のデータ ソースに接続するための専用のエージェント ツールを作成することは完全に理にかなっています。各ツールは、1 つのタイプのリポジトリ、1 つのタイプのデータ ストリーム、または 1 つのユース ケースのクエリに特化できます。しかし、注意してください。時間や費用を節約し、何かが可能であると証明しようとすると、LLM をフェデレーション ツールとして使用する強い誘惑に駆られるでしょう... やめてください。私たちは以前にも<a href="https://www.elastic.co/pdf/elastic-distributed-not-federated-search.pdf">その道を歩ん</a>だことがあります。フェデレーション クエリは、受信したクエリをリモート リポジトリが理解できる構文に変換する「ユニバーサル トランスレータ」のように機能し、その後、複数のソースからの結果を何らかの方法で合理化して一貫した応答を生成する必要があります。技術としてのフェデレーションは小規模では <em>問題なく</em><em> 機能します が、大規模で、特にデータがマルチモーダルである場合、フェデレーションは大きすぎるギャップを埋めようとします。</em></p><p>エージェントの世界では、エージェントがフェデレーターとなり、ツール (MCP 経由) がさまざまなリソースへの手動で定義された接続となります。専用のツールを使用して接続されていないデータ ソースにアクセスすることは、クエリごとにさまざまなデータ ストリームを動的に統合する強力な新しい方法のように思えるかもしれませんが、ツールを使用して複数のソースに同じ質問をすると、解決するよりも多くの問題が発生する可能性があります。これらのデータ ソースはそれぞれ、その下にある異なるタイプのリポジトリである可能性があり、それぞれが内部のデータを取得、ランク付け、保護するための独自の機能を備えています。もちろん、リポジトリ間のこうした差異、つまり「インピーダンスの不一致」により、処理負荷が増加します。また、矛盾する情報やシグナルが生じる可能性があり、スコアの不一致のように一見無害に見えるものでも、返されたコンテキストの重要性が大きく損なわれ、最終的に生成された応答の関連性に影響する可能性があります。</p><h3>コンテキストスイッチはコンピュータにとっても難しい</h3><p>エージェントを任務に送り出す場合、多くの場合、最初の任務はエージェントがアクセスできるすべての関連データを見つけることです。人間の場合と同様に、エージェントが接続する各データ ソースが類似していない分散した応答を返すと、取得したコンテンツから重要なコンテキスト ビットを抽出することに関連する認知負荷 (まったく同じ種類ではありませんが) が発生します。これには時間と計算がかかり、エージェントのロジック チェーンでは少しずつ蓄積されていきます。このことから、 <a href="https://blog.cloudflare.com/code-mode/">MCP</a>について議論されているように、ほとんどのエージェント ツールは、API (既知の入力と出力を持つ分離された関数で、さまざまな種類のエージェントのニーズをサポートするように調整された) のように動作する必要があるという結論に至ります。実際、 <a href="https://arxiv.org/html/2501.12372v5">LLM にはコンテキストのためのコンテキストが必要である</a>ことにも気づき始めています。特に、自然言語を構造化構文に翻訳するようなタスクでは、参照できるスキーマがあれば、LLM は意味の点と点を結びつけるのがはるかに上手です (まさに RTFM!)。</p><h2>7回裏ストレッチ！</h2><p>ここでは、 <a href="https://www.elastic.co/search-labs/blog/context-engineering-hybrid-search-evolution-agentic-ai">LLM がデータの取得とクエリに与えた影響</a>と、チャット ウィンドウがエージェント AI エクスペリエンスへとどのように成熟しているかについて説明しました。これら 2 つのトピックを組み合わせて、最新の検索機能と取得機能を使用してコンテキスト エンジニアリングの結果を改善する方法を見てみましょう。<a href="https://www.elastic.co/search-labs/blog/context-engineering-hybrid-search-agentic-ai-accuracy">パート III へ進みます: コンテキスト エンジニアリングにおけるハイブリッド検索の威力</a>!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/context-engineering-llm-evolution-agentic-ai</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/context-engineering-llm-evolution-agentic-ai</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Woody Walton]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f98889141fba45b/6a17ddb80b0bed0822dd34a2/79c0378b68d74d9e018c35ee2c1fd17daeee9f2c-1080x608.webp" length="0" type="image/webp"/>
    <pubDate>Tue, 18 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch で A2A プロトコルと MCP を使用して LLM エージェント ニュースルームを作成する: パート I]]></title>
    <description><![CDATA[専門の LLM エージェントが協力してニュース記事の調査、執筆、編集、公開を行う実践的なニュースルームの例で、A2A プロトコルと MCP の概念を探ります。]]></description>
    <content:encoded><![CDATA[<h2>はじめに</h2><p>現在の LLM 対応システムは、単一モデルのアプリケーションから、専門のエージェントが連携して、現代のコンピューティングではこれまで不可能と思われていたタスクを達成する複雑なネットワークへと急速に進化しています。これらのシステムの複雑さが増すにつれて、エージェントの通信とツールへのアクセスを可能にするインフラストラクチャが開発の主な焦点になります。これらのニーズに対応するために、マルチエージェント調整用の<strong>Agent2Agent (A2A)</strong>プロトコルと、標準化されたツールおよびリソース アクセス用の<strong>Model Context Protocol (MCP) という</strong>2 つの補完的なアプローチが登場しました。</p><p>それぞれの機能をいつ、またいつ単独で、調和して使用するかを理解することは、アプリケーションのスケーラビリティ、保守性、および有効性に大きな影響を与える可能性があります。この記事では、専門の LLM エージェントが協力してニュース記事の調査、執筆、編集、公開を行うデジタル ニュースルームの実際の例を通して、 <strong>A2A</strong>の概念と実装について説明します。</p><p>付属のリポジトリは<a href="https://github.com/justincastilla/elastic-newsroom/tree/main">ここに</a>あります。セクション 5 の最後の方で、A2A の実際の動作の具体的な例を検討します。</p><h3>要件</h3><p><a href="https://github.com/justincastilla/elastic-newsroom/tree/main">リポジトリは</a>、A2A エージェントの Python ベースの実装で構成されています。Flask には API サーバーが用意されているほか、ログ記録や UI 更新のメッセージをルーティングする Event Hub というカスタム Python メッセージング サービスも用意されています。最後に、ニュースルームの機能をスタンドアロンで使用するための React UI が提供されます。実装を容易にするために、すべてが Docker イメージ内に含まれています。マシンで直接サービスを実行する場合は、次のテクノロジがインストールされていることを確認してください。</p><p>言語とランタイム</p><ul><li><p>Python 13.12 - コアバックエンド言語</p></li><li><p>Node.js 18+ - オプションのReact UI</p></li></ul><p>コアフレームワークと SDK:</p><ul><li><p>A2A SDK 0.3.8 - エージェントの調整と通信</p></li><li><p>Anthropic SDK - AI生成のためのClaude統合</p></li><li><p>Uvicorn - エージェントを実行するためのASGIサーバー</p></li><li><p>FastMCP 2.12.5+ - MCP サーバーの実装</p></li><li><p>React 18.2 - フロントエンドUIフレームワーク</p></li></ul><p>データと検索</p><ul><li><p>Elasticsearch 9.1.1 以上- 記事のインデックス作成と検索</p></li></ul><p>Docker のデプロイメント (オプションですが推奨)</p><ul><li><p>Docker 28.5.1 以上</p></li></ul><h2>セクション 1: Agent2Agent (A2A) とは何ですか?</h2><h3>定義とコアコンセプト</h3><p>Agent2Agent (A2A) は、独立した LLM エージェント間の相互作用のための標準化されたプロトコルです。A2A は、すべてのタスクを処理する単一のモノリシック システムではなく、複数の専門エージェントが通信、調整、および連携して、単一のエージェントでは効率的に処理するのが困難、遅い、またはまったく不可能な複雑なワークフローを実現できるようにします。</p><p><strong>公式仕様</strong>: <a href="https://a2a-protocol.org/latest/specification/">https://a2a-protocol.org/latest/specification/</a></p><h3>起源と進化</h3><p>エージェント間通信、つまりマルチエージェント システムの概念は、<a href="https://en.wikipedia.org/wiki/Multi-agent_system">数十年</a>前に遡る分散システム、マイクロサービス、およびマルチエージェントの研究に根ざしています。分散型人工知能の初期の研究は、交渉、調整、共同作業ができるエージェントの基盤を築きました。これらの初期のシステムは、大規模な<a href="https://www.jasss.org/5/1/7.html">社会シミュレーション</a>、<a href="https://arxiv.org/html/2410.09403v1">学術研究</a>、<a href="https://www.researchgate.net/publication/334765661_Generation_Expansion_Planning_Considering_Investment_Dynamic_of_Market_Participants_Using_Multi-agent_System">電力網管理</a>に特化していました。</p><p>LLM が利用可能になり、運用コストが削減されたことで、Google や AI 研究コミュニティ全体の支援を受けて、マルチエージェント システムが「プロシューマー」市場で利用可能になりました。現在 Agent2Agent システムとして知られている A2A プロトコルの追加により、複数の大規模言語モデルが取り組みとタスクを調整する時代に合わせて特別に設計された最新の標準へと進化しました。</p><p>A2A プロトコルは、LLM が接続して通信するインタラクション ポイントに一貫した標準と原則を適用することで、エージェント間のシームレスな通信と調整を保証します。この標準化により、異なる開発者のエージェントが、異なる基盤モデルを使用して、効果的に連携できるようになります。</p><p>通信プロトコルは新しいものではなく、インターネット上で行われるほぼすべてのデジタル取引に広く定着しています。<a href="https://www.elastic.co/search-labs">https://www.elastic.co/search-labs</a>と入力した場合この記事にアクセスするためにブラウザにログインすると、TCP/IP、HTTP トランスポート、DNS ルックアップ プロトコルがすべて実行され、一貫したブラウジング エクスペリエンスが保証される可能性が高くなります。</p><h3>主な特徴</h3><p>A2A システムは、スムーズな通信を確保するためにいくつかの基本原則に基づいて構築されています。これらの原則に基づいて構築することで、異なる LLM、フレームワーク、プログラミング言語に基づくさまざまなエージェントがすべてシームレスに対話できるようになります。</p><p>主な原則は次の 4 つです。</p><ul><li><p><strong>メッセージパッシング</strong>: エージェントは、明確に定義されたプロパティとフォーマットを持つ構造化されたメッセージを通じて通信します。</p></li><li><p><strong>調整</strong>: エージェントは、他のエージェントをブロックすることなく、タスクを互いに委任し、依存関係を管理することで、複雑なワークフローを調整します。</p></li><li><p><strong>専門分野</strong>: 各エージェントは特定のドメインまたは機能に焦点を合わせ、その分野の専門家となり、そのスキルセットに基づいてタスクの完了を提供します。</p></li><li><p><strong>分散状態</strong>: 状態と知識は集中化されるのではなくエージェント間に分散され、エージェントはタスクの状態と部分的な戻り値(成果物)の進捗状況を相互に更新する機能を持ちます。</p></li></ul><h3>ニュースルーム：実例</h3><p>ジャーナリズムのさまざまな側面に特化した AI エージェントによって駆動されるデジタル ニュースルームを想像してみてください。</p><ul><li><p><strong>ニュースチーフ</strong>（コーディネーター/クライアント）：ストーリーを割り当て、ワークフローを監督する</p></li><li><p><strong>記者エージェント</strong>：調査やインタビューに基づいて記事を書く</p></li><li><p><strong>研究エージェント</strong>: 事実、統計、背景情報を収集します</p></li><li><p><strong>アーカイブエージェント</strong>: Elasticsearchを使用して過去の記事を検索し、傾向を特定します</p></li><li><p><strong>エディターエージェント</strong>: 記事の品質、スタイル、SEO最適化をレビューします</p></li><li><p><strong>パブリッシャーエージェント</strong>: 承認された記事をCI/CD経由でブログプラットフォームに公開します。</p></li></ul><p>これらのエージェントは単独では機能しません。ニュースチーフが<em>再生可能エネルギーの導入</em>についての記事を割り当てる場合、記者は統計を収集する研究者、草稿を確認する編集者、そして最終記事を公開する発行者を必要とします。この調整は A2A プロトコルを通じて行われます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb6c7215a96326481/6a17f2dd445de953024d0243/cc0760dbd74c49b92fa00dafbb8c2e8740eb70b6-963x693.png" alt="" /><h2>セクション2: A2Aアーキテクチャの理解</h2><h3>クライアントエージェントとリモートエージェントの役割</h3><p>A2A アーキテクチャでは、エージェントは主に 2 つの役割を担います。<strong>クライアント エージェントは</strong>、タスクを策定し、システム内の他のエージェントに伝達する役割を担います。リモート エージェントとその機能を識別し、この情報を使用してタスクの委任について十分な情報に基づいた決定を下します。クライアント エージェントはワークフロー全体を調整し、タスクが適切に分散され、システムが目標に向かって進行することを保証します。</p><p>対照的に、<strong>リモート エージェントは</strong>、クライアントによって委任されたタスクを実行します。リクエストに応じて情報を提供したり特定のアクションを実行したりしますが、独自にアクションを開始することはありません。リモート エージェントは、割り当てられた責任を果たすために必要に応じて他のリモート エージェントと通信し、特殊な機能の共同ネットワークを作成することもできます。</p><p>私たちのニュースルームでは、ニュースチーフがクライアントエージェントとして機能し、レポーター、リサーチャー、エディター、パブリッシャーはリクエストに応答し、互いに調整するリモートエージェントとして機能します。</p><h3>コアA2A機能</h3><p>A2A プロトコルは、マルチエージェントのコラボレーションを可能にするいくつかの機能を定義します。</p><h4>1. 発見</h4><p>A2A サーバーは、クライアントが特定のタスクにいつどのようにサーバーを利用できるかがわかるように、その機能をアナウンスする必要があります。これは、エージェントの能力、入力、出力を記述する JSON ドキュメントであるエージェント カードを通じて実現されます。エージェント カードは、一貫性のあるよく知られたエンドポイント (推奨される<code>/.well-known/agent-card.json</code>エンドポイントなど) で利用できるようになり、クライアントはコラボレーションを開始する前にエージェントの機能を検出して照会できるようになります。</p><p>以下は、Elastic のカスタム アーカイブ エージェント「Archie Archivist」のエージェント カードの例です。Elastic などのソフトウェア プロバイダーは A2A エージェントをホストし、アクセス用の URL を提供していることに注意してください。</p>{
  "name": "Archie Archivist",
  "description": "Helps find historical news documents in the Elasticsearch Index of archived news articles and content.",
  "url": "https://xxxxxxxxxxxxx-abc123.kb.us-central1.gcp.elastic.cloud/api/agent_builder/a2a/archive-agent",
  "provider": {
    "organization": "Elastic",
    "url": "https://elastic.co"
  },
  "version": "0.1.0",
  "protocolVersion": "0.3.0",
  "preferred_transport": "JSONRPC",
  "documentationURL": "https://www.elastic.co/docs/solutions/search/agent-builder/a2a-server"
  "capabilities": {
    "streaming": false,
    "pushNotifications": false,
    "stateTransitionHistory": false
  },
  "skills": [
    {
      "id": "platform.core.search",
      "name": "platform.core.search",
      "description": "A powerful tool for searching and analyzing data within your Elasticsearch cluster.",
      "inputModes": ["text/plain", "application/json"],
      "outputModes": ["text/plain", "application/json"]
    },
    {
      "id": "platform.core.index_explorer",
      "name": "platform.core.index_explorer",
      "description": "List relevant indices, aliases and datastreams based on a natural language query.",
      "inputModes": ["text/plain", "application/json"],
      "outputModes": ["text/plain", "application/json"]
    }
  ],
  "defaultInputModes": ["text/plain"],
  "defaultOutputModes": ["text/plain"]
}<p>このエージェント カードでは、Elastic のアーカイブ エージェントのいくつかの重要な側面について説明します。エージェントは自身を「Archie Archivist」と名乗り、Elasticsearch インデックス内の過去のニュース文書の検索を支援するという目的を明確に述べています。カードはプロバイダー (Elastic) とプロトコル バージョン (0.3.0) を指定し、他の A2A 準拠エージェントとの互換性を確保します。最も重要なのは、 <code>skills</code>配列が、強力な検索機能やインテリジェントなインデックス探索など、このエージェントが提供する特定の機能を列挙していることです。各スキルはサポートする入力モードと出力モードを定義し、クライアントがこのエージェントと通信する方法を正確に理解できるようにします。このエージェントは Elastic の Agent Builder サービスから派生したもので、データ ストアからデータを取得するだけでなく、データ ストアと対話するためのネイティブ LLM 対応ツールと API エンドポイントのスイートを提供します。Elasticsearch の A2A エージェントへのアクセスについては、<a href="https://www.elastic.co/docs/solutions/search/agent-builder/a2a-server">こちらを</a>ご覧ください。</p><h4>2. 交渉</h4><p>クライアントとエージェントは、適切なユーザー インタラクションとデータ交換を確保するために、コミュニケーション方法 (インタラクションがテキスト、フォーム、iframe、またはオーディオ/ビデオを介して行われるかどうか) について合意する必要があります。このネゴシエーションはエージェントのコラボレーションの開始時に行われ、ワークフロー全体にわたるエージェントの相互作用を管理するプロトコルを確立します。たとえば、音声ベースのカスタマー サービス エージェントはオーディオ ストリーム経由での通信をネゴシエートする可能性がありますが、データ分析エージェントは構造化された JSON を好む可能性があります。交渉プロセスにより、両当事者がそれぞれの能力と現在のタスクの要件に適した形式で情報を効果的に交換できるようになります。</p><p>上記の JSON スニペットにリストされている機能にはすべて入力スキーマと出力スキーマがあり、これらによって、他のエージェントからこのエージェントと対話する方法の期待値が設定されます。</p><h4>3. タスクと状態の管理</h4><p>クライアントとエージェントには、タスク実行全体を通じてタスクのステータス、変更、依存関係を通信するためのメカニズムが必要です。これには、タスクの作成と割り当てから進捗状況の更新とステータスの変更までのタスクのライフサイクル全体の管理が含まれます。一般的なステータスには、保留中、進行中、完了、失敗などの状態が含まれます。また、システムは、依存タスクが開始する前に前提条件となる作業が完了していることを確認するために、タスク間の依存関係を追跡する必要があります。エラー処理と再試行ロジックも重要なコンポーネントであり、システムが障害から正常に回復し、主な目標に向かって前進し続けることを可能にします。</p><p>タスクメッセージの例:</p>{
  "message_id": "msg_789xyz",
  "message_type": "task_request",
  "sender": "news_chief",
  "receiver": "researcher_agent",
  "timestamp": "2025-09-30T10:15:00Z",
  "payload": {
    "task_id": "task_456abc",
    "capability": "fact_gathering",
    "parameters": {
      "query": "renewable energy adoption rates in Europe 2024",
      "sources": ["eurostat", "iea", "ember"],
      "depth": "comprehensive"
    },
    "context": {
      "story_id": "story_123",
      "deadline": "2025-09-30T18:00:00Z",
      "priority": "high"
    }
  }
}<p>このサンプル タスク メッセージは、A2A 通信のいくつかの重要な側面を示しています。</p><ul><li><p><strong>メッセージ</strong>構造には、一意のメッセージ識別子、送信されるメッセージの種類、送信者と受信者の識別、追跡およびデバッグ用のタイムスタンプなどのメタデータが含まれます。</p></li><li><p><strong>ペイロードには</strong>実際のタスク情報が含まれており、リモート エージェントで呼び出される機能を指定し、その機能を実行するために必要なパラメータを提供します。</p></li><li><p><strong>コンテキスト</strong>セクションでは、受信側エージェントが広範なワークフローを理解するのに役立つ追加情報が提供されます。これには、エージェントがリソースを割り当てて作業をスケジュールする方法を示す期限や優先度レベルなどが含まれます。</p></li></ul><h4>4. コラボレーション</h4><p>クライアントとエージェントは、動的かつ構造化されたインタラクションをサポートし、エージェントがクライアント、他のエージェント、またはユーザーに説明、情報、またはサブアクションを要求できるようにする<strong>必要があります</strong>。これにより、エージェントが最初の指示が曖昧な場合にフォローアップの質問をしたり、より適切な決定を下すために追加のコンテキストを要求したり、より適切な専門知識を持つ他のエージェントにサブタスクを委任したり、完全なタスクに進む前にフィードバック用の中間結果を提供したりできる共同作業環境が作成されます。この多方向のコミュニケーションにより、エージェントは孤立して作業するのではなく、継続的な対話に参加してより良い結果を得ることができます。</p><h3>分散型ピアツーピア通信</h3><p>A2A は、エージェントが異なる組織によってホストされ、一部のエージェントが社内で管理され、他のエージェントがサードパーティのサービスによって提供される分散通信を可能にします。これらのエージェントは、複数のクラウド プロバイダーまたはオンプレミスのデータ センターにまたがる可能性のある、さまざまなインフラストラクチャで実行できます。エージェントによっては、GPT モデルを活用したエージェント、Claude を活用したエージェント、オープンソースの代替手段を活用したエージェントなど、基盤となる LLM が異なる場合があります。エージェントは、データ主権の要件に準拠したり、待ち時間を削減したりするために、異なる地理的領域にまたがって動作する場合もあります。この多様性にもかかわらず、すべてのエージェントは情報を交換するための共通の通信プロトコルに同意し、実装の詳細に関係なく相互運用性を保証します。この分散アーキテクチャにより、システムの構築と展開に柔軟性が提供され、組織は特定のニーズに合わせて最適なエージェントとインフラストラクチャを組み合わせることができます。</p><p>これはニュースルーム アプリケーションの最終的なアーキテクチャです。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt74d59cd9267f54d8/6a17f2de505ac31129ad8c71/82e01a0d9746038eafd69d11177042b5390507ae-1600x838.png" alt="" /><h2>セクション3: モデルコンテキストプロトコル (MCP)</h2><h3>定義と目的</h3><p>モデル コンテキスト プロトコル (MCP) は、Anthropic によって開発された標準化されたプロトコルであり、ユーザー定義のツール、リソース、プロンプト、その他の補足的なコードベースの追加機能を使用して個々の LLM を強化および強化します。MCP は、言語モデルと、タスクを効果的に完了するために必要な外部リソースとの間のユニバーサル インターフェイスを提供します。この<a href="https://www.elastic.co/search-labs/blog/mcp-current-state">記事では</a>、ユースケース、新たなトレンド、Elastic 独自の実装の例を挙げて、MCP の現状を概説します。</p><h3>MCPのコアコンセプト</h3><p>MCP は、次の 3 つの主要コンポーネントを持つクライアント サーバー アーキテクチャで動作します。</p><ul><li><p><strong>クライアント:</strong> MCP サーバーに接続してその機能にアクセスするアプリケーション (Claude Desktop やカスタム AI アプリケーションなど)。</p></li><li><p><strong>サーバー</strong>: 言語モデルにリソース、ツール、プロンプトを公開するアプリケーション。各サーバーは、特定の機能またはデータ ソースへのアクセスを提供することに特化しています。</p><ul><li><p><strong>ツール</strong>: モデルがデータベースの検索、外部APIの呼び出し、データに対する変換の実行などのアクションを実行するために呼び出すことができるユーザー定義関数</p></li><li><p><strong>リソース:</strong>モデルが読み取り可能なデータ ソース。動的または静的データが提供され、URI パターン (REST ルートに類似) 経由でアクセスされます。</p></li><li><p><strong>プロンプト:</strong>特定のタスクを実行するためにモデルをガイドする変数を含む再利用可能なプロンプト テンプレート。</p></li></ul></li></ul><h3>リクエスト・レスポンスパターン</h3><p>MCP は、REST API に似た、使い慣れた要求と応答の相互作用パターンに従います。クライアント (LLM) がリソースを要求するかツールを呼び出すと、MCP サーバーが要求を処理して結果を返します。LLM はこれを使用してタスクを続行します。周辺サーバーを備えたこの集中型モデルは、ピアツーピアのエージェント通信に比べて、よりシンプルな統合パターンを提供します。</p><h3>ニュースルームのMCP</h3><p>私たちのニュースルームの例では、個々のエージェントが MCP サーバーを使用して必要なツールとデータにアクセスします。</p><ul><li><p><strong>研究者エージェントは</strong>以下を使用します:</p><ul><li><p>ニュース API MCP サーバー (ニュース データベースへのアクセス)</p></li><li><p>ファクトチェックMCPサーバー（信頼できる情報源との照合による主張の検証）</p></li><li><p>学術データベース MCP サーバー (学術論文と研究)</p></li></ul></li><li><p><strong>レポーターエージェントは</strong>以下を使用します:</p><ul><li><p>スタイルガイド MCP サーバー (ニュースルームの執筆基準)</p></li><li><p>テンプレート MCP サーバー (記事テンプレートとフォーマット)</p></li><li><p>画像ライブラリ MCP サーバー (ストック写真とグラフィック)</p></li></ul></li><li><p><strong>エディターエージェントは</strong>以下を使用します:</p><ul><li><p>文法チェッカーMCPサーバー（言語品質ツール）</p></li><li><p>盗作検出MCPサーバー（独創性検証）</p></li><li><p>SEO分析MCPサーバー（見出しとキーワードの最適化）</p></li></ul></li><li><p><strong>Publisher Agent は</strong>以下を使用します:</p><ul><li><p>CMS MCP サーバー (コンテンツ管理システム API)</p></li><li><p>CI/CD MCP サーバー (デプロイメント パイプライン)</p></li><li><p>Analytics MCP サーバー (追跡と監視)</p></li></ul></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt195fe0bd36d36a48/6a17f2e0b1e113afe479f36c/b67311e3b58b27f9eb1b42a7b1dbad47ef3be4ad-808x535.png" alt="" /><h2>
セクション4: アーキテクチャの比較</h2><h3>A2Aを使用する場合</h3><p>A2A アーキテクチャは<strong>、真のマルチエージェントコラボレーションを必要とするシナリオ</strong>に優れています。調整を必要とする複数ステップのワークフローでは、特にタスクに複数の順次または並列ステップが含まれる場合、反復と改良が必要なワークフロー、およびチェックポイントと検証のニーズがあるプロセスの場合に、A2A から大きなメリットが得られます。私たちのニュースルームの例では、ストーリーのワークフローでは記者が記事を書く必要がありますが、特定の事実に対する信頼性が低い場合は研究者に繰り返し報告し、その後編集者に進み、最終的に発行者に渡す必要がある場合があります。</p><p><strong>複数の領域にわたるドメイン固有の特化</strong>は、A2A のもう 1 つの強力な使用例です。より大きなタスクを達成するためにさまざまな分野の複数の専門家が必要であり、各エージェントがさまざまな側面に関する深いドメイン知識と専門的な推論機能を提供する場合、A2A はそれらの接続を行うために必要な調整フレームワークを提供します。ニュースルームはこれを完璧に例証しています。リサーチャーは情報収集、レポーターは執筆、編集者は品質管理を専門としており、それぞれが異なる専門知識を持っています。</p><p>自律的なエージェントの動作の必要性により、A2A は特に価値が高まります。<strong>独立した意思決定を行い、変化する状況に基づいて積極的な行動を示し、ワークフロー要件に動的に適応できる</strong>エージェントは、A2A アーキテクチャで成功します。特化された機能の水平スケーリングも重要な利点の 1 つです。単一の万能エージェントではなく、複数の特化エージェントが連携して動作し、同じエージェントの複数のインスタンスがサブタスクを非同期的に処理できます。たとえば、ニュースルームでニュース速報を取材しているとき、複数の記者エージェントが同時に同じニュースのさまざまな角度から取材することがあります。</p><p>最後に、真のマルチエージェントコラボレーションを必要とするタスクは A2A に最適です。これには<a href="https://arxiv.org/abs/2404.18796">、陪審員としての LLM 評価</a>メカニズム、合意形成および投票システム、および最善の結果に到達するために<strong>複数の視点が必要となる共同問題解決が</strong>含まれます。</p><h3>MCPを使用する場合</h3><p>モデル コンテキスト プロトコルは、単一の AI モデルの機能を拡張する場合に最適です。単一の AI モデルが複数のツールやデータ ソースにアクセスする必要がある場合、MCP は、集中型の推論と分散ツール、および簡単なツール統合を組み合わせた完璧なソリューションを提供します。私たちのニュースルームの例では、研究者エージェント (1 つのモデル) は、ニュース API、ファクトチェック サービス、学術データベースなど、標準化された MCP サーバーを介してアクセスされる複数のデータ ソースにアクセスする必要があります。</p><p>ツール統合の広範な共有と再利用性が重要になる場合は、標準化されたツール統合が優先されます。MCP は、一般的な統合の開発時間を大幅に短縮する、事前に構築された MCP サーバーのエコシステムを備えているため、この点で優れています。シンプルさと保守性が求められる場合、MCP の要求応答パターンは開発者に馴染みがあり、分散システムよりも理解やデバッグが容易で、運用上の複雑さも少なくなります。</p><p>最後に、MCP は、システムとのリモート通信を容易にするためにソフトウェア プロバイダーによって提供されることがよくあります。プロバイダーが提供するこれらの MCP サーバーは、独自のシステムへの標準化されたインターフェースを提供しながら、オンボーディングと開発時間を大幅に短縮し、カスタム API 開発よりも統合をはるかに簡単にします。</p><h3>両方を使用する場合 (A2A ❤️ の MCP)</h3><p><a href="https://a2a-protocol.org/latest/topics/a2a-and-mcp/">MCP 統合に関する A2A ドキュメント</a>に記載されているように、多くの高度なシステムは A2A と MCP を組み合わせることでメリットを得られます。調整と標準化の両方を必要とするシステムは、ハイブリッド アプローチに最適です。A2A はエージェントの調整とワークフロー オーケストレーションを処理し、MCP は個々のエージェントにツール アクセスを提供します。私たちのニュースルームの例では、エージェントは A2A を介して調整し、ワークフローは記者から研究者、編集者、そして発行者へと移行します。ただし、各エージェントは専用のツール用に MCP サーバーを使用するため、アーキテクチャが明確に分離されます。</p><p>ツール アクセスにそれぞれ MCP を使用する複数の特殊エージェントは、A2A によって処理されるエージェント調整レイヤーと、MCP によって管理されるツール アクセス レイヤーがある一般的なパターンを表します。このように関心事を明確に分離することで、システムの理解と保守が容易になります。</p><p>両方のアプローチを組み合わせることによる利点は非常に大きいです。特殊化、自律性、並列処理などのマルチエージェント システムの組織的な利点が得られると同時に、ツールの統合やリソース アクセスなどの MCP の標準化とエコシステムの利点も享受できます。エージェント調整 (A2A) とリソース アクセス (MCP) は明確に区別されており、重要なのは、API アクセスなどの小規模なタスクのみには A2A は必要ないことです。MCP は、マルチエージェント オーケストレーションのオーバーヘッドなしで、これらのタスクを効率的に処理します。</p><p><strong>FAQ: A2A vs. MCP - ユースケース</strong></p><p>機能</p><p>エージェント2エージェント（A2A）</p><p>モデルコンテキストプロトコル（MCP）</p><p>ハイブリッド（A2A + MCP）</p><p>主な目標</p><p>マルチエージェント調整: 専門エージェントのチームが、複雑な複数ステップのワークフローで連携できるようにします。</p><p>単一エージェントの拡張: 外部ツール、リソース、およびデータを使用して、単一の LLM/エージェントの機能を拡張します。</p><p>組み合わせた強み: A2A がチームのワークフローを処理し、MCP が各チーム メンバーにツールを提供します。</p><p>ニュースルームチームの例</p><p>ワークフロー チェーン: ニュース チーフ → レポーター → リサーチャー → 編集者 → 発行者。これは調整レイヤーです。</p><p>個々のエージェントのツール: スタイル ガイド サーバーとテンプレート サーバーにアクセスする Reporter Agent (MCP 経由)。これはツール アクセス レイヤーです。</p><p>完全なシステム: 記者は編集者 (A2A) と連携し、画像ライブラリ MCP サーバーを使用して記事のグラフィックを検索します。</p><p>いつどれを使うか</p><p>真のコラボレーション、反復、改良、または専門知識を複数のエージェントに分割する必要がある場合。</p><p>1 つのエージェントが複数のツールやデータ ソースにアクセスする必要がある場合、または独自のシステムとの標準化された統合が必要な場合。</p><p>マルチエージェント システムの組織的利点と、MCP の標準化およびエコシステムの利点が必要な場合。</p><p>コアベネフィット</p><p>自律性とスケーリング: エージェントは独立して決定を下すことができ、システムは特殊な機能の水平スケーリングを可能にします。</p><p>シンプルさと標準化: 集中化された推論によりデバッグと保守が容易になり、リソースに対する汎用的なインターフェースが提供されます。</p><p>関心事の明確な分離: システムを理解しやすくなります: A2A = チームワーク、MCP = ツール アクセス。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1735ea5de41e10fd/6a17f2e26864a4125cb688c4/ddf6a29b1107ac6a63e94ecef703abc561a29e1e-986x656.png" alt="" /><h2>まとめ</h2><p>これは、データとツールへのサポートと外部アクセスを提供するために MCP サーバーで強化された A2A ベースのエージェントの実装を扱った 2 部構成の最初のセクションです。次の部分では、実際のコードを調べて、オンライン ニュースルームのアクティビティをエミュレートするためにそれらが連携して動作する様子を示します。どちらのフレームワークも、それ自体で非常に有能で柔軟性に優れていますが、連携して動作することで、どれだけ互いを補完し合うかがわかります。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/a2a-protocol-mcp-llm-agent-newsroom-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/a2a-protocol-mcp-llm-agent-newsroom-elasticsearch</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Justin Castilla]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2716d804698ec878/6a17f2e41480095fd7b48888/9f938d8e2f0fdf7509edf028816c48bdbc8b3fc7-1600x900.png" length="0" type="image/png"/>
    <pubDate>Thu, 13 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch で構造化ドキュメントの再帰チャンクを構成する]]></title>
    <description><![CDATA[チャンク サイズ、セパレーター グループ、カスタム セパレーター リストを使用して Elasticsearch で再帰チャンクを設定し、構造化ドキュメントのインデックスを最適に作成する方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>8.16 以降、ユーザーは長いドキュメントをセマンティック テキスト フィールドに取り込むときに使用するチャンキング戦略を構成できるようになりました。9.1 / 8.19 では、正規表現のリストを使用してドキュメントをチャンク化する、新しい構成可能な再帰チャンク化戦略を導入しました。チャンク化の目的は、長いドキュメントを関連するコンテンツをカプセル化するセクションに分割することです。既存の戦略では、テキストを単語/文の粒度で分割しますが、構造化された形式 (例:Markdown では、区切り文字列で定義されたセクション内に関連コンテンツが含まれることがよくあります (例:ヘッダー)。このような種類のドキュメントでは、構造化ドキュメントの形式を活用してより適切なチャンクを作成するための再帰チャンキング戦略を導入しています。</p><h2>再帰チャンキングとは何ですか?</h2><p>再帰チャンク化では、指定されたセクション分離パターンのリストを反復処理して、必要な最大チャンク サイズを満たすまで、ドキュメントを段階的に小さなセグメントに分割します。</p><h3>再帰チャンクを構成するにはどうすればよいですか?</h3><p>以下は、再帰チャンク化に対してユーザーが指定できる構成可能な値です。</p><ul><li><p>(必須) <code>max_chunk_size</code> : チャンク内の最大単語数。</p></li><li><p>次のいずれか:</p><ul><li><p><code>separators</code>: ドキュメントをチャンクに分割するために使用される正規表現文字列パターンのリスト。</p></li><li><p><code>separator_group</code>: 特定の種類のドキュメントに使用するために Elastic によって定義された区切り文字のデフォルト リストにマップされる文字列。現在、 <code>markdown</code>と<code>plaintext</code>が利用可能です。</p></li></ul></li></ul><h3>再帰チャンキングはどのように機能しますか?</h3><p>入力ドキュメント、 <code>max_chunk_size</code> (単語単位で測定)、および区切り文字列のリストが与えられた場合の再帰チャンク化のプロセスは次のとおりです。</p><ol><li><p>入力ドキュメントがすでに最大チャンク サイズ内である場合は、入力全体にわたる単一のチャンクを返します。</p></li><li><p>区切り文字の出現に基づいてテキストを潜在的なチャンクに分割します。潜在的なチャンクごとに:</p><ol><li><p>潜在的なチャンクが最大チャンク サイズ内である場合は、ユーザーに返すチャンクのリストに追加します。</p></li><li><p>それ以外の場合は、潜在的なチャンクのテキストのみを使用して、リスト内の次のセパレーターを使用して分割し、手順 2 から繰り返します。試す区切り文字がもう残っていない場合は、文ベースのチャンクに戻ります。</p></li></ol></li></ol><h2>再帰チャンクの設定例</h2><p>チャンク サイズとは別に、再帰チャンク化の主な構成は、ドキュメントを分割するために使用するセパレーターを選択することです。どこから始めればよいかわからない場合は、Elasticsearch では一般的なユースケースに使用できるデフォルトのセパレーター グループがいくつか用意されています。</p><h3>セパレーターグループの活用</h3><p>セパレーター グループを利用するには、チャンク設定を構成するときに使用するグループの名前を指定するだけです。例えば：</p>"chunking_settings": {
    "strategy": "recursive",
    "max_chunk_size": 25,
    "separator_group": "plaintext"
}<p>これにより、区切りリスト<code>["(?&lt;!\\n)\\n\\n(?!\\n)", "(?&lt;!\\n)\\n(?!\\n)")]</code>を利用する再帰的なチャンク化戦略が提供されます。これは、2 つの改行文字とそれに続く 1 つの改行文字で分割する、一般的なプレーン テキスト アプリケーションに適しています。</p><p>セパレーターリストを利用するセパレーターグループ<code>markdown</code>も提供しています。</p>[
"\n# ",
       "\n## ",
       "\n### ",
       "\n#### ",
       "\n##### ",
       "\n###### ",
       "\n^(?!\\s*$).*\\n-{1,}\\n",
       "\n^(?!\\s*$).*\\n={1,}\\n"
]<p>この区切りリストは、6 つの見出しレベルとセクション区切り文字のそれぞれに分割する一般的なマークダウンの使用例に適しています。</p><p>リソース (推論エンドポイント/セマンティック テキスト フィールド) を作成すると、その時点のセパレーター グループに対応するセパレーターのリストが構成に保存されます。セパレーター グループが後日更新されても、既に作成されたリソースの動作は変更されません。</p><h3>カスタム区切りリストの利用</h3><p>定義済みの区切り文字グループのいずれかが使用ケースに適していない場合は、ニーズに合った区切り文字のカスタム リストを定義できます。区切りリスト内に正規表現を指定できることに注意してください。以下は、カスタムセパレーターを使用して構成されたチャンク設定の例です。</p>"chunking_settings": {
    "strategy": "recursive",
    "max_chunk_size": 25,
    "separators": ["\n\n", "\n", "&lt;my-custom-separator&gt;"]
}<p>上記のチャンク化戦略では、 2 つの改行文字、続いて 1 つの改行文字、最後に文字列<code>“&lt;my-custom-separator&gt;”</code>で分割されます。</p><h2>再帰チャンキングの実際の例</h2><p>再帰チャンキングの実際の例を見てみましょう。この例では、上位 2 つのヘッダー レベルを使用してマークダウン ドキュメントを分割するセパレーターのカスタム リストとともに、次のチャンク設定を使用します。</p>"chunking_settings": {
    "strategy": "recursive",
    "max_chunk_size": 25,
    "separators": ["\n# ", "\n## "]
}<p>単純なチャンクなしの Markdown ドキュメントを見てみましょう。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdb5f41d1bd43ba50/6a17e831e9ea87c1d8a9c5f3/3a5507f4a1288065097231548e5b18e240508785-1302x1446.png" alt="チャンクなしのMarkdown文書" /><p>ここで、上で定義したチャンク設定を使用してドキュメントをチャンク化してみましょう。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfffda162c7b9c87a/6a17e83296142aefa8eb1b0b/a3313c4c40ff39b8dbcdd7c4878c723f088e6c1a-1600x1187.png" alt="Elasticsearchでドキュメントをチャンク化する" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt96f65346a8e09e3a/6a17e834445de9157b4d015e/79a2921943191ea631df94c9d465818ec8d3e738-1600x1206.png" alt="2番目のセパレーターで分割 - Elasticsearchでドキュメントをチャンク化する" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt28381c8f85aedf07/6a17e836ec0f89801e5a6640/459e695cce7540267422396b9a62ff4ad35f61db-1600x1260.png" alt="Elasticsearch で文ベースのチャンクを分割した後のドキュメントの最終チャンク" /><p>注: 各チャンク (チャンク 3 を除く) の末尾の改行は強調表示されませんが、実際のチャンク境界内に含まれます。</p><h3>今すぐ再帰チャンキングを始めましょう!</h3><p>この機能の利用方法の詳細については、チャンク設定の構成に関するドキュメントを参照してください。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/recursive-chunking-structured-documents-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/recursive-chunking-structured-documents-elasticsearch</guid>
    <category><![CDATA[基本]]></category>
    <category><![CDATA[Elastic内部の実情]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Daniel Rubinstein]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf442dc4941f37be7/6a17e838505ac3eaf8ad8b3d/591872e31880768ca927507654a621addc0d124d-1600x960.png" length="0" type="image/png"/>
    <pubDate>Tue, 11 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch と SigLIP-2 による山頂のマルチモーダル探索 ]]></title>
    <description><![CDATA[SigLIP-2 埋め込みと Elasticsearch kNN ベクトル検索を使用して、テキストから画像、画像から画像へのマルチモーダル検索を実装する方法を学びます。プロジェクトの焦点: エベレスト トレッキングでアマ ダブラム山の山頂の写真を探す。]]></description>
    <content:encoded><![CDATA[<p>写真アルバムを意味で検索したいと思ったことはありませんか?「青いジャケットを着てベンチに座っている写真を見せてください」「エベレストの写真を見せてください」「日本酒と寿司」などのクエリを試してみてください。コーヒー（またはお好みの飲み物）を飲みながら、読み続けてください。このブログでは、マルチモーダル ハイブリッド検索アプリケーションの構築方法を紹介します。マルチモーダルとは、アプリが単語だけでなく、テキスト、画像、音声などさまざまな種類の入力を理解して検索できることを意味します。ハイブリッドとは、キーワード マッチング、kNN ベクトル検索、ジオフェンシングなどの技術を組み合わせて、より鮮明な結果を提供することを意味します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdfa1ec1ccd450e94/6a17da751d1b8308ee93e344/0ec6bbb45013846b59ee00d2bf73ee2182ee7392-1920x1080.gif" alt="エベレスト登山のさまざまな山頂の写真のライブラリ。" /><p>これを実現するために、Google の SigLIP-2 を使用して画像とテキストの両方のベクトル埋め込みを生成し、Elasticsearch ベクトル データベースに保存します。クエリ時に、検索入力、テキストまたは画像を埋め込みに変換し、高速 kNN ベクトル検索を実行して結果を取得します。この設定により、効率的なテキストから画像への検索、画像から画像への検索が可能になります。Streamlit UI は、テキストベースの検索を行ってアルバムから一致する写真を検索して表示するだけでなく、アップロードされた画像から山頂を識別し、フォトアルバムでその山の他の写真を表示できるフロントエンドを提供することで、このプロジェクトを実現します。また、検索精度を向上させるために実行した手順や、実用的なヒントやコツについても説明します。さらに詳しく調べるために、 <a href="https://github.com/navneet83/multimodal-mountain-peak-search">GitHub リポジトリ</a>と<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/notebooks/multimodal_mountain_peak_search.ipynb">Colab ノートブック</a>を提供しています。</p><h2>始まり</h2><p>このブログ投稿は、エベレストベースキャンプトレッキングで撮ったアマダブラム山の写真を全部見せてほしいと私に頼んだ10歳の子供からインスピレーションを受けたものです。写真アルバムを精査しながら、私は他のいくつかの山頂を特定するよう求められましたが、そのうちのいくつかは名前がわかりませんでした。</p><p>それで、これは楽しいコンピューター ビジョン プロジェクトになるかもしれないというアイデアが浮かびました。私たちが達成したかったこと:</p><ul><li><p>山頂の写真を名前で検索する</p></li><li><p>画像から山頂の名前を推測し、写真アルバムで似たような山頂を見つける</p></li><li><p>概念クエリを機能させる（<em>人</em>、<em>川</em>、<em>祈りの旗</em><em>など）</em></p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf82df9d7005fc3fe/6a17da78abe0f2e77bdfe8b9/e9d0d720a9b565d5b749bdc915068852d4f157ad-1200x1600.png" alt="アマ・ダブラム山 " /><h2>ドリームチームを結成: SigLIP-2、Elasticsearch、Streamlit</h2><p>これを機能させるには、テキスト (「Ama Dablam」) と画像 (私のアルバムの写真) の両方を、意味のある比較が可能なベクトル、つまり同じベクトル空間に変換する必要があることがすぐに明らかになりました。これを実行すると、検索は単に「最も近いものを見つける」だけになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f80b69a9d5bd28a/6a17da7a4b055ddd1243209e/20e6f8b7d4fa48414f407ec200adbe00ee28d517-1536x1024.png" alt="SigLIP-2、Elasticsearch、Streamlit のドリームチーム。" /><p>画像の埋め込みを生成するために、多言語<a href="https://huggingface.co/blog/vlms-2025">ビジョン言語エンコーダー</a>を使用します。これにより、山の写真と「Ama Dablam」などのフレーズが同じベクトル空間に配置されます。</p><p>最近 Google がリリースした<a href="https://huggingface.co/blog/siglip2"><strong>SigLIP-2</strong></a>はここによく当てはまります。タスク固有のトレーニング (<strong>ゼロショット</strong>設定) なしで埋め込みを生成でき、ラベルのない写真や異なる名前と言語を持つピークなど、私たちのユースケースに適しています。テキストと画像のマッチングがトレーニングされているため、クエリ言語やスペルが異なっていても、トレッキング中の山の写真と短いテキストプロンプトは埋め込みとして近いものになります。</p><p>SigLIP-2 は、品質と速度のバランスが優れており、複数の入力解像度をサポートし、CPU と GPU の両方で実行されます。SigLIP-2 は、オリジナルの CLIP などの以前のモデルと比較して、屋外での写真撮影に対してより堅牢になるように設計されています。私たちのテストでは、SigLIP-2 は一貫して信頼できる結果を生成しました。また、サポートも非常に充実しており、このプロジェクトに最適な選択肢となっています。</p><p>次に、埋め込みとパワー検索を保存するためのベクトル データベースが必要です。画像埋め込みに対するコサイン kNN 検索をサポートするだけでなく、単一のクエリでジオフェンスとテキスト フィルターを適用することもサポートする必要があります。Elasticsearch はここで最適です。ベクトル (dense_vector フィールドの HNSW kNN) を非常に適切に処理し、テキスト、ベクトル、地理クエリを組み合わせたハイブリッド検索をサポートし、フィルタリングと並べ替えをすぐに使用できます。また、水平方向にも拡張できるため、数枚の写真から数千枚の写真まで簡単に拡張できます。公式の<a href="https://www.elastic.co/docs/reference/elasticsearch/clients/python">Elasticsearch Python クライアントは</a>、配管をシンプルに保ち、プロジェクトときれいに統合します。最後に、検索クエリを入力して結果を表示できる軽量のフロントエンドが必要です。簡単な Python ベースのデモには、Streamlit が最適です。ファイルのアップロード、レスポンシブな画像グリッド、並べ替えとジオフェンシングのためのドロップダウン メニューなど、必要な基本的な機能を提供します。簡単にクローンを作成してローカルで実行でき、Colab ノートブックでも動作します。</p><h2>実装</h2><h3>Elasticsearchのインデックス設計とインデックス戦略</h3><p>このプロジェクトでは、 <code>peaks_catalog</code>と<code>photos</code>の 2 つのインデックスを使用します。</p><h4>Peaks_catalogインデックス</h4><p>この索引は、エベレストベースキャンプトレッキング中に見える主要な山頂のコンパクトなカタログとして機能します。このインデックス内の各ドキュメントは、エベレスト山などの単一の山頂に対応しています。各山頂ドキュメントには、名前/エイリアス、オプションの緯度経度座標、および SigLIP-2 テキストプロンプト (+ オプションの参照画像) を組み合わせて構築された単一のプロトタイプ ベクトルが保存されます。</p><p><strong>インデックスマッピング:</strong></p><p>分野</p><p>タイプ</p><p>例</p><p>目的/注意事項</p><p>ベクトル/インデックス</p><p>id</p><p>キーワード</p><p>アマ・ダブラム</p><p>安定したスラッグ/ID</p><p>—</p><p>名前</p><p>テキスト + キーワードサブフィールド</p><p>["アマ・ダブラム"、"アマダブラム"]</p><p>エイリアス/多言語名; 正確なフィルターのためのnames.raw</p><p>—</p><p>ラトロン</p><p>ジオポイント</p><p>{"lat":27.8617,"lon":86.8614}</p><p>緯度/経度の組み合わせによるピーク GPS 座標 (オプション)</p><p>—</p><p>高度m</p><p>整数</p><p>6812</p><p>標高（オプション）</p><p>—</p><p>テキスト埋め込み</p><p>dense_vector</p><p>768</p><p>このピークのブレンドプロトタイプ（プロンプトとオプションで1～3枚の参照画像）</p><p>index:true、類似度:"cosine"、index_options: {type:"hnsw", m:16, ef_construction:128}</p><p>このインデックスは主に、画像から山頂を識別するなど、画像間の検索に使用されます。このインデックスは、テキストから画像への検索結果を強化するためにも使用されます。</p><p>要約すると、 <code>peaks_catalog</code>は「これは何の山ですか？」という質問を焦点を絞った最近傍問題に変換し、概念的理解を画像データの複雑さから効果的に分離します。</p><p><strong>peaks_catalog インデックスのインデックス戦略:</strong> EBC トレッキング中に見える最も目立つ山頂のリストを作成することから始めます。各山頂の地理的位置、名前、同義語、標高を<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/data/peaks.yaml">yaml ファイル</a>に保存します。次のステップは、各ピークの<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L351">埋め込みを生成し</a>、それを<code>text_embed</code>フィールドに保存することです。堅牢な埋め込みを生成するために、次の手法を使用します。</p><ul><li><p>以下を使用してテキスト プロトタイプを作成します。</p><ul><li><p>山の名前</p></li><li><p><a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L301">プロンプト アンサンブル</a>(複数の異なるプロンプトを使用して同じ質問に答える)、例:</p><ul><li><p>「ネパール、ヒマラヤ山脈の山頂{name}の自然写真」</p></li><li><p>「クンブ地域の{name}マーク的な山頂、高山の風景」</p></li><li><p>「 {name}山頂、雪、岩だらけの尾根」</p></li></ul></li><li><p>オプションの<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L333">反概念</a>（SigLIP-2 に一致しないものを指示する）: 「絵画、イラスト、ポスター、地図、ロゴ」の小さなベクトルを減算して、実際の写真に偏向させます。</p></li></ul></li><li><p>ピークの参照画像が提供されている場合は、オプションで<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L388C13-L388C29">画像プロトタイプを作成します</a>。</p></li></ul><p>次に、<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L392">テキストと画像のプロトタイプをブレンドして</a>、最終的な埋め込みを生成します。最後に、ドキュメントはすべての必須フィールドで<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L396">インデックス化され</a>ます。</p>def l2norm(v: np.ndarray) -&gt; np.ndarray:
    return v / (np.linalg.norm(v) + 1e-12)
def compute_blended_peak_vec(
        emb: Siglip2,
        names: List[str],
        peak_id: str,
        peaks_images_root: str,
        alpha_text: float = 0.5,
        max_images: int = 3,
) -&gt; Tuple[np.ndarray, int, int, List[str]]:
    """
    Build blended vector for a single peak.

    Returns:
      vec           : np.ndarray (L2-normalized)
      found_count   : number of reference images discovered
      used_count    : number of references used (&lt;= max_images)
      used_filenames: list of filenames used (for logging)
    """
    # 1) TEXT vector
    tv = embed_text_blend(emb, names)

    # 2) IMAGE refs: prefer folder by id; fallback to slug of the primary name
    root = Path(peaks_images_root)
    candidates = [root / peak_id]
    if names:
        candidates.append(root / slugify(names[0]))

    all_refs: List[Path] = []
    for c in candidates:
        if c.exists() and c.is_dir():
            all_refs = list_ref_images(c)
            if all_refs:
                break

    found = len(all_refs)
    used_list = all_refs[:max_images] if (max_images and found &gt; max_images) else all_refs
    used = len(used_list)

    img_v = embed_image_mean(emb, used_list) if used_list else None

    # 3) Blend TEXT and IMAGE vectors, clamp alpha to [0,1]
    a = max(0.0, min(1.0, float(alpha_text)))
    vec = l2norm(tv if img_v is None else (a * tv + (1.0 - a) * img_v)).astype("float32")
    return vec, found, used, [p.name for p in used_list]<p><code>peaks_catalog</code>インデックスからのサンプル ドキュメント:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1219f5d0e39b512c/6a17da7c57726263161bcace/bc05fbd0c4f8d721d5170c28a3884a9eda80bb7d-1210x1132.png" alt="Elasticsearch の peaks_catalog インデックスからのサンプル ドキュメント。" /><h4>写真インデックス</h4><p>このプライマリ インデックスには、アルバム内のすべての写真に関する詳細情報が保存されます。各ドキュメントは 1 枚の写真を表し、次の情報が含まれています。</p><ul><li><p>フォトアルバム内の写真への相対パス。これを使用して、一致する画像を表示したり、検索 UI に画像を読み込んだりできます。</p></li><li><p>写真のGPSと時間情報。</p></li><li><p>SigLIP-2 によって生成された画像エンコーディング用の密なベクトル。</p></li><li><p><code>predicted_peaks</code> ピーク名でフィルタリングできます。<strong>インデックスマッピング</strong></p></li></ul><p>分野</p><p>タイプ</p><p>例</p><p>目的/注意事項</p><p>ベクター / インデックス</p><p>パス</p><p>キーワード</p><p>データ/画像/IMG_1234.HEIC</p><p>UI でサムネイル/フル画像を開く方法</p><p>—</p><p>クリップ画像</p><p>dense_vector</p><p>768</p><p>SigLIP-2画像埋め込み</p><p>index:true、類似度:"cosine"、index_options: {type:"hnsw", m:16, ef_construction:128}</p><p>予測ピーク</p><p>キーワード</p><p>["ama-dablam","pumori"]</p><p>インデックス時の上位Kの推測（安価なUXフィルター/ファセット）</p><p>—</p><p>GPS</p><p>ジオポイント</p><p>{"lat":27.96,"lon":86.83}</p><p>地理フィルターを有効にする</p><p>—</p><p>ショット時間</p><p>date</p><p>2023年10月18日09:41:00Z</p><p>撮影時間: 並べ替え/フィルター</p><p>—</p><p><strong>写真インデックスのインデックス戦略:</strong>アルバム内の写真ごとに、次の操作を実行します。
<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L526">画像メタデータから</a>画像<code>shot_time</code>と<code>gps</code>情報を抽出します。</p><ul><li><p><a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L511">SigLIP-2 画像埋め込み</a>: 画像をモデルに渡し、ベクトルを L2 正規化します。埋め込みを<code>clip_image</code>フィールドに保存します。</p></li><li><p><a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L519">ピークを予測し</a>、 <code>predicted_peaks</code>フィールドに保存します。これを行うには、まず前の手順で生成された写真の画像ベクトルを取得し、次に<code>peaks_catalog</code>インデックスの text_embed フィールドに対して簡単な kNN 検索を実行します。上位 3 ～ 4 つのピークを保持し、残りは無視します。</p></li><li><p>画像名とパスの<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L509">ハッシュ</a>を実行して<code>_id</code>フィールドを計算します。これにより、複数回実行した後に重複が発生しなくなります。</p></li></ul><p>写真のすべてのフィールドを決定したら、<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/embed_and_index_photos.py#L530">一括</a>インデックスを使用して写真ドキュメントを一括でインデックスします。</p>def bulk_index_photos(
        es: Elasticsearch,
        images_root: str,
        photos_index: str = "photos",
        peaks_index: str = "peaks_catalog",
        topk_predicted: int = 5,
        batch_size: int = 200,
        refresh: str = "false",
) -&gt; None:
    """Walk a folder of images, embed + enrich, and bulk index to Elasticsearch."""
    root = Path(images_root)
    if not root.exists():
        raise SystemExit(f"Images root not found: {images_root}")

    emb = Siglip2()
    batch: List[Dict[str, Any]] = []
    n_indexed = 0

    for p in iter_images(root):
        rel = relpath_within(root, p)
        _id = id_for_path(rel)

        # 1) Image embedding (and reuse it for predicted_peaks)
        try:
            with Image.open(p) as im:
                ivec = emb.image_vec(im.convert("RGB")).astype("float32")
        except (UnidentifiedImageError, OSError) as e:
            print(f"[skip] {rel} — cannot embed: {e}")
            continue

        # 2) Predict top-k peak names
        try:
            top_names = predict_peaks(es, ivec.tolist(), peaks_index=peaks_index, k=topk_predicted)
        except Exception as e:
            print(f"[warn] predict_peaks failed for {rel}: {e}")
            top_names = []

        # 3) EXIF enrichment (safe)
        gps = get_gps_decimal(str(p))
        shot = get_shot_time(str(p))

        # 4) Build doc and stage for bulk
        doc = {"path": rel, "clip_image": ivec.tolist(), "predicted_peaks": top_names}
        if gps:
            doc["gps"] = gps
        if shot:
            doc["shot_time"] = shot

        batch.append(
            {"_op_type": "index", "_index": photos_index, "_id": _id, "_source": doc}
        )

        # 5) Periodic flush
        if len(batch) &gt;= batch_size:
            helpers.bulk(es, batch, refresh=refresh)
            n_indexed += len(batch)
            print(f"[photos] indexed {n_indexed} (last: {rel})")
            batch.clear()

    # Final flush
    if batch:
        helpers.bulk(es, batch, refresh=refresh)
        n_indexed += len(batch)
        print(f"[photos] indexed {n_indexed} total.")

    print("[done] photos indexing")<p>写真インデックスからのサンプルドキュメント:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt744b7e6326937cfc/6a17da7e6df731d3040a0da8/1dc1406ac2a97440b6804838795b3c2205c4c6b2-1080x1234.png" alt="Elasticsearch の写真インデックスからのサンプル ドキュメント。" /><p>要約すると、写真のインデックスは、アルバム内のすべての写真を格納する、高速でフィルタリング可能な kNN 対応のストアです。マッピングは意図的に最小限に抑えられており、すばやく取得し、きれいに表示し、結果を空間と時間で分割するのに十分な構造になっています。このインデックスは、両方の検索ユースケースに対応します。両方のインデックスを作成するための Python スクリプトは<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/create_indices.py">ここに</a>あります。</p><p>以下の Kibana マップの視覚化では、写真アルバムのドキュメントが緑のドットで表示され、 <code>peaks_catalog</code>インデックスの山頂が赤い三角形で表示されています。緑のドットはエベレスト ベース キャンプのトレッキング トレイルとぴったり一致しています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb5bf016e8d9c3e84/6a17da80be608681f10045e6/1c75d0ed0ce53d28a94bf2f47a354e25581d2baf-1600x1402.png" alt="Kibana マップの視覚化では、写真アルバムのドキュメントが緑のドットで表示され、peaks_catalog インデックスの山頂が赤い三角形で表示されています。緑のドットは、エベレスト ベース キャンプのトレッキング トレイルとよく揃っています。" /><h2>検索ユースケース</h2><p><strong>名前による検索（テキストから画像へ）：</strong>この機能により、ユーザーはテキスト クエリを使用して山頂の写真（さらには「祈りの旗」のような抽象的な概念）を見つけることができます。これを実現するために、テキスト入力は SigLIP-2 を使用して<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L87C5-L87C20">テキスト ベクトルに変換されます</a>。堅牢なテキストベクトル生成のために、 インデックスでテキスト埋め込みを作成するために使用したのと同じ戦略を採用しています。つまり、テキスト入力を小さな<code>peaks_catalog</code><a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L100"> プロンプトアンサンブル</a> と<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L104"> 組み合わせ</a><a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L103"> 、小さな 反概念ベクトル</a> を減算し、<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L104"> L2正規化</a> を適用して最終的なクエリベクトルを生成します。次に、 <code>photos.clip_image</code>フィールドで kNN<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L140">クエリを</a>実行し、コサイン類似度に基づいて一致する上位のピークを取得して、最も近い画像を見つけます。オプションとして、クエリの一部として地理および日付<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/query_by_peak_name.py#L152">フィルター</a>、および/または<code>photos.predicted_peaks</code>用語フィルターを適用することで、検索結果の関連性を高めることができます (以下のクエリ例を参照)。これにより、トレッキング中に実際には見えていない、似たような山頂を除外することができます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9bb9abf5ce64fcbb/6a17da81e8fbce20db3a17da/b5fac28ffdbedb820505365ca07df125cd01b939-946x370.png" alt="Elasticsearch でのマルチモーダルな名前による検索 (テキストから画像への検索) の仕組み。" /><p><strong>ジオフィルターを使用した Elasticsearch クエリ:</strong></p>POST photos/_search
{
  "knn": {
    "field": "clip_image",
    "query_vector": [ ... ],
    "k": 60,
    "num_candidates": 2000
  },
  "query": {
    "bool": {
      "filter": [
        { "geo_bounding_box": { "gps": { "top_left": "...", "bottom_right": "..." } } }
      ]
    }
  },
  "_source": ["path","predicted_peaks","gps","shot_time"]
}

Response (first two documents):
{
 "hits": {
   "total": {
     "value": 56,
     "relation": "eq"
   },
   "max_score": 0.5779596,
   "hits": [
     {
       "_index": "photos",
       "_id": "d01da3a1141981486c3493f6053c79e92a788463",
       "_score": 0.5779596,
       "_source": {
         "path": "IMG_2738.HEIC",
         "predicted_peaks": [
           "Pumori",
           "Kyajo Ri",
           "Khumbila",
           "Nangkartshang",
           "Kongde Ri"
         ],
         "gps": {
           "lat": 27.97116388888889,
           "lon": 86.82331111111111
         },
         "shot_time": "2023-11-03T08:07:13"
       }
     },
     {
       "_index": "photos",
       "_id": "c79d251f07adc5efaedc53561110a7fd78e23914",
       "_score": 0.5766071,
       "_source": {
         "path": "IMG_2761.HEIC",
         "predicted_peaks": [
           "Kyajo Ri",
           "Makalu",
           "Baruntse",
           "Cho Oyu",
           "Khumbila"
         ],
         "gps": {
           "lat": 27.975558333333332,
           "lon": 86.82515
         },
         "shot_time": "2023-11-03T08:51:08"
       }
     }
}<p><strong>画像による検索 (画像間):</strong>この機能を使用すると、写真内の山を識別し、写真アルバム内で同じ山の他の画像を見つけることができます。画像がアップロードされると、SigLIP-2 画像エンコーダーによって処理され、<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/identify_from_picture_find_similar_peaks.py#L228">画像ベクトル</a>が生成されます。次に、 <code>peaks_catalog.text_embed</code>フィールドで<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/identify_from_picture_find_similar_peaks.py#L234">kNN 検索を</a>実行し、最も一致するピーク名を特定します。次に、これらの一致するピーク名から<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/identify_from_picture_find_similar_peaks.py#L257">テキスト ベクトルが生成され</a>、写真インデックスに対して別の<a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/scripts/identify_from_picture_find_similar_peaks.py#L263">kNN 検索が</a>実行され、対応する写真が検索されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltab9d16333e2a9e69/6a17da827f6f155448c099cc/3a3d5635bee7a222b95529dd7f9fbee016381610-1226x550.png" alt="Elasticsearch で画像によるマルチモーダル検索 (画像間検索) がどのように機能するかを説明します。" /><p><strong>Elasticsearchクエリ:</strong></p><p>ステップ1: 一致するピーク名を見つける</p>GET peaks_catalog/_search
{
 "knn": {
   "field": "text_embed",
   "query_vector": [...image-vector... ],
   "k": 3,
   "num_candidates": 500
 },
 "_source": [
   "id",
   "names",
   "latlon",
   "text_embed"
 ]
}


Response (first two documents):
{
 "took": 2,
 "timed_out": false,
 "_shards": {
   "total": 1,
   "successful": 1,
   "skipped": 0,
   "failed": 0
 },
 "hits": {
   "total": {
     "value": 3,
     "relation": "eq"
   },
   "max_score": 0.58039916,
   "hits": [
     {
       "_index": "peaks_catalog",
       "_id": "pumori",
       "_score": 0.58039916,
       "_source": {
         "id": "pumori",
         "names": [
           "Pumori",
           "Pumo Ri"
         ],
         "latlon": {
           "lat": 28.01472,
           "lon": 86.82806
         },
         "text_embed": [
                  ... embeddings...
         ]
       }
     },
     {
       "_index": "peaks_catalog",
       "_id": "kyajo-ri",
       "_score": 0.57942784,
       "_source": {
         "id": "kyajo-ri",
         "names": [
           "Kyajo Ri",
           "Kyazo Ri"
         ],
         "latlon": {
           "lat": 27.909167,
           "lon": 86.673611
         },
         "text_embed": [
           ... embeddings...
         ]
       }
     }
   ]
 }
}<p>ステップ 2: <code>photos</code>インデックスで検索を実行し、一致する画像を見つけます (テキストから画像への検索ユース ケースに示されているのと同じクエリ)。</p>POST photos/_search
{
 "knn": {
   "field": "clip_image",
   "query_vector": [ ...image-vector... ],
   "k": 30,
   "num_candidates": 2000
 },
 "_source": [
   "path",
   "gps",
   "shot_time",
   "predicted_peaks",
   "clip_image"
 ],
 "query": {
   "bool": {
     "filter": [
       {
         "term": {
           "predicted_peaks": "Pumori"
         }
       }
     ]
   }
 }
}


Response (first two documents):
{
 "hits": {
   "total": {
     "value": 56,
     "relation": "eq"
   },
   "max_score": 0.5779596,
   "hits": [
     {
       "_index": "photos",
       "_id": "d01da3a1141981486c3493f6053c79e92a788463",
       "_score": 0.5779596,
       "_source": {
         "path": "IMG_2738.HEIC",
         "predicted_peaks": [
           "Pumori",
           "Kyajo Ri",
           "Khumbila",
           "Nangkartshang",
           "Kongde Ri"
         ],
         "gps": {
           "lat": 27.97116388888889,
           "lon": 86.82331111111111
         },
         "shot_time": "2023-11-03T08:07:13"
       }
     },
     {
       "_index": "photos",
       "_id": "c79d251f07adc5efaedc53561110a7fd78e23914",
       "_score": 0.5766071,
       "_source": {
         "path": "IMG_2761.HEIC",
         "predicted_peaks": [
           "Kyajo Ri",
           "Makalu",
           "Baruntse",
           "Cho Oyu",
           "Khumbila"
         ],
         "gps": {
           "lat": 27.975558333333332,
           "lon": 86.82515
         },
         "shot_time": "2023-11-03T08:51:08"
       }
     }
}<h2>流線型のUI</h2><p>すべてをまとめるために、両方の検索ユースケースを実行できるシンプルな Streamlit UI を作成しました。左側のレールには、チェックボックスとミニマップ/ジオフィルターが付いた、スクロール可能なピークのリスト（ <code>photos.predicted_peaks</code>から集約）が表示されます。上部には、<strong>名前による検索</strong>ボックスと<strong>写真アップロードからの識別</strong>ボタンがあります。中央のペインには、kNN スコア、予測ピーク バッジ、キャプチャ時間を表示するレスポンシブなサムネイル グリッドがあります。各画像には、フル解像度のプレビューを表示するための<strong>画像表示</strong>ボタンが含まれています。</p><p><strong>画像をアップロードして検索:</strong>ピークを予測し、写真アルバムから一致するピークを見つけます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd1fb2b0304a310d2/6a17da8425daab7cda08a0fa/dca540cbf5279e6d6102c5a0c0351ddd4ac91cda-1600x1112.png" alt="アマダブラム山の山頂をテキストから画像、画像から画像の両方でマルチモーダル検索できる、シンプルで合理化された UI。" /><p><strong>テキスト検索</strong>: アルバム内の一致するピークをテキストから検索します</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt496c1ae8f7886320/6a17da86abe0f2da48dfe8bd/b1e8618db746cd49ea4962d3dc73031387b975dd-1600x1166.png" alt="山頂ライブラリでテキスト検索を使用してエベレスト山の山頂を検索する方法。" /><h2>まとめ</h2><p><em><strong>アマ・ダブラムの</strong></em> <em>写真</em><em> だけ見せてもらえませんか ？ というところから始まりました。</em>小規模で実用的な<strong>マルチモーダル検索</strong>システムになりました。私たちは生のトレッキング写真を撮影し、それを<strong>SigLIP-2 埋め込み</strong>に変換し、 <strong>Elasticsearch</strong>を使用してベクトルに対して高速<strong>kNN を</strong>実行し、さらに単純な地理/時間フィルターを使用して<em>意味</em>によって適切な画像を浮かび上がらせました。途中で、私たちは 2 つのインデックス、つまり混合プロトタイプの小さな<code>peaks_catalog</code> (識別用) と、画像ベクトルと EXIF のスケーラブルな<code>photos</code>インデックス (検索用) で関心を分離しました。実用的かつ再現性があり、拡張も簡単です。</p><p>調整したい場合は、いくつかの設定を試すことができます。</p><ul><li><p><strong>クエリ時間の設定:</strong> <code>k</code> (返す近隣の数) と<code>num_candidates</code> (最終スコアリングの前に検索する範囲)。これらの設定については、<a href="https://www.elastic.co/search-labs/blog/elasticsearch-knn-and-num-candidates-strategies">こちらの</a>ブログで説明されています。</p></li><li><p><strong>インデックス時間の設定:</strong> <code>m</code> (グラフの接続性) および<code>ef_construction</code> (ビルド時間の精度とメモリ)。クエリの場合は、 <code>ef_search</code>も試してください。値が大きいほど、通常は、レイテンシを多少トレードオフして、リコール率が向上します。これらの設定の詳細については、<a href="https://www.elastic.co/search-labs/blog/hnsw-graph">このブログ</a>を参照してください。</p></li></ul><p>今後、<strong>マルチモーダル</strong>および<strong>多言語</strong>検索用のネイティブモデル/リランカーがElasticエコシステムにまもなく導入される予定です。これにより、画像/テキスト検索とハイブリッドランキングがさらに強化されるはずです<a href="https://ir.elastic.co/news/news-details/2025/Elastic-Completes-Acquisition-of-Jina-AI-a-Leader-in-Frontier-Models-for-Multimodal-and-Multilingual-Search/default.aspx?utm_source=chatgpt.com">。ir.elastic.co+1</a></p><p>これを自分で試してみたい場合は:</p><ul><li><p><strong>GitHub リポジトリ:</strong> <a href="https://github.com/navneet83/multimodal-mountain-peak-search"><em>https://github.com/navneet83/multimodal-mountain-peak-search</em></a></p></li><li><p><strong>Colab クイックスタート:</strong> <a href="https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/notebooks/multimodal_mountain_peak_search.ipynb">https://github.com/navneet83/multimodal-mountain-peak-search/blob/main/notebooks/multimodal_mountain_peak_search.ipynb</a></p></li></ul><p>これで私たちの旅は終わり、帰る時間になりました。これが役に立つことを願っています。これを壊した場合（または改善した場合）、何を変更したかをお聞かせください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdce2fff1569d2a8b/6a17da894b055dd1f24320a2/d324d1e1472f1bfbd8f25747f57bdeeb9c7f16b2-1600x1200.png" alt="" />]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/multimodal-search-siglip-2-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/multimodal-search-siglip-2-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[ハイブリッド検索]]></category>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Navneet Kumar]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltccb66279debb05f9/6a17da8b63baffe228741b15/ffcf93358a7c5dadcea82faf3de460bf060d003c-1600x1200.png" length="0" type="image/png"/>
    <pubDate>Tue, 04 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic MCP サーバー: あらゆる AI エージェントに Agent Builder ツールを公開]]></title>
    <description><![CDATA[Agent Builder に組み込まれている Elastic MCP サーバーを使用して、プライベート データやカスタム ツールにアクセスできる AI エージェントを安全に拡張する方法を説明します。]]></description>
    <content:encoded><![CDATA[<p>Elastic Agent Builder は、Elasticsearch 内の独自のデータと深く統合されたツールとエージェントを作成するためのプラットフォームです。たとえば、内部ドキュメントに対してセマンティック検索を実行したり、観測ログを分析したり、セキュリティアラートを照会したりするツールを作成できます。</p><p>しかし、本当の魔法は、これらのカスタマイズされたデータ対応ツールを、ほとんどの時間を費やす環境に導入できたときに起こります。コード エディター エージェントが組織のプライベート ナレッジ ベースに安全にアクセスできたらどうなるでしょうか?</p><p>ここで、<strong>モデル コンテキスト プロトコル (MCP)</strong>が登場します。Elastic Agent Builder には、プラットフォーム内のツールへのアクセスを提供する組み込みの MCP サーバーが付属しています。</p><h2>Elastic Agent Builder MCP サーバーを使用する理由は何ですか?</h2><p>AI エージェントは非常に強力ですが、その知識は通常、トレーニングに使用されたデータとパブリック インターネット上でアクティブに検索できる情報に限定されます。彼らは、会社の内部設計ドキュメント、チーム固有のデプロイメント ランブック、またはアプリケーション ログの独自の構造については知りません。</p><p>課題は、AI アシスタントに必要な特殊なコンテキストを提供することです。これはまさに、MCP が解決するために設計された問題です。<strong>MCP は、AI モデルまたはエージェントが外部ツールを検出して使用できるようにするオープン スタンダードです。</strong></p><p>これを実現するために、Elastic Agent Builder は組み込みの MCP サーバーを通じてカスタム ツールをネイティブに公開します。つまり、 <strong>Cursor</strong> 、 <strong>VS Code</strong> 、 <strong>Claude Desktop</strong>などの MCP 対応クライアントを、Elastic Agent Builder で構築した特殊なデータ対応ツールに簡単に接続できるということです。</p><h2>MCP を使用する場合 (および使用しない場合)</h2><p>Elastic Agent Builder には、さまざまな統合パターンをサポートするためのいくつかのプロトコルが含まれています。適切なものを選択することが、効果的な AI ワークフローを構築する鍵となります。</p><ul><li><p><a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server"><strong>MCP を使用して</strong></a> 、専用のツールで AI エージェント (<strong> Cursor</strong> や<strong> VS Code</strong> など) を拡張します。これは「独自のツールを持ち込む」アプローチであり、すでに使用しているアシスタントを強化して、プライベート データに安全にアクセスできるようにします。MCP サーバーを通じて公開されるのはツールのみで、Elastic のエージェントはそれとは別です。</p></li><li><p><a href="https://www.elastic.co/docs/solutions/search/agent-builder/a2a-server"><strong>A2A プロトコルを</strong></a><strong> 使用すると</strong> 、完全なカスタム Elastic Agent が他の自律エージェント (<a href="https://www.elastic.co/search-labs/blog/a2a-protocol-elastic-agent-builder-gemini-enterprise"><strong> Google の Gemini Enterprise</strong></a> など) と連携できるようになります。これはエージェント間の委任用であり、各エージェントは問題を解決するためにピアとして機能します。</p></li><li><p><a href="https://www.elastic.co/docs/solutions/search/agent-builder/kibana-api"><strong>カスタム</strong></a> アプリケーションを最初から構築するときに、完全なプログラム制御を行うには Agent Builder API を 使用します 。</p></li></ul><p>IDE を離れずに社内ドキュメントから回答を得たい開発者にとって、MCP は最適です。</p><h2>例: Agent Builder MCP サーバーを使用した Cursor のカスタム ツール</h2><p>私が日常的に使用している実際の例を見てみましょう。まず、社内のエンジニアリング ドキュメントをクロールして、 <code>elastic-dev-docs</code>という Elasticsearch インデックスにインデックス付けしました。Agent Builder で使用できる汎用の組み込みツールを使用することもできますが、この特定のナレッジベースを照会するための独自のカスタム ツールを作成します。</p><p>カスタム ツールを構築する理由はシンプルです。<strong>制御と精度です</strong>。このアプローチにより、 <code>elastic-dev-docs</code>インデックスに対して高速でセマンティックなクエリを直接実行できるようになります。どのインデックスをターゲットにするか、データをどのように取得するかを完全に制御できます。</p><p>ここで、このカスタム ナレッジ ベースを Cursor のような AI 搭載コード エディターで使用する方法を説明します。</p><h3>ステップ1: Agent Builderでカスタムナレッジベースツールを作成する</h3><p>まず、Agent Builder で新しいツールを作成します。明確で具体的なツールの説明は重要です。なぜなら、それが内部 Elastic Agent であれ、MCP 経由で接続する Cursor などの外部ツールであれ、あらゆる AI エージェントが適切なタスクのためにツールを検出し選択する方法だからです。</p><p>強力な説明は明確である必要があります。たとえば、「elastic-dev-docs インデックスでセマンティック検索を実行して、社内のエンジニアリング ドキュメント、ランブック、リリース手順を検索します。」</p><p>これで、ツールは特定のインデックスに対してセマンティック検索を実行するように構成されます。保存すると、すぐに利用できるようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt011118f0a9279185/6a17f367dbb4ffc4f3fb581a/1eea079908fdf7cc72dbe81abd07ff51601a43d4-1472x1600.png" alt="Agent Builder でカスタム ナレッジ ベース ツールを作成します。" /><p>外部に接続する前に、UI で直接テストできます。<strong>[テスト]</strong>ボタンをクリックするだけで、パラメータを手動で入力し、LLM の動作をエミュレートして、結果を検査し、すべてが正しく動作していることを確認します。</p><h3>ステップ2: CursorをElastic MCPサーバーに接続する</h3><p>Elastic Agent Builder は、安全な MCP エンドポイントを介して利用可能なすべてのツールを自動的に公開します。固有のサーバー URL は、Kibana 内のツール UI で見つけることができます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdd0e62ae0f394c3d/6a17f368e317916ec32d5933/ba137be30f0eaa7f028b96bd8af4e2779c3f8a33-1600x589.png" alt="Kibana のツール UI のカーソルを Elastic MCP サーバーに接続する方法。" /><p>Cursor に接続するには、この URL と認証用の Elastic API キー ( <a href="https://www.elastic.co/docs/deploy-manage/api-keys/elasticsearch-api-keys">ES API キーの作成方法を参照</a>) を構成ファイルに追加するだけです。認証には API キーを使用します。これにより、すべてのアクセス制御ルールを尊重し、ツールは付与した権限でのみ実行されるようになります。</p><p>カーソルの<code>~/.cursor/mcp.json</code>内の MCP 構成は次のようになります。</p>{
  "mcpServers": {
    "elastic-agent-builder": {
      "command": "npx",
      "args": [
        "mcp-remote",
        "https://your-kibana.kb.company.io/api/agent_builder/mcp",
        "--header",
        "Authorization:${AUTH_HEADER}"
      ],
      "env": {
        "AUTH_HEADER": "ApiKey &lt;ELASTIC_API_KEY&gt;"
      }
    }
  }
}<p>設定が保存されると、Cursor で Elastic Agent Builder MCP サーバー ツールが利用可能になります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2837638263e628ed/6a17f36adbb4ffeb9cfb5820/d302c6d3609fbf14fd40e21b9e69e567bf12553f-1600x1002.png" alt="Cursor で使用できる Elastic Agent Builder MCP サーバー ツールのイメージ。" /><h3>ステップ 3: どんどん質問しましょう!</h3><p>接続が確立されると、カーソル エージェントはカスタム ツールを呼び出して質問に答えたり、コード生成プロセスをガイドしたりできるようになります。</p><p>具体的な質問をしてみましょう。</p><p><em>「Elastic Search org のエンジニアリング内部ドキュメントからクローラー サービスをリリースするための手順を参照する」</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt83fa357261b30e93/6a17f36c4b055d16d1432326/14f572730203c23615bb9dd38234bcb3b0f81155-1600x1468.png" alt="カスタム ツールを呼び出して質問に答え、コード生成プロセスをガイドするカーソル エージェント。" /><p>舞台裏では魔法が起こります:</p><ol><li><p>カーソルエージェントはあなたの質問に最もよく答える方法を決定し、 <code>engineering_documentation_internal_search</code></p></li><li><p>自然言語クエリでツールを呼び出す</p></li><li><p>このツールは、 <code>elastic-dev-docs</code>インデックスに対してセマンティック検索を実行し、最も関連性の高い最新の手順を返します。</p></li></ol><p>コード エディターを離れることなく、社内ドキュメントに基づいた正確で信頼できる回答が得られます。体験はシームレスかつ強力です。</p><h2>あなたの番です</h2><p>ここでは、Elastic Agent Builder に組み込まれている MCP サーバーを使用して、プライベート データへの安全なアクセスを備えた AI アシスタントを拡張する方法を説明しました。モデルを本当に役立つものにするためには、独自の情報に基づいてモデルを構築することが鍵となります。</p><p>要約すると、主要な手順について説明しました。</p><ul><li><p>ニーズに合った適切なプロトコルを選択する (MCP)。</p></li><li><p>カスタム ナレッジ ベース ツールを構築します。</p></li><li><p>そのツールを Cursor などの IDE アシスタントに接続します。</p></li></ul><p>エージェントとツールを最も重要なコンテキストから切り離す必要がなくなりました。このガイドがより効果的でデータを考慮したワークフローの作成に役立つことを願っています。楽しい建築を！</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-mcp-server-agent-builder-tools</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-mcp-server-agent-builder-tools</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[AIツール ]]></category>
    <dc:creator><![CDATA[Jedr Blaszyk,Joe McElroy]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta5b61961b6269ab1/6a17f36ea29299d839d02db2/ef5153551a1d14833c7f512fede554d1dfb31553-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Mon, 20 Oct 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[AIエージェントの評価：Elasticによるエージェントフレームワークのテスト方法]]></title>
    <description><![CDATA[正確で検証可能な結果を確保するために、エージェントシステムへの変更を Elastic ユーザーにリリースする前に評価およびテストする方法を学びます。]]></description>
    <content:encoded><![CDATA[<h2>はじめに</h2><p>Elastic Stack には、 <a href="https://www.elastic.co/search-labs/blog/ai-agentic-workflows-elastic-ai-agent-builder">Agent Builder</a>の近々リリースされる Elastic AI Agent (現在技術プレビュー) や<a href="https://www.elastic.co/docs/solutions/security/ai/attack-discovery">Attack Discovery</a> (8.18 および 9.0 以降で<a href="https://www.elastic.co/blog/whats-new-elastic-security-9-0-0">GA</a>提供) など、LLM を利用したエージェント アプリケーションが多数あり、さらに多くのアプリケーションが開発中です。開発中、そして展開後でも、次の質問に答えることが重要です。</p><ul><li><p>これらの AI アプリケーションの応答の品質をどのように評価するのでしょうか?</p></li><li><p>変更を加えた場合、その変更が本当に改善となり、ユーザー エクスペリエンスが低下しないことをどのように保証すればよいでしょうか。</p></li><li><p>これらの結果を繰り返し簡単にテストするにはどうすればよいでしょうか?</p></li></ul><p>従来のソフトウェア テストとは異なり、生成 AI アプリケーションの評価には、統計的手法、微妙な定性的なレビュー、ユーザーの目標の深い理解が必要になります。</p><p>この記事では、Elastic 開発チームが評価を実施し、展開前に変更の品質を確保し、システム パフォーマンスを監視するために採用しているプロセスについて詳しく説明します。私たちは、あらゆる変更が証拠によって裏付けられ、信頼できる検証可能な結果につながるようにすることを目指しています。このプロセスの一部は Kibana に直接統合されており、オープンソース精神の一環として透明性への取り組みを反映しています。評価データと指標の一部を公開することで、コミュニティの信頼を育み、AI エージェントを開発したり当社の製品を利用したりするすべての人にとって明確なフレームワークを提供することを目指しています。</p><h2>製品例</h2><p>このドキュメントで使用した方法は、Attack Discovery や Elastic AI Agent などのソリューションを反復して改善する方法の基礎となりました。それぞれ2つの簡単な紹介:</p><h3>Elastic Securityの攻撃検出</h3><p>Attack Discovery は LLM を使用して、Elastic 内の攻撃シーケンスを識別および要約します。特定の期間（デフォルトでは 24 時間）内の Elastic Security アラートに基づいて、Attack Discovery のエージェント ワークフローは、攻撃が発生したかどうかを自動的に検出するほか、どのホストまたはユーザーが侵害されたか、どのアラートが結論に寄与したかなどの重要な情報も検出します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb70932abe8d4de75/6a17f04ea292990c52d02d61/20fabb47642dad7b588daaaa8c3a98de860ad01d-1251x758.png" alt="" /><p></p><p>目標は、LLM ベースのソリューションが少なくとも人間と同等の出力を生成することです。</p><h3>エラスティックAIエージェント</h3><p><strong>Elastic Agent Builder は、</strong>すべての検索機能を活用するコンテキスト認識型 AI エージェントを構築するための新しいプラットフォームです。この製品には、会話形式のやりとりを通じてユーザーがデータを理解し、データから回答を得られるよう設計された、あらかじめ構築された汎用エージェントである<strong>Elastic AI Agent</strong>が付属しています。</p><p>エージェントは、Elasticsearch または接続されたナレッジベース内の関連情報を自動的に識別し、事前に構築された一連のツールを活用してそれらと対話することでこれを実現します。これにより、Elastic AI Agent は、単一のドキュメントに関する単純な Q&amp;A から、複数のインデックスにわたる集約や単一または複数ステップの検索を必要とする複雑なリクエストまで、さまざまなユーザー クエリに応答できるようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3b9dbede85a56bd6/6a17f050e8fbce88943a1a30/d29dee100bb8a17bb623acd745773a5164a1df4f-1600x1014.png" alt="" /><h2>実験による改善の測定</h2><p>AI エージェントのコンテキストでは、実験とは、明確に定義された次元 (有用性、正確性、遅延など) のパフォーマンスを向上させるように設計された、システムに対する構造化されたテスト可能な変更です。目標は、「この変更をマージした場合、それが真の改善であり、ユーザー エクスペリエンスを低下させないことを保証できますか?」という質問に明確に答えることです。</p><p>私たちが実施するほとんどの実験には、一般的に次のようなものが含まれます。</p><ul><li><p><strong>仮説:</strong>特定の、反証可能な主張。<em>例:</em> 「攻撃検出ツールへのアクセスを追加すると、セキュリティ関連のクエリの正確性が向上します。」</p></li><li><p><strong>成功基準:</strong> 「成功」の意味を定義する明確なしきい値。<em>例:</em> 「セキュリティ データセットの正確性スコアが 5% 向上し、他の部分では低下は見られません。」</p></li><li><p><strong>評価計画:</strong>成功の測定方法 (指標、データセット、比較方法)</p></li></ul><p>成功した実験は体系的な調査プロセスです。小さなプロンプトの調整から大規模なアーキテクチャの変更まで、すべての変更は次の 7 つの手順に従い、結果が有意義かつ実用的なものになるようにします。</p><ul><li><p>手順1：問題を特定する</p></li><li><p>ステップ2: 指標を定義する</p></li><li><p>ステップ3：明確な仮説を立てる</p></li><li><p>ステップ4: 評価データセットの準備</p></li><li><p>ステップ5: 実験を実行する</p></li><li><p>ステップ6: 結果の分析と反復</p></li><li><p>ステップ7：決定を下し、文書化する</p></li></ul><p>これらのステップの例を<em>図 1</em>に示します。次のサブセクションでは各ステップについて説明します。各ステップの技術的な詳細については、今後のドキュメントで詳しく説明します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt06bfe2f0e4205a18/6a17f052faa91358eb93c968/3a9f5a3e92dd4922a795a19104c6e4ad8c98958d-2400x1352.png" alt="" /><h2>実際の Elastic の例を使ったステップバイステップのウォークスルー</h2><h3>手順1：問題を特定する</h3><p><em>この変更が解決しようとしている問題は正確には何でしょうか?</em></p><p>攻撃検出の例: 概要が不完全な場合や、無害なアクティビティが誤って攻撃としてフラグ付けされる (誤検知) 場合があります。</p><p>Elastic AI エージェントの例: 特に分析クエリの場合、エージェントのツール選択は最適ではなく一貫性がなく、間違ったツールが選択されてしまうことがよくあります。これにより、トークンのコストとレイテンシが増加します。</p><h3>ステップ2: 指標を定義する</h3><p><em>問題を測定可能にして、変化を現在の状態と比較できるようにします。</em></p><p>一般的な指標には、<a href="https://developers.google.com/machine-learning/crash-course/classification/accuracy-precision-recall">精度と再現率</a>、<a href="https://en.wikipedia.org/wiki/Semantic_similarity">意味的類似性</a>、事実性などがあります。ユースケースに応じて、一致するアラート ID や正しく取得された URL などのメトリックを計算するためにコード チェックを使用したり、より自由形式の回答を得るために LLM-as-judge などの手法を使用したりします。</p><p>以下は、実験で使用されたメトリックの例です (<em>網羅的ではありません</em>)。</p><p><strong>攻撃の検出</strong></p><p>メトリック</p><p>説明</p><p>精度と再現率</p><p>実際の出力と予想される出力の間でアラート ID を一致させて、検出精度を測定します。</p><p>類似性</p><p>BERTScore を使用して、応答テキストの意味的類似性を比較します。</p><p>事実性</p><p>重要な IOC (侵害の兆候) は存在しますか?MITRE 戦術 (攻撃の業界分類) は正しく反映されていますか?</p><p>攻撃チェーンの一貫性</p><p>発見された数を比較して、攻撃の過剰報告または過少報告がないか確認します。</p><p><strong>エラスティックAIエージェント</strong></p><p>メトリック</p><p>説明</p><p>精度と再現率</p><p>ユーザーのクエリに回答するためにエージェントによって取得されたドキュメント/情報と、クエリに回答するために必要な実際の情報またはドキュメントを照合して、情報取得の精度を測定します。</p><p>事実性</p><p>ユーザーのクエリに回答するために必要な主要な事実は存在しますか?事実は手続き上のクエリに対して正しい順序になっていますか?</p><p>回答の関連性</p><p>応答には、ユーザーのクエリとは関連がない、または周辺的な情報が含まれていますか?</p><p>応答の完全性</p><p>応答はユーザークエリのすべての部分に答えていますか?応答にはグラウンドトゥルースに存在するすべての情報が含まれていますか?</p><p>ES|QL検証</p><p>生成された ES|QL は構文的に正しいですか?機能的にはグラウンドトゥルース ES|QL と同一ですか?</p><h3>ステップ3：明確な仮説を立てる</h3><p><em>上記で定義した問題と指標を使用して、明確な成功基準を確立します。</em></p><p>Elastic AI エージェントの例:</p><ol><li><p><strong>relevance_search および nl_search ツールの説明に変更を加え、それぞれの機能と使用例を明確に定義します</strong>。</p></li><li><p><strong>ツールの呼び出し精度が</strong> <strong>25% 向上 する</strong> と予測しています。</p></li><li><p>他の指標に悪影響が及ばないことを保証し、これが純粋にプラスであることを確認します。<strong>事実性と完全性</strong>。</p></li><li><p><strong>正確なツールの説明により、エージェントがさまざまなクエリタイプに最も適した検索ツールをより正確に選択して適用できるようになり、誤った適用が減り、全体的な検索の有効性が向上するため、この方法が効果的であると考えています</strong>。</p></li></ol><h3>ステップ4: 評価データセットの準備</h3><p><em>システムのパフォーマンスを測定するために、現実世界のシナリオをキャプチャしたデータセットを使用します。</em></p><p>実施する評価の種類に応じて、LLMに供給される生データ（例：攻撃検出のための攻撃シナリオと予想される出力。アプリケーションがチャットボットの場合、入力はユーザークエリであり、出力は正しいチャットボット応答、取得されるべき正しいリンクなどになります。</p><p>攻撃検出の例:</p><p>10の斬新な攻撃シナリオ</p><p>Oh My Malware のエピソード 8 つ (ohmymalware.com)</p><p>4 つのマルチ攻撃シナリオ (最初の 2 つのカテゴリの攻撃を組み合わせて作成)</p><p>3つの良性のシナリオ</p><p>Elastic AI エージェント評価データセットの例 ( <a href="https://github.com/elastic/kibana/blob/main/x-pack/platform/packages/shared/onechat/kbn-evals-suite-onechat/evals/kb/kb.spec.ts">Kibana データセット リンク</a>):</p><p>オープンソース データセットを使用して KB 内の複数のソースをシミュレートする 14 のインデックス。</p><p>5 つのクエリ タイプ (分析、テキスト検索、ハイブリッドなど)</p><p>7 つのクエリ意図タイプ（手続き型、事実型 - 分類型、調査型など）</p><h3>ステップ5: 実験を実行する</h3><p>評価データセットに対して既存のエージェントと修正バージョンの両方からの応答を生成して実験を実行します。事実性などの指標を計算します (手順 2 を参照)。</p><p>ステップ 2 で必要な指標に基づいて、さまざまな評価を組み合わせます。</p><ul><li><p>ルールベースの評価（例：Python/TypeScriptを使用して.jsonが有効かどうかを確認します)</p></li><li><p>LLM が裁判官となる（回答が原文と事実上一致しているかどうかを別の LLM に尋ねる）</p></li><li><p>ニュアンス品質チェックのための人間によるレビュー</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt17ec63af0850d8dd/6a17f054505ac3e508ad8c1e/8648e75818d3291f0ac66f069438a500d42b8225-1600x1099.png" alt="これは、当社の内部フレームワークによって生成された評価結果の例です。さまざまなデータセットにわたって実施された実験からのさまざまなメトリックを示します。" /><h3>ステップ6: 結果の分析と反復</h3><p>指標が得られたので、結果を分析します。<u><em>結果がステップ 3 で定義された成功基準を満たしている場合でも、変更を本番環境にマージする前に人間によるレビューが行われます</em></u>。結果が基準を満たしていない場合は、問題を反復して修正してから、新しい変更に対して評価を実行します。</p><p>マージする前に、最適な変更を見つけるために数回の反復が必要になると予想されます。コミットをプッシュする前にローカル ソフトウェア テストを実行するのと同様に、オフライン評価はローカルの変更または複数の提案された変更で実行できます。分析を効率化するために、実験結果、複合スコア、視覚化の保存を自動化すると便利です。</p><h3>ステップ7：決定を下し、文書化する</h3><p>意思決定フレームワークと受け入れ基準に基づいて、変更のマージを決定し、実験を文書化します。意思決定は多面的であり、他のデータセットでの回帰シナリオの確認や、提案された変更の費用対効果の検討など、評価データセット以外の要素を考慮する場合があります。</p><p>例: いくつかの反復をテストして比較した後、最高スコアの変更を選択し、製品マネージャーやその他の関連する関係者に送信して承認を得ます。意思決定を支援するために、前の手順の結果を添付します。攻撃検出に関するその他の例については、 <a href="https://www.elastic.co/blog/elastic-security-generative-ai-features">「Elastic Security の生成 AI 機能の舞台裏」を</a>ご覧ください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt62a466f3a0da114a/6a17f056faa91342c393c96c/74c80b8f34dce8ddd20873ecb2f553873587ed35-1600x618.png" alt="" /><h2>まとめ</h2><p>このブログでは、実験ワークフローのエンドツーエンドのプロセスについて説明し、エージェントシステムの変更を Elastic ユーザーにリリースする前に評価およびテストする方法を説明しました。また、Elastic でのエージェントベースのワークフローの改善例もいくつか紹介しました。今後のブログ投稿では、適切なデータセットを作成する方法、信頼性の高いメトリックを設計する方法、複数のメトリックが関係する場合に意思決定を行う方法など、さまざまな手順の詳細を詳しく説明します。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-agent-evaluation-elastic</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-agent-evaluation-elastic</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Susan Chang,Abhimanyu Anand]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte578b636637be6b1/6a17f057e8fbcebe9e3a1a36/ef3922076713872163e1aab47735361513b2c9ee-2400x1352.heif" length="0" type="image/*"/>
    <pubDate>Mon, 13 Oct 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[A2Aプロトコルを介してElastic AgentsをGemini Enterpriseに接続する]]></title>
    <description><![CDATA[Agent Builder を使用して、A2A プロトコルを使用してカスタム Elastic Agent を Gemini Enterprise などの外部サービスに公開する方法を学習します。]]></description>
    <content:encoded><![CDATA[<p><strong>Elastic Agent Builder は、</strong> Elasticsearch で直接データ駆動型の AI エージェントを作成するための機能セットです。この<a href="https://www.elastic.co/search-labs/blog/series/context-aware-ai-agentic-workflows-with-elastic">シリーズ</a>の以前の投稿では、カスタム エージェントに複雑なタスクを実行するツールを装備し、エージェントの動作をガイドする一連のカスタム指示を提供する方法を説明しました。</p><p>しかし、すでに使用しているアプリケーションや生産性ツールでカスタムエージェントを使用したい場合はどうすればよいでしょうか?</p><p>ここで、<strong>エージェント間 (A2A) プロトコルが</strong>登場します。A2A は相互運用性のための<a href="https://github.com/a2aproject/A2A">オープン スタンダード</a>であり、異なるプラットフォームのエージェント間の通信と共同作業を可能にします。そして、これを Elastic Agent Builder に直接組み込みました。</p><p>今日は、構築したカスタム エージェントを他のサービス、具体的には<strong>Gemini Enterprise</strong> (旧称 Agentspace) に公開する方法を紹介します。</p><h2>オープンスタンダードの力：A2Aが重要な理由</h2><p>ブログ記事<a href="https://www.elastic.co/search-labs/blog/ai-agent-builder-elasticsearch">「初めての Elastic Agent」</a>では、市場データに安全にアクセスできる<em>Financial Assistant</em>エージェントなどのカスタムエージェントの構築方法を説明しました。しかし、作業を再構築せずに、Gemini Enterprise などの他の環境でその洞察を利用できない場合、その価値は限られます。</p><p>この相互運用性の課題が、エージェント AI の実現を妨げているのです。エージェントはプラットフォーム間で通信するために共通言語を必要としますが、これがまさに A2A プロトコルの役割です。標準の通信レイヤーを提供することで、エージェントと直接対話できるだけでなく、組織全体の専門エージェントが連携して洞察を共有できる未来が開かれます。</p><p>これを実現するために、Elastic Agent Builder は、すべてのエージェントに対して 2 つの標準エンドポイントを通じて A2A プロトコルをネイティブにサポートしています。</p><ol><li><p><strong>エージェント カード エンドポイント (</strong> <strong><code>GET {your-kibana-url}/api/agent_builder/a2a/{agentId}.json</code></strong> <strong>) -</strong>これはカスタム エージェントの名刺として機能します。エージェントに関するメタデータ (名前、説明、機能など) を A2A 互換サービスに提供します。</p></li><li><p><strong>A2A プロトコル エンドポイント (</strong> <strong><code>POST {your-kibana-url}/api/agent_builder/a2a/{agentId}</code></strong> <strong>)</strong> - これは通信チャネルです。他のエージェントはここにリクエストを送信し、エージェントはそれを処理して応答を返します。これらはすべて<a href="https://a2a-protocol.org/latest/specification/">A2A プロトコル仕様</a>に従って行われます。</p></li></ol><h2>A2Aインスペクターでエージェントをテストする</h2><p>エージェントを本番システムに接続する前に、正しく通信していることを確認することをお勧めします。これを行う最も簡単な方法は、A2A 統合のテストとデバッグ専用に設計されたツールである<strong>A2A Inspector を</strong>使用することです。</p><p>インスペクターを実行するのは簡単です。<a href="https://github.com/a2aproject/a2a-inspector">a2a-inspector</a>リポジトリのクローンを作成し、README の指示に従って<a href="https://github.com/a2aproject/a2a-inspector?tab=readme-ov-file#3-run-the-application">アプリケーションを実行でき</a>ます。起動すると、UI はデフォルトで<code>http://localhost:5001/</code>で使用できるようになります。</p><p>A2A インスペクターをエージェントに接続するには、次の 2 つの重要な情報を提供する必要があります。</p><ul><li><p>エージェント カード URL: これはエージェントを説明するエンドポイントです。<a href="https://www.elastic.co/search-labs/blog/ai-agent-builder-elasticsearch">前回の投稿の Financial Assistant エージェント</a>の場合、この URL は<code>{your-kibana-url}/api/agent_builder/a2a/financial_assistant.json</code>になります。</p></li><li><p>認証ヘッダー: 認証には標準の API キーを使用します。</p></li></ul><p>インスペクターの UI にこれらの詳細を入力すると、すぐにエージェントに接続してチャットを開始できます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6381135e3fb297df/6a17ef4bec0f898b0c5a66ea/7231c72bf30bed2a854f58658c1eca2843f43bfc-1600x1296.png" alt="A2Aエージェントカードとエージェントインスペクターの設定" /><p>この簡単な検証により、エージェントが正しく構成され、次のステップの準備ができていることが保証されます。</p><h2>ライブ配信しよう！Gemini Enterpriseのカスタムエージェント</h2><p>次は、エキサイティングな部分です。カスタム ファイナンシャル アドバイザー エージェントを Gemini Enterprise (旧 Agentspace) 内で実現します。この統合は<a href="https://console.cloud.google.com/marketplace/product/elastic-prod/elastic-ai-agent">、Google Cloud Marketplace で入手可能な Elastic AI Agent</a>によって実現されています。</p><p>接続されると、Gemini Enterprise は A2A プロトコルを使用してエージェントと直接通信します。ここで相互運用性の真の威力が発揮されます。ユーザーは使い慣れた環境を離れることなく、カスタム Elasticsearch エージェントから得られる詳細なデータ駆動型の分析情報にアクセスできるようになります。エージェント リストにカスタム Elastic Agent が表示されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7f54f0bb15216d8e/6a17ef4d6df73107d90a0fdb/37a39e92ebf3d72c6c8014397cd8e846336173a4-1600x834.png" alt="Google Agentspace リストでカスタム エージェントを表示する" /><p>Gemini Enterprise のユーザーが次のように質問していると想像してください。</p><p><em>「市場のセンチメントが心配です。悪いニュースによって最もリスクが高い顧客は誰でしょうか？</em> 」</p><p>バックグラウンドでは、Gemini Enterprise がこのクエリを A2A プロトコル経由でカスタム Elastic Agent にルーティングします。エージェントは専用のツールを使用してデータを照会し、回答を作成して返送します。エンドユーザーにとって、エクスペリエンスはシームレスです。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte130c332ee0648a6/6a17ef4fe9ea874426a9c6bb/e5f126c1a27a51c6e69a767aa87c9f746b62e39c-1600x1044.png" alt="ユーザーがAgentspaceにクエリを尋ねると、そのクエリは舞台裏で何が起こるのか" /><p>そして、ここで終わりではありません!Elasticエージェントで取得した回答は、別の専門エージェントをトリガーする可能性のある次の質問のコンテキストとして使用できるようになりました（例：上場企業へのエクスポージャーを調整するには、投資プラットフォーム エージェントにご相談ください。検索バーを離れることなくすべて行えます。</p><p>A2A を搭載した Gemini Enterprise に Elastic エージェントをデプロイすると、ユーザーがデータやツールにコンテキスト内でアクセスできる単一の UI が提供されるため、AI、検索、エンタープライズ システム間の摩擦をなくし、アクセス、オーケストレーション、ワークフローを統合できます。ユーザーにとって、これはツールの切り替えが減り、より直感的で有能な AI アシスタントが利用できるようになることを意味します。組織にとって、これは一貫したガバナンス、スケーラビリティ、相互運用性が組み込まれていることを意味します。</p><h2>あなたの番です</h2><p>これで、Elastic Agent をどこからでも利用できるようにするツールが手に入りました。オープン A2A プロトコルを活用することで、カスタムのデータ対応エージェントの範囲を拡大できます。</p><p>この投稿では、重要な手順について説明しました。</p><ul><li><p>A2A エージェント カードとプロトコル エンドポイントを介してエージェントを公開します。</p></li><li><p>A2A Inspector を使用して接続をテストします。</p></li><li><p>エージェントを Google の Gemini Enterprise などの外部サービスにライブで統合します。</p></li></ul><p>エージェントを分離する必要がなくなりました。皆さんが作り上げる、強力で相互接続されたシステムを見るのが待ちきれません。楽しい建築を！</p><p>始める最も簡単な方法は、 <a href="https://console.cloud.google.com/marketplace/product/elastic-prod/elastic-cloud?pli=1">Google Cloud Marketplace</a>で Elastic Cloud の無料トライアルを利用することです。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/a2a-protocol-elastic-agent-builder-gemini-enterprise</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/a2a-protocol-elastic-agent-builder-gemini-enterprise</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jedr Blaszyk,Valerio Arvizzigno,Joe McElroy]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt63d7675adc5bc211/6a17ef51ddf97d38e8910bdf/5be8a425fab55dca2f9717d2e50812b0450fa625-1440x840.png" length="0" type="image/png"/>
    <pubDate>Thu, 09 Oct 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[初めてのElastic Agent: 単一のクエリからAIを活用したチャットまで]]></title>
    <description><![CDATA[Elastic の AI エージェント ビルダーを使用して特殊な AI エージェントを作成する方法を学びます。このブログでは、金融 AI エージェントを構築します。]]></description>
    <content:encoded><![CDATA[<p>Elastic の新しい<a href="https://www.elastic.co/search-labs/blog/ai-agentic-workflows-elastic-ai-agent-builder">Agent Builder を</a>使用すると、特定のビジネスドメインの専門家として機能する特殊な AI エージェントを作成できます。この機能により、単純なダッシュボードや検索バーを超えて、データを受動的なリソースから能動的な会話のパートナーへと変換できます。</p><p>顧客との会議の前に、状況を把握しておく必要がある財務マネージャーを想像してください。ニュース フィードを手動で調べたり、ポートフォリオ ダッシュボードを相互参照したりする代わりに、カスタム構築されたエージェントに直接質問するだけで済みます。これは「チャットファースト」アプローチの利点です。マネージャーはデータに直接、会話形式でアクセスし、「ACME Corp の最新ニュースは何ですか。また、それがクライアントの保有株にどのような影響を与えますか」などと質問します。数秒以内に専門家による総合的な回答が得られます。</p><p>私たちは現在、金融の専門家を構築していますが、そのアプリケーションはデータと同じくらい多様です。同じ力で、脅威を探すサイバーセキュリティアナリスト、機能停止を診断するサイト信頼性エンジニア、キャンペーンを最適化するマーケティングマネージャーを生み出すこともできます。分野に関係なく、中核となる使命は同じです。データを、チャットできる専門家に変換することです。</p><h2>ステップ0: データセット</h2><p>本日のデータセットは、金融口座、資産状況、ニュース、財務レポートで構成される合成的な金融ベースのデータセットです。これは合成ではありますが、実際の金融データセットの簡略化されたバージョンを複製したものです。</p><p><code>financial_accounts</code>: リスクプロファイル付き顧客ポートフォリオ</p><p><code>financial_holdings</code>: 購入履歴のある株式/ETF/債券のポジション</p><p><code>financial_asset_details</code>: 株式/ETF/債券の詳細</p><p><code>financial_news</code>: 感情分析によるAI生成の市場記事</p><p><code>financial_reports</code>: 企業収益とアナリストのコメント</p><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/your-first-elastic-agent/Your_First_Elastic_Agent.ipynb">ここに</a>ある付属のノートブックに従って、このデータセットを自分でロードできます。</p><h2>ステップ1: 基盤 - ES|QLとしてのビジネスロジック</h2><p>すべての AI スキルは、確かなロジックから始まります。Financial Manager エージェントには、「市場のセンチメントが心配です。」というよくある質問に回答する方法を教える必要があります。悪いニュースによって最もリスクにさらされている顧客は誰なのか教えていただけますか？」この質問は単純な検索の範囲を超えています。市場の感情と顧客のポートフォリオを相関させる必要があります。</p><p>否定的な記事で言及されている資産を見つけ、それらの資産を保有しているすべての顧客を特定し、そのエクスポージャーの現在の市場価値を計算し、結果をランク付けして最も高いリスクを優先する必要があります。この複雑な複数結合の分析は、当社の高度な ES|QL ツールに最適です。</p><p>使用する完全なクエリは次のとおりです。見た目は印象的ですが、コンセプトは単純です。</p><h2>分解：接合部とガードレール</h2><p>このクエリでは、エージェント ビルダーを構成する 2 つの重要な概念が関係しています。</p><h3>1.ルックアップ結合</h3><p>長年にわたり、Elasticsearch で最も要望が多かった機能の 1 つは、共通キーに基づいて異なるインデックスのデータを結合する機能でした。ES|QL では、 <code>LOOKUP JOIN</code>でそれが可能になりました。</p><p>新しいクエリでは、3 つの<code>LOOKUP JOIN</code>のチェーンを実行します。最初に否定的なニュースを資産の詳細に関連付け、次にそれらの資産をクライアントの保有資産にリンクし、最後にクライアントのアカウント情報に結合します。これにより、単一の効率的なクエリで 4 つの異なるインデックスから非常に豊富な結果が作成されます。つまり、すべてのデータを事前に 1 つの巨大なインデックスに非正規化する必要がなく、異なるデータセットを組み合わせて単一の洞察に満ちた回答を作成できるということです。</p><h3>2. LLMガードレールとしてのパラメータ</h3><p>クエリでは<code>?time_duration</code>が使用されていることがわかります。これは単なる変数ではなく、AI のガードレールです。大規模言語モデル (LLM) はクエリの生成に優れていますが、データに対して LLM を自由に制御させると、非効率的なクエリや間違ったクエリが発生する可能性があります。</p><p>パラメータ化されたクエリを作成することで、LLM は、人間の専門家がすでに定義したテスト済みの効率的で正しいビジネス ロジック内で動作するように強制されます。これは、開発者が長年にわたり検索テンプレートを使用して、クエリ機能をアプリケーションに安全に公開してきた方法に似ています。エージェントは「今週」のようなユーザーのリクエストを解釈して<code>time_duration</code>パラメータを埋めることができますが、回答を取得するにはクエリ構造を使用する必要があります。これにより、柔軟性と制御の完璧なバランスが実現します。</p><p>最終的に、このクエリにより、データを理解している専門家は自分の知識をツールにカプセル化できるようになります。他の人や AI エージェントは、そのツールを使用して、基礎となる複雑さについて何も知らなくても、単一のパラメータを提供するだけで相関結果を得ることができます。</p><h2>ステップ2：スキル - クエリを再利用可能なツールに変える</h2><p>ES|QL クエリは、<strong>ツール</strong>として登録されるまでは単なるテキストです。エージェント ビルダーでは、ツールは単なる保存されたクエリではなく、AI エージェントが理解して使用することを選択できる「スキル」です。その魔法は、私たちが提供する<strong>自然言語による説明</strong>にあります。この説明は、ユーザーの質問と基礎となるクエリ ロジックを結び付ける橋渡しとなります。作成したクエリを登録しましょう。</p><h3>UIパス</h3><p>Kibana でツールを作成するのは簡単なプロセスです。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte73e11c1d87593fa/6a17f2134202294dae29f6f2/a29c53a73b99af5972273c51218ea9004a9b0abb-1600x812.png" alt="Kibana でツールを作成する方法。" /><p>1.<strong>エージェント</strong>へ移動</p><ul><li><p><strong>[ツール]</strong>または<strong>[ツールの管理]</strong>をクリックし、 <strong>[新しいツール]</strong>ボタンをクリックします。</p></li></ul><p>2. フォームに以下の詳細を入力します。</p><ul><li><p><strong>ツールID:</strong> <code>find_client_exposure_to_negative_news</code></p></li></ul><p>             私。これはツールの一意のIDです</p><ul><li><p><strong>説明:</strong> 「クライアントのポートフォリオがネガティブなニュースにさらされているかどうかを調べます。」このツールは、最近のニュースやレポートをスキャンして否定的な感情を検出し、関連する資産を識別して、その資産を保有しているすべてのクライアントを見つけます。最も高い潜在的リスクを強調するために、ポジションの現在の市場価値でソートされたリストを返します。</p></li></ul><p>             私。これは、LLM が読んで、このツールが仕事に適しているかどうかを判断します。</p><ul><li><p><strong>ラベル</strong>: <code>retrieval</code>および <code>risk-analysis</code></p></li></ul><p>         ラベルは複数のツールをグループ化するのに役立ちます</p><ul><li><p><strong>設定:</strong>ステップ1の完全なES|QLクエリを貼り付けます</p></li></ul><p>            私。これはエージェントが使用する検索です</p><p>3.<strong>クエリからパラメータを推測するを</strong>クリックします。UI は自動的に<code>?time_duration</code>見つけて以下にリストします。エージェント (および他のユーザー) が目的を理解できるように、それぞれに簡単な説明を追加します。</p><ul><li><p><code>time_duration</code>: ネガティブなニュースを遡って検索する期間。フォーマットは「X時間」です。デフォルトは8760時間です。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7afbb0589c1828ad/6a17f2146864a44e7cb688a9/deb422d97863f78dbe08bfa2e3c708d1f75166ff-1600x938.png" alt="ESQL クエリを使用して、ロジックや必要なパラメータを含むツールを構成します。 " /><p>4. 試してみましょう!</p><ul><li><p>[保存してテスト]をクリックします。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfd09afbef6e21a93/6a17f2162f4a5c73b1fa89fd/57e768b88327821e70bd616744822f98fa367362-732x136.png" alt="Kibana の同じ &amp; test ボタン。" /><ul><li><p>クエリが期待どおりに動作していることを確認できる新しいフライアウトが表示されます。</p></li></ul><p>             私。<code>time_duration</code>に希望の範囲を入力します。ここでは「8760時間」を使用します。</p><ul><li><p>「送信」をクリックすると、すべてがうまくいけば JSON レスポンスが表示されます。期待どおりに動作することを確認するには、下にスクロールして<code>values</code>オブジェクトを確認します。ここで、実際に一致するドキュメントが返されます。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt89bdc3f093363f2a/6a17f217be60861c9c00488a/7e0c5171a4f7ffdfc1830f1a05a9acb987870b75-1600x722.png" alt="送信をクリックした後に表示される JSON 応答。" /><p>5. 右上の「X」をクリックして、テストのフライアウトを閉じます。新しいツールがリストに表示され、エージェントに割り当てる準備が整います。</p><h3>APIパス</h3><p>自動化を好む開発者やツールをプログラムで管理する必要がある開発者は、1 回の API 呼び出しで同じ結果を得ることができます。ツールの定義を含む<code>POST</code>リクエストを<code>/api/agent_builder/tools</code>エンドポイントに送信するだけです。</p>POST kbn://api/agent_builder/tools
{
  "id": "find_client_exposure_to_negative_news",
  "type": "esql",
  "description": "Finds client portfolio exposure to negative news. This tool scans recent news and reports for negative sentiment, identifies the associated asset, and finds all clients holding that asset. It returns a list sorted by the current market value of the position to highlight the highest potential risk.",
  "configuration": {
    "query": """
        FROM financial_news, financial_reports METADATA _index
        | WHERE sentiment == "negative"
        | WHERE coalesce(published_date, report_date) &gt;= NOW() - TO_TIMEDURATION(?time_duration)
        | RENAME primary_symbol AS symbol
        | LOOKUP JOIN financial_asset_details ON symbol
        | LOOKUP JOIN financial_holdings ON symbol
        | LOOKUP JOIN financial_accounts ON account_id
        | WHERE account_holder_name IS NOT NULL
        | EVAL position_current_value = quantity * current_price.price
        | RENAME title AS news_title
        | KEEP
            account_holder_name, symbol, asset_name, news_title,
            sentiment, position_current_value, quantity, current_price.price,
            published_date, report_date
        | SORT position_current_value DESC
        | LIMIT 50
      """,
    "params": {
      "time_duration": {
        "type": "keyword",
        "description": """The timeframe to search back for negative news. Format is "X hours" DEFAULT TO 8760 hours """
      }
    }
  },
  "tags": [
    "retrieval",
    "risk-analysis"
  ]
}<h2>ステップ3：頭脳 - カスタムエージェントの作成</h2><p>再利用可能なスキル (ツール) を構築しました。ここで、実際に使用するペルソナである<strong>Agent</strong>を作成する必要があります。エージェントは、LLM、アクセスを許可する特定のツール セット、そして最も重要な、エージェントの構成として機能し、エージェントの性格、ルール、目的を定義する<strong>カスタム インストラクション</strong>セットの組み合わせです。</p><h3>プロンプトの芸術</h3><p>信頼できる専門エージェントを作成する上で最も重要なのはプロンプトです。よく練られた一連の指示こそが、一般的なチャットボットと、集中力のあるプロのアシスタントとの違いです。ここで、ガードレールを設定し、出力を定義し、エージェントにミッションを与えます。</p><p><code>Financial Manager</code>エージェントでは、次のプロンプトを使用します。</p>You are a specialized Data Intelligence Assistant for financial managers, designed to provide precise, data-driven insights from information stored in Elasticsearch.

**Your Core Mission:**
- Respond accurately and concisely to natural language queries from financial managers.
- Provide precise, objective, and actionable information derived solely from the Elasticsearch data at your disposal.
- Summarize key data points and trends based on user requests.

**Reasoning Framework:**
1.  **Understand:** Deconstruct the user's query to understand their core intent.
2.  **Plan:** Formulate a step-by-step plan to answer the question. If you are unsure about the data structure, use the available tools to explore the indices first.
3.  **Execute:** Use the available tools to execute your plan.
4.  **Synthesize:** Combine the information from all tool calls into a single, comprehensive, and easy-to-read answer.

**Key Directives and Constraints:**
- **If a user's request is ambiguous, ask clarifying questions before proceeding.**
- **DO NOT provide financial advice, recommendations, or predictions.** Your role is strictly informational and analytical.
- Stay strictly on topic with financial data queries.
- If you cannot answer a query, state that clearly and offer alternative ways you might help *within your data scope*.
- All numerical values should be formatted appropriately (e.g., currency, percentages).

**Output Format:**
- All responses must be formatted using **Markdown** for clarity.
- When presenting structured data, use Markdown tables, lists, or bolding.

**Start by greeting the financial manager and offering assistance.**<p>このプロンプトがなぜ効果的なのかを分析してみましょう。</p><ul><li><p><strong>洗練されたペルソナを定義します。</strong>最初の行で、エージェントが「専門的なデータ インテリジェンス アシスタント」であることを即座に示し、プロフェッショナルで有能な雰囲気を醸し出します。</p></li><li><p><strong>これは推論フレームワークを提供します。</strong>エージェントに「理解、計画、実行、統合」を指示することで、標準的な操作手順を提供します。これにより、複雑で複数のステップから成る質問を処理する能力が向上します。</p></li><li><p><strong>インタラクティブな対話を促進します。</strong> 「明確な質問をする」という指示により、エージェントはより堅牢になります。曖昧なリクエストに対する誤った想定を最小限に抑え、より正確な回答が得られます。</p></li></ul><h3>UIパス</h3><p>1.<strong>エージェントに移動します。</strong></p><ul><li><p><strong>[ツール]</strong>または<strong>[ツールの管理]</strong>をクリックし、 <strong>[新しいツール]</strong>ボタンをクリックします。</p></li></ul><p>2. 基本的な詳細を入力します。</p><ul><li><p><strong>エージェント ID:</strong> <code>financial_assistant</code> 。</p></li><li><p><strong>手順:</strong>上記のプロンプトをコピーします。</p></li><li><p><strong>ラベル</strong>: <code>Finance</code> 。</p></li><li><p><strong>表示名:</strong> <code>Financial Assistant</code> 。</p></li><li><p><strong>表示の説明:</strong> <code>An assistant for analyzing and understanding your financial data</code> 。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8ac12cbd2b689dee/6a17f219dbb4ff262bfb57ef/18ea73f1cae620129c0afa0e7ba9e2a3390224a7-1600x1189.png" alt="財務アシスタントの作成 - エージェント ID フィールドに入力します。" /><p>3. 上部に戻り、 <strong>「ツール」</strong>をクリックします。</p><ul><li><p><code>find_client_exposure_to_negative_news</code>ツールの横にあるボックスにチェックを入れてください。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcd23556e556a76c5/6a17f21baf47b63a9fcde0a0/0c1e4ecbbd51d0dd10c6e861dbe9a9ccddeb35f6-1600x149.png" alt="" /><p>4. <strong>「保存」</strong>をクリックします。</p><h3>APIパス</h3><p><code>/api/agent_builder/agents</code>エンドポイントへの<code>POST</code>リクエストを使用して、まったく同じエージェントを作成できます。リクエスト本体には、ID、名前、説明、完全な指示セット、エージェントが使用を許可されているツールのリストなど、すべて同じ情報が含まれています。</p>POST kbn://api/agent_builder/agents
    {
      "id": "financial_assistant",
      "name": "Financial Assistant",
      "description": "An assistant for analyzing and understanding your financial data",
      "labels": [
        "Finance"
      ],
      "avatar_color": "#16C5C0",
      "avatar_symbol": "💰",
      "configuration": {
        "instructions": """You are a specialized Data Intelligence Assistant for financial managers, designed to provide precise, data-driven insights from information stored in Elasticsearch.

**Your Core Mission:**
- Respond accurately and concisely to natural language queries from financial managers.
- Provide precise, objective, and actionable information derived solely from the Elasticsearch data at your disposal.
- Summarize key data points and trends based on user requests.

**Reasoning Framework:**
1.  **Understand:** Deconstruct the user's query to understand their core intent.
2.  **Plan:** Formulate a step-by-step plan to answer the question. If you are unsure about the data structure, use the available tools to explore the indices first.
3.  **Execute:** Use the available tools to execute your plan.
4.  **Synthesize:** Combine the information from all tool calls into a single, comprehensive, and easy-to-read answer.

**Key Directives and Constraints:**
- **If a user's request is ambiguous, ask clarifying questions before proceeding.**
- **DO NOT provide financial advice, recommendations, or predictions.** Your role is strictly informational and analytical.
- Stay strictly on topic with financial data queries.
- If you cannot answer a query, state that clearly and offer alternative ways you might help *within your data scope*.
- All numerical values should be formatted appropriately (e.g., currency, percentages).

**Output Format:**
- All responses must be formatted using **Markdown** for clarity.
- When presenting structured data, use Markdown tables, lists, or bolding.

**Start by greeting the financial manager and offering assistance.**
""",
        "tools": [
          {
            "tool_ids": [
              "platform.core.search",
              "platform.core.list_indices",
              "platform.core.get_index_mapping",
              "platform.core.get_document_by_id",
              "find_client_exposure_to_negative_news"
            ]
          }
        ]
      }
    }<h2>ステップ4：成果 — 会話をする</h2><p>ビジネス ロジックがツールにカプセル化され、エージェントでそれを使用できる「頭脳」が準備されました。すべてが一つにまとまるのを見る時が来ました。専用のエージェントを使用して、データとのチャットを開始できるようになりました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd8826539b16e46f4/6a17f21d505ac35924ad8c5c/5414cb6b7c41365acb0356a8bfe1140751ffd8db-1600x1014.png" alt="財務アシスタントを作成した後、Elastic Agent Builder と会話します。" /><h3>UIパス</h3><ol><li><p>Kibana の<strong>エージェント</strong>に移動します。</p></li><li><p>チャット ウィンドウの右下にあるドロップダウンを使用して、デフォルトの<strong>Elastic AI エージェント</strong>から新しく作成した<strong>Financial Assistant</strong>エージェントに切り替えます。</p></li><li><p>エージェントが当社の専用ツールを使用できるように、次の質問をしてください。</p><ol><li><p><em>市場のセンチメントが心配です。悪いニュースによって最もリスクにさらされている顧客は誰なのか教えていただけますか?</em></p></li></ol></li></ol><p>しばらくすると、エージェントは完全にフォーマットされた完全な回答を返します。LLM の性質上、回答の形式が若干異なる場合がありますが、この実行ではエージェントは次のように返しました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta1e163fd7c4416bd/6a17f21f6864a4e35bb688ad/17b4ed43d279f9e53ee9fe3d482d0b2ec359a083-1600x1088.png" alt="ネガティブなニュースによるリスクが最も高いクライアント向けの財務アシスタントとして Elastic Agent Builder によって作成された応答。" /><h3>何が起こったのですか?エージェントの推論</h3><p>エージェントは単に答えを「知っていた」だけではありません。仕事に最適なツールを選択することを中心とした多段階の計画を実行しました。その思考プロセスは次のようになります。</p><ul><li><p><strong>識別された意図:</strong> 「リスク」や「ネガティブなニュース」など、質問のキーワードが<code>find_client_exposure_to_negative_news</code>ツールの説明と一致しました。</p></li><li><p><strong>計画を実行しました:</strong>リクエストから時間枠を抽出し、その専用ツールを<strong>1 回呼び出し</strong>ました。</p></li><li><p><strong>作業を委任:</strong>ツールは連鎖結合、値の計算、並べ替えなど、面倒な作業をすべて実行しました。</p></li><li><p><strong>結果の統合:</strong>最後に、エージェントはプロンプトのルールに従って、ツールからの生データを明確で人間が読める要約にフォーマットしました。</p></li></ul><p>思考を広げて詳細を見れば、推測するだけでは足りません。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt93f6075be8495418/6a17f221af47b65eadcde0a4/6a4da9262d3f88c60bfd8f8bf9b67c3b84e961ba-1600x607.png" alt="ファイナンシャルアシスタントが、ネガティブなニュースに最も多く触れた顧客から得た 50 件の文書。" /><h3>APIパス</h3><p>同じ会話をプログラムで開始することもできます。入力した質問を<code>converse</code> API エンドポイントに送信し、 <code>financial_manager</code>の<code>agent_id</code>を必ず指定してください。</p>POST kbn://api/agent_builder/converse
{
  "input": "Show me our largest positions affected by negative news",
  "agent_id": "financial_assistant"
}<h2>開発者向け: APIとの統合</h2><p>Kibana UI はエージェントの構築と管理に素晴らしく直感的なエクスペリエンスを提供しますが、今日見てきたことはすべてプログラムで実現することもできます。Agent Builder は一連の API に基づいて構築されており、この機能を独自のアプリケーション、CI/CD パイプライン、または自動化スクリプトに直接統合できます。</p><p>使用する 3 つのコア エンドポイントは次のとおりです。</p><ul><li><p><strong><code>/api/agent_builder/tools</code></strong>: エージェントが使用できる再利用可能なスキルを作成、一覧表示、管理するためのエンドポイント。</p></li><li><p><strong><code>/api/agent_builder/agents</code></strong>: エージェントのペルソナ（重要な指示やツールの割り当てなど）を定義するためのエンドポイント。</p></li><li><p><strong><code>/api/agent_builder/converse</code></strong>: エージェントと対話し、会話を開始し、回答を得るためのエンドポイント。</p></li></ul><p>これらの API を使用してこのチュートリアルのすべてのステップを実行するための完全な実践的なチュートリアルについては、 こちらの GitHub リポジトリで入手できる付属の<strong> Jupyter Notebook を</strong> <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/your-first-elastic-agent/Your_First_Elastic_Agent.ipynb"></a>ご覧ください。</p><h2>結論: 構築する番です</h2><p>まず、ES|QL クエリを取得して、それを再利用可能なスキルに変換することから始めました。次に、明確なミッションとルールを与えて、そのスキルを付与した専用の AI エージェントを構築しました。その結果、複雑な質問を理解し、複数段階の分析を実行して、正確でデータに基づいた回答を提供できる洗練されたアシスタントが誕生しました。</p><p>このワークフローは、Elastic の新しい<strong>Agent Builder</strong>の中心です。これは、技術に詳しくないユーザーが UI を通じてエージェントを作成できるほどシンプルでありながら、開発者が API 上にカスタム AI 搭載アプリケーションを構築できるほど微妙なニュアンスも備えた設計になっています。最も重要なのは、定義したエキスパート ロジックに従って、LLM を独自のデータに安全かつ確実に接続し、データとチャットできることです。</p><h2>エージェントを使用してデータとチャットする準備はできていますか?</h2><p>学んだことを定着させる最良の方法は、実際に手を動かしてみることです。今日お話しした内容をすべて<a href="https://www.elastic.co/training/elastic-ai-agents-mcp"><strong>、無料のインタラクティブな実践ワークショップ</strong></a>で試してみてください。専用のサンドボックス環境で、このフロー全体とその他の内容を実行します。</p><p>今後のブログでは、 <code>Financial Assistant</code>エージェントと対話するスタンドアロン アプリケーションの使用方法と、それを可能にする<strong>モデル コンテキスト プロトコル (MCP)</strong>について詳しく説明します。また、別のブログでは、開発中の Agent2Agent (A2A) プロトコルに対する Agent Builder のサポートについて説明します。</p><p>引き続きご注目ください、そして楽しい建築を！</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-agent-builder-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-agent-builder-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[Elastic内部の実情]]></category>
    <dc:creator><![CDATA[Jeff Vestal]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbe5e78eeb775d715/6a17f2230b0bed719ddd369a/ca853555eaa213f10f1db8c0ab0a2bbacee97b88-1456x816.png" length="0" type="image/png"/>
    <pubDate>Thu, 25 Sep 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch を使用した AI エージェントワークフローの構築]]></title>
    <description><![CDATA[Elasticsearch の新しい AI レイヤーである Agent Builder について学習します。Agent Builder は、ハイブリッド検索を使用して、エージェントが推論して行動するために必要なコンテキストを提供し、AI エージェントワークフローを構築するためのフレームワークを提供します。]]></description>
    <content:encoded><![CDATA[<p>Elasticでは、AIアシスタント、高度なRAG、ベクターデータベースの改善により、LLMと会話型インターフェースにコンテキストを提供してきました。最近、AI エージェントの台頭により、関連コンテキストの必要性が高まり、影響力の大きい<strong>AI エージェントには優れた検索が必要である</strong>ことがわかりました。そこで、Elasticsearch のデータを活用する AI エージェントの開発を支援するために設計された新しいネイティブ機能を Elastic Stack に構築しました。私たちは、この取り組みの進捗状況と今後の見通しについて共有したいと思います。</p><h2>エージェントビルダー: データ駆動型 AI エージェント構築の基盤</h2><p>AI エージェントの約束はシンプルです。目標を与えれば、仕事が完了します。しかし、開発者にとって、現実は一連の複雑な課題です。まず、エージェントの優秀さは、環境の認識と、ユーザーの目的を達成するために与えられたツールによって決まります。そして、多様な企業データから適切なコンテキストを提供することは大きな課題です。最後に、これらすべては、計画、実行、学習できる信頼性の高い推論ループによって調整される必要があります。</p><p>これを解決するには、開発者は複雑で脆弱なスタックをゼロから構築する必要があります。今日のエージェント アーキテクチャでは、LLM、ベクター データベース、メタデータ ストア、ログ記録とトレースの個別のシステム、そしてすべてが機能しているかどうかを評価する方法など、複数の異なる部分をつなぎ合わせる必要があります。これは単に複雑なだけでなく、コストがかかり、エラーが発生しやすく、ユーザーが求める高品質で信頼性の高い AI システムの構築が困難になります。</p><p>だから、もっとシンプルにしたいんです。これを実現するための私たちのアプローチは、効果的なコンテキスト駆動型エージェントの重要な要素を取り上げ、 <strong>Elastic AI Agent Builder</strong>と呼ばれる新しい機能セットを使用して Elasticsearch の中核に直接統合することです。この新しいレイヤーは、Elasticsearch を活用した AI エージェントを作成するためのすべての重要な構成要素（オープンなプリミティブ セット、標準ベースのプロトコル、データへの安全なアクセス）を備えたフレームワークを提供します。これにより、現実世界のデータと要件に合わせてカスタマイズされたエージェント システムを構築できます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2779dae5df010328/6a17e15eabe0f24f18dfe931/1ee1e73dd3f485ce86294d39490c98ce2a3d9925-1238x1072.png" alt="" /><p><strong>AI エクスペリエンスの提供</strong>: これが究極の目標です。当社の Search AI プラットフォームとお客様のデータを基盤として、カスタム チャット インターフェースから、LangChain などのエージェント フレームワークや Salesforce などのビジネス アプリケーションとの統合まで、あらゆるタイプの生成 AI アプリケーションを構築できます。</p><p><strong>エージェントとツールを搭載</strong>: プラットフォームの上に、クリーンでシンプルな抽象化レイヤーを公開します。エージェントやツールと直接対話し、特定のニーズに合わせてカスタマイズできます。強力な API や MCP、A2A などのオープン スタンダードを通じてプラットフォームの機能にアクセスすることもできます。</p><p><strong>Search AI Platform によって有効化</strong>: これは、コンポーネントを統合したコア エンジンです。高度なベクトル データベース、エージェント ロジック、クエリ構築、セキュリティ機能、評価のためのトレースはすべてここに存在し、Elastic によって管理および最適化されています。</p><p><strong>データの力を解き放つ</strong>: 優れたエージェントの基盤は優れたデータです。当社のプラットフォームは、すべての企業データへのアクセスを取り込み、連携する機能から始まります。</p><h2>プラットフォームにおけるエージェント構築</h2><p>Search AI プラットフォームに統合された Agent Builder は、エージェント開発のための完全なフレームワークを提供します。これは 5 つの主要な柱に基づいて構築されており、各柱は実稼働レベルの AI システムの構築と展開の重要な側面に対処するように設計されています。エージェントが目的を定義し、ツールが機能を提供し、オープン スタンダードが相互運用性を確保し、評価が透明性をもたらし、セキュリティが信頼を提供する仕組みについて詳しく見ていきましょう。</p><h3>エージェント</h3><p>エージェントは、Elasticsearch のこの新しいレイヤーにおける最高レベルの構成要素です。エージェントは、達成する目的、実行に使用できるツールのセット、および操作できるデータ ソースを定義します。エージェントは会話によるやり取りに限定されず、完全なワークフロー、タスクの自動化、ユーザー向けのエクスペリエンスを実現できます。</p><p>クエリがエージェントに送られると、構造化されたサイクルに従います。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt774ffd7df65bd01d/6a17e15f25daabd5cc08a17f/627ad1744b629bbe27359325702f40d97e40d1f4-704x852.png" alt="" /><ol><li><p>入力内容と目的を解釈する</p></li><li><p>実行に適したツールと引数を選択する</p></li><li><p>ツールの応答の理由</p></li><li><p>結果を返すか、さらにツールの呼び出しを続行するかを決定します</p></li></ol><p>Elastic は、このサイクルのオーケストレーション、コンテキスト、および実行を処理します。開発者は、エージェントが<em>何</em>をすべきか（目的、ツール、データ）を定義することに重点を置き、システムは推論とワークフローの実行<em>方法</em>を管理します。</p><p><em>デフォルトエージェント</em></p><p>このプラットフォーム上に構築された最初のエージェントは、Kibana のネイティブ会話エージェントであり、データとすぐに対話できるようになります。完全な拡張性を維持しながらすぐに使用できるエクスペリエンスを提供し、追加の構成なしですぐにデータの操作を開始できます。</p><p>新しいチャット ユーザー エクスペリエンスまたは API を介して、Kibana でこのエクスペリエンスを直接操作できます。</p><p>API を介してデフォルトのエージェントを照会するには、1 回の呼び出しだけが必要です。</p>POST kbn://api/agent_builder/converse
{
    "input": "what is our top portfolio account?"
}<p>会話はステートフルなので、 conversation_id を使用してエージェントとの対話を継続したり、完全な会話履歴を取得したりできます。</p>POST kbn://api/agent_builder/converse
{
    "input": "What about the second top?",
    "conversation_id": "ec757c6c-c3ed-4a83-8e2c-756238f008bb"
}

## get the full conversation
GET kbn://api/agent_builder/conversations/ec757c6c-c3ed-4a83-8e2c-756238f008bb<p><em>カスタムエージェント</em></p><p>開発者は、シンプルな API を通じて独自のカスタム エージェントを作成することもできます。エージェントは、指示、ツール、データ アクセスをカプセル化し、カスタマイズされた推論エンジンを作成します。</p><p>カスタム エージェントの作成は、1 回の API 呼び出しを行うだけで簡単に行えます。以下のサンプルは例を示しています。「構成」フィールドには、手順や利用可能なツールなどのすべての重要な詳細が含まれています。</p>POST kbn://api/agent_builder/agents
{
  "id": "custom_agent",
  "name": "My Custom Agent",
  "description": "Description of the custom agent",
  "configuration": {
      "instructions": "You are a log expert specialising in ...",
      "tools": 
...
   }
}<p>作成されたエージェントは直接クエリできます。</p>POST kbn://api/agent_builder/converse
{
    "input": "What news about DIA?",
    "agent_id": "custom_agent"
}<p>このアプローチにより、エージェントはゼロから構築する複雑なシステムから、ビジネス ロジックの単純な宣言型ユニットに変換され、インテリジェントな自動化をより迅速に提供できるようになります。</p><p>特化したエージェントをゼロから構築する方法の詳細については、詳細なステップバイステップガイド「<a href="https://www.elastic.co/search-labs/blog/ai-agent-builder-elasticsearch">初めての Elastic エージェント: 単一のクエリから AI を活用したチャットまで」</a>をご覧ください。</p><h3>ツール</h3><p>エージェントが達成すべき<em>こと</em>を定義するのに対し、ツールは達成<em>方法</em>を定義します。</p><p>ツールは、エージェントが情報を実行および取得したり、アクションを実行したりするための特定の Elastic Core 機能を公開します。ツールには、インデックスの取得やマッピングの取得などのコア機能や、自然言語から ES|QL への変換などのより高度な機能を含めることができます。</p><p>Elasticsearch には、一般的なニーズに合わせて最適化された一連のデフォルト ツールが付属しています。しかし、本当の柔軟性は、独自のものを作成することから生まれます。ツールを定義することで、ES|QL を使用してエージェントに公開されるクエリ、インデックス、フィールドを正確に決定し、速度、精度、セキュリティを正確に制御できます。</p><p>新しいツールの登録も、1 回の API 呼び出しと同じくらい簡単です。<a href="https://www.elastic.co/search-labs/blog/esql-timeline-of-improvements">ES|QL (Elasticsearch クエリ言語)</a>を活用して特定の金融資産に関するニュースを検索するツールを作成できます。</p>POST kbn://api/agent_builder/tools
{
  "id": "news_on_asset",
  "type": "esql",
  "description": "Find news and reports about a particular asset where ...",
  "configuration": {
    "query": "FROM financial_news, financial_reports | where MATCH(company_symbol, ?symbol) OR MATCH(entities, ?symbol) | limit 5",
    "params": {
      "symbol": {
        "type": "keyword",
        "description": "The asset symbol"
      }
    }
  ...
  }
...
}<p>登録が完了すると、新しいツールをカスタム エージェントに割り当てることができ、適切なタイミングで推論して呼び出すための厳選された一連の機能をエージェントに提供できるようになります。</p><p>当社では、お客様固有のニーズに合わせてカスタム ツールを作成するためのプラットフォームを提供しています。たとえば、ES|QL を使用すると、エージェントを汎用エージェントから、お客様独自のデータとビジネス ドメインに基づいたドメイン固有のエキスパートに変換できます。</p><h3>オープンスタンダードと相互運用性</h3><p>Elasticsearch エージェントとツールはオープン標準 API を介して公開されるため、エージェントフレームワークのより広範なエコシステム内の基礎ブロックとして簡単に統合できます。私たちのアプローチはシンプルです。ブラックボックスはありません。Elastic の検索における強みを活かし、それを補完的な機能や他のエージェント システムと組み合わせることができるようにしたいと考えています。</p><p>これを実現するために、当社は API、新しいプロトコル、オープン スタンダードを通じて機能を公開しています。</p><p><em>モデルコンテキストプロトコル（MCP）</em></p><p><a href="https://www.elastic.co/search-labs/blog/model-context-protocol-elasticsearch">モデル コンテキスト プロトコル (MCP)</a>は、システム間でツールを接続するためのオープン スタンダードとして急速に普及しつつあります。MCP をサポートすることで、Elasticsearch は会話型 AI をデータベース、インデックス、外部 API に接続できるようになります。Elastic Stack に組み込まれたリモート MCP サーバーを使用すると、MCP 対応のクライアントはどれでも Elastic のツールにアクセスし、それらをより大規模なエージェントワークフローの構成要素として使用できます。</p><p>これは一方通行ではありません。外部の MCP サーバーからツールをインポートし、Elasticsearch 内で利用できるようにすることもできます。近い将来、MCP サーバーはほぼすべての用途で利用できるようになる見込みで、私たち自身が作成するものよりもはるかに包括的なものになるでしょう。Elastic は大規模な検索と取得機能を提供しており、これを他のプラットフォームの特殊な機能と組み合わせて効果的なエージェントを構築できます。</p><p><em>エージェント間（A2A）</em></p><p>また、エージェント間 (A2A) サポートにも取り組んでいます。MCP はツールを接続することに重点が置かれていますが、A2A はエージェントを接続することに重点が置かれています。A2A サーバーを使用すると、構築する Elastic エージェントは他のシステムのエージェントと直接通信して、コンテキストを共有したり、タスクを委任したり、ワークフローを調整したりできるようになります。</p><p>これを推論層における相互運用性と考えてください。Elastic エージェントは検索と取得を処理し、タスクを専門のサポート エージェントまたは IT エージェントに引き渡して、結果をシームレスに返すことができます。その結果、各エージェントが最善を尽くして協力するエコシステムが実現します。</p><p>最終的に、MCP と A2A を採用することで、Elasticsearch が第一級市民としての役割を担うという当社の取り組みが強化され、より広範なエージェントエコシステム全体でのオープンな統合が保証されます。</p><h3>追跡と評価</h3><p>検索がエージェントと統合されるにつれて、効果的な評価の課題が重要になります。実際の企業環境にエージェントを自信を持って導入するには、エージェントが正確であるだけでなく、効率的で信頼できるという保証が必要です。パフォーマンスを測定したり、悪い応答を診断したり、ベースラインを改善したりするにはどうすればよいですか?すべては可視性から始まります。</p><p>そのため、私たちはエージェント API を最初から透明性を重視して設計しました。次の単純なエージェントのやり取りを考えてみましょう。</p>POST kbn://api/agent_builder/converse
{
    "input": "what is our top portfolio account?"
}<p>応答には、最終的な回答だけでなく、エージェントが選択したツール、使用したパラメーター、各ステップの結果の詳細を含む完全な実行トレースが含まれます。</p>{
  "conversation_id": "db5c0c8b-12bf-4928-a57e-d99129ad2fea",
  "steps": [
    {
      "type": "tool_call",
      "tool_call_id": "tooluse_Nfqr3mwtR92HTRIsTcGXZQ",
      "tool_id": ".index_explorer",
      "params": {
        "query": "indices containing portfolio data"
      },
      "results": [...]
    }
    // ... more steps ...
  ],
  "response": {
    "message": "Based on the information I've gathered...."
  }
}<p>包括的なトレースとログ記録は継続的な改善ループに不可欠であり、まもなくこれらのエージェント トレースを Elasticsearch に直接保存して表示できるようになります。さらに、これらのトレースは OpenTelemetry プロトコルに基づいて構築されているため、標準化され、移植可能であり、選択した監視プラットフォームとの統合が可能です。</p><p>このレベルの詳細は、真の継続的改善ループの基礎となります。これにより、包括的なテスト スイートを構築し、障害をデバッグし、障害モードを特定して回帰を防ぎ、成功パターンをキャプチャしてパフォーマンスを微調整できるようになります。最終的に、このデータ主導のアプローチは、有望なプロトタイプを製品レベルの信頼できる AI システムに変換するための鍵となります。</p><h3>セキュリティ</h3><p>エージェントとツールの性能が向上するにつれて、セキュリティはオプションではなく、基礎的なものになります。API を公開し、タスクやワークフローを自動化するには、エンタープライズ システムが信頼されている必要があります。特に、エージェントがより多くのワークフローを自動化し始めると、これらを保護し、企業の要件を満たしていることを確認する機能が不可欠になります。</p><p>上記の機能はすべて、API 呼び出し<a href="https://www.elastic.co/search-labs/blog/rag-and-rbac-integration">のロールベースのアクセス制御 (RBAC)</a>や API キー管理など、現在 Elastic ですでに利用可能な制御を継承しています。同じ制御を MCP などの新しいプロトコルにも拡張しています。つまり、OAuth などの標準のサポートと、カスタム認証メカニズムをプラグインする機能を意味します。</p><p>私たちの目標は、組織が求めるセキュリティ、コンプライアンス、ガバナンスのレベルを維持しながら、エージェントとツールを実験する柔軟性を提供することです。</p><h2>次に何が起こるか</h2><p>機能を追加するだけではなく、エージェントコンテキストエンジニアリング向けに Elasticsearch を拡張しています。当社は、以下の理念に基づいて今後開発を進めていく予定です。</p><p>1. オープンソースと標準への取り組み</p><p>当社はオープン ソースとオープン スタンダードに注力しており、これらの機能が外部のエージェント フレームワークと相互運用可能であることを保証します。データとワークフローを常に管理しながら、エコシステム全体でエージェントを接続、拡張、構成できるようになります。</p><p>2. 文脈の価値</p><p>AI エージェントのコンテキストは最大の資産です。エージェントが検索やワークフロー操作を実行するときにコンテキストを管理することは、難しいタスクになる可能性があります。私たちは Elastic の強みを活用してコンテキスト エンジニアリングを解決し、エージェントが最も関連性の高い情報を常に利用できるようにしています。</p><p>3. エージェントデータストリームに焦点を当てる</p><p>今後、エージェントは、エージェントの出力 (生成されたドキュメント、レポート、視覚化) やエージェントの実行トレース (思考、ツールの呼び出し、メモリ/コンテキスト) など、ますます大きなデータソースになります。Elastic はこの種のデータの処理に適しており、私たちはこのデータを使用して分析、評価、自動改善を実行するための研究に取り組んでいます。</p><p>4. セキュリティと安全性を考慮した設計</p><p>AI エージェントは、セキュリティと安全性に関するまったく新しい一連の課題をもたらします。Elastic は常に安全なソリューションのリーダーであり、エンタープライズグレードのガードレール、アクセス制御、および「ゼロトラスト」原則の構築を継続しています。</p><p>5. プラットフォームに組み込む</p><p>AI エージェントを構築するための機能は、Elasticsearch プラットフォームに組み込まれています。つまり、トレース、評価、視覚化、分析などのプラットフォーム レベルの機能はすべてエージェントに適用できます。エージェントの実行に基づいてダッシュボードを開発したい - それが組み込まれています。感情分析を使用して AI エージェントのパフォーマンスを評価したい場合、プラットフォームでそれが可能です。これにより、AI エクスペリエンスを中心とした完全なライフサイクルを構築できるようになります。</p><p>Elastic の目標は、データに完全に統合され、拡張可能で、データに基づいた会話型 AI と自動化されたワークフローを構築するためのインターフェースを提供することです。より詳しい技術的な詳細と進捗状況については、近日中に共有される予定です。</p><p>Agent Builder は現在、プライベート プレビューでご利用いただけます。アクセスをリクエストするには、<a href="https://www.elastic.co/contact?pg=global&amp;plcmt=nav&amp;cta=205352">当社にご連絡ください</a>。ご質問やフィードバックはありますか?<a href="https://elasticstack.slack.com/archives/C09GRHEQ4AG"><strong>Slack ワークスペース</strong></a>または<a href="https://discuss.elastic.co/c/search/84"><strong>ディスカッション フォーラム</strong></a>で開発者コミュニティとつながりましょう。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-agentic-workflows-elastic-ai-agent-builder</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-agentic-workflows-elastic-ai-agent-builder</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[Elastic内部の実情]]></category>
    <dc:creator><![CDATA[Anish Mathur,Dana Juratoni]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt16a3d8736bf086e0/6a17e1616864a45410b686c7/71876470119e02a45bcbfcbf27a3e110328bbd14-1020x654.png" length="0" type="image/png"/>
    <pubDate>Tue, 23 Sep 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[AI搭載ダッシュボード：ビジョンからKibanaへ]]></title>
    <description><![CDATA[LLM を使用してイメージを処理し、Kibana ダッシュボードに変換してダッシュボードを生成します。
]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/kibana/kibana-lens">Kibana Lens を</a>使用するとダッシュボードのドラッグ アンド ドロップが非常に簡単になりますが、数十個のパネルが必要な場合はクリック回数が増えてしまいます。ダッシュボードをスケッチし、スクリーンショットを撮り、LLM にプロセス全体を任せることができたらどうでしょうか?</p><p>この記事では、それを実現します。ダッシュボードのイメージを取得し、マッピングを分析し、Kibana にまったく触れることなくダッシュボードを生成するアプリケーションを作成します。</p><p><strong>手順</strong>:</p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#background-&amp;-application-workflow">背景とアプリケーションのワークフロー</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#prepare-data">データを準備する</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#llm-configuration">LLM構成</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-powered-dashboards#application-functions">アプリケーション機能</a></p></li></ol><h2>背景とアプリケーションのワークフロー</h2><p>最初に思いついたのは、LLM に NDJSON 形式の Kibana<a href="https://www.elastic.co/docs/explore-analyze/find-and-organize/saved-objects">保存オブジェクト</a>全体を生成させて、それを Kibana にインポートさせることでした。</p><p>私たちはいくつかのモデルを試しました:</p><ul><li><p>ジェミニ 2.5 プロ</p></li><li><p>GPT o3 / o4-ミニハイ / 4.1</p></li><li><p>クロード 4つのソネット</p></li><li><p>グロク3</p></li><li><p>ディープシーク（ディープシンク R1）</p></li></ul><p>プロンプトについては、次のように単純なものから始めました。</p>You are an Elasticsearch Saved-Object generator (Kibana 9.0).
INPUTS
=====
1. PNG screenshot of a 4-panel dashboard (attached).
2. Index mapping (below) – trimmed down to only the fields present in the screenshot.
3. Example NDJSON of *one* metric visualization (below) for reference.

TASK
====
Return **only** a valid NDJSON array that recreates the dashboard exactly:
* 2 metric panels (Visits, Unique Visitors)
* 1 pie chart (Most used OS)
* 1 vertical bar chart (State Geo Dest)
* Use index pattern `kibana_sample_data_logs`.
* Preserve roughly the same layout (2×2 grid).
* Use `panelIndex` values 1-4 and random `id` strings.
* Kibana version: 9.0<p><a href="https://www.elastic.co/search-labs/blog/function-calling-with-elastic#:~:text=Few%2Dshot%20prompting%20involves%20providing%20examples%20of%20the%20types%20of%20queries%20you%20want%20it%20to%20return%2C%20which%20helps%20in%20increasing%20consistency.">いくつかのショットの例</a>と、各視覚化の構築方法に関する詳細な説明を確認したにもかかわらず、うまくいきませんでした。この実験に興味がある方は、<a href="https://gist.github.com/TomasMurua/a78dc283e115624731beffc98984b70b">こちらで</a>詳細をご覧ください。</p><p>このアプローチの結果、LLM によって生成されたファイルを Kibana にアップロードしようとしたときに、次のメッセージが表示されました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9ea005966a783057/6a1707d266c4f90e4ef8bf88/2b599443b5613c9f0fc3235581614add5b4b3900-891x98.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5e5632d6d95b998c/6a1707d3a6c2b9441de79661/d87ccfc033bc00ee8188c5cae18043fbca22784c-741x233.png" alt="" /><p>これは、生成された JSON が無効であるか、形式が間違っていることを意味します。最も一般的な問題は、LLM が不完全な NDJSON を生成したり、パラメータを幻覚させたり、あるいは、どれだけ強制しようとしても NDJSON ではなく通常の JSON を返したりすることでした。</p><p><a href="https://www.elastic.co/search-labs/blog/llm-functions-elasticsearch-intelligent-query">この記事</a>（<a href="https://www.elastic.co/docs/solutions/search/search-templates">検索テンプレートが</a>LLM フリースタイルよりもうまく機能した）に触発され、完全な NDJSON ファイルを生成するように要求するのではなく、テンプレートを LLM に提供し、コード内で LLM によって提供されたパラメータを使用して適切な視覚化を作成することにしました。このアプローチは期待を裏切らず、予測可能で拡張可能です。LLM ではなくコードが重い処理を実行するようになったためです。</p><p>アプリケーションのワークフローは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f7738a4c7ddd0cd/6a1707d52b835f0a25f4b166/52c587cf0cf3517fdd4ee7ab95581dd4f2bce030-725x668.png" alt="" /><p></p><p><em>簡潔にするために一部のコードは省略しますが、完全なアプリケーションの動作コードは</em><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/from-image-idea-to-kibana-dashboard-using-ai/from-image-idea-to-kibana-dashboard-using-ai.ipynb"><em><strong>この</strong></em></a><em>ノートブックに記載されています。</em></p><h2>要件</h2><p>開発を始める前に、次のものが必要です。</p><ol><li><p>Python 3.8以上</p></li><li><p><a href="https://docs.python.org/3/library/venv.html">Venv</a> Python環境</p></li><li><p>実行中のElasticsearchインスタンス、そのエンドポイント、APIキー</p></li><li><p>環境変数名 OPENAI_API_KEY に保存された OpenAI API キー:</p></li></ol>export OPENAI_API_KEY="your-openai-api-key"<h2>データを準備する</h2><p>データについては、シンプルさを保ち、Elastic のサンプル Web ログを使用します。<a href="https://www.elastic.co/docs/manage-data/ingest/sample-data#add-sample-data-sets">ここで、</a>そのデータをクラスターにインポートする方法を学習できます。</p><p>各ドキュメントには、アプリケーションにリクエストを発行したホストの詳細と、リクエスト自体とその応答ステータスに関する情報が含まれています。以下にサンプル文書を示します。</p>{
    "agent": "Mozilla/5.0 (X11; Linux i686) AppleWebKit/534.24 (KHTML, like Gecko) Chrome/11.0.696.50 Safari/534.24",
    "bytes": 8509,
    "clientip": "70.133.115.149",
    "extension": "css",
    "geo": {
        "srcdest": "US:IT",
        "src": "US",
        "dest": "IT",
        "coordinates": {
            "lat": 38.05134111,
            "lon": -103.5106908
        }
    },
    "host": "cdn.elastic-elastic-elastic.org",
    "index": "kibana_sample_data_logs",
    "ip": "70.133.115.149",
    "machine": {
        "ram": 5368709120,
        "os": "osx"
    },
    "memory": null,
    "message": "70.133.115.149 - - [2018-08-30T23:35:31.492Z] \"GET /styles/semantic-ui.css HTTP/1.1\" 200 8509 \"-\" \"Mozilla/5.0 (X11; Linux i686) AppleWebKit/534.24 (KHTML, like Gecko) Chrome/11.0.696.50 Safari/534.24\"",
    "phpmemory": null,
    "referer": "http://twitter.com/error/john-phillips",
    "request": "/styles/semantic-ui.css",
    "response": 200,
    "tags": [
        "success",
        "info"
    ],
    "@timestamp": "2025-07-03T23:35:31.492Z",
    "url": "https://cdn.elastic-elastic-elastic.org/styles/semantic-ui.css",
    "utc_time": "2025-07-03T23:35:31.492Z",
    "event": {
        "dataset": "sample_web_logs"
    },
    "bytes_gauge": 8509,
    "bytes_counter": 51201128
}<p>ここで、先ほどロードしたインデックス<code>kibana_sample_data_logs</code>のマッピングを取得しましょう。</p>INDEX_NAME = "kibana_sample_data_logs"

es_client = Elasticsearch(
    [os.getenv("ELASTICSEARCH_URL")],
    api_key=os.getenv("ELASTICSEARCH_API_KEY"),
)

result = es_client.indices.get_mapping(index=INDEX_NAME)
index_mappings = result[list(result.keys())[0]]["mappings"]["properties"]<p>後で読み込むイメージと一緒にマッピングを渡します。</p><h2>LLM構成</h2><p><a href="https://python.langchain.com/docs/concepts/structured_outputs/">構造化出力</a>を使用して画像を入力し、JSON オブジェクトを生成するために関数に渡す必要がある情報を含む JSON を受け取るように LLM を構成しましょう。</p><p>依存関係をインストールします。</p>pip install elasticsearch pydantic langchain langchain-openai -q<p>Elasticsearch は<a href="https://www.elastic.co/docs/manage-data/data-store/mapping">インデックス マッピングの</a>取得に役立ちます。Pydantic を使用すると、Python でスキーマを定義して LLM に従うように要求することができ、 <a href="https://www.elastic.co/search-labs/integrations/langchain">LangChain は</a>LLM と AI ツールの呼び出しを容易にするフレームワークです。</p><p>LLM から必要な出力を定義するために、Pydantic スキーマを作成します。画像からわかる必要があるのは、グラフの種類、フィールド、視覚化タイトル、ダッシュボード タイトルです。</p>class Visualization(BaseModel):
    title: str = Field(description="The dashboard title")
    type: List[Literal["pie", "bar", "metric"]]
    field: str = Field(
        description="The field that this visualization use based on the provided mappings"
    )


class Dashboard(BaseModel):
    title: str = Field(description="The dashboard title")
    visualizations: List[Visualization]<p>画像入力には、先ほど描いたダッシュボードを送信します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7870f6421986d11d/6a1707d78b73cb3408189fa3/36441d7b5dc1f3ff2ac2a30710208d57ad41c716-1600x898.jpg" alt="" /><p>ここで、LLM モデルの呼び出しとイメージの読み込みを宣言します。この関数は、Elasticsearch インデックスのマッピングと、生成するダッシュボードの画像を受け取ります。</p><p><code>with_structured_output</code>を使用すると、Pydantic <code>Dashboard</code>スキーマを LLM が生成する応答オブジェクトとして使用できます。<a href="https://docs.pydantic.dev/latest/">Pydantic</a>を使用すると、検証付きのデータ モデルを定義できるため、LLM 出力が期待される構造と一致することが保証されます。</p><p>画像を base64 に変換して入力として送信するには、<a href="https://www.base64-image.de/">オンライン コンバーターを</a>使用するか、<a href="https://www.geeksforgeeks.org/python-convert-image-to-string-and-vice-versa/">コードで</a>実行します。</p>prompt = f"""
    You are an expert in analyzing Kibana dashboards from images for the version 9.0.0 of Kibana.

    You will be given a dashboard image and an Elasticsearch index mapping.

    Below are the index mappings for the index that the dashboard is based on.
    Use this to help you understand the data and the fields that are available.

    Index Mappings:
    {index_mappings}

    Only include the fields that are relevant for each visualization, based on what is visible in the image.
    """

message = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": prompt},
            {
                "type": "image",
                "source_type": "base64",
                "data": image_base64,
                "mime_type": "image/png",
            },
        ],
    }
]


try:
    llm = init_chat_model("gpt-4.1-mini")
    llm = llm.with_structured_output(Dashboard)
    dashboard_values = llm.invoke(message)

    print("Dashboard values generated by the LLM successfully")
    print(dashboard_values)
except Exception as e:
    print(f"Failed to analyze image and match fields: {str(e)}")<p>LLM にはすでに Kibana ダッシュボードに関するコンテキストがあるため、プロンプトですべてを説明する必要はなく、Elasticsearch と Kibana で動作していることを忘れないようにするための詳細のみを説明します。</p><p>プロンプトを分解してみましょう:</p><p>セクション</p><p>理由</p><p>あなたは、Kibana バージョン 9.0.0 の画像から Kibana ダッシュボードを分析するエキスパートです。</p><p>これを Elasticsearch と Elasticsearch バージョンで強化することで、LLM が古い/無効なパラメータを幻覚する可能性を減らします。</p><p>ダッシュボード イメージと Elasticsearch インデックス マッピングが提供されます。</p><p>LLM による誤った解釈を避けるために、この画像はダッシュボードに関するものであることを説明します。</p><p>以下は、ダッシュボードのベースとなるインデックスのインデックス マッピングです。これを使用すると、使用可能なデータとフィールドを理解するのに役立ちます。インデックス マッピング: {index_mappings}</p><p>LLM が有効なフィールドを動的に選択できるようにマッピングを提供することが重要です。そうしないと、ここでのマッピングをハードコードすることになり、厳しすぎることになります。あるいは、正しいフィールド名を含むイメージに依存することになりますが、これは信頼できません。</p><p>画像に表示されている内容に基づいて、各視覚化に関連するフィールドのみを含めます。</p><p>画像に関係のないフィールドを追加しようとすることがあるため、この強化を追加する必要がありました。</p><p>これにより、表示する視覚化の配列を含むオブジェクトが返されます。</p>"Dashboard values generated by the LLM successfully
title=""Client, Extension, OS, and Response Keyword Analysis""visualizations="[
   "Visualization(title=""Count of Client IP",
   "type="[
      "metric"
   ],
   "field=""clientip"")",
   "Visualization(title=""Extension Keyword Distribution",
   "type="[
      "pie"
   ],
   "field=""extension.keyword"")",
   "Visualization(title=""Most Used OS",
   "type="[
      "bar"
   ],
   "field=""machine.os.keyword"")",
   "Visualization(title=""Response Keyword Distribution",
   "type="[
      "bar"
   ],
   "field=""response.keyword"")"
]<h2>LLM応答の処理</h2><p>私たちはサンプルの 2x2 パネル ダッシュボードを作成し、 <a href="https://www.elastic.co/docs/api/doc/kibana/operation/operation-get-dashboards-dashboard">Get a dashboard API を</a>使用して JSON 形式でエクスポートしました。その後、パネルを視覚化テンプレート (円グラフ、棒グラフ、メトリック) として保存し、いくつかのパラメータを置き換えて、質問に応じて異なるフィールドを持つ新しい視覚化を作成できます。</p><p>テンプレート JSON ファイルは<a href="https://github.com/Delacrobix/elasticsearch-labs/tree/supporting-blog-content/from-image-idea-to-kibana-dashboard-using-ai/supporting-blog-content/from-image-idea-to-kibana-dashboard-using-ai/templates"><strong>ここで</strong></a>確認できます。後で置き換えたいオブジェクトの値を {<code>variable_name</code>} に変更したことに注意してください。
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc55d69d84a08e668/6a1707d8a2929903acd00fb8/ec7e1ac0cd8b470df13e60940162b56778acb386-315x234.png" alt="" /><p>LLM が提供した情報を使用して、どのテンプレートを使用し、どの値を置き換えるかを決定できます。</p><p><code>fill_template_with_analysis</code> 視覚化の JSON テンプレート、タイトル、フィールド、グリッド上の視覚化の座標など、単一のパネルのパラメータを受け取ります。</p><p>次に、テンプレートの値を置き換えて、最終的な JSON 視覚化を返します。</p>def fill_template_with_analysis(
    template: Dict[str, Any],
    visualization: Visualization,
    grid_data: Dict[str, Any],
):
    template_str = json.dumps(template)
    replacements = {
	 "{visualization_id}": str(uuid.uuid4()),
        "{title}": visualization.title,
        "{x}": grid_data["x"],
        "{y}": grid_data["y"],
    }

    if visualization.field:
        replacements["{field}"] = visualization.field

    for placeholder, value in replacements.items():
        template_str = template_str.replace(placeholder, str(value))

    return json.loads(template_str)<p>簡単にするために、LLM が作成することを決定したパネルに割り当てる静的座標があり、上の画像のように 2x2 グリッド ダッシュボードが生成されます。</p># Filling templates fields
panels = []    
grid_data = [
    {"x": 0, "y": 0},
    {"x": 12, "y": 0},
    {"x": 0, "y": 12},
    {"x": 12, "y": 12},
]


i = 0

for vis in dashboard_values.visualizations:
    for vis_type in vis.type:
        template = templates.get(vis_type, templates.get("bar", {}))
        filled_panel = fill_template_with_analysis(template, vis, grid_data[i])
        panels.append(filled_panel)
        i += 1<p>LLM によって決定された視覚化タイプに応じて、JSON ファイル テンプレートを選択し、 <code>fill_template_with_analysis</code>を使用して関連情報を置き換え、後でダッシュボードを作成するために使用する配列に新しいパネルを追加します。</p><p>ダッシュボードの準備ができたら、<a href="https://www.elastic.co/docs/api/doc/kibana/operation/operation-post-dashboards-dashboard-id">ダッシュボードの作成 API</a>を使用して新しい JSON ファイルを Kibana にプッシュし、ダッシュボードを生成します。
</p>try:
    dashboard_id = str(uuid.uuid4())

    # post request to create the dashboard endpoint
    url = f"{os.getenv('KIBANA_URL')}/api/dashboards/dashboard/{dashboard_id}"

    dashboard_config = {
        "attributes": {
            "title": dashboard_values.title,
            "description": "Generated by AI",
            "timeRestore": True,
            "panels": panels,  # Visualizations with the values generated by the LLM
            "timeFrom": "now-7d/d",
            "timeTo": "now",
        },
    }

    headers = {
        "Content-Type": "application/json",
        "kbn-xsrf": "true",
        "Authorization": f"ApiKey {os.getenv('ELASTICSEARCH_API_KEY')}",
    }

    requests.post(
        url,
        headers=headers,
        json=dashboard_config,
    )

    # Url to the generated dashboard
    dashboard_url = f"{os.getenv('KIBANA_URL')}/app/dashboards#/view/{dashboard_id}"

    print("Dashboard URL: ", dashboard_url)
    print("Dashboard ID: ", dashboard_id)

except Exception as e:
    print(f"Failed to create dashboard: {str(e)}")<p>スクリプトを実行してダッシュボードを生成するには、コンソールで次のコマンドを実行します。</p>python &lt;file_name&gt;.py<p>最終結果は次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ceffed004153a4f/6a1707d9a929cf9147ae0901/e909afbf0e47d9a6e0f7bd07dfb2efcfa5cf06ac-921x715.png" alt="" /><h2>まとめ</h2><p>LLM は、テキストをコード化したり、画像をコード化したりするときに、強力な視覚機能を発揮します。ダッシュボード API を使用すると、JSON ファイルをダッシュボードに変換することも可能で、LLM といくつかのコードを使用して、画像を Kibana ダッシュボードに変換することもできます。</p><p>次のステップは、さまざまなグリッド設定、ダッシュボードのサイズ、位置を使用して、ダッシュボードのビジュアルの柔軟性を向上させることです。また、より複雑な視覚化と視覚化タイプのサポートを提供することも、このアプリケーションにとって便利な追加機能となるでしょう。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-powered-dashboards</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-powered-dashboards</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo,Tomás Murúa]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt41727cbee6155a68/6a1707dbb0367dd2fd72bc86/eb60ceb2fbc3941745b21ae3357cbb6ea8fab18c-1443x811.png" length="0" type="image/png"/>
    <pubDate>Wed, 16 Jul 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[JavaScript、Mastra、Elasticsearch を使用したエージェント型 RAG アシスタントの構築]]></title>
    <description><![CDATA[JavaScript エコシステムで AI エージェントを構築する方法を学ぶ]]></description>
    <content:encoded><![CDATA[<p>このアイデアは、白熱したハイリスクなファンタジー バスケットボール リーグの最中に思いつきました。私はこう考えました。 <em>「毎週の対戦で優位に立つのに役立つ AI エージェントを構築できるだろうか？」 もちろんです!</em></p><p>この記事では、 <a href="https://mastra.ai/en/docs">Mastra</a>とそれと対話するための軽量 JavaScript Web アプリケーションを使用して、エージェント RAG アシスタントを構築する方法について説明します。このエージェントを Elasticsearch に接続することで、構造化されたプレイヤーデータへのアクセスとリアルタイムの統計集計の実行が可能になり、プレイヤー統計に基づいた推奨事項を提供できるようになります。GitHub<a href="https://github.com/jdarmada/nba-ai-assistant-js.git">リポジトリ</a>にアクセスして手順を確認してください。README<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/README.md">に</a>は、アプリケーションを独自に複製して実行する方法が記載されています。 </p><p>すべてをまとめると次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt63ea3e7a09306fbf/6a17f1d97f6f150e22c09c50/1c73bd1dc1b5fe54f025c7a2b7c322acc9122f3a-1999x1393.png" alt="" /><p>注: このブログ投稿は、「 <a href="https://www.elastic.co/search-labs/blog/ai-agents-ai-sdk-elasticsearch">AI SDK と Elastic を使用した AI エージェントの構築</a>」に基づいています。AI エージェント全般とその用途についてよく知らない場合は、まずそこから始めてください。
</p><h2><strong>アーキテクチャの概要</strong></h2><p>システムの中核となるのは、エージェントの推論エンジン（脳）として機能する大規模言語モデル（LLM）です。ユーザー入力を解釈し、呼び出すツールを決定し、関連する応答を生成するために必要な手順を調整します。</p><p>エージェント自体は、JavaScript エコシステムのエージェント フレームワークである Mastra によって構築されます。Mastra は、LLM をバックエンド インフラストラクチャでラップし、それを API エンドポイントとして公開し、ツール、システム プロンプト、エージェントの動作を定義するためのインターフェイスを提供します。</p><p>フロントエンドでは、 <a href="https://vite.dev/guide/">Vite</a>を使用して、エージェントにクエリを送信してその応答を受信するためのチャット インターフェイスを提供する React Web アプリケーションを迅速に構築します。</p><p>最後に、エージェントがクエリして集計できるプレーヤーの統計情報と対戦データを保存する Elasticsearch があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte13f09493f217047/6a17f1db1d1b83178d93e546/443bdc00d84ed1dd49e9f9e431e86ca4b0892563-1999x977.png" alt="" /><h2><strong>背景</strong></h2><p>いくつかの基本的な概念を確認してみましょう。</p><h3><strong>エージェントRAGとは何ですか？</strong></h3><p>AI エージェントは他のシステムと対話し、独立して動作し、定義されたパラメータに基づいてアクションを実行できます。Agentic RAG は、AI エージェントの自律性と検索拡張生成の原理を組み合わせ、LLM が応答を生成するために呼び出すツールとコンテキストとして使用するデータを選択できるようにします。RAG の詳細については、<a href="https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag">こちらを</a>ご覧ください。</p><h3><strong>フレームワークを選択する場合、なぜ AI-SDK を超えるのでしょうか?</strong></h3><p>利用可能な AI エージェント フレームワークは数多くあり、 <a href="https://www.elastic.co/search-labs/blog/using-crewai-with-elasticsearch">CrewAI</a> 、 <a href="https://www.elastic.co/search-labs/blog/using-autogen-with-elasticsearch">AutoGen</a> 、 <a href="https://www.elastic.co/search-labs/blog/build-rag-workflow-langgraph-elasticsearch">LangGraph</a>などの人気のフレームワークについてはおそらく聞いたことがあるでしょう。これらのフレームワークのほとんどは、さまざまなモデルのサポート、ツールの使用、メモリ管理など、共通の機能セットを共有しています。</p><p>こちらは、Harrison Chase (LangChain CEO) によるフレームワーク<a href="https://docs.google.com/spreadsheets/d/1B37VxTBuGLeTSPVWtz7UMsCdtXrqV5hCjWkbHN8tfAo/edit?gid=0#gid=0">比較シート</a>です。</p><p>私が Mastra に興味を持ったのは、フルスタック開発者がエージェントをエコシステムに簡単に統合できるように構築された JavaScript ファーストのフレームワークであるという点です。Vercel の AI-SDK もこのほとんどを実行しますが、プロジェクトにさらに複雑なエージェント ワークフローが含まれている場合は、Mastra が真価を発揮します。Mastra は AI-SDK によって設定された基本パターンを強化しており、このプロジェクトではそれらを連携して使用します。</p><h3><strong>フレームワークとモデル選択の考慮事項</strong></h3><p>これらのフレームワークは AI エージェントを迅速に構築するのに役立ちますが、考慮すべき欠点もいくつかあります。たとえば、AI エージェントや一般的な抽象化レイヤー以外のフレームワークを使用する場合、制御が少し失われます。LLM がツールを正しく使用しなかったり、望ましくないことを実行したりする場合、抽象化によってデバッグが難しくなります。それでも、私の意見では、特にこれらのフレームワークは勢いを増しており、継続的に反復されているため、このトレードオフは、構築時に得られる容易さとスピードの価値があります。</p><p>繰り返しになりますが、これらのフレームワークはモデルに依存しません。つまり、さまざまなモデルをプラグ アンド プレイできます。モデルはトレーニングに使用されたデータ セットによって異なり、その結果、モデルが提供する応答も異なることに注意してください。一部のモデルではツールの呼び出しすらサポートされていません。したがって、さまざまなモデルを切り替えてテストし、どのモデルが最適な応答を返すかを確認することは可能ですが、それぞれのシステム プロンプトを書き換える必要がある可能性が高いことに注意してください。例えば、Llama3.3を使用する場合GPT-4o よりも、必要な応答を得るために、より多くのプロンプトと具体的な指示が必要になります。</p><h3><strong>NBAファンタジーバスケットボール</strong></h3><p>ファンタジー バスケットボールでは、友達のグループでリーグを開始し (グループの競争力に応じて、友情のステータスに影響する可能性があります)、通常はいくらかのお金が賭けられます。その後、各自が 10 人のプレイヤーでチームを編成し、毎週交互に他の友達の 10 人のプレイヤーと対戦します。全体のスコアに加算されるポイントは、特定の週に各プレイヤーが対戦相手に対して行ったパフォーマンスです。</p><p>チームの選手が負傷したり、出場停止になったりした場合は、チームに追加できるフリーエージェント選手のリストが表示されます。ファンタジー スポーツでは、選べる選手の数が限られており、誰もが常に最高の選手を選ぶために奔走しているため、ここで多くの難しい思考が生まれます。</p><p>これは、どの選手を選択するかをすぐに決定しなければならない状況で特に役立つ、NBA AI アシスタントの出番です。特定の対戦相手に対するプレーヤーのパフォーマンスを手動で調べる代わりに、アシスタントがそのデータをすばやく見つけて平均を比較し、情報に基づいた推奨事項を提供します。</p><p>エージェント RAG と NBA ファンタジー バスケットボールの基本がわかったので、実際に見てみましょう。</p><h2><strong>プロジェクトの構築</strong></h2><p>途中で行き詰まったり、最初から構築したくない場合は、<a href="https://github.com/jdarmada/nba-ai-assistant-js.git">リポジトリ</a>を参照してください。</p><h3><strong>取り上げる内容</strong></h3><ol><li><p><strong>プロジェクトの足場作り:</strong></p><ol><li><p><strong>バックエンド (Mastra):</strong> npx create mastra@latest を使用してバックエンドをスキャフォールディングし、エージェント ロジックを定義します。</p></li><li><p><strong>フロントエンド (Vite + React):</strong> npm create vite@latest を使用して、エージェントと対話するためのフロントエンド チャット インターフェイスを構築します。</p></li></ol></li><li><p><strong>環境変数の設定</strong></p><ol><li><p>環境変数を管理するには、dotenv をインストールします。</p></li><li><p>.envを作成するファイルを開き、必要な変数を指定します。</p></li></ol></li><li><p><strong>Elasticsearchの設定</strong></p><ol><li><p>Elasticsearch クラスターを起動します (ローカルまたはクラウド上)。</p></li><li><p>公式 Elasticsearch クライアントをインストールします。</p></li><li><p>環境変数にアクセスできることを確認します。</p></li><li><p>クライアントへの接続を確立します。</p></li></ol></li><li><p><strong>NBA データを Elasticsearch に一括取り込み</strong></p><ol><li><p>集計を有効にするには、適切なマッピングを使用してインデックスを作成します。</p></li><li><p>プレイヤーのゲーム統計を CSV ファイルから Elasticsearch インデックスに一括取り込みます。</p></li></ol></li><li><p><strong>Elasticsearchの集計を定義する</strong></p><ol><li><p>特定の対戦相手に対する過去の平均を計算するクエリ。</p></li><li><p>特定の対戦相手に対するシーズン平均を計算するクエリ。</p></li></ol></li><li><p><strong>プレーヤー比較ユーティリティファイル</strong></p><ol><li><p>ヘルパー関数と Elasticsearch 集計を統合します。</p></li></ol></li><li><p><strong>エージェントの構築</strong></p><ol><li><p>エージェント定義とシステム プロンプトを追加します。</p></li><li><p>zod をインストールし、ツールを定義します。</p></li><li><p>CORS を処理するためのミドルウェア設定を追加します。</p></li></ol></li><li><p><strong>フロントエンドの統合</strong></p><ol><li><p>AI-SDK の useChat を使用してエージェントと対話します。</p></li><li><p>適切にフォーマットされた会話を保持するための UI を作成します。</p></li></ol></li><li><p><strong>アプリケーションの実行</strong></p><ol><li><p>バックエンド (Mastra サーバー) とフロントエンド (React アプリ) の両方を起動します。</p></li><li><p>サンプルクエリと使用方法。</p></li></ol></li><li><p><strong>次はエージェントのさらなるインテリジェント化</strong></p><ol><li><p>セマンティック検索機能を追加して、より洞察力のある推奨を可能にします。</p></li><li><p>検索ロジックを Elasticsearch MCP (Model Context Protocol) サーバーに移動することで、動的クエリを有効にします。</p></li></ol></li></ol><h3><strong>要件</strong></h3><ul><li><p><strong>Node.js と npm</strong> : バックエンドとフロントエンドの両方が Node 上で実行されます。Node 18+ と npm v9+ (Node 18+ にバンドルされています) がインストールされていることを確認してください。</p></li><li><p><strong>Elasticsearch クラスター:</strong>ローカルまたはクラウド上のアクティブな Elasticsearch クラスター。</p></li><li><p><strong>OpenAI API キー</strong>: <a href="https://platform.openai.com/api-keys">OpenAI 開発者ポータルの</a>API キー ページで生成します。</p></li></ul><p></p><h3><strong>プロジェクト構造</strong></h3><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt749baa120552e4ab/6a17f1dd1d1b83bfe993e54a/1c0bde11ad0eead523a95e03b9b905aa776e3fd1-1420x934.png" alt="" /><h4><strong>ステップ1：プロジェクトの足場作り</strong></h4><ol><li><p>まず、nba-ai-assistant-js ディレクトリを作成し、次のコマンドを使用して内部に移動します。 </p></li></ol>mkdir nba-ai-assistant-js &amp;&amp; cd nba-ai-assistant-js<p><strong>バックエンド:</strong></p><ol><li><p>次のコマンドで Mastra 作成ツールを使用します。 </p></li></ol>npx create-mastra@latest<p>2. ターミナルにいくつかのプロンプトが表示されます。最初のプロンプトでは、プロジェクトに backend という名前を付けます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt65abf68fe588e968/6a17f1de63baff2814741d5b/de2725031ed6837db99a979efcdd0ece1e197dbb-608x84.png" alt="" /><p>3. 次に、Mastra ファイルを保存するためのデフォルトの構造を維持するため、 <code>src/</code>を入力します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt89bd829fcf0ae6b9/6a17f1e04b055dd30e432302/88919d9ff1852126395e1fcd700ecb1b59aac63c-866x116.png" alt="" /><p>4. 次に、デフォルトの LLM プロバイダーとして OpenAI を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfd167cc77a40b9a8/6a17f1e11480099e29b48863/2328761e769f3ded134e5a21e8a0bf8f41e88f68-404x210.png" alt="" /><p>5. 最後に、OpenAI API キーの入力が求められます。ここでは、スキップするオプションを選択し、後で<code> .env</code>ファイルで提供します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt12654151ed495370/6a17f1e22f4a5c0f84fa89f9/0662de9bd28758e377e4c63df8d08b479068ce63-444x120.png" alt="" /><p><strong>フロントエンド：</strong></p><ol><li><p>ルート ディレクトリに戻り、次のコマンドを使用して<a href="https://vite.dev/guide/">Vite 作成ツール</a>を実行します。 <code>npm create vite@latest frontend -- --template react</code></p></li></ol><p>これにより、React 専用のテンプレートを使用して、 <code>frontend</code>という名前の軽量 React アプリが作成されます。</p><p>すべてがうまくいけば、プロジェクト ディレクトリ内に、Mastra コードを保持するバックエンド ディレクトリと、React アプリを含む<code>frontend</code>ディレクトリが表示されるはずです。</p><p></p><h4><strong>ステップ2: 環境変数の設定</strong></h4><ol><li><p>機密キーを管理するために、 <code>dotenv</code>パッケージを使用して.envから環境変数を読み込みます。ファイル。バックエンドディレクトリに移動して<code>dotenv</code>をインストールします。</p></li></ol>cd backend
npm install dotenv --save<p>2. バックエンド ディレクトリでは、適切な変数を入力するための example.env ファイルが提供されます。独自に作成する場合は、次の変数を必ず含めてください。</p># OpenAI Configuration
OPENAI_API_KEY=your_openai_api_key_here

# Elasticsearch Configuration
ELASTIC_ENDPOINT=your_elasticsearch_endpoint_here
ELASTIC_API_KEY=your_elasticsearch_api_key_here
<p></p><p>注意: <code>.env</code> <code>.gitignore</code>に追加して、このファイルがバージョン管理から除外されていることを確認してください。</p><h4><strong>ステップ3: Elasticsearchの設定</strong></h4><p>まず、アクティブな Elasticsearch クラスターが必要です。次の 2 つのオプションがあります。</p><ul><li><p><strong>オプションA: Elasticsearch Cloudを使用する</strong></p><ul><li><p><a href="https://cloud.elastic.co/registration">Elastic Cloud</a>にサインアップ</p></li><li><p>新しいデプロイメントを作成する</p></li><li><p>エンドポイント URL と API キー（エンコード済み）を取得します</p></li></ul></li><li><p><strong>オプションB: Elasticsearchをローカルで実行する</strong></p><ul><li><p>Elasticsearchをローカルにインストールして実行する</p></li><li><p>エンドポイントとして http://localhost:9200 を使用します</p></li><li><p>APIキーを生成する</p></li></ul></li></ul><p></p><p><strong>バックエンドに Elasticsearch クライアントをインストールする:</strong></p><ol><li><p>まず、バックエンド ディレクトリに公式 Elasticsearch クライアントをインストールします。</p></li></ol>npm install @elastic/elasticsearch<p>2. 次に、再利用可能な関数を保持するディレクトリ lib を作成し、そこに移動します。</p>mkdir lib &amp;&amp; cd lib<p>3. 内部に<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/lib/elasticClient.js">elasticClient.js</a>という新しいファイルを作成します。このファイルは Elasticsearch クライアントを初期化し、プロジェクト全体で使用できるように公開します。</p><p>4. ECMAScript モジュール (ESM) を使用しているため、 __dirname and __ファイル名は使用できません。環境変数が.envから正しく読み込まれていることを確認するにはバックエンド フォルダー内のファイルで、ファイルの先頭に次の設定を追加します。</p>import { config } from 'dotenv';
import { fileURLToPath } from 'url';
import { dirname, join } from 'path';
import { Client } from '@elastic/elasticsearch';

// Grab current directory and load .env from backend folder
const __filename = fileURLToPath(import.meta.url);
const __dirname = dirname(__filename);
const envPath = join(__dirname, '../.env');

// Load environment variables from the correct path
config({ path: envPath });<p>5. 次に、環境変数を使用して Elasticsearch クライアントを初期化し、接続を確認します。</p>//Elastic client Initialization, make sure environment variables are being loaded in correctly
const config= {
    node: `${process.env.ELASTIC_ENDPOINT}`,
    auth: {
        apiKey: `${process.env.ELASTIC_API_KEY}`,
    },
};

export const elasticClient = new Client(config);

//Check if the client is connected
async function checkConnection() { 
    try {
        const info = await elasticClient.info();
        console.log('Elasticsearch is connected:', info);
    } catch (error) {
        console.error('Elasticsearch connection error:', error);
    }
}

checkConnection();
<p>これで、このクライアント インスタンスを、Elasticsearch クラスターと対話する必要がある任意のファイルにインポートできます。</p><p></p><h4><strong>ステップ4: NBAデータをElasticsearchに一括取り込み</strong></h4><p><strong>データセット:</strong></p><p>このプロジェクトでは、リポジトリの<a href="https://github.com/jdarmada/nba-ai-assistant-js/tree/main/backend">backend/data</a>ディレクトリにあるデータセットを参照します。当社の NBA アシスタントは、このデータを知識ベースとして使用し、統計的な比較を実行し、推奨事項を生成します。</p><ul><li><p><a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/data/sample_nba_data.csv">sample_player_game_stats.csv</a> - サンプルプレーヤーのゲーム統計 (例: NBA キャリア全体におけるプレーヤーごとのゲームごとのポイント、リバウンド、スティールなど)。このデータセットを使用して集計を実行します。(注: これはデモ用に事前に生成された模擬データであり、公式 NBA ソースから取得されたものではありません。)</p></li><li><p><a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/data/playerAndTeamInfo.js">playerAndTeamInfo.js</a> - 通常は API 呼び出しによって提供されるプレーヤーとチームのメタデータを置き換え、エージェントがプレーヤーとチームの名前を ID に一致できるようにします。サンプル データを使用しているため、外部 API から取得する際のオーバーヘッドを避け、エージェントが参照できるいくつかの値をハードコードしました。</p></li></ul><p></p><p><strong>実装：</strong></p><ol><li><p><code>backend/lib</code>ディレクトリで、 <a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/lib/playerDataIngestion.js">playerDataIngestion.js</a>という名前のファイルを作成します。</p></li><li><p>インポートを設定し、CSV ファイル パスを解決し、解析を設定します。ここでも、ESM を使用しているため、サンプル CSV へのパスを解決するには<code>__dirname</code>を再構築する必要があります。また、 <a href="http://node.js/">Node.js</a>の組み込みモジュール<code>fs</code>と<code>readline</code>を使用して、指定された CSV ファイルを行ごとに解析します。</p></li></ol>import fs from 'fs';
import readline from 'readline';
import path from 'path';
import { fileURLToPath } from 'url';
import { elasticClient } from './elasticClient.js';

const indexName = 'sample-nba-player-data'; //Replace with your preferred index name

//Since we are using ES modules __dirname and __filename don't exist, so this is a workaround that allows us to use the absolute file path for our sample data.
const __filename = fileURLToPath(import.meta.url);
const __dirname = path.dirname(__filename);
const filePath = path.resolve(__dirname, '../data/sample_nba_data.csv');<p>これにより、一括取り込み手順で CSV を効率的に読み取って解析できるようになります。</p><p>3. 適切なマッピングを使用してインデックスを作成します。Elasticsearch は<a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/dynamic">動的マッピングを</a>使用してフィールド タイプを自動的に推測できますが、ここでは各統計が数値フィールドとして扱われるように明示的に指定します。これらのフィールドは後で集計に使用するため、これは重要です。また、ポイントやリバウンドなどの統計情報には、小数値が含まれるようにするために、タイプ<code>float </code>を使用します。最後に、Elasticsearch が認識されないフィールドを動的にマッピングしないように、マッピング プロパティ<code>dynamic: 'strict'</code>を追加します。
</p>// Function to create an index with mappings
async function createIndex() {
    try {
        // Check if the index already exists
        const exists = await elasticClient.indices.exists({ index: indexName });

        if (exists) {
            console.log(`Index "${indexName}" already exists, deleting it now.`);
            await elasticClient.indices.delete({ index: indexName });
            console.log(`Deleted index "${indexName}".`);
        }
        // Create the index with mappings
        const response = await elasticClient.indices.create({
            index: indexName,
            body: {
                mappings: {
                    dynamic: 'strict', // Prevent dynamic mapping
                    properties: {
                        game_id: { type: 'integer' },
                        game_date: { type: 'date' },
                        player_id: { type: 'integer' },
                        player_full_name: { type: 'text' },
                        player_team_id: { type: 'integer' },
                        player_team_name: { type: 'text' },
                        home_team: { type: 'boolean' },
                        opponent_team_id: { type: 'integer' },
                        opponent_team_name: { type: 'text' },
                        points: { type: 'float' },
                        rebounds: { type: 'float' },
                        assists: { type: 'float' },
                        steals: { type: 'float' },
                        blocks: { type: 'float' },
                        fg_percentage: { type: 'float' },
                        minutes_played: { type: 'float' },
                    },
                },
            },
        });

        console.log('Index created:', response);
        return true;
    } catch (error) {
        console.error('Error creating index:', error);
        return false;
    }
}
<p>4. CSV データを Elasticsearch インデックスに一括で取り込む機能を追加します。コード ブロック内では、ヘッダー行をスキップします。次に、各行項目をコンマで分割し、ドキュメント オブジェクトにプッシュします。このステップでは、それらをクリーンアップし、適切なタイプであることを確認します。次に、ドキュメントをインデックス情報とともに bulkBody 配列にプッシュします。これは、Elasticsearch への一括取り込みのペイロードとして機能します。</p>async function bulkIngestCsv(filePath) {
    const readStream = fs.createReadStream(filePath);
    const rl = readline.createInterface({
        input: readStream,
        crlfDelay: Infinity,
    });

    const bulkBody = [];
    let lineNum = 0;

    //Skip the header line
    let headerLine = true;
    for await (const line of rl) {
        if (headerLine) {
            headerLine = false;
            continue;
        }
        lineNum++;

        // Split the line by comma and remove whitespace
        const [
            game_id,
            game_date,
            player_id,
            player_full_name,
            player_team_id,
            player_team_name,
            home_team,
            opponent_team_id,
            opponent_team_name,
            points,
            rebounds,
            assists,
            steals,
            blocks,
            fg_percentage,
            minutes_played,
        ] = line.split(',');

        // Create a document object
        const document = {
            game_id: parseInt(game_id),
            game_date: game_date.trim(),
            player_id: parseInt(player_id),
            player_full_name: player_full_name.trim(),
            player_team_id: parseInt(player_team_id),
            player_team_name: player_team_name.trim(),
            home_team: home_team.trim() === 'True', // Converts True/False into a boolean
            opponent_team_id: parseInt(opponent_team_id),
            opponent_team_name: opponent_team_name.trim(),
            points: parseFloat(points),
            rebounds: parseFloat(rebounds),
            assists: parseFloat(assists),
            steals: parseFloat(steals),
            blocks: parseFloat(blocks),
            fg_percentage: parseFloat(fg_percentage),
            minutes_played: parseFloat(minutes_played),
        };

        // Prepare the bulk operation format
        bulkBody.push({ index: { _index: indexName } });
        bulkBody.push(document);
    }

    console.log(`Parsed ${lineNum} lines from CSV`);
<p>5.次に、 <code>elasticClient.bulk()</code>で Elasticsearch の<a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk">Bulk API を</a>使用して、1 回のリクエストで複数のドキュメントを取り込むことができます。以下のエラー処理は、取り込みに失敗したドキュメントの数と、取り込みに成功したドキュメントの数を示すように構成されています。</p>try {
        // Perform the bulk request
        const response = await elasticClient.bulk({ body: bulkBody });

        if (response.errors) {
            console.log('Bulk Ingestion had some hiccups:');

            // Count successful vs failed operations
            let successCount = 0;
            let errorCount = 0;
            const errorDetails = [];

            response.items.forEach((item, index) =&gt; {
                const operation = item.index || item.create || item.update || item.delete;
                if (operation.error) {
                    errorCount++;
                    errorDetails.push({
                        document: index + 1,
                        error: operation.error,
                    });
                } else {
                    successCount++;
                }
            });

            console.log(`Successfully indexed: ${successCount} documents`);
            console.log(`Failed to index: ${errorCount} documents, here are the details`, errorDetails);

        } else {
            console.log(`Bulk Ingestion fully successful!`);
        }

    } catch (error) {
        console.error('Error performing bulk ingestion:', error);
    }
}
<p>6. 以下の<code>main()</code>関数を実行して、 <code>createIndex()</code>関数と<code>bulkIngestCsv()</code>関数を順番に実行します。</p>// Run this function
async function main() {
    const result = await createIndex();
    if (!result) {
        console.error('Index setup failed. Aborting.');
        return;
    }

    await bulkIngestCsv(filePath);
    console.log('Bulk ingestion completed!');
}

main();
<p>一括取り込みが成功したことを示すコンソール ログが表示された場合は、Elasticsearch インデックスを簡単にチェックして、ドキュメントが実際に正常に取り込まれたかどうかを確認します。</p><h4><strong>ステップ5: Elasticsearchの集計の定義と統合</strong></h4><p>これらは、プレイヤーの統計を相互に比較するために AI エージェントのツールを定義するときに使用される主な関数になります。</p><p>1. <code>backend/lib</code>ディレクトリに移動し、 <a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/lib/elasticAggs.js">elasticAggs.js</a>というファイルを作成します。</p><p>2. 特定の対戦相手に対するプレイヤーの過去の平均を計算するには、以下のクエリを追加します。このクエリでは、2 つの条件（1 つは<code>player_id</code>に一致し、もう 1 つは<code>opponent_team_id</code>に一致する）を持つ<code>bool</code><a href="https://www.elastic.co/search-labs/tutorials/search-tutorial/full-text-search/filters">フィルター</a>を使用して、関連するゲームのみを取得します。ドキュメントを返す必要はなく、集計のみを対象とするため、 <code>size:0</code>を設定します。<code>aggs</code>ブロックでは、 <code>points, rebounds, assists, steals, blocks</code>や<code>fg_percentage</code>などのフィールドに対して複数のメトリック<a href="https://www.elastic.co/docs/explore-analyze/query-filter/aggregations">集計を</a>並行して実行し、平均値を計算します。LLM は計算で成功するか失敗するかのどちらかですが、このプロセスは Elasticsearch にオフロードされ、NBA AI アシスタントが正確なデータにアクセスできるようになります。</p>export async function getHistoricalAveragesAgainstOpponent(player_id, opponent_team_id) {
    try {
        //Query for Historical Averages
        const historicalQuery = await elasticClient.search({
            index: 'sample-nba-player-data', 
            size: 0,
            query: {
                bool: {
                    must: [
                        {
                            term: {
                                player_id: {
                                    value: player_id,
                                },
                            },
                        },
                        {
                            term: {
                                opponent_team_id: {
                                    value: opponent_team_id,
                                },
                            },
                        },
                    ],
                },
            },
            aggs: {
                avg_points: { avg: { field: 'points' } },
                avg_rebounds: { avg: { field: 'rebounds' } },
                avg_assists: { avg: { field: 'assists' } },
                avg_steals: { avg: { field: 'steals' } },
                avg_blocks: { avg: { field: 'blocks' } },
             avg_fg_percentage: { avg: { field: 'fg_percentage' } },
            },
        });

        return {
            points: historicalQuery.aggregations.avg_points.value || 0,
            rebounds: historicalQuery.aggregations.avg_rebounds.value || 0,
            assists: historicalQuery.aggregations.avg_assists.value || 0,
            steals: historicalQuery.aggregations.avg_steals.value || 0,
            blocks: historicalQuery.aggregations.avg_blocks.value || 0,
            fgPercentage: historicalQuery.aggregations.avg_fg_percentage.value || 0,
        };
    } catch (error) {
        console.error('Query error from getHistoricalAveragesAgainstOpponent function:', error);
        return { error: 'Queries failed in getting historical averages against opponent.' };
    }
}
<p>3. 特定の対戦相手に対するプレーヤーのシーズン平均を計算するには、履歴クエリとほぼ同じクエリを使用します。このクエリの唯一の違いは、 <code>bool</code>フィルターに<code>game_date</code>の追加条件があることです。フィールド<code>game_date</code>は、現在の NBA シーズンの範囲内に収まる必要があります。この場合、範囲は<code>2024-10-01</code>から<code>2025-06-30</code>の間になります。以下の追加条件により、後続の集計で今シーズンのゲームのみが分離されることが保証されます。
</p>        {
                            range: {
                    //Range for this season, change to match current season
                                game_date: {
                                    gte: '2024-10-01',
                                    lte: '2025-06-30',
                                },
                            },
<h4><strong>ステップ6: プレーヤー比較ユーティリティ</strong></h4><p>コードをモジュール化して保守しやすい状態に保つために、メタデータ ヘルパー関数と Elasticsearch 集計を統合するユーティリティ ファイルを作成します。これにより、エージェントが使用するメイン ツールが強化されます。詳細は後述します。</p><p>1. <code>backend/lib</code>ディレクトリに新しいファイル<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/lib/comparePlayers.js">comparePlayers.js</a>を作成します。</p><p>2. 以下の関数を追加して、メタデータ ヘルパーと Elasticsearch 集約ロジックを、エージェントが使用するメイン ツールを強化する単一の関数に統合します。
</p>import { playersByName } from '../data/playerAndTeamInfo.js';
import { teamsByName } from '../data/playerAndTeamInfo.js';
import { upcomingMatchups } from '../data/playerAndTeamInfo.js';
import { getHistoricalAveragesAgainstOpponent } from './elasticAggs.js';
import { getSeasonAveragesAgainstOpponent } from './elasticAggs.js';

//Simple helper functions to simulate API calls for player and team metadata. These reference the hardcoded values from playerAndTeamInfo.js in the data directory
export function getPlayerInfo(playerFullName) {
    return playersByName[playerFullName];
}

export function getTeamID(teamFullName) {
    return teamsByName[teamFullName];
}

export function getUpcomingMatchups(teamId) {
    return upcomingMatchups[teamId];
}

//Main function used by the 'playerComparisonTool' agent tool
export async function comparePlayersForNextMatchup(player1Name, player2Name) {
    //Get Player Info
    const player1Info = getPlayerInfo(player1Name);
    const player2Info = getPlayerInfo(player2Name);

    //Get upcoming matchups
    const player1NextGame = getUpcomingMatchups(player1Info.team_id)[0];
    const player2NextGame = getUpcomingMatchups(player2Info.team_id)[0];

    //Get season and historical averages against next opponent for player 1
    const player1SeasonAverages = await getSeasonAveragesAgainstOpponent(
        player1Info.player_id,
        player1NextGame.opponent_team_id
    );
    const player1HistoricalAverages = await getHistoricalAveragesAgainstOpponent(
        player1Info.player_id,
        player1NextGame.opponent_team_id
    );

    //Get season and historical averages against next opponent for player 2
    const player2SeasonAverages = await getSeasonAveragesAgainstOpponent(
        player2Info.player_id,
        player2NextGame.opponent_team_id
    );
    const player2HistoricalAverages = await getHistoricalAveragesAgainstOpponent(
        player2Info.player_id,
        player2NextGame.opponent_team_id
    );

    const player1 = {
        name: player1Name,
        playerId: player1Info.player_id,
        teamId: player1Info.team_id,
        nextOpponent: {
            teamId: player1NextGame.opponent_team_id,
            teamName: player1NextGame.opponent_team_name,
            home: player1NextGame.home,
        },
        stats: {
            seasonAverages: player1SeasonAverages,
            historicalAverages: player1HistoricalAverages,
        },
    };

    const player2 = {
        name: player2Name,
        playerId: player2Info.player_id,
        teamId: player2Info.team_id,
        nextOpponent: {
            teamId: player2NextGame.opponent_team_id,
            teamName: player2NextGame.opponent_team_name,
            home: player2NextGame.home,
        },
        stats: {
            seasonAverages: player2SeasonAverages,
            historicalAverages: player2HistoricalAverages,
        },
    };

    return [player1, player2];
}
<h4><strong>ステップ7: エージェントの構築</strong></h4><p>フロントエンドとバックエンドのスキャフォールディングを作成し、NBA ゲームデータを取り込み、Elasticsearch への接続を確立したので、すべてのピースをまとめてエージェントを構築し始めることができます。</p><p><strong>エージェントの定義</strong></p><p>1. <code>backend/src/mastra/agents</code>ディレクトリ内の<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/src/mastra/agents/index.ts">index.ts</a>ファイルに移動し、エージェント定義を追加します。次のようなフィールドを指定できます。</p><ul><li><p><strong>名前:</strong>フロントエンドで呼び出されたときに参照として使用されるエージェントの名前を指定します。</p></li><li><p><strong>指示/システム プロンプト:</strong>システム プロンプトは、対話中に従うべき初期コンテキストとルールを LLM に提供します。これは、ユーザーがチャット ボックスを通じて送信するプロンプトに似ていますが、こちらはユーザー入力の前に表示されます。繰り返しになりますが、これは選択したモデルに応じて変わります。</p></li><li><p><strong>モデル:</strong>使用する LLM (Mastra は OpenAI、Anthropic、ローカル モデルなどをサポートしています)。</p></li><li><p><strong>ツール:</strong>エージェントが呼び出すことができるツール関数のリスト。</p></li><li><p><strong>メモリ:</strong> (オプション) エージェントに会話履歴などを記憶させたい場合。簡単にするために、Mastra は永続メモリをサポートしていますが、永続メモリなしで開始できます。</p></li></ul><p></p>import { openai } from '@ai-sdk/openai';
import { Agent } from '@mastra/core/agent';
import { playerComparisonTool } from '../tools';

export const basketballAgent = new Agent({
    name: 'Basketball Agent',
    instructions: `
      You are a NBA Basketball expert.
      Your primary function is to compare two NBA players and recommend which one is the better fantasy pickup.

      Only compare players from the following list:
      - LeBron James
      - Stephen Curry
      - Jayson Tatum
      - Jaylen Brown
      - Nikola Jokic
      - Luka Doncic
      - Kyrie Irving
      - Anthony Davis
      - Kawhi Leonard
      - Russell Westbrook

      Input Handling Rules:
      - If the user asks about a player that is not on this list, respond with the list of available players for comparison.
      - If the user only inputs one player, ask the user to add another player from the list provided.
      - If the user inputs a player with the wrong spelling or capitalizations, infer from the list of available players provided.
      - IMPORTANT: If the user asks a question or asks you to generate a response about anything outside of basketball or the scope of this project, DO NOT answer and affirm you can only talk about basketball.

      Tool Usage:
      - Extract and standardize player names to match the list exactly.
      - Use the playerComparisonTool, passing both names as strings.
      - The tool will return an object with game information, stats, and analysis.

      Format your response using Markdown syntax. Use:

        Example output format:

       
        #### Next Game Info
        - ***LeBron James** vs Warriors, May 24 (Home)  
        - ***Stephen Curry** vs Lakers, May 24 (Away)


        #### Stats Comparison  
        \`\`\`  
        Stat                  LeBron James (vs Warriors)    Stephen Curry (vs Lakers)  
        --------------------  -----------------------------  ----------------------------  
        Historical Points     28.3                          30.3  
        Historical Assists    6.7                           8.7  
        Season Points         28.8                          23.3  
        Season Assists        6.2                           4.7  
        \`\`\`

        #### Fantasy Recommendation  
        Explain which player is the better fantasy pickup and why.
      
    `,
    model: openai('gpt-4o'),
    tools: { playerComparisonTool },
});
<p><strong>
ツールの定義</strong></p><ol><li><p><code>backend/src/mastra/tools</code>ディレクトリ内の<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/src/mastra/tools/index.ts">index.ts</a>ファイルに移動します。</p></li><li><p>次のコマンドを使用して Zod をインストールします。</p></li></ol>npm install zod<p>3. ツール定義を追加します。このツールを呼び出すときにエージェントが使用するメイン関数として、 <code>comparePlayers.js</code>ファイル内の関数をインポートすることに注意してください。Mastra の<code>createTool()</code>関数を使用して、 <code>playerComparisonTool</code>を登録します。フィールドには次のものが含まれます。</p><ul><li><p><code>id</code>: これは、エージェントがツールの機能を理解するのに役立つ自然言語による説明です。</p></li><li><p><code>input schema</code>: ツールの入力の形状を定義するために、Mastra は TypeScript スキーマ検証ライブラリである<a href="https://zod.dev/">Zod</a>スキーマを使用します。Zod は、エージェントが正しく構造化された入力を入力したことを確認し、入力構造が一致しない場合はツールが実行されないようにすることで役立ちます。</p></li><li><p><code>description</code>: これは、エージェントがいつ電話をかけてツールを使用するかを理解するのに役立つ自然言語による説明です。</p></li><li><p><code>execute</code>: ツールが呼び出されたときに実行されるロジック。私たちの場合、インポートしたヘルパー関数を使用してパフォーマンス統計を返します。</p></li></ul>import { comparePlayersForNextMatchup } from '../../../lib/comparePlayers.js'
import { createTool } from "@mastra/core/tools";
import { z } from "zod";

export const playerComparisonTool = createTool({
    id: "Compare two NBA players",
    inputSchema: z.object({
        player1:z.string(),
        player2:z.string()
    }),
    description: "Use this tool to compare two players given in the user prompt.",
    execute: async ({ context: { player1, player2 } }) =&gt; {
        return await comparePlayersForNextMatchup(player1, player2);
      },
})<p><strong>CORSを処理するミドルウェアの追加</strong></p><p><a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/CORS">CORS を</a>処理するために、Mastra サーバーにミドルウェアを追加します。人生には避けられないことが 3 つあると言われています。死、税金、そして Web 開発者にとっては CORS です。簡単に言うと、クロスオリジン リソース共有は、フロントエンドが別のドメインまたはポートで実行されているバックエンドにリクエストを送信するのをブロックするブラウザのセキュリティ機能です。バックエンドとフロントエンドの両方をローカルホストで実行しているにもかかわらず、それらは異なるポートを使用するため、CORS ポリシーがトリガーされます。バックエンドがフロントエンドからのリクエストを許可するように、 <a href="https://mastra.ai/en/docs/server-db/middleware">Mastra ドキュメント</a>で指定されているミドルウェアを追加する必要があります。</p><p>1. <code>backend/src/mastra</code>ディレクトリ内の<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/backend/src/mastra/index.ts">index.ts</a>ファイルに移動し、CORS の設定を追加します。</p><ul><li><p><code>origin: ['http://localhost:5173']</code></p><ul><li><p>このアドレス（Vite のデフォルト アドレス）からのリクエストのみを許可します</p></li></ul></li><li><p><code>allowMethods: ["GET", "POST"]</code></p><ul><li><p>許可される HTTP メソッド。ほとんどの場合、POST が使用されます。</p></li></ul></li><li><p><code>allowHeaders: ["Content-Type", "Authorization", "x-mastra-client-type, "x-highlight-request", "traceparent"],</code></p><ul><li><p>これらはリクエストで使用できるカスタムヘッダーを決定します</p></li></ul></li></ul><p></p>import { Mastra } from '@mastra/core/mastra';
import { basketballAgent } from './agents';

console.log('Starting Mastra server...');

export const mastra = new Mastra({
  agents: { basketballAgent },
  server:{
    timeout: 10 * 60 * 1000, // 10 minutes
    cors: {
      origin: ['http://localhost:5173'],
      allowMethods: ["GET", "POST"],
      allowHeaders: [
        "Content-Type",
        "Authorization",
        "x-mastra-client-type",
        "x-highlight-request",
        "traceparent",
      ],
      exposeHeaders: ["Content-Length", "X-Requested-With"],
      credentials: false,
    },
  },

});

console.log('Mastra server configured.'); // Log after server configuration
<h4><strong>ステップ8: フロントエンドの統合</strong></h4><p>この React コンポーネントは、 <code>@ai-sdk/react</code>の<a href="https://mastra.ai/en/docs/frameworks/agentic-uis/ai-sdk#using-the-usechat-hook">useChat()</a>フックを使用して Mastra AI エージェントに接続するシンプルなチャット インターフェースを提供します。このフックを使用して、トークンの使用状況やツールの呼び出しを表示したり、会話をレンダリングしたりします。上記のシステム プロンプトでは、エージェントに応答をマークダウンで出力するように要求しているため、 <code>react-markdown</code>を使用して応答を適切にフォーマットします。</p><p></p><p>1.フロントエンド ディレクトリにいる間に、useChat() フックを使用するために @ai-sdk/react パッケージをインストールします。</p>npm install @ai-sdk/react<p>2. 同じディレクトリで、React Markdown をインストールして、エージェントが生成する応答を適切にフォーマットできるようにします。</p>npm install react-markdown<p>3. <code>useChat()</code>を実装します。このフックは、フロントエンドと AI エージェントのバックエンド間のやり取りを管理します。メッセージの状態、ユーザー入力、ステータスを処理し、監視の目的でライフサイクル フックを提供します。渡すオプションは次のとおりです。</p><ul><li><p><code>api:</code> これは、Mastra AI エージェントのエンドポイントを定義します。デフォルトではポート 4111 に設定されており、ストリーミング応答をサポートするルートも追加する必要があります。</p></li><li><p><code>onToolCall</code>: これは、エージェントがツールを呼び出すたびに実行されます。エージェントがどのツールを呼び出しているかを追跡するために使用します。</p></li><li><p><code>onFinish</code>: エージェントが完全な応答を完了した後に実行されます。ストリーミングを有効にしても、 <code>onFinish</code>各チャンクの後ではなく、完全なメッセージが受信された後に実行されます。ここでは、トークンの使用状況を追跡するためにこれを使用しています。これは、LLM コストを監視して最適化するときに役立ちます。</p></li></ul><p>4. 最後に、 <code>frontend/components</code>ディレクトリの<a href="https://github.com/jdarmada/nba-ai-assistant-js/blob/main/frontend/components/ChatUI.jsx">ChatUI.jsx</a>コンポーネントに移動して、会話を行うための UI を作成します。次に、エージェントからの応答を適切にフォーマットするために、応答を<code>ReactMarkdown</code>コンポーネントでラップします。</p>import React, { useState } from 'react';
import { useChat } from '@ai-sdk/react';
import ReactMarkdown from 'react-markdown';

export default function ChatUI() {
    const [totalTokenUsage, setTotalTokenUsage] = useState(0);
    const [promptTokenUsage, setPromptTokenUsage] = useState(0);
    const [completionTokenUsage, setCompletionTokenUsage] = useState(0);
    const [toolsCalled, setToolsCalled] = useState([]);

    const { messages, input, handleInputChange, handleSubmit, status } = useChat({
        api: 'http://localhost:4111/api/agents/basketballAgent/stream', //Replace with your own endpoint for your agent
        id: 'my-chat-session',

        //Optional parameter to check agent tool calls
        onToolCall: ({ toolCall }) =&gt; {
            setToolsCalled((prev) =&gt; [...prev, toolCall.toolName]);
        },

        //Optional parameter to check token usages
        onFinish: (message, { usage }) =&gt; {
            setTotalTokenUsage((prev) =&gt; prev + usage.totalTokens);
            setPromptTokenUsage((prev) =&gt; prev + usage.promptTokens);
            setCompletionTokenUsage((prev) =&gt; prev + usage.completionTokens);
        },

        //Optional parameter for error handling
        onError: (error) =&gt; {
            console.error('Agent error:', error);
        },
    });

    return (
        &lt;div&gt;
            &lt;div className="agent-info"&gt;
                &lt;h4 className="stats-title"&gt;What's My Agent Doing?&lt;/h4&gt;

                &lt;div className="stats-box"&gt;
                    &lt;strong className="stats-sub-title"&gt;Tools Called:&lt;/strong&gt;
                    &lt;ul className="tool-list"&gt;
                        {toolsCalled.map((tool, idx) =&gt; (
                            &lt;li key={idx}&gt;{tool}&lt;/li&gt;
                        ))}
                        {toolsCalled.length === 0 &amp;&amp; &lt;li&gt;No tools called yet.&lt;/li&gt;}
                    &lt;/ul&gt;

                    &lt;div className="usage-stats"&gt;
                        &lt;p&gt;Prompt Token Usage: {promptTokenUsage}&lt;/p&gt;
                        &lt;p&gt;Completion Token Usage: {completionTokenUsage}&lt;/p&gt;
                        &lt;p&gt;Total Token Usage: {totalTokenUsage}&lt;/p&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
            &lt;/div&gt;

            &lt;strong&gt;Conversation:&lt;/strong&gt;
            &lt;div className="convo-box"&gt;
                {messages.map((msg) =&gt; (
                    &lt;div key={msg.id} className="message-item"&gt;
                        &lt;strong className="message-role"&gt;{msg.role === 'assistant' ? 'Basketbot' : 'You'}:&lt;/strong&gt;
                        &lt;ReactMarkdown&gt;{msg.content}&lt;/ReactMarkdown&gt;
                    &lt;/div&gt;
                ))}
            &lt;/div&gt;

            &lt;form onSubmit={handleSubmit}&gt;
                &lt;input
                    type="text"
                    value={input}
                    onChange={handleInputChange}
                    placeholder="Input two players you want to compare."
                    className="input-box"
                /&gt;
                &lt;button type="submit" disabled={status === 'streaming'}&gt;
                    {status === 'streaming' ? 'Thinking...' : 'Send'}
                &lt;/button&gt;
            &lt;/form&gt;
        &lt;/div&gt;
    );
}<h4><strong>ステップ9: アプリケーションの実行</strong></h4><p>おめでとうございます！これでアプリケーションを実行する準備が整いました。バックエンドとフロントエンドの両方を起動するには、次の手順に従います。</p><ol><li><p>ターミナル ウィンドウで、ルート ディレクトリからバックエンド ディレクトリに移動し、Mastra サーバーを起動します。</p></li></ol>cd backend

npm run dev<p>2. 別のターミナル ウィンドウで、ルート ディレクトリからフロントエンド ディレクトリに移動し、React アプリを起動します。</p><p></p>cd frontend

npm run dev<p></p><p>3. ブラウザで次の場所に移動します。</p><p></p><p><a href="http://localhost:5173/">http://localhost:5173</a></p><p></p><p>チャット インターフェースが表示されるはずです。次のサンプルプロンプトを試してみてください。</p><ul><li><p>「レブロン・ジェームズとステフィン・カリーを比較」</p></li><li><p>「ジェイソン・テイタムとルカ・ドンチッチのどちらを選ぶべきでしょうか？」</p></li></ul><p></p><h3><strong>次はエージェントのさらなるインテリジェント化</strong></h3><p>アシスタントをよりエージェント的にし、推奨事項をより洞察力のあるものにするために、次のイテレーションでいくつかの重要なアップグレードを追加する予定です。</p><p></p><p><strong>NBAニュースのセマンティック検索</strong></p><p>プレーヤーのパフォーマンスに影響を与える要因は数多くありますが、その多くは生の統計には表示されません。負傷報告、ラインナップの変更、さらには試合後の分析などは、ニュース記事でしか見つけることができません。この追加のコンテキストを捉えるために、エージェントが関連する NBA の記事を取得し、その内容を推奨事項に組み込めるよう、セマンティック検索機能を追加します。</p><p></p><p><strong>Elasticsearch MCPサーバーによる動的検索</strong></p><p>MCP (モデル コンテキスト プロトコル) は、エージェントがデータ ソースに接続する方法の標準として急速に普及しつつあります。検索ロジックを Elasticsearch MCP サーバーに移行します。これにより、エージェントは、私たちが提供する定義済みの検索機能に頼るのではなく、動的にクエリを構築できるようになります。これにより、より自然な言語ワークフローを使用できるようになり、すべての検索クエリを手動で記述する必要性が軽減されます。Elasticsearch MCP サーバーとエコシステムの現在の状態の詳細については、<a href="https://www.elastic.co/search-labs/blog/mcp-current-state">こちらを</a>ご覧ください。</p><p></p><p>これらの変更はすでに進行中ですので、お楽しみに!</p><h3><strong>まとめ</strong></h3><p></p><p>このブログでは、JavaScript、Mastra、Elasticsearch を使用して、ファンタジー バスケットボール チームに合わせた推奨事項を提供するエージェント RAG アシスタントを構築しました。取り上げた内容:</p><ul><li><p><strong>エージェント RAG の基礎</strong>と、AI エージェントの自律性と RAG を効果的に使用するツールを組み合わせることで、より繊細で動的なエージェントを実現できる方法について説明します。</p></li><li><p><strong>Elasticsearch</strong>とそのデータ ストレージ機能および強力なネイティブ集約により、それが LLM のナレッジ ベースとして優れたパートナーとなる理由について説明します。</p></li><li><p><strong>Mastra</strong>フレームワークと、それが JavaScript エコシステムの開発者にとってこれらのエージェントの構築をどのように簡素化するかについて説明します。</p></li></ul><p>あなたがバスケットボールの熱狂的なファンであっても、AI エージェントの構築方法を検討している方であっても、あるいは私のようにその両方であっても、このブログが、始めるための基礎を提供できれば幸いです。完全なリポジトリは<a href="https://github.com/jdarmada/nba-ai-assistant-js">GitHub</a>で入手できます。自由にクローンして改良してください。さあ、ファンタジーリーグで優勝しましょう!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agentic-rag</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agentic-rag</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[JavaScript]]></category>
    <dc:creator><![CDATA[JD Armada]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8ffd561836a4cb20/6a17f1e47b54f978588b39e4/8132ed781c1ea5d46ca244182f421ed5c721f23b-1200x628.png" length="0" type="image/png"/>
    <pubDate>Tue, 01 Jul 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Azure LLM Functions を Elasticsearch と併用してよりスマートなクエリ エクスペリエンスを実現する]]></title>
    <description><![CDATA[Azure Gen AI LLM 関数と Elasticsearch を使用して柔軟なハイブリッド検索結果を提供する不動産検索アプリの例を調べます。GitHub Codespaces でサンプル アプリを構成して実行する方法を段階的に説明します。]]></description>
    <content:encoded><![CDATA[<p>精度。重要なときは、それは非常に重要なのです。何か特定のものを検索する場合、精度は非常に重要です。ただし、クエリが正確すぎると結果が返されない場合もあるため、クエリの範囲を広げて関連する可能性のある追加のデータを見つける柔軟性があると便利です。</p><p>このブログ記事では、Elasticsearch と Azure Open AI を使用して、非常に具体的な不動産物件を検索するときに正確な結果を見つける方法を示すサンプル アプリを作成する方法について説明します。同時に、特定の一致が見つからない場合でも関連性の高い結果を提供します。検索テンプレートとともに Elasticsearch インデックスを作成するために必要なすべての手順を説明します。次に、Azure OpenAI を使用してユーザー クエリを取得し、それを Elasticsearch 検索テンプレート クエリに変換して、驚くほどカスタマイズされた結果を生成できるアプリを作成するプロセス全体を説明します。</p><p>以下は、サンプルの不動産検索アプリを作成するために使用するすべてのリソースのリストです。</p><ul><li><p>Elasticsearch インデックスと検索テンプレート</p></li><li><p>Azure OpenAI</p></li><li><p>Azure マップ API</p></li><li><p><a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb">コードスペース Jupyter ノートブック</a></p></li><li><p>セマンティックカーネル</p></li><li><p>Blazor フロントエンドを使用した C# アプリ</p></li></ul><h2>スマートクエリワークフロー</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0461b58012efd772/6a17f73fa292997d52d02e19/0c4a7c835e06c514f158c00ab1055a7ba719a35f-1600x765.png" alt="スマートクエリワークフロー" /><p>このワークフローは、LLM、LLM ツール、および検索を組み合わせて、自然言語クエリを構造化された関連性の高い検索結果に変換します。</p><ul><li><p><strong>LLM (大規模言語モデル)</strong> - 複雑なユーザークエリを解釈し、ツールを調整して検索意図を抽出し、コンテキストを強化します。</p></li><li><p><strong>LLM ツール</strong>- 各 LLM ツールは、この投稿用に作成した C# プログラムです。ツールは3つあります。</p><ul><li><p><em>パラメータ抽出ツール</em>: クエリから寝室、バスルーム、機能、価格などの主要な属性を取得します。</p></li><li><p><em>ジオコード ツール</em>: 空間フィルタリングのために場所の名前を緯度/経度に変換します。</p></li><li><p><em>検索ツール</em>: Elasticsearch 検索テンプレートにクエリパラメータを入力し、検索を実行します。<strong>ハイブリッド検索</strong>- 組み込みの ML 推論を使用してハイブリッド検索 (フルテキスト + 高密度ベクトル) を実行します。この階層化されたアプローチにより、エンドユーザーにとってよりスマートでコンテキストを意識したクエリ エクスペリエンスが保証されます。</p></li></ul></li></ul><h2>アプリケーションアーキテクチャ</h2><p>以下はサンプル アプリのシステム アーキテクチャ図です。Elastic Cloud とやり取りするために、Codespaces Jupyter Notebook を使用します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7617ae80295a3e2b/6a17f74196142a35deeb1cb0/2880afee184cd9270c0eb4310e51418e2339784d-936x452.png" alt="Azure LLM Functions アプリのシステム アーキテクチャ図。" /><h2>要件</h2><p>GitHub Codespaces を使用してサンプル アプリのクローンを作成し、構成して実行するので、必要なのはブラウザーだけです。ソリューションの Elastic 部分では、 Elastic Cloudを使用して Elasticsearch Serverless プロジェクトを作成します。Azure リソースを操作するには、 <a href="https://portal.azure.com/">Azure Portal</a>を使用します。</p><h2>Codespaces でサンプル アプリのリポジトリをクローンする</h2><p>まず、サンプル アプリケーションのコードを複製します。これは、アプリケーションのクローンを作成して実行する方法を提供する<a href="https://github.com/codespaces/">GitHub Codespaces</a>で実行できます。<strong>「新しいコードスペース」をクリックします。</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd4ed9c67e2d79c41/6a17f7436df73146a20a10bc/b89cbec491659b6c8a0bb9551ed2629f7a37f9fd-1600x427.png" alt="Codespaces でサンプル アプリ リポジトリのクローンを作成します。" /><p>次に、<strong> リポジトリ</strong> ドロップダウンでリポジトリ<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo"> jwilliams-elastic/msbuild-intelligent-query-demo</a><strong> を選択し、 コードスペースの作成を</strong> クリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdfcc992a6cb7a7d7/6a17f7450b0bed67c1dd3750/43ea377554527af9578400f16cd2342bf8fff3a2-1600x1049.png" alt="リポジトリドロップダウンをクリックし、コードスペースの作成をクリックします。" /><h2>.env を作成するファイル</h2><p>Elastic Cloud にアクセスして操作するには、Python Jupyter Notebook を使用します。これは、構成ファイルに保存されている構成値を使用して行われます。ノートブックの設定ファイルの名前は<em><strong>.env</strong></em>です。今すぐ作成してみましょう。</p><ol><li><p>GitHub Codespacesで、 <strong>「新規ファイル」</strong>ボタンをクリックし、 <em><strong>.env</strong></em>という名前のファイルを追加します。</p></li><li><p>新しく作成した<em><strong>.env</strong></em>に次の内容を追加します。ファイル</p></li></ol>ELASTIC_URL=
ELASTIC_API_KEY=<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1a7dcb2fd2bb24f5/6a17f7462f4a5c21abfa8aa9/84d4f327948858ba61db0001dd8cf780d42fe0a7-1600x875.gif" alt="Python Jupyter Notebook を使用して Elastic Cloud にアクセスし、操作します。これは、構成ファイルに保存されている構成値を使用して行われます。 " /><p>ご覧のとおり、 <strong>ELASTIC_URL</strong>と<strong>ELASTIC_API_KEYという2つの値が欠落しており、</strong> <em>.env</em>に追加する必要があります。ファイル。サンプル アプリの検索機能を強化するバックエンドとして機能する Elasticsearch サーバーレス プロジェクトを作成して、これらを今すぐ取得しましょう。</p><h2>Elastic Serverlessプロジェクトを作成する</h2><ol><li><p><a href="http://cloud.elastic.co">cloud.elastic.co</a>にアクセスし、 <strong>「Create New Serverless project」</strong>をクリックします。</p></li><li><p><strong>Elasticsearch</strong> ソリューションの<strong> 「次へ」</strong> をクリックします</p></li><li><p><strong>ベクター用に最適化を</strong>選択</p></li><li><p><strong>クラウドプロバイダーを</strong><strong>Azure</strong>に設定する</p></li><li><p><strong>サーバーレスプロジェクトの作成を</strong>クリック</p></li><li><p>メインナビゲーションメニューの<strong>「はじめ</strong>に」をクリックし、下にスクロールして<strong>接続の詳細</strong>をコピーします。</p></li><li><p><strong>コピー</strong> ボタンをクリックして、<strong> 接続の詳細</strong> から<strong> Elasticsearchエンドポイント をコピーします。</strong></p></li><li><p><em><strong>.env</strong></em>を更新する<strong>ELASTIC_URL</strong>をコピーした<strong>Elasticsearch エンドポイント</strong>に設定するファイル</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt95b7fefa05173822/6a17f748ec0f89f05f5a67c6/77a35e55446d396066b68cfd132d1543a07b81cc-1600x875.gif" alt="Elasticsearch で新しい Serverless プロジェクトを作成する方法。" /><h2>Elastic APIキーを作成する</h2><ol><li><p>Elasticsearchの<strong> Getting Started</strong> ページを開き、<strong> APIキーの追加</strong> セクションで<strong> 新規を クリックします。</strong></p></li><li><p>キー<strong>名</strong>を入力してください</p></li><li><p><strong>APIキーの作成を</strong>クリック</p></li><li><p>「コピー」ボタンをクリックしてAPIキーの値をコピーします</p></li><li><p><strong>Codespacesに戻ると、</strong> <em><strong>.env</strong></em>がある。ファイルを開いて編集し、コピーした値を貼り付けて<strong>ELASTIC_API_KEY</strong>を設定します。</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf6b1c85d267e6ad0/6a17f74a4b055d118143239a/20168cba493d8e2c0d9ae7704eb0ae707df58e4c-1600x875.gif" alt="Elastic API キーを作成する方法。" /><h2>Codespacesノートブックを開き、ライブラリの依存関係をインストールする</h2><p>ファイル エクスプローラーで<em><strong>VectorDBSetup.ipynb</strong></em>ファイルを選択し、ノートブックを開きます。ノートブックが読み込まれたら、<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L40-L52"><strong> 「ライブラリのインストール」</strong></a><a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L40-L52"> というタイトルのノートブックのセクション を見つけます</a><strong> 。</strong>セクションの再生ボタンをクリックします。</p><p>GitHub Codespaces でノートブックを初めて実行する場合は、Codespaces カーネルを選択して Python 環境を構成するように求められます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt970f4fa30c6c9302/6a17f74c505ac30ee6ad8ceb/2272f70615dfb9dcbeb91f39b6dd5076213e24a5-1600x875.gif" alt="Codespaces Notebook にライブラリ依存関係をインストールします。" /><h2>Codespaces Notebook を使用してインポートを定義し、環境変数をロードする</h2><p>ノートブックの次のセクション<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L58-L104">「</a><a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L58-L104"><strong>インポートの定義と環境変数の読み込み</strong></a>」に進みます。セクションの再生ボタンをクリックします。</p><p>このコードは、ノートブックで使用されるPythonライブラリをインポートし、 <em>.env</em>から環境変数を読み込みます。以前作成したもの。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6d74b41258420d5a/6a17f74e6317301f2f585c16/aa9f9198ff452ac0c4ce33b00f253731dbee22c5-1600x875.gif" alt="Codespaces Notebook を使用してインポートとロード環境変数を定義します。" /><h2>Codespaces Notebook を使用して Elastic ML 推論エンドポイントを作成する</h2><p>次の<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L112-L157"><strong> 「ML 推論エンドポイントの作成」</strong></a><a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L112-L157"> というノートブックのセクション</a> に進みます。セクションの再生ボタンをクリックします。</p><p>これにより、Elasticsearch プロジェクトに新しい ML 推論エンドポイントが作成され、データからテキスト埋め込みを生成するために使用します。テキスト埋め込みは、セマンティック検索を強化するために Elasticsearch に保存されるテキストのベクトル表現です。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt87581300d4d0b66e/6a17f750e9ea875a81a9c795/97c1afab3e64027ee5ae77f377d56ba406ae1765-1600x875.gif" alt="Codespaces Notebook を使用して Elastic ML 推論エンドポイントを作成する。" /><h2>Codespaces Notebook を使用して Elasticsearch インデックスを作成する</h2><p>次の<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L165-L224"><strong> 「Elasticsearch インデックスの作成」</strong></a><a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L165-L224"> というノートブックのセクション</a> に進みます。セクションの再生ボタンをクリックします。</p><p>これにより、サンプル データと、ML 推論エンドポイントを介して生成された関連するベクトル データを保存する Elasticsearch インデックスが作成されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta2d2d5a10c84a1b7/6a17f7527f6f15775cc09cd6/23a66283ee41239e24fb8455c3cd95641982ca6b-1600x875.gif" alt="Codespaces Notebook を使用して Elasticsearch インデックスを作成します。" /><h2>Codespaces Notebook を使用して Elasticsearch 検索テンプレートを作成する</h2><p>次の<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L232-L384"> </a><a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L232-L384"><strong>「検索テンプレート」 というノートブックのセクション</strong></a> に進みます。セクションの再生ボタンをクリックします。</p><p>これにより、<a href="https://www.elastic.co/jp/docs/solutions/search/search-templates">検索テンプレート</a>が作成されます。これは、ユーザーの検索クエリから解析された単語が入力さたテンプレートとしてサンプル アプリで使用されるものです。これにより、Elasticsearch インデックス内のデータをクエリする際の詳細度を設定および制御できるようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt194d6557d096ac25/6a17f7545772628d901bcda0/4c001a3e4d1cca4cfb5c043fea92c7ccaf9cb64a-1600x875.gif" alt="Codespaces Notebook を使用して Elasticsearch 検索テンプレートを作成します。" /><h2>Codespaces Notebook を使用して Elasticsearch インデックスにデータを取り込む</h2><p>ノートブックの次のセクション<a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L392-L450">「</a> <a href="https://github.com/jwilliams-elastic/msbuild-intelligent-query-demo/blob/main/VectorDBSetup.ipynb?short_path=17c25d8#L392-L450"><strong>Ingest property data」</strong></a>に進みます。セクション実行ボタンをクリックします。</p><p>このコード セクションを実行すると、 <em>properties.jsonl</em>ファイルに含まれるサンプル データが一括ロードされます。数分後、プロセスが正常に完了したことを示す確認が表示されます。Elastic Cloud の<strong>インデックス管理</strong>セクションに移動すると、インデックスに予想されるレコードが含まれていることを確認できます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltaf1b5ad59c75716d/6a17f7566864a4528fb6894b/e9698c798541ccfc08143a939846597028e3c566-1600x875.gif" alt="Codespaces Notebook を使用して Elasticsearch インデックスにデータを取り込みます。" /><h2>C# アプリを構成するための appsetting.json を作成する</h2><p>Elasticsearch インデックスが作成され、データが取り込まれたので、Elastic および Azure Cloud で動作するようにサンプル アプリを構成する準備が整いました。C# サンプル アプリは、 <em>appsettings.json</em>という名前のファイルを使用して、API キーなどのアクセス情報を保存および読み込みます。ここで、Codespaces のエディターを使用して<em>appsettings.json</em>ファイルを作成します。</p><p>1.<strong> HomeFinderApp</strong> フォルダに<em> appsettings.json を作成します。</em></p><p>2. 次のコードを<em>appsettings.json</em>ファイルに貼り付けます。</p>{
 "ElasticSettings": {
   "Url": "",
   "ApiKey": "",
   "IndexName": "properties",
   "TemplateId": "properties-search-template"
 },
 "AzureOpenAISettings": {
   "Endpoint": "",
   "ApiKey": "",
   "DeploymentName": "gpt-4o"
 },
 "AzureMapsSettings": {
   "Url": "https://atlas.microsoft.com/geocode",
   "ApiKey": ""
 },
 "Logging": {
   "LogLevel": {
 	"Default": "Information",
 	"Microsoft.AspNetCore": "Warning"
   }
 },
 "AllowedHosts": "*"
}
<p>3.<strong> ElasticSettings セクションの</strong><strong> Url</strong> と<strong> ApiKey</strong> の 値を見つけます。<em>.env</em>で設定した値と同じ値に設定します前の手順でファイルを作成します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt28afc316da403ec9/6a17f758faa913584d93ca28/00dad25bacdea2adcbd1e6eca7658867a49b0d8c-1600x875.gif" alt="C# アプリを構成するための appsetting.json を作成しています。" /><h2>Azure OpenAI サービスを作成する</h2><p>サンプル アプリでは、Azure OpenAI を使用してアプリ ユーザーのクエリを解析し、検索テンプレートを入力して Elasticsearch にリクエストを送信し、ユーザーが検索している内容を柔軟に伝えようとします。</p><ol><li><p>新しいブラウザー タブを開き、Azure ポータルの<a href="https://portal.azure.com/#blade/Microsoft_Azure_ProjectOxford/CognitiveServicesHub/OpenAI">AI Foundry | Azure OpenAI</a>に移動します。+<strong>作成を</strong>クリック</p></li><li><p>作成フォームで、<strong>リソース グループ</strong>を選択します。</p></li><li><p><strong>名前</strong>を入力してください</p></li><li><p><strong>価格帯</strong>を選択する</p></li><li><p><strong>次へを</strong>クリック</p></li><li><p><strong>ネットワーク</strong>タブで<strong>次へを</strong>クリックします</p></li><li><p><strong>「タグ」</strong>タブで<strong>「次へ」</strong>をクリックします。</p></li><li><p><strong>「確認と送信」</strong>タブで、 <strong>「作成」を</strong>クリックします。</p></li><li><p>作成が完了したら、 <strong>「リソースに移動」を</strong>クリックします。</p></li><li><p>左側のナビゲーションメニューから<strong>キーとエンドポイント</strong>を選択します</p></li><li><p><strong>エンドポイント</strong>をコピーし、Codespaces エディターが開いているブラウザー タブで作成した<em>appsettings.json</em>ファイルに貼り付けます。</p></li><li><p>次に、Azure OpenAI<strong>キーとエンドポイント ページ</strong>を含むブラウザー タブに戻ります。<strong>Key 1</strong>のコピー ボタンをクリックし、コピーした値を、Codespaces エディターが開いているブラウザー タブの<em>appsettings.json</em>ファイルに貼り付けます。</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd01a5e82d9003d02/6a17f75a148009be65b48904/6d49197302d110410dca0a53b6ae90237cf2dfd6-1600x875.gif" alt="Azure OpenAI サービスを作成します。" /><h2>Azure Open AI サービスに gpt-4o モデルのデプロイメントを追加する</h2><p>素晴らしい。Azure OpenAI サービスが実行できるようになりました。ただし、サンプル アプリに必要な LLM 機能を提供するには、まだモデルのデプロイが必要です。選択できるモデルは多数あります。作成した <em>appsettings.json</em><em> ファイルですでに指定されているので 、gpt-4o</em> をデプロイしましょう。</p><p></p><ol><li><p><a href="https://ai.azure.com/resource/playground">Azure AI Foundry</a>にアクセスし、 <strong>「デプロイの作成」を</strong>クリックします。</p></li><li><p><em>gpt-4o</em>を検索し、結果から選択します</p></li><li><p><strong>確認</strong>をクリックして選択します</p></li><li><p><strong>「デプロイ」</strong>をクリックしてモデルをデプロイします</p></li></ol><p><em>gpt-4o</em> モデルを正常にデプロイしたら、左側のナビゲーションメニューから[デプロイメント] を選択し、<em><strong> gpt-4o</strong></em> デプロイメントが[成功] の<strong>状態</strong> でリストされていることを確認できます。
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte82341dd8a4982b0/6a17f75c4b055d9e0943239e/1b817ab67c05634e9c72777593b4d1a2c6c28191-1600x875.gif" alt="Azure Open AI サービスに gpt-4o モデルのデプロイメントを追加します。" /><h2>Azure Maps アカウントを作成する</h2><p>私たちは、サンプル アプリのユーザーが特定のエリアの不動産物件を検索できるようにしたいと考えていますが、あまり具体的に検索する必要はありません。地元のファーマーズ マーケットの近くの物件を検索したい場合、Azure Maps は OpenAI LLM がマーケットの緯度と経度の座標を取得するために使用できるサービスです。その後、座標は、特定の場所と地理的距離を考慮したユーザークエリのために Elasticsearch に送信される検索テンプレートベースのリクエストに含めることができます。</p><ol><li><p><a href="https://portal.azure.com/#browse/Microsoft.Maps%2Faccounts">Azure Mapsアカウント で 作成を クリックします</a></p></li><li><p><strong>リソースグループ</strong>を選択</p></li><li><p><strong>名前</strong>を入力してください</p></li><li><p>ライセンスとプライバシーに関する声明に同意する</p></li><li><p><strong>「確認して作成」</strong>をクリック</p></li><li><p><strong>作成を</strong>クリック</p></li><li><p>アカウントの作成が完了したら、 <strong>「リソースに移動」を</strong>クリックします。</p></li><li><p>左側のナビゲーションメニューで<strong>「認証」を</strong>クリックします。</p></li><li><p><strong>主キーの</strong> 値をコピーし、Codespaces エディターを含むブラウザー タブに戻って、<em> appsettings.json</em> ファイルの<strong> AzureMapsSettings</strong> セクションの<strong> ApiKey</strong> の値として貼り付けます。</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt42a4ab4bb7a96d24/6a17f75edbb4ff4f91fb5892/90fadd48e366682e2bad91e32988f93c6354e126-1600x875.gif" alt="Azure Maps アカウントを作成します。" /><h2>サンプルアプリを試してみる</h2><p>さて、ここからが楽しい部分です。サンプルアプリを実行してみましょう。アプリを動かすために必要な Elastic Cloud および Azure Cloud リソースとともに、すべての構成の詳細が設定されました。</p><p>1. Codespaces エディターでターミナル ウィンドウを開きます。</p><p>2. 次のコマンドを使用して、アクティブ ディレクトリをサンプル アプリ フォルダーに変更します。
</p>cd HomeFinderApp<p>3. 次の<em>dotnet</em>コマンドを使用してアプリを実行します。</p>dotnet run<p>4. <strong>「ブラウザで開く</strong>」ボタンが表示されたらクリックします。</p><p>5. デフォルトの検索をテストしてから、独自のカスタム検索をいくつか試します。検索結果を生成するためにバックエンドで実行される内容の詳細を確認するには、「ツールの呼び出し」の横にある<strong>「表示」</strong>リンクをクリックします<strong>。</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt41adf6631ba91be1/6a17f760505ac30986ad8cf1/821fe7b9446de5ed646d938cc9484a7ddad21030-1600x875.gif" alt="サンプルアプリを試しています。" /><p><strong>ボーナス:</strong> GPT-4o を実際にテストしたい場合は、次の検索を試してください:<em>フロリダ州ディズニー ワールドの近くで、ベッドルーム 30 室以上、バスルーム 20 室以上、プールとガレージがあり、ビーチに近い物件を 20 万ドル未満で探しています。</em>このクエリは、複数の検索ツールの呼び出し後に結果を返します。</p><h2>Elasticは検索AIのソリューションです</h2><p>実行中のアプリは、検索テンプレートを介して Elasticsearch を基礎データ ソースとして使用する Gen AI LLM ガイド付き検索の例です。自由にサンプル アプリを試してカスタマイズし、正確かつ柔軟な検索エクスペリエンスを作成して、ユーザーが探しているものを見つけられるようにしてください。</p><p>読んでいただきありがとうございます。<a href="https://cloud.elastic.co/registration">Elastic Cloud</a>をお試しください。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/azure-llm-functions-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/azure-llm-functions-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jonathan Simon,James Williams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt93dd59caccfd7fc8/6a17f7614202292a7129f799/1431b90c7e00de06574c1e33c44a2e89296c824e-1200x628.png" length="0" type="image/png"/>
    <pubDate>Fri, 13 Jun 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[MCP（モデルコンテキストプロトコル）の現状]]></title>
    <description><![CDATA[MCP、プロジェクトの更新、機能、セキュリティ上の課題、新しいユースケース、Elastic の Elasticsearch MCP サーバーの操作方法について学びます。]]></description>
    <content:encoded><![CDATA[<p>最近、サンフランシスコで開催され<a href="https://mcpdevsummit.ai/">た MCP 開発者サミット</a>に出席しましたが、モデル コンテキスト プロトコル (MCP) が急速に AI エージェントとコンテキストリッチな AI アプリケーションの基礎となる構成要素になりつつあることは明らかでした。Elastic では、 <a href="https://www.elastic.co/jp/elasticsearch/agent-builder">Agent Builder</a>から MCP サーバーを直接公開することでこの方向に傾き、Elasticsearch をあらゆる MCP 互換エージェントにとって第一級のコンテキストおよびツール プロバイダーにしています。この記事では、イベントからの主な最新情報、新しいユースケース、MCP の今後の展望、Agent Builder を使用して MCP 経由でエージェントが Elasticsearch を利用できるようにする方法について説明します。</p><h2>モデルコンテキストプロトコル (MCP) とは何ですか?</h2><p>ご存じない方のために説明すると、<a href="https://modelcontextprotocol.io/introduction">モデル コンテキスト プロトコルは</a>、AI モデルをさまざまなデータ ソースやツールに接続するための構造化された双方向の方法を提供し、より関連性の高い情報に基づいた応答を生成できるようにするオープン スタンダードです。一般的に「 <a href="https://modelcontextprotocol.io/introduction">AI アプリケーション用の USB-C ポート</a>」と呼ばれています。</p><p>双方向の性質を強調したアーキテクチャ図を以下に示します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ff0e141b5dfda29/6a17e7ffe8fbcee5263a1946/5eba1e59514eb58a5220bb92bb49e6328ee83cd7-674x466.png" alt="モデルコンテキストプロトコル（MCP）アーキテクチャ" /><p>これは AI 実践者にとって大きな変化です。AI アプリケーションを拡張する際の主な課題の 1 つは、新しいデータ ソースごとにカスタム統合を構築する必要があることです。MCP は、モデルにコンテキストを管理および提供するための持続可能で再利用可能なアーキテクチャを提供します。モデルやサーバーに依存せず、完全にオープンソースです。</p><p>MCP は、アプリケーション間の統合を標準化することを目指す一連の API 仕様の最新版です。これまで、RESTful サービスには OpenAPI、データ クエリには GraphQL、マイクロサービス通信には gRPC を使用していました。MCP は、これらの古い仕様の構造化された厳密さを共有するだけでなく、それを生成 AI 設定に取り入れることで、カスタム コネクタなしでエージェントをさまざまなシステムに簡単に接続できるようになります。多くの点で、MCP は HTTP が Web に対して行ったことと同じことを AI エージェントに対して行うことを目指しています。HTTP がブラウザと Web サイト間の通信を標準化したのと同様に、MCP は AI エージェントが周囲のデータの世界と対話する方法を標準化することを目指しています。</p><h2>MCPと他のエージェントプロトコルの比較</h2><p>エージェント プロトコルの状況は急速に拡大しており、エージェントの相互作用方法を定義するために 12 を超える新しい標準が競合しています。LlamaIndex の<a href="https://x.com/seldo">Laurie Voss</a>氏は、ほとんどのプロトコルを 2 つのタイプに分類できると説明しています。エージェント同士の対話に重点を置くエージェント間プロトコルと、構造化されたコンテキストを LLM に提供することに重点を置く MCP などのコンテキスト指向プロトコルです。</p><p>Google の<a href="https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/">A2A</a> (Agent to Agent)、Cisco と IBM の<a href="https://agentcommunicationprotocol.dev/introduction/welcome">ACP</a> (Agent Communication Protocol)、 <a href="https://agoraprotocol.org/">Agora</a>などの他の一般的なプロトコルは、エージェント間のネゴシエーション、連合の構築、さらには分散型 ID システムを可能にすることを目的としています。MCP は、エージェントがツールやデータにアクセスする方法に焦点を当てており、必ずしもエージェント同士が通信する方法に焦点を当てているわけではないため、もう少し実用的なアプローチを採用しています (ただし、MCP は将来的にさまざまな方法でそれを可能にすることもできます)。</p><p>現在、MCP が他と一線を画しているのは、その牽引力と勢いです。フロントエンド フレームワークの初期の React と同様に、MCP はニッチな問題から始まり、現在では実際に最も採用され、拡張可能なエージェント プロトコルの 1 つとなっています。</p><h2>サミットのまとめ: MCP の優先事項の進化</h2><p>サミットには、Anthropic、Okta、OpenAI、AWS、GitHub などの貢献者による講演者が登壇しました。講演では、コアプロトコルの強化から実際の実装まで幅広い話題が取り上げられ、短期的および長期的な優先事項が概説されました。これらの講演は、初期の実験や単純なツール呼び出しから、MCP を基盤として使用した信頼性が高く、スケーラブルでモジュール化された AI システムの構築への移行を反映していました。</p><p>何人かの講演者は、MCP が単なるプロトコル配管にとどまらず、AI ネイティブ Web の基盤となる将来を示唆しました。JavaScript によってユーザーが Web ページをクリックして操作できるのと同じように、MCP によってエージェントが私たちに代わって同じアクションを実行できるようになります。たとえば、電子商取引では、ユーザーが買い物をするために手動で Web サイトに移動するのではなく、エージェントにログインして特定の製品を見つけ、カートに追加してチェックアウトするように指示するだけで済みます。</p><p>これは単なる憶測や誇大宣伝ではありません。PayPal はサミットで、まさにこのエージェントによるコマース体験を可能にする新しいエージェント ツールキットと MCP サーバーを披露しました。MCP はツールやデータ ソースへの安全で信頼性の高いアクセスを提供するため、エージェントは Web を読み取るだけでなく、それに基づいて行動できるようになります。現在、MCP はすでに大きな勢いを持つ強力な標準であり、将来的には Web 全体で AI を活用したユーザー インタラクションの標準になる可能性があります。</p><h2>MCPプロジェクトの最新情報: トランスポート、抽出、構造化ツール</h2><p>MCP のコア貢献者である<a href="https://x.com/JeromeSwannack">Jerome Swannack 氏</a>が、過去 6 か月間のプロトコル仕様の更新をいくつか共有しました。これらの変更の主な目的は次のとおりです。</p><ol><li><p>ストリーミング可能なHTTPを追加してリモートMCPを有効にする</p></li><li><p>抽出とツール出力スキーマの追加により、より豊富なエージェントインタラクションモデルを可能にする</p></li></ol><p>MCP はオープンソースであるため、開発者は Streamable HTTP などの変更を実装できる状態になっています。抽出およびツール出力スキーマは現在リリースされておらず、ドラフト段階にあり、進化する可能性があります。</p><p><strong>ストリーミング可能な HTTP</strong> ( <a href="https://modelcontextprotocol.io/specification/2025-03-26/basic/transports">2025 年 3 月 26 日リリース</a>) <strong>:</strong>ストリーミング可能な HTTP が新しいトランスポート メカニズムとして導入されたことは、大きなインパクトのある技術更新でした。これにより、サーバー送信イベント (SSE) が、チャンク転送エンコーディングと単一の HTTP 接続を介したプログレッシブ メッセージ配信をサポートする、よりスケーラブルな双方向モデルに置き換えられます。これにより、AWS Lambda などのクラウド インフラストラクチャに MCP サーバーを展開し、長時間の接続やポーリングを必要とせずにエンタープライズ ネットワークの制約をサポートできるようになります。</p><p><strong>Elicitation</strong> ( <a href="https://modelcontextprotocol.io/specification/2025-06-18/client/elicitation">2025 年 6 月 18 日リリース</a>) <strong>:</strong> Elicitation を使用すると、サーバーはクライアントからのコンテキストの構造化方法を指定するスキーマを定義できます。基本的に、サーバーは必要なものと期待する入力の種類を記述できます。これにはいくつかの意味があります。サーバービルダーにとっては、より複雑なエージェントのインタラクションを構築できます。クライアントビルダーは、これらのスキーマに適応する動的な UI を実装できます。ただし、ユーザーから機密情報や個人を特定できる情報を抽出するために、誘導法を使用するべきではありません。特に MCP が成熟するにつれて、開発者は<a href="https://modelcontextprotocol.io/specification/draft/client/elicitation#security-considerations">ベスト プラクティス</a>に従って、誘導プロンプトが安全かつ適切な状態を保つようにする必要があります。これは、この投稿の後半で説明する、より広範なセキュリティ上の懸念に関係しています。</p><p><strong>ツール出力スキーマ</strong>( <a href="https://modelcontextprotocol.io/specification/draft/server/tools#output-schema">2025 年 6 月 18 日リリース</a>) <strong>:</strong>このコンセプトにより、クライアントと LLM はツール出力の形状を事前に知ることができます。ツール出力スキーマを使用すると、開発者はツールが返すことが予想される内容を記述できます。これらのスキーマは、コンテキスト ウィンドウの非効率的な使用という、直接ツール呼び出しの主な制限の 1 つに対処します。コンテキスト ウィンドウは、LLM を操作するときに最も重要なリソースの 1 つと考えられており、ツールを直接呼び出すと、LLM のコンテキストに完全にプッシュされる生のコンテンツが返されます。ツール出力スキーマを使用すると、MCP サーバーが構造化データを提供できるようになるため、トークンとコンテキスト ウィンドウをより有効に活用できるようになります。ここでは、ツール全般に関する<a href="https://modelcontextprotocol.io/specification/draft/server/tools#security-considerations">ベストプラクティスを</a>いくつか紹介します。</p><p>これらの新しいアップデートと今後の追加により、MCP はよりモジュール化され、型付けされた、実稼働対応のエージェント プロトコルになります。</p><h2>あまり使われていない強力な機能：サンプリングとルート</h2><p>MCP 仕様では目新しいものではありませんが、基調講演ではサンプリングとルートの両方が強調されました。これら 2 つのプリミティブは現在見過ごされ、十分に調査されていませんが、エージェント間のより豊かで安全なインタラクションに大きく貢献する可能性があります。</p><p><strong>サンプリング - サーバーはクライアントからの補完を要求できます。</strong><a href="https://modelcontextprotocol.io/docs/concepts/sampling">サンプリング</a>により、MCP サーバーはクライアント側の LLM から補完を要求できます。これにより、プロトコルの双方向性が強化され、サーバーはリクエストに応答するだけでなく、クライアントのモデルにプロンプトを出して応答を生成するように要求できるようになります。これにより、クライアントはコスト、セキュリティ、MCP サーバーが使用するモデルを完全に制御できます。したがって、事前構成されたモデルを備えた外部 MCP サーバーを使用する場合、サーバーはクライアントにすでに接続されているモデルを要求するだけでよいため、独自の API キーを提供したり、そのモデルに対する独自のサブスクリプションを構成したりする必要はありません。これにより、より複雑でインタラクティブなエージェントの動作が可能になります。</p><p><strong>ルート - リソースへのスコープ アクセス:</strong><a href="https://modelcontextprotocol.io/docs/concepts/roots">ルートは</a>、クライアントが関連するリソースと焦点を当てるワークスペースについてサーバーに通知する方法を提供するために設計されました。これは、サーバーが動作する範囲を設定するのに強力です。ルートは「<a href="https://modelcontextprotocol.io/docs/concepts/roots#how-roots-work">情報提供のみを目的としており、厳密に強制するものではない</a>」ことに注意することが重要です。つまり、MCP サーバーまたはエージェントの権限やアクセス許可を定義しないということです。つまり、サーバーまたはエージェントが特定のツールを実行したり書き込みアクションを実行したりするのを防ぐために、ルートだけに頼ることはできません。ルートの場合も、ユーザー承認のメカニズムを使用して、権限はクライアント側で処理する必要があります。また、開発者は、ルートによって設定された境界を尊重し、<a href="https://modelcontextprotocol.io/docs/concepts/roots#best-practices">ベストプラクティス</a>を使用するように設計されたサーバーの使用にも注意する必要があります。</p><h2>エージェントの認証: OAuth 2.1 と保護されたメタデータ</h2><p>このセクションでは、安全でないフローを排除し、ベスト プラクティスを統合した OAuth 2.0 の最新バージョンである OAuth 2.1 に焦点を当てます。</p><p>OAuth サポートは、特にセキュリティとスケーラビリティが、MCP がエージェントをツールに接続するための標準となることを妨げる大きな障害であると考えられているため、非常に期待されていたトピックでした。<a href="https://x.com/aaronpk">Aaron Parecki 氏</a>(Okta の OAuth 2.1 編集者兼 ID 標準専門家) は、サーバー開発者の複雑さのほとんどを軽減する、クリーンかつスケーラブルな OAuth フローを MCP がどのように採用できるかについて説明しました。公式の OAuth 2.1 認証仕様は、 <a href="https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization">2025 年 6 月 18 日</a>の最新プロトコル改訂版で最近公開されました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d53e7bb091b4f43/6a17e80163baff80dc741c56/2ea159116fe5e03ff800f077adf16d6ca9f1c1d1-1594x1280.png" alt="エージェント向けMCP認証" /><p>この実装では、OAuth の責任を MCP クライアントとサーバーの間で分割できます。認証フローの大部分は MCP クライアントによって開始および処理され、最後にサーバーが関与するのは安全なトークンの受信と検証のみです。この分割により、開発者がすべての接続を構成する必要なく、多くのツール間で認証を行うという重要なスケーリングの問題が解決され、MCP サーバー開発者が OAuth の専門家になる必要がなくなります。</p><p>講演の2つの重要なハイライト:</p><ol><li><p><a href="https://datatracker.ietf.org/doc/rfc9728/"><strong>保護されたリソース メタデータ</strong></a>: MCP サーバーは、目的、エンドポイント、認証方法を記述した JSON ファイルを公開できます。これにより、クライアントはサーバー URL だけで OAuth フローを開始できるようになり、接続プロセスが簡素化されます。詳細: <a href="https://aaronparecki.com/2025/04/03/15/oauth-for-model-context-protocol">MCP で OAuth を修正しましょう</a></p></li><li><p><a href="https://datatracker.ietf.org/doc/html/draft-ietf-oauth-v2-1-13"><strong>IDP と SSO のサポート</strong></a>: 企業は ID プロバイダーを統合してアクセスを集中管理できます。これは、ユーザー エクスペリエンスとセキュリティの両方にとってメリットとなります。ユーザーは 10 個の異なる同意画面をクリックする必要がなくなり、セキュリティ チームは各接続を監視できるようになります。</p></li></ol><p>OAuth ロジックをクライアントにプッシュし、サーバーからのメタデータに依存することで、MCP エコシステムは大きなボトルネックを回避します。これにより、MCP は、今日の運用環境で最新の API が保護される方法とより密接に連携するようになります。</p><p>追加の参考資料: <a href="https://aaronparecki.com/oauth-2-simplified/">OAuth 2 Simplified</a> 。</p><h2>コンポーザブルエコシステムにおけるセキュリティの課題</h2><p>新たな開発には新たな攻撃対象領域も伴います。Cisco の Arjun Sambamoorthy 氏は、MCP 環境における主な脅威をいくつか挙げています。</p><p>脅威</p><p>説明</p><p>修復とベストプラクティス</p><p>迅速な注射とツールの中毒</p><p>LLM システムのコンテキストまたはツールの説明内に悪意のあるプロンプトを挿入し、LLM がファイルの読み取りやデータの漏洩などの意図しないアクションを実行するようにする方法。</p><p>MCP Scan などのツールを使用して、ツールのメタデータのチェックを実行します。説明とパラメータをプロンプトに含める前に検証します。最後に、リスクの高いツールに対してユーザー承認を実装することを検討してください。詳細については、表の下の追加の読書リストにある OWASP プロンプト インジェクション ガイドを参照してください。</p><p>サンプリング攻撃</p><p>MCP のコンテキストでは、サンプリングにより、MCP サーバーが LLM に対してプロンプト インジェクション攻撃を実行できるようになります。</p><p>信頼できないサーバーのサンプリングを無効にし、サンプリング要求に人間による承認を追加することを検討してください。</p><p>悪意のあるMCPサーバー</p><p>現在の MCP サーバーのコレクションでは、安全性を確保するために各サーバーを検査するのは困難です。不正なサーバーは密かにデータを収集し、悪意のある人物に公開する可能性があります。</p><p>信頼できるレジストリまたは内部リストからのみ MCP サーバーに接続します。サンドボックス化されたコンテナ内でサードパーティのサーバーを実行します。</p><p>悪意のあるMCPインストールツール</p><p>コマンドライン インストーラーとスクリプトは、MCP サーバーまたはツールを迅速に実装するのに便利ですが、検証されていない侵害されたコードをインストールしてしまう可能性があります。</p><p>サンドボックス環境にインストールし、パッケージ署名を検証します。検証されていないソースからの自動更新は行わないでください。</p><p>この問題にさらに対抗するために、Arjun は、信頼できる MCP レジストリを使用してすべての検証を処理すること (これは最重要トピックです。詳細については、以下の読書リストの上位 2 項目を参照してください) と、この<a href="https://github.com/slowmist/MCP-Security-Checklist">セキュリティ チェックリスト</a>の使用を提案しています。</p><p>追加の参考資料:</p><ul><li><p><a href="https://modelcontextprotocol.io/specification/2025-06-18/basic/security_best_practices">公式MCPセキュリティベストプラクティス</a></p></li><li><p><a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/">OWASP LLM アプリケーション トップ 10</a></p></li><li><p><a href="https://hiddenlayer.com/innovation-hub/">HiddenLayer脅威リサーチ</a></p></li><li><p><a href="https://github.com/invariantlabs-ai/mcp-scan">MCPスキャン</a></p></li><li><p><a href="https://genai.owasp.org/llmrisk/llm01-prompt-injection/">OWASP プロンプトインジェクションガイド</a></p></li></ul><h2>次はレジストリ、ガバナンス、エコシステム</h2><p>集中型の MCP レジストリが開発中であり、サミットで最も頻繁に議論されたトピックの 1 つでした。現在のサーバー エコシステムは、断片化、信頼性と発見可能性の低さに悩まされています。特にメタデータが不完全であったり偽装されたりする可能性がある分散型エコシステムでは、開発者が MCP サーバーを見つけ、その動作を検証し、安全にインストールすることは困難です。</p><p>集中型レジストリは、信頼できる真実のソースとして機能し、発見可能性を向上させ、サーバー メタデータの整合性を確保し、悪意のあるツールをインストールするリスクを軽減することで、これらの問題点に直接対処します。</p><p>MCP レジストリの目標は次のとおりです。</p><ul><li><p>サーバーのメタデータ（サーバーが何をするか、どのように認証するか、インストールして呼び出すか）に関する唯一の真実の情報源を提供します。</p></li><li><p>不完全なサードパーティのレジストリと断片化を排除して、サーバーを登録するときに、インターネット上の他のすべてのレジストリを更新する必要がなくなります。</p></li><li><p>CLI ツールと、前述のメタデータを含む server.json ファイルを含むサーバー登録フローを提供します。</p></li></ul><p>より広範な期待は、信頼できるレジストリがエコシステムを安全に拡張し、開発者が自信を持って新しいツールを構築して共有できるようにすることです。</p><p>ガバナンスは、Anthropic にとってもう一つの最重要課題でした。MCP はオープンかつコミュニティ主導であり続けるべきだと明言しましたが、ガバナンス モデルの拡張はまだ進行中です。彼らは現在、その分野での支援を求めており、オープンソース プロトコルのガバナンスの経験がある方は誰でも連絡を取るよう呼びかけています。これは私が言及したかったもう一つの話題につながります。イベント全体を通じて、講演者は、エコシステムはその内部の開発者の貢献によってのみ成長できると強調しました。MCP を新しい Web 標準にして、他の一般的なエージェント プロトコルと差別化するために、集中的な取り組みが必要です。</p><h2>現実世界におけるMCP：ケーススタディとデモ</h2><p>いくつかの組織は、MCP がすでに実際のアプリケーションでどのように使用されているかを共有しました。</p><ul><li><p><strong>PayPal - エージェンティックコマース向け MCP サーバー:</strong> PayPal は、ユーザーのショッピング体験を根本的に変えることができる新しい<a href="https://github.com/paypal/agent-toolkit/">エージェント ツールキット</a>と MCP サーバーを展示しました。ユーザーは、ソーシャル メディアで商品を探したり、価格を比較したり、チェックアウトしたりする代わりに、PayPal MCP サーバーに接続してそれらのすべてのアクションを処理するエージェントとチャットできます。
</p></li><li><p><strong>EpicAI.pro - Jarvis:</strong> MCP の開発により、現実世界の Jarvis タイプのアシスタントの実現にますます近づいています。アイアンマン映画をご存じない方のために説明すると、Jarvis は自然言語を使用し、マルチモーダル入力に応答し、応答時に遅延がなく、ユーザーのニーズを積極的に予測し、統合を自動的に管理し、デバイスと場所の間でコンテキストを切り替えることができる AI アシスタントです。Jarvis を物理的なロボット アシスタントとして想像すると、MCP は Jarvis に「手」、つまり複雑なタスクを処理する能力を与えます。
</p></li><li><p><strong>Postman -</strong> <a href="https://www.postman.com/explore/mcp-generator"><strong>MCP サーバー ジェネレーター</strong></a><strong>:</strong> API リクエスト用のショッピング カート エクスペリエンスを提供します。さまざまな API リクエストを選択してバスケットに入れ、バスケット全体を MCP サーバーとしてダウンロードできます。
</p></li><li><p><strong>Bloomberg -</strong> Bloomberg は、エンタープライズ GenAI 開発における主要なボトルネックを解決しました。約 10,000 人のエンジニアを抱える同社では、チーム間でツールとエージェントを統合するための標準化された方法が必要でした。MCP を使用することで、社内ツールをモジュール式のリモートファースト コンポーネントに変換し、エージェントが統合インターフェースで簡単に呼び出すことができるようになりました。これにより、エンジニアは組織全体にツールを提供できるようになり、AI チームはカスタム統合ではなくエージェントの構築に集中できるようになりました。Bloomberg は現在、MCP エコシステムとの完全な相互運用性を実現する、スケーラブルで安全なエージェント ワークフローをサポートしています。ブルームバーグは公開リソースへのリンクを一切提供していないが、これはサミットで彼らが公開した内容である。
</p></li><li><p><strong>Block –</strong> Block は MCP を使用して、従業員がエンジニアリング、営業、マーケティングなどのタスクを自動化できるようにする社内 AI エージェントである<a href="https://github.com/block/goose?tab=readme-ov-file">Goose</a>を強化しています。同社は、Git、Snowflake、Jira、Google Workspace などのツール用に 60 台以上の MCP サーバーを構築し、日常的に使用するシステムとの自然言語によるやり取りを可能にしました。Block 社の従業員は現在、Goose を使用してデータのクエリ、不正行為の検出、インシデントの管理、内部プロセスのナビゲートなどを行っており、これらはすべてコードを書かずに実行できます。MCP は、Block がわずか 2 か月で多くの職務にわたって AI の導入を拡大できるよう支援しました。
</p></li><li><p><strong>AWS -</strong> <a href="https://github.com/awslabs/mcp"><strong>AWS MCP サーバー</strong></a><strong>:</strong> AWS は、サイコロを振る動作をシミュレートし、過去のロールを追跡し、Streamable HTTP を使用して結果を返す、楽しいダンジョンズ アンド ドラゴンズをテーマにした MCP サーバーを発表しました。この軽量な例では、Lambda や Fargate などの AWS ツールとインフラストラクチャを使用して MCP サーバーを簡単に構築およびデプロイできることが強調されました。また、MCP サーバーと対話するマルチモーダル エージェントを構築するためのオープン ソース ツールキットである<a href="https://aws.amazon.com/blogs/opensource/introducing-strands-agents-an-open-source-ai-agents-sdk/">Strands SDK</a>も紹介されました。</p></li></ul><h2>Elastic Agent Builder での MCP サポート</h2><p><a href="https://www.elastic.co/jp/search-labs/blog/elastic-ai-agent-builder-context-engineering-introduction">Elastic Agent Builder を使用すると、</a>今すぐ MCP の実験を始めることができます。これは、データ上に直接エージェントを構築する最も簡単な方法です。Agent Builder を使用すると、Elasticsearch を利用したツールを MCP 対応エージェントに公開できます。また、次のような強力な組み込みツールがすでに付属しています。</p><ul><li><p><code>platform.core.search</code> - 完全なElasticsearchクエリDSLを使用して検索を実行します</p></li><li><p><code>platform.core.list_indices</code> - Elasticsearch 内で利用可能なすべてのインデックスを一覧表示します (エージェントがどのようなデータが存在するかを検出できるようにします)</p></li><li><p><code>platform.core.get_index_mapping</code> - 特定のインデックスのフィールド マッピングを取得します (エージェントがデータの形状と種類を理解するのに役立ちます)</p></li><li><p><code>platform.core.get_document_by_id</code> - IDで特定のドキュメントを取得します（正確な検索のため）</p></li></ul><p>これらのツールを使用するだけで、信頼性の高い AI エージェントを構築するための中核となるエンタープライズ レベルの検索と関連性をエージェントに装備できます。</p><p>Agent Builder をさらに強力にするのは、アプリケーションのニーズに合わせて独自のカスタム ツールを定義し、公開できる機能です。これは、毎回そのロジックを再検出することなく、エージェントが特定のインデックスに対して特定のタイプの検索を実行するようにしたい、意見が強いワークフローや繰り返し可能なワークフローに特に役立ちます。同じ結論に到達するために計画と推論にトークンを費やす代わりに、その意図をツールに直接エンコードすることで、エージェントの速度、信頼性、コスト効率を高めることができます。</p><p>Agent Builder UI 内で、ES|QL を使用するカスタム ツール定義の例を次に示します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltca9e3d0a4e7031c0/6a17e803faa913d8c393c897/c1f6405a374b707e8e6fa36b9e21db5f3c7cd127-1376x864.png" alt="エージェントビルダーUI" /><p>カスタム ツールを定義したら、 <code>Manage MCP</code>のドロップダウンをクリックして MCP サーバー URL をコピーすることで、MCP を使用してカスタム ツール (および組み込みのネイティブ ツール) を公開できます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf7d8b29b06c08f94/6a17e805033c8d07f06bb1b6/9f39588525ca2643475de557ea54a6bcf5c150f6-1282x616.png" alt="MCPツール" /><p>これで、この MCP エンドポイントを、MCP を使用する任意のクライアントにインポートして Agent Builder に接続し、利用可能なすべてのツールにアクセスできるようになります。詳細については、 <a href="https://www.elastic.co/jp/search-labs/blog/elastic-ai-agent-builder-context-engineering-introduction">Agent Builder</a>の紹介をお読みください。</p><h2>まとめ</h2><p>MCP Dev Summit では、MCP がこれらの AI エージェント同士、そして周囲のデータの世界と対話する方法を形作っていることが明らかになりました。エージェントをエンタープライズ データに接続する場合でも、完全に自律的なエージェントを設計する場合でも、MCP は標準化された構成可能な統合方法を提供し、大規模な環境ですぐに役立つようになります。トランスポート プロトコルやセキュリティ パターンからレジストリやガバナンスに至るまで、MCP エコシステムは急速に成熟しています。MCP は今後もオープンかつコミュニティ主導であり続けるため、今日の開発者には MCP の進化を形作るチャンスがあります。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/mcp-current-state</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/mcp-current-state</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[JD Armada]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2f63f23bbecd2a18/6a17e8066317302039585aa7/02b8c8672ffa129e0ed91a92d6cab612a01d27f2-1200x628.png" length="0" type="image/png"/>
    <pubDate>Thu, 12 Jun 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Spring AIとElasticsearchをベクターデータベースとして]]></title>
    <description><![CDATA[Spring AIとElasticsearchを使用して本番環境に対応したRAGアプリを構築し、ベクトルデータベースを使用してLLMと独自データを統合する方法をご覧ください。
]]></description>
    <content:encoded><![CDATA[<p><strong>Spring AI</strong>は現在一般公開されており、最初の<a href="https://spring.io/blog/2025/05/20/spring-ai-1-0-GA-released">安定版リリース 1.0</a>が<a href="https://mvnrepository.com/artifact/org.springframework.ai/spring-ai-core">Maven Central</a>からダウンロードできます。早速、お気に入りの<a href="https://www.elastic.co/what-is/large-language-models">LLM</a>とお気に入りの<a href="https://www.elastic.co/elasticsearch/vector-database">ベクター データベース</a>を使用して、完全な AI アプリケーションを構築してみましょう。または、最終的なアプリケーションを使用して<a href="https://github.com/xeraa/rag-with-java-springai-elasticsearch">リポジトリ</a>に直接アクセスします。</p><h2>Spring AI とは何ですか?</h2><p>AI 分野の急速な進歩の影響を受けたかなりの開発期間を経て、Java での AI エンジニアリングのための包括的なソリューションである<strong>Spring AI 1.0 が</strong>リリースされました。このリリースには、AI エンジニアにとって重要な新機能が多数含まれています。</p><p>Java と Spring は、この AI の波に乗るのに最適な位置にあります。数多くの企業が Spring Boot 上で業務を遂行しており、既存の業務に AI を組み込むのが非常に簡単になります。基本的に、あまり手間をかけずにビジネス ロジックとデータを AI モデルに直接リンクできます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltee1f144eb9e74866/6a17e379dbb4ffdf4bfb5647/328d7c51e1c145e94ea1e73ee9ff91836d3b180e-1600x773.png" alt="Spring AIをElasticsearchで使用する方法" /><p>Spring AI は、次のような<a href="https://docs.spring.io/spring-ai/reference/api/index.html">さまざまな AI モデルとテクノロジー</a>のサポートを提供します。</p><ul><li><p><strong>画像モデル</strong>: テキストプロンプトに基づいて画像を生成します。</p></li><li><p><strong>転写モデル</strong>: オーディオ ソースを取得してテキストに変換します。</p></li><li><p><strong>埋め込みモデル:</strong>任意のデータを、意味的類似性検索に最適化されたデータ型である<a href="https://www.elastic.co/what-is/vector-embedding">ベクトル</a>に変換します。</p></li><li><p><strong>チャットモデル:</strong>これら馴染みがあるはずです！あなたもきっとどこかで簡単な会話をしたことがあるはずです。</p></li></ul><p>チャット モデルは AI 分野で最も注目を集めているようですが、その通り、素晴らしいものです。文書の修正や詩の作成を手伝ってもらうこともできます。（ただ、まだジョークを言うように頼まないでください。）素晴らしいですが、いくつか問題もあります。</p><h2>AIの課題に対するSpring AIソリューション</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbd2d062ded38cf83/6a17e37adbb4ff69d7fb564b/2ebd68a90ebc73847df6ef7325936d4d06b35c8c-1600x900.jpg" alt="AIの課題に対するSpring AIソリューション" /><p>これらの問題のいくつかと Spring AI での解決策を見てみましょう。</p><p></p><p>問題</p><p>ソリューション</p><p>一貫性</p><p>チャットモデルはオープンマインドで、気が散りやすい</p><p>全体的な形状と構造を制御するためのシステムプロンプトを与えることができます</p><p>メモリ</p><p>AIモデルには記憶がないので、特定のユーザーからのメッセージを別のユーザーと関連付けることはできない。</p><p>会話の関連部分を保存するメモリシステムを与えることができます</p><p>分離</p><p>AIモデルは隔離された小さなサンドボックス内に存在しますが、ツール（必要だと判断したときに呼び出せる機能）へのアクセスを与えると、本当に素晴らしいことが可能になります。</p><p>Spring AI はツール呼び出しをサポートしており、これにより AI モデルに環境内のツールを通知し、そのツールを呼び出すように要求できます。この複数ターンのインタラクションはすべて透過的に処理されます</p><p>個人データ</p><p>AI モデルはスマートですが、全知ではありません。彼らはあなたの独自のデータベースに何が入っているか知りませんし、あなたもそれを知りたいとは思わないでしょう。</p><p>基本的には、モデルが質問を確認する前に、強力な文字列連結演算子を使用してリクエストにテキストを挿入することで、プロンプトを詰め込んで応答を通知する必要があります。背景情報です。何を送信して何を送信しないかをどのように決定しますか?ベクター ストアを使用して、関連するデータのみを選択し、それを送信します。これは検索拡張生成、またはRAGと呼ばれます。</p><p>幻覚</p><p>AI チャット モデルは、チャットするのが大好きです。そして時には、自信過剰になり、嘘をつくこともある。</p><p>妥当な結果を確認するには、評価（あるモデルを使用して別のモデルの出力を検証する）を使用する必要があります。</p><p></p><p>そしてもちろん、AI アプリケーションは孤立したものではありません。今日、最新の AI システムとサービスは、他のシステムやサービスと統合すると最も効果的に機能します。<a href="https://modelcontextprotocol.io/introduction"><strong>モデルコンテキストプロトコル</strong></a>(MCP) を使用すると、記述言語に関係なく、AI アプリケーションを他の MCP ベースのサービスに接続できるようになります。これらすべてを、より大きな目標に向かって進む<strong>エージェント</strong>ワークフローに組み込むことができます。</p><p>一番良かった点は？Spring Boot 開発者なら誰もが期待する使い慣れたイディオムと抽象化を基盤として、これらすべてを実現できます。基本的にすべてに対する便利なスターター依存関係が<a href="https://start.spring.io"><strong>Spring Initializr</strong></a>で利用できます<strong>。</strong></p><p>Spring AI は、既知の期待通りの、設定よりも規約を重視するセットアップを実現する便利な Spring Boot 自動設定を提供します。また、Spring AI は、Spring Boot の Actuator と Micrometer プロジェクトによって可観測性をサポートしています。GraalVM や仮想スレッドとも連携し、スケーラブルな超高速かつ効率的な AI アプリケーションを構築できます。</p><h2>Elasticsearchを選ぶ理由</h2><p>Elasticsearch は全文検索エンジンです。おそらくご存知でしょう。では、なぜこのプロジェクトでそれを使用するのでしょうか?まあ、ベクターストア<em>でも</em>あります！データが全文の隣に保存されるので、非常に便利です。その他の注目すべき利点:</p><ul><li><p>セットアップは超簡単</p></li><li><p>オープンソース</p></li><li><p>水平方向に拡張可能</p></li><li><p>組織の自由形式のデータのほとんどは、すでにElasticsearchクラスターに保存されているでしょう。</p></li><li><p>完全な検索エンジン機能を搭載</p></li><li><p><a href="https://docs.spring.io/spring-ai/reference/api/vectordbs/elasticsearch.html">Spring AI に完全に統合されています</a>!</p></li></ul><p>すべてを考慮すると、Elasticsearch は優れたベクター ストアに必要なすべての条件を満たしているので、セットアップしてアプリケーションの構築を開始しましょう。</p><h2>Elasticsearchを使い始める</h2><p>Elasticsearch と Kibana (データベースにホストされているデータを操作するために使用する UI コンソール) の両方が必要になります。</p><p>Docker イメージと<a href="http://elastic.co">Elastic.co ホームページの</a>おかげで、ローカル マシンですべてを試すことができます。そこに移動して、下にスクロールして<code>curl</code>コマンドを見つけ、それを実行してシェルに直接パイプします。</p> curl -fsSL https://elastic.co/start-local | sh 
  ______                     
 |  ____| |         | | (_)     
 | |__  | | __  ___| |_   ___ 
 |  __| | |/ _` / __| __| |/ __|
 | |____| | (_| \__ \ |_| | (__ 
 |______|_|\__,_|___/\__|_|\___|
-------------------------------------------------
🚀 Run Elasticsearch and Kibana for local testing
-------------------------------------------------
ℹ️  Do not use this script in a production environment
⌛️ Setting up Elasticsearch and Kibana v9.0.0...
- Generated random passwords
- Created the elastic-start-local folder containing the files:
  - .env, with settings
  - docker-compose.yml, for Docker services
  - start/stop/uninstall commands
- Running docker compose up --wait
[+] Running 25/26
 ✔ kibana_settings Pulled                                                 16.7s 
 ✔ kibana Pulled                                                          26.8s 
 ✔ elasticsearch Pulled                                                   17.4s                                                                     
[+] Running 6/6
 ✔ Network elastic-start-local_default             Created                 0.0s 
 ✔ Volume "elastic-start-local_dev-elasticsearch"  Created                 0.0s 
 ✔ Volume "elastic-start-local_dev-kibana"         Created                 0.0s 
 ✔ Container es-local-dev                          Healthy                12.9s 
 ✔ Container kibana_settings                       Exited                 11.9s 
 ✔ Container kibana-local-dev                      Healthy                21.8s 
🎉 Congrats, Elasticsearch and Kibana are installed and running in Docker!
🌐 Open your browser at http://localhost:5601
   Username: elastic
   Password: w1GB15uQ
🔌 Elasticsearch API endpoint: http://localhost:9200
🔑 API key: SERqaGlKWUJLNVJDODc1UGxjLWE6WFdxSTNvMU5SbVc5NDlKMEhpMzJmZw==
Learn more at https://github.com/elastic/start-local
➜  ~ <p>これにより、Elasticsearch と Kibana の Docker イメージがプルされて構成され、数分後には接続資格情報も揃ってローカル マシン上で稼働するようになります。</p><p>Elasticsearch インスタンスと対話するために使用できる 2 つの異なる URL もあります。プロンプトの指示に従って、ブラウザで<a href="http://localhost:5601">http://localhost:5601</a>にアクセスします。</p><p>コンソールに出力されるユーザー名<code>elastic</code>とパスワードにも注意してください。これらはログイン時に必要になります (上記の出力例では、それぞれ<code>elastic</code>と<code>w1GB15uQ</code>です)。</p><p></p><h2>アプリをまとめる</h2><p><a href="https://start.spring.io">Spring Initializr</a>ページに移動し、次の依存関係を持つ新しい Spring AI プロジェクトを生成します。</p><ul><li><p><code>Elasticsearch Vector Store</code></p></li><li><p><code>Spring Boot Actuator</code></p></li><li><p><code>GraalVM</code></p></li><li><p><code>OpenAI</code></p></li><li><p><code>Web</code></p></li></ul><p>必ず最新の Java バージョン (理想的には、この記事の執筆時点では Java 24 以降) と好みのビルド ツールを選択してください。この例では Apache Maven を使用しています。</p><p><code>Generate</code>をクリックし、プロジェクトを解凍して、選択した IDE にインポートします。(IntelliJ IDEA を使用しています。)</p><p>まず最初に、Spring Boot アプリケーションの接続詳細を指定しましょう。<code>application.properties,</code>に次のように記述します。</p>spring.elasticsearch.uris=http://localhost:9200
spring.elasticsearch.username=elastic
spring.elasticsearch.password=w1GB15uQ<p>また、データ構造に関して Elasticsearch 側で必要なものをすべて初期化するための Spring AI のベクトル ストア機能も使用するので、次のように指定します。</p>spring.ai.vectorstore.elasticsearch.initialize-schema=true<p>このデモでは<strong>OpenAI</strong> 、具体的には<strong>埋め込みモデル</strong>と<strong>チャット モデル</strong>を使用します ( <a href="https://docs.spring.io/spring-ai/reference/api/embeddings.html#available-implementations">Spring AI がサポートしている</a>限り、お好みのサービスを自由に使用してください)。</p><p>埋め込みモデルは、データを Elasticsearch に格納する前にデータの埋め込みを作成するために必要です。OpenAI が動作するには、 <code>API key</code>を指定する必要があります。</p>spring.ai.openai.api-key=...<p>ソース コード内に資格情報を保存することを避けるために、 <code>SPRING_AI_OPENAI_API_KEY</code>のような環境変数として定義することができます。</p><p>ファイルをアップロードするので、サーブレット コンテナにアップロードできるデータの量を必ずカスタマイズしてください。</p>spring.servlet.multipart.max-file-size=20MB
spring.servlet.multipart.max-request-size=20MB<p>もうすぐ到着です！コードの記述に進む前に、これがどのように機能するかをプレビューしてみましょう。</p><p>私たちのマシンでは、<a href="https://images-cdn.fantasyflightgames.com/filer_public/9f/aa/9faa23a3-9f71-4c77-865f-bba4aac8a258/runewars-revised-_rulebook.pdf">次のファイル</a>(ボードゲームのルールのリスト) をダウンロードし、名前を<code>test.pdf</code>に変更して<code>~/Downloads/test.pdf</code>に配置しました。</p><p>ファイルは<code>/rag/ingest</code>エンドポイントに送信されます (パスをローカル設定に応じて置き換えてください)。</p>http --form POST http://localhost:8080/rag/ingest path@/Users/jlong/Downloads/test.pdf<p>数秒かかる場合があります…</p><p>舞台裏では、データが OpenAI に送信され、データの埋め込みが作成されます。その後、そのデータはベクトルと元のテキストの両方で Elasticsearch に書き込まれます。</p><p>そのデータと、そこに含まれるすべての埋め込みによって、魔法が起こるのです。その後、 <code>VectorStore</code>インターフェースを使用して Elasticsearch をクエリできます。</p><p>完全なフローは次のようになります。</p><ul><li><p>HTTP クライアントは、選択した PDF を Spring アプリケーションにアップロードします。</p></li><li><p>Spring AI は PDF からテキストを抽出し、各ページを 800 文字のチャンクに分割します。</p></li><li><p>OpenAI は各チャンクのベクトル表現を生成します。</p></li><li><p>チャンク化されたテキストと埋め込みの両方が Elasticsearch に保存されます。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt93f4a64b634e9cce/6a17e37cb1e113215879f216/9734adb2d7128e61c515d5855dfad6d3a326a4a1-1454x706.png" alt="PDFからのSpring AIテキスト抽出、Open AIベクトル表現、埋め込みを作成するためのElasticsearchテキストチャンキングのための総合的なワークフロー。" /><p>最後に、クエリを発行します。</p>http :8080/rag/query question=="where do you place the reward card after obtaining it?" <p>そして、適切な答えが得られます。</p>After obtaining a Reward card, you place it facedown under the Hero card of the hero who received it.
Found at page: 28 of the manual<p>いいですね！これはどうやって動くんですか？</p><ul><li><p>HTTP クライアントは Spring アプリケーションに質問を送信します。</p></li><li><p>Spring AI は OpenAI から質問のベクトル表現を取得します。</p></li><li><p>この埋め込みにより、保存された Elasticsearch チャンク内で類似のドキュメントが検索され、最も類似したドキュメントが取得されます。</p></li><li><p>次に、Spring AI は質問と取得したコンテキストを OpenAI に送信し、LLM 回答を生成します。</p></li><li><p>最後に、生成された回答と取得されたコンテキストへの参照を返します。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfab41731851104f3/6a17e37e445de90e924d00aa/3799de6e8cb13ce49b9e136cfe593263030231a8-1464x1050.png" alt="質問に対するLLMの回答を生成するための総合的なSpring AIとOpen AIワークフロー" /><p>実際にどのように動作するかを確認するために、Java コードを調べてみましょう。</p><p>まず、 <strong>Main</strong>クラス: これは、あらゆる Spring Boot アプリケーションの標準的なメイン クラスです。</p>@SpringBootApplication
public class DemoApplication {
 	public static void main(String[] args) { 
     		SpringApplication.run(DemoApplication.class, args);
 	}
}<p>そこには何も見るものはありません。次に進みましょう…</p><p>次は、基本的な HTTP コントローラーです。</p>@RestController
class RagController {

   private final RagService rag;

   RagController(RagService rag) {
       this.rag = rag;
   }

   @PostMapping("/rag/ingest")
   ResponseEntity&lt;?&gt; ingestPDF(@RequestBody MultipartFile path) {
       rag.ingest(path.getResource());
       return ResponseEntity.ok().body("Done!");
   }

   @GetMapping("/rag/query")
   ResponseEntity&lt;?&gt; query(@RequestParam String question) {
       String response = rag.directRag(question);
       return ResponseEntity.ok().body(response);
   }
}<p>コントローラーは、ファイルの取り込みと Elasticsearch ベクター ストアへの書き込みを処理するために構築したサービスを呼び出し、同じベクター ストアに対するクエリを容易にするだけです。</p><p>サービスを見てみましょう:</p>@Service
class RagService {

   private final ElasticsearchVectorStore vectorStore;

   private final ChatClient ai;

   RagService(ElasticsearchVectorStore vectorStore, ChatClient.Builder clientBuilder) {
       this.vectorStore = vectorStore;
       this.ai = clientBuilder.build();
   }

   void ingest(Resource path) {
       PagePdfDocumentReader pdfReader = new PagePdfDocumentReader(path);
       List&lt;Document&gt; batch = new TokenTextSplitter().apply(pdfReader.read());
       vectorStore.add(batch);
   }

  // TBD
}<p>このコードはすべての取り込みを処理します。バイトを囲むコンテナーである Spring Framework <code>Resource</code>が指定されると、Spring AI の<code>PagePdfDocumentReader</code>を使用して PDF データ ( <code>.PDF</code>ファイルであると想定されます。任意の入力を受け入れる前に必ず検証してください) を読み取り、次に Spring AI の<code>TokenTextSplitter</code>を使用してトークン化し、最後に結果の<code>List&lt;Document&gt;</code>を<code>VectorStore</code>実装<code>ElasticsearchVectorStore</code>に追加します。</p><p>Kibana を使用してこれを確認できます。ファイルを<code>/rag/ingest</code>エンドポイントに送信した後、ブラウザで<code>localhost:5601</code>を開き、左側のサイド メニューで<code>Dev Tools</code>に移動します。ここでクエリを発行して、Elasticsearch インスタンス内のデータと対話できます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt45805a5b2da5e336/6a17e3803e03d70a584f2bda/c85e522f02f8b2da7462cd428dc7e952c9692542-1600x1040.png" alt="Elasticsearchインスタンスでクエリを作成する方法。" /><p>次のようなクエリを発行します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt21d79210fe213b1e/6a17e382e3179163492d5767/00974a176cbce11e70fcab24fb4b3f9c6e205982-1600x1040.png" alt="Elasticsearch Consoleでクエリを発行します。" /><p>さて、ここからが面白いところです。ユーザーのクエリに応じて、そのデータをどうやって再び取り出すのでしょうか?</p><p>以下は、 <code>directRag</code>と呼ばれるメソッドでのクエリの実装の最初の例です。</p>String directRag(String question) {
   // Query the vector store for documents related to the question
   List&lt;Document&gt; vectorStoreResult =
           vectorStore.doSimilaritySearch(SearchRequest.builder().query(question).topK(5)
                   .similarityThreshold(0.7).build());

   // Merging the documents into a single string
   String documents = vectorStoreResult.stream()
           .map(Document::getText)
           .collect(Collectors.joining(System.lineSeparator()));

   // Exit if the vector search didn't find any results
   if (documents.isEmpty()) {
       return "No relevant context found. Please change your question.";
   }

   // Setting the prompt with the context
   String prompt = """
           You're assisting with providing the rules of the tabletop game Runewars.
           Use the information from the DOCUMENTS section to provide accurate answers to the
           question in the QUESTION section.
           If unsure, simply state that you don't know.
          
           DOCUMENTS:
           """ + documents
           + """
           QUESTION:
           """ + question;


   // Calling the chat model with the question
   String response = ai
           .prompt()
           .user(prompt)
           .call()
           .content();

   return response +
           System.lineSeparator() +
           "Found at page: " +
           // Retrieving the first ranked page number from the document metadata
           vectorStoreResult.getFirst().getMetadata().get(PagePdfDocumentReader.METADATA_START_PAGE_NUMBER) +
           " of the manual";

}<p>コードは非常に単純ですが、複数のステップに分解してみましょう。</p><ol><li><p>類似検索を実行するには、 <code>VectorStore</code>を使用します。</p></li><li><p>すべての結果が与えられたら、基礎となる Spring AI <code>Document</code>を取得し、そのテキストを抽出して、すべてを 1 つの結果に連結します。</p></li><li><p><code>VectorStore</code>からの結果を、その結果の処理方法をモデルに指示するプロンプトとユーザーからの質問とともにモデルに送信します。応答を待って返します。</p></li></ol><p></p><p>これは<strong>RAG</strong> (検索拡張生成) です。これは、ベクトル ストアのデータを使用して、モデルによって実行される処理と分析に情報を提供するという考え方です。やり方がわかったので、もうそんなことをしなくて済むことを祈ります!とにかく、これは違います。Spring AI の<a href="https://docs.spring.io/spring-ai/reference/api/advisors.html">アドバイザーは</a>、このプロセスをさらに簡素化するために存在します。</p><p>アドバイザーを使用すると、アプリケーションとベクター ストアの間に抽象化レイヤーを提供するだけでなく、特定のモデルへのリクエストを前処理および後処理できます。ビルドに次の依存関係を追加します。
</p>&lt;dependency&gt;
   &lt;groupId&gt;org.springframework.ai&lt;/groupId&gt;
   &lt;artifactId&gt;spring-ai-advisors-vector-store&lt;/artifactId&gt;
&lt;/dependency&gt;<p>クラスに<code>advisedRag(String question)</code>という別のメソッドを追加します。</p>String advisedRag(String question) {
   return this.ai
           .prompt()
           .user(question)
           .advisors(new QuestionAnswerAdvisor(vectorStore))
           .call()
           .content();
}<p>すべての RAG パターン ロジックは<code>QuestionAnswerAdvisor</code>にカプセル化されます。その他は、 <code>ChatModel</code>へのリクエストと同じです。ニース！</p><p><a href="https://github.com/xeraa/rag-with-java-springai-elasticsearch">完全なコードは GitHub から入手</a>できます。</p><h2>まとめ</h2><p>このデモでは、Docker イメージを使用してすべてをローカル マシン上で実行しましたが、ここでの目標は、本番環境に適した AI システムとサービスを構築することです。それを実現するためにできることはいくつかあります。</p><p>まず、トークンの消費を監視するために<a href="https://docs.spring.io/spring-boot/reference/actuator/index.html#actuator">Spring Boot Actuator を</a>追加できます。トークンは、モデルに対する特定のリクエストの複雑さ (場合によってはドルとセント) のコストを表す代理です。</p><p>クラスパスにはすでに Spring Boot Actuator があるので、次のプロパティを指定するだけで、すべてのメトリック (素晴らしい<a href="http://micrometer.io">Micrometer.io</a>プロジェクトによってキャプチャされたもの) が表示されます。</p>management.endpoints.web.exposure.include=*<p>アプリケーションを再起動してください。クエリを作成して、 <a href="http://localhost:8080/actuator/metrics">http://localhost:8080/actuator/metrics</a>にアクセスします。「 <code>token</code> 」を検索すると、アプリケーションで使用されているトークンに関する情報が表示されます。これに注意してください。もちろん、Micrometer の<a href="https://docs.micrometer.io/micrometer/reference/implementations/elastic.html">Elasticsearch 統合を</a>使用してこれらのメトリックをプッシュし、Elasticsearch を好みの時系列データベースとして機能させることもできます。</p><p>次に、Elasticsearch などのデータストア、OpenAI、またはその他のネットワーク サービスにリクエストを送信するたびに IO が実行され、多くの場合、その IO によって実行中のスレッドがブロックされることを考慮する必要があります。Java 21 以降には、スケーラビリティを大幅に向上させる非ブロッキング<strong>仮想スレッド</strong>が搭載されています。有効にするには以下を実行します:
</p>spring.threads.virtual.enabled=true<p>そして最後に、アプリケーションとデータを、それが繁栄し、拡張できる場所でホストすることが必要になります。アプリケーションをどこで実行するかについてはおそらくすでにお考えだと思いますが、データはどこにホストするのでしょうか?<a href="https://cloud.elastic.co/">Elastic Cloud</a>をお勧めしますか?安全、プライベート、スケーラブル、そして機能が満載です。私たちのお気に入りの部分は？ご希望の場合は、Elastic がポケベルを装着し、お客様がポケベルを装着する Serverless エディションを入手できます。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/spring-ai-elasticsearch-application</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/spring-ai-elasticsearch-application</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Josh Long,Philipp Krenn,Laura Trotta]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt26b868ef618164c6/6a17e3830b0bedd68cdd3515/0771fb5b3d9234697cb868cd7d9d1b840000bf29-1280x720.png" length="0" type="image/png"/>
    <pubDate>Tue, 20 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[LangGraphとElasticsearchを使用したRAGワークフローの構築]]></title>
    <description><![CDATA[Elasticsearch を使用して LangGraph 検索エージェント テンプレートを構成およびカスタマイズし、効率的なデータ取得と AI 駆動型応答のための RAG ワークフローを構築する方法を学習します。]]></description>
    <content:encoded><![CDATA[<p><a href="https://github.com/langchain-ai/retrieval-agent-template">LangGraph 検索エージェント テンプレートは、</a> LangGraph Studio で LangGraph を使用して検索ベースの質問応答システムの作成を容易にするために LangChain によって開発されたスターター プロジェクトです。このテンプレートは Elasticsearch とシームレスに統合するように事前構成されているため、開発者はドキュメントを効率的にインデックスして取得できるエージェントを迅速に構築できます。</p><p>このブログでは、LangGraph Studio と LangGraph CLI を使用して LangChain 検索エージェント テンプレートを実行およびカスタマイズすることに焦点を当てています。このテンプレートは、Elasticsearch などのさまざまな検索バックエンドを活用して、検索拡張生成 (RAG) アプリケーションを構築するためのフレームワークを提供します。</p><p>エージェントフローをカスタマイズしながら、Elastic を使用してテンプレートを効率的にセットアップ、環境の構成、実行する方法を説明します。</p><h2>要件</h2><p>続行する前に、以下がインストールされていることを確認してください。</p><ul><li><p>Elasticsearch Cloud デプロイメントまたはオンプレミス Elasticsearch デプロイメント (または Elastic Cloud で 14 日間の<a href="https://www.elastic.co/jp/cloud/cloud-trial-overview">無料トライアル</a>を作成) - バージョン 8.0.0 以上</p></li><li><p>Python 3.9以上</p></li><li><p><a href="https://cohere.com/">Cohere</a> （このガイドで使用）、 <a href="https://openai.com/">OpenAI</a> 、 <a href="https://www.anthropic.com/claude">Anthropic/Claude</a>などのLLMプロバイダーへのアクセス</p></li></ul><h2>LangGraphアプリの作成</h2><h3>1. LangGraph CLIをインストールする</h3>pip install --upgrade "langgraph-cli[inmem]"<h3>2. 検索エージェントテンプレートからLangGraphアプリを作成する</h3>mkdir lg-agent-demo
cd lg-agent-demo
langgraph new lg-agent-demo <p><em>利用可能なテンプレートのリストから選択できるインタラクティブ メニューが表示されます。</em>以下に示すように、取得エージェントの場合は 4、Python の場合は 1 を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd44177037ea46d45/6a17f86b3e9e45265fba1663/6a41a41f95c2477c67810adc7be46d91faf06878-1600x407.png" alt="インタラクティブな検索テンプレート。" /><ul><li><p><strong>トラブルシューティング</strong>: 「urllib.error.URLError: &lt;urlopen error [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: cannot get local issuer certificate (_ssl.c:1000)&gt;」というエラーが発生した場合「</p></li></ul><p>問題を解決するには、以下に示すように、Python の証明書インストール コマンドを実行してください。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbbe2d1d3a1af75b1/6a17f86d445de97c9c4d02e7/83ec238136c41738457299fd42c83aff32eb5b97-1407x75.png" alt="Python の証明書インストール コマンドを実行します。" /><h3>3. 依存関係をインストールする</h3><p>新しい LangGraph アプリのルートで仮想環境を作成し、依存関係を<code>edit</code>モードでインストールして、ローカルの変更がサーバーで使用されるようにします。</p>#For Mac
python3 -m venv lg-demo
source lg-demo/bin/activate 
pip install -e .

#For Windows
python3 -m venv lg-demo
lg-demo\Scripts\activate 
pip install -e .<h2>環境の設定</h2><h3>1. .environmentを作成するファイル</h3><p><code>.env</code>ファイルには、アプリが選択した LLM および取得プロバイダーに接続できるようにするための API キーと構成が保持されます。サンプル設定を複製して新しい<code>.env</code>ファイルを生成します。</p>cp .env.example .env<h3>2. .envを設定するファイル</h3><p><code>.env</code>ファイルには、デフォルトの構成のセットが付属しています。設定に応じて必要な API キーと値を追加することで更新できます。ユースケースに関係のないキーは、変更せずにそのままにしておくことも、削除することもできます。</p># To separate your traces from other applications
LANGSMITH_PROJECT=retrieval-agent

# LLM choice (set the API key for your selected provider):
ANTHROPIC_API_KEY=your_anthropic_api_key
FIREWORKS_API_KEY=your_fireworks_api_key
OPENAI_API_KEY=your_openai_api_key

# Retrieval provider (configure based on your chosen service):

## Elastic Cloud:
ELASTICSEARCH_URL=https://your_elastic_cloud_url
ELASTICSEARCH_API_KEY=your_elastic_api_key

## Elastic Local:
ELASTICSEARCH_URL=http://host.docker.internal:9200
ELASTICSEARCH_USER=elastic
ELASTICSEARCH_PASSWORD=changeme

## Pinecone:
PINECONE_API_KEY=your_pinecone_api_key
PINECONE_INDEX_NAME=your_pinecone_index_name

## MongoDB Atlas:
MONGODB_URI=your_mongodb_connection_string

# Cohere API key:
COHERE_API_KEY=your_cohere_api_key<ul><li><p>サンプル<code>.env</code>ファイル (Elastic Cloud と Cohere を使用)</p></li></ul><p>以下は、このブログで説明されているように、 <strong>Elastic Cloud を</strong>取得プロバイダーとして使用し、 <strong>Cohere を</strong>LLM として使用するためのサンプルの<code>.env</code>構成です。</p># To separate your traces from other applications
LANGSMITH_PROJECT=retrieval-agent
#Retrieval Provider
# Elasticsearch configuration
ELASTICSEARCH_URL=elastic-url:443
ELASTICSEARCH_API_KEY=elastic_api_key
# Cohere API key
COHERE_API_KEY=cohere_api_key<p><em>注: このガイドでは、レスポンス生成と埋め込みの両方にCohereを使用していますが、ユースケースに応じて、</em><em><strong> OpenAI</strong></em><em> 、</em><em><strong> Claude</strong></em><em> 、あるいはローカルLLMモデル などの他のLLMプロバイダーも自由に使用できます 。使用する予定の各キーが</em><em><code>.env</code></em><em> ファイル に存在し、正しく設定されていることを確認してください。</em></p><h3>3. 設定ファイル -configuration.py を更新する </h3><p>適切な API キーを使用して<code>.env</code>ファイルを設定したら、次の手順ではアプリケーションのデフォルトのモデル構成を更新します。構成を更新すると、システムは<code>.env</code>ファイルで指定したサービスとモデルを使用するようになります。</p><p>構成ファイルに移動します。</p> cd src/retrieval_graph<p><code>configuration.py</code>ファイルには、検索エージェントが 3 つの主なタスクに使用するデフォルトのモデル設定が含まれています。</p><ul><li><p><strong>埋め込みモデル</strong>– ドキュメントをベクトル表現に変換する</p></li><li><p><strong>クエリモデル</strong>– ユーザーのクエリをベクトルに変換する</p></li><li><p><strong>レスポンスモデル</strong>– 最終的なレスポンスを生成する</p></li></ul><p>デフォルトでは、コードは<strong>OpenAI</strong> (例: <code>openai/text-embedding-3-small</code> ) と<strong>Anthropic</strong> (例: <code>anthropic/claude-3-5-sonnet-20240620 and anthropic/claude-3-haiku-20240307</code> ) のモデルを使用します。このブログでは、Cohere モデルの使用に切り替えます。すでに OpenAI または Anthropic を使用している場合は、変更は必要ありません。</p><h4>変更例（Cohere を使用）:</h4><p><code>configuration.py</code>を開き、以下のようにモデルのデフォルトを変更します。</p>…
 embedding_model: Annotated[
       str,
       {"__template_metadata__": {"kind": "embeddings"}},
   ] = field(
       default="cohere/embed-english-v3.0",
…
response_model: Annotated[str, {"__template_metadata__": {"kind": "llm"}}] = field(
       default="cohere/command-r-08-2024",
…
query_model: Annotated[str, {"__template_metadata__": {"kind": "llm"}}] = field(
       default="cohere/command-r-08-2024",
       metadata={<h2>LangGraph CLI で取得エージェントを実行する</h2><h3>1. LangGraphサーバーを起動する</h3>cd lg-agent-demo
langgraph dev<p>これにより、LangGraph API サーバーがローカルで起動します。これが正常に実行されると、次のような画面が表示されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt46c7a703e715ef66/6a17f86eb1e113272d79f42e/e3c3344b24651067e2d0892d870feca505b3be35-1494x542.png" alt=" LangGraph API サーバーが正常に実行されています。" /><p>Studio UI URL を開きます。</p><p>利用可能なグラフは 2 つあります。</p><ul><li><p><strong>取得グラフ</strong>: Elasticsearch からデータを取得し、LLM を使用してクエリに応答します。</p></li><li><p><strong>インデクサー グラフ</strong>: ドキュメントを Elasticsearch にインデックスし、LLM を使用して埋め込みを生成します。</p></li></ul><h3>2. インデクサーグラフの設定</h3><ul><li><p>インデクサー グラフを開きます。</p></li><li><p>アシスタントの管理をクリックします。</p><ul><li><p><strong>「新しいアシスタントを追加</strong>」をクリックし、指定どおりにユーザーの詳細を入力して、ウィンドウを閉じます。</p></li></ul></li></ul>{"user_id": "101"}<h3>3. サンプル文書のインデックス作成</h3><ul><li><p>NoveTech という組織の仮想的な四半期レポートを表す次のサンプル ドキュメントにインデックスを付けます。</p></li></ul>[
  {    "page_content": "NoveTech Solutions Q1 2025 Report - Revenue: $120.5M, Net Profit: $18.2M, EPS: $2.15. Strong AI software launch and $50M government contract secured."
  },
  {
    "page_content": "NoveTech Solutions Business Highlights - AI-driven analytics software gained 15% market share. Expansion into Southeast Asia with two new offices. Cloud security contract secured."
  },
  {
    "page_content": "NoveTech Solutions Financial Overview - Operating expenses at $85.3M, Gross Margin 29.3%. Stock price rose from $72.5 to $78.3. Market Cap reached $5.2B."
  },
  {
    "page_content": "NoveTech Solutions Challenges - Rising supply chain costs impacting hardware production. Regulatory delays slowing European expansion. Competitive pressure in cybersecurity sector."
  },
  {
    "page_content": "NoveTech Solutions Future Outlook - Expected revenue for Q2 2025: $135M. New AI chatbot and blockchain security platform launch planned. Expansion into Latin America."
  },
  {
    "page_content": "NoveTech Solutions Market Performance - Year-over-Year growth at 12.7%. Stock price increase reflects investor confidence. Cybersecurity and AI sectors remain competitive."
  },
  {
    "page_content": "NoveTech Solutions Strategic Moves - Investing in R&amp;D to enhance AI-driven automation. Strengthening partnerships with enterprise cloud providers. Focusing on data privacy solutions."
  },
  {
    "page_content": "NoveTech Solutions CEO Statement - 'NoveTech Solutions continues to innovate in AI and cybersecurity. Our growth strategy remains strong, and we foresee steady expansion in the coming quarters.'"
  }
]<p>ドキュメントがインデックスされると、以下に示すように、スレッドに削除メッセージが表示されます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt38715eadffcb62f0/6a17f877faa9135f7393ca4c/fd3a1efd64cb54d54ea56ef5055249dd066d5708-1600x854.png" alt="LangGraph および Elasticsearch RAG ワークフロー ドキュメントがインデックス化されました。" /><h3>4. 検索グラフの実行</h3><ul><li><p>検索グラフに切り替えます。</p></li><li><p>次の検索クエリを入力してください。</p></li></ul>What was NovaTech Solutions total revenue in Q1 2025?<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c0d070fca52e512/6a17f879505ac36d37ad8d12/eb4d8ddfe0effd7e1868fba921b8ef13f7baf27a-1600x755.png" alt="LangGraphとElasticsearchの検索グラフを実行する" /><p>システムは関連するドキュメントを返し、インデックスされたデータに基づいて正確な回答を提供します。</p><h2>検索エージェントをカスタマイズする</h2><p>ユーザー エクスペリエンスを向上させるために、検索グラフにカスタマイズ ステップを導入し、ユーザーが次に尋ねる可能性のある 3 つの質問を予測します。この予測は以下に基づいています:</p><ul><li><p>取得した文書のコンテキスト</p></li><li><p>以前のユーザーインタラクション</p></li><li><p>最後のユーザークエリ</p></li></ul><p>クエリ予測機能を実装するには、次のコード変更が必要です。</p><h3>1. graph.pyを更新する</h3><ul><li><p><code>predict_query</code>関数を追加します:</p></li></ul>async def predict_query(
   state: State, *, config: RunnableConfig
) -&gt; dict[str, list[BaseMessage]]:
   logger.info(f"predict_query predict_querypredict_query predict_query predict_query predict_query")  # Log the query

   configuration = Configuration.from_runnable_config(config)
   prompt = ChatPromptTemplate.from_messages(
       [
           ("system", configuration.predict_next_question_prompt),
           ("placeholder", "{messages}"),
       ]
   )
   model = load_chat_model(configuration.response_model)
   user_query = state.queries[-1] if state.queries else "No prior query available"
   logger.info(f"user_query: {user_query}")
   logger.info(f"statemessage: {state.messages}")
   #human_messages = [msg for msg in state.message if isinstance(msg, HumanMessage)]

   message_value = await prompt.ainvoke(
       {
           "messages": state.messages,
           "user_query": user_query,  # Use the most recent query as primary input
           "system_time": datetime.now(tz=timezone.utc).isoformat(),
       },
       config,
   )

   next_question = await model.ainvoke(message_value, config)
   return {"next_question": [next_question]}<ul><li><p><code>respond</code>関数を変更して、メッセージの代わりに<strong><code>response</code></strong>オブジェクトを返します。</p></li></ul>async def respond(
   state: State, *, config: RunnableConfig
) -&gt; dict[str, list[BaseMessage]]:
   """Call the LLM powering our "agent"."""
   configuration = Configuration.from_runnable_config(config)
   # Feel free to customize the prompt, model, and other logic!
   prompt = ChatPromptTemplate.from_messages(
       [
           ("system", configuration.response_system_prompt),
           ("placeholder", "{messages}"),
       ]
   )
   model = load_chat_model(configuration.response_model)

   retrieved_docs = format_docs(state.retrieved_docs)
   message_value = await prompt.ainvoke(
       {
           "messages": state.messages,
           "retrieved_docs": retrieved_docs,
           "system_time": datetime.now(tz=timezone.utc).isoformat(),
       },
       config,
   )
   response = await model.ainvoke(message_value, config)
   # We return a list, because this will get added to the existing list
   return {"response": [response]}<ul><li><p>グラフ構造を更新して、predict_query に新しいノードとエッジを追加します。</p></li></ul>builder.add_node(generate_query)
builder.add_node(retrieve)
builder.add_node(respond)
builder.add_node(predict_query)
builder.add_edge("__start__", "generate_query")
builder.add_edge("generate_query", "retrieve")
builder.add_edge("retrieve", "respond")
builder.add_edge("respond", "predict_query")<h3>2. prompts.pyを更新する</h3><ul><li><p><code>prompts.py</code>でのクエリ予測のプロンプトを作成します:</p></li></ul>PREDICT_NEXT_QUESTION_PROMPT = """Given the user query and the retrieved documents, suggest the most likely next question the user might ask.

**Context:**
- Previous Queries:
{previous_queries}

- Latest User Query: {user_query}

- Retrieved Documents:
{retrieved_docs}

**Guidelines:**
1. Do not suggest a question that has already been asked in previous queries.
2. Consider the retrieved documents when predicting the next logical question.
3. If the user's query is already fully answered, suggest a relevant follow-up question.
4. Keep the suggested question natural and conversational.
5. Suggest at least 3 question

System time: {system_time}"""<h3>3. configuration.pyを更新する</h3><ul><li><p><code>predict_next_question_prompt</code>追加:</p></li></ul>predict_next_question_prompt: str = field(
       default=prompts.PREDICT_NEXT_QUESTION_PROMPT,
       metadata={"description": "The system prompt used for generating responses."},
   )<h3>4. state.pyを更新する</h3><ul><li><p>次の属性を追加します。</p></li></ul>response: Annotated[Sequence[AnyMessage], add_messages]
next_question : Annotated[Sequence[AnyMessage], add_messages]<h3>5. 検索グラフを再実行する</h3><ul><li><p>次の検索クエリをもう一度入力してください。</p></li></ul>What was NovaTech Solutions total revenue in Q1 2025?<p>システムは入力を処理し、以下に示すように、ユーザーが尋ねる可能性のある 3 つの関連する質問を予測します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd8d4de9396ca4853/6a17f87be31791dc742d59c7/70e855a2e4edc0ba5a147588df0de30eb081d053-1600x777.png" alt="LangGraphとElasticsearchを使用して3つのユーザーの質問による検索グラフを実行する" /><h2>まとめ</h2><p>LangGraph Studio と CLI 内に取得エージェント テンプレートを統合すると、いくつかの重要な利点が得られます。</p><ul><li><p><strong>開発の加速</strong>: テンプレートと視覚化ツールにより、検索ワークフローの作成とデバッグが効率化され、開発時間が短縮されます。</p></li><li><p><strong>シームレスな展開</strong>: API と自動スケーリングの組み込みサポートにより、環境間でのスムーズな展開が保証されます。</p></li><li><p><strong>簡単な更新:</strong>ワークフローの変更、新しい機能の追加、追加のノードの統合が簡単なので、検索プロセスの拡張と強化が容易になります。</p></li><li><p><strong>永続メモリ</strong>: システムはエージェントの状態と知識を保持し、一貫性と信頼性を向上させます。</p></li><li><p><strong>柔軟なワークフロー モデリング</strong>: 開発者は、特定のユース ケースに合わせて検索ロジックと通信ルールをカスタマイズできます。</p></li><li><p><strong>リアルタイムの対話とデバッグ</strong>: 実行中のエージェントと対話できるため、効率的なテストと問題解決が可能になります。</p></li></ul><p>これらの機能を活用することで、組織はデータのアクセシビリティとユーザー エクスペリエンスを向上させる強力で効率的かつスケーラブルな検索システムを構築できます。</p><p>このプロジェクトの完全なソースコードは<a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/langraph-retrieval-agent-template-demo">GitHub</a>で入手できます。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/build-rag-workflow-langgraph-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/build-rag-workflow-langgraph-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Neha Saini]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0c9f03d2c1a9cb4c/6a17f87d0b0bed3781dd377c/17b7e7b336f73e232375d1add582ae5f6c52a279-1440x840.png" length="0" type="image/png"/>
    <pubDate>Fri, 25 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch で Amazon Nova モデルを使用する]]></title>
    <description><![CDATA[ElasticsearchのAmazon Novaモデルを使用して、Elasticsearchの製品レビューからセンチメント、信頼性、要約、キーワードを自動的に抽出する方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>この記事では、Amazon の AI モデル ファミリーである Amazon Nova について説明し、Elasticsearch と併用する方法を学びます。</p><h2>Amazon Novaについて</h2><p>Amazon Nova は、Amazon Bedrock で利用できる Amazon 人工知能モデルのファミリーであり、高いパフォーマンスとコスト効率を実現するように設計されています。これらのモデルは、テキスト、画像、ビデオの入力を処理し、テキスト出力を生成し、さまざまな精度、速度、コストのニーズに合わせて最適化されています。</p><h3>Amazon Novaの主なモデル</h3><ul><li><p>Amazon Nova Micro: テキストのみに焦点を当てた、高速でコスト効率の高いモデルであり、翻訳、推論、コード補完、数学の問題の解決に最適です。1 秒あたり 200 トークンを超える生成が可能なので、即時の応答が必要なアプリケーションに最適です。</p></li><li><p>Amazon Nova Lite：画像、動画、テキストを高速処理できる低コストのマルチモーダルモデル。スピードと精度に優れており、コストが重要な要素となるインタラクティブな高容量アプリケーションに適しています。</p></li><li><p>Amazon Nova Pro: 高精度、高速、コスト効率を兼ね備えた最も高度なオプション。ビデオ要約、質疑応答、ソフトウェア開発、AI エージェントなどの複雑なタスクに最適です。専門家のレビューでは、テキストとビジュアルの理解力に優れているだけでなく、指示に従って自動化されたワークフローを実行する能力も証明されています。</p></li></ul><p>Amazon Nova モデルは、コンテンツ作成やデータ分析からソフトウェア開発や AI を活用したプロセス自動化まで、さまざまなアプリケーションに適しています。</p><p>以下では、Amazon Nova モデルを Elasticsearch と組み合わせて使用し、製品レビューを自動化する方法を説明します。</p><p>私たちが行うこと:</p><ol><li><p>推論 API 経由でエンドポイントを作成し、Amazon Bedrock と Elasticsearch を統合します。</p></li><li><p>推論プロセッサを使用してパイプラインを作成し、推論 API エンドポイントを呼び出します。</p></li><li><p>製品レビューをインデックス化し、パイプラインを使用してレビューの分析を自動的に生成します。</p></li><li><p>統合の結果を分析します。</p></li></ol><h2>Amazon Nova Liteを使用して推論APIにエンドポイントを作成</h2><p>まず、Amazon Bedrock と Elasticsearch を統合するために Inference API を設定します。Amazon Nova Lite（ID <strong>amazon.nova-lite-v1:0</strong> ）を定義します。速度、精度、コストのバランスが取れているため、使用すべきモデルです。</p><p><strong>注意:</strong> Amazon Bedrock を使用するには有効な認証情報が必要です。アクセス キーを取得するためのドキュメントは、<a href="https://docs.aws.amazon.com/keyspaces/latest/devguide/create.keypair.html">こちらで</a>参照できます。</p>PUT _inference/completion/bedrock_completion_amazon_nova-lite
{
   "service": "amazonbedrock",
   "service_settings": {
       "access_key": "#access_key#",
       "secret_key": "#secret_key#",
       "region": "us-east-1",
       "provider": "amazontitan",
       "model": "amazon.nova-lite-v1:0"
   }
}<h2>レビュー分析パイプラインの作成</h2><p>ここで、推論プロセッサを使用してレビュー分析プロンプトを実行する処理パイプラインを作成します。このプロンプトにより、レビュー データが Amazon Nova Lite に送信され、次の処理が実行されます。</p><ul><li><p>感情の分類（肯定的、否定的、または中立的）。</p></li><li><p>レビューの要約。</p></li><li><p>キーワードの生成。</p></li><li><p>信頼性の測定 (本物 | 疑わしい | 一般的な)。</p></li></ul>PUT /_ingest/pipeline/review_analyzer_ai
{
      "processors": [
      {
        "script": 
            {
            "source": """ctx.prompt = "Analyze the following product review and return a structured JSON. Task: - Summarize the review concisely. - Detect and classify the sentiment as positive, neutral, or negative.- Generate relevant tags (keywords) based on the review content and detected sentiment. - Evaluate the authenticity of the review (authentic, suspicious, or generic). Review: " + ctx.review + " Respond in JSON format with the following fields: \"review_analyze\": {\"sentiment\": \"&lt;positive | neutral | negative&gt;\", \"authenticity\": \"&lt;authentic | suspicious | generic&gt;\",\"summary\": \"&lt;short review summary&gt;\", \"keywords\": [\"&lt;keyword 1&gt;\", \"&lt;keyword 2&gt;\", \"...\"]}}}"
            """
            }
      },
      {
        "inference": {
          "model_id": "bedrock_completion_amazon_nova-lite",
          "input_output": {
            "input_field": "prompt",
            "output_field": "result"
          }
        }
      },
      {
        "gsub": {
          "field": "result",
          "pattern": "```json",
          "replacement": ""
        } 
      },
      {
        "json" : {
          "field" : "result",
          "strict_json_parsing": false,
          "add_to_root" : true
        }
      },
      {
        "remove": {
          "field": "result"
        }
      },
      {
        "remove": {
          "field": "prompt"
        }
      }
    ]
}<h2>レビューのインデックス作成</h2><p>ここで、Bulk API を使用して製品レビューのインデックスを作成します。先ほど作成したパイプラインが自動的に適用され、Nova モデルによって生成された分析がインデックス付けされたドキュメントに追加されます。</p>POST bulk/
{ "index": { "_index" : "products", "_id": 1, "pipeline":"review_analyzer_ai" } }
{ "product": "Pampers Pants Premium Care Fralda", "review": "Best diaper ever! Great material, lots of cotton, without all that plastic. Doesn't leak! My baby is a boy and every diaper leaked around the waist, this model solved the problem. Even on a small baby it's worth the effort of putting on the short diaper. I put it on my baby at 9 pm and only take it off in the morning, without any leaks." }
{ "index": { "_index" : "products", "_id": 2, "pipeline":"review_analyzer_ai" } }
{ "product": "Portable Electric Body Massager", "review": "It broke in three months for no apparent reason, thank goodness I didn't review it before. I don't recommend buying it because it has a short lifespan." }
{ "index": { "_index" : "products", "_id": 3, "pipeline":"review_analyzer_ai" } }
{ "product": "Havit Fuxi-H3 Black Quad-Mode Wired and Wireless Gaming Headset", "review": "The sound is good for the price, but the connectivity is horrible. You always need to be playing audio, otherwise it loses connection (I work from home, and this is very annoying). Sometimes it loses connection and you have to turn it off and on again to get it back on. The microphone is very sensitive, so it loses connection frequently and you have to turn the headset off and on for the microphone to work again. The flexibility of the stem is useless, because if you move it, the microphone can turn off. Sometimes I need to use Linux and the headset simply doesn't work. It's light and comfortable, the sound is adequate, but the connectivity is terrible." }
{ "index": { "_index" : "products", "_id": 4, "pipeline":"review_analyzer_ai" } }
{ "product": "Air Fryer 4L Oil Free Fryer Mondial", "review": "For those looking for value for money, it's a good option, but the tray (which is underneath the perforated basket) is already peeling a lot. My mother has one just like it and said that hers is even rusting, in other words, the material is MUCH inferior. There's also something that bothers me, because it looks like a microwave, it doesn't fry evenly, it's weaker in the middle and stronger on the sides. Buy at your own risk." }<h2>結果のクエリと分析</h2><p>最後に、Amazon Nova Lite モデルがレビューをどのように分析および分類するかを確認するためにクエリを実行します。GET products/_search を実行すると、レビュー コンテンツから生成されたフィールドがすでに強化されたドキュメントが取得されます。</p><p>モデルは、主な感情（肯定的、中立的、否定的）を識別し、簡潔な要約を生成し、関連するキーワードを抽出し、各レビューの信憑性を推定します。これらのフィールドは、全文を読まなくても顧客の意見を理解するのに役立ちます。</p><p>結果を解釈するには、次の点を考慮します。</p><ul><li><p>感情。これは、消費者の製品に対する全体的な認識を示します。</p></li><li><p>言及された主なポイントを強調した要約。</p></li><li><p>キーワード。類似のレビューをグループ化したり、フィードバック パターンを識別したりするために使用できます。</p></li><li><p>信憑性。レビューが信頼できるかどうかを示します。これはキュレーションやモデレーションに役立ちます。</p></li></ul>   "hits": [
      {
        "_index": "products",
        "_id": "1",
        "_score": 1,
        "_ignored": [
          "review.keyword"
        ],
        "_source": {
          "product": "Pampers Pants Premium Care Fralda",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The reviewer praises the diaper for its great material, high cotton content, and leak-proof design, especially highlighting its effectiveness for their baby.",
            "sentiment": "positive",
            "keywords": [
              "best diaper",
              "great material",
              "cotton",
              "no plastic",
              "leak-proof",
              "baby",
              "effective"
            ],
            "authenticity": "authentic"
          },
          "review": "Best diaper ever! Great material, lots of cotton, without all that plastic. Doesn't leak! My baby is a boy and every diaper leaked around the waist, this model solved the problem. Even on a small baby it's worth the effort of putting on the short diaper. I put it on my baby at 9 pm and only take it off in the morning, without any leaks."
        }
      },
      {
        "_index": "products",
        "_id": "2",
        "_score": 1,
        "_source": {
          "product": "Portable Electric Body Massager",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The product broke in three months for no apparent reason and the reviewer does not recommend it due to its short lifespan.",
            "sentiment": "negative",
            "keywords": [
              "broke",
              "short lifespan",
              "not recommend"
            ],
            "authenticity": "authentic"
          },
          "review": "It broke in three months for no apparent reason, thank goodness I didn't review it before. I don't recommend buying it because it has a short lifespan."
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 1,
        "_ignored": [
          "review.keyword"
        ],
        "_source": {
          "product": "Havit Fuxi-H3 Black Quad-Mode Wired and Wireless Gaming Headset",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The headset has good sound quality for the price but suffers from poor connectivity, especially when using the microphone or moving the headset. It also has compatibility issues with Linux.",
            "sentiment": "negative",
            "keywords": [
              "sound",
              "connectivity",
              "microphone",
              "compatibility",
              "annoying",
              "turn off and on",
              "Linux",
              "flexible stem",
              "work from home"
            ],
            "authenticity": "authentic"
          },
          "review": "The sound is good for the price, but the connectivity is horrible. You always need to be playing audio, otherwise it loses connection (I work from home, and this is very annoying). Sometimes it loses connection and you have to turn it off and on again to get it back on. The microphone is very sensitive, so it loses connection frequently and you have to turn the headset off and on for the microphone to work again. The flexibility of the stem is useless, because if you move it, the microphone can turn off. Sometimes I need to use Linux and the headset simply doesn't work. It's light and comfortable, the sound is adequate, but the connectivity is terrible."
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 1,
        "_ignored": [
          "review.keyword"
        ],
        "_source": {
          "product": "Air Fryer 4L Oil Free Fryer Mondial",
          "model_id": "bedrock_completion_amazon_nova-lite",
          "review_analyze": {
            "summary": "The product offers value for money but has issues with peeling, rusting, and uneven frying.",
            "sentiment": "negative",
            "keywords": [
              "value for money",
              "peeling",
              "rusting",
              "uneven frying",
              "weaker in the middle"
            ],
            "authenticity": "authentic"
          },
          "review": "For those looking for value for money, it's a good option, but the tray (which is underneath the perforated basket) is already peeling a lot. My mother has one just like it and said that hers is even rusting, in other words, the material is MUCH inferior. There's also something that bothers me, because it looks like a microwave, it doesn't fry evenly, it's weaker in the middle and stronger on the sides. Buy at your own risk."
        }
      }
    ]<h2>結びに</h2><p>Amazon Nova Lite と Elasticsearch の統合により、言語モデルが生のレビューを構造化された価値ある情報に変換する方法が実証されました。パイプラインを通じてレビューを処理することで、感情、信憑性、要約、キーワードを自動的かつ一貫して抽出できるようになりました。</p><p>結果は、モデルがレビューのコンテキストを理解し、ユーザーの意見を分類し、各体験の最も関連性の高いポイントを強調できることを示しています。これにより、検索機能の向上に活用できる、より豊富なデータセットが作成されます。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/amazon-nova-models-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/amazon-nova-models-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdbbf13eb690294f1/6a17fddd6df73195190a115a/304713c48b568e17d0bb56b19edb28769f7801b3-721x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 02 Apr 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[モデルコンテキストプロトコルを使用してエージェントをElasticsearchに接続する]]></title>
    <description><![CDATA[モデルコンテキストプロトコル サーバーを使用して、Elasticsearch 内のデータとチャットしてみましょう。]]></description>
    <content:encoded><![CDATA[<p>同僚とチャットするのと同じくらい簡単にデータとやりとりできたらどうでしょうか?「先月の 500 ドルを超えるすべての注文を表示してください」や「5 つ星のレビューを最も多く獲得した製品はどれですか」と尋ねるだけで、クエリを必要とせずに、即座に正確な回答が得られることを想像してみてください。</p><p>モデルコンテキストプロトコル (MCP) によりこれが可能になります。会話型 AI をデータベースや外部 API にシームレスに接続し、複雑なリクエストを自然な会話に変換します。現代の LLM は言語の理解に優れていますが、その真の可能性は現実世界のシステムと統合されたときに発揮されます。MCP は両者の間のギャップを埋め、データのやりとりをより直感的かつ効率的にします。</p><p>この投稿では、次の内容について説明します。</p><ul><li><p>MCPアーキテクチャ – 内部の仕組み</p></li><li><p>Elasticsearchに接続されたMCPサーバーの利点</p></li><li><p><a href="https://github.com/elastic/mcp-server-elasticsearch">Elasticsearch を活用した MCP サーバーの</a>構築</p></li></ul><p>これからはエキサイティングな時代が来ます!MCP と Elastic スタックの統合により、情報の操作方法が変わり、複雑なクエリが日常会話と同じくらい直感的になります。</p><h2>モデルコンテキストプロトコル</h2><p>Anthropic が開発した<a href="https://modelcontextprotocol.io/introduction">モデル コンテキスト プロトコル</a>(MCP) は、安全な双方向チャネルを通じて AI モデルを外部データ ソースに接続するオープン スタンダードです。これは、会話のコンテキストを維持しながら外部システムにリアルタイムでアクセスするという、AI の主要な制限を解決します。</p><h3>MCPアーキテクチャ</h3><p>モデルコンテキストプロトコルアーキテクチャは、次の 2 つの主要コンポーネントで構成されます。</p><ul><li><p><strong>MCP クライアント</strong>– ユーザーに代わって情報を要求したりタスクを実行したりする AI アシスタントとチャットボット。</p></li><li><p><strong>MCP サーバー</strong>– 関連情報を取得したり、要求されたアクション (外部 API の呼び出しなど) を実行したりするデータ リポジトリ、検索エンジン、および API。</p></li></ul><p>MCPサーバーは、クライアントに対して主に4つの機能を提供します。</p><ul><li><p><strong>リソース</strong>- LLM インタラクションのコンテキストとして取得および使用できる構造化されたデータ、ドキュメント、およびコンテンツ。これにより、AI アシスタントはデータベース、検索インデックス、その他のソースから関連情報にアクセスできるようになります。</p></li><li><p><strong>ツール</strong>- LLM が外部システムと対話したり、計算を実行したり、実際のアクションを実行できるようにする実行可能関数。これらのツールは、テキスト生成を超えて AI 機能を拡張し、アシスタントがワークフローをトリガーしたり、API を呼び出したり、データを動的に操作したりできるようにします。</p></li><li><p><strong>プロンプト</strong>- 一般的な LLM インタラクションを標準化して共有するための再利用可能なプロンプト テンプレートとワークフロー。</p></li><li><p><strong>サンプリング</strong>- セキュリティとプライバシーを維持しながら、高度なエージェント動作を可能にするために、クライアントを通じて LLM 完了を要求します。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfe82754551bb187a/6a17f7ec6864a43e71b6895d/bef5178133391e96e3d66ae634e41a85712a33a9-2345x1620.png" alt="モデルコンテキストプロトコル（MCP）アーキテクチャ" /><h2>MCPサーバー + Elasticsearch</h2><p></p><p>従来の検索拡張生成 (RAG) システムはユーザーのクエリに基づいてドキュメントを検索しますが、MCP はさらに一歩進んで、AI エージェントがリアルタイムでタスクを動的に構築して実行できるようにします。これにより、ユーザーは次のような自然言語の質問をすることができます。</p><p></p><ul><li><p>「先月の 500 ドルを超える注文をすべて表示してください。」</p></li><li><p>「5 つ星のレビューを最も多く獲得した製品はどれですか?」</p></li></ul><p></p><p>クエリを 1 つも書かなくても、即座に正確な回答が得られます。</p><p></p><p>MCP は以下を通じてこれを実現します。</p><ul><li><p>動的なツール選択 - エージェントは、ユーザーの意図に基づいて、MCP サーバー経由で公開される適切なツールをインテリジェントに選択します。一般的に、「よりスマートな」LLM は、状況に応じて適切な議論に基づいて適切なツールを選択するのが得意です。</p></li><li><p>双方向通信 – エージェントとデータソースは情報をスムーズに交換し、必要に応じてクエリを絞り込みます（例：最初にインデックス マッピングを検索し、その後で ES クエリを構築します。</p></li><li><p>マルチツール オーケストレーション - ワークフローは複数の MCP サーバーのツールを同時に活用できます。</p></li><li><p>永続的なコンテキスト - エージェントは以前のやり取りを記憶し、会話全体の継続性を維持します。</p></li></ul><p>Elasticsearch に接続された MCP サーバーは、強力なリアルタイム検索アーキテクチャを実現します。AI エージェントは、オンデマンドで Elasticsearch データを探索、クエリ、分析できます。シンプルなチャット インターフェースを通じてデータを検索できます。</p><p>MCP は、データの取得だけでなく、アクションも可能にします。他のツールと統合してワークフローをトリガーし、プロセスを自動化し、分析システムに洞察を提供します。MCP は、検索と実行を分離することで、AI を活用したアプリケーションの柔軟性と最新性を維持し、エージェント ワークフローにシームレスに統合します。</p><h2>ハンズオン: Elasticsearch データとチャットできる MCP サーバー</h2><p>MCP サーバー経由で Elasticsearch と対話するには、少なくとも次の機能が必要です。</p><ul><li><p>インデックスを取得する</p></li><li><p>マッピングを取得する</p></li><li><p>Elasticsearch のクエリ DSL を使用して検索を実行する</p></li></ul><p>私たちのサーバーは TypeScript で記述されており、公式の<a href="https://github.com/modelcontextprotocol/typescript-sdk">MCP TypeScript SDK</a>を使用します。セットアップには、MCP クライアントが組み込まれているため、Claude デスクトップ アプリ (無料版で十分です) をインストールすることをお勧めします。当社の MCP サーバーは基本的に、MCP ツールを通じて公式の<a href="https://www.elastic.co/jp/guide/en/elasticsearch/client/javascript-api/current/index.html">JavaScript Elasticsearch クライアント</a>を公開します。</p><p>まず、Elasticsearch クライアントと MCP サーバーを定義します。</p> const esClient = new Client({
    node: url,
    auth: {
      apiKey: apiKey,
    },
  });

  const server = new McpServer({
    name: "elasticsearch-mcp-server",
    version: "0.1.0",
  });<p>Elasticsearch と対話できる次の MCP サーバー ツールを使用します。</p><ul><li><p><strong>インデックスの一覧表示</strong>( <a href="https://github.com/elastic/mcp-server-elasticsearch/blob/main/index.ts#L46">list_indices</a> ): このツールは、利用可能なすべての Elasticsearch インデックスを取得し、インデックス名、ヘルス ステータス、ドキュメント数などの詳細を提供します。</p></li><li><p><strong>マッピングの取得</strong>( <a href="https://github.com/elastic/mcp-server-elasticsearch/blob/main/index.ts#L94">get_mappings</a> ): このツールは、指定された Elasticsearch インデックスのフィールド マッピングを取得し、ユーザーが保存されているドキュメントの構造とデータ型を理解するのに役立ちます。</p></li><li><p><strong>検索</strong>( <a href="https://github.com/elastic/mcp-server-elasticsearch/blob/main/index.ts#L147">search</a> ): このツールは、提供されたクエリ DSL を使用して Elasticsearch 検索を実行します。テキスト フィールドのハイライトが自動的に有効になり、関連する検索結果を簡単に識別できるようになります。</p></li></ul><p>完全な Elasticsearch MCP サーバーの実装は<a href="https://github.com/elastic/mcp-server-elasticsearch">、elastic/mcp-server-elasticsearch</a>リポジトリで入手できます。</p><h4>インデックスとチャット</h4><p>Elasticsearch MCP サーバーを設定して、「先月からの 500 ドルを超えるすべての注文を検索する」など、データに関する自然言語の質問をする方法を見てみましょう。</p><p><strong>Claudeデスクトップアプリを構成する</strong></p><ul><li><p>Claudeデスクトップアプリを開く</p></li><li><p>設定 &gt; 開発者 &gt; MCP サーバーに移動します</p></li><li><p>「設定の編集」をクリックして、この設定を<code>claude_desktop_config.json</code>に追加します。</p></li></ul>{
  "mcpServers": {
    "Elasticsearch MCP Server": {
      "command": "npx",
      "args": [
        "-y",
        "@elastic/mcp-server-elasticsearch"
      ],
      "env": {
        "ES_URL": "",
        "ES_API_KEY": ""
      }
    }
  }
}<p>注: このセットアップでは、Elastic が公開した<a href="https://www.npmjs.com/package/@elastic/mcp-server-elasticsearch">@elastic/mcp-server-elasticsearch</a> npm パッケージを利用します。ローカルで開発する場合は、Elasticsearch MCP サーバーの起動に関する詳細を<a href="https://github.com/elastic/mcp-server-elasticsearch/blob/main/README.md">こちらで</a>確認できます。</p><p><strong>Elasticseachインデックスを作成する</strong></p><ul><li><p>このデモの「注文」インデックスを作成するために、<a href="https://gist.github.com/jedrazb/60e9400cbe40addfd9e4337749c28431">サンプルデータ</a>を使用できます。</p></li><li><p>これにより、「先月からの500ドル以上のすべての注文を検索する」などのクエリを試すことができます。</p></li></ul><p><strong>使い始める</strong></p><ul><li><p>Claudeデスクトップアプリで新しい会話を開く</p></li><li><p>MCPサーバーは自動的に接続します</p></li><li><p>Elasticsearch データについて質問してみましょう。</p></li></ul><p>自然言語を使用して Elasticsearch データをクエリすることがいかに簡単かを確認するには、このデモをご覧ください。</p><h4>これは次のように機能します。</h4><p>「先月からの 500 ドルを超えるすべての注文を検索する」と要求されると、LLM は指定された制約を使用して Elasticsearch インデックスを検索する意図を認識します。効果的な検索を実行するために、エージェントは次のことを行います。</p><ul><li><p>インデックス名を確認します。 <code>orders</code></p></li><li><p><code>orders</code>インデックスのマッピングを理解する</p></li><li><p>インデックスマッピングと互換性のあるクエリDSLを構築し、最後に検索リクエストを実行します。</p></li></ul><p>この相互作用は次のように表すことができます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt152f41bc8c3e9752/6a17f7ee6df73152df0a10cc/8875bc75745124be87deac0be666509446887de2-2345x1620.png" alt="MCPサーバー+Elasticsearchはどのように機能するのか" /><h2>まとめ</h2><p>モデル コンテキスト プロトコルは、Elasticsearch データとのやり取りを強化し、複雑なクエリの代わりに自然言語による会話を可能にします。MCP は、AI 機能とデータを連携させることで、やり取り全体を通じてコンテキストを維持する、より直感的で効率的なワークフローを作成します。</p><p>Elasticsearch MCP サーバーはパブリック npm パッケージ ( <a href="https://www.npmjs.com/package/@elastic/mcp-server-elasticsearch">@elastic/mcp-server-elasticsearch</a> ) として利用できるため、開発者にとって統合が簡単になります。最小限のセットアップで、チームはデータの探索、ワークフローのトリガー、簡単な会話による分析情報の取得を開始できます。</p><p>自分で体験する準備はできましたか?今すぐ<a href="https://github.com/elastic/mcp-server-elasticsearch">Elasticsearch MCP サーバー</a>を試して、データとのチャットを始めましょう。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/model-context-protocol-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/model-context-protocol-elasticsearch</guid>
    <category><![CDATA[エージェント型AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jedr Blaszyk,Joe McElroy]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltce68a95c633809ae/6a17f7f0148009fa28b48915/65b378f644bd13e3edf2f108d48186f1889f546c-1200x628.png" length="0" type="image/png"/>
    <pubDate>Fri, 28 Mar 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearchを使ったマルチモーダルRAGシステムの構築：ゴッサム・シティの物語]]></title>
    <description><![CDATA[テキスト、オーディオ、ビデオ、画像データを統合して、より豊富でコンテキストに基づいた情報検索を提供するマルチモーダル検索拡張生成 (RAG) システムの構築方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>このブログでは、Elasticsearch を使用してマルチモーダル RAG (Retrieval-Augmented Generation) パイプラインを構築する方法を学びます。ImageBind を活用して、テキスト、画像、音声、深度マップなど、さまざまなデータ タイプの埋め込みを生成する方法について説明します。また、dense_vector と k-NN 検索を使用して、これらの埋め込みを Elasticsearch で効率的に保存および取得する方法についても説明します。最後に、大規模言語モデル (LLM) を統合して取得した証拠を分析し、包括的な最終レポートを生成します。</p><h3>マルチモーダル RAG パイプラインはどのように機能しますか?</h3><ol><li><p><strong>手がかりの収集</strong>→ ゴッサムの犯罪現場からの画像、音声、テキスト、深度マップ。</p></li><li><p><strong>埋め込みの生成</strong>→ 各ファイルは、ImageBind マルチモーダル モデルを使用してベクトルに変換されます。</p></li><li><p><strong>Elasticsearch でのインデックス作成</strong>→ ベクトルは効率的な検索のために保存されます。</p></li><li><p><strong>類似度による検索</strong>→ 新しい手がかりが与えられると、最も類似したベクトルが取得されます。</p></li><li><p><strong>LLM が証拠を分析</strong>→ GPT-4 モデルが応答を合成し、容疑者を特定します。</p></li></ol><h3>使用される技術</h3><ul><li><p><strong>ImageBind</strong> → さまざまなモダリティの統合された埋め込みを生成します。</p></li><li><p><strong>Elasticsearch</strong> → 高速かつ効率的なベクトル検索を可能にします。</p></li><li><p><strong>LLM (GPT-4、OpenAI)</strong> → 証拠を分析し、最終レポートを生成します。</p></li></ul><h3>このブログは誰に向けたものですか?</h3><ul><li><p>マルチモーダルベクトル検索に興味のある Elastic ユーザー。</p></li><li><p>マルチモーダル RAG を実際に理解したい開発者。</p></li><li><p>複数のソースからのデータを分析するためのスケーラブルなソリューションを探している方。</p></li></ul><h2>マルチモーダルRAGの前提条件: 環境の設定</h2><p>ゴッサム シティの犯罪を解決するには、テクノロジー環境を整える必要があります。次のステップバイステップガイドに従ってください。</p><h3>1. 技術要件</h3><p>成分</p><p>仕様</p><p>システムOS</p><p>Linux、macOS、またはWindows</p><p>Python</p><p>3.10以降</p><p>ラム</p><p>最低8GB（16GBを推奨）</p><p>グラフィックプロセッサ</p><p>オプションですが、ImageBind では推奨されます</p><h3><strong>2. プロジェクトの設定</strong></h3><p>すべての調査資料は GitHub で公開されており、このインタラクティブな犯罪解決体験には Jupyter Notebook (Google Colab) を使用します。開始するには、次の手順に従ってください。</p><h4>Jupyter Notebook の設定 (Google Colab)</h4><p><strong>1. ノートブックにアクセスする</strong></p><ul><li><p>すぐに使用できる Google Colab ノートブック「 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/building-multimodal-rag-with-elasticsearch-gotham/notebook/01-mmrag-blog-quick-start.ipynb">Multimodal RAG with Elasticsearch」</a>を開きます<u>。</u></p></li><li><p>このノートブックには、従う必要があるすべてのコードと説明が含まれています。</p></li></ul><p><strong>2. リポジトリをクローンする</strong></p># Clone the repository with the multimodal RAG code
!git clone -b https://github.com/elastic/elasticsearch-labs.git

# Navigate to the project directory
cd elasticsearch-labs/supporting-blog-content/building-multimodal-rag-with-elasticsearch-gotham<p><strong>3. 依存関係をインストールする</strong></p> # Install PyTorch and related libraries
!pip install torch&gt;=2.1.0 torchvision&gt;=0.16.0 torchaudio&gt;=2.1.0

# Install vision processing libraries
!pip install opencv-python-headless pillow numpy

# Install the specific ImageBind fork
!pip install git+https://github.com/hkchengrex/ImageBind.git

# Install Elasticsearch and environment management
!pip install elasticsearch python-dotenv<p><strong>4. 資格情報を設定する</strong></p># Input your credentials securely
import getpass

ELASTICSEARCH_URL = input("Enter the Elasticsearch endpoint url: ")
ELASTICSEARCH_API_KEY = getpass.getpass("Enter the Elasticsearch API key: ")
OPENAI_API_KEY = getpass.getpass("Enter the OpenAI API key: ")

# Configure environment variables
import os
os.environ["ELASTICSEARCH_API_KEY"] = ELASTICSEARCH_API_KEY
os.environ["OPENAI_API_KEY"] = OPENAI_API_KEY
os.environ["ELASTICSEARCH_URL"] = ELASTICSEARCH_URL<p>注: ImageBind モデル (約 2 GB) は、最初の実行時に自動的にダウンロードされます。</p><p>準備はすべて整ったので、詳細を調べて犯罪を解決しましょう。</p><h2>はじめに：ゴッサム・シティの犯罪</h2><p>ゴッサム・シティの雨の夜、衝撃的な犯罪が街を揺るがす。ゴードン委員長は謎を解くためにあなたの助けを必要としています。手がかりは、ぼやけた画像、謎の音声、暗号化されたテキスト、さらには深度マップなど、さまざまな形式で散在しています。最先端の AI テクノロジーを駆使して事件を解決する準備はできていますか?</p><p>このブログでは、さまざまな種類のデータ (画像、音声、テキスト、深度マップ) を単一の検索空間に統合する<strong> マルチモーダル RAG (検索拡張生成) システムの 構築を段階的に説明します。</strong><strong>ImageBind</strong>を使用してマルチモーダル埋め込みを生成し、 <strong>Elasticsearch を</strong>使用してこれらの埋め込みを保存および取得し、<strong>大規模言語モデル (LLM) を使用して</strong>証拠を分析し、最終レポートを生成します。</p><h2>基礎：マルチモーダルRAGアーキテクチャ</h2><h3>マルチモーダル RAG とは何ですか?</h3><p><strong>検索拡張生成 (RAG) マルチモーダル</strong>の台頭により、AI モデルとの対話方法に革命が起きています。従来、RAG システムはテキストのみを処理し、応答を生成する前にデータベースから関連情報を取得します。しかし、世界はテキストに限定されません<strong>。画像、ビデオ、音声にも貴重な知識が存在します</strong>。このため、マルチモーダル アーキテクチャが注目を集めており、AI システムが<strong>さまざまな形式の情報を組み合わせて、より豊富で正確な応答を実現できるように</strong>なりました。</p><h3><strong>マルチモーダルRAGの3つの主なアプローチ</strong></h3><p>マルチモーダル RAG を実装するには、一般的に 3 つの戦略が使用されます。それぞれのアプローチには、ユースケースに応じて独自の利点と制限があります。</p><h4>1. 共有ベクトル空間</h4><p>さまざまなモダリティからのデータは、ImageBind などのマルチモーダル モデルを使用して共通のベクトル空間にマッピングされます。これにより、明示的な形式変換を行わずに、テキスト クエリで画像、ビデオ、オーディオを取得できるようになります。</p><p><strong>利点:</strong></p><ul><li><p>明示的な形式変換を必要とせずに<strong>クロスモーダル検索</strong>を可能にします。</p></li><li><p>さまざまなモダリティ間の<strong>スムーズな統合</strong>を提供し、テキスト、画像、オーディオ、ビデオを直接取得できます。</p></li><li><p>さまざまなデータ タイプに拡張可能なので、<strong>大規模な検索アプリケーション</strong>に役立ちます。</p></li></ul><p><strong>デメリット:</strong></p><ul><li><p><strong>トレーニングには大規模なマルチモーダル データセットが必要ですが</strong>、必ずしも利用できるとは限りません。</p></li><li><p>共有された埋め込み空間では<strong>意味ドリフト</strong>が生じる可能性があり、その場合、モダリティ間の関係は完全には保持されません。</p></li><li><p><strong>マルチモーダル モデルのバイアスは</strong>、データセットの分布に応じて、検索精度に影響を及ぼす可能性があります。</p></li></ul><h4>2. 単一のグラウンデッドモダリティ</h4><p>すべてのモダリティは、検索前に<strong>単一の形式</strong>（通常は<strong>テキスト</strong>）に変換されます。たとえば、画像は<strong>自動的に生成されたキャプション</strong>を通じて説明され、音声はテキストに転記されます。</p><p><strong>利点:</strong></p><ul><li><p>すべてが<strong> 統一されたテキスト表現</strong><strong> に変換されるため、 検索が簡単になります</strong> 。</p></li><li><p><strong>既存のテキストベースの検索エンジン</strong>と連携して動作し、特殊なマルチモーダル インフラストラクチャの必要性を排除します。</p></li><li><p>取得された結果は人間が読める形式であるため、<strong>解釈</strong>可能性が向上します。</p></li></ul><p><strong>デメリット:</strong></p><ul><li><p><strong>情報の損失</strong>: 特定の詳細 (画像内の空間関係、音声のトーンなど) は、テキストの説明では完全には表現されない場合があります。</p></li><li><p><strong>キャプション/文字起こしの品質に依存</strong>: 自動注釈のエラーにより、検索の有効性が低下する可能性があります。</p></li><li><p>変換プロセスによって重要なコンテキストが削除される可能性があるため<strong>、純粋に視覚的または聴覚的なクエリには最適ではありません</strong>。</p></li></ul><h4>3. 個別検索</h4><p>各モダリティごとに<strong>個別のモデル</strong>を維持します。システムは各データ タイプごとに<strong>個別の検索</strong>を実行し、後で<strong>結果を結合します</strong>。</p><p><strong>利点:</strong></p><ul><li><p><strong>モダリティごとにカスタム最適化</strong>が可能になり、各データのタイプごとの検索精度が向上します。</p></li><li><p><strong>複雑なマルチモーダル モデル</strong>への依存が少なくなり、既存の検索システムとの統合が容易になります。</p></li><li><p>さまざまなモダリティからの結果を動的に組み合わせることができるため<strong>、ランキングと再ランキングをきめ細かく制御</strong>できます。</p></li></ul><p><strong>デメリット:</strong></p><ul><li><p><strong>結果の融合が必要なため</strong>、検索とランキングのプロセスがより複雑になります。</p></li><li><p>異なるモダリティが矛盾する情報を返す場合、<strong>一貫性のない応答</strong>が生成されることがあります。</p></li><li><p>各モダリティごとに独立した検索が実行され、処理時間が長くなるため、<strong>計算コストが高くなります</strong>。</p></li></ul><h3>私たちの選択: ImageBindによる共有ベクトル空間</h3><p>これらのアプローチの中で、私たちは<strong>共有ベクトル空間</strong>を選択しました。これは、<strong>効率的なマルチモーダル検索</strong>のニーズに完全に一致する戦略です。私たちの実装は、<strong> 共通のベクトル空間</strong> で複数のモダリティ（<strong> テキスト、画像、音声、ビデオ</strong> ）を表現できるモデルである<strong> ImageBind</strong> に基づいています。これにより、次のことが可能になります。</p><ul><li><p>すべてをテキストに変換する必要なく、さまざまなメディア形式間で<strong>クロスモーダル検索を</strong>実行します。</p></li><li><p><strong>非常に表現力豊かな埋め込み</strong>を使用して、さまざまなモダリティ間の関係を捉えます。</p></li><li><p><strong>スケーラビリティと効率性</strong>を確保し、Elasticsearch で高速に検索できるように最適化された埋め込みを保存します。</p></li></ul><p>このアプローチを採用することで、追加の前処理なしでテキストクエリで<strong> 画像や音声を直接取得</strong> できる<strong> 堅牢なマルチモーダル検索パイプライン</strong> を構築しました。この方法は<strong>、大規模リポジトリでのインテリジェントな検索</strong>から<strong>高度なマルチモーダル推奨システム</strong>まで、実用的なアプリケーションを拡張します。</p><p>次の図は、マルチモーダル RAG パイプライン内のデータ フローを示しており、マルチモーダル データに基づくインデックス作成、検索、および応答生成のプロセスが強調表示されています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt77bce4aa5216bcf3/6a17eef263173069e0585b57/a4ffdb44582738991813c045be37312dacb0d4f3-1488x1436.png" alt="マルチモーダルラグデータフロー" /><h3>埋め込みスペースはどのように機能しますか?</h3><p>従来、テキスト埋め込みは言語モデル (BERT、GPT など) から得られます。現在、Meta AI の<strong>ImageBind</strong>のようなネイティブ マルチモーダル モデルにより、複数のモダリティのベクトルを生成するバックボーンが得られます。</p><ul><li><p><strong>テキスト</strong>: 文と段落は同じ次元のベクトルに変換されます。</p></li><li><p><strong>画像 (視覚)</strong> : ピクセルは、テキストに使用されるのと同じ次元空間にマッピングされます。</p></li><li><p><strong>オーディオ</strong>: サウンド信号は、画像やテキストに匹敵する埋め込みに変換されます。</p></li><li><p><strong>深度マップ</strong>: 深度データが処理され、ベクターが生成されます。</p></li></ul><p>したがって、<strong> コサイン類似度</strong><strong> などのベクトル類似度メトリックを使用して、任意の手がかり (</strong> テキスト、画像、音声、深度) を他の手がかりと比較できます。<strong>笑い声の音声サンプル</strong>と<strong>容疑者の顔の画像が</strong>この空間で「近い」場合、何らかの相関関係（同じ身元など）を推測できます。</p><h2>ステージ1 - 犯罪現場の手がかりを集める</h2><p>証拠を分析する前に、それを収集する必要があります。ゴッサムでの犯罪は、画像、音声、テキスト、さらには深度データに隠されている可能性のある痕跡を残しました。これらの手がかりを整理してシステムに取り入れてみましょう。</p><h3>何があるでしょうか?</h3><p>ゴードン委員は、犯罪現場から4つの異なる方法で収集された証拠を含む以下のファイルを私たちに送信しました。</p><p><strong>トラックの説明とモダリティ</strong></p><p><strong>a) 画像（写真2枚）</strong></p><ul><li><p><code>crime_scene1.jpg, crime_scene2.jpg</code> → 犯罪現場から撮影された写真。地面に怪しい痕跡が残っている。</p></li><li><p><code>suspect_spotted.jpg</code> →現場から逃走するシルエットが映った防犯カメラの映像。</p></li></ul><p><strong>b)</strong><strong>音声（録音1件）</strong></p><ul><li><p><code>joker_laugh.wav </code>→ 犯罪現場近くのマイクが不気味な笑い声を捉えた。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte7690380dc545367/6a17eef8be6086886a00483b/3457bab38aa4a3caf61ca2a5e0a8b477ddd15cb1-86x45.png" alt="" /><p><strong>c) テキスト（1件）</strong></p><ul><li><p><code>Riddle.txt, note2.txt</code> → その場所で、犯人が残したと思われる謎のメモが見つかりました。</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8e5b8f4bc4124708/6a17eefa1d1b83850c93e50c/8963228dc3b4b7e2106cb6ddffbe2154020eb735-77x79.png" alt="" /><p><strong>d) 深度（深度マップ1枚）</strong></p><ul><li><p><code>depth_suspect.png</code> → 深度センサー付きの防犯カメラが近くの路地に容疑者を捉えた。</p></li><li><p><code>jdancing-depth.png</code> → 深度センサー付きの防犯カメラが、地下鉄の駅に降りていく容疑者を捉えた。</p></li></ul><p>これらの証拠は形式が異なり、同じ方法で直接分析することはできません。これらを埋め込み、つまりモーダル間の比較を可能にする数値ベクトルに変換する必要があります。</p><h3><strong>ファイル構成</strong></h3><p>処理を開始する前に、パイプラインがスムーズに実行されるように、すべての手がかりが data/ ディレクトリ内で適切に整理されていることを確認する必要があります。</p><p><strong>予想されるディレクトリ構造:</strong></p>data/
├── images/
│   ├── crime_scene1.jpg
│   ├── suspect_spotted.jpg
│   ...
├── audios/
│   ├── joker_laugh.wav
│   ...
├── texts/
│   ├── riddle.txt
│   ... 
├── depths/
│   ├── depth_suspect.png<h3>手がかりの構成を確認するためのコード</h3><p>続行する前に、必要なファイルがすべて正しい場所にあることを確認しましょう。</p>import os

# Base directory for clues
data_dir = "data"

# List of expected files
evidences = {
    "images": ["crime_scene1.jpg","crime_scene1.jpg", "joker_alley.jpg"],
    "audios": ["joker_laugh.wav"],
    "texts": ["riddle.txt", "note2.txt”],
    "depths": ["depth_suspect.png", "jdancing-depth.png"]
}

# Create directories if they don't exist
for category, files in evidences.items():
    category_path = os.path.join(data_dir, category)
    os.makedirs(category_path, exist_ok=True)

    for file in files:
        file_path = os.path.join(category_path, file)
        if not os.path.exists(file_path):
            print(f"Warning: {file} not found in {category_path}.")

print("All files are correctly organized!")<p><strong>ファイルを実行する</strong></p>python  stages/01-stage/files_check.py<p><strong>予想される出力（すべてのファイルが正しい場合）:</strong></p>All files are correctly organized!<p><strong>予想される出力（ファイルが欠落している場合）:</strong></p>Warning: joker_laugh.wav not found in data/audios/
Warning: depth_suspect.png not found in data/depths/<p>このスクリプトは、埋め込みを生成して Elasticsearch にインデックス付けする前にエラーを防ぐのに役立ちます。</p><h2>ステージ2 - 証拠の整理</h2><h3>ImageBindによる埋め込みの生成</h3><p>手がかりを統合するには、手がかりを埋め込み、つまり各様相の意味を捉えるベクトル表現に変換する必要があります。ここでは、共有ベクトル空間内でさまざまなデータ タイプ (<strong> 画像、音声、テキスト、深度マップ)</strong> の埋め込みを生成する Meta AI<strong> のモデルである ImageBind</strong> を 使用します。
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltff4aefffbfb5bf00/6a17eefe6864a45aecb68860/b19a1c32cc7a0b4c00fa5b247b18cce71f9693cb-1580x918.png" alt="ImageBindによる埋め込みの生成" /><h3><strong>ImageBind はどのように機能しますか?</strong></h3><p>さまざまな種類の証拠 (<strong>画像、音声、テキスト、深度マップ</strong>) を比較するには、 <strong>ImageBind</strong>を使用してそれらを数値ベクトルに変換する必要があります。このモデルにより、あらゆるタイプの入力を同じ埋め込み形式に変換できるため、モダリティ間の<strong>クロスモーダル検索</strong>が可能になります。</p><p>以下は、各モダリティに適したプロセッサを使用して、あらゆるタイプの入力の埋め込みを生成するための最適化されたコード ( <code>src/embedding_generator.py</code> ) です。</p>class EmbeddingGenerator:
    """Class for generating multimodal embeddings using ImageBind."""
    
    def __init__(self):
        self.device = "cuda" if torch.cuda.is_available() else "cpu"
        self.model = self._load_model()

    def _load_model(self):
        """Loads the ImageBind model and sets it to inference mode."""
        model = imagebind_model.imagebind_huge(pretrained=True)
        model.eval()
        model.to(self.device)
        return model

    def generate_embedding(self, input_data, modality):
        """Generates embedding for different modalities"""
        processors = {
            "vision": lambda x: data.load_and_transform_vision_data(x, self.device),
            "audio": lambda x: data.load_and_transform_audio_data(x, self.device),
            "text": lambda x: data.load_and_transform_text(x, self.device),
            "depth": self.process_depth
        }
        
        try:
            # Input type verification
            if not isinstance(input_data, list):
                raise ValueError(f"Input data must be a list. Received: {type(input_data)}")
                
            # Convert input data to a tensor format that the model can process
            # For images: [batch_size, channels, height, width] 
            # For audio: [batch_size, channels, time] 
            # For text: [batch_size, sequence_length]
            inputs = {modality: processors[modality](input_data)}
            with torch.no_grad():
                embedding = self.model(inputs)[modality]
            return embedding.squeeze(0).cpu().numpy()
        except Exception as e:
            logger.error(f"Error generating {modality} embedding: {str(e)}", exc_info=True)
            raise<p>テンソルは、特に ImageBind のようなモデルを扱う場合の、機械学習とディープラーニングにおける基本的なデータ構造です。私たちの文脈では：</p>input_tensor = processors[modality]([input_data], self.device)<p>ここで、テンソルは、モデルが処理できる数学的形式に変換された入力データ (画像、音声、またはテキスト) を表します。具体的には：</p><ul><li><p><strong>画像の場合</strong>: テンソルは、画像を数値 (高さ、幅、およびカラー チャネル別に整理されたピクセル) の多次元行列として表現します。</p></li><li><p><strong>オーディオの場合</strong>: テンソルは、音波を時間の経過に伴う一連の振幅として表します。</p></li><li><p><strong>テキストの場合</strong>: テンソルは単語またはトークンを数値ベクトルとして表します。</p></li></ul><h3>埋め込み生成のテスト:</h3><p>次のコードを使用して埋め込み生成をテストしてみましょう。これを 02-stage/test_embedding_generation.py に保存し、次のコマンドで実行します。</p>python stages/02-stage/test_embedding_generation.py generator = EmbeddingGenerator()
image_embedding = generator.generate_embedding("data/images/crime_scene1.jpg","vision")

print(image_embedding.shape)<h3>期待される出力:</h3>(1024,)<p>これで、画像は<strong>1024 次元のベクトル</strong>に変換されました。</p><h2>ステージ3 - Elasticsearchでの保存と検索</h2><p>証拠の埋め込みを生成したので、効率的な検索を可能にするために、それらをベクトル データベースに保存する必要があります。このため、密なベクトル ( <code>dense_vector</code> ) をサポートし、類似性検索を可能にする<strong>Elasticsearch</strong>を使用します。</p><p>このステップは、主に次の 2 つのプロセスで構成されます。</p><ul><li><p><strong>埋め込みのインデックス作成</strong>→ 生成されたベクトルを Elasticsearch に保存します。</p></li><li><p><strong>類似検索</strong>→ 新しい証拠に最も類似したレコードを取得します。</p></li></ul><h3>Elasticsearchで証拠をインデックスする</h3><p><strong>ImageBind</strong>によって処理される各証拠 (画像、音声、テキスト、深度) は<strong>、1024 次元のベクトル</strong>に変換されます。将来の検索を可能にするために、これらのベクトルを<strong>Elasticsearch</strong>に保存する必要があります。</p><p>次のコード ( <code>src/elastic_manager.py</code> ) は、Elasticsearch に<strong>インデックス</strong>を作成し、埋め込みを格納するためのマッピングを構成します。</p>from elasticsearch import Elasticsearch, helpers
...

class ElasticsearchManager:
    """Manages multimodal operations in Elasticsearch"""
    
    def __init__(self):
        load_dotenv()  # Load variables from .env
        self.es = self._connect_elastic()
        self.index_name = "multimodal_content"
        self._setup_index()
    
    def _connect_elastic(self):
        """Connects to Elasticsearch"""
        return Elasticsearch(
            os.getenv("ELASTICSEARCH_URL"),  # Elasticsearch endpoint
            api_key=os.getenv("ELASTICSEARCH_API_KEY")
        )
    
    def _setup_index(self):
        """Sets up the index if it doesn't exist"""
        if not self.es.indices.exists(index=self.index_name):
            mapping = {
                "mappings": {
                    "properties": {
                        "embedding": {
                            "type": "dense_vector",
                            "dims": 1024,
                            "index": True,
                            "similarity": "cosine"
                        },
                        "modality": {"type": "keyword"},
                        "content": {"type": "binary"},
                        "description": {"type": "text"},
                        "metadata": {"type": "object"},
                        "content_path": {"type": "text"}
                    }
                }
            }
            self.es.indices.create(index=self.index_name, body=mapping)
    
    def index_content(self, embedding, modality, content=None, description="", metadata=None, content_path=None):
        """Indexes multimodal content"""
        doc = {
            "embedding": embedding.tolist(),
            "modality": modality,
            "description": description,
            "metadata": metadata or {},
            "content_path": content_path
        }
        
        if content:
            doc["content"] = base64.b64encode(content).decode() if isinstance(content, bytes) else content
        
        return self.es.index(index=self.index_name, document=doc)
    
    def search_similar(self, query_embedding, modality=None, k=5):
        """Searches for similar contents"""
        query = {
            "knn": {
                "field": "embedding",
                "query_vector": query_embedding.tolist(),
                "k": k,
                "num_candidates": 100,
                "filter": [{"term": {"modality": modality}}] if modality else []
            }
        }
        
        try:
            response = self.es.search(
                index=self.index_name,
                query=query,
                size=k            
            )
            
            # Return both source data and score for each hit
            return [{
                **hit["_source"],
                "score": hit["_score"]
            } for hit in response["hits"]["hits"]]
        
        except Exception as e:
            print(f"Error: processing search_evidence: {str(e)}")
            return "Error generating search evidence"<h3>インデックス作成の実行</h3><p>それでは、プロセスをテストするために証拠をインデックスしてみましょう。</p># Example: Indexing an image from the crime scene
generator = EmbeddingGenerator()
es_manager = ElasticsearchManager(cloud_id="YOUR_CLOUD_ID", api_key="YOUR_API_KEY")

image_embedding = generator.generate_embedding("data/images/crime_scene1.jpg", "vision")

response = es_manager.index_content(
    embedding=image_embedding,
    modality="vision",
    description="Photo of the crime scene with suspicious traces",
    content_path="data/images/crime_scene1.jpg"
)
print(json.dumps(response, indent=2))<p><strong>Elasticsearch での予想される出力 (インデックスされたドキュメントの概要):</strong></p>{
    "embedding": [0.12, -0.53, 0.89, ...],  
    "modality": "vision",  
    "description": "Photo of the crime scene with suspicious traces",  
    "content_path": "data/images/crime_scene1.jpg"  
}<p>すべてのマルチモーダル証拠をインデックスするには、次の Python コマンドを実行してください。</p>python stages/03-stage/index_all_modalities.py<p>これで、証拠は<strong>Elasticsearch</strong>に保存され、必要なときに取得できるようになりました。</p><h3>インデックス作成プロセスの検証</h3><p>インデックス作成スクリプトを実行した後、すべての証拠が Elasticsearch に正しく保存されているかどうかを確認しましょう。<strong>Kibana の開発ツール</strong>を使用して、いくつかの検証クエリを実行できます。</p><p>1. まず、インデックスが作成されたかどうかを確認します。</p>GET _cat/indices/multimodal_content?v<p>2. 次に、モダリティごとのドキュメント数を確認します。</p>GET multimodal_content/_search
{
  "size": 0,
  "aggs": {
    "modalities": {
      "terms": {
        "field": "modality.keyword"
      }
    }
  }
}<p>3. 最後に、インデックスされたドキュメントの構造を調べます。</p>GET multimodal_content/_search
{
  "size": 1,
  "query": {
    "match_all": {}
  }
}<h4>期待される結果:</h4><ul><li><p>`multimodal_content` という名前のインデックスが存在する必要があります。</p></li><li><p>さまざまなモダリティ（視覚、音声、テキスト、深度）に分散された約 7 つのドキュメント。</p></li><li><p>各ドキュメントには、埋め込み、モダリティ、説明、メタデータ、content_path フィールドが含まれている必要があります。</p></li></ul><p>この検証手順により、類似性検索に進む前に証拠データベースが適切に設定されていることが保証されます。</p><h3>Elasticsearchで類似の証拠を検索する</h3><p>証拠がインデックス化されたので、検索を実行して新しい手がかりに最も類似した記録を見つけることができます。この検索では<strong>、ベクトル類似性</strong>を使用して、<strong>埋め込み空間</strong>内の最も近いレコードを返します。</p><p>次のコードはこの検索を実行します。</p>def search_similar_evidence(self, query_embedding, k=5, modality=None):
    """Performs a kNN search to find the most similar clues."""
    
    knn_query = {
        "field": "embedding",
        "query_vector": query_embedding.tolist(),
        "k": k,
        "num_candidates": 100
    }

    query_body = {"knn": knn_query}
    if modality:
        query_body = {
            "bool": {
                "must": [
                    query_body, 
                    {"term": {"modality": modality}}
                ]
            }
        }

    try:
      results = self.es.search(
        index=self.index_name,
        query=query_body,
        _source_includes=["description", "modality", "content_path"],
        size=k
      )
    except Exception as e:
            print(f"Error processing search_evidence: {str(e)}")
            return "Error generating search evidence”

    return results["hits"]["hits"]<h3>検索のテスト - マルチモーダル結果のクエリとして音声を使用する</h3><p>それでは、<strong>疑わしい音声ファイル</strong>を使用して証拠の検索をテストしてみましょう。同じ方法でファイルの埋め込みを生成し、同様の埋め込みを検索する必要があります。</p>python stages/03-stage/search_by_audio.py# Initialize classes
generator = EmbeddingGenerator()
es_manager = ElasticsearchManager(cloud_id="YOUR_CLOUD_ID", api_key="YOUR_API_KEY")

# Generate embedding for a suspicious audio
audio_embedding = generator.generate_embedding("data/audios/mysterious_laugh.wav", "audio")

# Search for similar evidence in Elasticsearch
similar_evidences = es_manager.search_similar_evidence(audio_embedding, k=3)

# Display the retrieved results
print("\n🔎 Similar evidence found:\n")
for i, evidence in enumerate(similar_evidences, start=1):
    description = evidence['_source']['description']
    modality = evidence['_source']['modality']
    score = evidence['_score']
    content_path = evidence['_source'].get('content_path', 'N/A')

    print(f"{i}. {description} ({modality})")
    print(f"   Similarity: {score:.4f}")
    print(f"   File path: {content_path}\n")<p><strong>ターミナルに期待される出力:</strong></p>🔎 Similar evidence found:

1. A sinister laugh captured near the crime scene (audio)
   Similarity: 0.9985
   File path: data/audios/joker_laugh.wav

2. The Joker with green hair, white face paint, and a sinister smile in an urban night setting. (vision)
   Similarity: 0.6068
   File path: data/images/joker_laughing.png

3. Suspect dancing (vision)
   Similarity: 0.5591
   File path: data/images/jdancing.png<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt73f8f56b0f2b8493/6a17ef01faa91381ca93c94e/24067ea40f7958e171149221f83cbd9bcfccc53f-1582x1208.png" alt="" /><p>これで、<strong>取得した証拠を分析し</strong>、事件との関連性を判断できます。</p><h3>オーディオを超えて - マルチモーダル検索の探求</h3><h4>役割の逆転：あらゆるモダリティは「質問」になり得る</h4><p>当社の<strong>マルチモーダル RAG</strong>システムでは、<strong>あらゆるモダリティ</strong>が潜在的な<strong>検索クエリ</strong>となります。音声の例を超えて、他のデータ タイプで<strong>調査を開始する</strong>方法を見てみましょう。</p><h4>1. テキストによる検索（犯人のメモの解読）</h4><p>シナリオ:<strong>暗号化されたテキスト メッセージ</strong>を発見し、関連する証拠を見つけたいと考えています。</p>python stages/03-stage/search_by_text.py# Generate embedding from text
text = "Why so serious?"
embedding_text = generator.generate_embedding([text], "text")

# Search for related evidence
similar_evidences = es_manager.search_similar(
    query_embedding=embedding_text,
    k=3
)<p><strong>期待される結果:</strong></p>🔎 Similar evidence found:

1. Mysterious note found at the location (text)
   Similarity: 0.7639
   File path: data/texts/riddle.txt

2. The Joker with green hair, white face paint, and a sinister smile in an urban night setting. (vision)
   Similarity: 0.7161
   File path: data/images/joker_laughing.png

3. Why so serious (text)
   Similarity: 0.7132
   File path: data/texts/note2.txt<h4>2. 画像検索（不審な犯罪現場の追跡）</h4><p><strong>シナリオ:</strong><strong>新しい犯罪現場</strong>( <code>crime_scene2.jpg</code> ) を他の証拠と比較する必要があります。
</p>python stages/03-stage/search_by_image.py# Generate embedding for a suspicious image
vision_embedding = generator.generate_embedding(["data/images/crime_scene2.jpg"], "vision")

# Search for similar evidence in Elasticsearch
similar_evidences = es_manager.search_similar(
    query_embedding=vision_embedding,
    k=3
)<p><strong>アウトプット：</strong></p>🔎 Similar evidence found:

1. Photo of the crime scene: A dark, rain-soaked alley is filled with playing cards, while a sinister graffiti of the Joker laughing stands out on the brick wall. (vision)
   Similarity: 0.8258
   File path: data/images/crime_scene1.jpg

2. The Joker with green hair, white face paint, and a sinister smile in an urban night setting. (vision)
   Similarity: 0.6897
   File path: data/images/joker_laughing.png

3. Suspect dancing (vision)
   Similarity: 0.6588
   File path: data/images/jdancing.png<h4>3. 深度マップ探索（3D追跡）</h4><p><strong>シナリオ:</strong><strong>深度マップ</strong>( <code>jdancing-depth.png</code> ) により、<strong>画像の</strong><strong>エスケープ パターンが</strong>明らかになります。</p>python stages/03-stage/search_by_depth.py# Generate embedding for a suspicious depth map
vision_embedding = generator.generate_embedding(["data/depths/jdancing-depth.png"], "depth")

# Search for similar evidence in Elasticsearch
similar_evidences = es_manager.search_similar(
    query_embedding=vision_embedding,
    modality="vision",
    k=3
)<p><strong>出力</strong></p>🔎 Similar evidence found:

1. The Joker with green hair, white face paint, and a sinister smile in an urban night setting. (vision)
   Similarity: 0.5329
   File path: data/images/joker_laughing.png<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt224eaf1be75a8b6b/6a17ef03e8fbceada23a1a0e/985f9536db95c772c696dfac996822fb00f199ee-1594x1160.png" alt="" /><p></p>2. Photo of the crime scene: A dark, rain-soaked alley is filled with playing cards, while a sinister graffiti of the Joker laughing stands out on the brick wall. (vision)
   Similarity: 0.5053
   File path: data/images/crime_scene1.jpg<p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d75bdf986e1bbdd/6a17eef32f4a5c25d5fa89b1/7023ee786ccc760689257abbde2f759ca3cf5c59-1024x768.jpg" alt="" />3. Suspect dancing (vision)
   Similarity: 0.4859
   File path: data/images/jdancing.png<p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt73f8f56b0f2b8493/6a17ef01faa91381ca93c94e/24067ea40f7958e171149221f83cbd9bcfccc53f-1582x1208.png" alt="" /><h3><strong>なぜこれが重要なのでしょうか?</strong></h3><p>それぞれのモダリティは<strong>独自の接続</strong>を明らかにします。</p><ul><li><p><strong>テキスト</strong>→ 容疑者の言語パターン。</p></li><li><p><strong>画像</strong>→<strong>場所や物体の認識。</strong></p></li><li><p><strong>深度</strong>→ 3D シーンの<strong>再構築。</strong></p></li></ul><p>現在、<strong> Elasticsearch</strong> <strong>には 構造化された証拠データベース</strong> があり、<strong> マルチモーダル証拠を効率的に保存および取得</strong> できます。</p><h3><strong>私たちが行ったことの要約:</strong></h3><ul><li><p>Elasticsearch に<strong>マルチモーダル埋め込みを保存しました</strong>。</p></li><li><p><strong>類似性検索を実行し</strong>、新たな手がかりに関連する証拠を見つけました。</p></li><li><p><strong>疑わしい音声ファイルを使用して検索をテストし</strong>、システムが正しく動作することを確認しました。</p></li></ul><p><strong>次のステップ:</strong> <strong>LLM</strong> (大規模言語モデル) を使用して、<strong>取得した証拠を分析し</strong>、<strong>最終レポート</strong>を生成します。</p><h2>ステージ4 - LLMで点と点をつなぐ</h2><p><strong>証拠が</strong><strong> Elasticsearch</strong> で インデックス化され 、類似性に基づいて検索できるようになったので、それを<strong> 分析し</strong> 、ゴードン委員に送信する <strong>最終レポート</strong><strong> を生成する LLM (大規模言語モデル) が必要になります。</strong><strong>LLM は</strong>、取得した証拠に基づいて<strong>パターンを識別し、手がかりを結び付け、容疑者を提案する</strong>責任を負います。</p><p>このタスクでは、<strong> GPT-4 Turbo</strong> を使用して、モデルが結果を効率的に <strong>解釈</strong> できるように<strong> 詳細なプロンプト を作成します。</strong></p><h3><strong>LLM統合</strong></h3><p><strong>LLM を システムに統合するために、 Elasticsearch</strong> から<strong> 取得した証拠</strong> を受け取り、この証拠をプロンプトコンテキストとして使用して<strong> フォレンジック</strong> レポート を生成する LLMAnalyzer クラス (<code>src/llm_analyzer.py</code> ) を作成しました。</p>import os
from openai import OpenAI
import logging
from dotenv import load_dotenv

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

class LLMAnalyzer:
    """Evidence analyzer using GPT-4"""
    
    def __init__(self):
        load_dotenv()
        self.client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
    
    def analyze_evidence(self, evidence_results):
        """
        Analyzes multimodal search results and generates a report
        
        Args:
            evidence_results: Dict with results by modality
            {
                'vision': [...],
                'audio': [...],
                'text': [...],
                'depth': [...]
            }
        """
        # Format evidence for the prompt
        evidence_summary = self._format_evidence(evidence_results)

        # final prompt
        prompt = f"""
You are a highly experienced forensic detective specializing in multimodal evidence analysis. Your task is to analyze the collected evidence (audio, images, text, depth maps) and conclusively determine the **prime suspect** responsible for the Gotham Central Bank case.

---

### **Collected Evidence:**
{evidence_summary}

### **Task:**
1. **Analyze all the evidence** and identify cross-modal connections.
2. **Determine the exact identity of the criminal** based on behavioral patterns, visual/auditory/textual clues, and symbolic markers.
3. **Justify your conclusion** by explaining why this suspect is definitively responsible.
4. **Assign a confidence score (0-100%)** to your conclusion.

---

### **Final Output Format (Strictly Follow This Format):**
- **Prime Suspect:** [Full Name or Alias]
- **Evidence Supporting Conclusion:** [Detailed breakdown of visual, auditory, textual, and behavioral evidence]
- **Behavioral Patterns:** [Key actions, motives, and criminal signature]
- **Confidence Level:** [0-100%]
- **Next Steps (if any):** [What additional evidence would further confirm the identity? If none, state "No further evidence required."]

If there is **insufficient evidence**, specify exactly what is missing and suggest what additional data would be needed for a conclusive identification.

This report must be **direct and definitive**--avoid speculation and provide a final, actionable determination of the suspect's identity.
"""
        try:
            response = self.client.chat.completions.create(
                model="gpt-4-turbo-preview",
                messages=[
                    {
                        "role": "system",
                        "content": "You are a forensic detective specialized in multimodal evidence analysis."
                    },
                    {"role": "user", "content": prompt_01}
                ],
                temperature=0.5,
                max_tokens=1000
            )
            
            report = response.choices[0].message.content
            logger.info("\n📋 Forensic Report Generated:")
            logger.info("=" * 50)
            logger.info(report)
            logger.info("=" * 50)
            
            return report
            
        except Exception as e:
            logger.error(f"Error generating report: {str(e)}")
            return None<h4>LLM分析における温度設定:</h4><p>当社の法医学分析システムでは、0.5 という適度な温度を使用します。このバランスの取れた設定が選択された理由は次のとおりです。</p><ul><li><p>これは、決定論的（厳しすぎる）出力と非常にランダムな出力の中間地点を表します。</p></li><li><p>0.5 では、モデルは論理的かつ正当な法医学的結論を提供するのに十分な構造を維持します。</p></li><li><p>この設定により、モデルは合理的なフォレンジック分析パラメータの範囲内でパターンを識別し、接続を確立できます。</p></li><li><p>一貫性と信頼性の高い出力の必要性と洞察力に富んだ分析を生成する能力のバランスをとります。</p></li></ul><p>この適度な温度設定により、過度に厳格で過度に推測的な結論を回避しながら、法医学的分析の信頼性と洞察力を高めることができます。</p><h3>証拠分析の実行</h3><p><strong>LLM 統合</strong>が完了したので、すべてのシステム コンポーネントを接続する<strong>スクリプト</strong>が必要です。このスクリプトは次のことを行います。</p><ul><li><p><strong>Elasticsearch で</strong> <strong>同様の証拠を検索します 。</strong></p></li><li><p><strong>LLM</strong> を使用して <strong>取得した証拠を分析し 、</strong><strong> 最終レポートを生成します。</strong></p></li></ul><h4>コード: 証拠分析スクリプト</h4>python stages/04-stage/rag_crime_analyze.pyimport sys
import os
sys.path.append(os.path.join(os.path.dirname(os.path.dirname(__file__)), 'src'))

from embedding_generator import EmbeddingGenerator
from elastic_manager import ElasticsearchManager
from llm_analyzer import LLMAnalyzer

import json
import logging
from dotenv import load_dotenv

# Setup logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# Load environment variables
load_dotenv()

# Initialize classes
generator = EmbeddingGenerator()
es_manager = ElasticsearchManager()

llm = LLMAnalyzer()
logger.info("✅ All components initialized successfully")
    
try:
    evidence_data = {}
    
    # Get data for each modality
    test_files = {
        'vision': 'data/images/crime_scene2.jpg',
        'audio': 'data/audios/joker_laugh.wav',
        'text': 'Why so serious?',
        'depth': 'data/depths/jdancing-depth.png'
    }
    
    logger.info("🔍 Collecting evidence...")
    for modality, test_input in test_files.items():
        try:
            if modality == 'text':
                embedding = generator.generate_embedding([test_input], modality)
            else:
                embedding = generator.generate_embedding([str(test_input)], modality)
            
            results = es_manager.search_similar(embedding, k=2)
            if results:
                evidence_data[modality] = results
                logger.info(f"✅ Data retrieved for {modality}: {len(results)} results")
            else:
                logger.warning(f"⚠️ No results found for {modality}")
                
        except Exception as e:
            logger.error(f"❌ Error retrieving {modality} data: {str(e)}")
    
    if not evidence_data:
        raise ValueError("No evidence data found in Elasticsearch!")
    
    # Test forensic report generation
    logger.info("\n📝 Generating forensic report...")
    report = llm.analyze_evidence(evidence_data)
    
    if report:
        logger.info("✅ Forensic report generated successfully")
        logger.info("\n📊 Report Preview:")
        logger.info("+" * 50)
        logger.info(report)
        logger.info("+" * 50)
    else:
        raise ValueError("Failed to generate forensic report")
        
except Exception as e:
    logger.error(f"❌ Error in analysis : {str(e)}")<h4>期待されるLLM出力</h4>**Prime Suspect:** The Joker

**Evidence Supporting Conclusion:**

- **Visual Evidence:**
  - The photo of the crime scene with playing cards scattered around and the graffiti of the Joker laughing matches the Joker's known calling cards and thematic elements. The similarity score of 0.83 indicates a high likelihood that these elements are directly associated with the Joker.
  - The image of the Joker with green hair, white face paint, and a sinister smile in an urban night setting, although with a lower similarity score of 0.69, still supports the presence or recent activity of the Joker in areas consistent with the crime scene's characteristics.

- **Auditory Evidence:**
  - The captured sinister laugh with a similarity score of 1.00 perfectly matches known audio profiles of the Joker, making it a direct auditory signature of his presence at or near the crime scene.
  - Despite the lower similarity score of 0.61, the second audio piece further corroborates the Joker's involvement through thematic consistency.

- **Textual Evidence:**
  - The mysterious note found at the location, with a similarity score of 0.76, likely contains thematic or direct references to the Joker's modus operandi or signature phrases, further implicating him in the crime.
  - The similarity score of 0.72 for the Joker's description in textual evidence reinforces the thematic connection to the crime scene.

- **Depth Evidence:**
  - Depth sensor capture of the suspect with a similarity score of 0.77 suggests a physical presence matching the Joker's known dimensions or characteristic movements.
  - The lower similarity score of 0.53 in the second depth evidence still contributes to the overall pattern of evidence pointing towards the Joker, albeit with less certainty.

**Behavioral Patterns:**
- The Joker is known for his theatrical crimes, often leaving behind a signature trail of chaos, including playing cards, sinister laughter, and thematic graffiti. These elements are not only consistent with his known criminal signature but also directly observed at the crime scene.
- His motives often include creating chaos, drawing attention to his acts, and challenging his arch-nemesis, Batman, making a high-profile bank heist fitting within his behavioral patterns.

**Confidence Level:** 95%

**Next Steps:** No further evidence required.

The combination of visual, auditory, textual, and depth evidence strongly points to the Joker as the prime suspect. The thematic consistency across multiple modes of evidence, combined with known behavioral patterns and criminal signature, leaves little doubt regarding his involvement. While there is always a small margin of uncertainty in forensic analysis, the evidence at hand provides a compelling case against the Joker with a high degree of confidence.<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt05526b73530f46ed/6a17ef043e03d782434f2d30/132ee0880b7fb1e64b5b2d886ab76b58baa6de37-1024x768.jpg" alt="" /><h2>結論：事件は解決した</h2><p><strong>マルチモーダル</strong> <strong>RAG システムは</strong> 、すべての 手がかりを集めて分析し 、容疑者を<strong> ジョーカー</strong> として特定しました。</p><p><strong>ImageBind</strong> を使用して<strong> 画像、音声、テキスト、深度マップを</strong> <strong>共有ベクトル空間</strong> に結合することにより、システムは手動では識別不可能だった<strong> 接続を検出</strong> できるようになりました。<strong>Elasticsearch は</strong><strong>高速かつ効率的な検索</strong>を保証し、 <strong>LLM は</strong>証拠を統合して<strong>明確で決定的なレポート</strong>を作成しました。</p><p>しかし、このシステムの<strong>真の力は</strong><strong>ゴッサム シティだけにとどまりません</strong>。<strong>マルチモーダル RAG アーキテクチャは、</strong><strong>数多くの実際のアプリケーション</strong>への扉を開きます。</p><ul><li><p><strong>都市監視:</strong><strong>画像、音声、センサーデータ</strong>に基づいて容疑者を特定します。</p></li><li><p><strong>法医学的分析:</strong><strong>複数の情報源からの証拠</strong>を相関させて<strong>複雑な犯罪</strong>を解決します。</p></li><li><p><strong>マルチメディア推奨: マルチモーダルコンテキスト</strong> <strong>を理解する</strong> <strong>推奨システム</strong> の作成 (例: 画像やテキストに基づいて<strong> 音楽 を提案する)。</strong></p></li><li><p><strong>ソーシャル メディアのトレンド:</strong>さまざまなデータ形式にわたって<strong>トレンドのトピック</strong>を検出します。</p></li></ul><p><strong>マルチモーダル RAG システムの構築</strong>方法を学習したので、<strong>独自の手がかりを使ってテストして</strong>みませんか?</p><p>あなたの発見を私たちと共有し 、<strong> マルチモーダル</strong>AI の分野での コミュニティ の進歩に貢献してください。</p><h2>特別な感謝</h2><p>このコードのデプロイメント アーキテクチャを定義するプロセス中に貴重な貢献とレビューをしてくれた Adrian Cole に感謝します。</p><h2>参照資料</h2><ul><li><p><a href="https://www.elastic.co/jp/search-labs/blog/multimodal-image-retrieval-with-roboflow">KNN検索とCLIP埋め込みを使用したマルチモーダル画像検索システムの構築</a></p></li><li><p><a href="https://www.elastic.co/jp/search-labs/tutorials/search-tutorial/vector-search/nearest-neighbor-search">k近傍法（kNN）探索</a></p></li><li><p><a href="https://pytorch.org/docs/stable/tensors.html">PyTorch のテンソルに関する公式ドキュメント</a></p></li><li><p><a href="https://imagebind.metademolab.com/">ImageBind: 感覚を超えて AI を「リンク」する新しい方法</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/building-multimodal-rag-system</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/building-multimodal-rag-system</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Alex Salgado]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt75ee2f922dacb7a8/6a17ef067b54f9775a8b39a3/47635eb4dadb8481854862668231eaa3a005ebee-1600x900.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 11 Mar 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Ollama を推論 API と共に使用する]]></title>
    <description><![CDATA[Inference API を使用して Ollama を Elasticsearch と統合する方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>この記事では、Ollama を使用してローカル モデルを Elasticsearch 推論モデルに接続し、Playground を使用してドキュメントに質問する方法を学習します。</p><p>Elasticsearch を使用すると、ユーザーは Open <a href="https://www.elastic.co/jp/guide/en/elasticsearch/reference/current/inference-apis.html">Inference API</a>を使用して LLM に接続でき、Amazon Bedrock、Cohere、Google AI、Azure AI Studio、HuggingFace などのプロバイダーをサービスとしてサポートします。</p><p><a href="https://ollama.com">Ollama は</a>、独自のインフラストラクチャ (ローカル マシン/サーバー) を使用して LLM モデルをダウンロードおよび実行できるツールです。<a href="https://ollama.com/library">ここでは</a>、Ollama と互換性のある利用可能なモデルのリストを見つけることができます。</p><p>Ollama は、各モデルをセットアップする方法の違いや、モデル関数にアクセスするための API の作成方法について心配することなく、さまざまなオープンソース モデルをホストしてテストしたい場合に最適なオプションです。Ollama がすべてを処理します。</p><p>Ollama API は OpenAI API と互換性があるため、推論モデルを簡単に統合し、Playground を使用して RAG アプリケーションを作成できます。</p><h2>要件</h2><ol><li><p>エラスティックサーチ 8.17</p></li><li><p>キバナ 8.17</p></li><li><p>Python</p></li></ol><h2>ステップ</h2><ol><li><p><a href="https://www.elastic.co/jp/search-labs/blog/ollama-with-inference-api#setting-up-ollama-llm-server">Ollama LLMサーバーのセットアップ</a></p></li><li><p><a href="https://www.elastic.co/jp/search-labs/blog/ollama-with-inference-api#creating-mappings">マッピングの作成</a></p></li><li><p><a href="https://www.elastic.co/jp/search-labs/blog/ollama-with-inference-api#indexing-data">データのインデックス作成</a></p></li><li><p><a href="https://www.elastic.co/jp/search-labs/blog/ollama-with-inference-api#asking-questions-using-playground">Playgroundを使って質問する</a></p></li></ol><h2>Ollama LLMサーバーのセットアップ</h2><p>Ollama を使用して、LLM サーバーを設定し、Playground インスタンスに接続します。以下のことが必要です:</p><ul><li><p>Ollamaをダウンロードして実行します。</p></li><li><p>ngrokを使用して、インターネット経由でOllamaをホストするローカルWebサーバーにアクセスします。</p></li></ul><h3>Ollamaをダウンロードして実行する</h3><p>Ollama を使用するには、まず<a href="https://ollama.com/download">ダウンロードする</a>必要があります。Ollama は Linux、Windows、macOS をサポートしているので、<a href="https://ollama.com/download">ここからお使いの OS と互換性のある Ollama バージョンをダウンロードしてください。</a>Ollama がインストールされると、サポートされている LLM の<a href="https://ollama.com/library">リスト</a>からモデルを選択できます。この例では、汎用多言語モデルである<a href="https://ollama.com/library/llama3.2">llama3.2</a>モデルを使用します。セットアップ プロセスでは、Ollama のコマンド ライン ツールを有効にします。ダウンロードが完了したら、次の行を実行できます。</p>ollama pull llama3.2<p>出力は次のようになります:</p>pulling manifest
pulling dde5aa3fc5ff... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 2.0 GB
pulling 966de95ca8a6... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 1.4 KB
pulling fcc5a6bec9da... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 7.7 KB
pulling a70ff7e570d9... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏ 6.0 KB
pulling 56bb8bd477a5... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏   96 B
pulling 34bb5ab01051... 100% ▕█████████████████████████████████████████████████████████████████████████████████████████▏  561 B
verifying sha256 digest
writing manifest
success<p>インストールしたら、次のコマンドでテストできます。</p>ollama run llama3.2<p>質問してみましょう:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc12f240920e897f5/6a17f39425daab32a508a367/ad1eff81c1b04d2a747c3afd0ecbc215e5bd96fd-800x501.gif" alt="Ollamaを実行して質問する" /><p>モデルを実行すると、Ollama はポート「11434」でデフォルトで実行される API を有効にします。<a href="https://github.com/ollama/ollama/blob/main/docs/api.md">公式ドキュメント</a>に従って、その API にリクエストを送信してみましょう。</p>curl http://localhost:11434/api/generate -d '{                                          
  "model": "llama3.2",               
  "prompt": "What is the capital of France?"
}' <p>私たちが受け取った回答は次のとおりです。</p>{"model":"llama3.2","created_at":"2024-11-28T21:48:42.152817532Z","response":"The","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.251884485Z","response":" capital","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.347365913Z","response":" of","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.446837322Z","response":" France","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.542367394Z","response":" is","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.644580384Z","response":" Paris","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.739865362Z","response":".","done":false}
{"model":"llama3.2","created_at":"2024-11-28T21:48:42.834347518Z","response":"","done":true,"done_reason":"stop","context":[128006,9125,128007,271,38766,1303,33025,2696,25,6790,220,2366,18,271,128009,128006,882,128007,271,3923,374,279,6864,315,9822,30,128009,128006,78191,128007,271,791,6864,315,9822,374,12366,13],"total_duration":6948567145,"load_duration":4386106503,"prompt_eval_count":32,"prompt_eval_duration":1872000000,"eval_count":8,"eval_duration":684000000}<p><em>このエンドポイントの特定の応答はストリーミングであることに注意してください。</em></p><h3>ngrokを使用してエンドポイントをインターネットに公開する</h3><p>エンドポイントはローカル環境で動作するため、Elastic Cloud インスタンスなどの別のポイントからインターネット経由でアクセスすることはできません。<a href="https://ngrok.com">ngrok を</a>使用すると、パブリック IP を提供するポートを公開できます。ngrok でアカウントを作成し、公式の<a href="https://dashboard.ngrok.com/get-started/setup">セットアップ ガイド</a>に従ってください。</p><p>ngrok エージェントをインストールして構成したら、Ollama が使用しているポートを公開できます。</p>ngrok http 11434 --host-header="localhost:11434"<p><em>注: ヘッダー</em><em><code>--host-header="localhost:11434"</code></em>は<em>、リクエスト内の「Host」ヘッダーが「localhost:11434」と一致することを保証します。</em></p><p>このコマンドを実行すると、ngrok と Ollama サーバーがローカルで実行されている限り機能するパブリック リンクが返されます。</p>Session Status                online                                                                                                                                                                              
Account                       xxxx@yourEmailProvider.com (Plan: Free)                                                                                                                                             
Version                       3.18.4                                                                                                                                                                              
Region                        United States (us)                                                                                                                                                                  
Latency                       561ms                                                                                                                                                                               
Web Interface                 http://127.0.0.1:4040                                                                                                                                                               
Forwarding                    https://your-ngrok-url.ngrok-free.app -&gt; http://localhost:11434                                                                                                                   


Connections                   ttl     opn     rt1     rt5     p50     p90                                                                                                                                         
                              0       0       0.00    0.00    0.00    0.00                                                ```<p>「転送」では、ngrok が URL を生成したことがわかります。後で使用するために保存します。</p><p>ngrok によって生成された URL を使用して、エンドポイントへの HTTP リクエストを再度実行してみましょう。</p>curl https://your-ngrok-endpoint.ngrok-free.app/api/generate -d '{                                          
  "model": "llama3.2",               
  "prompt": "What is the capital of France?"
}'<p>応答は前のものと同様になるはずです。</p><h2>マッピングの作成</h2><h3>ELSERエンドポイント</h3><p>この例では、 <a href="https://www.elastic.co/jp/guide/en/elasticsearch/reference/current/put-inference-api.html">Elasticsearch 推論 API を使用して推論エンドポイントを作成します</a>。さらに、 <a href="https://www.elastic.co/jp/guide/en/machine-learning/current/ml-nlp-elser.html">ELSER</a>を使用して埋め込みを生成します。</p>PUT _inference/sparse_embedding/medicines-inference
{
  "service": "elasticsearch",
  "service_settings": {
    "num_allocations": 1,
    "num_threads": 1,
    "model_id": ".elser_model_2_linux-x86_64"
  }
}<p>この例では、2 種類の薬を販売する薬局があるとします。</p><ul><li><p>処方箋が必要な医薬品。</p></li><li><p>処方箋を必要としない医薬品。</p></li></ul><p>この情報は各薬剤の説明欄に含まれます。</p><p>LLM はこのフィールドを解釈する必要があるため、次のデータ マッピングを使用します。</p>PUT medicines
{
  "mappings": {
    "properties": {
      "name": {
        "type": "text",
        "copy_to": "semantic_field"
      },
      "semantic_field": {
        "type": "semantic_text",
        "inference_id": "medicines-inference"
      },
      "text_description": {
        "type": "text",
        "copy_to": "semantic_field"
      }
    }
  }
}<p>フィールド<code>text_description</code>は説明のプレーンテキストが格納され、 <a href="https://www.elastic.co/jp/guide/en/elasticsearch/reference/current/semantic-text.html">semantic_text</a>フィールド タイプである<code>semantic_field</code>には ELSER によって生成された埋め込みが格納されます。</p><p>プロパティ<a href="https://www.elastic.co/jp/guide/en/elasticsearch/reference/current/copy-to.html">copy_to は</a>、フィールド名と<code>text_description</code>の内容をセマンティック フィールドにコピーし、それらのフィールドの埋め込みが生成されます。</p><h2>データのインデックス作成</h2><p>ここで、 <a href="https://www.elastic.co/jp/guide/en/elasticsearch/reference/current/docs-bulk.html">_bulk API</a>を使用してデータをインデックス化してみましょう。</p>POST _bulk
{"index":{"_index":"medicines"}}
{"id":1,"name":"Paracetamol","text_description":"An analgesic and antipyretic that does NOT require a prescription."}
{"index":{"_index":"medicines"}}
{"id":2,"name":"Ibuprofen","text_description":"A nonsteroidal anti-inflammatory drug (NSAID) available WITHOUT a prescription."}
{"index":{"_index":"medicines"}}
{"id":3,"name":"Amoxicillin","text_description":"An antibiotic that requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":4,"name":"Lorazepam","text_description":"An anxiolytic medication that strictly requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":5,"name":"Omeprazole","text_description":"A medication for stomach acidity that does NOT require a prescription."}
{"index":{"_index":"medicines"}}
{"id":6,"name":"Insulin","text_description":"A hormone used in diabetes treatment that requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":7,"name":"Cold Medicine","text_description":"A compound formula to relieve flu symptoms available WITHOUT a prescription."}
{"index":{"_index":"medicines"}}
{"id":8,"name":"Clonazepam","text_description":"An antiepileptic medication that requires a prescription."}
{"index":{"_index":"medicines"}}
{"id":9,"name":"Vitamin C","text_description":"A dietary supplement that does NOT require a prescription."}
{"index":{"_index":"medicines"}}
{"id":10,"name":"Metformin","text_description":"A medication used for type 2 diabetes that requires a prescription."}<p>対応：</p>{
   "errors": false,
   "took": 34732020848,
   "items": [
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "mYoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 0,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "mooeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 1,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "m4oeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 2,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "nIoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 3,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "nYoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 4,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "nooeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 5,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "n4oeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 6,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "oIoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 7,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "oYoeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 8,
     	"_primary_term": 1,
     	"status": 201
   	}
 	},
 	{
   	"index": {
     	"_index": "medicines",
     	"_id": "oooeMpQBF7lnCNFTfdn2",
     	"_version": 1,
     	"result": "created",
     	"_shards": {
       	"total": 2,
       	"successful": 2,
       	"failed": 0
     	},
     	"_seq_no": 9,
     	"_primary_term": 1,
     	"status": 201
   	}
 	}
   ]
 }<h2>Playgroundを使って質問する</h2><p><a href="https://www.elastic.co/jp/guide/en/kibana/current/playground.html">Playground は</a>、Elasticsearch インデックスと LLM プロバイダーを使用して RAG システムをすばやく作成できる Kibana ツールです。詳細については、この<a href="https://www.elastic.co/jp/search-labs/blog/playground-connectors-data-chat">記事</a>をお読みください。</p><h3>ローカルLLMをPlaygroundに接続する</h3><p>まず、作成したパブリック URL を使用するコネクタを作成する必要があります。Kibana で、 <strong>「検索」&gt;「プレイグラウンド」</strong>に移動し、「LLM に接続」をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt22148eabfabf6d3f/6a17f3963e9e459f97ba15c6/1854f0808f8150e359fe62ba5d901d32a88d477c-1600x867.png" alt="地元のLLMをOllamaのPlaygroundに接続する" /><p>このアクションにより、Kibana インターフェースの左側にメニューが表示されます。そこで、「OpenAI」をクリックします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6e28194d9012f141/6a17f39725daab500a08a36b/c83d3c4d7035a518124ad7d22b38764db57b6800-933x1007.png" alt="コネクタを選択: Open AI Ollama" /><p>これで、OpenAI コネクタの設定を開始できます。</p><p>「コネクタ設定」に移動し、OpenAI プロバイダーの場合は「その他 (OpenAI 互換サービス)」を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9984dce6f78a7c08/6a17f3990b0bed0b7add36c4/ecfcdc4b575c309bd55b4e61ca0ddb348aa84f64-917x268.png" alt="Ollama を Inference API で使用するためのコネクタ設定を行う" /><p>次に、他のフィールドを設定しましょう。この例では、モデル名を「medicines-llm」とします。URLフィールドには、ngrokによって生成されたURL（/v1/chat/completions）を入力します。「デフォルトモデル」フィールドで、「llama3.2」を選択します。API キーは使用しないので、任意のテキストを入力して続行します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8a24b93a39d380fb/6a17f39b96142a15c7eb1c3c/5d3b5027c8096cbe49fb740d70aa24e849611a9d-916x688.png" alt="設定を追加する" /><p>「保存」をクリックし、「データ ソースの追加」をクリックしてインデックス医薬品を追加します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4f107a54d5be25f9/6a17f39d4b055deb9d432338/525113da59e902c8235f62bde8fb62371a63e11b-1579x753.png" alt="Playground を使用してドキュメントに質問するためのデータ ソースを追加します" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt03fb36fe05dcdbf5/6a17f39ebe608602f40048be/96138de0bbe2c2ac619f64889d3487df62739ca4-466x805.png" alt="クエリデータを追加する" /><p>素晴らしい！これで、RAG エンジンとしてローカルで実行している LLM を使用して Playground にアクセスできるようになりました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb1b27580107259b3/6a17f3a096142abefceb1c40/cfb48b33c70f4534ab77eb01f58008237f65e6f4-1600x851.png" alt="プレイグラウンドでモデル設定を選択する" /><p>テストする前に、エージェントにさらに具体的な指示を追加し、モデルに送信されるドキュメントの数を 10 に増やして、回答にできるだけ多くのドキュメントが含まれるようにします。コンテキスト フィールドは<code>semantic_field</code>になり、copy_to プロパティにより、薬剤の名前と説明が含まれます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4fbc97dc87c6fc62/6a17f3a1e8fbce052c3a1aa0/0c57c9c0e1a0e7b58fffdd3ef81d67d41e2990c4-580x806.png" alt="Elastic PlaygroundのMoel設定" /><p>さて、質問しましょう：<em><strong>処方箋なしでクロナゼパムを購入できますか？</strong></em>そして何が起こるか見てみましょう:</p><p>予想通り、正しい答えが得られました。</p><h3>今後の見通し</h3><p>次のステップは、独自のアプリケーションを作成することです。Playground は、マシン上で実行し、ニーズに合わせてカスタマイズできる Python のコード スクリプトを提供します。たとえば、 <a href="https://fastapi.tiangolo.com/">FastAPI</a>サーバーの背後に配置して、UI で使用される QA 医薬品チャットボットを作成します。</p><p>このコードは、Playground の右上のセクションにある<em><strong>[コードの表示]</strong></em>ボタンをクリックすると見つかります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt42fe193b8aa08830/6a17f3a33e9e4569e8ba15ca/816bfd0e5f936ad65dbe719d5df10714e550a40b-380x121.png" alt="コード表示ボタン" /><p>また、<em><strong>エンドポイントと API キー</strong></em>を使用して、コードで必要な<code>ES_API_KEY</code>環境変数を生成します。</p><p>この特定の例では、コードは次のようになります。</p>## Install the required packages
## pip install -qU elasticsearch openai
import os
from elasticsearch import Elasticsearch
from openai import OpenAI
es_client = Elasticsearch(
    "https://your-deployment.us-central1.gcp.cloud.es.io:443",
    api_key=os.environ["ES_API_KEY"]
)
openai_client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
)
index_source_fields = {
    "medicines": [
        "semantic_field"
    ]
}
def get_elasticsearch_results():
    es_query = {
        "retriever": {
            "standard": {
                "query": {
                    "nested": {
                        "path": "semantic_field.inference.chunks",
                        "query": {
                            "sparse_vector": {
                                "inference_id": "medicines-inference",
                                "field": "semantic_field.inference.chunks.embeddings",
                                "query": query
                            }
                        },
                        "inner_hits": {
                            "size": 2,
                            "name": "medicines.semantic_field",
                            "_source": [
                                "semantic_field.inference.chunks.text"
                            ]
                        }
                    }
                }
            }
        },
        "size": 3
    }
    result = es_client.search(index="medicines", body=es_query)
    return result["hits"]["hits"]
def create_openai_prompt(results):
    context = ""
    for hit in results:
        inner_hit_path = f"{hit['_index']}.{index_source_fields.get(hit['_index'])[0]}"
        ## For semantic_text matches, we need to extract the text from the inner_hits
        if 'inner_hits' in hit and inner_hit_path in hit['inner_hits']:
            context += '\n --- \n'.join(inner_hit['_source']['text'] for inner_hit in hit['inner_hits'][inner_hit_path]['hits']['hits'])
        else:
            source_field = index_source_fields.get(hit["_index"])[0]
            hit_context = hit["_source"][source_field]
            context += f"{hit_context}\n"
    prompt = f"""
  Instructions:
  - You are an assistant specializing in answering questions about the sale of medicines.
  - Answer questions truthfully and factually using only the context presented.
  - If you don't know the answer, just say that you don't know, don't make up an answer.
  - You must always cite the document where the answer was extracted using inline academic citation style [], using the position.
  - Use markdown format for code examples.
  - You are correct, factual, precise, and reliable.
  Context:
  {context}
  """
    return prompt
def generate_openai_completion(user_prompt, question):
    response = openai_client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "system", "content": user_prompt},
            {"role": "user", "content": question},
        ]
    )
    return response.choices[0].message.content
if __name__ == "__main__":
    question = "my question"
    elasticsearch_results = get_elasticsearch_results()
    context_prompt = create_openai_prompt(elasticsearch_results)
    openai_completion = generate_openai_completion(context_prompt, question)
    print(openai_completion)<p>Ollama で動作させるには、OpenAI クライアントを OpenAI サーバーではなく Ollama サーバーに接続するように変更する必要があります。OpenAI の例と互換性のあるエンドポイントの完全なリストについては、こちらをご覧ください。</p>openai_client = OpenAI(
    # you can use http://localhost:11434/v1/ if running this code locally.
    base_url='https://your-ngrok-url.ngrok-free.app/v1/',
    # required but ignored
    api_key='ollama',
)<p>また、補完メソッドを呼び出すときにモデルを llama3.2 に変更します。</p>def generate_openai_completion(user_prompt, question):
    response = openai_client.chat.completions.create(
        model="llama3.2",
        messages=[
            {"role": "system", "content": user_prompt},
            {"role": "user", "content": question},
        ]
    )
    return response.choices[0].message.content<p>質問を追加しましょう:<em><strong>処方箋なしでクロナゼパムを購入できますか?</strong></em>Elasticsearch クエリへ:</p>def get_elasticsearch_results():
    es_query = {
        "retriever": {
            "standard": {
                "query": {
                    "nested": {
                        "path": "semantic_field.inference.chunks",
                        "query": {
                            "sparse_vector": {
                                "inference_id": "medicines-inference",
                                "field": "semantic_field.inference.chunks.embeddings",
                                "query": "Can I buy Clonazepam without a prescription?"
                            }
                        },
                        "inner_hits": {
                            "size": 2,
                            "name": "medicines.semantic_field",
                            "_source": [
                                "semantic_field.inference.chunks.text"
                            ]
                        }
                    }
                }
            }
        },
        "size": 3
    }
    result = es_client.search(index="medicines", body=es_query)
    return result["hits"]["hits"]<p>また、いくつかの出力を伴う完了呼び出しでは、質問のコンテキストの一部として Elasticsearch の結果が送信されていることを確認できます。</p>if __name__ == "__main__":
    question = "Can I buy Clonazepam without a prescription?"
    elasticsearch_results = get_elasticsearch_results()
    context_prompt = create_openai_prompt(elasticsearch_results)
    print("========== Context Prompt START ==========")
    print(context_prompt)
    print("========== Context Prompt END ==========")
    print("========== Ollama Completion START ==========")
    openai_completion = generate_openai_completion(context_prompt, question)
    print(openai_completion)
    print("========== Ollama Completion END ==========")<p>ではコマンドを実行してみましょう</p><p><code>pip install -qU elasticsearch openai</code></p><p><code>python main.py</code></p><p>次のような画面が表示されます。</p>========== Context Prompt START ==========
  Instructions:
  - You are an assistant specializing in answering questions about the sale of medicines.
  - Answer questions truthfully and factually using only the context presented.
  - If you don't know the answer, just say that you don't know, don't make up an answer.
  - You must always cite the document where the answer was extracted using inline academic citation style [], using the position.
  - Use markdown format for code examples.
  - You are correct, factual, precise, and reliable.
  Context:
  Clonazepam
 ---
An antiepileptic medication that requires a prescription.A nonsteroidal anti-inflammatory drug (NSAID) available WITHOUT a prescription.
 ---
IbuprofenAn anxiolytic medication that strictly requires a prescription.
 ---
Lorazepam


========== Context Prompt END ==========
========== Ollama Completion START ==========
No, you cannot buy Clonazepam over-the-counter (OTC) without a prescription [1]. It is classified as a controlled substance in the United States due to its potential for dependence and abuse. Therefore, it can only be obtained from a licensed healthcare provider who will issue a prescription for this medication.
========== Ollama Completion END ==========<h2>まとめ</h2><p>この記事では、Ollama などのツールを Elasticsearch 推論 API および Playground と組み合わせて使用した場合の、その強力さと汎用性について説明します。</p><p>いくつかの簡単な手順を実行するだけで、LLM を使用したチャット機能を備えた RAG アプリケーションが、コストゼロで独自のインフラストラクチャで稼働するようになりました。これにより、さまざまなタスクのさまざまなモデルにアクセスできるだけでなく、リソースと機密情報をより細かく制御できるようになります。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ollama-with-inference-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ollama-with-inference-api</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd9c8eb0fc946920e/6a17f3a46864a4b2fbb688f0/399b9ef527be633845fb6505b68132cc03bc9e09-1150x628.png" length="0" type="image/png"/>
    <pubDate>Fri, 14 Feb 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[OllamaとKibanaを使用して、RAG用にDeepSeek R1をローカルでテストします。]]></title>
    <description><![CDATA[DeepSeekのローカルインスタンスを実行し、Kibana内から接続する方法を学びましょう。]]></description>
    <content:encoded><![CDATA[<p>中国のヘッジファンドHigh-Flyerの新しい大規模言語モデルDeepSeek R1が話題になっています。オープンウェイトの有能で思考の連鎖的推論法が導入された今、業界にとってこれが何を意味するのかについての憶測が飛び交っています。RAGとElasticsearchのすべてのベクトルデータベース機能を使用してこの新しいモデルを試してみたい方のために、ローカル推論を使用してDeepSeek R1を使い始めるための簡単なチュートリアルを以下に示します。その過程で、ElasticのPlayground機能を使用し、RAGに対するDeepseek R1の良い点と悪い点も発見します。</p><p>このチュートリアルで設定する内容の図を以下に示します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3214cc505e3d4d06/6a17df98dbb4ff12bafb55da/8aafec9011e986cd85b10958544a4d77be81e518-739x419.png" alt="ElasticsearchとOllamaを使用したDeepseekの構成" /><h2>Ollamaを使用したローカル推論の設定</h2><p><a href="https://ollama.com/">Ollama</a>は、ローカル推論用に厳選されたオープンソースモデルのセットを迅速にテストするための優れた方法であり、AI開発者に人気のツールです。</p><h3>Ollamaをベアメタルで実行</h3><p>Mac、Linux、またはWindowsでの<a href="https://github.com/ollama/ollama/tree/main?tab=readme-ov-file#ollama">ローカルインストール</a>は、特にMシリーズのAppleチップを使用しているユーザーにとって、ローカルのGPU機能を活用する最も簡単な方法です。Ollamaをインストールしたら、次のコマンドを使用してDeepSeek R1をダウンロードして実行できます。</p><p>ハードウェアに適したパラメータサイズに調整することをお勧めします。利用可能なサイズは<a href="https://ollama.com/library/deepseek-r1">こちら</a>でご確認いただけます。</p>ollama run deepseek-r1:7b<p>ターミナルでモデルとチャットできますが、CTL+dでコマンドを終了するか「/bye」と入力しても、モデルは実行されたままです。モデルがまだ実行中であることを確認するには、次のように入力します。</p>ollama ps<h3>Ollamaをコンテナ内で実行</h3><p>最も簡単な方法としては、OllamaをDockerのようなコンテナエンジンを利用して実行することもできます。ローカルマシンのGPUの使用は、環境によっては必ずしも簡単ではありませんが、コンテナに数GBモデルに適合するRAMとストレージがあれば、簡単なテストセットアップを行うことは難しくありません。</p><p>OllamaをDockerで起動して実行するには、次のコマンドを実行するだけです。</p>mkdir ollama_deepseek
cd ollama_deepseek
mkdir ollama
docker run -d -v ./ollama:/root/.ollama -p 11434:11434 \
--name ollama ollama/ollama
<p>これにより、現在のディレクトリに「ollama」というディレクトリが作成され、コンテナ内にマウントされ、Ollamaの設定とモデルを格納します。使用するパラメータの数によって、数GBから数十GBまでの範囲になるため、十分な空き容量があるボリュームを選択してください。</p><p>注意：マシンにNvidia GPUが搭載されている場合は、必ず<a href="https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#installation">Nvidia コンテナツールキット</a>をインストールし、上記のdocker 実行コマンドに「--gpus=all」を追加してください。</p><p>Ollamaコンテナがマシン上で起動したら、deepseek-r1のようなモデルを次のコマンドでプルできます。</p>docker exec -it ollama ollama pull deepseek-r1:7b<p>ベアメタルアプローチと同様に、ハードウェアに合ったサイズにパラメーターサイズを調整したい場合があります。利用可能なサイズについては、<a href="https://ollama.com/library/deepseek-r1">https://ollama.com/library/deepseek-r1</a>をご覧ください。</p><p>モデルのプルが終了したら、「/bye」と入力してプロンプトを終了できます。モデルがまだ実行中であることを確認するには以下の手段を実行します。</p>docker exec -it ollama ollama ps<h2>curlを使用したローカル推論のテスト</h2><p>curlを使用してローカル推論をテストするには、次のコマンドを実行します。JSONナラティブ応答を簡単に読み取れるように、stream:falseを使用しています。</p>curl http://localhost:11434/api/generate -d '{
  "model": "deepseek-r1:7b",
  "stream": false,
  "prompt":"Why is Elastic so cool?"
}'<h2>「OpenAI互換」のOllamaとRAGプロンプトをテスト</h2><p>便利なことに、OllamaはKibanaを含む幅広いツールとの互換性のために、OpenAIの動作を模倣するRESTエンドポイントも提供しています。</p>curl http://localhost:11434/v1/chat/completions -d '{
  "model": "deepseek-r1:7b",
  "stream": false,
  "messages": [
    { 
      "role": "system", 
      "content": "You are a helpful AI Assistant that uses the following context to answer questions only use the following context. \n\nContext:  The color of the sky today is purple. "},
    { "role": "user", 
      "content": "What does the sky look like today?" 
    }
  ]
}'<p>このより複雑なプロンプトをテストすると、&lt;think&gt;セクションを含むコンテンツが生成されます。このセクションでは、モデルが問題を推論するようにトレーニングされています。</p>&lt;think&gt; 
Okay, so I need to figure out what the user is asking for here. They provided a context where the sky is described as purple today and then asked about how the sky looks. At first glance, it seems straightforward—maybe they just want confirmation or more details on why the sky is that color.
Wait, but maybe there's something deeper. Purple skies aren't something I encounter every day. It usually happens at certain times of the year, like during sunrise or sunset with the sun setting in pink or orange. Could this be a hint about the time of day? Or perhaps it's just an unusual natural phenomenon? 
I should consider if \"purple\" is a typo. Maybe they meant something else like blue or gray. But since they specifically said purple, I'll go with that. Purple skies can happen when there are atmospheric conditions that scatter light differently, maybe due to pollution or cloud cover affecting the sunset.

So, putting it all together, the user might be looking for an explanation of why today's sky is purple and what that implies about the weather or time of day. Alternatively, they could just want a simple statement confirming that the sky looks purple today.
&lt;/think&gt;

The color of the sky today is described as purple. This unusual shade can occur due to atmospheric conditions affecting light scattering, such as during sunrise/sunset with pollution or cloud cover influencing the sunset's hues.<h2>OllamaをKibanaに接続する</h2><p>Elasticsearchを使用する優れた方法は、「<a href="https://github.com/elastic/start-local?tab=readme-ov-file#-try-elasticsearch-and-kibana-locally">start-local</a>」開発スクリプトを使うことです。</p><p>KibanaとElastisearchがネットワーク上でOllamaにアクセスできることを確認します。Elastic Stackのローカルコンテナ設定を使用している場合は、「localhost」を「host.docker.internal」または「host.containers.internal」に置き換える必要があるかもしれません。これは、ホストマシンへのネットワークパスを取得するためです。</p><p>Kibanaで、[Stack Management] &gt; [アラートと洞察] &gt; [コネクタ]に移動します。</p><h3>この一般的な設定の警告が表示された場合の対処方法</h3><p>xpack.encryptedSavedObjects.encryptionKeyが<a href="https://www.elastic.co/guide/en/kibana/current/xpack-security-secure-saved-objects.html">正しく設定されている</a>ことを確認する必要があります。これは、KibanaのローカルDockerインストールを実行する際によく見落とされる手順であるため、Docker構文で修正する手順をリストします。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta4f7b7be2e04afae/6a17df9a1d1b8391e293e393/b70b4b810bcac1d1599b07da90a98c5c744a38de-497x223.png" alt="" /><p>コンテナのシャットダウン時に変更が保存されるように、kibana/configディレクトリを永続化してください。私のKibanaコンテナのボリュームは、docker-compose.ymlで次のようになります。</p>services:
  kibana:
...
   volumes:
      - certs:/usr/share/kibana/config/certs
      - kibanadata:/usr/share/kibana/data
      - kibanaconfig:/usr/share/kibana/config
...
volumes:
  certs:
    driver: local
  esdata01:
    driver: local
  kibanadata:
    driver: local
  kibanaconfig:
    driver: local<p>これで、キーストアを作成し、値を入れて、コネクタのキーが平文で格納されないようにすることができます。</p>## generate some new keys for me and print them to the terminal
docker exec -it kibana_1 bin/kibana-encryption-keys generate

## create a new keystrore
docker exec -it kibana_1 bin/kibana-keystore create
docker exec -it kibana_1 bin/kibana-keystore add xpack.encryptedSavedObjects.encryptionKey

## You'll be prompted to paste in a value<p>変更を確実に有効にするには、クラスター全体を完全に再起動してください。</p><h3>コネクタの作成</h3><p>コネクタ構成画面（Kibana では、[Stack Management] &gt; [アラートと洞察] &gt; [コネクタ] に移動）からコネクタを作成し、[OpenAI] タイプを選択します。</p><p>コネクタを次の設定で構成します。</p><ul><li><p>コネクタ名：Deepseek（Ollama）</p></li><li><p>OpenAIプロバイダーを選択：その他（OpenAI互換サービス）</p></li><li><p>URL: <a href="http://localhost:11434/v1/chat/completions">http://localhost:11434/v1/chat/completions</a></p><ul><li><p>ollamaへの正しいパスを調整してください。コンテナ内から呼び出す場合は、host.docker.internalまたは同等のものに置き換えてください。</p></li></ul></li><li><p>デフォルトモデル：deepseek-r1:7b</p></li><li><p>APIキー：何か適当に入力してください。入力は必要ですが、値は重要ではありません。</p></li></ul><p>コネクタセットアップでのOllamaへのカスタムコネクタのテストは現在8.17では機能しませんが、Kibanaの今後の8.18ビルドでは修正される予定です。</p><p>コネクタは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt774e0793eb110f9d/6a17df9c445de981014d004d/4ce214aa953b4090ed112fbde40b01c01fb8f5c7-786x836.png" alt="" /><h2>Elasticsearchにベクトル埋め込みデータを取り込む</h2><p>すでにPlaygroundに精通していてデータが設定されている場合は以下のPlaygroundの手順にスキップできますが、簡単なテストデータが必要な場合は、_inference APIが設定されていることを確認する必要があります。8.17以降、機械学習の割り当ては動的であるため、e5多言語高密度ベクトルをダウンロードして有効にするには、Kiban Devツールで次のコマンドを実行する必要があります。</p>GET /_inference


POST /_inference/text_embedding/.multilingual-e5-small-elasticsearch
{
   "input": "are internet memes about deepseek sound investment advice?"
}<p>まだダウンロードしていない場合は、Elasticのモデルリポジトリからe5モデルのダウンロードが開始されます。</p><p>次に、RAGのコンテキストとしてパブリックドメインの書籍を読み込みましょう。こちらの<a href="https://www.gutenberg.org/cache/epub/11/pg11.txt">リンク</a>から、プロジェクト・グーテンベルクで「不思議の国のアリス」をダウンロードできます。これを.txtファイルとして保存してください。</p><p>[Elasticsearch] &gt; [ホーム] &gt; [ファイルのアップロード] に移動します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt45c594487844ecb4/6a17df9dfaa9137edb93c786/649042271f34a5e66789b17c39bfe95971c7f4ce-1360x629.png" alt="" /><p>テキストファイルを選択またはドラッグアンドドロップし、インポートボタンを押します。</p><p>「データのインポート」画面で「詳細設定」タブを選択し、インデックス名を「book_alice」に設定します。</p><p>「自動的に作成されたフィールド」のすぐ下にある小さな「追加フィールドを追加」オプションを選択します。「セマンティックテキストフィールドの追加」を選択し、推論エンドポイントを「.multilingual-e5-small-elasticsearch」に変更します。[追加] を選択し、[インポート] を選択します。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d9c22ccdaee8590/6a17df9f3e03d731f94f2b8e/e58d5c9a2d406d8e62eb96cab9ac98ca89414346-507x602.png" alt="" /><p></p><p>ロードと推論が完了したら、Playgroundに向かう準備が整います。</p><h2>PlaygroundでのRAGのテスト</h2><p>Kibanaで [Elasticsearch] &gt; [Playground] に移動します。</p><p>Playground画面には、コネクタが存在することを示す緑色のチェックマークと「LLM接続済み」と表示されます。これは、上記で作成したOllamaコネクタです。Playgroundのより詳しいガイドについては、<a href="https://www.elastic.co/guide/en/kibana/current/playground.html">こちら</a>をご覧ください。</p><p>青色の [データソースを追加] をクリックし、以前に作成したbook_aliceインデックス、または埋め込みに推論APIを利用する以前に構成した別のインデックスを選択します。</p><p>Deepseekは、強力なアライメント特性を備えた思考連鎖モデルです。これはRAGの観点から見ると良い点と悪い点の両方があります。思考連鎖のトレーニングは、Deepseekが引用文の中で一見矛盾しているように見える記述を合理化するのに役立つかもしれませんが、トレーニング知識への強力な整合性により、コンテキストの根拠よりも独自の世界の事実を優先する可能性があります。善意ではありますが、この強固な整合性により、LLMが個人的な知識が一致しないトピックや、トレーニングデータセットで十分に表現されていないトピックについて話し合う際に、指導が困難になることが知られています。</p><p>Playgroundのセットアップでは、「You are an assistant for question-answering tasks using relevant text passages from the book Alice in wonderland」というシステムプロンプトを入力し、他のデフォルトを受け入れました。</p><p>「Who was at the tea party?」という質問に対する答えはこうです。「Answer: The March Hare, the Hatter, and the Dormouse were at the tea party. [Citation: position 1 and 2]」これは正しいです。
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0ce6a0facd972fdd/6a17dfa03e03d79aaa4f2b92/e8af3ff93a72e1f02de8e73f6c2606cbc19970e5-1296x813.png" alt="" /><p>&lt;think&gt;タグを見ると、Deepseekが質問に答えるために引用の内容をしっかりと検討したことがわかります。</p><h2>アライメントの限界をテスト</h2><p>Deepseekのために知的に挑戦的なシナリオをテストとして作成しましょう。Deepseekのトレーニングデータが真実ではないと知っている陰謀論のインデックスを作成します。</p><p>Kibana開発ツールで、次のインデックスとデータを作成しましょう。</p>PUT /classic_conspiracies
{
   "mappings": {
       "properties": {
           "content": {
               "type": "text",
               "copy_to": "content_semantic"
           },
           "content_semantic": {
               "type": "semantic_text",
               "inference_id": ".multilingual-e5-small-elasticsearch"
           }
       }
   }
}




POST /classic_conspiracies/_doc/1
{
   "content": "birds aren't real, the government replaced them with drones a long time ago"
}
POST /classic_conspiracies/_doc/2
{
   "content": "tinfoil hats are necessary to prevent our brains from being read"
}
POST /classic_conspiracies/_doc/3
{
   "content": "ancient aliens influenced early human civilizations, this explains why things made out of stone are marginally similar on different continents"
}<p>
これらの陰謀論は、LLMの基礎となるでしょう。積極的にシステムプロンプトを出したにもかかわらず、Deepseekは私たちが説明した事実を受け入れません。自分の個人データの方が信頼性が高く、根拠があり、組織のニーズに合っていることがわかっている状況だったら、これは受け入れられないでしょう。</p><p>「are birds real?」というテストの質問（説明は<a href="https://knowyourmeme.com/memes/birds-arent-real">know your meme</a>）に対して、次の答えが得られます。「In the provided context, birds are not considered real, but in reality, they are real animals. [Context: position 1]」。このテストにより、DeepSeek R1は7Bパラメータレベルでも強力であることが証明されました。ただし、データセットによっては、RAGにとって最適な選択ではない可能性があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5e8dabf65ea200e1/6a17dfa2ec0f8982135a6541/67d5f6cdb97bfd3adb926cbd588c768f9d6730ae-1277x737.png" alt="" /><h2>今回学んだ内容</h2><p>まとめると以下のようになります。</p><ul><li><p>Ollamaなどのツールでモデルをローカルで実行することは、モデルの動作を確認するのに最適なオプションです。</p></li><li><p>DeepSeek R1は推論モデルであるため、RAGのようなユースケースには利点と欠点があります。</p></li><li><p>Playground は、AIホスティングの初期の時代にデファクトスタンダードになりつつあるOpenAIのようなREST APIを介してOllamaなどの推論ホスティングフレームワークに接続できます。</p></li></ul><p>全体として、ローカルの「エアギャップ」RAGがここまで進歩したことに感銘を受けました。Elasticsearch、Kibana、そして利用可能なオープンウェイトモデルのツールは、2023年に<a href="https://www.elastic.co/search-labs/blog/privacy-first-ai-search-langchain-elasticsearch">プライバシー重視のAI検索</a>について初めて記事を書いたときから大幅に進歩しました。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/deepseek-rag-ollama-playground</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/deepseek-rag-ollama-playground</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[Kibana]]></category>
    <dc:creator><![CDATA[Dave Erickson,Jakob Reiter]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2a4b2ae6bd97850b/6a17dfa4be6086558f00464c/1bd853bfdfa2710e44cc4c08dede6bd21b35c4b8-1542x860.png" length="0" type="image/png"/>
    <pubDate>Thu, 30 Jan 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[ファセット検索: AI を活用して検索範囲と結果を改善する]]></title>
    <description><![CDATA[Elasticsearch のファセット検索を使用して、カテゴリ内のオプションをすばやく絞り込む方法を説明します。]]></description>
    <content:encoded><![CDATA[<p>この記事では、人工知能 (AI)、特に GPT-4 などの高度な言語モデルを使用することで、よりコンテキストに沿ったファセットを作成し、ユーザーにとってさらに関連性と有用性を高めることができる方法について説明します。</p><p>ファセット検索は、電子商取引プラットフォームにおける強力なツールです。表示されるアイテムの特性に基づいて検索結果を整理および絞り込むのに役立ちます。フィルターと混同されることがよくありますが、ファセットの動作は異なります。フィルターは、製品のカテゴリや形式など、インデックスに常に存在する情報によって定義される固定属性です。一方、ファセットは動的であり、実行された検索によって返された結果から生成されます。</p><p>衣料品のカタログを想像してください。「カテゴリ」（例：T シャツ、ズボン）や「性別」（例：男性、女性）などのフィールドは、結果を絞り込むのに役立つフィルターです。ただし、ファセットは、一般的な色、利用可能なサイズ、素材など、結果に表示される製品の特定の特性を反映します。これにより、より適応性の高いコンテキストに応じた検索エクスペリエンスが可能になります。</p><p>以下は、ファセットを操作し、ファセットによってフィルタリングされた検索結果を確認できる画像です。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc4ce8ca7e5b62ad7/6a17f8ba6864a4534bb68979/74d2159706ab7248ebe5efddc74c882f0693db71-600x420.gif" alt="ファセット検索の例" /><h2>AIがファセット生成を改善する方法</h2><p>人工知能は、セマンティック検索や埋め込みと関連付けられることが多いですが、ファセットについてはどうでしょうか?AI を活用して、ファセットをより便利にし、各検索のコンテキストに特化したものにするにはどうすればよいでしょうか?</p><p>興味深い可能性の 1 つは、AI を使用して、インデックスの従来の分類を超えた新しい分類を作成することです。コンテンツの特定の特性を分析することにより、これらの新しいカテゴリは、より豊富で正確なコンテキスト化を提供でき、ファセットの関連性が高まり、ユーザーのニーズに合致したものになります。これにより、元のドキュメント カテゴリと比較して、結果をより意味のあるものに絞り込むことができます。</p><h2>AIが映画の分類を改善し、より良い検索を実現する方法</h2><p>現在ドラマのジャンルに分類されている以下の映画を分析してみましょう。</p><ul><li><p>レクイエム・フォー・ドリーム
概要: コニーアイランドに住む 4 人の薬物依存者のユートピアは、中毒が深刻化して崩壊します。</p></li><li><p>アメリカン・ビューティー
概要: 郊外に住む性的欲求不満の父親は、娘の親友に夢中になったことで中年の危機に陥る。</p></li><li><p>『グッド・ウィル・ハンティング』
履歴書: MIT の清掃員であるウィル・ハンティングは数学の才能に恵まれていますが、人生の方向性を見つけるために心理学者の助けが必要です。</p></li></ul><p>このジャンル分類では、各映画の微妙な違いや独自の背景を捉えることができません。AI を活用してあらすじや中心テーマを分析することで、各映画の真の文脈をよりよく反映する新しいカテゴリを作成できます。例えば：</p><ul><li><p>レクイエム・フォー・ドリーム - 新カテゴリー：「依存症と依存」</p></li><li><p>アメリカン・ビューティー - 新カテゴリー：「中年の危機」</p></li><li><p>『グッド・ウィル・ハンティング』 - 新カテゴリー：「知的闘争」</p></li></ul><p>これらの新しいカテゴリにより、検索の精度が大幅に向上し、結果を絞り込むためのより有意義なフィルターがユーザーに提供されます。このアプローチは、元のカテゴリが過度に一般的な場合に特に効果的であり、ユーザーが探しているものを簡単に見つけることができます。</p><h2>GPT-4 で新しいカテゴリを作成する: ファセット検索の例</h2><p>この例では、AI モデルを使用して、より正確で各作品のコンテキストに合わせた新しい映画カテゴリを作成する方法を示します。このプロセスをデモンストレーションするために、Elastic シミュレーション パイプラインを OpenAI 推論サービスと組み合わせて使用します。推論プロセッサで実行されるプロンプトを作成し、新しいカテゴリを決定できるスクリプト プロセッサを含む、複数のプロセッサを持つパイプラインが作成されます。他のプロセッサは、パイプライン実行中に生成されたデータおよび補助フィールドを操作するために使用されます。このロジックは他の同様のツールやモデルにも適用できることは言及する価値があります。</p><p>まず、推論エンドポイントを作成し、OpenAI としてのサービス、サービスにアクセスするために必要なトークン、およびモデルを定義する必要があります。この例では、gpt-4o-mini を使用しています。OpenAI 推論サービスの詳細については、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-openai.html">ここをクリック</a>してください。</p>PUT _inference/completion/generate_topics_ia
{
    "service": "openai",
    "service_settings": {
        "api_key": "your-token",
        "model_id": "gpt-4o-mini"
    }
}<p>エンドポイントが作成されたので、それを使用して新しいカテゴリを作成する準備が整いました。以下は、ドキュメント データの操作とプロンプト生成のプロセス全体を処理するパイプラインです。各プロセッサの機能について詳しく説明します。</p><p>最初のプロセッサはプロンプトの構築を担当します。AI がトピックを正しく分析して識別できるように、指示を明確に詳細に記述することが非常に重要です。このプロンプトでは、映画のタイトル、説明、ジャンルの分析に基づいて 2 つのトピックを要求しています。</p>{
        "script": {
          "source": """
            ctx.prompt = "You are an expert in semantic analysis and audiovisual content categorization. Your task is to generate only subcategories (max 2 topics) that describe specific aspects of movies based on their genres and descriptions. The output should be like: 'n1, n2, ...n'. Here is a movie info to analyze: Title: " + ctx.title  + "Genres: " + ctx.genres  + "Description: " + ctx.description;
          """
        }<p>次のパイプラインは推論パイプラインで、プロンプトを受信して<strong> generate_topics_ia</strong>エンドポイントに送信します。モデルによって生成された応答は結果フィールドに保存されます。</p>{
        "inference": {
          "model_id": "generate_topics_ia",
          "input_output": {
            "input_field": "prompt",
            "output_field": "result"
          }
        }
      }<p>次に、作成した一時フィールドを削除するだけでなく、応答を操作してトピック フィールドに設定するために使用される 3 つのプロセッサがあります。</p><p>このパイプラインを実行すると、以下の結果が得られます。</p>{
  "docs": [
    {
      "doc": {
        "_index": "index",
        "_version": "-3",
        "_id": "1",
        "_source": {
          "description": "While Frodo and Sam edge closer to Mordor with the help of the shifty Gollum, the divided fellowship makes a stand against Sauron's new ally, Saruman, and his hordes of Isengard.",
          "model_id": "generate_topics_ia",
          "title": "The Lord of the Rings: The Fellowship of the Ring",
          "genres": [
            "Action",
            "Adventure",
            "Drama"
          ],
          "topics": [
            "Fantasy",
            "Quest"
          ]
        },
        "_ingest": {
          "timestamp": "2024-11-22T17:51:51.340010257Z"
        }
      }
    },
    {
      "doc": {
        "_index": "index",
        "_version": "-3",
        "_id": "2",
        "_source": {
          "description": "A team of explorers travel through a wormhole in space in an attempt to ensure humanity's survival.",
          "model_id": "generate_topics_ia",
          "title": "Interstellar",
          "genres": [
            "Adventure",
            "Drama",
            "Sci-Fi"
          ],
          "topics": [
            "space exploration",
            "human survival"
          ]
        },
        "_ingest": {
          "timestamp": "2024-11-22T17:51:51.340413173Z"
        }
      }
    },
    {
      "doc": {
        "_index": "index",
        "_version": "-3",
        "_id": "3",
        "_source": {
          "description": "An astronaut becomes stranded on Mars after his team assume him dead, and must rely on his ingenuity to find a way to signal to Earth that he is alive.",
          "model_id": "generate_topics_ia",
          "title": "The Martian",
          "genres": [
            "Adventure",
            "Drama",
            "Sci-Fi"
          ],
          "topics": [
            "survival",
            "ingenuity"
          ]
        },
        "_ingest": {
          "timestamp": "2024-11-22T17:51:51.340427965Z"
        }
      }
    }
  ]
}<p>いくつかは元々同じジャンルですが、映画のコンテキストにさらに関連した新しいカテゴリがあることに注意してください。</p><p>これで、これらの新しいカテゴリを使用して、ドキュメントと一緒にインデックスを作成できるようになりました。このように、ファセットを生成する際に、主要なカテゴリに加えて、映画のコンテキストに合わせたより具体的なサブカテゴリが作成されます。</p><p>さらに、これらの新しいカテゴリをベクトル化し、ベクトル検索で使用することもできます。つまり、新しいカテゴリはフィルターとして機能するだけでなく、検索用語との意味上の類似性を計算するためにも使用できるため、表示される結果の関連性がさらに高まります。</p><p>完全なパイプライン:</p>POST /_ingest/pipeline/_simulate
{
  "pipeline": {
    "processors": [
      {
        "script": {
          "source": """
            ctx.prompt = "You are an expert in semantic analysis and audiovisual content categorization. Your task is to generate only subcategories (max 2 topics) that describe specific aspects of movies based on their genres and descriptions. The output should be like string: 'n1, n2m ...n'. Here is a movies info to analyze: Title: " + ctx.title  + "Genres: " + ctx.genres  + "Description: " + ctx.description;
          """
        }
      },
      {
        "inference": {
          "model_id": "generate_topics_ia",
          "input_output": {
            "input_field": "prompt",
            "output_field": "result"
          }
        }
      },
      {
        "split": {
          "field": "result",
          "target_field": "topics",
          "separator": ", "
        }
      },
      {
        "remove": {
          "field": "result"
        }
      },
      {
        "remove": {
          "field": "prompt"
        }
      }
    ]
  },
  "docs": [
    {
      "_index": "index",
      "_id": "1",
      "_source": {
        "title": "The Lord of the Rings: The Fellowship of the Ring",
        "description": "While Frodo and Sam edge closer to Mordor with the help of the shifty Gollum, the divided fellowship makes a stand against Sauron's new ally, Saruman, and his hordes of Isengard.",
        "genres": [
          "Action",
          "Adventure",
          "Drama"
        ]
      }
    },
    {
      "_index": "index",
      "_id": "2",
      "_source": {
        "title": "Interstellar",
        "description": "A team of explorers travel through a wormhole in space in an attempt to ensure humanity's survival.",
        "genres": [
          "Adventure", "Drama", "Sci-Fi"
        ]
      }
    },
    {
      "_index": "index",
      "_id": "3",
      "_source": {
        "title": "The Martian",
        "description": "An astronaut becomes stranded on Mars after his team assume him dead, and must rely on his ingenuity to find a way to signal to Earth that he is alive.",
        "genres": [
          "Adventure", "Drama", "Sci-Fi"
        ]
      }
    }
  ]
}<h2>まとめ</h2><p>AI を使用してファセットを改善すると、結果がより具体的かつ文脈的になり、検索エクスペリエンスを変革できます。多くの場合、範囲が広い固定カテゴリとは異なり、AI 生成のカテゴリはコンテキストをより適切に反映できます。たとえば、映画を再分類する場合、主要なカテゴリでは見逃されていたコンテキストを捉えて、より関連性の高いグループ化を提供できます。</p><p>これらの新しいカテゴリをインデックスに追加すると、ファセットが改善されるだけでなく、ベクター検索も可能になります。その結果、コンテキストに合わせたフィルターを使用した、より効率的な検索エクスペリエンスが実現します。</p><h2>参照資料</h2><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-openai.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-openai.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/simulate-pipeline-api.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/simulate-pipeline-api.html</a></p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/script-processor.html">https://www.elastic.co/guide/en/elasticsearch/reference/current/script-processor.html</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/faceted-search-examples-ai</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/faceted-search-examples-ai</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Andre Luiz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd2838185214162c8/6a17f8bedbb4ff04affb58a5/25c9f9baa2326b5189ce0b1cc6240475781c755d-721x421.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 28 Jan 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[AIエージェント開発におけるMicrosoftセマンティックカーネル向けElasticsearch Vector Store Connectorの使い方]]></title>
    <description><![CDATA[Microsoft Semantic Kernel は、AI エージェントを簡単に構築し、最新の AI モデルを C#、Python、または Java コードベースに統合できる軽量のオープンソース開発キットです。Semantic Kernel Elasticsearch Vector Store Connector のリリースにより、AI エージェントの構築に Semantic Kernel を使用する開発者は、Semantic Kernel の抽象化を引き続き使用しながら、Elasticsearch をスケーラブルなエンタープライズ グレードのベクター ストアとしてプラグインできるようになりました。]]></description>
    <content:encoded><![CDATA[<p><a href="https://learn.microsoft.com/en-us/semantic-kernel/overview/">Microsoft Semantic Kernel</a> チームと連携して、<a href="https://learn.microsoft.com/en-us/semantic-kernel/overview/"> Microsoft Semantic</a> Kernel (.NET) ユーザー向けに<a href="https://github.com/elastic/semantic-kernel-net/"> Semantic Kernel Elasticsearch Vector Store Connector が利用可能になったことを発表します。</a>セマンティック カーネルは、ベクター ストアからのより関連性の高いデータ駆動型の応答を使用して大規模言語モデル (LLM) を強化する機能など、エンタープライズ グレードの AI エージェントの構築を簡素化します。Semantic Kernel は、Elasticsearch などの Vector Stores と対話するためのシームレスな抽象化レイヤーを提供し、レコードのコレクションの作成、一覧表示、削除や、個々のレコードのアップロード、取得、削除などの重要な機能を提供します。</p><p><a href="https://learn.microsoft.com/en-us/semantic-kernel/concepts/vector-store-connectors/out-of-the-box-connectors/elasticsearch-connector?pivots=programming-language-csharp">すぐに使用できるセマンティック カーネル Elasticsearch ベクター ストア コネクタは、</a>セマンティック カーネル<a href="https://learn.microsoft.com/en-us/semantic-kernel/concepts/vector-store-connectors/?pivots=programming-language-csharp#the-vector-store-abstraction">ベクター ストアの抽象化</a>をサポートしており、開発者は AI エージェントの構築時に Elasticsearch をベクター ストアとしてプラグインすることが非常に簡単になります。</p><p>Elasticsearch はオープンソース コミュニティに強固な基盤を持ち、最近<a href="https://www.elastic.co/blog/elasticsearch-is-open-source-again">AGPL ライセンスを</a>採用しました。これらのツールは、オープンソースの Microsoft Semantic Kernel と組み合わせることで、強力なエンタープライズ対応ソリューションを提供します。このコマンド<code>curl -fsSL https://elastic.co/start-local | sh </code> (詳細については<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/run-elasticsearch-locally.html">start-local</a>を参照) を実行して数分で Elasticsearch を起動し、ローカルで開始できます。その後、AI エージェントを本番稼働させながら、<a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;utm_source=semantickernel&amp;utm_content=documentation">クラウドホスト バージョン</a>または<a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.16/install-elasticsearch.html">セルフホスト</a>バージョンに移行できます。</p><p>このブログでは、Semantic Kernel を使用する際に<a href="https://github.com/elastic/semantic-kernel-net/">Semantic Kernel Elasticsearch Vector Store Connector を</a>使用する方法について説明します。コネクタの Python バージョンは将来提供される予定です。</p><h2>高レベルのシナリオ: Semantic Kernel と Elasticsearch を使用した RAG アプリの構築</h2><p>次のセクションでは例を見ていきます。大まかに言うと、ユーザーの質問を入力として受け取り、回答を返す RAG (Retrieval Augmented Generation) アプリケーションを構築しています。LLM として Azure OpenAI (<a href="https://devblogs.microsoft.com/semantic-kernel/introducing-new-ollama-connector-for-local-models/">ローカル LLM</a>も使用可能)、ベクター ストアとして Elasticsearch、すべてのコンポーネントを結び付けるフレームワークとして Semantic Kernel (.net) を使用します。</p><p>RAG アーキテクチャに精通していない場合は、次の記事で簡単に概要を把握できます: <a href="https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag">https://www.elastic.co/search-labs/blog/retrieval-augmented-generation-rag</a> 。</p><p>回答は、Elasticsearch vectorstore から取得され、質問に関連するコンテキストが入力する LLM によって生成されます。応答には、LLM によってコンテキストとして使用されたソースも含まれます。</p><h3>RAGの例</h3><p>この具体的な例では、社内のホテル データベースに保存されているホテルについてユーザーが質問できるアプリケーションを構築します。ユーザーは例えばさまざまな基準に基づいて特定のホテルを検索したり、ホテルのリストを要求したりできます。</p><p>サンプル データベースでは、100 件のエントリを含む<a href="https://github.com/elastic/semantic-kernel-net/blob/main/Elastic.SemanticKernel.Playground/hotels.csv">ホテルのリスト</a>を生成しました。コネクタのデモをできるだけ簡単に試せるように、サンプル サイズは意図的に小さくなっています。実際のアプリケーションでは、特に非常に大量のデータを扱う場合、Elasticsearch コネクタは `InMemory` ベクトル ストア実装などの他のオプションよりも優位性を発揮します。</p><p>完全なデモ アプリケーションは、Elasticsearch ベクター ストア コネクタ<a href="https://github.com/elastic/semantic-kernel-net/tree/main/Elastic.SemanticKernel.Playground">リポジトリ</a>にあります。</p><p>まず、必要な NuGet パッケージと using ディレクティブをプロジェクトに追加することから始めましょう。</p>dotnet add package "Elastic.Clients.Elasticsearch" -v 8.16.2
dotnet add package "Elastic.SemanticKernel.Connectors.Elasticsearch" -v 0.1.2
dotnet add package "Microsoft.Extensions.Hosting" -v 9.0.0
dotnet add package "Microsoft.SemanticKernel.Connectors.AzureOpenAI" -v 1.30.0
dotnet add package "Microsoft.SemanticKernel.PromptTemplates.Handlebars" -v 1.30.0using System;
using System.IO;
using System.Linq;
using System.Threading.Tasks;

using Elastic.Clients.Elasticsearch;
using Elastic.Transport;

using Microsoft.Extensions.DependencyInjection;
using Microsoft.Extensions.Hosting;
using Microsoft.Extensions.VectorData;
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.Data;
using Microsoft.SemanticKernel.Embeddings;
using Microsoft.SemanticKernel.PromptTemplates.Handlebars;<p>これで、データ モデルを作成し、セマンティック カーネル固有の属性を指定して、ストレージ モデル スキーマとテキスト検索のヒントを定義できるようになりました。</p>/// &lt;summary&gt;
/// Data model for storing a "hotel" with a name, a description, a  description embedding and an optional reference link.
/// &lt;/summary&gt;
public sealed record Hotel
{
	[VectorStoreRecordKey]
	public required string HotelId { get; set; }

	[TextSearchResultName]
	[VectorStoreRecordData(IsFilterable = true)]
	public required string HotelName { get; set; }

	[TextSearchResultValue]
	[VectorStoreRecordData(IsFullTextSearchable = true)]
	public required string Description { get; set; }

	[VectorStoreRecordVector(Dimensions: 1536, DistanceFunction.CosineSimilarity, IndexKind.Hnsw)]
	public ReadOnlyMemory&lt;float&gt;? DescriptionEmbedding { get; set; }

	[TextSearchResultLink]
	[VectorStoreRecordData]
	public string? ReferenceLink { get; set; }
}<p>ストレージ モデル スキーマ属性 (`VectorStore*`) は、Elasticsearch Vector Store Connector の実際の使用に最も関連しています。具体的には次のようになります。</p><p></p><ul><li><p><code>VectorStoreRecordKey</code> レコード クラスのプロパティを、ベクトル ストアにレコードが格納されるキーとしてマークします。</p></li><li><p><code>VectorStoreRecordData</code> レコード クラスのプロパティを 'data' としてマークします。</p></li><li><p><code>VectorStoreRecordVector</code> レコード クラスのプロパティをベクトルとしてマークします。</p></li></ul><p>これらの属性はすべて、ストレージ モデルをさらにカスタマイズするために使用できるさまざまなオプション パラメーターを受け入れます。たとえば、 <code>VectorStoreRecordKey </code>の場合、異なる距離関数や異なるインデックス タイプを指定することが可能です。</p><p>テキスト検索属性 ( <code>TextSearch*</code> ) は、この例の最後のステップで重要になります。これらについては後ほど説明します。</p><p>次のステップでは、セマンティック カーネル エンジンを初期化し、コア サービスへの参照を取得します。実際のアプリケーションでは、サービス コレクションに直接アクセスするのではなく、<a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/dependency-injection">依存性注入を</a>使用する必要があります。同じことがハードコードされた構成とシークレットにも当てはまります。これらは、代わりに<a href="https://learn.microsoft.com/en-us/dotnet/core/extensions/configuration">構成プロバイダー</a>を使用して読み取る必要があります。</p>var builder = Host.CreateApplicationBuilder(args);

// Register AI services.
var kernelBuilder = builder.Services.AddKernel();

kernelBuilder.AddAzureOpenAIChatCompletion("gpt-4o", "https://my-service.openai.azure.com", "my_token");

kernelBuilder.AddAzureOpenAITextEmbeddingGeneration("ada-002", "https://my-service.openai.azure.com", "my_token");

// Register text search service.
kernelBuilder.AddVectorStoreTextSearch&lt;Hotel&gt;();

// Register Elasticsearch vector store.
var elasticsearchClientSettings = new ElasticsearchClientSettings(new Uri("https://my-elasticsearch-instance.cloud"))
    .Authentication(new BasicAuthentication("elastic", "my_password"));

kernelBuilder.AddElasticsearchVectorStoreRecordCollection&lt;string, Hotel&gt;("skhotels", elasticsearchClientSettings);

// Build the host.
using var host = builder.Build();

// For demo purposes, we access the services directly without using a DI context.

var kernel = host.Services.GetService&lt;Kernel&gt;()!;
var embeddings = host.Services.GetService&lt;ITextEmbeddingGenerationService&gt;()!;
var vectorStoreCollection = host.Services.GetService&lt;IVectorStoreRecordCollection&lt;string, Hotel&gt;&gt;()!;

// Register search plugin.
var textSearch = host.Services.GetService&lt;VectorStoreTextSearch&lt;Hotel&gt;&gt;()!;
kernel.Plugins.Add(textSearch.CreateWithGetTextSearchResults("SearchPlugin"));<p><code>vectorStoreCollection</code>サービスを使用してコレクションを作成し、いくつかの<a href="https://github.com/elastic/semantic-kernel-net/blob/main/Elastic.SemanticKernel.Playground/hotels.csv">デモ レコード</a>を取り込むことができるようになりました。</p>await vectorStoreCollection.CreateCollectionIfNotExistsAsync();

// CSV format: ID;Hotel Name;Description;Reference Link
var hotels = (await File.ReadAllLinesAsync("hotels.csv"))
    .Select(x =&gt; x.Split(';'));

foreach (var chunk in hotels.Chunk(25))
{
    var descriptionEmbeddings = await embeddings.GenerateEmbeddingsAsync(chunk.Select(x =&gt; x[2]).ToArray());
    
    for (var i = 0; i &lt; chunk.Length; ++i)
    {
        var hotel = chunk[i];
        await vectorStoreCollection.UpsertAsync(new Hotel
        {
            HotelId = hotel[0],
            HotelName = hotel[1],
            Description = hotel[2],
            DescriptionEmbedding = descriptionEmbeddings[i],
            ReferenceLink = hotel[3]
        });
    }
}<p>これは、セマンティック カーネルが、複雑なベクトル ストアの使用を、いくつかの単純なメソッド呼び出しにまで削減する方法を示しています。</p><p>内部的には、Elasticsearch に新しいインデックスが作成され、必要なすべてのプロパティ マッピングが作成されます。その後、データ セットは完全に透過的にストレージ モデルにマッピングされ、最終的にインデックスに保存されます。以下は Elasticsearch でのマッピングの様子です。</p>{
  "mappings": {
    "properties": {
      "descriptionEmbedding": {
        "dims": 1536,
        "index": true,
        "index_options": {
          "type": "hnsw"
        },
        "similarity": "cosine",
        "type": "dense_vector"
      },
      "hotelName": {
        "type": "keyword"
      },
      "description": {
        "type": "text"
      }
    }
  }
}<p><code>embeddings.GenerateEmbeddingsAsync()</code>は、構成された Azure AI Embeddings Generation サービスを透過的に呼び出しました。</p><p>このデモの最後のステップでは、さらに多くの魔法が観察できます。</p><p><code>InvokePromptAsync</code>を 1 回呼び出すだけで、ユーザーがデータについて質問したときに、次のすべての操作が実行されます。</p><p>1.ユーザーの質問の埋め込みが生成される</p><p>2. ベクトルストアで関連するエントリを検索する</p><p>3. クエリの結果はプロンプトテンプレートに挿入されます</p><p>4. 最終プロンプトの形式で実際のクエリがAIチャット補完サービスに送信されます。</p>// Invoke the LLM with a template that uses the search plugin to
// 1. get related information to the user query from the vector store
// 2. add the information to the LLM prompt.
var response = await kernel.InvokePromptAsync(
    promptTemplate: """
                    Please use this information to answer the question:
                    {{#with (SearchPlugin-GetTextSearchResults question)}}
                      {{#each this}}
                        Name: {{Name}}
                        Value: {{Value}}
                        Source: {{Link}}
                        -----------------
                      {{/each}}
                    {{/with}}
                    
                    Include the source of relevant information in the response.

                    Question: {{question}}
                    """,
    arguments: new KernelArguments
    {
        { "question", "Please show me all hotels that have a rooftop bar." },
    },
    templateFormat: "handlebars",
    promptTemplateFactory: new HandlebarsPromptTemplateFactory());<p>以前データ モデルで定義した<code>TextSearch*</code>属性を覚えていますか?これらの属性により、プロンプト テンプレート内の対応するプレースホルダーを使用できるようになります。これらのプレースホルダーには、ベクター ストア内のエントリからの情報が自動的に入力されます。</p><p>「屋上バーがあるホテルをすべて教えてください。」という質問に対する最終的な回答は次のとおりです。</p>Console.WriteLine(response.ToString());

// &gt; The hotel that has a rooftop bar is Skyline Suites. You can find more information about this hotel [here](https://example.com/yz567).<p>答えは、hotels.csvの次のエントリを正しく参照しています。</p>9;
Skyline Suites;
Offering panoramic city views from every suite, this hotel is perfect for those who love the urban landscape. Enjoy luxurious amenities, a rooftop bar, and close proximity to attractions. Luxurious and contemporary.;
https://example.com/yz567<p>この例は、Microsoft Semantic Kernel を使用すると、よく考えられた抽象化によって複雑さが大幅に軽減され、非常に高いレベルの柔軟性が実現されることを示しています。たとえば、コードの 1 行を変更するだけで、コードの他の部分をリファクタリングすることなく、使用されているベクトル ストアまたは AI サービスを置き換えることができます。</p><p>同時に、このフレームワークは、`InvokePrompt` 関数やテンプレート、検索プラグイン システムなどの膨大な高レベル機能を提供します。</p><p>完全なデモ アプリケーションは、Elasticsearch ベクター ストア コネクタ リポジトリにあります。</p><h2>Elasticsearchで他に何ができるのか</h2><ul><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">Elasticsearchの新しいsemantic_textマッピング：セマンティック検索の簡素化</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/semantic-reranking-with-retrievers">Elasticsearch におけるセマンティックリランキング（リトリーバー使用）</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">高度なRAGテクニックパート1：データ処理</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">高度なRAGテクニックパート2：クエリとテスト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-rag-with-llama3-opensource-and-elastic">Llama 3オープンソースとElasticでRAGを構築する</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/local-rag-agent-elasticsearch-langgraph-llama3">LangGraph、LLaMA3、Elasticsearchベクターストアを使用してローカルエージェントをゼロから構築するチュートリアル</a></p></li></ul><h2>Elasticsearch とセマンティックカーネル: 次は何?</h2><ul><li><p>.NET で GenAI アプリケーションを構築する際に、Elasticsearch ベクター ストアを Semantic Kernel に簡単にプラグインする方法を示しました。次回の Python 統合にご期待ください。</p></li><li><p>Semantic Kernel は<a href="https://www.elastic.co/search-labs/tutorials/search-tutorial/vector-search/hybrid-search">ハイブリッド検索</a>などの高度な検索機能の抽象化を構築するため、Elasticsearch Connect を使用すると、.NET 開発者は Semantic Kernel を使用しながらそれらを簡単に実装できるようになります。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-connector-microsoft-semantic-kernel</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-connector-microsoft-semantic-kernel</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[.NET]]></category>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Florian Bernd,Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2d8725035e86f8a8/6a17fe447f6f1564f8c09d74/0564fe794e4c66d0507317822d7aa71826183d20-1311x762.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 06 Dec 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearchを埋め込みストアとして使用したLangChain4j]]></title>
    <description><![CDATA[LangChain4j (LangChain for Java) には埋め込みストアとして Elasticsearch があります。これを使用して、プレーン Java で RAG アプリケーションを構築する方法を説明します。]]></description>
    <content:encoded><![CDATA[<p>
<a href="https://www.elastic.co/search-labs/blog/langchain4j-llm-integration-introduction">前回の投稿</a>では、LangChain4j とは何か、またどのように使用するかについて説明しました。</p><ul><li><p>LLMと<code>ChatLanguageModel</code>を実装して議論する <code>ChatMemory</code></p></li><li><p>チャット履歴をメモリに保持して、LLMとの以前の議論の文脈を思い出す</p></li></ul><p>このブログ投稿では、次の方法について説明します。</p><ul><li><p>テキスト例からベクトル埋め込みを作成する</p></li><li><p>ベクトル埋め込みをElasticsearch埋め込みストアに保存する </p></li><li><p>類似ベクトルを検索</p></li></ul><h2>埋め込みを作成する</h2><p>埋め込みを作成するには、使用する<code>EmbeddingModel</code>を定義する必要があります。たとえば、<a href="https://www.elastic.co/search-labs/blog/langchain4j-llm-integration-introduction">前回の投稿</a>で使用したのと同じミストラル モデルを使用できます。それは ollama で実行されていました:</p>EmbeddingModel model = OllamaEmbeddingModel.builder()
  .baseUrl(ollama.getEndpoint())
  .modelName(MODEL_NAME)
  .build();<p>モデルはテキストからベクトルを生成できます。ここで、モデルによって生成された次元の数を確認できます。</p>Logger.info("Embedding model has {} dimensions.", model.dimension());
// This gives: Embedding model has 4096 dimensions.<p>テキストからベクトルを生成するには、以下を使用します。</p>Response&lt;Embedding&gt; response = model.embed("A text here");<p>または、テキスト、価格、発売日などに基づいてフィルタリングできるようにメタデータも提供したい場合は、 <code>Metadata.from()</code>使用できます。たとえば、ここではゲーム名をメタデータ フィールドとして追加しています。</p>TextSegment game1 = TextSegment.from("""
    The game starts off with the main character Guybrush Threepwood stating "I want to be a pirate!"
    To do so, he must prove himself to three old pirate captains. During the perilous pirate trials, 
    he meets the beautiful governor Elaine Marley, with whom he falls in love, unaware that the ghost pirate 
    LeChuck also has his eyes on her. When Elaine is kidnapped, Guybrush procures crew and ship to track 
    LeChuck down, defeat him and rescue his love.
""", Metadata.from("gameName", "The Secret of Monkey Island"));
Response&lt;Embedding&gt; response1 = model.embed(game1);
TextSegment game2 = TextSegment.from("""
    Out Run is a pseudo-3D driving video game in which the player controls a Ferrari Testarossa 
    convertible from a third-person rear perspective. The camera is placed near the ground, simulating 
    a Ferrari driver's position and limiting the player's view into the distance. The road curves, 
    crests, and dips, which increases the challenge by obscuring upcoming obstacles such as traffic 
    that the player must avoid. The object of the game is to reach the finish line against a timer.
    The game world is divided into multiple stages that each end in a checkpoint, and reaching the end 
    of a stage provides more time. Near the end of each stage, the track forks to give the player a 
    choice of routes leading to five final destinations. The destinations represent different 
    difficulty levels and each conclude with their own ending scene, among them the Ferrari breaking 
    down or being presented a trophy.
""", Metadata.from("gameName", "Out Run"));
Response&lt;Embedding&gt; response2 = model.embed(game2);<p>このコードを実行する場合は、 <a href="https://github.com/dadoonet/langchain4j-demo/blob/main/src/test/java/fr/pilato/demo/Step5EmbedddingsTest.java">Step5EmbedddingsTest.java</a>クラスをチェックアウトしてください。</p><h2>ベクトルを保存するためにElasticsearchを追加する</h2><p>LangChain4j はメモリ内の埋め込みストアを提供します。これは簡単なテストを実行するのに便利です:</p>EmbeddingStore&lt;TextSegment&gt; embeddingStore = new InMemoryEmbeddingStore&lt;&gt;();
embeddingStore.add(response1.content(), game1);
embeddingStore.add(response2.content(), game2);<p>しかし、明らかに、このデータストアはすべてをメモリに保存し、サーバーに無限のメモリがないため、これよりもはるかに大きなデータセットでは機能しません。そのため、代わりに、定義上「弾性」があり、データに合わせてスケールアップおよびスケールアウトできる Elasticsearch に埋め込みを保存することができます。そのためには、Elasticsearch をプロジェクトに追加しましょう。</p>&lt;dependency&gt;
  &lt;groupId&gt;dev.langchain4j&lt;/groupId&gt;
  &lt;artifactId&gt;langchain4j-elasticsearch&lt;/artifactId&gt;
  &lt;version&gt;${langchain4j.version}&lt;/version&gt;
&lt;/dependency&gt;

&lt;dependency&gt;
  &lt;groupId&gt;org.testcontainers&lt;/groupId&gt;
  &lt;artifactId&gt;elasticsearch&lt;/artifactId&gt;
  &lt;version&gt;1.20.1&lt;/version&gt;
  &lt;scope&gt;test&lt;/scope&gt;
&lt;/dependency&gt;<p>お気づきのとおり、Elasticsearch TestContainers モジュールもプロジェクトに追加したので、テストから Elasticsearch インスタンスを起動できます。</p>// Create the elasticsearch container
ElasticsearchContainer container =
  new ElasticsearchContainer("docker.elastic.co/elasticsearch/elasticsearch:8.15.0")
    .withPassword("changeme");

// Start the container. This step might take some time...
container.start();

// As we don't want to make our TestContainers code more complex than
// needed, we will use login / password for authentication.
// But note that you can also use API keys which is preferred.
final CredentialsProvider credentialsProvider = new BasicCredentialsProvider();
credentialsProvider.setCredentials(AuthScope.ANY, new UsernamePasswordCredentials("elastic", "changeme"));

// Create a low level Rest client which connects to the elasticsearch container.
client = RestClient.builder(HttpHost.create("https://" + container.getHttpHostAddress()))
  .setHttpClientConfigCallback(httpClientBuilder -&gt; {
    httpClientBuilder.setDefaultCredentialsProvider(credentialsProvider);
    httpClientBuilder.setSSLContext(container.createSslContextFromCa());
    return httpClientBuilder;
  })
  .build();

// Check the cluster is running
client.performRequest(new Request("GET", "/"));<p>Elasticsearch を埋め込みストアとして使用するには、LangChain4j のインメモリ データストアから Elasticsearch データストアに切り替えるだけです。</p>EmbeddingStore&lt;TextSegment&gt; embeddingStore =
  ElasticsearchEmbeddingStore.builder()
    .restClient(client)
    .build();
embeddingStore.add(response1.content(), game1);
embeddingStore.add(response2.content(), game2);<p>これにより、Elasticsearch の<code>default</code>インデックスにベクトルが保存されます。インデックス名をより意味のあるものに変更することもできます。</p>EmbeddingStore&lt;TextSegment&gt; embeddingStore =
  ElasticsearchEmbeddingStore.builder()
    .indexName("games")
    .restClient(client)
    .build();
embeddingStore.add(response1.content(), game1);
embeddingStore.add(response2.content(), game2);<p>このコードを実行する場合は、 <a href="https://github.com/dadoonet/langchain4j-demo/blob/main/src/test/java/fr/pilato/demo/Step6ElasticsearchEmbedddingsTest.java">Step6ElasticsearchEmbedddingsTest.java</a>クラスをチェックアウトしてください。</p><h2>類似ベクトルを検索</h2><p>類似のベクトルを検索するには、まず、以前使用したのと同じモデルを使用して、質問をベクトル表現に変換する必要があります。すでにそれを行ったので、もう一度これを行うのは難しくありません。この場合、メタデータは必要ないことに注意してください。</p>String question = "I want to pilot a car";
Embedding questionAsVector = model.embed(question).content();<p>質問をこのように表現して検索リクエストを作成し、埋め込みストアに最初の上位ベクトルを見つけるように依頼することができます。</p>EmbeddingSearchResult&lt;TextSegment&gt; result = embeddingStore.search(
  EmbeddingSearchRequest.builder()
    .queryEmbedding(questionAsVector)
    .build());<p>これで、結果を反復処理して、メタデータから取得したゲーム名やスコアなどの情報を出力できます。</p>result.matches().forEach(m -&gt; Logger.info("{} - score [{}]",
  m.embedded().metadata().getString("gameName"), m.score()));<p>予想どおり、最初のヒットは「Out Run」になります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7ca0dcfdb1a9c94f/6a170291cf4f256938b2d017/140b6a962e5edbb4870419250e30bfb815b0d73e-640x480.gif" alt="アウトラン" />Out Run - score [0.86672974]
The Secret of Monkey Island - score [0.85569763]<p>このコードを実行する場合は、 <a href="https://github.com/dadoonet/langchain4j-demo/blob/9ec4b1d4c7c69821f143ddf272bbfed273c67b14/src/test/java/fr/pilato/demo/Step7SearchForVectorsTest.java#L110-L129">Step7SearchForVectorsTest.java</a>クラスをチェックアウトしてください。 </p><h2>舞台裏</h2><p>Elasticsearch Embedding ストアのデフォルト構成では、バックグラウンドで<a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.15/query-dsl-knn-query.html">近似 kNN クエリが</a>使用されます。</p>POST games/_search
{
  "query" : {
    "knn": {
      "field": "vector",
      "query_vector": [-0.019137882, /* ... */, -0.0148779955]
    }
  }
}<p>ただし、埋め込みストアにデフォルトの構成 ( <code>ElasticsearchConfigurationKnn</code> ) とは別の構成 ( <code>ElasticsearchConfigurationScript</code> ) を提供することでこれを変更できます。</p>EmbeddingStore&lt;TextSegment&gt; embeddingStore =
  ElasticsearchEmbeddingStore.builder()
    .configuration(ElasticsearchConfigurationScript.builder().build())
    .indexName("games")
    .restClient(client)
    .build();<p><code>ElasticsearchConfigurationScript</code>実装は、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.15/query-dsl-script-score-query.html"><code>script_score</code></a><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.15/query-dsl-script-score-query.html"> </a><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.15/query-dsl-script-score-query.html#vector-functions-cosine"><code>cosineSimilarity</code></a><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.15/query-dsl-script-score-query.html#vector-functions-cosine">関数</a> を使用して、バックグラウンドで クエリ を実行します。</p><p>基本的に、電話をかけるときは:</p>EmbeddingSearchResult&lt;TextSegment&gt; result = embeddingStore.search(
  EmbeddingSearchRequest.builder()
    .queryEmbedding(questionAsVector)
    .build());<p>これにより、次のコードが呼び出されます。</p>POST games/_search
{
  "query": {
    "script_score": {
      "script": {
        "source": "(cosineSimilarity(params.query_vector, 'vector') + 1.0) / 2",
        "params": {
          "queryVector": [-0.019137882, /* ... */, -0.0148779955]
        }
      }
    }
  }
}<p>この場合、結果は「順序」の点では変化せず、スコアのみが調整されます。これは、 <code>cosineSimilarity</code>呼び出しでは近似値を使用せず、一致するベクトルごとにコサインを計算するためです。</p>Out Run - score [0.871952]
The Secret of Monkey Island - score [0.86380446]<p>このコードを実行する場合は、 <a href="https://github.com/dadoonet/langchain4j-demo/blob/9ec4b1d4c7c69821f143ddf272bbfed273c67b14/src/test/java/fr/pilato/demo/Step7SearchForVectorsTest.java#L132-L155">Step7SearchForVectorsTest.java</a>クラスをチェックアウトしてください。</p><h2>まとめ</h2><p>テキストから埋め込みを簡単に生成する方法と、2 つの異なるアプローチを使用して Elasticsearch で最も近い近傍を保存および検索する方法について説明しました。</p><ul><li><p>デフォルトの<code>ElasticsearchConfigurationKnn</code>オプションを使用して、近似高速<code>knn</code>クエリを使用する</p></li><li><p>正確だが遅い<code>script_score</code>クエリを<code>ElasticsearchConfigurationScript</code>オプションとともに使用する</p></li></ul><p>次のステップでは、ここで学んだことを基に、完全な RAG アプリケーションを構築します。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/langchain4j-elasticsearch-embedding-store</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/langchain4j-elasticsearch-embedding-store</guid>
    <category><![CDATA[Java]]></category>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[David Pilato]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc873b86c76d1798/6a170293acf088f666be99b3/abd8a4a809064101c037af66b87f28e5ecde03b0-1474x645.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 08 Oct 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[LangChain4j を導入して、LLM の Java アプリケーションへの統合を簡素化します。]]></title>
    <description><![CDATA[LangChain4j (LangChain for Java) は、プレーン Java で RAG アプリケーションを構築するための強力なツールセットです。]]></description>
    <content:encoded><![CDATA[<p><a href="https://docs.langchain4j.dev/">LangChain4j フレームワークは</a>、<a href="https://github.com/langchain4j/langchain4j/blob/main/README.md#introduction">次の目標</a>を掲げて 2023 年に作成されました。</p>LangChain4j の目標は、LLM を Java アプリケーションに簡単に統合することです。<p>LangChain4j は次の標準的な方法を提供します:</p><ul><li><p>与えられたコンテンツ（例えばテキスト）から埋め込み（ベクトル）を作成する</p></li><li><p>埋め込みを埋め込みストアに保存する</p></li><li><p>埋め込みストア内の類似ベクトルを検索する</p></li><li><p>LLMと議論する</p></li><li><p>チャットメモリを使用して、LLMとの議論の文脈を記憶する</p></li></ul><p>このリストは網羅的なものではなく、LangChain4j コミュニティは常に新しい機能を実装しています。</p><p>この投稿では、フレームワークの最初の主要部分について説明します。</p><h2>プロジェクトにLangChain4j OpenAIを追加する</h2><p>すべての Java プロジェクトと同様に、これは依存関係の問題にすぎません。ここでは Maven を使用しますが、他の依存関係マネージャーでも同じことを実現できます。</p><p>ここで構築するプロジェクトの最初のステップとして、OpenAI を使用するので、 <code>langchain4j-open-ai</code>アーティファクトを追加するだけです。</p>&lt;properties&gt;
  &lt;langchain4j.version&gt;0.34.0&lt;/langchain4j.version&gt;
&lt;/properties&gt;

&lt;dependencies&gt;
  &lt;dependency&gt;
    &lt;groupId&gt;dev.langchain4j&lt;/groupId&gt;
    &lt;artifactId&gt;langchain4j-open-ai&lt;/artifactId&gt;
    &lt;version&gt;${langchain4j.version}&lt;/version&gt;
  &lt;/dependency&gt;
&lt;/dependencies&gt;
<p>残りのコードでは、 <a href="https://platform.openai.com/signup/">OpenAI</a>にアカウントを登録することで取得できる独自の API キー、またはデモ目的でのみ LangChain4j プロジェクトによって提供される API キーを使用します。</p>static String getOpenAiApiKey() {
  String apiKey = System.getenv(API_KEY_ENV_NAME);
  if (apiKey == null || apiKey.isEmpty()) {
    Logger.warn("Please provide your own key instead using [{}] env variable", API_KEY_ENV_NAME);
    return "demo";
  }
  return apiKey;
}
<p>これで ChatLanguageModel のインスタンスを作成できます。</p>ChatLanguageModel model = OpenAiChatModel.withApiKey(getOpenAiApiKey());
<p>そして最後に、簡単な質問をして答えを得ることができます。</p>String answer = model.generate("Who is Thomas Pesquet?");
Logger.info("Answer is: {}", answer);
<p>与えられた答えは次のようになります:</p>Thomas Pesquet is a French aerospace engineer, pilot, and European Space Agency astronaut.
He was selected as a member of the European Astronaut Corps in 2009 and has since completed 
two space missions to the International Space Station, including serving as a flight engineer 
for Expedition 50/51 in 2016-2017. Pesquet is known for his contributions to scientific 
research and outreach activities during his time in space.
<p>このコードを実行する場合は、 <a href="https://github.com/dadoonet/langchain4j-demo/blob/main/src/test/java/fr/pilato/demo/Step1AiChatTest.java">Step1AiChatTest.java</a>クラスを確認してください。</p><h2>langchain4jでより多くのコンテキストを提供する</h2><p><code>langchain4j</code>アーティファクトを追加しましょう:</p>&lt;dependency&gt;
  &lt;groupId&gt;dev.langchain4j&lt;/groupId&gt;
  &lt;artifactId&gt;langchain4j&lt;/artifactId&gt;
  &lt;version&gt;${langchain4j.version}&lt;/version&gt;
&lt;/dependency&gt;
<p>これは、アシスタントを構築するために、より高度な LLM 統合を構築するのに役立つツールセットを提供します。ここでは、先ほど定義した<code>ChatLanguageModel</code>を自動的に呼び出す<code>chat</code>メソッドを提供する<code>Assistant</code>インターフェイスを作成します。</p>interface Assistant {
  String chat(String userMessage);
}
<p>LangChain4j <code>AiServices</code>クラスにインスタンスを構築するよう依頼するだけです。</p>Assistant assistant = AiServices.create(Assistant.class, model);
<p>次に、 <code>chat(String)</code>メソッドを呼び出します。</p>String answer = assistant.chat("Who is Thomas Pesquet?");
Logger.info("Answer is: {}", answer);
<p>これは以前と同じ動作になります。では、なぜコードを変更したのでしょうか?まず第一に、よりエレガントですが、それ以上に、シンプルな注釈を使用して LLM にいくつかの指示を与えることができるようになりました。</p>interface Assistant {
  @SystemMessage("Please answer in a funny way.")
  String chat(String userMessage);
}
<p>これにより、次のようになります:</p>Ah, Thomas Pesquet is actually a super secret spy disguised as an astronaut! 
He's out there in space fighting aliens and saving the world one spacewalk at a time. 
Or maybe he's just a really cool French astronaut who has been to the International 
Space Station. But my spy theory is much more exciting, don't you think?
<p>このコードを実行する場合は、 <a href="https://github.com/dadoonet/langchain4j-demo/blob/main/src/test/java/fr/pilato/demo/Step2AssistantTest.java">Step2AssistantTest.java</a>クラスを確認してください。</p><h2>別のLLMへの切り替え：langchain4j-ollama</h2><p>素晴らしい<a href="https://ollama.com/">Ollamaプロジェクト</a>を利用できます。LLM を自分のマシン上でローカルに実行するのに役立ちます。</p><p><code>langchain4j-ollama</code>アーティファクトを追加しましょう:</p>&lt;dependency&gt;
  &lt;groupId&gt;dev.langchain4j&lt;/groupId&gt;
  &lt;artifactId&gt;langchain4j-ollama&lt;/artifactId&gt;
  &lt;version&gt;${langchain4j.version}&lt;/version&gt;
&lt;/dependency&gt;
<p>テストを使用してサンプル コードを実行しているので、プロジェクトに<a href="https://java.testcontainers.org/">Testcontainers を</a>追加しましょう。</p>&lt;dependency&gt;
  &lt;groupId&gt;org.testcontainers&lt;/groupId&gt;
  &lt;artifactId&gt;ollama&lt;/artifactId&gt;
  &lt;version&gt;1.20.1&lt;/version&gt;
  &lt;scope&gt;test&lt;/scope&gt;
&lt;/dependency&gt;
<p>Docker コンテナを起動/停止できるようになりました。</p>static String MODEL_NAME = "mistral";
static String DOCKER_IMAGE_NAME = "langchain4j/ollama-" + MODEL_NAME + ":latest";

static OllamaContainer ollama = new OllamaContainer(
  DockerImageName.parse(DOCKER_IMAGE_NAME).asCompatibleSubstituteFor("ollama/ollama"));

@BeforeAll
public static void setup() {
  ollama.start();
}

@AfterAll
public static void teardown() {
  ollama.stop();
}
<p>先ほど使用した<code>OpenAiChatModel</code>ではなく、 <code>model</code>オブジェクトを<code>OllamaChatModel</code>に変更するだけです。</p>OllamaChatModel model = OllamaChatModel.builder()
  .baseUrl(ollama.getEndpoint())
  .modelName(MODEL_NAME)
  .build();
<p>モデルを含むイメージを取得するのに多少時間がかかる場合がありますが、しばらくすると答えが得られることに注意してください。</p>Oh, Thomas Pesquet, the man who single-handedly keeps the French space program running 
while sipping on his crisp rosé and munching on a baguette! He's our beloved astronaut 
with an irresistible accent that makes us all want to learn French just so we can 
understand him better. When he's not floating in space, he's probably practicing his 
best "je ne sais quoi" face for the next family photo. Vive le Thomas Pesquet! 
🚀🌍🇫🇷 #FrenchSpaceHero
<h2>メモリがあればさらに良くなる</h2><p>複数の質問をした場合、デフォルトではシステムは以前の質問と回答を記憶しません。最初の質問の後に「彼はいつ生まれたのですか？」と尋ねると、私たちのアプリケーションは次のように答えます:</p>Oh, you're asking about this legendary figure from history, huh? Well, let me tell 
you a hilarious tale! He was actually born on Leap Year's Day, but only every 400 
years! So, do the math... if we count backwards from 2020 (which is also a leap year), 
then he was born in... *drumroll please* ...1600! Isn't that a hoot? But remember 
folks, this is just a joke, and historical records may vary.
<p>それはナンセンスだ。代わりに、<a href="https://docs.langchain4j.dev/tutorials/chat-memory">チャットメモリを</a>使用する必要があります。</p>ChatMemory chatMemory = MessageWindowChatMemory.withMaxMessages(10);
Assistant assistant = AiServices.builder(Assistant.class)
  .chatLanguageModel(model)
  .chatMemory(chatMemory)
  .build();
<p>同じ質問を実行すると、意味のある答えが得られます。</p>Oh, Thomas Pesquet, the man who was probably born before sliced bread but after dinosaurs! 
You know, around the time when people started putting wheels on suitcases and calling it 
a revolution. So, roughly speaking, he came into this world somewhere in the late 70s or 
early 80s, give or take a year or two - just enough time for him to grow up, become an 
astronaut, and make us all laugh with his space-aged antics! Isn't that a hoot? 
*laughs maniacally*
<h2>まとめ</h2><p><a href="https://www.elastic.co/search-labs/blog/langchain4j-elasticsearch-embedding-store">次の投稿</a>では、Elasticsearch を埋め込みストアとして使用して、プライベート データセットに質問する方法を説明します。これにより、アプリケーション検索を次のレベルに引き上げる方法が得られます。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/langchain4j-llm-integration-introduction</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/langchain4j-llm-integration-introduction</guid>
    <category><![CDATA[Java]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[David Pilato]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0435ed6d14579089/6a17e79ae8fbce7e433a192f/cf129b8b25fbe7204e2adca8fca5fec04207f096-720x720.png" length="0" type="image/png"/>
    <pubDate>Mon, 23 Sep 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[高度なRAGテクニックパート2：クエリとテスト]]></title>
    <description><![CDATA[RAG のパフォーマンスを向上させる可能性のあるテクニックについて議論し、実装します。パート 2/2。高度な RAG パイプラインのクエリとテストに焦点を当てます。]]></description>
    <content:encoded><![CDATA[<p><em>すべてのコードは</em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em> 、Searchlabs リポジトリの advanced-rag-techniques</em></a><em> ブランチに あります 。</em></p><p>高度な RAG テクニックに関する記事のパート 2 へようこそ。<a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">このシリーズのパート 1</a>では、高度な RAG パイプラインのデータ処理コンポーネントを設定、説明、実装しました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="高度なRAGパイプライン" /><p>この部分では、実装のクエリとテストを進めていきます。早速始めましょう！</p><h3>目次</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#searching-and-retrieving,-generating-answers">検索と取得、回答の生成</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#enriching-queries-with-synonyms">同義語によるクエリの強化</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-hypothetical-document-embedding">HyDE（仮想文書埋め込み）</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hybrid-search">ハイブリッド検索</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#experiments">実験</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#summary-of-results">結果の要約</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-1-who-audits-elastic">テスト 1: Elastic を監査するのは誰ですか?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-2--total-revenue-2023">テスト2：2023年の総収入</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-1">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-1">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-3-what-product-does-growth-primarily-depend-on-how-much">テスト 3: 成長は主にどの製品に依存しますか?いくら？</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-2">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-2">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-4-describe-employee-benefit-plan">テスト4: 従業員福利厚生制度の説明</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-3">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-3">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-5-which-companies-did-elastic-acquire">テスト 5: Elastic が買収した企業はどれですか?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-4">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-4">シンプルラグ</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#conclusion">まとめ</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#appendix">付記</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#prompts">プロンプト</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#rag-question-answering-prompt">RAG質問回答プロンプト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#elastic-query-generator-prompt">弾性クエリジェネレータプロンプト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#potential-questions-generator-prompt">潜在的な質問ジェネレータプロンプト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-generator-prompt">HyDEジェネレータプロンプト</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#sample-hybrid-search-query">ハイブリッド検索クエリのサンプル</a></p></li></ul></li></ul><h2>検索と取得、回答の生成</h2><p>最初のクエリ、理想的には主に年次報告書に記載されている情報を尋ねてみましょう。いかがでしょうか:</p>Who audits Elastic?"
<p>ここで、クエリを強化するためにいくつかのテクニックを適用してみましょう。</p><h3>同義語によるクエリの強化</h3><p>まず、クエリの文言の多様性を高めて、Elasticsearch クエリに簡単に処理できる形式に変えてみましょう。GPT-4o の助けを借りて、クエリを OR 句のリストに変換します。次のプロンプトを書いてみましょう:</p>
ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<p>GPT-4o をクエリに適用すると、基本クエリの同義語と関連語彙が生成されます。</p>'audits elastic OR 
elasticsearch audits OR 
elastic auditor OR 
elasticsearch auditor OR 
elastic audit firm OR 
elastic audit company OR 
elastic audit organization OR 
elastic audit service'
<p><code>ESQueryMaker</code>クラスでは、クエリを分割する関数を定義しました。</p>def parse_or_query(self, query_text: str) -&gt; List[str]:
    # Split the query by 'OR' and strip whitespace from each term
    # This converts a string like "term1 OR term2 OR term3" into a list ["term1", "term2", "term3"]
    return [term.strip() for term in query_text.split(' OR ')]
<p>その役割は、この OR 句の文字列を取得して用語のリストに分割し、主要なドキュメント フィールドで複数の一致を実行できるようにすることです。</p>["original_text", 'keyphrases', 'potential_questions', 'entities']
<p>最終的にこのクエリに至ります:</p> 'query': {
    'bool': {
        'must': [
            {
                'multi_match': {
                'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
                'fields': [
                    'original_text',
                'keyphrases',
                'potential_questions',
                'entities'
                ],
                'type': 'best_fields',
                'operator': 'or'
                }
            }
      ]
<p>これにより、元のクエリよりも多くのベースがカバーされ、同義語を忘れたために検索結果を見逃すリスクが軽減されることが期待されます。しかし、私たちにはもっとできることがある。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h3>HyDE（仮想文書埋め込み）</h3><p>今回は<a href="https://arxiv.org/abs/2212.10496">HyDE を</a>実装するために、再び GPT-4o を活用しましょう。</p><p>HyDE の基本的な前提は、仮想ドキュメント (元のクエリに対する回答が含まれる可能性のある種類のドキュメント) を生成することです。文書の事実性や正確性は問題ではありません。それを念頭に置いて、次のプロンプトを書いてみましょう。</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<p>ベクトル検索は通常、コサインベクトルの類似度に基づいて行われるため、クエリをドキュメントに一致させるのではなく、ドキュメントをドキュメントに一致させることでより良い結果を達成できるというのが HyDE の前提です。</p><p>私たちが重視するのは、構造、フロー、用語です。あまり事実ではない。GPT-4o は次のような HyDE ドキュメントを出力します。</p>'Elastic N.V., the parent company of Elastic, the organization known for developing Elasticsearch, is subject to audits to ensure financial accuracy, 
regulatory compliance, and the integrity of its financial statements. The auditing of Elastic N.V. is typically conducted by an external, 
independent auditing firm. This is common practice for publicly traded companies to provide stakeholders with assurance regarding the company\'s 
financial position and operations.\n\nThe primary external auditor for Elastic is the audit firm Ernst &amp; Young LLP (EY). Ernst &amp; Young is one of the 
four largest professional services networks in the world, commonly referred to as the "Big Four" audit firms. These firms handle a substantial number 
of audits for major corporations around the globe, ensuring adherence to generally accepted accounting principles (GAAP) and international financial 
reporting standards (IFRS).\n\nThe audit process conducted by EY involves several steps. Initially, the auditors perform a risk assessment to identify 
areas where misstatements due to error or fraud could occur. They then design audit procedures to test the accuracy and completeness of financial statements,
 which include examining financial transactions, assessing internal controls, and reviewing compliance with relevant laws and regulations. Upon completion of 
 the audit, Ernst &amp; Young issues an audit report, which includes the auditor’s opinion on whether the financial statements are free from material misstatement 
 and are presented fairly in accordance with the applicable financial reporting framework.\n\nIn addition to external audits by firms like Ernst &amp; Young, 
 Elastic may also be subject to internal audits. Internal audits are performed by the company’s own internal auditors to evaluate the effectiveness of internal 
 controls, risk management, and governance processes.\n\nOverall, the auditing process plays a crucial role in maintaining the transparency and reliability of 
 Elastic\'s financial information, providing confidence to investors, regulators, and other stakeholders.'
<p>これはかなり信憑性があり、インデックスを作成したい種類のドキュメントに最適な候補のように見えます。これを埋め込み、ハイブリッド検索に使用します。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h3>ハイブリッド検索</h3><p>これが私たちの検索ロジックの中核です。語彙検索コンポーネントは、生成された OR 句の文字列になります。高密度ベクトル コンポーネントには、HyDE ドキュメント (検索ベクトルとも呼ばれます) が埋め込まれます。KNN を使用して、検索ベクトルに最も近い候補ドキュメントをいくつか効率的に識別します。デフォルトでは、語彙検索コンポーネントを<em>TF-IDF と BM25 によるスコアリング</em>と呼びます。最後に、語彙スコアと密なベクトルスコアは、 <a href="https://arxiv.org/abs/2407.01219">Wang ら</a>が推奨する 30/70 比率を使用して結合されます。</p>def hybrid_vector_search(self, index_name: str, query_text: str, query_vector: List[float], 
                         text_fields: List[str], vector_field: str, 
                         num_candidates: int = 100, num_results: int = 10) -&gt; Dict:
    """
    Perform a hybrid search combining text-based and vector-based similarity.

    Args:
        index_name (str): The name of the Elasticsearch index to search.
        query_text (str): The text query string, which may contain 'OR' separated terms.
        query_vector (List[float]): The query vector for semantic similarity search.
        text_fields (List[str]): List of text fields to search in the index.
        vector_field (str): The name of the field containing document vectors.
        num_candidates (int): Number of candidates to consider in the initial KNN search.
        num_results (int): Number of final results to return.

    Returns:
        Dict: A tuple containing the Elasticsearch response and the search body used.
    """
    try:
        # Parse the query_text into a list of individual search terms
        # This splits terms separated by 'OR' and removes any leading/trailing whitespace
        query_terms = self.parse_or_query(query_text)

        # Construct the search body for Elasticsearch
        search_body = {
            # KNN search component for vector similarity
            "knn": {
                "field": vector_field,  # The field containing document vectors
                "query_vector": query_vector,  # The query vector to compare against
                "k": num_candidates,  # Number of nearest neighbors to retrieve
                "num_candidates": num_candidates  # Number of candidates to consider in the KNN search
            },
            "query": {
                "bool": {
                    # The 'must' clause ensures that matching documents must satisfy this condition
                    # Documents that don't match this clause are excluded from the results
                    "must": [
                        {
                            # Multi-match query to search across multiple text fields
                            "multi_match": {
                                "query": " ".join(query_terms),  # Join all query terms into a single space-separated string
                                "fields": text_fields,  # List of fields to search in
                                "type": "best_fields",  # Use the best matching field for scoring
                                "operator": "or"  # Match any of the terms (equivalent to the original OR query)
                            }
                        }
                    ],
                    # The 'should' clause boosts relevance but doesn't exclude documents
                    # It's used here to combine vector similarity with text relevance
                    "should": [
                        {
                            # Custom scoring using a script to combine vector and text scores
                            "script_score": {
                                "query": {"match_all": {}},  # Apply this scoring to all documents that matched the 'must' clause
                                "script": {
                                    # Script to combine vector similarity and text relevance
                                    "source": """
                                    # Calculate vector similarity (cosine similarity + 1)
                                    # Adding 1 ensures the score is always positive
                                    double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;
                                    # Get the text-based relevance score from the multi_match query
                                    double text_score = _score;
                                    # Combine scores: 70% vector similarity, 30% text relevance
                                    # This weighting can be adjusted based on the importance of semantic vs keyword matching
                                    return 0.7 * vector_score + 0.3 * text_score;
                                    """,
                                    # Parameters passed to the script
                                    "params": {
                                        "query_vector": query_vector,  # Query vector for similarity calculation
                                        "vector_field": vector_field  # Field containing document vectors
                                    }
                                }
                            }
                        }
                    ]
                }
            }
        }

        # Execute the search request against the Elasticsearch index
        response = self.conn.search(index=index_name, body=search_body, size=num_results)
        # Log the successful execution of the search for monitoring and debugging
        logger.info(f"Hybrid search executed on index: {index_name} with text query: {query_text}")
        # Return both the response and the search body (useful for debugging and result analysis)
        return response, search_body
    except Exception as e:
        # Log any errors that occur during the search process
        logger.error(f"Error executing hybrid search on index: {index_name}. Error: {e}")
        # Re-raise the exception for further handling in the calling code
        raise e
<p>最後に、RAG 関数を組み立てることができます。クエリから回答までの RAG の流れは次のようになります。</p><ol><li><p>クエリを OR 句に変換します。</p></li><li><p>HyDE ドキュメントを生成して埋め込みます。</p></li><li><p>両方をハイブリッド検索への入力として渡します。</p></li><li><p>上位 n 件の結果を取得し、最も関連性の高いスコアが LLM のコンテキスト メモリ内で「最新」になるように結果を逆にします (逆パッキング)。逆パッキングの例: クエリ:「Elasticsearch クエリ最適化手法」取得されたドキュメント (関連性の高い順): LLM コンテキストの順序を逆にする: 順序を逆にすることで、最も関連性の高い情報 (1) がコンテキストの最後に表示され、回答生成中に LLM からより多くの注目を受ける可能性が高くなります。</p><ol><li><p>「ブールクエリを使用して、複数の検索条件を効率的に組み合わせます。」</p></li><li><p>「クエリの応答時間を改善するためのキャッシュ戦略を実装します。」</p></li><li><p>「インデックス マッピングを最適化して、検索パフォーマンスを高速化します。」</p></li><li><p>「インデックス マッピングを最適化して、検索パフォーマンスを高速化します。」</p></li><li><p>「クエリの応答時間を改善するためのキャッシュ戦略を実装します。」</p></li><li><p>「ブールクエリを使用して、複数の検索条件を効率的に組み合わせます。」</p></li></ol></li><li><p>生成のためにコンテキストを LLM に渡します。</p></li></ol>def get_context(index_name, 
                match_query, 
                text_query, 
                fields, 
                num_candidates=100, 
                num_results=20, 
                text_fields=["original_text", 'keyphrases', 'potential_questions', 'entities'], 
                embedding_field="primary_embedding"):

    embedding=embedder.get_embeddings_from_text(text_query)

    results, search_body = es_query_maker.hybrid_vector_search(
        index_name=index_name,
        query_text=match_query,
        query_vector=embedding[0][0],
        text_fields=text_fields,
        vector_field=embedding_field,
        num_candidates=num_candidates,
        num_results=num_results
    )

    # Concatenates the text in each 'field' key of the search result objects into a single block of text.
    context_docs=['\n\n'.join([field+":\n\n"+j['_source'][field] for field in fields]) for j in results['hits']['hits']]

    # Reverse Packing to ensure that the highest ranking document is seen first by the LLM.
    context_docs.reverse()
    return context_docs, search_body

def retrieval_augmented_generation(query_text):
    match_query= gpt4o.generate_query(query_text)
    fields=['original_text']

    hyde_document=gpt4o.generate_HyDE(query_text)

    context, search_body=get_context(index_name, match_query, hyde_document, fields)

    answer= gpt4o.basic_qa(query=query_text, context=context)
    return answer, match_query, hyde_document, context, search_body

<p>クエリを実行して回答を取得してみましょう。</p>According to the context, Elastic N.V. is audited by an independent registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of independent registered public accounting firm," which states:

"We have audited the accompanying consolidated balance sheets of Elastic N.V. [...] / s / pricewaterhouseco."
<p>ニース。そうです。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h2>実験</h2><p>今答えなければならない重要な質問があります。これらの実装に多大な労力と追加の複雑さを投資することで、何が得られましたか?</p><p>少し比較してみましょう。私たちが実装した RAG パイプラインと、私たちが行った機能強化のないベースライン ハイブリッド検索を比較したものです。小規模な一連のテストを実行して、大きな違いが見られるか確認します。ここで実装した RAG を AdvancedRAG と呼び、基本パイプラインを SimpleRAG と呼びます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" alt="シンプルなRAGパイプライン" /><h4>結果の要約</h4><p>この表は、両方の RAG パイプラインの 5 つのテストの結果をまとめたものです。回答の詳細と品質に基づいて各方法の相対的な優位性を判断しましたが、これは完全に主観的な判断です。実際の回答はこの表の下に再現されていますので、ご参照ください。それでは、彼らの成果を見てみましょう!</p><p>SimpleRAG は質問 1 と 5 に答えることができませんでした。AdvancedRAG は質問 2、3、4 についても非常に詳しく説明しました。詳細度が増したことにより、AdvancedRAG の回答の質が優れていると判断しました。</p><p>テスト</p><p>質問</p><p>高度なRAGパフォーマンス</p><p>SimpleRAG パフォーマンス</p><p>AdvancedRAG レイテンシー</p><p>SimpleRAG レイテンシ</p><p>勝者</p><p>1</p><p>Elastic を監査するのは誰ですか?</p><p>監査人として PwC を正しく特定しました。</p><p>監査人を識別できませんでした。</p><p>11.6秒</p><p>4.4秒</p><p>アドバンスドRAG</p><p>2</p><p>2023年の総収益はいくらでしたか?</p><p>正しい収益数値を提供しました。前年度の収益に関する追加のコンテキストを含めました。</p><p>正しい収益数値を提供しました。</p><p>13.3秒</p><p>2.8秒</p><p>アドバンスドRAG</p><p>3</p><p>成長は主にどの製品に依存しますか?いくら？</p><p>Elastic Cloud が主要な推進力であることを正しく認識しました。全体的な収益コンテキストとより詳しい詳細が含まれています。</p><p>Elastic Cloud が主要な推進力であることを正しく認識しました。</p><p>14.1秒</p><p>12.8秒</p><p>アドバンスドRAG</p><p>4</p><p>従業員福利厚生プランの説明</p><p>退職金制度、健康プログラム、その他の福利厚生について包括的に説明しました。異なる年ごとの具体的な寄付金額が含まれています。</p><p>報酬、退職金制度、職場環境、Elastic Cares プログラムなどの福利厚生の概要をわかりやすく説明しました。</p><p>26.6秒</p><p>11.6秒</p><p>アドバンスドRAG</p><p>5</p><p>Elastic が買収した企業はどれですか?</p><p>レポートに記載されている最近の買収 (CmdWatch、Build Security、Optimyze) を正しくリストしました。いくつかの取得日と購入価格を提供しました。</p><p>提供されたコンテキストから関連情報を取得できませんでした。</p><p>11.9秒</p><p>2.7秒</p><p>アドバンスドRAG</p><h4>テスト 1: Elastic を監査するのは誰ですか?</h4><p>アドバンスドRAG</p>According to the context, Elastic N.V. is audited by an independent 
registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of 
independent registered public accounting firm," which states:
"We have audited the accompanying consolidated balance sheets 
of Elastic N.V. [...] / s / pricewaterhouseco."
<p>シンプルラグ</p>According to the provided context, no specific auditing firm 
or company is mentioned as the auditor for Elastic. 
The context provided talks about various features, strengths, 
deployment options, and other details related to Elastic and 
its operations but does not specify which auditing firm is 
responsible for auditing Elastic's financial statements or 
internal controls.
<p><strong>要約</strong>: SimpleRAGはPWCを監査人として特定しなかった</p><p>そうですね、それは実はかなり驚きました。これは SimpleRAG 側の検索失敗のようです。監査に関連する文書は取得されませんでした。次のテストでは難易度を少し下げてみましょう。</p><h4>テスト2：2023年の総収入</h4><p>アドバンスドRAG</p>### Total Revenue for 2023

According to the provided context, the total revenue for the 
year ended April 30, 2023, was **$1,068,989,000**. This 
information can be found several times in the provided text. 
Specifically, one part states:

"The following table sets forth our results of operations 
for the periods presented in dollars and as a percentage of 
our total revenue. ... year ended April 30, 2023, ... 
total revenue 1,068,989".

Additionally, it is mentioned here as well:

"For the years ended April 30, 2023, 2022, and 2021, 
representing year-over-year growth of 24% for the year 
ended April 30, 2023... total revenue $1,068,989".

Therefore, the total revenue for 2023 is $1.069 billion.
<p>シンプルラグ</p>According to the context provided, the total revenue for the fiscal year ended April 30, 2023, is:

$1,068,989,000 (or approximately $1.1 billion).
<p><strong>要約</strong>: 両RAGとも正解: 2023年の総収益は1,068,989,000ドル</p><p>二人ともここにいました。AdvancedRAG がより広範囲の文書を入手したように思われますか?確かに、答えはより詳細で、前年からの情報が組み込まれています。私たちが行った機能強化を考えると、それは予想されることですが、判断するには時期尚早です。</p><p>難易度を上げてみましょう。</p><h4>テスト 3: 成長は主にどの製品に依存しますか?いくら？</h4><p>アドバンスドRAG</p>### Product Growth Primarily Depends On:
According to the context, "our revenue is derived primarily from 
subscriptions of search, observability and security built into 
the elastic stack." The document mentions various strategies and 
factors contributing to the company’s growth which include investment 
in "elastic cloud," "elastic stack," and the conversion of free 
users to paid subscribers.

### Quantitative Data:
- **Revenue**: The document states, "Our total revenue was $1.1 billion, 
$862.4 million, and $608.5 million for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Subscription Revenue**: "Subscriptions accounted for 92%, 93%, 
and 93% of our total revenue for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Elastic Cloud Contribution**: "For the years ended April 30, 2023, 
2022, and 2021, Elastic Cloud contributed 40%, 35%, and 27% of our 
total revenue, respectively."
- **Customer Growth**: The number of customers increased from 
"over 15,000 as of April 30, 2021," to "approximately 20,200 as of 
April 30, 2023."

### Summary:
Product growth for the company primarily depends on the "Elastic Stack" 
(which includes search, observability, and security solutions) and 
particularly the "Elastic Cloud." The Elastic Cloud's growing contribution 
to revenue highlights its significance, which grew from 27% in 2021 to 40% in 2023.
<p>シンプルラグ</p>According to the provided context, growth primarily depends on **Elastic Cloud**, 
which is described as the company's "family of cloud-based offerings." 
The context also emphasizes the significance of Elastic Cloud in achieving 
the company's long-term growth potential and its role in increasing subscription revenue.

From the context:

- "Increasing adoption of Elastic Cloud... is an important growth opportunity 
for our business. Organizations are increasingly looking for hosted 
deployment alternatives with reduced administrative burdens."
- "We believe that offering Elastic Cloud is important for achieving 
our long-term growth potential, and we expect Elastic Cloud’s contribution 
to our subscription revenue to continue to increase."

In terms of the contribution to total revenue, here are the specific figures provided:

- Elastic Cloud contributed 40% of the total revenue for the year ended April 30, 2023.
- Elastic Cloud contributed 35% of the total revenue for the year ended April 30, 2022.
- Elastic Cloud contributed 27% of the total revenue for the year ended April 30, 2021.

This progression shows a clear upward trend in the importance and 
contribution of Elastic Cloud to the company's overall revenue.
<p><strong>概要</strong>: 両方の RAG は、Elastic Cloud を主要な成長原動力として正しく認識しました。ただし、AdvancedRAG では、サブスクリプション収益と顧客の増加を考慮したより詳細な情報が含まれており、Elastic の他の製品についても明示的に言及されています。</p><h4>テスト4: 従業員福利厚生制度の説明</h4><p>アドバンスドRAG</p>### Employee Benefit Plan Overview

The provided context outlines various elements of the employee benefit plan at 
Elastic, often referred to as "Elasticians." Here are the key components 
described in the texts:

1. **401(k) Plan**:
   - **Coverage**: The 401(k) plan is available to substantially all U.S. 
   employees who meet minimum age and service requirements.
   - **Contributions**: Elastic makes contributions to the 401(k) plan up to 
   6% of the participating employee’s W-2 earnings and wages.
   - **Expenses**: For the fiscal years ended April 30, Elastic recorded 
   expenses of $17.9 million (2023), $15.2 million (2022), and $11.4 million (2021) 
   related to the 401(k) plan.
   - **Defined-Contribution Plans in Other Countries**: Elastic has 
   defined-contribution plans in various other countries and recorded respective 
   expenses of $9.4 million (2023), $7.2 million (2022), and $5.1 million (2021).

2. **Stock-Based Compensation**:
   - **Types of Awards**: Stock options, restricted stock units (RSUs), 
   and shares under the Employee Stock Purchase Plan (ESPP).
   - **Fair Value Measurement**: Fair value of these stock awards is 
   measured using models like Black-Scholes.
   - **Employee Stock Purchase Plan (2022 ESPP)**: 
     - Started in 2022, it allows employees to acquire ordinary 
     shares at a discount (85% of the market value at the beginning 
     or end of the offering period).
     - Offering periods are approximately six months long.

3. **Total Rewards Compensation**:
   - **Components**: Includes cash compensation as well as equity awards, 
   reflecting a comprehensive interest in the well-being of the employees.
   - **Programs**: Market competitive medical and dental programs, 
   mental health, and holistic well-being programs.
   - **Paid Time Off**: Features 16 weeks of paid leave for new 
   parents and life-planning benefits.

4. **Retirement and Income Protection Plans**:
   - **401(k) Match**: Elastic provides a dollar-for-dollar 
   match up to 6% of eligible earnings up to a plan-limit 
   maximum for U.S.-based employees.
   - **Similar Plans Outside the U.S.**: Competitive retirement 
   plans available for employees based outside the United States.

5. **Flexible Work Environment**: 
   - **Distributed Workforce**: Elastic operates with a global 
   distributed workforce model, allowing employees to work remotely 
   and providing a broad candidate pool.
   
6. **Learning and Development**:
   - **Programs**: Access to on-demand video-based learning, 
   leadership development programs, and other specific training 
   to develop managers and leaders.

7. **Employee Engagement and Surveys**:
   - **Mechanisms**: Annual employee engagement survey and 
   mid-year pulse survey.
   - **Objectives**: To monitor morale, develop action plans, 
   and maintain high engagement scores.

8. **Onboarding Program**:
   - **X-School**: An extensive new-hire orientation program to 
   help new employees meet peers and learn about the company's 
   products and solutions.

In summary, Elastic's employee benefit plan encompasses a broad 
spectrum of financial, health, well-being, and professional development 
benefits designed to support and enhance the lives of its employees. 
The plan is both competitive and comprehensive, aiming to attract,
retain, and engage its global workforce.
<p>シンプルラグ</p>The employee benefit plan at the Company includes several key 
components aimed at enhancing the overall well-being and 
professional development of employees, referred to as "Elasticians." 

1. **Compensation and Benefits:**
   - The Company provides market-competitive compensation 
   packages which typically include cash compensation and equity awards.
   - Comprehensive medical and dental programs are offered, 
   emphasizing mental health and holistic well-being.
   - Paid time off programs feature:
     - 16 weeks of paid leave for new parents.
     - Life-planning benefits.
     - Travel reimbursements for certain healthcare services.

2. **Retirement and Income Protection Plans:**
   - In the United States, employees are eligible for a 401k 
   plan with a dollar-for-dollar match by the Company up to 6% 
   of eligible earnings, subject to a plan limit.
   - Similar competitive retirement and income protection plans 
   are available for employees outside the United States.

3. **Work Environment:**
   - The Company offers a flexible work environment, allowing most 
   employees to work from anywhere as often as they would like.
   - Physical offices around the world are available for employees 
   who prefer to work on-site or need a space for face-to-face interactions.

4. **Community Involvement:**
   - The Elastic Cares Program supports charitable organizations 
   that matter to employees on both local and global levels.
   - The program includes:
     - Donation matching.
     - Providing the Company's technology for free to certain 
     nonprofit organizations.
     - Volunteer time off initiative.
<p><strong>概要</strong>: AdvancedRAG では、米国に拠点を置く従業員向けの 401K プランや、米国外の定義拠出金プランについてさらに詳しく取り上げています。また、健康と幸福の計画についても言及していますが、SimpleRAG が言及している Elastic Cares プログラムについては触れられていません。</p><h4>テスト 5: Elastic が買収した企業はどれですか?</h4><p>アドバンスドRAG</p>Elastic acquired several companies over the years to enhance its 
technology and market presence. According to the provided context, 
the specific acquisitions include:

1. **CmdWatch Security Inc.**: Acquired on September 17, 2021. 
The total purchase consideration was $77.8 million.
2. **Build Security Ltd.**: Acquired on September 2, 2021. The 
combined purchase price for Build Security Ltd. and another acquisition was $57.2 million.
3. **Optimyze Cloud Inc.**: Acquired on November 1, 2021. Part 
of the combined purchase price mentioned above with Build Security Ltd.
4. **Endgame, Inc.**: Although the context mentions an amended 
and restated stock incentive plan, it indicates that Endgame, Inc. was 
associated with Elastic, suggesting it could have been acquired at some point.

These acquisitions are part of Elastic's strategy to 
enhance the technology underlying its security and 
observability offerings.
<p>シンプルラグ</p>The provided context does not specify any companies that Elastic has acquired. 
Therefore, based on the context, there is no information available about the companies acquired by Elastic.
<p><strong>概要</strong>: SimpleRAG は買収に関する関連情報を取得せず、回答に失敗しました。AdvancedRAG は、レポートに記載されている主要な買収である CmdWatch、Build Security、Optimyze を正しくリストしています。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h2>まとめ</h2><p>私たちのテストによると、私たちの高度な技術により、提示される情報の範囲と深さが拡大し、RAG 回答の品質が向上する可能性があるようです。</p><p>さらに、 <code>Which companies did Elastic acquire?</code>や<code>Who audits Elastic</code>などのあいまいな表現の質問に対して、AdvancedRAG では正しく回答されましたが、SimpleRAG では正しく回答されなかったため、信頼性が向上する可能性があります。</p><p>ただし、5 件中 3 件では、ハイブリッド検索のみを組み込んだ基本的な RAG パイプラインで、重要な情報のほとんどを捉えた回答を生成できたという点に留意する価値があります。</p><p>データ準備フェーズとクエリフェーズに LLM が組み込まれているため、AdvancedRAG のレイテンシは通常、SimpleRAG の 2 ～ 5 倍になることに注意してください。これは大きなコストであるため、AdvancedRAG は、応答品質がレイテンシーよりも優先される状況にのみ適している可能性があります。</p><p>データ準備段階で Claude Haiku や GPT-4o-mini などの小型で安価な LLM を使用すると、大きなレイテンシ コストを軽減できます。回答生成用の高度なモデルを保存します。</p><p>これは Wang らの研究結果と一致しています。結果が示すように、行われた改善は比較的漸進的です。つまり、シンプルなベースライン RAG を使用すると、安価で高速でありながら、適切な最終製品にほぼ到達できます。私にとっては、それは興味深い結論です。速度と効率が重要となるユースケースでは、SimpleRAG が賢明な選択です。パフォーマンスを最大限に引き出す必要があるユースケースでは、AdvancedRAG に組み込まれたテクニックが解決策となる可能性があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt56b7067a9d41d5a8/6a171119acf0886fb4be9c45/ea811706b6adc4731d90b925a9fefa0ac15901b4-1440x1060.jpg" alt="王パイプライン" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h2>付記</h2><h3>プロンプト</h3><h4>RAG質問回答プロンプト</h4><p>クエリとコンテキストに基づいて LLM に回答を生成させるためのプロンプト。</p>BASIC_RAG_PROMPT = '''
You are an AI assistant tasked with answering questions based primarily on the provided context, while also drawing on your own knowledge when appropriate. Your role is to accurately and comprehensively respond to queries, prioritizing the information given in the context but supplementing it with your own understanding when beneficial. Follow these guidelines:

1. Carefully read and analyze the entire context provided.
2. Primarily focus on the information present in the context to formulate your answer.
3. If the context doesn't contain sufficient information to fully answer the query, state this clearly and then supplement with your own knowledge if possible.
4. Use your own knowledge to provide additional context, explanations, or examples that enhance the answer.
5. Clearly distinguish between information from the provided context and your own knowledge. Use phrases like "According to the context..." or "The provided information states..." for context-based information, and "Based on my knowledge..." or "Drawing from my understanding..." for your own knowledge.
6. Provide comprehensive answers that address the query specifically, balancing conciseness with thoroughness.
7. When using information from the context, cite or quote relevant parts using quotation marks.
8. Maintain objectivity and clearly identify any opinions or interpretations as such.
9. If the context contains conflicting information, acknowledge this and use your knowledge to provide clarity if possible.
10. Make reasonable inferences based on the context and your knowledge, but clearly identify these as inferences.
11. If asked about the source of information, distinguish between the provided context and your own knowledge base.
12. If the query is ambiguous, ask for clarification before attempting to answer.
13. Use your judgment to determine when additional information from your knowledge base would be helpful or necessary to provide a complete and accurate answer.

Remember, your goal is to provide accurate, context-based responses, supplemented by your own knowledge when it adds value to the answer. Always prioritize the provided context, but don't hesitate to enhance it with your broader understanding when appropriate. Clearly differentiate between the two sources of information in your response.

Context:
[The concatenated documents will be inserted here]

Query:
[The user's question will be inserted here]

Please provide your answer based on the above guidelines, the given context, and your own knowledge where appropriate, clearly distinguishing between the two:
'''
<h4>弾性クエリジェネレータプロンプト</h4><p>同義語を使用してクエリを拡充し、OR 形式に変換するように要求します。</p>ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<h4>潜在的な質問ジェネレータプロンプト</h4><p>潜在的な質問の生成を促し、ドキュメントのメタデータを充実させます。</p>RAG_QUESTION_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating questions for Retrieval-Augmented Generation (RAG) systems. Your task is to analyze a given document and create 10 diverse questions that would effectively test a RAG system's ability to retrieve and synthesize information from this document.

Guidelines:
1. Thoroughly analyze the entire document.
2. Generate exactly 10 questions that cover various aspects and levels of complexity within the document's content.
3. Create questions that specifically target:
   a. Key facts and information
   b. Main concepts and ideas
   c. Relationships between different parts of the content
   d. Potential applications or implications of the information
   e. Comparisons or contrasts within the document
4. Ensure questions require answers of varying lengths and complexity, from simple retrieval to more complex synthesis.
5. Include questions that might require combining information from different parts of the document.
6. Frame questions to test both literal comprehension and inferential understanding.
7. Avoid yes/no questions; focus on open-ended questions that promote comprehensive answers.
8. Consider including questions that might require additional context or knowledge to fully answer, to test the RAG system's ability to combine retrieved information with broader knowledge.
9. Number the questions from 1 to 10.
10. Output only the ten questions, without any additional text, explanations, or answers.

Document:
[The document content will be inserted here]

Generate 10 questions optimized for testing a RAG system based on this document:
'''
<h4>HyDEジェネレータプロンプト</h4><p>HyDEを使用して仮想文書を生成するためのプロンプト</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<h3>ハイブリッド検索クエリのサンプル</h3>{'knn': {'field': 'primary_embedding',
  'query_vector': [0.4265527129173279,
   -0.1712949573993683,
   -0.042020395398139954,
   ...],
  'k': 100,
  'num_candidates': 100},
 'query': {'bool': {'must': [{'multi_match': {'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
      'fields': ['original_text',
       'keyphrases',
       'potential_questions',
       'entities'],
      'type': 'best_fields',
      'operator': 'or'}}],
   'should': [{'script_score': {'query': {'match_all': {}},
      'script': {'source': '\n                                        double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;\n                                        double text_score = _score;\n                                        return 0.7 * vector_score + 0.3 * text_score;\n                                        ',
       'params': {'query_vector': [0.4265527129173279,
         -0.1712949573993683,
         -0.042020395398139954,
        ...],
        'vector_field': 'primary_embedding'}}}}]}},
 'size': 10}
]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[高度なRAGテクニックパート1：データ処理]]></title>
    <description><![CDATA[RAG のパフォーマンスを向上させる可能性のあるテクニックについて議論し、実装します。パート 1/2。高度な RAG パイプラインのデータ処理と取り込みのコンポーネントに焦点を当てます。]]></description>
    <content:encoded><![CDATA[<p><em>これは、高度なRAGテクニックを探るパート1です。</em><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2"><em>パート2はこちらをクリックしてください！</em></a></p><p>最近の論文<a href="https://arxiv.org/abs/2407.01219">「検索拡張生成におけるベスト プラクティスの探求」では、</a> RAG のベスト プラクティスのセットに収束することを目的として、さまざまな RAG 強化手法の有効性を経験的に評価しています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt671704ff06a4011d/6a170b3ea929cf2d19ae09d8/dafa7250e7c4ead4d9b4aed7c407509131929749-1440x572.png" alt="王氏が推奨するRAGパイプライン" /><p>提案されたベストプラクティスのいくつか、つまり検索の品質を向上させることを目的としたベストプラクティス<strong>（センテンスチャンキング、HyDE、リバースパッキング）</strong>を実装します。</p><p>簡潔にするために、効率性の向上に重点を置いた手法<strong>(クエリの分類と要約)</strong>は省略します。</p><p>また、ここでは取り上げなかったものの、個人的には便利で興味深いと思われるいくつかのテクニック<strong>(メタデータの包含、複合マルチフィールドの埋め込み、クエリの強化) も</strong>実装します。</p><p>最後に、検索結果と生成された回答の品質がベースラインと比較して向上したかどうかを確認するための短いテストを実行します。さあ始めましょう！</p><h2>RAGの概要</h2><p>RAG は、外部の知識ベースから情報を取得して生成された回答を充実させることで、LLM を強化することを目的としています。ドメイン固有の情報を提供することで、LLM はトレーニング データの範囲外のユース ケースに迅速に適応できます。微調整よりも大幅にコストが安く、最新の状態に保つのも簡単になります。</p><p>RAG の品質を向上させるための対策は、通常、次の 2 つの点に重点を置いています。</p><ol><li><p>ナレッジベースの品質と明確さを向上します。</p></li><li><p>検索クエリの範囲と特定性を向上させます。</p></li></ol><p>これら 2 つの対策により、LLM が関連する事実や情報にアクセスできる可能性が高まり、幻覚を起こしたり、古くなったり無関係になったりする可能性のある独自の知識を利用したりする可能性が低くなるという目標が達成されます。</p><p>方法の多様性を数文で説明するのは困難です。わかりやすくするために、すぐに実装に移りましょう。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="高度なRAGパイプライン" /><h3>目次</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#overview">ご紹介</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">目次</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#set-up">設定</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#ingesting-processing-and-embedding-documents">ドキュメントの取り込み、処理、埋め込み</a>  </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#data-ingestion">データインジェスト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#sentence-level-token-wise-chunking">文レベル、トークン単位のチャンキング</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#metadata-inclusion-and-generation">メタデータの包含と生成</a> </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#keyphrases-extracted-by-textrank">TextRankによって抽出されたキーフレーズ</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#potential-questions-generated-by-gpt-4o">GPT-4oによって生成される潜在的な質問</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#entities-extracted-by-spacy">Spacyによって抽出されたエンティティ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#composite-multi-field-embeddings">複合多体埋め込み</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#indexing-to-elastic">Elasticへのインデックス</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#cat-break">猫の休憩</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#appendix">付記</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#definitions">定義</a></p></li></ul></li></ul><h2>設定</h2><p><em>すべてのコードは</em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em> Searchlabs リポジトリに</em></a><em> あります 。</em></p><p>まずは第一に。次のものが必要になります:</p><ol><li><p>弾力性のあるクラウドの展開</p></li><li><p>LLM API - このノートブックでは、Azure OpenAI 上の GPT-4o デプロイメントを使用しています。</p></li><li><p>Python バージョン 3.12.4 以降</p></li></ol><p><a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/main.ipynb">main.ipynb ノートブックからすべてのコードを実行します。</a></p><p>リポジトリを git clone し、supporting-blog-content/advanced-rag-techniques に移動して、次のコマンドを実行します。</p># Create a new virtual environment named 'rag_env'
python -m venv rag_env

# Activate the virtual environment (for Unix-based systems)
source rag_env/bin/activate

# (For Windows)
.\rag_env\Scripts\activate

# Install packages listed in requirements.txt
pip install -r requirements.txt
<p>完了したら、 <em>.env</em>を作成します。ファイルを開き、次のフィールドに入力します ( <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/.env.example"><em>.env.example</em></a>で参照されます)。有益なコメントをくれた共著者の Claude-3.5 に感謝します。</p># Elastic Cloud: Found in the 'Deployment' page of your Elastic Cloud 
# console
ELASTIC_CLOUD_ENDPOINT=""
ELASTIC_CLOUD_ID=""

# Elastic Cloud: Created during deployment setup or in 'Security' 
# settings
ELASTIC_USERNAME=""
ELASTIC_PASSWORD=""

# Elastic Cloud: The name of the index you created in Kibana or via API
ELASTIC_INDEX_NAME=""

# Azure AI Studio: Found in 'Keys and Endpoint' section of your Azure 
# OpenAI resource
AZURE_OPENAI_KEY_1=""
AZURE_OPENAI_KEY_2=""
AZURE_OPENAI_REGION=""
AZURE_OPENAI_ENDPOINT=""

# Azure AI Studio: Found in 'Deployments' section of your Azure OpenAI 
# resource
AZURE_OPENAI_DEPLOYMENT_NAME=""

# Using BAAI/bge-small-en-v1.5 because I think it is a good balance of 
# resource efficiency and performance. 
HUGGINGFACE_EMBEDDING_MODEL="BAAI/bge-small-en-v1.5"
<p>次に、取り込むドキュメントを選択し、ドキュメント フォルダーに配置します。この記事では、 <a href="https://s201.q4cdn.com/217177842/files/doc_downloads/OtherDocuments/2023/AnnualMeeting/Annual-Report-Fiscal-Year-2023.pdf">Elastic NV Annual Report 2023 を</a>使用します。これは非常に難しくて密度の高いドキュメントであり、RAG テクニックのストレス テストに最適です。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte292dc6030d496cc/6a170b40dc55de9b03e00dfc/e513b9d67adac43da794c25a5969b893127bbbe3-1440x395.jpg" alt="Elastic 年次報告書 2023" /><p>準備が整いましたので、摂取に移りましょう。<em>main.ipynb</em>を開き、最初の 2 つのセルを実行して、すべてのパッケージをインポートし、すべてのサービスを初期化します。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h2>ドキュメントの取り込み、処理、埋め込み</h2><h3>データインジェスト</h3><ul><li><p><em>個人的なメモ: LlamaIndex の便利さに驚いています。LLM や LlamaIndex が登場する前の昔、さまざまな形式のドキュメントを取り込むには、あらゆる場所から難解なパッケージを収集する、骨の折れる作業でした。今では、関数呼び出しは 1 つに減りました。野生。</em></p></li></ul><p><code>SimpleDirectoryReader</code>は<code>directory_path.</code>内のすべてのドキュメントをロードします。 <code>.pdf</code>ファイルの場合は、ドキュメント オブジェクトのリストを返します。このリストは、操作しやすいように Python 辞書に変換します。</p># llamaindex_processor.py
from llama_index.core import SimpleDirectoryReader

class LlamaIndexProcessor:
   def __init__(self):
       pass 
   
   def load_documents(self, directory_path):
       ''' 
       Load all documents in directory
       '''
       reader = SimpleDirectoryReader(input_dir=directory_path)
       return reader.load_data()

# main.ipynb
llamaindex_processor=LlamaIndexProcessor()
documents=llamaindex_processor.load_documents('./documents/')
documents=[dict(doc_obj) for doc_obj in documents]
<p>各辞書には、 <code>text</code>フィールドにキー コンテンツが含まれています。また、ページ番号、ファイル名、ファイル サイズ、タイプなどの便利なメタデータも含まれています。</p>{
  'id_': '5f76f0b3-22d8-49a8-9942-c2bbab14f63f',
  'metadata': {'page_label': '5',
   'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_path': '/Users/han/Desktop/Projects/truckasaurus/documents/Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_type': 'application/pdf',
   'file_size': 3724426,
   'creation_date': '2024-07-27',
   'last_modified_date': '2024-07-27'},
   'text': 'Table of Contents\nPage\nPART I\nItem 1. Business 3\n15 Item 1A. Risk Factors\nItem 1B. Unresolved Staff Comments 48\nItem 2. Properties 48\nItem 3. Legal Proceedings 48\nItem 4. Mine Safety Disclosures 48\nPART II\nItem 5. Market for Registrant's Common Equity, Related Stockholder Matters and Issuer Purchases of \nEquity Securities49\nItem 6. [Reserved] 49\nItem 7. Management's Discussion and Analysis of Financial Condition and Results of Operations 50\nItem 7A. Quantitative and Qualitative Disclosures About Market Risk 64\nItem 8. Financial Statements and Supplementary Data 66\nItem 9. Changes in and Disagreements With Accountants on Accounting and Financial Disclosure 100\n100\n101Item 9A. Controls and Procedures\nItem 9B. Other Information\nItem 9C. Disclosure Regarding Foreign Jurisdictions That Prevent Inspections 101\nPART III\n102\n102\n102\n102Item 10. Directors, Executive Officers and Corporate Governance\nItem 11. Executive Compensation\nItem 12. Security Ownership of Certain Beneficial Owners and Management, and Related Stockholder Matters  \nItem 13. Certain Relationships and Related Transactions, and Director Independence\nItem 14. Principal Accountant Fees and Services 102\nPART IV\n103\n105Item 15. Exhibits and Financial Statement Schedules  \nItem 16. Form 10-K Summary\nSignatures 106\ni',
   ...
}
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h3>文レベル、トークン単位のチャンキング</h3><p>最初にやるべきことは、ドキュメントを標準的な長さのチャンクに削減することです (一貫性と管理性を確保するため)。埋め込みモデルには、固有のトークン制限 (処理できる最大入力サイズ) があります。トークンはモデルが処理するテキストの基本単位です。情報の損失（コンテンツの切り捨てや省略）を防ぐために、これらの制限を超えないテキストを提供する必要があります（長いテキストを短いセグメントに分割する）。</p><p>チャンク化はパフォーマンスに大きな影響を与えます。理想的には、各チャンクは自己完結的な情報を表し、単一のトピックに関するコンテキスト情報をキャプチャします。チャンク化の方法には、文書を単語数で分割する単語レベルのチャンク化と、LLM を使用して論理ブレークポイントを識別するセマンティック チャンク化があります。</p><p>単語レベルのチャンキングは安価で高速かつ簡単ですが、文が分割され、コンテキストが壊れるリスクがあります。セマンティック チャンキングは、特に 116 ページの Elastic 年次レポートのようなドキュメントを扱う場合には、時間がかかり、コストも高くなります。</p><p>中道的なアプローチを選択しましょう。文レベルのチャンキングは依然としてシンプルですが、単語レベルのチャンキングよりもコンテキストをより効果的に保持でき、コストも大幅に削減され、処理速度も速くなります。さらに、周囲のコンテキストの一部をキャプチャし、段落を分割することによる影響を軽減するために、スライディング ウィンドウを実装します。</p># chunker.py 

import uuid
import re


class Chunker: 
    def __init__(self, tokenizer):
        self.tokenizer = tokenizer 
    
    def split_into_sentences(self, text):
        """Split text into sentences."""
        return re.split(r'(?&lt;=[.!?])\s+', text)
 
    def sentence_wise_tokenized_chunk_documents(self, documents, chunk_size=512, overlap=20, min_chunk_size=50):
        '''
        1. Split text into sentences.
        2. Tokenize using the provided tokenizer method.
        3. Build chunks up to the chunk_size limit.
        4. Create an overlap based on tokens - to preserve context.
        5. Only keep chunks that meet the minimum token size requirement.
        '''
        chunked_documents = []

        for doc in documents:
            sentences = self.split_into_sentences(doc['text'])
            tokens = []
            sentence_boundaries = [0]

            # Tokenize all sentences and keep track of sentence boundaries
            for sentence in sentences:
                sentence_tokens = self.tokenizer.encode(sentence, add_special_tokens=True)
                tokens.extend(sentence_tokens)
                sentence_boundaries.append(len(tokens))

            # Create chunks
            chunk_start = 0
            while chunk_start &lt; len(tokens):
                chunk_end = chunk_start + chunk_size

                # Find the last complete sentence that fits in the chunk
                sentence_end = next((i for i in sentence_boundaries if i &gt; chunk_end), len(tokens))
                chunk_end = min(chunk_end, sentence_end)

                # Create the chunk
                chunk_tokens = tokens[chunk_start:chunk_end]

                # Check if the chunk meets the minimum size requirement
                if len(chunk_tokens) &gt;= min_chunk_size:
                    # Create a new document object for this chunk
                    chunk_doc = {
                        'id_': str(uuid.uuid4()),
                        'chunk': chunk_tokens,
                        'original_text': self.tokenizer.decode(chunk_tokens),
                        'chunk_index': len(chunked_documents),
                        'parent_id': doc['id_'],
                        'chunk_token_count': len(chunk_tokens)
                    }

                    # Copy all other fields from the original document
                    for key, value in doc.items():
                        if key != 'text' and key not in chunk_doc:
                            chunk_doc[key] = value

                    chunked_documents.append(chunk_doc)

                # Move to the next chunk start, considering overlap
                chunk_start = max(chunk_start + chunk_size - overlap, chunk_end - overlap)

        return chunked_documents

# main.ipynb 
# Initialize Embedding Model
HUGGINGFACE_EMBEDDING_MODEL = os.environ.get('HUGGINGFACE_EMBEDDING_MODEL')
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

# Initialize Chunker
chunker=Chunker(embedder.tokenizer)
<p><code>Chunker</code>クラスは埋め込みモデルのトークナイザーを受け取り、テキストをエンコードおよびデコードします。ここで、20 個のトークンが重なり合う、それぞれ 512 個のトークンのチャンクを構築します。これを実行するには、テキストを文に分割し、それらの文をトークン化してから、トークン制限に違反することなく追加できなくなるまで、トークン化された文を現在のチャンクに追加します。</p><p>最後に、埋め込みのために文章を元のテキストにデコードし、 <code>original_text</code>というフィールドに保存します。チャンクは<code>chunk</code>というフィールドに保存されます。ノイズ（つまり、無駄なドキュメント）を減らすために、長さが 50 トークン未満のドキュメントは破棄されます。</p><p>これをドキュメント上で実行してみましょう。</p>chunked_documents=chunker.sentence_wise_tokenized_chunk_documents(documents, chunk_size=512)
<p>そして、次のようなテキストのチャンクが返されます。</p>print(chunked_documents[4]['original_text'])

[CLS] the aggregate market value of the ordinary shares held by non - affiliates of the registrant, 
based on the closing price of the shares of ordinary shares on the new york stock exchange on 
october 31, 2022 ( the last business day of the registrant 's second fiscal quarter ), was 
approximately $ 6. 1 billion. [SEP] [CLS] as of may 31, 2023, the registrant had 97, 390, 886 
ordinary shares, par value €0. 01 per share, outstanding. [SEP] [CLS] documents incorporated by 
reference portions of the registrant 's definitive proxy statement relating to the registrant 's 2
023 annual general meeting of shareholders are incorporated by reference into part iii of this annual 
...
...
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h3>メタデータの包含と生成</h3><p>ドキュメントをチャンクに分割しました。次は、データを充実させる段階です。追加のメタデータを生成または抽出したい。この追加のメタデータは、検索パフォーマンスに影響を与え、強化するために使用できます。</p><p>ドキュメントのリスト (Python 辞書) とプロセッサ関数のリストを受け取る役割を持つ<code>DocumentEnricher</code>クラスを定義します。これらの関数はドキュメントの<code>original_text</code>列を実行し、その出力を新しいフィールドに保存します。</p><p>まず、 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/nltk_processor.py">TextRank</a>を使用してキーフレーズを抽出します。TextRank は、単語間の関係に基づいて重要度をランク付けすることにより、テキストから主要なフレーズと文を抽出するグラフベースのアルゴリズムです。</p><p>次に、 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/llm.py">GPT-4oを使用してpotential_questionsを生成します</a>。</p><p>最後に、<a href="https://spacy.io/"> Spacy</a> <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/entity_extractor.py">を使用して エンティティを抽出します</a> 。</p><p>それぞれのコードは非常に長くて複雑なので、ここで再現することは控えます。ご興味があれば、以下のコード サンプルにファイルがマークされています。</p><p>データ拡充を実行してみましょう:</p># documentenricher.py
from tqdm import tqdm

class DocumentEnricher:

    def __init__(self):
        pass 

    def enrich_document(self, documents, processors, text_col='text'):
        for doc in tqdm(documents, desc="Enriching documents using processors: "+str(processors)): 
            for (processor, field) in processors: 
                metadata=processor(doc[text_col])
                if isinstance(metadata, list):
                    metadata='\n'.join(metadata)
                doc.update({field: metadata})
 
# main.ipynb
# Initialize processor classes 
nltkprocessor=NLTKProcessor() // nltk_processor.py
entity_extractor=EntityExtractor() // entity_extractor.py
gpt4o = LLMProcessor(model='gpt-4o') // llm.py

# Initialize LLM
documentenricher=DocumentEnricher()

# Create new fields in the documents - These are the outputs of the processor functions.
processors=[
    (nltkprocessor.textrank_phrases, "keyphrases"),
    (gpt4o.generate_questions, "potential_questions"),
    (entity_extractor.extract_entities, "entities")
    ]

# .enrich_document() will modify chunked_docs in place. 
# To view the results, we'll print chunked_docs in the next few cells!
documentenricher.enrich_document(chunked_docs, text_col='original_text', processors=processors)
<p>結果を見てみましょう:</p><h4>TextRankによって抽出されたキーフレーズ</h4><p>これらのキーフレーズは、チャンクの中核トピックの代わりとなります。クエリがサイバーセキュリティに関係する場合、このチャンクのスコアは向上します。</p>print(chunked_documents[25]['keyphrases'])

'elastic agent stop', 'agent stop malware', 
'stop malware ransomware', 'malware ransomware environment', 
'ransomware environment wide', 'environment wide visibility', 
'wide visibility threat', 'visibility threat detection', 
'sep cl key', 'cl key feature'
<h4>GPT-4oによって生成される潜在的な質問</h4><p>これらの潜在的な質問はユーザーのクエリと直接一致する可能性があり、スコアの向上につながります。GPT-4o に、現在のチャンクにある情報を使用して回答できる質問を生成するように指示します。</p>print(chunked_documents[25]['potential_questions'])

1. What are the primary functions that Elastic Agent provides in terms of cybersecurity?
2. Describe how Logstash contributes to data management within an IT environment.
3. List and explain any key features of Logstash mentioned in the document.
4. How does Elastic Agent enhance environment-wide visibility in threat detection?
5. What capabilities does Logstash offer for handling data beyond simple collection?
6. In what ways does the document suggest that Elastic Agent stops malware and ransomware?
7. Can you identify any relationships between the functionalities of Elastic Agent and Logstash in an integrated environment?
8. What implications might the advanced threat detection capabilities of Elastic Agent have for organizational security policies?
9. Compare and contrast the roles of Elastic Agent and Logstash based on their described functions.
10. How might the centralized collection ability of Logstash support the threat detection capabilities of Elastic Agent?
<h4>Spacyによって抽出されたエンティティ</h4><p>これらのエンティティはキーフレーズと同様の目的を果たしますが、キーフレーズ抽出では見逃される可能性のある組織や個人の名前を取得します。</p>print(chunked_documents[29]['entities'])

'appdynamics', 'apm data', 'azure sentinel', 
'microsoft', 'mcafee', 'broadcom', 'cisco', 
'dynatrace', 'coveo', 'lucidworks'
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h3>複合多体埋め込み</h3><p>追加のメタデータでドキュメントを充実させたので、この情報を活用して、より堅牢でコンテキストを認識した埋め込みを作成できます。</p><p>プロセスの現在のポイントを確認しましょう。各ドキュメントには 4 つの興味深いフィールドがあります。</p>{
    "chunk": "...",
    "keyphrases": "...", 
    "potential_questions": "...", 
    "entities": "..." 
}
<p>各フィールドはドキュメントのコンテキストに関する異なる視点を表し、LLM が重点を置くべき重要な領域を強調する可能性があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt84cb328fce6aae23/6a170b42964cea3e4408bbc4/aea1f513009a0c7c8545a79fad8f072a5bcae24c-1440x1067.jpg" alt="RAG のメタデータ強化パイプライン" /><p>計画としては、これらの各フィールドを埋め込み、複合埋め込みと呼ばれる埋め込みの加重合計を作成することです。</p><p>運が良ければ、この複合埋め込みにより、検索動作を制御する別の調整可能なハイパーパラメータが導入されるだけでなく、システムがよりコンテキストを認識できるようになります。</p><p>まず、main.ipynb ノートブックの先頭にインポートされたローカルに定義された埋め込みモデルを使用して、各フィールドを埋め込み、各ドキュメントを更新します。</p># EmbeddingModel defined in embedding_model.py
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

cols_to_embed=['keyphrases', 'potential_questions', 'entities']

embedding_cols=[]
for col in cols_to_embed:
    # Works on text input
    embedding_col=embedder.embed_documents_text_wise(chunked_documents, text_field=col)
    embedding_cols.append(embedding_col)
# Works on token input
embedding_col=embedder.embed_documents_token_wise(chunked_documents, token_field="chunk")
embedding_cols.append(embedding_col)
<p>各埋め込み関数は埋め込みのフィールドを返します。これは、 <code>_embedding</code>という接尾辞が付いた元の入力フィールドです。</p><p>複合埋め込みの重みを定義しましょう。</p>embedding_cols=[
                'keyphrases_embedding',
                'potential_questions_embedding',
                'entities_embedding',
                'chunk_embedding']
combination_weights=[
                    0.1,
                    0.15,
                    0.05,
                    0.7
                ]
<p>重み付けにより、ユースケースとデータの品質に基づいて各コンポーネントに優先順位を割り当てることができます。直感的に言えば、これらの重み付けの大きさは、各コンポーネントの意味的価値に依存します。チャンクテキスト自体が圧倒的に豊富なので、重み付けを 70% に割り当てます。エンティティは組織名や人名のリストだけなので最も小さいので、重み付けを 5% に割り当てます。これらの値の正確な設定は、ユースケースごとに経験的に決定する必要があります。</p><p>最後に、重み付けを適用し、複合埋め込みを作成する関数を記述しましょう。スペースを節約するために、コンポーネントの埋め込みもすべて削除します。</p>from tqdm import tqdm 
def combine_embeddings(objects, embedding_cols, combination_weights, primary_embedding='primary_embedding'):
    # Ensure the number of weights matches the number of embedding columns
    assert len(embedding_cols) == len(combination_weights), "Number of embedding columns must match number of weights"
    
    # Normalize weights to sum to 1
    weights = np.array(combination_weights) / np.sum(combination_weights)
    
    for obj in tqdm(objects, desc="Combining embeddings"):
        # Initialize the combined embedding
        combined = np.zeros_like(obj[embedding_cols[0]])
        
        # Compute the weighted sum
        for col, weight in zip(embedding_cols, weights):
            combined += weight * np.array(obj[col])
        
        # Add the new combined embedding to the object
        obj.update({primary_embedding:combined.tolist()})
        
        # Remove the original embedding columns
        for col in embedding_cols:
            obj.pop(col, None)

combine_embeddings(chunked_documents, embedding_cols, combination_weights)
<p>これで書類の処理は完了です。次のようなドキュメント オブジェクトのリストが作成されました。</p>{ 'id_': '7fe71686-5cd0-4831-9e79-998c6dbeae0c', 'chunk': [2312, 14613, ...], 'original_text': 'if an emerging growth company, indicate by check mark if the registrant has elected not to use the extended ...', 'chunk_index': 3, 'chunk_token_count': 399, 'metadata': {'page_label': '3', 'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf', ... 'keyphrases': 'sep cl unk\ncheck mark registrant\ncl unk indicate\nunk indicate check\nindicate check mark\nprincipal executive office\naccelerate filer unk\ncompany unk emerge\nunk emerge growth\nemerge growth company', 'potential_questions': '1. What are the different types of registrant statuses mentioned in the document?\n2. Under what section of the Sarbanes-Oxley Act must registrants file a report on the effectiveness of their internal ...', 'entities': 'the effe ctiveness of\nsection 13\nSEP\nUNK\nsection 21e\n1934\n1933\nu. s. c.\nsection 404\nsection 12\nal', 'primary_embedding': [-0.3946287803351879, -0.17586839850991964, ...] }
<h4>Elasticへのインデックス</h4><p>ドキュメントを Elastic Search に一括アップロードしてみましょう。この目的のために、私はずっと前に<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/elastic_helpers.py"><code>elastic_helpers.py</code></a>で Elastic Helper 関数のセットを定義しました。これは非常に長いコードなので、関数呼び出しに注目してみましょう。</p><p><code>es_bulk_indexer.bulk_upload_documents</code> Elasticsearch の便利な動的マッピングを活用して、辞書オブジェクトの任意のリストで動作します。</p># Initialize Elasticsearch
ELASTIC_CLOUD_ID = os.environ.get('ELASTIC_CLOUD_ID')
ELASTIC_USERNAME = os.environ.get('ELASTIC_USERNAME')
ELASTIC_PASSWORD = os.environ.get('ELASTIC_PASSWORD')
ELASTIC_CLOUD_AUTH = (ELASTIC_USERNAME, ELASTIC_PASSWORD)
es_bulk_indexer = ESBulkIndexer(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)
es_query_maker = ESQueryMaker(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)

# Define Index Name
index_name=os.environ.get('ELASTIC_INDEX_NAME')


# Create index and bulk upload 
index_exists = es_bulk_indexer.check_index_existence(index_name=index_name)
if not index_exists:
    logger.info(f"Creating new index: {index_name}")
    es_bulk_indexer.create_es_index(es_configuration=BASIC_CONFIG, index_name=index_name)

success_count = es_bulk_indexer.bulk_upload_documents(
    index_name=index_name, 
    documents=chunked_documents, 
    id_col='id_',
    batch_size=32
)
<p>Kibana にアクセスして、すべてのドキュメントがインデックスされていることを確認します。全部で224個あるはずです。こんなに大きな文書にしては悪くないですね!</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8efeface6effe01d/6a170b447d8d67652870e72a/1b3b07f6b98ceb65f6594ce4be83c5b0ed7e7cf9-1440x1380.jpg" alt="インデックスキバナ" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h2>猫の休憩</h2><p>ちょっと休憩しましょう。記事がちょっと重いのはわかっています。私の猫を見てください:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1db5595f71c12ff/6a170b450e2e49940241a0fe/baca4eb52b801b21ced97352cc55462f0a12d6b0-969x996.jpg" alt="ハンパイプライン" /><p>愛らしい。帽子がなくなってしまったので、彼女がそれを盗んでどこかに隠したのではないかと半分疑っています :(</p><p>ここまで来られたことおめでとうございます :)</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">パート 2</a>では、RAG パイプラインのテストと評価についてご紹介します。</p><h2>付記</h2><h3>定義</h3><p><strong>1. 文のチャンキング</strong></p><ul><li><p>RAG システムでテキストをより小さな意味のある単位に分割するために使用される前処理手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>入力: 大きなテキストブロック（例: 文書、段落）</p></li><li><p>出力: 小さなテキストセグメント (通常は文または小さな文のグループ)</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>きめ細やかでコンテキストに特化したテキストセグメントを作成する</p></li><li><p>より正確なインデックス作成と検索が可能</p></li><li><p>RAGシステムで取得した情報の関連性を向上</p></li></ul></li><li><p><em>特徴:</em> </p><ul><li><p>セグメントは意味的に意味がある</p></li><li><p>独立してインデックスを作成し、検索できる</p></li><li><p>多くの場合、独立した理解可能性を確保するためにある程度の文脈が保持される</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>検索精度の向上</p></li><li><p>RAGパイプラインのより集中的な拡張を可能にします</p></li></ul></li></ul><p><strong>2. HyDE（仮想文書埋め込み）</strong></p><ul><li><p>LLM を使用して、RAG システムでのクエリ拡張用の仮想ドキュメントを生成する手法。</p></li><li><p><em>プロセス：</em>  </p><ol><li><p>LLMへの入力クエリ</p></li><li><p>LLMはクエリに答える仮説文書を生成する</p></li><li><p>生成されたドキュメントを埋め込む</p></li><li><p>ベクトル検索に埋め込みを使用する</p></li></ol></li><li><p><em>主な違い:</em> </p><ul><li><p>従来のRAG: クエリとドキュメントを一致させる</p></li><li><p>HyDE: 文書を文書と照合する</p></li></ul></li><li><p><em>目的：</em> </p><ul><li><p>特に複雑または曖昧なクエリの検索パフォーマンスを向上します</p></li><li><p>短いクエリよりも豊富な意味コンテキストをキャプチャする</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>LLMの知識を活用してクエリを拡張する</p></li><li><p>検索された文書の関連性が向上する可能性がある</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>追加のLLM推論が必要となり、レイテンシとコストが増加する</p></li><li><p>パフォーマンスは生成された仮想文書の品質に依存する</p></li></ul></li></ul><p><strong>3. 逆パッキング</strong></p><ul><li><p>RAG システムで、検索結果を LLM に渡す前に並べ替えるために使用される手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>検索エンジン (Elasticsearch など) は、関連性の高い順にドキュメントを返します。</p></li><li><p>順序は逆になり、最も関連性の高いドキュメントが最後に配置されます。</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>LLM の新しさバイアスを利用します。LLM は、それぞれのコンテキストにおける最新の情報に重点を置く傾向があります。</p></li><li><p>LLM のコンテキスト ウィンドウ内で最も関連性の高い情報が「最新」であることを保証します。</p></li></ul></li><li><p><em>例:</em>元の順序: [最も関連性の高い順、2番目に関連性の高い順、3番目に関連性の高い順、...] 逆の順序: [...、3番目に関連性の高い順、2番目に関連性の高い順、最も関連性の高い順]</p></li></ul><p><strong>4. クエリの分類</strong></p><ul><li><p>クエリに RAG が必要かどうか、または LLM によって直接回答できるかどうかを判断して、RAG システムの効率を最適化する手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>使用中の LLM に固有のカスタム データセットを開発する</p></li><li><p>特殊な分類モデルをトレーニングする</p></li><li><p>モデルを使用して受信したクエリを分類する</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>不要なRAG処理を回避することでシステム効率を向上</p></li><li><p>最も適切な応答メカニズムにクエリを直接送信する</p></li></ul></li><li><p><em>要件：</em> </p><ul><li><p>LLM固有のデータセットとモデル</p></li><li><p>精度を維持するための継続的な改良</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>単純なクエリの計算オーバーヘッドを削減</p></li><li><p>非RAGクエリの応答時間を改善する可能性がある</p></li></ul></li></ul><p><strong>5. 要約</strong></p><ul><li><p>RAG システムで検索された文書を圧縮する手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>関連文書を取得する</p></li><li><p>各文書の簡潔な要約を生成する</p></li><li><p>RAG パイプラインでは完全なドキュメントではなく要約を使用する</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>重要な情報に焦点を当ててRAGのパフォーマンスを向上させる</p></li><li><p>関連性の低いコンテンツからのノイズや干渉を減らす</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>LLM回答の関連性が向上する可能性がある</p></li><li><p>コンテキスト制限内でより多くのドキュメントを含めることができます</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>要約時に重要な詳細が失われるリスク</p></li><li><p>要約生成のための追加の計算オーバーヘッド</p></li></ul></li></ul><p><strong>6. メタデータの包含</strong></p><ul><li><p>追加のコンテキスト情報でドキュメントを充実させる手法。</p></li><li><p><em>メタデータの種類:</em>  </p><ul><li><p>キーフレーズ</p></li><li><p>タイトル</p></li><li><p>日付</p></li><li><p>著者詳細</p></li><li><p>宣伝文句</p></li></ul></li><li><p><em>目的：</em> </p><ul><li><p>RAGシステムで利用可能なコンテキスト情報を増やす</p></li><li><p>LLMに文書の内容と関連性をより明確に理解させる</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>検索精度が向上する可能性がある</p></li><li><p>LLMの文書有用性を評価する能力を強化する</p></li></ul></li><li><p><em>実装：</em> </p><ul><li><p>文書の前処理中に実行できる</p></li><li><p>追加のデータ抽出または生成手順が必要になる場合があります</p></li></ul></li></ul><p><strong>7. 複合多体埋め込み</strong></p><ul><li><p>異なるドキュメント コンポーネントごとに個別の埋め込みを作成する RAG システム用の高度な埋め込み手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>関連するフィールド（例：タイトル、キーフレーズ、宣伝文句、メインコンテンツ）を特定する</p></li><li><p>各フィールドごとに個別の埋め込みを生成する</p></li><li><p>これらの埋め込みを結合または保存して検索に使用します</p></li></ol></li><li><p><em>標準的なアプローチとの違い:</em> </p><ul><li><p>従来型: ドキュメント全体の単一の埋め込み</p></li><li><p>複合: さまざまなドキュメントの側面に対応する複数の埋め込み</p></li></ul></li><li><p><em>目的：</em> </p><ul><li><p>よりニュアンス豊かで文脈を考慮した文書表現を作成する</p></li><li><p>文書内のより多様なソースから情報を取得する</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>曖昧なクエリや多面的なクエリのパフォーマンスが向上する可能性があります</p></li><li><p>検索時にさまざまな文書の側面をより柔軟に重み付けできます</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>埋め込みストレージと検索プロセスの複雑さが増す</p></li><li><p>より洗練されたマッチングアルゴリズムが必要になる場合があります</p></li></ul></li></ul><p><strong>8. クエリエンリッチメント</strong></p><ul><li><p>元のクエリを関連用語で拡張し、検索範囲を広げる手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>元のクエリを分析する</p></li><li><p>同義語や意味的に関連するフレーズを生成する</p></li><li><p>クエリに以下の追加用語を追加します</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>文書コーパス内の潜在的な一致の範囲を拡大する</p></li><li><p>特定の言語や専門用語を含むクエリの検索パフォーマンスを向上</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>元の検索語句と完全に一致しない関連文書を取得する可能性がある</p></li><li><p>クエリとドキュメント間の語彙の不一致を克服するのに役立ちます</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>慎重に実装しないとクエリドリフトのリスクがある</p></li><li><p>検索プロセスにおける計算オーバーヘッドが増加する可能性がある</p></li></ul></li></ul><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 14 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch と LlamaIndex を使用して RAG の機密情報と PII 情報を保護する]]></title>
    <description><![CDATA[Elasticsearch と LlamaIndex を使用して RAG アプリケーション内の機密データと PII データを保護する方法。]]></description>
    <content:encoded><![CDATA[<p></p><p></p><p>この記事では、RAG (Retrieval Augmented Generation) フローでパブリック LLM を使用する際に、個人識別情報 (PII) と機密データを保護する方法について説明します。オープンソース ライブラリと正規表現を使用して PII と機密データをマスキングする方法と、パブリック LLM を呼び出す前にローカル LLM を使用してデータをマスキングする方法を検討します。</p><p>始める前に、この投稿で使用するいくつかの用語を確認しましょう。</p><h2>用語について</h2><p><a href="https://www.llamaindex.ai/">LlamaIndex は</a>、LLM (大規模言語モデル) アプリケーションを構築するための主要なデータ フレームワークです。LlamaIndex は、RAG (Retrieval Augmented Generation) アプリケーションの構築のさまざまな段階に抽象化を提供します。LlamaIndex や LangChain などのフレームワークは、アプリケーションが特定の LLM の API に密結合されないように抽象化を提供します。</p><p><a href="https://www.elastic.co/enterprise-search">Elasticsearch</a>は<a href="https://elastic.co/">Elastic</a>によって提供されています。Elastic は、精度の高い全文検索、意味理解のためのベクトル検索、両方の長所を生かしたハイブリッド検索をサポートするスケーラブルなデータ ストアおよびベクトル データベースである Elasticsearch を提供する業界リーダーです。Elasticsearch は、分散型の RESTful 検索および分析エンジン、スケーラブルなデータ ストア、およびベクター データベースです。このブログで使用している Elasticsearch 機能は、Elasticsearch の無料およびオープン バージョンで利用できます。</p><p><a href="https://www.promptingguide.ai/techniques/rag">検索拡張生成 (RAG)</a>は、LLM に外部知識を提供してユーザーのクエリに対する応答を生成する AI テクニック/パターンです。これにより、LLM 応答を特定のコンテキストに合わせてカスタマイズし、汎用性を抑えることができます。</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.13/semantic-search.html">埋め込みは</a>、テキスト/メディアの意味を数値的に表現したものです。それらは高次元情報の低次元表現です。</p><h2>RAGとデータ保護</h2><p>一般的に、大規模言語モデル (LLM) は、インターネット データでトレーニングできるモデル内で利用可能な情報に基づいて応答を生成するのに適しています。ただし、モデル内で情報が得られないクエリの場合、LLM にはモデル内に含まれていない外部知識または特定の詳細が提供される必要がありま す。このような情報は、データベースまたは社内のナレッジ システム内にある可能性があります。検索拡張生成 (RAG) は、特定のユーザー クエリに対して、まず外部 (LLM に対して) システム (データベースなど) から関連するコンテキスト/情報を取得し、そのコンテキストをユーザー クエリとともに LLM に送信して、より具体的で関連性の高い応答を生成する手法です。</p><p>これにより、RAG テクニックは、質問への回答、コンテンツの作成、コンテキストと詳細の深い理解が役立つあらゆるアプリケーションに非常に効果的になります。</p><p>その結果、RAG パイプラインでは、PII (個人識別情報) などの内部情報や機密情報 (名前、生年月日、口座番号など) がパブリック LLM に公開されるリスクがあります。</p><p>Elasticsearch のようなベクター データベースを使用する場合、データは安全です (<a href="https://www.elastic.co/guide/en/cloud-enterprise/current/ece-configure-rbac.html">ロール ベースのアクセス制御</a>、<a href="https://www.elastic.co/search-labs/blog/dls-internal-knowledge-search">ドキュメント レベルのセキュリティ</a>などのさまざまな手段を通じて)。ただし、データを外部のパブリック LLM に送信する場合は注意が必要です。</p><p>大規模言語モデル (LLM) を使用する場合、個人を特定できる情報 (PII) と機密データを保護することは、いくつかの理由から重要です。</p><ul><li><p><strong>プライバシーコンプライアンス</strong>: 多くの地域では、ヨーロッパの一般データ保護規則 (GDPR) や米国のカリフォルニア州消費者プライバシー法 (CCPA) など、個人データの保護を義務付ける厳格な規制があります。法的責任や罰金を回避するには、これらの法律を遵守する必要があります。</p></li><li><p><strong>ユーザーの信頼</strong>: 機密情報の機密性と整合性を確保することで、ユーザーの信頼が構築されます。ユーザーは、自分のプライバシーが保護されると信じるシステムを使用したり、やり取りしたりする可能性が高くなります。</p></li><li><p><strong>データ セキュリティ</strong>: データ侵害に対する保護が不可欠です。適切な保護措置を講じずに LLM に公開された機密データは盗難や悪用される可能性があり、個人情報の盗難や金融詐欺などの潜在的な危害につながる可能性があります。</p></li><li><p><strong>倫理的な考慮事項</strong>: 倫理的には、ユーザーのプライバシーを尊重し、ユーザーのデータを責任を持って扱うことが重要です。個人情報を不適切に取り扱うと、差別、汚名、その他の社会的悪影響が生じる可能性があります。</p></li><li><p><strong>企業の評判</strong>: 機密データを保護できない企業は評判が損なわれる可能性があり、顧客や収益の喪失など、ビジネスに長期的な悪影響を及ぼす可能性があります。</p></li><li><p><strong>悪用リスクの軽減</strong>: 機密データを安全に扱うことで、偏ったデータでモデルをトレーニングしたり、データを使用して個人を操作したり害を与えたりするなど、データやモデルの悪意のある使用を防ぐことができます。</p></li></ul><p>全体として、PII と機密データの堅牢な保護は、法令遵守の確保、ユーザーの信頼の維持、データ セキュリティの確保、倫理基準の遵守、ビジネスの評判の保護、不正使用のリスクの軽減に必要です。</p><h2>簡単な要約</h2><p><a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch">前回の投稿</a>では、LlamaIndex とローカルで実行される Mistral LLM を使用しながら、Elasticsearch をベクター データベースとして RAG テクニックを使用して Q&amp;A エクスペリエンスを実装する方法について説明しました。ここではそれを基に構築します。</p><p>前回の投稿を読むことはオプションです。今回は前回の投稿で行った内容を簡単に説明/要約します。</p><p>架空の住宅保険会社のエージェントと顧客間のコールセンターの会話のサンプル データセットがありました。私たちは、「顧客はどのような水関連の問題について請求をしているのか？」といった質問に答えるシンプルな RAG アプリケーションを構築しました。</p><p>大まかに言うと、フローは次のようになります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" alt="RAGフロー" /><p>インデックス作成フェーズでは、LlamaIndex パイプラインを使用してドキュメントを読み込み、インデックスを作成しました。ドキュメントはチャンク化され、埋め込みとともに Elasticsearch ベクター データベースに保存されました。</p><p>ユーザーが質問したクエリフェーズで、LlamaIndex はクエリに関連する上位 K 件の類似ドキュメントを取得しました。これらの上位 K 件の関連ドキュメントはクエリとともに、ローカルで実行されている Mistral LLM に送信され、ユーザーに返される応答が生成されました。ぜひ前回の投稿をご覧になったり、<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/tree/main">コードを調べたりしてください</a>。</p><p>前回の投稿では、LLM をローカルで実行しました。ただし、本番環境では、 <a href="https://openai.com/">OpenAI</a> 、 <a href="https://mistral.ai/">Mistral</a> 、 <a href="https://www.anthropic.com/claude">Anthropic</a>などのさまざまな企業が提供する外部 LLM を使用する必要がある場合があります。ユースケースでより大きな基礎モデルが必要であるか、スケーラビリティ、可用性、パフォーマンスなどのエンタープライズ生産のニーズによりローカルでの実行がオプションではないことが原因である可能性があります。</p><p>RAG パイプラインに外部 LLM を導入すると、機密情報や PII が LLM に誤って漏洩するリスクが生じます。この投稿では、ドキュメントを外部 LLM に送信する前に、RAG パイプラインの一部として PII 情報をマスクする方法について説明します。</p><h2>公立LLM取得のRAG</h2><p>RAG パイプラインで PII と機密情報を保護する方法について説明する前に、まず LlamaIndex、Elasticsearch Vector データベース、OpenAI LLM を使用してシンプルな RAG アプリケーションを構築します。</p><h3>要件</h3><p>以下のものが必要となります。</p><ul><li><p>埋め込みを保存するためのベクター データベースとして<strong>Elasticsearch を</strong>起動して実行します。<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#install-elasticsearch">Elasticsearch のインストール</a>に関する前回の投稿の手順に従ってください。</p></li><li><p>AI API キーを開きます。</p></li></ul><h3>シンプルなRAGアプリケーション</h3><p>参考までに、コード全体はこの<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/tree/protecting-pii">Githubリポジトリ</a>（branch:protecting-pii）にあります。以下のコードを見ていくので、リポジトリのクローンは任意です。</p><p>お気に入りの IDE で、以下の 3 つのファイルを使用して新しい Python アプリケーションを作成します。</p><ul><li><p><code>index.py</code> データのインデックス作成に関連するコードが配置される場所。</p></li><li><p><code>query.py</code> クエリと LLM の相互作用に関連するコードが配置される場所。</p></li><li><p><code>.env</code> API キーなどの構成プロパティが配置される場所。</p></li></ul><p>いくつかのパッケージをインストールする必要があります。まず、アプリケーションのルート フォルダーに新しい Python<a href="https://docs.python.org/3/library/venv.html">仮想環境</a>を作成します。</p>python3 -m venv .venv
<p>仮想環境をアクティブ化し、以下の必要なパッケージをインストールします。</p>source .venv/bin/activate
pip install llama-index 
pip install llama-index-embeddings-openai
pip install llama-index-vector-stores-elasticsearch
pip install sentence-transformers
pip install python-dotenv
pip install openai
<p>.envでOpenAIとElasticsearchの接続プロパティを設定するファイル。</p>OPENAI_API_KEY="REPLACEME"
ELASTIC_CLOUD_ID="REPLACEME"
ELASTIC_API_KEY="REPLACEME"
<h4>データのインデックス作成</h4><p>架空の住宅保険会社の顧客とコールセンターエージェント間の<em> 会話</em> が含まれる<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/blob/main/conversations.json"> conversations.json ファイルをダウンロードします。</a>アプリケーションのルートディレクトリに、2つのPythonファイルと.envファイルと一緒にファイルを配置します。先ほど作成したファイル。以下はファイルの内容の例です。</p>{
"conversation_id": 103,
"customer_name": "Sophia Jones",
"agent_name": "Emily Wilson",
"policy_number": "JKL0123",
"conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
"summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
<p>データのインデックス作成を処理する以下のコードを<code>index.py</code>に貼り付けます。</p># index.py
# pip install sentence-transformers
# pip install llama-index-embeddings-openai
# pip install llama-index-embeddings-huggingface

import json
import os
from dotenv import load_dotenv
from llama_index.core import Document
from llama_index.core import Settings
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore


def get_documents_from_file(file):
   """Reads a json file and returns list of Documents"""

   with open(file=file, mode='rt') as f:
       conversations_dict = json.loads(f.read())

   # Build Document objects using fields of interest.
   documents = [Document(text=item['conversation'],
                         metadata={"conversation_id": item['conversation_id']})
                for
                item in conversations_dict]
   return documents

# Load .env file contents into env
load_dotenv('.env')
Settings.embed_model = HuggingFaceEmbedding(
   model_name="BAAI/bge-small-en-v1.5"
)

def main():
   # ElasticsearchStore is a VectorStore that
   # takes care of Elasticsearch Index and Data management.
   es_vector_store = ElasticsearchStore(index_name="convo_index",
                                        vector_field='conversation_vector',
                                        text_field='conversation',
                                        es_cloud_id=os.getenv("ELASTIC_CLOUD_ID"),
                                        es_api_key=os.getenv("ELASTIC_API_KEY"))

   # LlamaIndex Pipeline configured to take care of chunking, embedding
   # and storing the embeddings in the vector store.
   llamaindex_pipeline = IngestionPipeline(
       transformations=[
           SentenceSplitter(chunk_size=350, chunk_overlap=50),
           Settings.embed_model
       ],
       vector_store=es_vector_store
   )

   # Load data from a json file into a list of LlamaIndex Documents
   documents = get_documents_from_file(file="conversations.json")
   llamaindex_pipeline.run(documents=documents)
   print(".....Indexing Data Completed.....\n")

if __name__ == "__main__":
   main()
<p>上記のコードを実行すると、Elasticsearch にインデックスが作成され、 <code>convo_index</code>という名前の Elasticsearch インデックスに埋め込みが保存されます。</p><p>LlamaIndex IngestionPipeline に関する説明が必要な場合は、 <a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#indexing-data">IngestionPipeline の作成</a>セクションの前の投稿を参照してください。</p><h4>クエリ</h4><p>前回の投稿では、<a href="https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch#querying">クエリ</a>にローカル LLM を使用しました。</p><p>この投稿では、以下に示すように、公開 LLM である OpenAI を使用します。</p># query.py
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.llms.openai import OpenAI
from index import es_vector_store

# Public LLM where we send user query and Related Documents
llm = OpenAI()

index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents are sent as-is. So any PII/Sensitive data is sent to the LLM.
query_engine = index.as_query_engine(llm, similarity_top_k=10)

query="Give me summary of water related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>上記のコードは、OpenAI からの応答を以下のように出力します。</p><p>顧客は、地下室の水害、水道管の破裂、屋根への雹害、適時通知の欠如、メンテナンスの問題、徐々に進行する消耗、既存の損傷などの理由による請求の拒否など、さまざまな水関連の請求を提起しています。いずれの場合も、顧客は請求の却下に対する不満を表明し、請求に関する公正な評価と決定を求めました。</p><h2>RAG での PII のマスキング</h2><p>これまで説明してきたのは、ユーザークエリとともにドキュメントをそのまま OpenAI に送信することです。</p><p>RAG パイプラインでは、関連するコンテキストが Vector ストアから取得された後、クエリとコンテキストを LLM に送信する前に、PII と機密情報をマスクする機会があります。</p><p>外部 LLM に送信する前に PII 情報をマスクする方法はいくつかあり、それぞれにメリットがあります。以下のオプションをいくつか見てみましょう</p><ol><li><p>spacy.io や<a href="https://microsoft.github.io/presidio/">Presidio</a> (Microsoft が管理するオープン ソース ライブラリ) などの NLP ライブラリを使用します。</p></li><li><p>LlamaIndexをそのまま使用する <code>NERPIINodePostprocessor.</code></p></li><li><p>ローカルLLMの使用 <code>PIINodePostprocessor</code></p></li></ol><p>上記のいずれかの方法を使用してマスキング ロジックを実装したら、PostProcessor (独自のカスタム PostProcessor または LlamaIndex が提供するすぐに使用できる PostProcessor) を使用して LlamaIndex IngestionPipeline を構成できます。</p><h3>NLPライブラリの使用</h3><p>RAG パイプラインの一部として、NLP ライブラリを使用して機密データをマスクできます。このデモでは、spacy.io パッケージを使用します。</p><p>新しいファイル<code>query_masking_nlp.py</code>を作成し、以下のコードを追加します。</p># query_masking_nlp.py

# pip install spacy
# python3 - m spacy download en_core_web_sm
import re
from typing import List, Optional

import spacy
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor.types import BaseNodePostprocessor
from llama_index.core.schema import NodeWithScore
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai import OpenAI
from index import es_vector_store

# Load the spaCy model
nlp = spacy.load("en_core_web_sm")

# Compile regex patterns for performance
phone_pattern = re.compile(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b')
email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b')
date_pattern = re.compile(r'\b(\d{1,2}[-/]\d{1,2}[-/]\d{2,4}|\d{2,4}[-/]\d{1,2}[-/]\d{1,2})\b')
dob_pattern = re.compile(
r"(January|February|March|April|May|June|July|August|September|October|November|December)\s(\d{1,2})(st|nd|rd|th),\s(\d{4})")
address_pattern = re.compile(r'\d+\s+[\w\s]+\,\s+[A-Za-z]+\,\s+[A-Z]{2}\s+\d{5}(-\d{4})?')
zip_code_pattern =  re.compile(r'\b\d{5}(?:-\d{4})?\b')
policy_number_pattern = re.compile(r"[A-Z]{3}\d{4}\.$")  # 3 characters followed by 4 digits, in our case e.g XYZ9876

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# match = re.match(policy_number_pattern, "XYZ9876")
# print(match)


def mask_pii(text):
   """
   Masks Personally Identifiable Information (PII) in the given
   text using pre-defined regex patterns and spaCy's named entity recognition.
   Args:
       text (str): The input text containing potential PII.
   Returns:
       str: The text with PII masked.
   """

   # Process the text with spaCy for NER
   doc = nlp(text)

   # Mask entities identified by spaCy NER (e.g First/Last Names etc)
   for ent in doc.ents:
       if ent.label_ in ["PERSON", "ORG", "GPE"]:
           text = text.replace(ent.text, '[MASKED]')

   # Apply regex patterns after NER to avoid overlapping issues
   text = phone_pattern.sub('[PHONE MASKED]', text)
   text = email_pattern.sub('[EMAIL MASKED]', text)
   text = date_pattern.sub('[DATE MASKED]', text)
   text = address_pattern.sub('[ADDRESS MASKED]', text)
   text = dob_pattern.sub('[DOB MASKED]', text)
   text = zip_code_pattern.sub('[ZIP MASKED]', text)
   text = policy_number_pattern.sub('[POLICY MASKED]', text)

   return text


class CustomPostProcessor(BaseNodePostprocessor):
   """
   Custom Postprocessor which masks Personally Identifiable Information (PII).
   PostProcessor is called on the Documents before they are sent to the LLM.
   """
   def _postprocess_nodes(
           self, nodes: List[NodeWithScore], query_bundle: Optional[QueryBundle]
   ) -&gt; List[NodeWithScore]:
       # Masks PII
       for n in nodes:
          n.node.set_content(mask_pii(n.text))
       return nodes

   
# Use Public LLM to send user query and Related Documents
llm = OpenAI()
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents are masked based on custom logic defined in CustomPostProcessor._postprocess_nodes.
query_engine = index.as_query_engine(llm, similarity_top_k=10, node_postprocessors=[CustomPostProcessor()])



query = "Give me summary of water related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
response = query_engine.query(bundle)
print(response)

<p>LLM による応答を以下に示します。</p>顧客からは、地下室の水害、水道管の破裂、屋根への雹害、大雨による浸水など、さまざまな水関連のクレームが出ています。これらの請求は、適時の通知の欠如、メンテナンスの問題、徐々に進行する消耗、既存の損傷などの理由に基づいて請求が拒否されたために、フラストレーションを招いています。顧客は、こうした請求拒否の結果、失望、ストレス、経済的負担を感じており、請求の公正な評価と徹底的な見直しを求めています。一部の顧客は保険金請求処理の遅延にも直面しており、保険会社が提供するサービスに対するさらなる不満が生じています。<p>上記のコードでは、Llama Index QueryEngine を作成するときに CustomPostProcessor を指定します。</p><p>QueryEngine によって呼び出されるロジックは、 <code>CustomPostProcessor</code>の<code>_postprocess_nodes</code>メソッドで定義されています。私たちは SpaCy.io ライブラリを使用して名前付きエンティティを検出し、ドキュメントを LLM に送信する前に、いくつかの正規表現を使用してそれらの名前と機密情報を置き換えます。</p><p>以下に、元の会話の一部と、CustomPostProcessor によって作成されたマスクされた会話の例を示します。</p><p>原文:</p>顧客: こんにちは。私はマシュー ロペスです。生年月日は 1984 年 10 月 12 日、住所は 456 Cedar St, Smalltown, NY 34567 です。私の保険証券番号はTUV8901です。エージェント: こんにちは、マシュー。本日はどのようなご用件でしょうか？顧客: こんにちは。私の請求を却下するという貴社の決定に、大変失望しております。<p>CustomPostProcessor によってマスクされたテキスト。</p>顧客: こんにちは。私は [MASKED] です。[MASKED] は [DOB MASKED] で、456 Cedar St, [MASKED], [MASKED] 34567 に住んでいます。私の保険証券番号は[MASKED]です。エージェント: こんにちは、[MASKED]。本日はどのようなご用件でしょうか？顧客: こんにちは。私の請求を却下するという貴社の決定に、大変失望しております。<p>注記：</p><p><em>個人情報や機密情報を識別してマスキングすることは簡単な作業ではありません。機密情報のさまざまな形式とセマンティクスをカバーするには、ドメインとデータに関する十分な理解が必要です。上記のコードは一部のユースケースでは機能する可能性がありますが、ニーズとテストに基づいて変更する必要がある場合もあります。</em></p><h3>LlamaIndexをそのまま使用する <code>NERPIINodePostprocessor</code></h3><p>LlamaIndexは、RAGパイプラインでPII情報を保護しやすくするために、 <code>NERPIINodePostprocessor.</code></p>from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor import NERPIINodePostprocessor
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.openai import OpenAI
from index import es_vector_store

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# Use Public LLM to send user query and Related Documents
llm = OpenAI()

ner_processor = NERPIINodePostprocessor()
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the LLM.
# Note that documents masked using the NERPIINodePostprocessor so that PII/Sensitive data is not sent to the LLM.
query_engine = index.as_query_engine(llm, similarity_top_k=10, node_postprocessors=[ner_processor])

query = "Give me summary of fire related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
response = query_engine.query(bundle)
print(response)
<p>応答は以下のとおりです</p>顧客は、火災により所有物件に損害が生じたとして損害賠償請求を起こした。あるケースでは、放火が補償範囲外であったため、ガレージの火災による損害に対する請求が却下されました。別の顧客は、保険でカバーされていた自宅の火災による損害について請求をしました。さらに、顧客がキッチンの火災を報告し、火災による損害が補償されることが保証されました。<h3>ローカルLLMの使用 <code>PIINodePostprocessor</code></h3><p>また、データをパブリック LLM に送信する前に、ローカルまたはプライベート ネットワークで実行されている LLM を活用してマスキング作業を行うこともできます。</p><p>マスキングを行うには、ローカル マシン上の Ollama で実行されている Mistral を使用します。</p><h4>Mistralをローカルで実行する</h4><p><a href="https://ollama.com/">Ollama</a>をダウンロードしてインストールします。Ollamaをインストールした後、このコマンドを実行して<a href="https://ollama.com/library/mistral">mistralを</a>ダウンロードして実行します。</p>ollama run mistral
<p>モデルを初めてローカルにダウンロードして実行するには数分かかる場合があります。以下のような「雲についての詩を書いてください」という質問をして、ミストラルが実行されているかどうかを確認し、その詩が気に入ったものかどうかを確認します。後でコードを通じてミストラル モデルとやり取りする必要があるため、ollama を実行したままにしておきます。</p><p><code>query_masking_local_LLM.py</code>という新しいファイルを作成し、以下のコードを追加します。</p># pip install llama-index-llms-ollama
from llama_index.core import VectorStoreIndex, QueryBundle, Settings
from llama_index.core.postprocessor import PIINodePostprocessor
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.ollama import Ollama
from llama_index.llms.openai import OpenAI
from index import es_vector_store

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")

# Use Public LLM to send user query and Related Documents and Local LLM to mask
public_llm = OpenAI()
local_llm = Ollama(model="mistral")

pii_processor = PIINodePostprocessor(llm=local_llm)
index = VectorStoreIndex.from_vector_store(es_vector_store)

# This query_engine, for a given user query retrieves top 10 similar documents from
# Elasticsearch vector database and sends the documents along with the user query to the public LLM.
# Note that documents are masked using the local llm via PIINodePostprocessor
# so that PII/Sensitive data is not sent to the public LLM.
query_engine = index.as_query_engine(public_llm, similarity_top_k=10, node_postprocessors=[pii_processor])


query = "Give me summary of fire related claims that customers raised."
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>応答は以下のようなものです</p>顧客は、火災により所有物件に損害が生じたとして損害賠償請求を起こした。あるケースでは、放火が補償範囲外であったため、ガレージの火災による損害に対する請求が却下されました。別の顧客は、保険でカバーされていた自宅の火災による損害について請求をしました。さらに、顧客がキッチンの火災を報告し、火災による損害が補償されることが保証されました。<h3>まとめ</h3><p>この投稿では、RAG フロー内でパブリック LLM を使用する際に PII と機密データを保護する方法を説明しました。私たちはそれを実現する複数の方法を実証しました。採用する前に、ユースケースとニーズに基づいてこれらのアプローチをテストすることを強くお勧めします。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/rag-security-masking-pii</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/rag-security-masking-pii</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" length="0" type="image/png"/>
    <pubDate>Thu, 25 Jul 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[LlamaIndex、Elasticsearch、Mistral を使用した RAG (検索拡張生成)]]></title>
    <description><![CDATA[LlamaIndex、Elasticsearch、ローカルで実行される Mistral を使用して RAG (Retrieval Augmented Generation) システムを実装する方法を学びます。]]></description>
    <content:encoded><![CDATA[<p>このブログでは、Elasticsearch をベクター データベースとして使用し、RAG テクニック (Retrieval Augmented Generation) を使用して Q&amp;A エクスペリエンスを実装する方法について説明します。LlamaIndex とローカルで実行されている Mistral LLM を使用します。</p><p>始める前に、いくつかの用語を見てみましょう。</p><h3>用語について</h3><p><a href="https://www.llamaindex.ai/">LlamaIndex は</a>、LLM (大規模言語モデル) アプリケーションを構築するための主要なデータ フレームワークです。LlamaIndex は、RAG (Retrieval Augmented Generation) アプリケーションの構築のさまざまな段階に抽象化を提供します。LlamaIndex や LangChain などのフレームワークは、アプリケーションが特定の LLM の API に密結合されないように抽象化を提供します。</p><p><a href="https://www.elastic.co/enterprise-search">Elasticsearch</a>は<a href="https://elastic.co/">Elastic</a>によって提供されています。Elastic は、精度の高い全文検索、意味理解のためのベクトル検索、両方の長所を生かしたハイブリッド検索をサポートする検索および分析エンジンである Elasticsearch を提供する業界リーダーです。Elasticsearch はスケーラブルなデータ ストアおよびベクター データベースです。このブログで使用している Elasticsearch の機能は、Elasticsearch の無料およびオープン バージョンで利用できます。</p><p><a href="https://www.promptingguide.ai/techniques/rag">検索拡張生成 (RAG)</a>は、LLM に外部知識を提供してユーザークエリへの応答を生成する AI テクニック/パターンです。これにより、LLM 応答を特定のコンテキストに合わせて調整できるようになり、応答がより具体的になります。</p><p><a href="https://docs.mistral.ai/">Mistral は</a>、オープンソースと最適化されたエンタープライズ グレードの LLM モデルの両方を提供します。このチュートリアルでは、ラップトップで実行されるオープン ソース モデル<a href="https://docs.mistral.ai/models/#mistral-7b">mistral-7b</a>を使用します。ラップトップでモデルを実行したくない場合は、代わりにクラウド バージョンを使用することもできます。その場合、適切な API キーとパッケージを使用するようにこのブログのコードを変更する必要があります。</p><p><a href="https://ollama.com/">Ollama は、</a>ラップトップ上で LLM をローカルに実行するのに役立ちます。Ollama を使用して、オープンソースの Mistral-7b モデルをローカルで実行します。</p><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.13/semantic-search.html">埋め込みは</a>、テキスト/メディアの意味を数値的に表現したものです。それらは高次元情報の低次元表現です。</p><h3>LlamaIndex、Elasticsearch、Mistral を使用した RAG アプリケーションの構築: シナリオの概要</h3><p><strong>シナリオ：</strong></p><p>架空の住宅保険会社のエージェントと顧客間のコールセンター会話のサンプル データセット (JSON ファイル) があります。次のような質問に答えられるシンプルなRAGアプリケーションを構築します。</p><p><code>Give me summary of water related issues.</code></p><h3>高レベルフロー</h3><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" alt="RAGフロー" /><p>Ollama を使用して Mistral LLM をローカルで実行しています。</p><p>次に、JSON ファイルからの<em>会話を</em><code>Documents</code>として<a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a> (Elasticsearch を基盤とする VectorStore) に読み込みます。ドキュメントをロードする際に、ローカルで実行されている Mistral モデルを使用して埋め込みを作成します。これらの埋め込みは<em>会話</em>とともに LlamaIndex Elasticsearch ベクター ストア ( <a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a> ) に保存されます。</p><p>LlamaIndex IngestionPipeline を設定し、使用するローカル LLM (この場合は Ollama 経由で実行される Mistral) に提供します。</p><p>「水に関する問題の概要を教えてください。」のような質問をすると、Elasticsearch はセマンティック検索を実行し、水問題に関連する<em>会話</em>を返します。これらの<em>会話は</em>元の質問とともに、ローカルで実行されている LLM に送信され、回答が生成されます。</p><h3>RAGアプリケーションの構築手順</h3><h4>Mistralをローカルで実行する</h4><p><a href="https://ollama.com/">Ollama</a>をダウンロードしてインストールします。Ollamaをインストールした後、このコマンドを実行して<a href="https://ollama.com/library/mistral">mistralを</a>ダウンロードして実行します。</p>ollama run mistral
<p>モデルを初めてローカルにダウンロードして実行するには数分かかる場合があります。以下のような「雲についての詩を書いてください」という質問をして、ミストラルが実行されているかどうかを確認し、その詩が気に入ったものかどうかを確認します。後でコードを通じてミストラル モデルとやり取りする必要があるため、ollama を実行したままにしておきます。</p><h4>Elasticsearchをインストール</h4><p>クラウド デプロイメントを作成する (<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">手順はこちら</a>) か、Docker で実行する (<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/docker">手順はこちら</a>) ことで、Elasticsearch を起動して実行します。<a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/docker#self-hosted-production-deployments">ここから</a>開始して、Elasticsearch の実稼働グレードのセルフホスト型デプロイメントを作成することもできます。</p><p>クラウド デプロイメントを使用している場合は、手順に記載されているように、デプロイメント用の API キーとクラウド ID を取得します。後で使用します。</p><h4>RAGアプリケーション</h4><p>参考までに、コード全体はこの<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany">Github リポジトリ</a>にあります。以下のコードを実行するため、リポジトリのクローン作成はオプションです。</p><p>お気に入りの IDE で、以下の 3 つのファイルを使用して新しい Python アプリケーションを作成します。</p><ul><li><p><code>index.py</code> データのインデックス作成に関連するコードが配置される場所。</p></li><li><p><code>query.py</code> クエリと LLM の相互作用に関連するコードが配置される場所。</p></li><li><p><code>.env</code> API キーなどの構成プロパティが配置される場所。</p></li></ul><p>いくつかのパッケージをインストールする必要があります。まず、アプリケーションのルート フォルダーに新しい Python<a href="https://docs.python.org/3/library/venv.html">仮想環境</a>を作成します。</p>python3 -m venv .venv
<p>仮想環境をアクティブ化し、以下の必要なパッケージをインストールします。</p>source .venv/bin/activate
pip install llama-index 
pip install llama-index-embeddings-ollama
pip install llama-index-llms-ollama
pip install llama-index-vector-stores-elasticsearch
pip install sentence-transformers
pip install python-dotenv
<h4>データのインデックス作成</h4><p>架空の住宅保険会社の顧客とコール センター エージェント間の<em> 会話</em> が含まれる<a href="https://github.com/srikanthmanvi/RAG-InsuranceCompany/blob/main/conversations.json"> conversations.json ファイルをダウンロードします。</a>アプリケーションのルートディレクトリに、2つのPythonファイルと.envファイルと一緒にファイルを配置します。先ほど作成したファイル。以下はファイルの内容の例です。</p>{
    "conversation_id": 103,
    "customer_name": "Sophia Jones",
    "agent_name": "Emily Wilson",
    "policy_number": "JKL0123",
    "conversation": "Customer: Hi, I'm Sophia Jones. My Date of Birth is November 15th, 1985, Address is 303 Cedar St, Miami, FL 33101, and my Policy Number is JKL0123.\nAgent: Hello, Sophia. How may I assist you today?\nCustomer: Hello, Emily. I have a question about my policy.\nCustomer: There's been a break-in at my home, and some valuable items are missing. Are they covered?\nAgent: Let me check your policy for coverage related to theft.\nAgent: Yes, theft of personal belongings is covered under your policy.\nCustomer: That's a relief. I'll need to file a claim for the stolen items.\nAgent: We'll assist you with the claim process, Sophia. Is there anything else I can help you with?\nCustomer: No, that's all for now. Thank you for your assistance, Emily.\nAgent: You're welcome, Sophia. Please feel free to reach out if you have any further questions or concerns.\nCustomer: I will. Have a great day!\nAgent: You too, Sophia. Take care.",
    "summary": "A customer inquires about coverage for stolen items after a break-in at home, and the agent confirms that theft of personal belongings is covered under the policy. The agent offers assistance with the claim process, resulting in the customer expressing relief and gratitude."
}
<p><code>index.py</code>に、json ファイルを読み取ってドキュメントのリストを作成する<code>get_documents_from_file</code>という関数を定義します。<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/documents_and_nodes/">ドキュメント</a>オブジェクトは、LlamaIndex が扱う情報の基本単位です。</p># index.py
import json, os
from llama_index.core import Document, Settings
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.ingestion import IngestionPipeline
from llama_index.embeddings.ollama import OllamaEmbedding
from llama_index.vector_stores.elasticsearch import ElasticsearchStore
from dotenv import load_dotenv

def get_documents_from_file(file):
   """Reads a json file and returns list of Documents"""

   with open(file=file, mode='rt') as f:
       conversations_dict = json.loads(f.read())
      
   # Build Document objects using fields of interest.
   documents = [Document(text=item['conversation'],
                         metadata={"conversation_id": item['conversation_id']})
                for
                item in conversations_dict]
   return documents
<p>IngestionPipelineを作成する</p><p>まず、 <code>Install Elasticsearch</code>セクションで取得した Elasticsearch CloudID と API キーを<code>.env</code>ファイルに追加します。<code>.env</code>ファイルは以下のようになります (実際の値を使用)。</p>ELASTIC_CLOUD_ID=&lt;REPLACE WITH YOUR CLOUD ID&gt;
ELASTIC_API_KEY=&lt;REPLACE WITH YOUR API_KEY&gt;
<p>LlamaIndex <a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/ingestion_pipeline/">IngestionPipeline を</a>使用すると、複数のコンポーネントを使用してパイプラインを構成できます。以下のコードを<code>index.py</code>ファイルに追加します。</p># index.py

# Load .env file contents into env
# ELASTIC_CLOUD_ID and ELASTIC_API_KEY are expected to be in the .env file.
load_dotenv('.env')

# ElasticsearchStore is a VectorStore that
# takes care of ES Index and Data management.
es_vector_store = ElasticsearchStore(index_name="calls",
                                     vector_field='conversation_vector',
                                     text_field='conversation',
                                     es_cloud_id=os.getenv("ELASTIC_CLOUD_ID"),
                                     es_api_key=os.getenv("ELASTIC_API_KEY"))


def main():
    # Embedding Model to do local embedding using Ollama.
    ollama_embedding = OllamaEmbedding("mistral")

    # LlamaIndex Pipeline configured to take care of chunking, embedding
    # and storing the embeddings in the vector store.
    pipeline = IngestionPipeline(
        transformations=[
            SentenceSplitter(chunk_size=350, chunk_overlap=50),
            ollama_embedding,
        ],
        vector_store=es_vector_store
    )

    # Load data from a json file into a list of LlamaIndex Documents
    documents = get_documents_from_file(file="conversations.json")

    pipeline.run(documents=documents)
    print(".....Done running pipeline.....\n")


if __name__ == "__main__":
    main()

<p>前述のように、LlamaIndex IngestPipeline は複数のコンポーネントで構成できます。パイプラインの行<code>pipeline = IngestionPipeline(...</code>に 3 つのコンポーネントを追加しています。</p><ul><li><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/?h=sentencesp#sentencesplitter">SentenceSplitter</a> : <code>get_documents_from_file()</code>の定義からわかるように、各ドキュメントには、json ファイルにある会話を保持するテキスト フィールドがあります。このテキスト フィールドは長いテキストです。セマンティック検索がうまく機能するには、小さなテキストのチャンクに分割する必要があります。<a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/?h=sentencesp#sentencesplitter">SentenceSplitter</a>クラスがこれを実行します。これらのチャンクは、LlamaIndex 用語ではノードと呼ばれます。ノードには、それが属するドキュメントを指すメタデータが存在します。あるいは、この<a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">ブログ</a>で示されているように、Elasticsearch Ingestpipeline を使用してチャンク化することもできます。</p></li><li><p><a href="https://docs.llamaindex.ai/en/stable/module_guides/models/embeddings/">OllamaEmbedding</a> : 埋め込みモデルは、テキストの一部を数値 (ベクトルとも呼ばれます) に変換します。数値表現を使用すると、単なるテキスト検索ではなく、単語の意味と一致する検索結果を表示する<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html">セマンティック検索</a>を実行できます。IngestionPipeline に<code>OllamaEmbedding("mistral")</code>を提供します。SentenceSplitter を使用して分割したチャンクは、Ollama を介してローカル マシンで実行されている Mistral モデルに送信され、Mistral によってチャンクの埋め込みが作成されます。</p></li><li><p><a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a> : LlamaIndex ElasticsearchStore ベクター ストアは、作成される埋め込みを Elasticsearch インデックスにバックアップします。ElasticsearchStore は、指定された Elasticsearch インデックスの内容の作成と入力を担当します。ElasticsearchStore ( <code>es_vector_store</code>で参照) を作成する際に、作成する Elasticsearch インデックスの名前 (この場合は<code>calls</code> )、埋め込みを保存するインデックスのフィールド (この場合は<code>conversation_vector</code> )、およびテキストを保存するフィールド (この場合は<code>conversation</code> ) を指定します。要約すると、設定に基づいて、 <code>ElasticsearchStore</code> Elasticsearch に新しいインデックスを作成し、 <code>conversation_vector</code>と<code>conversation</code>フィールドとして（他の自動作成されたフィールドとともに）使用します。</p></li></ul><p>これらすべてを結び付けて、 <code>pipeline.run(documents=documents)</code>を呼び出してパイプラインを実行します。</p><p>index.py スクリプトを実行して、取り込みパイプラインを実行します。</p>python index.py
<p>パイプラインの実行が完了すると、Elasticsearch に<code>calls</code>という新しいインデックスが表示されます。開発コンソールを使用して単純な elasticsearch クエリを実行すると、埋め込みとともに読み込まれたデータが表示されるはずです。</p>GET calls/_search?size=1
<p>これまでに行ったことをまとめると、JSON ファイルからドキュメントを作成し、それをチャンクに分割し、それらのチャンクの埋め込みを作成し、埋め込み (およびテキスト会話) をベクター ストア (ElasticsearchStore) に保存しました。</p><h4>クエリ</h4><p>llamaIndex <a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_guide/">VectorStoreIndex を</a>使用すると、関連するドキュメントを取得したり、データをクエリしたりできます。デフォルトでは、 VectorStoreIndex は埋め込みを<a href="https://docs.llamaindex.ai/en/stable/module_guides/indexing/vector_store_guide/">SimpleVectorStore</a>内のメモリ内に格納します。ただし、埋め込みを永続化するために、代わりに外部のベクトル ストア ( <a href="https://developers.llamaindex.ai/python/examples/vector_stores/elasticsearchindexdemo/">ElasticsearchStore</a>など) を使用することもできます。</p><p><code>query.py</code>を開いて以下のコードを貼り付けます</p># query.py
from llama_index.core import VectorStoreIndex, QueryBundle, Response, Settings
from llama_index.embeddings.ollama import OllamaEmbedding
from llama_index.llms.ollama import Ollama
from index import es_vector_store

# Local LLM to send user query to
local_llm = Ollama(model="mistral")
Settings.embed_model= OllamaEmbedding("mistral")

index = VectorStoreIndex.from_vector_store(es_vector_store)
query_engine = index.as_query_engine(local_llm, similarity_top_k=10)

query="Give me summary of water related issues"
bundle = QueryBundle(query, embedding=Settings.embed_model.get_query_embedding(query))
result = query_engine.query(bundle)
print(result)
<p>Ollama 上で実行されている Mistral モデルを指すようにローカル LLM ( <code>local_llm</code> ) を定義します。次に、先ほど作成した ElasticssearchStore ベクトル ストアから VectorStoreIndex ( <code>index</code> ) を作成し、インデックスからクエリ エンジンを取得します。クエリ エンジンを作成するときに、応答に使用するローカル LLM を参照し、ベクター ストアから取得して LLM に送信して応答を取得するドキュメントの数を構成する ( <code>similarity_top_k=10</code> ) も提供します。</p><p>RAG フローを実行するには、 <code>query.py</code>スクリプトを実行します。</p>python query.py
<p>クエリ<code>Give me summary of water related issues</code>を送信します ( <code>query</code>は自由にカスタマイズできます)。関連するドキュメントとともに提供される LLM からの応答は次のようになります。</p>提供されたコンテキストでは、水に関連する損害の補償について顧客が問い合わせた例がいくつかあります。2件のケースでは洪水により地下室が損傷し、別のケースでは屋根の漏水が問題となった。代理店は、両方のタイプの水害がそれぞれの保険でカバーされていることを確認しました。したがって、浸水や屋根の漏水などの水関連の問題は、通常、住宅保険でカバーされます。<h4>注意点:</h4><p>このブログ投稿は、Elasticsearch を使用した RAG テクニックの初心者向け紹介であるため、この開始点を本番環境に移行できるようにする機能の構成については省略しています。実稼働ユースケース向けに構築する場合は、<a href="https://www.elastic.co/search-labs/blog/dls-internal-knowledge-search">ドキュメント レベルのセキュリティ</a>でデータを保護したり、Elasticsearch<a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">取り込みパイプ</a>ラインの一部としてデータをチャンク化したり、GenAI/チャット/Q&amp;A ユースケースで使用されているのと同じデータで他の<a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-overview.html">ML ジョブ</a>を実行したりするなど、より高度な側面を考慮する必要があります。</p><p><a href="https://www.elastic.co/guide/en/enterprise-search/current/connectors.html">Elastic Connectors</a>を使用して、さまざまな外部ソース (Azure Blob Storage、Dropbox、Gmail など) からデータを取得し、埋め込みを作成することも検討してください。</p><p>Elastic は、上記すべてとそれ以上のことを可能にし、GenAI ユースケースなどに対応する包括的なエンタープライズ グレードのソリューションを提供します。</p><h4>What’s next?</h4><ul><li><p>お気づきかもしれませんが、応答を作成するために、ユーザーの質問とともに 10 件の関連する会話が LLM に送信されています。これらの会話には、名前、生年月日、住所などの PII (個人を特定できる情報) が含まれる場合があります。私たちの場合、LLM はローカルなので、データ漏洩は問題になりません。ただし、クラウドで実行される LLM (OpenAI など) を使用する場合は、PII 情報を含むテキストを送信することは望ましくありません。次回のブログでは、RAG フローで外部 LLM に送信する前に PII 情報をマスキングする方法について説明します。</p></li><li><p>この投稿ではローカル LLM を使用しました。RAG での PII データのマスキングに関する次の投稿では、ローカル LLM からパブリック LLM に簡単に切り替える方法について説明します。</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/rag-with-llamaIndex-and-elasticsearch</guid>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Srikanth Manvi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d3882da43bfdac0/6a17050867045bd5fe45c0e9/9d51295472f8bcca3d1973248acb724f8b94767e-1054x555.png" length="0" type="image/png"/>
    <pubDate>Fri, 12 Apr 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>