<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Priscilla Parodi - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Priscilla Parodi - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/jp/search-labs/author/priscilla-parodi</link>
    </image>
    <link>https://www.elastic.co/jp/search-labs/author/priscilla-parodi</link>
    <atom:link href="https://www.elastic.co/jp/search-labs/rss/author/priscilla-parodi.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[jp]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 12:38:41 GMT</lastBuildDate>
  <item>
    <title><![CDATA[AI盗作：Elasticsearchによる盗作検出]]></title>
    <description><![CDATA[ここでは、NLP モデルと Vector Search のユースケースに焦点を当て、Elasticsearch を使用して AI の盗用をチェックする方法を説明します。]]></description>
    <content:encoded><![CDATA[<p>盗作には、<strong>直接的な</strong>盗用（コンテンツの一部または全体をコピーする）と<strong>言い換えれば言い換え</strong>（著者の作品のいくつかの単語やフレーズを変更して言い換える）があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb6e139760e56ec98/6a171147dc55de0ad2e00edf/5d7073187fda829438aeec8d3a1194a5bea2ba57-1440x347.png" alt="" /><p>インスピレーションと言い換えには違いがあります。コンテンツを読んでインスピレーションを得て、たとえ同じような結論に達したとしても、自分の言葉でそのアイデアを探求することは可能です。</p><p>盗作は長い間議論されてきたテーマですが、コンテンツの制作と公開が加速するにつれて、盗作は関連性を保ち、継続的な課題となっています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc2af1678cb04506b/6a171149cf4f251c2ab2d257/a0b9a98d729db09dae0a79315c001e6763c12704-1400x1016.png" alt="" /><p>この課題は、盗作チェックが頻繁に行われる書籍、学術研究、司法文書に限定されません。それは新聞やソーシャルメディアにも及ぶ可能性があります。</p><p>情報が豊富で出版物に簡単にアクセスできるようになった今、スケーラブルなレベルで盗作を効果的にチェックするにはどうすればよいでしょうか?</p><p>大学、政府機関、企業はさまざまなツールを使用していますが、単純な<a href="https://www.elastic.co/search-labs/lexical-and-semantic-search-with-elasticsearch">語彙検索で</a>直接的な盗用を効果的に検出できる一方で、<strong>言い換えられたコンテンツを特定することが主な課題です。</strong></p><h2>生成AIによる盗作検出</h2><p>生成 AI によって新たな課題が生まれます。AI によって生成されたコンテンツは、コピーされた場合に盗作とみなされますか?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8a1ed1b6fa3f1a56/6a17114b4a531bc40736aa69/9345b28d6d27c37469bc38e823c41780b4eabfe5-1440x875.png" alt="" /><p>たとえば、 <a href="https://openai.com/">OpenAI の</a><a href="https://openai.com/policies/terms-of-use">利用規約</a>では、OpenAI はユーザー向けに API によって生成されたコンテンツに対する著作権を主張しないことが規定されています。この場合、生成 AI を使用する個人は、生成されたコンテンツを引用なしで自由に使用できます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbc2e08ab5b808e38/6a17114dab7f086cbadb9f93/e1f415f69a247666f02ddc81468944920c874cd7-968x814.png" alt="" /><p>しかし、効率性を向上させるために生成 AI を使用することが受け入れられるかどうかは、依然として議論の余地があります。</p><p>OpenAIは盗作検出に貢献しようと<a href="https://huggingface.co/roberta-base-openai-detector">検出モデル</a>を開発したが、後にその精度が十分に高くないことを認めた。</p><p><em>「これは単独の検出では精度が十分ではなく、より効果を上げるにはメタデータベースのアプローチ、人間の判断、一般の教育と組み合わせる必要があると考えています。」</em></p><p>課題は依然として残っていますが、利用できるツールが増えたことにより、言い換えられたコンテンツや AI コンテンツの場合でも盗作を検出するための選択肢が増えています。</p><h2>Elasticsearchで盗作を検出する</h2><p>これを認識して、このブログでは、メタデータ検索を超えて、自然言語処理 (NLP) モデルとベクトル検索、盗作検出を使用したもう 1 つのユース ケースを検討します。</p><p>これは<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/plagiarism-detection-with-elasticsearch/plagiarism_detection_es.ipynb"> Python の例</a> で実証されており、NLP 関連の記事を含む<a href="https://www.sbert.net/"> SentenceTransformers</a> の<a href="https://sbert.net/datasets/emnlp2016-2018.json"> データセット を活用しています。</a>以前に Elasticsearch にインポートされた<a href="https://huggingface.co/sentence-transformers/all-mpnet-base-v2">テキスト埋め込みモデル</a>で生成された「要約」埋め込みを考慮して「意味的テキスト類似性」を実行することにより、要約の盗用をチェックします。さらに、AI によって生成されたコンテンツ (AI 盗作) を識別するために、OpenAI によって開発された<a href="https://huggingface.co/roberta-base-openai-detector">NLP モデル</a>も Elasticsearch にインポートされました。</p><p>次の図はデータ フローを示しています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0464f88d3ed12070/6a17114fab7f084905db9f97/1ad89c98a2f42a497548ca3947749bad54ec1172-1440x880.png" alt="" /><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/inference-processor.html">推論プロセッサ</a> を使用した<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest.html"> 取り込みパイプラインの</a> 実行中に、「abstract」段落は 768 次元のベクトル「abstract_vector.predicted_value」にマッピングされます。</p><p>マッピング：</p>"abstract_vector.predicted_value": { # Inference results field
"type": "dense_vector", 
"dims": 768, # model embedding_size
"index": "true", 
"similarity": "dot_product" # When indexing vectors for approximate kNN search, you need to specify the similarity function for comparing the vectors.
<p>ベクトル表現間の類似性は、「類似性」<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html#dense-vector-params">パラメータ</a>を使用して定義されるベクトル類似性メトリックを使用して測定されます。</p><p><a href="https://en.wikipedia.org/wiki/Cosine_similarity">コサイン</a>はデフォルトの類似度メトリックであり、「(1 + cosine(クエリ、ベクトル)) / 2」として計算されます。元のベクトルを保持する必要があり、事前に正規化できない場合を除き、コサイン類似度を実行する最も効率的な方法は、すべてのベクトルを単位長さに正規化することです。これにより、検索中に余分なベクトルの長さの計算を実行することを回避でき、代わりに 'dot_product' を使用できます。</p><p>この同じパイプラインでは、<a href="https://huggingface.co/roberta-base-openai-detector">テキスト分類モデル</a>を含む別の推論プロセッサが、コンテンツがおそらく人間によって書かれた「本物」か、おそらく AI によって書かれた「偽物」かを検出し、各ドキュメントに「openai-detector.predicted_value」を追加します。</p><p>取り込みパイプライン:</p>client.ingest.put_pipeline( 
    id="plagiarism-checker-pipeline",
    processors = [
    {
      "inference": { #for ml models - to infer against the data that is being ingested in the pipeline
        "model_id": "roberta-base-openai-detector", #text classification model id
        "target_field": "openai-detector", # Target field for the inference results
        "field_map": { #Maps the document field names to the known field names of the model.
        "abstract": "text_field" # Field matching our configured trained model input. 
        }
      }
    },
    {
      "inference": {
        "model_id": "sentence-transformers__all-mpnet-base-v2", #text embedding model id
        "target_field": "abstract_vector", # Target field for the inference results
        "field_map": {
        "abstract": "text_field" # Field matching our configured trained model input. Typically for NLP models, the field name is text_field.
        }
      }
    }
    
  ]
)
<p>クエリ時に、同じテキスト埋め込みモデルを使用して、「query_vector_builder」オブジェクト内のクエリ「model_text」のベクトル表現も生成されます。</p><p>k 最近傍 (kNN) 検索は、類似度メトリックによって測定されたクエリ ベクトルに最も近い k ベクトルを見つけます。</p><p>各ドキュメントの _score は類似性から導き出され、スコアが大きいほどランキングが高くなります。これは、ドキュメントが意味的に類似していることを意味します。結果として、3 つの可能性を出力します。スコアが 0.9 を超える場合は「高い類似性」を考慮し、スコアが 0.7 未満の場合は「低い類似性」、それ以外の場合は「中程度の類似性」を考慮します。ユースケースに応じて、_score のどのレベルが盗作とみなされるかを決定するために、さまざまなしきい値を柔軟に設定できます。</p><p>さらに、テキスト分類が実行され、テキスト クエリ内の AI によって生成された要素もチェックされます。</p><p>クエリ:</p>from elasticsearch import Elasticsearch
from elasticsearch.client import MlClient

#duplicated text - direct plagiarism test

model_text = 'Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at http://hucvl.github.io/recipeqa.'

response = client.search(index='plagiarism-checker', size=1,
    knn={
        "field": "abstract_vector.predicted_value",
        "k": 9,
        "num_candidates": 974,
        "query_vector_builder": { #The 'all-mpnet-base-v2' model is also employed to generate the vector representation of the query in a 'query_vector_builder' object.
            "text_embedding": {
                "model_id": "sentence-transformers__all-mpnet-base-v2",
                "model_text": model_text
            }
        }
    }
)

for hit in response['hits']['hits']:
    score = hit['_score']
    title = hit['_source']['title']
    abstract = hit['_source']['abstract']
    openai = hit['_source']['openai-detector']['predicted_value']
    url = hit['_source']['url']

    if score &gt; 0.9:
        print(f"\nHigh similarity detected! This might be plagiarism.")
        print(f"\nMost similar document: '{title}'\n\nAbstract: {abstract}\n\nurl: {url}\n\nScore:{score}\n\n")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

    elif score &lt; 0.7:
        print(f"\nLow similarity detected. This might not be plagiarism.")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

    else:
        print(f"\nModerate similarity detected.")
        print(f"\nMost similar document: '{title}'\n\nAbstract: {abstract}\n\nurl: {url}\n\nScore:{score}\n\n")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

ml_client = MlClient(client)

model_id = 'roberta-base-openai-detector' #open ai text classification model

document = [
    {
        "text_field": model_text
    }
]

ml_response = ml_client.infer_trained_model(model_id=model_id, docs=document)

predicted_value = ml_response['inference_results'][0]['predicted_value']

if predicted_value == 'Fake':
    print("\nNote: The text query you entered may have been generated by AI.\n")
<p>アウトプット：</p>High similarity detected! This might be plagiarism.

Most similar document: 'RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes'

Abstract: Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at[ http://hucvl.github.io/recipeqa](http://hucvl.github.io/recipeqa).

url:[http://aclweb.org/anthology/D18-1166](http://aclweb.org/anthology/D18-1166)

Score:1.0
<p>この例では、データセットの「抽象」値の 1 つをテキスト クエリ「model_text」として使用した後、盗用が特定されました。類似度スコアは 1.0 であり、類似度が高いこと、<strong>つまり直接的な盗用であること</strong>を示しています。ベクトル化されたクエリとドキュメントは、予想どおり AI 生成コンテンツとして認識されませんでした。</p><p>クエリ:</p>#similar text - paraphrase plagiarism test 

model_text = 'Comprehending and deducing information from culinary instructions represents a promising avenue for research aimed at empowering artificial intelligence to decipher step-by-step text. In this study, we present CuisineInquiry, a database for the multifaceted understanding of cooking guidelines. It encompasses a substantial number of informative recipes featuring various elements such as headings, explanations, and a matched assortment of visuals. Utilizing an extensive set of automatically crafted question-answer pairings, we formulate a series of tasks focusing on understanding and logic that necessitate a combined interpretation of visuals and written content. This involves capturing the sequential progression of events and extracting meaning from procedural expertise. Our initial findings suggest that CuisineInquiry is poised to function as a demanding experimental platform.'
<p>アウトプット：</p>High similarity detected! This might be plagiarism.

Most similar document: 'RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes'

Abstract: Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at[ http://hucvl.github.io/recipeqa](http://hucvl.github.io/recipeqa).

url:[http://aclweb.org/anthology/D18-1166](http://aclweb.org/anthology/D18-1166)

Score:0.9302529

Note: The text query you entered may have been generated by AI.
<p>テキストクエリ「model_text」を、類似した単語の繰り返しを最小限に抑えながら同じメッセージを伝える AI 生成テキストで更新すると、検出された類似度は依然として高かったものの、スコアは 1.0 ではなく 0.9302529 になりました<strong>(言い換え盗用)</strong> 。AIによって生成されたこのクエリが検出されることも予想されました。</p><p>最後に、テキスト クエリ「model_text」を、これらのドキュメントの要約ではない Elasticsearch に関するテキストと見なすと、検出された類似度は 0.68991005 となり、考慮されたしきい値によると類似度が低いことが示されました。</p><p>クエリ:</p>#different text - not a plagiarism

model_text = 'Elasticsearch provides near real-time search and analytics for all types of data.'
<p>アウトプット：</p>Low similarity detected. This might not be plagiarism.
<p>AI によって生成されたテキスト クエリでは、言い換えや直接コピーされたコンテンツの場合と同様に、盗用が正確に識別されましたが、盗用検出の状況を把握するには、さまざまな側面を認識する必要があります。</p><p>AI 生成コンテンツの検出という観点から、私たちは価値ある貢献を果たすモデルを検討しました。ただし、単独の検出には固有の限界があり、精度を高めるには他の方法を組み込む必要があることを認識することが重要です。</p><p>テキスト埋め込みモデルの選択によってもたらされる変動性も、もうひとつの考慮事項です。異なるデータセットでトレーニングされたさまざまなモデルでは、さまざまなレベルの類似性が得られ、生成されたテキスト埋め込みの重要性が強調されます。</p><p>最後に、これらの例では、ドキュメントの要約を使用しました。ただし、盗作検出には大きな文書が関係することが多く、テキストの長さの問題に対処することが不可欠です。テキストがモデルのトークン制限を超えることはよくあり、埋め込みを構築する前にチャンクに分割する必要があります。これを処理する<a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.11/knn-search.html#nested-knn-search">実用的なアプローチ</a>としては、dense_vector を使用したネストされた構造を利用することが挙げられます。</p><h2>まとめ</h2><p>このブログでは、特に言い換えられたコンテンツや AI によって生成されたコンテンツにおける盗作の検出の課題と、この目的のために意味的なテキストの類似性とテキスト分類をどのように使用できるかについて説明しました。</p><p>これらの方法を組み合わせることで、AI によって生成されたコンテンツ、直接的な盗用と言い換えられた盗用を正常に識別した盗用検出の例を示しました。</p><p>主な目的は検出を簡素化するフィルタリング システムを確立することですが、検証には人間による評価が依然として不可欠です。</p><p>意味的テキスト類似性と NLP についてさらに詳しく知りたい場合は、次のリンクもご覧ください。</p><ul><li><p><a href="https://www.elastic.co/what-is/semantic-search">セマンティック検索とは？</a></p></li><li><p><a href="https://www.elastic.co/what-is/natural-language-processing">自然言語処理 (NLP) とは何ですか?</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch">Elasticsearch による語彙検索と意味検索</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">大規模なドキュメントをインジェストパイプラインでチャンク化し、ネストされたベクトルと組み合わせることで、簡単にパッセージを検索できます。</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-plagiarism-checker-with-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-plagiarism-checker-with-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt68a5bc2434a9b03b/6a1711510e2e49a09641a22a/83e05cd4f81799fbb7b7950ed87600e825ec81e9-1024x1024.png" length="0" type="image/png"/>
    <pubDate>Tue, 19 Dec 2023 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch による語彙検索と意味検索]]></title>
    <description><![CDATA[このブログでは、語彙検索と意味検索に焦点を当て、Elasticsearch を使用して情報を取得するためのさまざまなアプローチについて説明します。]]></description>
    <content:encoded><![CDATA[<p>検索は、検索クエリまたは複合クエリに基づいて最も関連性の高い情報を検索するプロセスであり、関連する検索結果はこれらのクエリに最も一致するドキュメントです。検索にはさまざまな課題と方法がありますが、最終的な目標は、<strong>質問に対する最適な回答を見つけることです</strong>。</p><p>この目標を考慮して、このブログ投稿では、Elasticsearch を使用して情報を取得するためのさまざまなアプローチを検討し、特にテキスト検索（<strong>語彙検索とセマンティック検索）に焦点を当てます。</strong></p><h2>要件</h2><p>これを実現するために、eコマース製品情報をシミュレートするために生成された<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/products-ecommerce.json">データセット</a>でのさまざまな検索シナリオを示す Python の例を提供します。</p><p>このデータセットには 2,500 を超える製品が含まれており、それぞれに説明が付いています。これらの製品は 76 の異なる製品カテゴリに分類されており、各カテゴリには以下に示すようにさまざまな数の製品が含まれています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt34415151c00ced2b/6a17d8710b0bed6c9bdd342c/4104466050f3024b6bcaf382da2a702650f62227-1440x708.png" alt="" /><p><em>ツリーマップの視覚化 - カテゴリ.キーワードの上位 22 個の値 (製品カテゴリ)</em></p><p>セットアップには次のものが必要です:</p><ul><li><p>Python 3.6以降</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/client/python-api/current/index.html">Elastic Pythonクライアント</a></p></li><li><p>Elastic 8.8 以降のデプロイメント、8GB メモリの機械学習ノード</p></li><li><p><a href="https://www.elastic.co/guide/en/machine-learning/8.9/ml-nlp-elser.html">Elastic Learned Sparse EncodeR</a>モデルは、Elastic にプリロードされており、デプロイメントにインストールされて起動します。</p></li></ul><p>Elastic Cloud を使用します。<a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">無料トライアルを</a>ご利用いただけます。</p><p>このブログ記事で提供されている検索クエリに加えて、 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/ecommerce_dense_sparse_project.ipynb">Python ノートブックで</a>は次のプロセスをガイドします。</p><ul><li><p>Pythonクライアントを使用してElasticデプロイメントへの接続を確立する</p></li><li><p>Elasticsearchクラスターにテキスト埋め込みモデルをロードする</p></li><li><p>特徴ベクトルと密ベクトルのインデックスを作成するためのマッピングを使用してインデックスを作成します。</p></li><li><p>テキスト埋め込みとテキスト拡張のための推論プロセッサを備えた取り込みパイプラインを作成する</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0e95406ece8f08d9/6a17d8731d1b83d32e93e2f4/54a9a490a3b0cf1b9c2228bee8eddd3f566bd435-1418x1102.png" alt="" /><h2>語彙検索 - スパース検索</h2><p>テキスト クエリに基づいて Elasticsearch がドキュメントの関連性をランク付けする従来の方法では、<strong> 語彙検索用のスパース</strong> モデルである<a href="https://en.wikipedia.org/wiki/Okapi_BM25"> BM25 モデルの</a> Lucene 実装が使用されます。この方法は、正確な用語の一致を探すという、テキスト検索の従来のアプローチに従います。</p><p>この検索を可能にするために、Elasticsearch はテキスト分析を実行して<strong>テキスト フィールド</strong>データを検索可能な形式に変換します。</p><p><strong>テキスト分析</strong>は、検索に関連するトークンを抽出するプロセスを管理する一連のルールである<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analyzer-anatomy.html">アナライザー</a>によって実行されます。アナライザーには、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenizers.html">トークナイザー</a>が 1 つだけ必要です。トークナイザーは文字のストリームを受け取り、それを個々のトークン (通常は個々の単語) に分割します。以下に例を示します。</p><h3>語彙検索のための文字列トークン化</h3>#Performs text analysis on a string and returns the resulting tokens.

# Define the text to be analyzed
text = "Comfortable furniture for a large balcony"

# Define the analyze request
request_body = {
  "analyzer": "standard",
  "text": text
}

# Perform the analyze request
response = client.indices.analyze(analyzer=request_body["analyzer"], text=request_body["text"])

# Extract and display the analyzed tokens
tokens = [token["token"] for token in response["tokens"]]
print("Analyzed Tokens:", tokens)
<p>出力</p>Analyzed Tokens: ['comfortable', 'furniture', 'for', 'a', 'large', 'balcony']
<p>この例では、デフォルトのアナライザーである<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-standard-analyzer.html">標準</a>アナライザーを使用しています。これは、英語の文法に基づいたトークン化を提供するため、ほとんどのユースケースに適しています。トークン化により、個々の用語での一致が可能になりますが、各トークンは文字どおりに一致します。</p><p>検索エクスペリエンスをカスタマイズしたい場合は、別の<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-analyzers.html">組み込みアナライザーを</a>選択できます。たとえば、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-stop-analyzer.html">ストップ アナライザーを</a>使用するようにコードを更新すると、ストップワードの削除がサポートされ、文字以外の文字でテキストがトークンに分割されます。</p>...
# Define the analyze request
request_body = {
  "analyzer": "stop",
  "text": text
}
...
<p>出力</p>Analyzed Tokens: ['comfortable', 'furniture', 'large', 'balcony']
<p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-charfilters.html">組み込みアナライザーがニーズを満たさない場合は、ゼロ個以上の 文字フィルター</a> 、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenizers.html"> トークナイザー</a> 、およびゼロ個以上の <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenfilters.html">トークン</a> フィルター の適切な組み合わせを使用する<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html"> カスタム アナライザー</a> を作成できます。</p>"analyzer":  {

  "my_analyzer": {

    "type": "custom", #For custom analyzers, use a type of custom or omit the type parameter.

    "tokenizer": "standard", #Built-in or customized tokenizer

    "filter": ["lowercase", "synonym"] #Built-in or customized token filters
  }
}
<p>トークナイザーとトークン フィルターを組み合わせた上記の例では、テキストは、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lowercase-tokenfilter.html"> </a><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-synonym-tokenfilter.html#:~:text=Elasticsearch%20will%20use%20the%20token,applied%20to%20the%20synonym%20entries.">同義語トークン フィルター によって処理される前に、</a> 小文字フィルター によって小文字化されます。</p><h2>語彙マッチング</h2><p><a href="https://www.elastic.co/blog/practical-bm25-part-2-the-bm25-algorithm-and-its-variables">BM25 は</a>、用語の頻度と重要度に基づいて、特定の検索クエリに対するドキュメントの関連性を測定します。</p><p>以下のコードは、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-match-query.html"> 一致</a> クエリを実行し、<em> 「ecommerce-search」</em> インデックスの<em> 「description」</em> <strong>フィールド値と検索クエリ</strong><em><strong> 「 Comfortable furniture for a large balcony 」</strong></em><strong> を考慮して最大</strong> 2 つのドキュメントを検索します。</p><p>このクエリに一致すると見なされるドキュメントの基準を絞り込むと、精度が向上します。ただし、より具体的な結果を得るには、変動に対する許容度が低くなるというデメリットがあります。</p># BM25

response = client.search(size=2,
index="ecommerce-search",
query= {
  "match": {
    "description" : {  
      "query": "Comfortable furniture for a large balcony",
      "analyzer": "stop"
    }
  }
}
)

hits = response['hits']['hits']

if not hits:
  print("No matches found")

else:
  for hit in hits:
    score = hit['_score']
    product = hit['_source']['product']
    category = hit['_source']['category']
    description = hit['_source']['description']
    print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>出力</p>Score: 15.607948
Product: Barbie Dreamhouse
Category: Toys
Description: is a classic Barbie playset with multiple rooms, furniture, a large balcony, a pool, and accessories. It allows kids to create their dream Barbie world.

Score: 9.137739
Product: Comfortable Rocking Chair
Category: Indoor Furniture
Description: enjoy relaxing moments with this comfortable rocking chair. Its smooth motion and cushioned seat make it an ideal piece of furniture for unwinding.
<p>出力を分析すると、最も関連性の高い結果は「<em> おもちゃ</em> 」カテゴリの「バービードリームハウス 」製品であり、その説明には「<em> 家具</em> 」、「<em> 大型」</em> 、「バルコニー 」という用語が含まれているため関連性が高く、説明に検索クエリに一致する用語が 3 つ含まれている唯一の製品であり、説明に「バルコニー」 という用語が含まれているのもこの製品のみです。</p><p>2 番目に関連性の高い製品は、「 屋内用家具」に分類される「快適なロッキングチェア 」で、その説明には「<em> 快適な</em> 」および「<em> 家具</em> 」という用語が含まれています。データセット内のこの検索クエリの少なくとも 2 つの用語に一致する製品は 3 つだけであり、この製品はそのうちの 1 つです。</p><p><em>「快適」という</em>言葉は 105 製品の説明に登場し、 <em>「家具」という</em>言葉は、<em>おもちゃ</em>、<em>屋内用家具、屋外用家具、および「犬と猫の用品とおもちゃ」という 4 つのカテゴリーの 4 製品の説明に登場しています。</em></p><p>ご覧のとおり、クエリを考慮すると最も関連性の高い製品はおもちゃであり、2 番目に関連性の高い製品は屋内用家具です。これらのドキュメントが一致する理由を知るためにスコア計算に関する詳細な情報が必要な場合は、 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-explain.html"><em>explain</em></a> __query パラメータを true に設定できます。</p><p>両方の結果が最も関連性の高いものであるにもかかわらず、このデータセット内のドキュメントの数と用語の出現の両方を考慮すると、クエリ「<em>大きなバルコニー用の快適な家具</em>」の背後にある意図は、おもちゃや室内用家具などを除いた、実際の大きなバルコニー用の家具を検索することです。</p><p>語彙検索は比較的<strong>単純かつ高速</strong>ですが、ユーザーの意図やクエリを必ずしも知らずにすべての用語と同義語を知ることが常に可能であるとは限らないため、限界があります。自然言語の使用においてよく見られる現象は<strong>語彙の不一致</strong>です。<a href="https://dl.acm.org/doi/abs/10.1145/32206.32212">調査</a>によると、平均して<strong>80% の確率で</strong>、異なる人々 (同じ分野の専門家) が同じものを異なる名前で呼ぶことがわかっています。</p><p>これらの制限により、意味的知識を組み込んだ他のスコアリング モデルを探すことになります。自然言語のような連続的な入力トークンの処理に優れたトランスフォーマーベースのモデルは、ドキュメントとクエリの両方の数学的表現を考慮して、検索の根本的な意味を捉えます。これにより、テキストの高密度でコンテキストを認識したベクトル表現が可能になり、関連するコンテンツを見つけるための洗練された方法である<strong>セマンティック検索</strong>が強化されます。</p><h2>セマンティック検索 - 高密度検索</h2><p>このコンテキストでは、データを意味のあるベクトル値に変換した後、 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html">k 最近傍 (kNN)</a>検索アルゴリズムを使用して、データセット内でクエリ ベクトルに最も類似したベクトル表現を検索します。Elasticsearch は、kNN 検索に、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#exact-knn">正確なブルート フォース kNN</a>と<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#approximate-knn">近似 kNN</a> (ANN とも呼ばれる) の 2 つの方法をサポートしています。</p><p>ブルートフォース kNN は正確な結果を保証しますが、大規模なデータセットでは適切に拡張できません。近似 kNN は、パフォーマンスを向上させるために精度をある程度犠牲にして、近似最近傍を効率的に見つけます。</p><p>Lucene の kNN 検索と高密度ベクトル インデックスのサポートにより、Elasticsearch は階層的ナビゲート可能スモール ワールド (HNSW) アルゴリズムを活用し、さまざまな<a href="http://ann-benchmarks.com/">ann ベンチマーク データセット</a>にわたって強力な検索パフォーマンスを発揮します。以下のサンプルコードを使用して、Python で近似 kNN 検索を実行できます。</p><h3>近似kNNによるセマンティック検索</h3># KNN - approximate kNN

response = client.search(index='ecommerce-search', size=2,
knn={
  "field": "description_vector.predicted_value",
  "k": 50, # Number of nearest neighbors to return as top hits.
#The optimal value of k is dependent on the data. It can vary in different scenarios.

  "num_candidates": 500, # Number of nearest neighbor candidates to consider per shard.

#Increasing num_candidates tends to improve the accuracy of the final k results.

  "query_vector_builder": { # Object indicating how to build a query_vector. kNN search enables you to perform semantic search by using a previously deployed text embedding model, the steps for this process are demonstrated in the Python notebook.
    "text_embedding": { 
      "model_id": "sentence-transformers__all-mpnet-base-v2", # Text embedding model id
      "model_text": "Comfortable furniture for a large balcony" # Query
    }
  }
}
)

for hit in response['hits']['hits']:
        
  score = hit['_score']
  product = hit['_source']['product']
  category = hit['_source']['category']
  description = hit['_source']['description']
  print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>このコード ブロックは、Elasticsearch の kNN を使用して、製品データセットの「<em> description</em><em> 」フィールドの埋め込みを考慮した「 大きなバルコニー用の快適な家具</em> 」のベクトル化されたクエリ (query_vector_build) に類似した説明を持つ最大 2 つの製品を返します。</p><p>製品の埋め込みは、以前はパイプラインに取り込まれたデータに対して推論するための<em>「</em> <a href="https://huggingface.co/sentence-transformers/all-mpnet-base-v2"><em>all-mpnet-base-v2</em></a> <em>」</em>テキスト埋め込みモデルを含む推論プロセッサを使用して、取り込みパイプラインで生成されていました。</p><p>このモデルは、 <em>「</em> <a href="https://github.com/UKPLab/sentence-transformers/blob/master/docs/package_reference/sentence_transformer/evaluation.md"><em>sentence_transformers.evaluation</em></a> <em>」</em>を使用した事前学習済みモデルの評価に基づいて選択されました。トレーニング中にさまざまなクラスを使用してモデルを評価します。「all-mpnet-base-v2」モデルは<a href="https://www.sbert.net/docs/pretrained_models.html">、Sentence-Transformers ランキング</a>で最高の平均パフォーマンスを示し、 <a href="https://huggingface.co/spaces/mteb/leaderboard">Massive Text Embedding Benchmark (MTEB)</a>リーダーボードでも好位置を獲得しました。このモデルは<a href="https://huggingface.co/microsoft/mpnet-base">、Microsoft/mpnet ベースの</a>モデルを事前トレーニングし、10 億の文のペアのデータセットで微調整されており、文を 768 次元の密なベクトル空間にマッピングします。</p><p>あるいは、特にドメイン固有のデータに合わせて微調整されたモデルなど、使用できる他のモデルも多数あります。</p><p>出力</p>Score: 0.79207325
Product: Patio Sofa Set with Ottoman
Category: Outdoor Furniture
Description: is a versatile and comfortable patio sofa set, including a sofa, ottoman, and coffee table, great for outdoor lounging.

Score: 0.7836937
Product: Patio Sofa Set with Canopy
Category: Outdoor Furniture
Description: is a luxurious and comfortable patio sofa set with a canopy, providing shade and style for outdoor lounging.
<p><em>出力は、選択したモデル、</em><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search-filter-example"><em> フィルター</em></a><em> 、および</em><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#tune-approximate-knn-for-speed-accuracy"><em> おおよその kNN チューニング</em></a><em> によって異なる場合があります 。</em></p><p>kNN 検索結果は両方とも「 <em>Outdoor Furniture</em> 」カテゴリにありますが、クエリの一部として「 <em>outdoor</em> 」という単語が明示的に言及されておらず、コンテキストにおけるセマンティクスの理解の重要性が強調されています。</p><p>高密度ベクトル検索にはいくつかの利点があります。</p><ul><li><p>セマンティック検索の有効化</p></li><li><p>非常に大規模なデータセットを処理できるスケーラビリティ</p></li><li><p>幅広いデータタイプを処理できる柔軟性</p></li></ul><p>しかし、<strong>高密度ベクトル探索には独自の課題もあります</strong>。</p><ul><li><p>ユースケースに適した埋め込みモデルの選択</p></li><li><p>モデルが選択されると、ドメイン固有のデータセットでのパフォーマンスを最適化するためにモデルを微調整する必要がある場合があり、このプロセスにはドメイン専門家の関与が求められる。</p></li><li><p>さらに、高次元ベクトルのインデックス作成は計算コストが高くなる可能性がある。</p></li></ul><h2>セマンティック検索 - 学習されたスパース検索</h2><p>別のアプローチとして、セマンティック検索を実行する別の方法である学習済みスパース検索を検討してみましょう。</p><p>スパースモデルとして、数十年にわたる最適化の恩恵を受けている Elasticsearch の Lucene ベースの転置インデックスを活用します。ただし、このアプローチは、BM25 などの語彙スコアリング関数を使用して同義語を単純に追加するだけにとどまりません。代わりに、より深い言語スケールの知識を使用して学習した関連性を組み込み、関連性を最適化します。</p><p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-elser.html">Elastic Learned Sparse Encoder は</a>、検索クエリを拡張して元のクエリには存在しない関連用語を含めることにより、以下の例に示すように、<strong>スパース ベクトル埋め込みを改善します</strong>。</p><h3>Elastic Learned Sparse Encoder によるスパースベクトル検索</h3># Elastic Learned Sparse Encoder

response = client.search(index='ecommerce-search', size=2,
query={
  "text_expansion": {
    "ml.tokens": {
      "model_id":"elser_model",
      "model_text":"Comfortable furniture for a large balcony"                
    }
  }
}
)

for hit in response['hits']['hits']:

  score = hit['_score']
  product = hit['_source']['product']
  category = hit['_source']['category']
  description = hit['_source']['description']
  print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>出力</p>Score: 14.405318
Product: Garden Lounge Set with Side Table
Category: Garden Furniture
Description: is a comfortable and stylish garden lounge set, including a sofa, chairs, and a side table for outdoor relaxation.

Score: 14.281318
Product: Rattan Patio Conversation Set
Category: Outdoor Furniture
Description: is a stylish and comfortable outdoor furniture set, including a sofa, two chairs, and a coffee table, all made of durable rattan material.
<p>この場合の結果には、「<em> 屋外用家具</em> 」に非常に類似した製品を提供する「<em> ガーデン家具</em> 」カテゴリが含まれます。</p><p>「ml.tokens」を分析すると、学習済みスパース検索によって生成されたトークンを含む「rank_features」フィールドを見ると、生成されたさまざまなトークンの中に、「<em>リラックス</em>」（快適）、「<em>ソファ</em>」（家具）、「<em>屋外</em>」（バルコニー）など、検索クエリの一部ではないものの、意味的には関連している用語があることがわかります。</p><p>以下の画像は、用語拡張ありとなしの両方で、クエリとともにこれらの用語の一部を強調表示しています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7e869985347b19a7/6a17d875e31791dc2c2d56a1/dd86607fce6137d843a3ec002390eaa988b432f9-1440x502.png" alt="" /><p>ご覧のとおり、このモデルはコンテキスト認識型の検索を提供し、より解釈しやすい結果を提供しながら語彙の不一致の問題を軽減するのに役立ちます。ドメイン固有の再トレーニングが適用されない場合でも、高密度ベクトル モデルよりも優れたパフォーマンスを発揮します。</p><h2>ハイブリッド検索: 語彙検索と意味検索を組み合わせた関連性の高い結果</h2><p>検索に関しては、普遍的な解決策はありません。これらの検索方法にはそれぞれ長所がありますが、課題もあります。ユースケースに応じて、最適なオプションは変わる場合があります。多くの場合、複数の検索方法間で最良の結果が補完的になります。したがって、関連性を高めるために、それぞれの方法の長所を組み合わせることを検討します。</p><p><strong>ハイブリッド検索を</strong>実装する方法は複数あります。線形結合、各スコアへの重み付け、重みの指定が不要な逆ランク融合 (RRF) などです。</p><h3>Elasticsearch: 語彙検索と意味検索の両方の長所を活用</h3># BM25 + Elastic Learned Sparse Encoder (Linear Combination)

response = client.search(index='ecommerce-search', size=2,

query= {
  "bool": {
    "should": [
    {
      "match": {
        "description" : {  
          "query": "A dining table and comfortable chairs for a large balcony",
          "boost": 1
        }
      }
    },                   
    {
      "text_expansion": {
        "ml.tokens": {
          "model_id": "elser_model",
          "model_text": "A dining table and comfortable chairs for a large balcony",
          "boost": 1
        }
      }
     }
    ]
  }
}
)

# The boost value is 1 for the text expansion and match query. This means that the relevance score of the results of these queries are not boosted. You can specify a boost value to give a weight to each score in the sum. The scores will be calculated as: score = boost value * match_score + boost value * text_expansion_score

for hit in response['hits']['hits']:

  score = hit['_score']
  product = hit['_source']['product']
  category = hit['_source']['category']
  description = hit['_source']['description']
  print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>このコードでは、「<em>大きなバルコニー用のダイニング テーブルと快適な椅子</em>」という値を持つ 2 つのクエリを使用してハイブリッド検索を実行しました。検索語として「<em>家具</em>」を使用する代わりに、探しているものを指定しており、両方の検索で同じフィールド値「説明」を考慮しています。ランキングは、BM25 スコアと ELSER スコアに等しい重み付けをした線形結合によって決定されます。</p><p>出力</p>Score: 31.628141
Product: Garden Dining Set with Swivel Rockers
Category: Garden Furniture
Description: is a functional and comfortable garden dining set, including a table and chairs with swivel rockers for easy movement.

Score: 31.334227
Product: Garden Dining Set with Swivel Chairs
Category: Garden Furniture
Description: is a functional and comfortable garden dining set, including a table and chairs with swivel seats for convenience.
<p>以下のコードでは、クエリに同じ値を使用しますが、逆ランク融合法を使用して BM25 (クエリ パラメータ) と kNN (knn パラメータ) のスコアを結合し、ドキュメントを結合してランク付けします。</p># BM25 + KNN (RRF)

response = client.search(index='ecommerce-search', size=2,
query={
  "bool": {
    "should": [
    {
      "match": {
        "description": {
        "query": "A dining table and comfortable chairs for a large balcony"
        }
      }
    }
    ]
  }
},
knn={
  "field": "description_vector.predicted_value",
  "k": 50,
  "num_candidates": 500,
  "query_vector_builder": {
    "text_embedding": {
      "model_id": "sentence-transformers__all-mpnet-base-v2",
      "model_text": "A dining table and comfortable chairs for a large balcony"
    }
  }
},
rank={
  "rrf": { # Reciprocal rank fusion
    "window_size": 50, # This value determines the size of the individual result sets per query.
    "rank_constant": 20 # This value determines how much influence documents in individual result sets per query have over the final ranked result set.
  }
}
)

for hit in response['hits']['hits']:
        
  rank = hit['_rank']
  category = hit['_source']['category']
  product = hit['_source']['product']
  description = hit['_source']['description']
  print(f"\nRank: {rank}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p><em>RRF 機能はテクニカル プレビュー段階です。GA の前に構文が変更される可能性があります。</em></p><p>出力</p>Rank: 1
Product: Patio Dining Set with Bench
Category: Outdoor Furniture
Description: is a spacious and functional patio dining set, including a dining table, chairs, and a bench for additional seating.

Rank: 2
Product: Garden Dining Set with Swivel Chairs
Category: Garden Furniture
Description: is a functional and comfortable garden dining set, including a table and chairs with swivel seats for convenience.
<p>ここでは、異なるフィールドと値を使用することもできます。これらの例のいくつかは、 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/ecommerce_dense_sparse_project.ipynb">Python ノートブック</a>で利用できます。</p><p>ご覧のとおり、Elasticsearch を使用すると、従来の語彙検索とベクトル検索（疎かでも密かでも）の両方の利点を活用でき、目標を達成し<strong>て質問に対する最適な回答を見つけることができます。</strong></p><p>ここで説明したアプローチについてさらに学習したい場合は、次のブログが役立ちます。</p><ul><li><p><a href="https://www.elastic.co/blog/improving-information-retrieval-elastic-stack-hybrid">Elastic Stackでの情報検索の改善：ハイブリッド検索</a></p></li><li><p><a href="https://www.elastic.co/blog/vector-search-elasticsearch-rationale">Elasticsearchのベクトル検索：設計の背後にある理論的根拠</a></p></li><li><p><a href="https://www.elastic.co/blog/lexical-ai-powered-search-elastic-vector-database">Elasticのベクターデータベースで語彙検索とAIを活用した検索を最大限に活用する方法</a></p></li><li><p><a href="https://www.elastic.co/blog/may-2023-launch-sparse-encoder-ai-model">Elastic Learned Sparse Encoder のご紹介: セマンティック検索のための Elastic の AI モデル</a></p></li><li><p><a href="https://www.elastic.co/blog/may-2023-launch-information-retrieval-elasticsearch-ai-model">Elastic Stackでの情報検索の改善: 新しい検索モデルElastic Learned Sparse Encoderのご紹介</a></p></li></ul><p>Elasticsearch は、ベクター検索を構築するために必要なすべてのツールとともに、ベクター データベースを提供します。</p><ul><li><p>Elasticsearch<a href="https://www.elastic.co/elasticsearch/vector-database">ベクターデータベース</a></p></li><li><p>Elasticによる<a href="https://www.elastic.co/enterprise-search/vector-search">ベクトル検索のユース</a>ケース</p></li></ul><h2>まとめ</h2><p>このブログ記事では、Elasticsearch を使用して情報を取得するためのさまざまなアプローチについて検討し、特にテキスト、語彙、意味の検索に焦点を当てました。これを実証するために、eコマース製品情報を含むデータセットを使用してさまざまな検索シナリオを紹介する Python の例を示しました。</p><p>BM25 を使用した従来の語彙検索をレビューし、語彙の不一致などの利点と課題について議論しました。この問題を克服するために、意味的知識を取り入れることの重要性を強調しました。さらに、セマンティック検索を可能にする高密度ベクトル検索について説明し、高次元ベクトルのインデックス作成時の計算コストなど、この検索方法に関連する課題についても説明しました。</p><p>一方、スパースベクトルは圧縮率が非常に高いことを説明しました。そこで、元のクエリには存在しない関連用語を含めるように検索クエリを拡張する Elastic の Learned Sparse Encoder について説明しました。</p><p>検索に関しては、万能の解決策は存在しません。それぞれの検索方法には長所と課題があります。そこで、ハイブリッド検索の概念についても議論しました。</p><p>ご覧のとおり、Elasticsearch を使用すると、従来の語彙検索とベクトル検索の両方の長所を活用できます。</p><p>始める準備はできましたか?利用可能な<a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/ecommerce_dense_sparse_project.ipynb">Python ノートブック</a>を確認し、 <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">Elastic Cloud の無料トライアル</a>を開始してください。</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Python]]></category>
    <category><![CDATA[クエリー言語]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbd7e961f594005e7/6a17d80f033c8d981f6bb009/d240bfef29e9d432069059b312dd044eb76eec6c-1440x840.png" length="0" type="image/png"/>
    <pubDate>Tue, 03 Oct 2023 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch の NLP とベクトル検索によるチャットボット機能の強化]]></title>
    <description><![CDATA[ベクトル検索と NLP がどのように機能してチャットボットの機能を強化するかを探り、Elasticsearch がどのようにプロセスを促進するかを確認します。]]></description>
    <content:encoded><![CDATA[<p>会話型インターフェースは以前から存在しており、顧客サービス、情報検索、タスク自動化など、さまざまなタスクを支援する手段としてますます人気が高まっています。通常、音声アシスタントやメッセージング アプリを通じてアクセスされるこれらのインターフェースは、ユーザーが質問をより効率的に解決できるように人間の会話をシミュレートします。</p><p>テクノロジーの進歩に伴い、チャットボットはユーザーにパーソナライズされたエクスペリエンスを提供しながら、より複雑なタスクを迅速に処理するために使用されるようになりました。自然言語処理 (NLP) により、チャットボットはユーザーの言語を処理し、メッセージの意図を識別し、そこから関連情報を抽出できるようになります。たとえば、固有表現抽出では、テキストを一連のカテゴリに分類して重要な情報を抽出します。感情分析は感情的なトーンを識別し、質問応答はクエリに対する「答え」を識別します。NLP の目標は、アルゴリズムが人間の言語を処理し、大量のテキストの中から関連する文章を見つけたり、テキストを要約したり、新しい独自のコンテンツを生成するなど、これまでは人間だけが実行できたタスクを実行できるようにすることです。</p><p>これらの高度な NLP 機能は、<a href="https://www.elastic.co/what-is/vector-search">ベクトル検索</a>と呼ばれるテクノロジーに基づいて構築されています。Elastic は、ベクトル検索、正確な<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search">k 最近傍 (kNN) 検索と近似 k 最近傍検索の</a>実行、および NLP をネイティブにサポートしており、Elasticsearch で直接カスタム モデルまたは<a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-model-ref.html#ml-nlp-model-ref">サードパーティ モデル</a>を使用できます。</p><p>このブログ記事では、ベクトル検索と NLP がどのように機能してチャットボットの機能を強化するかを説明し、Elasticsearch がどのようにそのプロセスを促進するかを説明します。まず、ベクトル検索の概要を簡単に説明します。</p><h2>ベクトル検索</h2><p>人間は書かれた言語の意味と文脈を理解できますが、機械は同じことはできません。ここでベクトルが登場します。テキストをベクトル表現（テキストの意味の数値表現）に変換することで、マシンはこの制限を克服できます。従来の検索と比較すると、キーワードや頻度に基づく語彙検索に依存するのではなく、ベクトルでは数値に対して定義された演算を使用してテキスト データの処理が可能になります。</p><p>これにより、ベクトル検索では、クエリベクトルが与えられた場合に類似性を表す「埋め込み空間」内の距離を使用して、類似の概念またはコンテキストを共有するデータを見つけることができます。データが類似している場合、対応するベクトルも同様になります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt53a615a8ba0ac931/6a17d795fbc5f8b257491910/08542abf8108aace288745b1aca8579b476ddc1b-1440x618.png" alt="" /><p>ベクトル検索は NLP アプリケーションで利用されるだけでなく、画像やビデオの処理など、非構造化データが関係するさまざまな他の領域でも使用されます。</p><p>チャットボット フローでは、ユーザーのクエリに対して複数のアプローチが可能であり、その結果、情報検索を改善してユーザー エクスペリエンスを向上させるさまざまな方法が存在します。それぞれの選択肢には独自の利点と欠点があるため、利用可能なデータとリソース、トレーニング時間（該当する場合）、および予想される精度を考慮することが重要です。次のセクションでは、質問応答 NLP モデルのこれらの側面について説明します。</p><h2>質問応答</h2><p>質問応答 (QA) モデルは、自然言語で尋ねられた質問に答えるように設計された NLP モデルの一種です。ユーザーが、ドキュメント内に既存のターゲット回答がなく、複数のリソースから回答を推測する必要がある質問がある場合、生成型 QA モデルが役立ちます。ただし、これらのモデルは計算コストが高く、ドメイン関連のトレーニングに大量のデータが必要になる可能性があり、この方法はドメイン外の質問を処理するのに特に価値があるにもかかわらず、状況によっては実用的ではない可能性があります。</p><p>一方、ユーザーが特定のトピックについて質問し、実際の回答がドキュメント内に存在する場合は、抽出型 QA モデルを使用できます。これらのモデルは、ソース ドキュメントから回答を直接抽出し、透明性と検証性に優れた結果を提供するため、質問にシンプルかつ効率的に回答したい企業や組織にとって、より実用的なオプションとなります。</p><p>以下の例は、 <a href="https://huggingface.co/deepset/minilm-uncased-squad2">Hugging Face で利用可能で</a>Elasticsearch にデプロイされた、事前トレーニング済みの抽出 QA モデルを使用して、特定のコンテキストから回答を抽出する方法を示しています。</p>POST _ml/trained_models/deepset__minilm-uncased-squad2/deployment/_infer
{
    "docs": [{"text_field": "Canvas is a data visualization and presentation application within Kibana. With Canvas, live data can be pulled directly from Elasticsearch and combined with colors, images, text, and other customized options to create dynamic, multi-page displays."}],
    "inference_config": {"question_answering": {"question": "What is Kibana Canvas?"}}
}


{
  "predicted_value": "a data visualization and presentation application",
  "start_offset": 10,
  "end_offset": 59,
  "prediction_probability": 0.28304219431376443
}
<p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-deploy-models.html">トレーニング済みのモデルをデプロイします。</a></p><p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-ner-example.html#ex-ner-ingest">推論取り込みパイプラインにモデルを追加します。</a></p><p>ユーザーのクエリを処理して情報を取得する方法はさまざまであり、非構造化データを扱う場合には、複数の言語モデルとデータ ソースを使用することが効果的な代替手段となります。これを説明するために、選択されたドキュメントから抽出されたデータを考慮してクエリに回答するために使用されるチャットボットのデータ処理の例を示します。</p><h2>チャットボットのデータ処理：NLPとベクトル検索</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta418a9c54bb16cf9/6a17d7975772624b371bca43/c2d1a2f110e937b1d3e5df0d5caac3c906c98fb0-1440x748.png" alt="" /><p>上記のように、チャットボットのデータ処理は 3 つの部分に分けられます。</p><ul><li><p><strong>ベクター処理:</strong>この部分はドキュメントをベクター表現に変換します。</p></li><li><p><strong>ユーザー入力処理:</strong>この部分では、ユーザークエリから関連情報を抽出し、セマンティック検索とハイブリッド検索を実行します。</p></li><li><p><strong>最適化:</strong>この部分には監視が含まれており、チャットボットの信頼性、最適なパフォーマンス、優れたユーザー エクスペリエンスを確保するために重要です。</p></li></ul><h2>ベクトル処理</h2><p><strong>処理</strong>部分では、まず各ドキュメントの構成要素を決定し、各要素をベクトル表現に変換します。これらの表現は、さまざまなデータ形式に対して作成できます。</p><p>埋め込みを計算するために使用できる方法は、事前トレーニング済みのモデルやライブラリなど、さまざまあります。</p><p>これらの表現に対する検索と取得の有効性は、既存のデータと、使用される方法の品質と関連性に依存することに注意することが重要です。</p><p>ベクトルが計算されると、 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html">dense_vector</a>フィールド タイプを使用して Elasticsearch に保存されます。</p>PUT &lt;target&gt;
{
  "mappings": {
    "properties": {
      "doc_part_vector": {
        "type": "dense_vector",
        "dims": 3
      },
      "doc_part" : {
        "type" : "keyword"
      }
    }
  }
}
<h2>チャットボットのユーザー入力処理</h2><p><strong>ユーザー</strong>側では、質問を受け取った後、先に進む前にそこから可能な限りすべての情報を抽出すると便利です。これはユーザーの意図を理解するのに役立ちます。この場合、それを支援するために<a href="https://huggingface.co/dslim/bert-base-NER">名前付きエンティティ認識モデル (NER)</a>を使用しています。NER は、名前付きエンティティを識別し、事前定義されたエンティティ カテゴリに分類するプロセスです。</p>POST _ml/trained_models/dslim__bert-base-ner/deployment/_infer
{
  "docs": { "text_field": "How many people work for Elastic?"}
}


{
  "predicted_value": "How many people work for [Elastic](ORG&amp;Elastic)?",
  "entities": [
    {
      "entity": "Elastic",
      "class_name": "ORG",
      "class_probability": 0.4993975435876747,
      "start_pos": 25,
      "end_pos": 32
    }
  ]
}
<p>必須のステップではありませんが、構造化データや上記または別の NLP モデルの結果を使用してユーザーのクエリを分類することで、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search-filter-example">フィルター</a>を使用して kNN 検索を制限できます。これにより、処理する必要があるデータの量が削減され、パフォーマンスと精度が向上します。</p>    "filter": {
      "term": {
        "org": "Elastic"
      }
    }
<h2>セマンティック検索とハイブリッド検索</h2><p>プロンプトはユーザーのクエリから生成され、チャットボットは変動性と曖昧性を持つ人間の言語を処理する必要があるため、<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#semantic-search">セマンティック検索は</a>最適です。Elasticsearch では、クエリ文字列と<a href="https://huggingface.co/sentence-transformers/msmarco-MiniLM-L-12-v3">埋め込みモデル</a>の ID を query_vector_builder オブジェクトに渡すことで、セマンティック検索を 1 ステップで実行できます。これにより、クエリがベクトル化され、kNN<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-search.html">検索</a>が実行され、クエリに意味的に最も近い上位 k 個の一致が取得されます。</p>POST /&lt;target&gt;/_search
{
  "knn": {
    "field": "doc_part_vector",
    "k": 5,
    "num_candidates": 20,
    "query_vector_builder": {
      "text_embedding": {
        "model_id": "&lt;text-embedding-model-id&gt;",
        "model_text": "&lt;query_string&gt;"
      }
    }
  }
 }
<p><a href="https://www.elastic.co/guide/en/machine-learning/8.7/ml-nlp-text-emb-vector-search-example.html">エンドツーエンドの例: テキスト埋め込みモデルを展開し、セマンティック検索に使用する方法。</a>Elasticsearch は、<strong>疎モデルである</strong>Okapi BM25 の Lucene 実装を使用してテキストクエリの関連性をランク付けし、<strong>密モデルはセマンティック検索</strong>に使用されます。ベクトル一致とテキスト <strong>クエリから取得された</strong><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#_combine_approximate_knn_with_other_features"><strong> 一致の両方の長所</strong></a><strong> を組み合わせる には 、</strong> ハイブリッド検索 を実行できます。</p>POST &lt;target&gt;/_search
{
  "query": {
          "match": {
            "content": {
              "query": "&lt;query_string&gt;"
            }
        }
  },
  "knn": {
    "field": "doc_part_vector",
    "query_vector_builder": {
      "text_embedding": {
    "model_id": "&lt;text-embedding-model-id&gt;",
     "model_text": "&lt;query_string&gt;"
      }
    },
    "filter": {
      "term": {
        "org": "Elastic"
      }
    }
  }
}
<h3>疎モデルと密モデルの両方を組み合わせると、多くの場合、最良の結果が得られます。</h3><p>通常、スパース モデルは短いクエリと特定の用語で優れたパフォーマンスを発揮しますが、密なモデルはコンテキストと関連付けを活用します。これらの方法がどのように比較され、相互に補完し合うかについて詳しく知りたい場合は、ここで、検索用に特別にトレーニングされた 2 つの高密度モデルに対して BM25 をベンチマークします。</p><p>最も関連性の高い結果は通常、ユーザーに最初に提供される回答になります。_scoreは、返されたドキュメントの<strong>関連性</strong>を判断するために使用される数値です。</p><h2>チャットボットの最適化</h2><p>チャットボットのユーザー エクスペリエンス、パフォーマンス、信頼性を向上させるには、ハイブリッド スコアリングを適用することに加えて、次のアプローチを組み込むことができます。<strong>感情分析:</strong>ダイアログが展開されるときにユーザーのコメントや反応を認識できるように、<a href="https://huggingface.co/distilbert-base-uncased-finetuned-sst-2-english">感情分析モデル</a>を組み込むことができます。</p>POST _ml/trained_models/distilbert-base-uncased-finetuned-sst-2-english/deployment/_infer
{
  "docs": { "text_field": "That was not my question!"}
}


{
  "predicted_value": "NEGATIVE",
  "prediction_probability": 0.980080439016437
}
<p><a href="https://www.elastic.co/blog/chatgpt-elasticsearch-openai-meets-private-data"><strong>GPT の機能</strong></a><strong>:</strong>全体的なエクスペリエンスを向上させる代わりに、Elasticsearch の検索関連性と OpenAI の GPT 質問応答機能を組み合わせ、 <a href="https://platform.openai.com/docs/guides/chat">Chat Completion API</a>を使用して、上位 k 件のドキュメントをコンテキストとして考慮し、ユーザーモデルが生成した応答を返すことができます。<em>プロンプト:&lt;user_question&gt; 「このドキュメント のみを使用して、この質問&lt;top_search_result&gt; に回答してください」</em></p><p><strong>可観測性:</strong>あらゆるチャットボットのパフォーマンスを確保することは非常に重要であり、これを実現するには監視が不可欠な要素です。チャットボットのやり取りをキャプチャするログに加えて、応答時間、待ち時間、その他の関連するチャットボットのメトリックを追跡することが重要です。これにより、パターンや傾向を特定し、異常を検出することさえ可能になります。Elastic <a href="https://www.elastic.co/blog/monitor-openai-api-gpt-models-opentelemetry-elastic">Observability</a>ツールを使用すれば、こうした情報を収集・分析できます。</p><h2>まとめ</h2><p>このブログ記事では、NLP とベクトル検索とは何かを説明し、ドキュメントのベクトル表現から抽出されたデータを考慮してユーザーのクエリに応答するために使用されるチャットボットの例を詳しく説明します。</p><p>実証されているように、NLP とベクトル検索を使用すると、チャットボットは構造化されたターゲット データを超えた複雑なタスクを実行できます。これには、複数のデータ ソースと形式をコンテキストとして使用して推奨事項を作成し、特定の製品またはビジネス関連のクエリに回答するとともに、パーソナライズされたユーザー エクスペリエンスを提供することも含まれます。</p><p>ユースケースは、顧客からの問い合わせに対応して顧客サービスを提供することから、開発者のクエリを支援し、ステップバイステップのガイダンスを提供したり、推奨事項を提案したり、タスクを自動化したりすることまで多岐にわたります。目標と既存のデータに応じて、他のモデルや方法を活用してさらに優れた結果を達成し、全体的なユーザー エクスペリエンスを向上させることもできます。</p><p>このトピックに関して役立つと思われるリンクをいくつか紹介します。</p><ol><li><p><a href="https://www.elastic.co/blog/how-to-deploy-natural-language-processing-nlp-getting-started">自然言語処理（NLP）の導入方法：はじめに</a></p></li><li><p><a href="https://www.elastic.co/blog/overview-image-similarity-search-in-elastic">Elasticsearchにおける画像類似検索の概要</a></p></li><li><p><a href="https://www.elastic.co/blog/chatgpt-elasticsearch-openai-meets-private-data">ChatGPTとElasticsearch: OpenAIとプライベートデータの出会い</a></p></li><li><p><a href="https://www.elastic.co/blog/monitor-openai-api-gpt-models-opentelemetry-elastic">Monitor OpenAI API and GPT models with OpenTelemetry and Elastic（OpenTelemetryとElasticを利用してOpenAI APIとGPTモデルを監視）</a></p></li><li><p><a href="https://www.elastic.co/blog/why-technology-leaders-need-vector-search">ITリーダーが検索エクスペリエンスを向上させるためにベクトル検索を必要とする5つの理由</a></p></li></ol><p>Elasticsearch に NLP とネイティブベクトル検索を組み込むことで、そのスピード、スケーラビリティ、検索機能を活用し、構造化データか非構造化データかを問わず大量のデータを処理できる、非常に効率的で効果的なチャットボットを作成できます。</p><p>始める準備はできましたか?<a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">Elastic Cloud の無料トライアル</a>を開始してください。</p><p><em>このブログ投稿では、それぞれの所有者が所有および運営するサードパーティ製の生成 AI ツールを使用したり、参照したりする場合があります。Elastic はサードパーティ製ツールを一切管理しておらず、そのコンテンツ、操作、使用、またそのようなツールの使用によって発生する損失や損害については一切責任を負いません。個人情報、機密情報、秘密情報を扱う AI ツールを使用する場合は注意してください。送信したデータは AI のトレーニングやその他の目的で使用される場合があります。お客様が提供する情報が安全に保管され、機密性が保たれるという保証はありません。生成 AI ツールを使用する前に、プライバシー慣行と利用規約をよく理解しておく必要があります。</em></p><p><em>Elastic、Elasticsearch および関連するマークは、米国およびその他の国における Elasticsearch NV の商標、ロゴ、または登録商標です。その他すべての会社名および製品名は、それぞれの所有者の商標、ロゴ、または登録商標です。</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/enhancing-chatbot-capabilities-with-nlp-and-vector-search-in-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/enhancing-chatbot-capabilities-with-nlp-and-vector-search-in-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c545fc80b6d79d6/6a170214839dfad776dcfd6f/d968e646240cd3ef7c79b5124d562a5f951d812b-1440x840.png" length="0" type="image/png"/>
    <pubDate>Wed, 21 Jun 2023 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>