<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/jp/search-labs/author/han-xiang-choong</link>
    </image>
    <link>https://www.elastic.co/jp/search-labs/author/han-xiang-choong</link>
    <atom:link href="https://www.elastic.co/jp/search-labs/rss/author/han-xiang-choong.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[jp]]></language>
    <lastBuildDate>Wed, 23 Sep 2026 07:09:18 GMT</lastBuildDate>
  <item>
    <title><![CDATA[高度なRAGテクニックパート2：クエリとテスト]]></title>
    <description><![CDATA[RAG のパフォーマンスを向上させる可能性のあるテクニックについて議論し、実装します。パート 2/2。高度な RAG パイプラインのクエリとテストに焦点を当てます。]]></description>
    <content:encoded><![CDATA[<p><em>すべてのコードは</em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em> 、Searchlabs リポジトリの advanced-rag-techniques</em></a><em> ブランチに あります 。</em></p><p>高度な RAG テクニックに関する記事のパート 2 へようこそ。<a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">このシリーズのパート 1</a>では、高度な RAG パイプラインのデータ処理コンポーネントを設定、説明、実装しました。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="高度なRAGパイプライン" /><p>この部分では、実装のクエリとテストを進めていきます。早速始めましょう！</p><h3>目次</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#searching-and-retrieving,-generating-answers">検索と取得、回答の生成</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#enriching-queries-with-synonyms">同義語によるクエリの強化</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-hypothetical-document-embedding">HyDE（仮想文書埋め込み）</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hybrid-search">ハイブリッド検索</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#experiments">実験</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#summary-of-results">結果の要約</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-1-who-audits-elastic">テスト 1: Elastic を監査するのは誰ですか?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-2--total-revenue-2023">テスト2：2023年の総収入</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-1">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-1">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-3-what-product-does-growth-primarily-depend-on-how-much">テスト 3: 成長は主にどの製品に依存しますか?いくら？</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-2">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-2">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-4-describe-employee-benefit-plan">テスト4: 従業員福利厚生制度の説明</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-3">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-3">シンプルラグ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-5-which-companies-did-elastic-acquire">テスト 5: Elastic が買収した企業はどれですか?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-4">アドバンスドRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-4">シンプルラグ</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#conclusion">まとめ</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#appendix">付記</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#prompts">プロンプト</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#rag-question-answering-prompt">RAG質問回答プロンプト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#elastic-query-generator-prompt">弾性クエリジェネレータプロンプト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#potential-questions-generator-prompt">潜在的な質問ジェネレータプロンプト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-generator-prompt">HyDEジェネレータプロンプト</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#sample-hybrid-search-query">ハイブリッド検索クエリのサンプル</a></p></li></ul></li></ul><h2>検索と取得、回答の生成</h2><p>最初のクエリ、理想的には主に年次報告書に記載されている情報を尋ねてみましょう。いかがでしょうか:</p>Who audits Elastic?"
<p>ここで、クエリを強化するためにいくつかのテクニックを適用してみましょう。</p><h3>同義語によるクエリの強化</h3><p>まず、クエリの文言の多様性を高めて、Elasticsearch クエリに簡単に処理できる形式に変えてみましょう。GPT-4o の助けを借りて、クエリを OR 句のリストに変換します。次のプロンプトを書いてみましょう:</p>
ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<p>GPT-4o をクエリに適用すると、基本クエリの同義語と関連語彙が生成されます。</p>'audits elastic OR 
elasticsearch audits OR 
elastic auditor OR 
elasticsearch auditor OR 
elastic audit firm OR 
elastic audit company OR 
elastic audit organization OR 
elastic audit service'
<p><code>ESQueryMaker</code>クラスでは、クエリを分割する関数を定義しました。</p>def parse_or_query(self, query_text: str) -&gt; List[str]:
    # Split the query by 'OR' and strip whitespace from each term
    # This converts a string like "term1 OR term2 OR term3" into a list ["term1", "term2", "term3"]
    return [term.strip() for term in query_text.split(' OR ')]
<p>その役割は、この OR 句の文字列を取得して用語のリストに分割し、主要なドキュメント フィールドで複数の一致を実行できるようにすることです。</p>["original_text", 'keyphrases', 'potential_questions', 'entities']
<p>最終的にこのクエリに至ります:</p> 'query': {
    'bool': {
        'must': [
            {
                'multi_match': {
                'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
                'fields': [
                    'original_text',
                'keyphrases',
                'potential_questions',
                'entities'
                ],
                'type': 'best_fields',
                'operator': 'or'
                }
            }
      ]
<p>これにより、元のクエリよりも多くのベースがカバーされ、同義語を忘れたために検索結果を見逃すリスクが軽減されることが期待されます。しかし、私たちにはもっとできることがある。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h3>HyDE（仮想文書埋め込み）</h3><p>今回は<a href="https://arxiv.org/abs/2212.10496">HyDE を</a>実装するために、再び GPT-4o を活用しましょう。</p><p>HyDE の基本的な前提は、仮想ドキュメント (元のクエリに対する回答が含まれる可能性のある種類のドキュメント) を生成することです。文書の事実性や正確性は問題ではありません。それを念頭に置いて、次のプロンプトを書いてみましょう。</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<p>ベクトル検索は通常、コサインベクトルの類似度に基づいて行われるため、クエリをドキュメントに一致させるのではなく、ドキュメントをドキュメントに一致させることでより良い結果を達成できるというのが HyDE の前提です。</p><p>私たちが重視するのは、構造、フロー、用語です。あまり事実ではない。GPT-4o は次のような HyDE ドキュメントを出力します。</p>'Elastic N.V., the parent company of Elastic, the organization known for developing Elasticsearch, is subject to audits to ensure financial accuracy, 
regulatory compliance, and the integrity of its financial statements. The auditing of Elastic N.V. is typically conducted by an external, 
independent auditing firm. This is common practice for publicly traded companies to provide stakeholders with assurance regarding the company\'s 
financial position and operations.\n\nThe primary external auditor for Elastic is the audit firm Ernst &amp; Young LLP (EY). Ernst &amp; Young is one of the 
four largest professional services networks in the world, commonly referred to as the "Big Four" audit firms. These firms handle a substantial number 
of audits for major corporations around the globe, ensuring adherence to generally accepted accounting principles (GAAP) and international financial 
reporting standards (IFRS).\n\nThe audit process conducted by EY involves several steps. Initially, the auditors perform a risk assessment to identify 
areas where misstatements due to error or fraud could occur. They then design audit procedures to test the accuracy and completeness of financial statements,
 which include examining financial transactions, assessing internal controls, and reviewing compliance with relevant laws and regulations. Upon completion of 
 the audit, Ernst &amp; Young issues an audit report, which includes the auditor’s opinion on whether the financial statements are free from material misstatement 
 and are presented fairly in accordance with the applicable financial reporting framework.\n\nIn addition to external audits by firms like Ernst &amp; Young, 
 Elastic may also be subject to internal audits. Internal audits are performed by the company’s own internal auditors to evaluate the effectiveness of internal 
 controls, risk management, and governance processes.\n\nOverall, the auditing process plays a crucial role in maintaining the transparency and reliability of 
 Elastic\'s financial information, providing confidence to investors, regulators, and other stakeholders.'
<p>これはかなり信憑性があり、インデックスを作成したい種類のドキュメントに最適な候補のように見えます。これを埋め込み、ハイブリッド検索に使用します。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h3>ハイブリッド検索</h3><p>これが私たちの検索ロジックの中核です。語彙検索コンポーネントは、生成された OR 句の文字列になります。高密度ベクトル コンポーネントには、HyDE ドキュメント (検索ベクトルとも呼ばれます) が埋め込まれます。KNN を使用して、検索ベクトルに最も近い候補ドキュメントをいくつか効率的に識別します。デフォルトでは、語彙検索コンポーネントを<em>TF-IDF と BM25 によるスコアリング</em>と呼びます。最後に、語彙スコアと密なベクトルスコアは、 <a href="https://arxiv.org/abs/2407.01219">Wang ら</a>が推奨する 30/70 比率を使用して結合されます。</p>def hybrid_vector_search(self, index_name: str, query_text: str, query_vector: List[float], 
                         text_fields: List[str], vector_field: str, 
                         num_candidates: int = 100, num_results: int = 10) -&gt; Dict:
    """
    Perform a hybrid search combining text-based and vector-based similarity.

    Args:
        index_name (str): The name of the Elasticsearch index to search.
        query_text (str): The text query string, which may contain 'OR' separated terms.
        query_vector (List[float]): The query vector for semantic similarity search.
        text_fields (List[str]): List of text fields to search in the index.
        vector_field (str): The name of the field containing document vectors.
        num_candidates (int): Number of candidates to consider in the initial KNN search.
        num_results (int): Number of final results to return.

    Returns:
        Dict: A tuple containing the Elasticsearch response and the search body used.
    """
    try:
        # Parse the query_text into a list of individual search terms
        # This splits terms separated by 'OR' and removes any leading/trailing whitespace
        query_terms = self.parse_or_query(query_text)

        # Construct the search body for Elasticsearch
        search_body = {
            # KNN search component for vector similarity
            "knn": {
                "field": vector_field,  # The field containing document vectors
                "query_vector": query_vector,  # The query vector to compare against
                "k": num_candidates,  # Number of nearest neighbors to retrieve
                "num_candidates": num_candidates  # Number of candidates to consider in the KNN search
            },
            "query": {
                "bool": {
                    # The 'must' clause ensures that matching documents must satisfy this condition
                    # Documents that don't match this clause are excluded from the results
                    "must": [
                        {
                            # Multi-match query to search across multiple text fields
                            "multi_match": {
                                "query": " ".join(query_terms),  # Join all query terms into a single space-separated string
                                "fields": text_fields,  # List of fields to search in
                                "type": "best_fields",  # Use the best matching field for scoring
                                "operator": "or"  # Match any of the terms (equivalent to the original OR query)
                            }
                        }
                    ],
                    # The 'should' clause boosts relevance but doesn't exclude documents
                    # It's used here to combine vector similarity with text relevance
                    "should": [
                        {
                            # Custom scoring using a script to combine vector and text scores
                            "script_score": {
                                "query": {"match_all": {}},  # Apply this scoring to all documents that matched the 'must' clause
                                "script": {
                                    # Script to combine vector similarity and text relevance
                                    "source": """
                                    # Calculate vector similarity (cosine similarity + 1)
                                    # Adding 1 ensures the score is always positive
                                    double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;
                                    # Get the text-based relevance score from the multi_match query
                                    double text_score = _score;
                                    # Combine scores: 70% vector similarity, 30% text relevance
                                    # This weighting can be adjusted based on the importance of semantic vs keyword matching
                                    return 0.7 * vector_score + 0.3 * text_score;
                                    """,
                                    # Parameters passed to the script
                                    "params": {
                                        "query_vector": query_vector,  # Query vector for similarity calculation
                                        "vector_field": vector_field  # Field containing document vectors
                                    }
                                }
                            }
                        }
                    ]
                }
            }
        }

        # Execute the search request against the Elasticsearch index
        response = self.conn.search(index=index_name, body=search_body, size=num_results)
        # Log the successful execution of the search for monitoring and debugging
        logger.info(f"Hybrid search executed on index: {index_name} with text query: {query_text}")
        # Return both the response and the search body (useful for debugging and result analysis)
        return response, search_body
    except Exception as e:
        # Log any errors that occur during the search process
        logger.error(f"Error executing hybrid search on index: {index_name}. Error: {e}")
        # Re-raise the exception for further handling in the calling code
        raise e
<p>最後に、RAG 関数を組み立てることができます。クエリから回答までの RAG の流れは次のようになります。</p><ol><li><p>クエリを OR 句に変換します。</p></li><li><p>HyDE ドキュメントを生成して埋め込みます。</p></li><li><p>両方をハイブリッド検索への入力として渡します。</p></li><li><p>上位 n 件の結果を取得し、最も関連性の高いスコアが LLM のコンテキスト メモリ内で「最新」になるように結果を逆にします (逆パッキング)。逆パッキングの例: クエリ:「Elasticsearch クエリ最適化手法」取得されたドキュメント (関連性の高い順): LLM コンテキストの順序を逆にする: 順序を逆にすることで、最も関連性の高い情報 (1) がコンテキストの最後に表示され、回答生成中に LLM からより多くの注目を受ける可能性が高くなります。</p><ol><li><p>「ブールクエリを使用して、複数の検索条件を効率的に組み合わせます。」</p></li><li><p>「クエリの応答時間を改善するためのキャッシュ戦略を実装します。」</p></li><li><p>「インデックス マッピングを最適化して、検索パフォーマンスを高速化します。」</p></li><li><p>「インデックス マッピングを最適化して、検索パフォーマンスを高速化します。」</p></li><li><p>「クエリの応答時間を改善するためのキャッシュ戦略を実装します。」</p></li><li><p>「ブールクエリを使用して、複数の検索条件を効率的に組み合わせます。」</p></li></ol></li><li><p>生成のためにコンテキストを LLM に渡します。</p></li></ol>def get_context(index_name, 
                match_query, 
                text_query, 
                fields, 
                num_candidates=100, 
                num_results=20, 
                text_fields=["original_text", 'keyphrases', 'potential_questions', 'entities'], 
                embedding_field="primary_embedding"):

    embedding=embedder.get_embeddings_from_text(text_query)

    results, search_body = es_query_maker.hybrid_vector_search(
        index_name=index_name,
        query_text=match_query,
        query_vector=embedding[0][0],
        text_fields=text_fields,
        vector_field=embedding_field,
        num_candidates=num_candidates,
        num_results=num_results
    )

    # Concatenates the text in each 'field' key of the search result objects into a single block of text.
    context_docs=['\n\n'.join([field+":\n\n"+j['_source'][field] for field in fields]) for j in results['hits']['hits']]

    # Reverse Packing to ensure that the highest ranking document is seen first by the LLM.
    context_docs.reverse()
    return context_docs, search_body

def retrieval_augmented_generation(query_text):
    match_query= gpt4o.generate_query(query_text)
    fields=['original_text']

    hyde_document=gpt4o.generate_HyDE(query_text)

    context, search_body=get_context(index_name, match_query, hyde_document, fields)

    answer= gpt4o.basic_qa(query=query_text, context=context)
    return answer, match_query, hyde_document, context, search_body

<p>クエリを実行して回答を取得してみましょう。</p>According to the context, Elastic N.V. is audited by an independent registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of independent registered public accounting firm," which states:

"We have audited the accompanying consolidated balance sheets of Elastic N.V. [...] / s / pricewaterhouseco."
<p>ニース。そうです。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h2>実験</h2><p>今答えなければならない重要な質問があります。これらの実装に多大な労力と追加の複雑さを投資することで、何が得られましたか?</p><p>少し比較してみましょう。私たちが実装した RAG パイプラインと、私たちが行った機能強化のないベースライン ハイブリッド検索を比較したものです。小規模な一連のテストを実行して、大きな違いが見られるか確認します。ここで実装した RAG を AdvancedRAG と呼び、基本パイプラインを SimpleRAG と呼びます。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" alt="シンプルなRAGパイプライン" /><h4>結果の要約</h4><p>この表は、両方の RAG パイプラインの 5 つのテストの結果をまとめたものです。回答の詳細と品質に基づいて各方法の相対的な優位性を判断しましたが、これは完全に主観的な判断です。実際の回答はこの表の下に再現されていますので、ご参照ください。それでは、彼らの成果を見てみましょう!</p><p>SimpleRAG は質問 1 と 5 に答えることができませんでした。AdvancedRAG は質問 2、3、4 についても非常に詳しく説明しました。詳細度が増したことにより、AdvancedRAG の回答の質が優れていると判断しました。</p><p>テスト</p><p>質問</p><p>高度なRAGパフォーマンス</p><p>SimpleRAG パフォーマンス</p><p>AdvancedRAG レイテンシー</p><p>SimpleRAG レイテンシ</p><p>勝者</p><p>1</p><p>Elastic を監査するのは誰ですか?</p><p>監査人として PwC を正しく特定しました。</p><p>監査人を識別できませんでした。</p><p>11.6秒</p><p>4.4秒</p><p>アドバンスドRAG</p><p>2</p><p>2023年の総収益はいくらでしたか?</p><p>正しい収益数値を提供しました。前年度の収益に関する追加のコンテキストを含めました。</p><p>正しい収益数値を提供しました。</p><p>13.3秒</p><p>2.8秒</p><p>アドバンスドRAG</p><p>3</p><p>成長は主にどの製品に依存しますか?いくら？</p><p>Elastic Cloud が主要な推進力であることを正しく認識しました。全体的な収益コンテキストとより詳しい詳細が含まれています。</p><p>Elastic Cloud が主要な推進力であることを正しく認識しました。</p><p>14.1秒</p><p>12.8秒</p><p>アドバンスドRAG</p><p>4</p><p>従業員福利厚生プランの説明</p><p>退職金制度、健康プログラム、その他の福利厚生について包括的に説明しました。異なる年ごとの具体的な寄付金額が含まれています。</p><p>報酬、退職金制度、職場環境、Elastic Cares プログラムなどの福利厚生の概要をわかりやすく説明しました。</p><p>26.6秒</p><p>11.6秒</p><p>アドバンスドRAG</p><p>5</p><p>Elastic が買収した企業はどれですか?</p><p>レポートに記載されている最近の買収 (CmdWatch、Build Security、Optimyze) を正しくリストしました。いくつかの取得日と購入価格を提供しました。</p><p>提供されたコンテキストから関連情報を取得できませんでした。</p><p>11.9秒</p><p>2.7秒</p><p>アドバンスドRAG</p><h4>テスト 1: Elastic を監査するのは誰ですか?</h4><p>アドバンスドRAG</p>According to the context, Elastic N.V. is audited by an independent 
registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of 
independent registered public accounting firm," which states:
"We have audited the accompanying consolidated balance sheets 
of Elastic N.V. [...] / s / pricewaterhouseco."
<p>シンプルラグ</p>According to the provided context, no specific auditing firm 
or company is mentioned as the auditor for Elastic. 
The context provided talks about various features, strengths, 
deployment options, and other details related to Elastic and 
its operations but does not specify which auditing firm is 
responsible for auditing Elastic's financial statements or 
internal controls.
<p><strong>要約</strong>: SimpleRAGはPWCを監査人として特定しなかった</p><p>そうですね、それは実はかなり驚きました。これは SimpleRAG 側の検索失敗のようです。監査に関連する文書は取得されませんでした。次のテストでは難易度を少し下げてみましょう。</p><h4>テスト2：2023年の総収入</h4><p>アドバンスドRAG</p>### Total Revenue for 2023

According to the provided context, the total revenue for the 
year ended April 30, 2023, was **$1,068,989,000**. This 
information can be found several times in the provided text. 
Specifically, one part states:

"The following table sets forth our results of operations 
for the periods presented in dollars and as a percentage of 
our total revenue. ... year ended April 30, 2023, ... 
total revenue 1,068,989".

Additionally, it is mentioned here as well:

"For the years ended April 30, 2023, 2022, and 2021, 
representing year-over-year growth of 24% for the year 
ended April 30, 2023... total revenue $1,068,989".

Therefore, the total revenue for 2023 is $1.069 billion.
<p>シンプルラグ</p>According to the context provided, the total revenue for the fiscal year ended April 30, 2023, is:

$1,068,989,000 (or approximately $1.1 billion).
<p><strong>要約</strong>: 両RAGとも正解: 2023年の総収益は1,068,989,000ドル</p><p>二人ともここにいました。AdvancedRAG がより広範囲の文書を入手したように思われますか?確かに、答えはより詳細で、前年からの情報が組み込まれています。私たちが行った機能強化を考えると、それは予想されることですが、判断するには時期尚早です。</p><p>難易度を上げてみましょう。</p><h4>テスト 3: 成長は主にどの製品に依存しますか?いくら？</h4><p>アドバンスドRAG</p>### Product Growth Primarily Depends On:
According to the context, "our revenue is derived primarily from 
subscriptions of search, observability and security built into 
the elastic stack." The document mentions various strategies and 
factors contributing to the company’s growth which include investment 
in "elastic cloud," "elastic stack," and the conversion of free 
users to paid subscribers.

### Quantitative Data:
- **Revenue**: The document states, "Our total revenue was $1.1 billion, 
$862.4 million, and $608.5 million for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Subscription Revenue**: "Subscriptions accounted for 92%, 93%, 
and 93% of our total revenue for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Elastic Cloud Contribution**: "For the years ended April 30, 2023, 
2022, and 2021, Elastic Cloud contributed 40%, 35%, and 27% of our 
total revenue, respectively."
- **Customer Growth**: The number of customers increased from 
"over 15,000 as of April 30, 2021," to "approximately 20,200 as of 
April 30, 2023."

### Summary:
Product growth for the company primarily depends on the "Elastic Stack" 
(which includes search, observability, and security solutions) and 
particularly the "Elastic Cloud." The Elastic Cloud's growing contribution 
to revenue highlights its significance, which grew from 27% in 2021 to 40% in 2023.
<p>シンプルラグ</p>According to the provided context, growth primarily depends on **Elastic Cloud**, 
which is described as the company's "family of cloud-based offerings." 
The context also emphasizes the significance of Elastic Cloud in achieving 
the company's long-term growth potential and its role in increasing subscription revenue.

From the context:

- "Increasing adoption of Elastic Cloud... is an important growth opportunity 
for our business. Organizations are increasingly looking for hosted 
deployment alternatives with reduced administrative burdens."
- "We believe that offering Elastic Cloud is important for achieving 
our long-term growth potential, and we expect Elastic Cloud’s contribution 
to our subscription revenue to continue to increase."

In terms of the contribution to total revenue, here are the specific figures provided:

- Elastic Cloud contributed 40% of the total revenue for the year ended April 30, 2023.
- Elastic Cloud contributed 35% of the total revenue for the year ended April 30, 2022.
- Elastic Cloud contributed 27% of the total revenue for the year ended April 30, 2021.

This progression shows a clear upward trend in the importance and 
contribution of Elastic Cloud to the company's overall revenue.
<p><strong>概要</strong>: 両方の RAG は、Elastic Cloud を主要な成長原動力として正しく認識しました。ただし、AdvancedRAG では、サブスクリプション収益と顧客の増加を考慮したより詳細な情報が含まれており、Elastic の他の製品についても明示的に言及されています。</p><h4>テスト4: 従業員福利厚生制度の説明</h4><p>アドバンスドRAG</p>### Employee Benefit Plan Overview

The provided context outlines various elements of the employee benefit plan at 
Elastic, often referred to as "Elasticians." Here are the key components 
described in the texts:

1. **401(k) Plan**:
   - **Coverage**: The 401(k) plan is available to substantially all U.S. 
   employees who meet minimum age and service requirements.
   - **Contributions**: Elastic makes contributions to the 401(k) plan up to 
   6% of the participating employee’s W-2 earnings and wages.
   - **Expenses**: For the fiscal years ended April 30, Elastic recorded 
   expenses of $17.9 million (2023), $15.2 million (2022), and $11.4 million (2021) 
   related to the 401(k) plan.
   - **Defined-Contribution Plans in Other Countries**: Elastic has 
   defined-contribution plans in various other countries and recorded respective 
   expenses of $9.4 million (2023), $7.2 million (2022), and $5.1 million (2021).

2. **Stock-Based Compensation**:
   - **Types of Awards**: Stock options, restricted stock units (RSUs), 
   and shares under the Employee Stock Purchase Plan (ESPP).
   - **Fair Value Measurement**: Fair value of these stock awards is 
   measured using models like Black-Scholes.
   - **Employee Stock Purchase Plan (2022 ESPP)**: 
     - Started in 2022, it allows employees to acquire ordinary 
     shares at a discount (85% of the market value at the beginning 
     or end of the offering period).
     - Offering periods are approximately six months long.

3. **Total Rewards Compensation**:
   - **Components**: Includes cash compensation as well as equity awards, 
   reflecting a comprehensive interest in the well-being of the employees.
   - **Programs**: Market competitive medical and dental programs, 
   mental health, and holistic well-being programs.
   - **Paid Time Off**: Features 16 weeks of paid leave for new 
   parents and life-planning benefits.

4. **Retirement and Income Protection Plans**:
   - **401(k) Match**: Elastic provides a dollar-for-dollar 
   match up to 6% of eligible earnings up to a plan-limit 
   maximum for U.S.-based employees.
   - **Similar Plans Outside the U.S.**: Competitive retirement 
   plans available for employees based outside the United States.

5. **Flexible Work Environment**: 
   - **Distributed Workforce**: Elastic operates with a global 
   distributed workforce model, allowing employees to work remotely 
   and providing a broad candidate pool.
   
6. **Learning and Development**:
   - **Programs**: Access to on-demand video-based learning, 
   leadership development programs, and other specific training 
   to develop managers and leaders.

7. **Employee Engagement and Surveys**:
   - **Mechanisms**: Annual employee engagement survey and 
   mid-year pulse survey.
   - **Objectives**: To monitor morale, develop action plans, 
   and maintain high engagement scores.

8. **Onboarding Program**:
   - **X-School**: An extensive new-hire orientation program to 
   help new employees meet peers and learn about the company's 
   products and solutions.

In summary, Elastic's employee benefit plan encompasses a broad 
spectrum of financial, health, well-being, and professional development 
benefits designed to support and enhance the lives of its employees. 
The plan is both competitive and comprehensive, aiming to attract,
retain, and engage its global workforce.
<p>シンプルラグ</p>The employee benefit plan at the Company includes several key 
components aimed at enhancing the overall well-being and 
professional development of employees, referred to as "Elasticians." 

1. **Compensation and Benefits:**
   - The Company provides market-competitive compensation 
   packages which typically include cash compensation and equity awards.
   - Comprehensive medical and dental programs are offered, 
   emphasizing mental health and holistic well-being.
   - Paid time off programs feature:
     - 16 weeks of paid leave for new parents.
     - Life-planning benefits.
     - Travel reimbursements for certain healthcare services.

2. **Retirement and Income Protection Plans:**
   - In the United States, employees are eligible for a 401k 
   plan with a dollar-for-dollar match by the Company up to 6% 
   of eligible earnings, subject to a plan limit.
   - Similar competitive retirement and income protection plans 
   are available for employees outside the United States.

3. **Work Environment:**
   - The Company offers a flexible work environment, allowing most 
   employees to work from anywhere as often as they would like.
   - Physical offices around the world are available for employees 
   who prefer to work on-site or need a space for face-to-face interactions.

4. **Community Involvement:**
   - The Elastic Cares Program supports charitable organizations 
   that matter to employees on both local and global levels.
   - The program includes:
     - Donation matching.
     - Providing the Company's technology for free to certain 
     nonprofit organizations.
     - Volunteer time off initiative.
<p><strong>概要</strong>: AdvancedRAG では、米国に拠点を置く従業員向けの 401K プランや、米国外の定義拠出金プランについてさらに詳しく取り上げています。また、健康と幸福の計画についても言及していますが、SimpleRAG が言及している Elastic Cares プログラムについては触れられていません。</p><h4>テスト 5: Elastic が買収した企業はどれですか?</h4><p>アドバンスドRAG</p>Elastic acquired several companies over the years to enhance its 
technology and market presence. According to the provided context, 
the specific acquisitions include:

1. **CmdWatch Security Inc.**: Acquired on September 17, 2021. 
The total purchase consideration was $77.8 million.
2. **Build Security Ltd.**: Acquired on September 2, 2021. The 
combined purchase price for Build Security Ltd. and another acquisition was $57.2 million.
3. **Optimyze Cloud Inc.**: Acquired on November 1, 2021. Part 
of the combined purchase price mentioned above with Build Security Ltd.
4. **Endgame, Inc.**: Although the context mentions an amended 
and restated stock incentive plan, it indicates that Endgame, Inc. was 
associated with Elastic, suggesting it could have been acquired at some point.

These acquisitions are part of Elastic's strategy to 
enhance the technology underlying its security and 
observability offerings.
<p>シンプルラグ</p>The provided context does not specify any companies that Elastic has acquired. 
Therefore, based on the context, there is no information available about the companies acquired by Elastic.
<p><strong>概要</strong>: SimpleRAG は買収に関する関連情報を取得せず、回答に失敗しました。AdvancedRAG は、レポートに記載されている主要な買収である CmdWatch、Build Security、Optimyze を正しくリストしています。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h2>まとめ</h2><p>私たちのテストによると、私たちの高度な技術により、提示される情報の範囲と深さが拡大し、RAG 回答の品質が向上する可能性があるようです。</p><p>さらに、 <code>Which companies did Elastic acquire?</code>や<code>Who audits Elastic</code>などのあいまいな表現の質問に対して、AdvancedRAG では正しく回答されましたが、SimpleRAG では正しく回答されなかったため、信頼性が向上する可能性があります。</p><p>ただし、5 件中 3 件では、ハイブリッド検索のみを組み込んだ基本的な RAG パイプラインで、重要な情報のほとんどを捉えた回答を生成できたという点に留意する価値があります。</p><p>データ準備フェーズとクエリフェーズに LLM が組み込まれているため、AdvancedRAG のレイテンシは通常、SimpleRAG の 2 ～ 5 倍になることに注意してください。これは大きなコストであるため、AdvancedRAG は、応答品質がレイテンシーよりも優先される状況にのみ適している可能性があります。</p><p>データ準備段階で Claude Haiku や GPT-4o-mini などの小型で安価な LLM を使用すると、大きなレイテンシ コストを軽減できます。回答生成用の高度なモデルを保存します。</p><p>これは Wang らの研究結果と一致しています。結果が示すように、行われた改善は比較的漸進的です。つまり、シンプルなベースライン RAG を使用すると、安価で高速でありながら、適切な最終製品にほぼ到達できます。私にとっては、それは興味深い結論です。速度と効率が重要となるユースケースでは、SimpleRAG が賢明な選択です。パフォーマンスを最大限に引き出す必要があるユースケースでは、AdvancedRAG に組み込まれたテクニックが解決策となる可能性があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt56b7067a9d41d5a8/6a171119acf0886fb4be9c45/ea811706b6adc4731d90b925a9fefa0ac15901b4-1440x1060.jpg" alt="王パイプライン" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">トップに戻る</a></p><h2>付記</h2><h3>プロンプト</h3><h4>RAG質問回答プロンプト</h4><p>クエリとコンテキストに基づいて LLM に回答を生成させるためのプロンプト。</p>BASIC_RAG_PROMPT = '''
You are an AI assistant tasked with answering questions based primarily on the provided context, while also drawing on your own knowledge when appropriate. Your role is to accurately and comprehensively respond to queries, prioritizing the information given in the context but supplementing it with your own understanding when beneficial. Follow these guidelines:

1. Carefully read and analyze the entire context provided.
2. Primarily focus on the information present in the context to formulate your answer.
3. If the context doesn't contain sufficient information to fully answer the query, state this clearly and then supplement with your own knowledge if possible.
4. Use your own knowledge to provide additional context, explanations, or examples that enhance the answer.
5. Clearly distinguish between information from the provided context and your own knowledge. Use phrases like "According to the context..." or "The provided information states..." for context-based information, and "Based on my knowledge..." or "Drawing from my understanding..." for your own knowledge.
6. Provide comprehensive answers that address the query specifically, balancing conciseness with thoroughness.
7. When using information from the context, cite or quote relevant parts using quotation marks.
8. Maintain objectivity and clearly identify any opinions or interpretations as such.
9. If the context contains conflicting information, acknowledge this and use your knowledge to provide clarity if possible.
10. Make reasonable inferences based on the context and your knowledge, but clearly identify these as inferences.
11. If asked about the source of information, distinguish between the provided context and your own knowledge base.
12. If the query is ambiguous, ask for clarification before attempting to answer.
13. Use your judgment to determine when additional information from your knowledge base would be helpful or necessary to provide a complete and accurate answer.

Remember, your goal is to provide accurate, context-based responses, supplemented by your own knowledge when it adds value to the answer. Always prioritize the provided context, but don't hesitate to enhance it with your broader understanding when appropriate. Clearly differentiate between the two sources of information in your response.

Context:
[The concatenated documents will be inserted here]

Query:
[The user's question will be inserted here]

Please provide your answer based on the above guidelines, the given context, and your own knowledge where appropriate, clearly distinguishing between the two:
'''
<h4>弾性クエリジェネレータプロンプト</h4><p>同義語を使用してクエリを拡充し、OR 形式に変換するように要求します。</p>ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<h4>潜在的な質問ジェネレータプロンプト</h4><p>潜在的な質問の生成を促し、ドキュメントのメタデータを充実させます。</p>RAG_QUESTION_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating questions for Retrieval-Augmented Generation (RAG) systems. Your task is to analyze a given document and create 10 diverse questions that would effectively test a RAG system's ability to retrieve and synthesize information from this document.

Guidelines:
1. Thoroughly analyze the entire document.
2. Generate exactly 10 questions that cover various aspects and levels of complexity within the document's content.
3. Create questions that specifically target:
   a. Key facts and information
   b. Main concepts and ideas
   c. Relationships between different parts of the content
   d. Potential applications or implications of the information
   e. Comparisons or contrasts within the document
4. Ensure questions require answers of varying lengths and complexity, from simple retrieval to more complex synthesis.
5. Include questions that might require combining information from different parts of the document.
6. Frame questions to test both literal comprehension and inferential understanding.
7. Avoid yes/no questions; focus on open-ended questions that promote comprehensive answers.
8. Consider including questions that might require additional context or knowledge to fully answer, to test the RAG system's ability to combine retrieved information with broader knowledge.
9. Number the questions from 1 to 10.
10. Output only the ten questions, without any additional text, explanations, or answers.

Document:
[The document content will be inserted here]

Generate 10 questions optimized for testing a RAG system based on this document:
'''
<h4>HyDEジェネレータプロンプト</h4><p>HyDEを使用して仮想文書を生成するためのプロンプト</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<h3>ハイブリッド検索クエリのサンプル</h3>{'knn': {'field': 'primary_embedding',
  'query_vector': [0.4265527129173279,
   -0.1712949573993683,
   -0.042020395398139954,
   ...],
  'k': 100,
  'num_candidates': 100},
 'query': {'bool': {'must': [{'multi_match': {'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
      'fields': ['original_text',
       'keyphrases',
       'potential_questions',
       'entities'],
      'type': 'best_fields',
      'operator': 'or'}}],
   'should': [{'script_score': {'query': {'match_all': {}},
      'script': {'source': '\n                                        double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;\n                                        double text_score = _score;\n                                        return 0.7 * vector_score + 0.3 * text_score;\n                                        ',
       'params': {'query_vector': [0.4265527129173279,
         -0.1712949573993683,
         -0.042020395398139954,
        ...],
        'vector_field': 'primary_embedding'}}}}]}},
 'size': 10}
]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[高度なRAGテクニックパート1：データ処理]]></title>
    <description><![CDATA[RAG のパフォーマンスを向上させる可能性のあるテクニックについて議論し、実装します。パート 1/2。高度な RAG パイプラインのデータ処理と取り込みのコンポーネントに焦点を当てます。]]></description>
    <content:encoded><![CDATA[<p><em>これは、高度なRAGテクニックを探るパート1です。</em><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2"><em>パート2はこちらをクリックしてください！</em></a></p><p>最近の論文<a href="https://arxiv.org/abs/2407.01219">「検索拡張生成におけるベスト プラクティスの探求」では、</a> RAG のベスト プラクティスのセットに収束することを目的として、さまざまな RAG 強化手法の有効性を経験的に評価しています。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt671704ff06a4011d/6a170b3ea929cf2d19ae09d8/dafa7250e7c4ead4d9b4aed7c407509131929749-1440x572.png" alt="王氏が推奨するRAGパイプライン" /><p>提案されたベストプラクティスのいくつか、つまり検索の品質を向上させることを目的としたベストプラクティス<strong>（センテンスチャンキング、HyDE、リバースパッキング）</strong>を実装します。</p><p>簡潔にするために、効率性の向上に重点を置いた手法<strong>(クエリの分類と要約)</strong>は省略します。</p><p>また、ここでは取り上げなかったものの、個人的には便利で興味深いと思われるいくつかのテクニック<strong>(メタデータの包含、複合マルチフィールドの埋め込み、クエリの強化) も</strong>実装します。</p><p>最後に、検索結果と生成された回答の品質がベースラインと比較して向上したかどうかを確認するための短いテストを実行します。さあ始めましょう！</p><h2>RAGの概要</h2><p>RAG は、外部の知識ベースから情報を取得して生成された回答を充実させることで、LLM を強化することを目的としています。ドメイン固有の情報を提供することで、LLM はトレーニング データの範囲外のユース ケースに迅速に適応できます。微調整よりも大幅にコストが安く、最新の状態に保つのも簡単になります。</p><p>RAG の品質を向上させるための対策は、通常、次の 2 つの点に重点を置いています。</p><ol><li><p>ナレッジベースの品質と明確さを向上します。</p></li><li><p>検索クエリの範囲と特定性を向上させます。</p></li></ol><p>これら 2 つの対策により、LLM が関連する事実や情報にアクセスできる可能性が高まり、幻覚を起こしたり、古くなったり無関係になったりする可能性のある独自の知識を利用したりする可能性が低くなるという目標が達成されます。</p><p>方法の多様性を数文で説明するのは困難です。わかりやすくするために、すぐに実装に移りましょう。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="高度なRAGパイプライン" /><h3>目次</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#overview">ご紹介</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">目次</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#set-up">設定</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#ingesting-processing-and-embedding-documents">ドキュメントの取り込み、処理、埋め込み</a>  </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#data-ingestion">データインジェスト</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#sentence-level-token-wise-chunking">文レベル、トークン単位のチャンキング</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#metadata-inclusion-and-generation">メタデータの包含と生成</a> </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#keyphrases-extracted-by-textrank">TextRankによって抽出されたキーフレーズ</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#potential-questions-generated-by-gpt-4o">GPT-4oによって生成される潜在的な質問</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#entities-extracted-by-spacy">Spacyによって抽出されたエンティティ</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#composite-multi-field-embeddings">複合多体埋め込み</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#indexing-to-elastic">Elasticへのインデックス</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#cat-break">猫の休憩</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#appendix">付記</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#definitions">定義</a></p></li></ul></li></ul><h2>設定</h2><p><em>すべてのコードは</em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em> Searchlabs リポジトリに</em></a><em> あります 。</em></p><p>まずは第一に。次のものが必要になります:</p><ol><li><p>弾力性のあるクラウドの展開</p></li><li><p>LLM API - このノートブックでは、Azure OpenAI 上の GPT-4o デプロイメントを使用しています。</p></li><li><p>Python バージョン 3.12.4 以降</p></li></ol><p><a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/main.ipynb">main.ipynb ノートブックからすべてのコードを実行します。</a></p><p>リポジトリを git clone し、supporting-blog-content/advanced-rag-techniques に移動して、次のコマンドを実行します。</p># Create a new virtual environment named 'rag_env'
python -m venv rag_env

# Activate the virtual environment (for Unix-based systems)
source rag_env/bin/activate

# (For Windows)
.\rag_env\Scripts\activate

# Install packages listed in requirements.txt
pip install -r requirements.txt
<p>完了したら、 <em>.env</em>を作成します。ファイルを開き、次のフィールドに入力します ( <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/.env.example"><em>.env.example</em></a>で参照されます)。有益なコメントをくれた共著者の Claude-3.5 に感謝します。</p># Elastic Cloud: Found in the 'Deployment' page of your Elastic Cloud 
# console
ELASTIC_CLOUD_ENDPOINT=""
ELASTIC_CLOUD_ID=""

# Elastic Cloud: Created during deployment setup or in 'Security' 
# settings
ELASTIC_USERNAME=""
ELASTIC_PASSWORD=""

# Elastic Cloud: The name of the index you created in Kibana or via API
ELASTIC_INDEX_NAME=""

# Azure AI Studio: Found in 'Keys and Endpoint' section of your Azure 
# OpenAI resource
AZURE_OPENAI_KEY_1=""
AZURE_OPENAI_KEY_2=""
AZURE_OPENAI_REGION=""
AZURE_OPENAI_ENDPOINT=""

# Azure AI Studio: Found in 'Deployments' section of your Azure OpenAI 
# resource
AZURE_OPENAI_DEPLOYMENT_NAME=""

# Using BAAI/bge-small-en-v1.5 because I think it is a good balance of 
# resource efficiency and performance. 
HUGGINGFACE_EMBEDDING_MODEL="BAAI/bge-small-en-v1.5"
<p>次に、取り込むドキュメントを選択し、ドキュメント フォルダーに配置します。この記事では、 <a href="https://s201.q4cdn.com/217177842/files/doc_downloads/OtherDocuments/2023/AnnualMeeting/Annual-Report-Fiscal-Year-2023.pdf">Elastic NV Annual Report 2023 を</a>使用します。これは非常に難しくて密度の高いドキュメントであり、RAG テクニックのストレス テストに最適です。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte292dc6030d496cc/6a170b40dc55de9b03e00dfc/e513b9d67adac43da794c25a5969b893127bbbe3-1440x395.jpg" alt="Elastic 年次報告書 2023" /><p>準備が整いましたので、摂取に移りましょう。<em>main.ipynb</em>を開き、最初の 2 つのセルを実行して、すべてのパッケージをインポートし、すべてのサービスを初期化します。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h2>ドキュメントの取り込み、処理、埋め込み</h2><h3>データインジェスト</h3><ul><li><p><em>個人的なメモ: LlamaIndex の便利さに驚いています。LLM や LlamaIndex が登場する前の昔、さまざまな形式のドキュメントを取り込むには、あらゆる場所から難解なパッケージを収集する、骨の折れる作業でした。今では、関数呼び出しは 1 つに減りました。野生。</em></p></li></ul><p><code>SimpleDirectoryReader</code>は<code>directory_path.</code>内のすべてのドキュメントをロードします。 <code>.pdf</code>ファイルの場合は、ドキュメント オブジェクトのリストを返します。このリストは、操作しやすいように Python 辞書に変換します。</p># llamaindex_processor.py
from llama_index.core import SimpleDirectoryReader

class LlamaIndexProcessor:
   def __init__(self):
       pass 
   
   def load_documents(self, directory_path):
       ''' 
       Load all documents in directory
       '''
       reader = SimpleDirectoryReader(input_dir=directory_path)
       return reader.load_data()

# main.ipynb
llamaindex_processor=LlamaIndexProcessor()
documents=llamaindex_processor.load_documents('./documents/')
documents=[dict(doc_obj) for doc_obj in documents]
<p>各辞書には、 <code>text</code>フィールドにキー コンテンツが含まれています。また、ページ番号、ファイル名、ファイル サイズ、タイプなどの便利なメタデータも含まれています。</p>{
  'id_': '5f76f0b3-22d8-49a8-9942-c2bbab14f63f',
  'metadata': {'page_label': '5',
   'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_path': '/Users/han/Desktop/Projects/truckasaurus/documents/Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_type': 'application/pdf',
   'file_size': 3724426,
   'creation_date': '2024-07-27',
   'last_modified_date': '2024-07-27'},
   'text': 'Table of Contents\nPage\nPART I\nItem 1. Business 3\n15 Item 1A. Risk Factors\nItem 1B. Unresolved Staff Comments 48\nItem 2. Properties 48\nItem 3. Legal Proceedings 48\nItem 4. Mine Safety Disclosures 48\nPART II\nItem 5. Market for Registrant's Common Equity, Related Stockholder Matters and Issuer Purchases of \nEquity Securities49\nItem 6. [Reserved] 49\nItem 7. Management's Discussion and Analysis of Financial Condition and Results of Operations 50\nItem 7A. Quantitative and Qualitative Disclosures About Market Risk 64\nItem 8. Financial Statements and Supplementary Data 66\nItem 9. Changes in and Disagreements With Accountants on Accounting and Financial Disclosure 100\n100\n101Item 9A. Controls and Procedures\nItem 9B. Other Information\nItem 9C. Disclosure Regarding Foreign Jurisdictions That Prevent Inspections 101\nPART III\n102\n102\n102\n102Item 10. Directors, Executive Officers and Corporate Governance\nItem 11. Executive Compensation\nItem 12. Security Ownership of Certain Beneficial Owners and Management, and Related Stockholder Matters  \nItem 13. Certain Relationships and Related Transactions, and Director Independence\nItem 14. Principal Accountant Fees and Services 102\nPART IV\n103\n105Item 15. Exhibits and Financial Statement Schedules  \nItem 16. Form 10-K Summary\nSignatures 106\ni',
   ...
}
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h3>文レベル、トークン単位のチャンキング</h3><p>最初にやるべきことは、ドキュメントを標準的な長さのチャンクに削減することです (一貫性と管理性を確保するため)。埋め込みモデルには、固有のトークン制限 (処理できる最大入力サイズ) があります。トークンはモデルが処理するテキストの基本単位です。情報の損失（コンテンツの切り捨てや省略）を防ぐために、これらの制限を超えないテキストを提供する必要があります（長いテキストを短いセグメントに分割する）。</p><p>チャンク化はパフォーマンスに大きな影響を与えます。理想的には、各チャンクは自己完結的な情報を表し、単一のトピックに関するコンテキスト情報をキャプチャします。チャンク化の方法には、文書を単語数で分割する単語レベルのチャンク化と、LLM を使用して論理ブレークポイントを識別するセマンティック チャンク化があります。</p><p>単語レベルのチャンキングは安価で高速かつ簡単ですが、文が分割され、コンテキストが壊れるリスクがあります。セマンティック チャンキングは、特に 116 ページの Elastic 年次レポートのようなドキュメントを扱う場合には、時間がかかり、コストも高くなります。</p><p>中道的なアプローチを選択しましょう。文レベルのチャンキングは依然としてシンプルですが、単語レベルのチャンキングよりもコンテキストをより効果的に保持でき、コストも大幅に削減され、処理速度も速くなります。さらに、周囲のコンテキストの一部をキャプチャし、段落を分割することによる影響を軽減するために、スライディング ウィンドウを実装します。</p># chunker.py 

import uuid
import re


class Chunker: 
    def __init__(self, tokenizer):
        self.tokenizer = tokenizer 
    
    def split_into_sentences(self, text):
        """Split text into sentences."""
        return re.split(r'(?&lt;=[.!?])\s+', text)
 
    def sentence_wise_tokenized_chunk_documents(self, documents, chunk_size=512, overlap=20, min_chunk_size=50):
        '''
        1. Split text into sentences.
        2. Tokenize using the provided tokenizer method.
        3. Build chunks up to the chunk_size limit.
        4. Create an overlap based on tokens - to preserve context.
        5. Only keep chunks that meet the minimum token size requirement.
        '''
        chunked_documents = []

        for doc in documents:
            sentences = self.split_into_sentences(doc['text'])
            tokens = []
            sentence_boundaries = [0]

            # Tokenize all sentences and keep track of sentence boundaries
            for sentence in sentences:
                sentence_tokens = self.tokenizer.encode(sentence, add_special_tokens=True)
                tokens.extend(sentence_tokens)
                sentence_boundaries.append(len(tokens))

            # Create chunks
            chunk_start = 0
            while chunk_start &lt; len(tokens):
                chunk_end = chunk_start + chunk_size

                # Find the last complete sentence that fits in the chunk
                sentence_end = next((i for i in sentence_boundaries if i &gt; chunk_end), len(tokens))
                chunk_end = min(chunk_end, sentence_end)

                # Create the chunk
                chunk_tokens = tokens[chunk_start:chunk_end]

                # Check if the chunk meets the minimum size requirement
                if len(chunk_tokens) &gt;= min_chunk_size:
                    # Create a new document object for this chunk
                    chunk_doc = {
                        'id_': str(uuid.uuid4()),
                        'chunk': chunk_tokens,
                        'original_text': self.tokenizer.decode(chunk_tokens),
                        'chunk_index': len(chunked_documents),
                        'parent_id': doc['id_'],
                        'chunk_token_count': len(chunk_tokens)
                    }

                    # Copy all other fields from the original document
                    for key, value in doc.items():
                        if key != 'text' and key not in chunk_doc:
                            chunk_doc[key] = value

                    chunked_documents.append(chunk_doc)

                # Move to the next chunk start, considering overlap
                chunk_start = max(chunk_start + chunk_size - overlap, chunk_end - overlap)

        return chunked_documents

# main.ipynb 
# Initialize Embedding Model
HUGGINGFACE_EMBEDDING_MODEL = os.environ.get('HUGGINGFACE_EMBEDDING_MODEL')
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

# Initialize Chunker
chunker=Chunker(embedder.tokenizer)
<p><code>Chunker</code>クラスは埋め込みモデルのトークナイザーを受け取り、テキストをエンコードおよびデコードします。ここで、20 個のトークンが重なり合う、それぞれ 512 個のトークンのチャンクを構築します。これを実行するには、テキストを文に分割し、それらの文をトークン化してから、トークン制限に違反することなく追加できなくなるまで、トークン化された文を現在のチャンクに追加します。</p><p>最後に、埋め込みのために文章を元のテキストにデコードし、 <code>original_text</code>というフィールドに保存します。チャンクは<code>chunk</code>というフィールドに保存されます。ノイズ（つまり、無駄なドキュメント）を減らすために、長さが 50 トークン未満のドキュメントは破棄されます。</p><p>これをドキュメント上で実行してみましょう。</p>chunked_documents=chunker.sentence_wise_tokenized_chunk_documents(documents, chunk_size=512)
<p>そして、次のようなテキストのチャンクが返されます。</p>print(chunked_documents[4]['original_text'])

[CLS] the aggregate market value of the ordinary shares held by non - affiliates of the registrant, 
based on the closing price of the shares of ordinary shares on the new york stock exchange on 
october 31, 2022 ( the last business day of the registrant 's second fiscal quarter ), was 
approximately $ 6. 1 billion. [SEP] [CLS] as of may 31, 2023, the registrant had 97, 390, 886 
ordinary shares, par value €0. 01 per share, outstanding. [SEP] [CLS] documents incorporated by 
reference portions of the registrant 's definitive proxy statement relating to the registrant 's 2
023 annual general meeting of shareholders are incorporated by reference into part iii of this annual 
...
...
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h3>メタデータの包含と生成</h3><p>ドキュメントをチャンクに分割しました。次は、データを充実させる段階です。追加のメタデータを生成または抽出したい。この追加のメタデータは、検索パフォーマンスに影響を与え、強化するために使用できます。</p><p>ドキュメントのリスト (Python 辞書) とプロセッサ関数のリストを受け取る役割を持つ<code>DocumentEnricher</code>クラスを定義します。これらの関数はドキュメントの<code>original_text</code>列を実行し、その出力を新しいフィールドに保存します。</p><p>まず、 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/nltk_processor.py">TextRank</a>を使用してキーフレーズを抽出します。TextRank は、単語間の関係に基づいて重要度をランク付けすることにより、テキストから主要なフレーズと文を抽出するグラフベースのアルゴリズムです。</p><p>次に、 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/llm.py">GPT-4oを使用してpotential_questionsを生成します</a>。</p><p>最後に、<a href="https://spacy.io/"> Spacy</a> <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/entity_extractor.py">を使用して エンティティを抽出します</a> 。</p><p>それぞれのコードは非常に長くて複雑なので、ここで再現することは控えます。ご興味があれば、以下のコード サンプルにファイルがマークされています。</p><p>データ拡充を実行してみましょう:</p># documentenricher.py
from tqdm import tqdm

class DocumentEnricher:

    def __init__(self):
        pass 

    def enrich_document(self, documents, processors, text_col='text'):
        for doc in tqdm(documents, desc="Enriching documents using processors: "+str(processors)): 
            for (processor, field) in processors: 
                metadata=processor(doc[text_col])
                if isinstance(metadata, list):
                    metadata='\n'.join(metadata)
                doc.update({field: metadata})
 
# main.ipynb
# Initialize processor classes 
nltkprocessor=NLTKProcessor() // nltk_processor.py
entity_extractor=EntityExtractor() // entity_extractor.py
gpt4o = LLMProcessor(model='gpt-4o') // llm.py

# Initialize LLM
documentenricher=DocumentEnricher()

# Create new fields in the documents - These are the outputs of the processor functions.
processors=[
    (nltkprocessor.textrank_phrases, "keyphrases"),
    (gpt4o.generate_questions, "potential_questions"),
    (entity_extractor.extract_entities, "entities")
    ]

# .enrich_document() will modify chunked_docs in place. 
# To view the results, we'll print chunked_docs in the next few cells!
documentenricher.enrich_document(chunked_docs, text_col='original_text', processors=processors)
<p>結果を見てみましょう:</p><h4>TextRankによって抽出されたキーフレーズ</h4><p>これらのキーフレーズは、チャンクの中核トピックの代わりとなります。クエリがサイバーセキュリティに関係する場合、このチャンクのスコアは向上します。</p>print(chunked_documents[25]['keyphrases'])

'elastic agent stop', 'agent stop malware', 
'stop malware ransomware', 'malware ransomware environment', 
'ransomware environment wide', 'environment wide visibility', 
'wide visibility threat', 'visibility threat detection', 
'sep cl key', 'cl key feature'
<h4>GPT-4oによって生成される潜在的な質問</h4><p>これらの潜在的な質問はユーザーのクエリと直接一致する可能性があり、スコアの向上につながります。GPT-4o に、現在のチャンクにある情報を使用して回答できる質問を生成するように指示します。</p>print(chunked_documents[25]['potential_questions'])

1. What are the primary functions that Elastic Agent provides in terms of cybersecurity?
2. Describe how Logstash contributes to data management within an IT environment.
3. List and explain any key features of Logstash mentioned in the document.
4. How does Elastic Agent enhance environment-wide visibility in threat detection?
5. What capabilities does Logstash offer for handling data beyond simple collection?
6. In what ways does the document suggest that Elastic Agent stops malware and ransomware?
7. Can you identify any relationships between the functionalities of Elastic Agent and Logstash in an integrated environment?
8. What implications might the advanced threat detection capabilities of Elastic Agent have for organizational security policies?
9. Compare and contrast the roles of Elastic Agent and Logstash based on their described functions.
10. How might the centralized collection ability of Logstash support the threat detection capabilities of Elastic Agent?
<h4>Spacyによって抽出されたエンティティ</h4><p>これらのエンティティはキーフレーズと同様の目的を果たしますが、キーフレーズ抽出では見逃される可能性のある組織や個人の名前を取得します。</p>print(chunked_documents[29]['entities'])

'appdynamics', 'apm data', 'azure sentinel', 
'microsoft', 'mcafee', 'broadcom', 'cisco', 
'dynatrace', 'coveo', 'lucidworks'
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h3>複合多体埋め込み</h3><p>追加のメタデータでドキュメントを充実させたので、この情報を活用して、より堅牢でコンテキストを認識した埋め込みを作成できます。</p><p>プロセスの現在のポイントを確認しましょう。各ドキュメントには 4 つの興味深いフィールドがあります。</p>{
    "chunk": "...",
    "keyphrases": "...", 
    "potential_questions": "...", 
    "entities": "..." 
}
<p>各フィールドはドキュメントのコンテキストに関する異なる視点を表し、LLM が重点を置くべき重要な領域を強調する可能性があります。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt84cb328fce6aae23/6a170b42964cea3e4408bbc4/aea1f513009a0c7c8545a79fad8f072a5bcae24c-1440x1067.jpg" alt="RAG のメタデータ強化パイプライン" /><p>計画としては、これらの各フィールドを埋め込み、複合埋め込みと呼ばれる埋め込みの加重合計を作成することです。</p><p>運が良ければ、この複合埋め込みにより、検索動作を制御する別の調整可能なハイパーパラメータが導入されるだけでなく、システムがよりコンテキストを認識できるようになります。</p><p>まず、main.ipynb ノートブックの先頭にインポートされたローカルに定義された埋め込みモデルを使用して、各フィールドを埋め込み、各ドキュメントを更新します。</p># EmbeddingModel defined in embedding_model.py
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

cols_to_embed=['keyphrases', 'potential_questions', 'entities']

embedding_cols=[]
for col in cols_to_embed:
    # Works on text input
    embedding_col=embedder.embed_documents_text_wise(chunked_documents, text_field=col)
    embedding_cols.append(embedding_col)
# Works on token input
embedding_col=embedder.embed_documents_token_wise(chunked_documents, token_field="chunk")
embedding_cols.append(embedding_col)
<p>各埋め込み関数は埋め込みのフィールドを返します。これは、 <code>_embedding</code>という接尾辞が付いた元の入力フィールドです。</p><p>複合埋め込みの重みを定義しましょう。</p>embedding_cols=[
                'keyphrases_embedding',
                'potential_questions_embedding',
                'entities_embedding',
                'chunk_embedding']
combination_weights=[
                    0.1,
                    0.15,
                    0.05,
                    0.7
                ]
<p>重み付けにより、ユースケースとデータの品質に基づいて各コンポーネントに優先順位を割り当てることができます。直感的に言えば、これらの重み付けの大きさは、各コンポーネントの意味的価値に依存します。チャンクテキスト自体が圧倒的に豊富なので、重み付けを 70% に割り当てます。エンティティは組織名や人名のリストだけなので最も小さいので、重み付けを 5% に割り当てます。これらの値の正確な設定は、ユースケースごとに経験的に決定する必要があります。</p><p>最後に、重み付けを適用し、複合埋め込みを作成する関数を記述しましょう。スペースを節約するために、コンポーネントの埋め込みもすべて削除します。</p>from tqdm import tqdm 
def combine_embeddings(objects, embedding_cols, combination_weights, primary_embedding='primary_embedding'):
    # Ensure the number of weights matches the number of embedding columns
    assert len(embedding_cols) == len(combination_weights), "Number of embedding columns must match number of weights"
    
    # Normalize weights to sum to 1
    weights = np.array(combination_weights) / np.sum(combination_weights)
    
    for obj in tqdm(objects, desc="Combining embeddings"):
        # Initialize the combined embedding
        combined = np.zeros_like(obj[embedding_cols[0]])
        
        # Compute the weighted sum
        for col, weight in zip(embedding_cols, weights):
            combined += weight * np.array(obj[col])
        
        # Add the new combined embedding to the object
        obj.update({primary_embedding:combined.tolist()})
        
        # Remove the original embedding columns
        for col in embedding_cols:
            obj.pop(col, None)

combine_embeddings(chunked_documents, embedding_cols, combination_weights)
<p>これで書類の処理は完了です。次のようなドキュメント オブジェクトのリストが作成されました。</p>{ 'id_': '7fe71686-5cd0-4831-9e79-998c6dbeae0c', 'chunk': [2312, 14613, ...], 'original_text': 'if an emerging growth company, indicate by check mark if the registrant has elected not to use the extended ...', 'chunk_index': 3, 'chunk_token_count': 399, 'metadata': {'page_label': '3', 'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf', ... 'keyphrases': 'sep cl unk\ncheck mark registrant\ncl unk indicate\nunk indicate check\nindicate check mark\nprincipal executive office\naccelerate filer unk\ncompany unk emerge\nunk emerge growth\nemerge growth company', 'potential_questions': '1. What are the different types of registrant statuses mentioned in the document?\n2. Under what section of the Sarbanes-Oxley Act must registrants file a report on the effectiveness of their internal ...', 'entities': 'the effe ctiveness of\nsection 13\nSEP\nUNK\nsection 21e\n1934\n1933\nu. s. c.\nsection 404\nsection 12\nal', 'primary_embedding': [-0.3946287803351879, -0.17586839850991964, ...] }
<h4>Elasticへのインデックス</h4><p>ドキュメントを Elastic Search に一括アップロードしてみましょう。この目的のために、私はずっと前に<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/elastic_helpers.py"><code>elastic_helpers.py</code></a>で Elastic Helper 関数のセットを定義しました。これは非常に長いコードなので、関数呼び出しに注目してみましょう。</p><p><code>es_bulk_indexer.bulk_upload_documents</code> Elasticsearch の便利な動的マッピングを活用して、辞書オブジェクトの任意のリストで動作します。</p># Initialize Elasticsearch
ELASTIC_CLOUD_ID = os.environ.get('ELASTIC_CLOUD_ID')
ELASTIC_USERNAME = os.environ.get('ELASTIC_USERNAME')
ELASTIC_PASSWORD = os.environ.get('ELASTIC_PASSWORD')
ELASTIC_CLOUD_AUTH = (ELASTIC_USERNAME, ELASTIC_PASSWORD)
es_bulk_indexer = ESBulkIndexer(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)
es_query_maker = ESQueryMaker(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)

# Define Index Name
index_name=os.environ.get('ELASTIC_INDEX_NAME')


# Create index and bulk upload 
index_exists = es_bulk_indexer.check_index_existence(index_name=index_name)
if not index_exists:
    logger.info(f"Creating new index: {index_name}")
    es_bulk_indexer.create_es_index(es_configuration=BASIC_CONFIG, index_name=index_name)

success_count = es_bulk_indexer.bulk_upload_documents(
    index_name=index_name, 
    documents=chunked_documents, 
    id_col='id_',
    batch_size=32
)
<p>Kibana にアクセスして、すべてのドキュメントがインデックスされていることを確認します。全部で224個あるはずです。こんなに大きな文書にしては悪くないですね!</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8efeface6effe01d/6a170b447d8d67652870e72a/1b3b07f6b98ceb65f6594ce4be83c5b0ed7e7cf9-1440x1380.jpg" alt="インデックスキバナ" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p><h2>猫の休憩</h2><p>ちょっと休憩しましょう。記事がちょっと重いのはわかっています。私の猫を見てください:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1db5595f71c12ff/6a170b450e2e49940241a0fe/baca4eb52b801b21ced97352cc55462f0a12d6b0-969x996.jpg" alt="ハンパイプライン" /><p>愛らしい。帽子がなくなってしまったので、彼女がそれを盗んでどこかに隠したのではないかと半分疑っています :(</p><p>ここまで来られたことおめでとうございます :)</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">パート 2</a>では、RAG パイプラインのテストと評価についてご紹介します。</p><h2>付記</h2><h3>定義</h3><p><strong>1. 文のチャンキング</strong></p><ul><li><p>RAG システムでテキストをより小さな意味のある単位に分割するために使用される前処理手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>入力: 大きなテキストブロック（例: 文書、段落）</p></li><li><p>出力: 小さなテキストセグメント (通常は文または小さな文のグループ)</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>きめ細やかでコンテキストに特化したテキストセグメントを作成する</p></li><li><p>より正確なインデックス作成と検索が可能</p></li><li><p>RAGシステムで取得した情報の関連性を向上</p></li></ul></li><li><p><em>特徴:</em> </p><ul><li><p>セグメントは意味的に意味がある</p></li><li><p>独立してインデックスを作成し、検索できる</p></li><li><p>多くの場合、独立した理解可能性を確保するためにある程度の文脈が保持される</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>検索精度の向上</p></li><li><p>RAGパイプラインのより集中的な拡張を可能にします</p></li></ul></li></ul><p><strong>2. HyDE（仮想文書埋め込み）</strong></p><ul><li><p>LLM を使用して、RAG システムでのクエリ拡張用の仮想ドキュメントを生成する手法。</p></li><li><p><em>プロセス：</em>  </p><ol><li><p>LLMへの入力クエリ</p></li><li><p>LLMはクエリに答える仮説文書を生成する</p></li><li><p>生成されたドキュメントを埋め込む</p></li><li><p>ベクトル検索に埋め込みを使用する</p></li></ol></li><li><p><em>主な違い:</em> </p><ul><li><p>従来のRAG: クエリとドキュメントを一致させる</p></li><li><p>HyDE: 文書を文書と照合する</p></li></ul></li><li><p><em>目的：</em> </p><ul><li><p>特に複雑または曖昧なクエリの検索パフォーマンスを向上します</p></li><li><p>短いクエリよりも豊富な意味コンテキストをキャプチャする</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>LLMの知識を活用してクエリを拡張する</p></li><li><p>検索された文書の関連性が向上する可能性がある</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>追加のLLM推論が必要となり、レイテンシとコストが増加する</p></li><li><p>パフォーマンスは生成された仮想文書の品質に依存する</p></li></ul></li></ul><p><strong>3. 逆パッキング</strong></p><ul><li><p>RAG システムで、検索結果を LLM に渡す前に並べ替えるために使用される手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>検索エンジン (Elasticsearch など) は、関連性の高い順にドキュメントを返します。</p></li><li><p>順序は逆になり、最も関連性の高いドキュメントが最後に配置されます。</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>LLM の新しさバイアスを利用します。LLM は、それぞれのコンテキストにおける最新の情報に重点を置く傾向があります。</p></li><li><p>LLM のコンテキスト ウィンドウ内で最も関連性の高い情報が「最新」であることを保証します。</p></li></ul></li><li><p><em>例:</em>元の順序: [最も関連性の高い順、2番目に関連性の高い順、3番目に関連性の高い順、...] 逆の順序: [...、3番目に関連性の高い順、2番目に関連性の高い順、最も関連性の高い順]</p></li></ul><p><strong>4. クエリの分類</strong></p><ul><li><p>クエリに RAG が必要かどうか、または LLM によって直接回答できるかどうかを判断して、RAG システムの効率を最適化する手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>使用中の LLM に固有のカスタム データセットを開発する</p></li><li><p>特殊な分類モデルをトレーニングする</p></li><li><p>モデルを使用して受信したクエリを分類する</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>不要なRAG処理を回避することでシステム効率を向上</p></li><li><p>最も適切な応答メカニズムにクエリを直接送信する</p></li></ul></li><li><p><em>要件：</em> </p><ul><li><p>LLM固有のデータセットとモデル</p></li><li><p>精度を維持するための継続的な改良</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>単純なクエリの計算オーバーヘッドを削減</p></li><li><p>非RAGクエリの応答時間を改善する可能性がある</p></li></ul></li></ul><p><strong>5. 要約</strong></p><ul><li><p>RAG システムで検索された文書を圧縮する手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>関連文書を取得する</p></li><li><p>各文書の簡潔な要約を生成する</p></li><li><p>RAG パイプラインでは完全なドキュメントではなく要約を使用する</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>重要な情報に焦点を当ててRAGのパフォーマンスを向上させる</p></li><li><p>関連性の低いコンテンツからのノイズや干渉を減らす</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>LLM回答の関連性が向上する可能性がある</p></li><li><p>コンテキスト制限内でより多くのドキュメントを含めることができます</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>要約時に重要な詳細が失われるリスク</p></li><li><p>要約生成のための追加の計算オーバーヘッド</p></li></ul></li></ul><p><strong>6. メタデータの包含</strong></p><ul><li><p>追加のコンテキスト情報でドキュメントを充実させる手法。</p></li><li><p><em>メタデータの種類:</em>  </p><ul><li><p>キーフレーズ</p></li><li><p>タイトル</p></li><li><p>日付</p></li><li><p>著者詳細</p></li><li><p>宣伝文句</p></li></ul></li><li><p><em>目的：</em> </p><ul><li><p>RAGシステムで利用可能なコンテキスト情報を増やす</p></li><li><p>LLMに文書の内容と関連性をより明確に理解させる</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>検索精度が向上する可能性がある</p></li><li><p>LLMの文書有用性を評価する能力を強化する</p></li></ul></li><li><p><em>実装：</em> </p><ul><li><p>文書の前処理中に実行できる</p></li><li><p>追加のデータ抽出または生成手順が必要になる場合があります</p></li></ul></li></ul><p><strong>7. 複合多体埋め込み</strong></p><ul><li><p>異なるドキュメント コンポーネントごとに個別の埋め込みを作成する RAG システム用の高度な埋め込み手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>関連するフィールド（例：タイトル、キーフレーズ、宣伝文句、メインコンテンツ）を特定する</p></li><li><p>各フィールドごとに個別の埋め込みを生成する</p></li><li><p>これらの埋め込みを結合または保存して検索に使用します</p></li></ol></li><li><p><em>標準的なアプローチとの違い:</em> </p><ul><li><p>従来型: ドキュメント全体の単一の埋め込み</p></li><li><p>複合: さまざまなドキュメントの側面に対応する複数の埋め込み</p></li></ul></li><li><p><em>目的：</em> </p><ul><li><p>よりニュアンス豊かで文脈を考慮した文書表現を作成する</p></li><li><p>文書内のより多様なソースから情報を取得する</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>曖昧なクエリや多面的なクエリのパフォーマンスが向上する可能性があります</p></li><li><p>検索時にさまざまな文書の側面をより柔軟に重み付けできます</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>埋め込みストレージと検索プロセスの複雑さが増す</p></li><li><p>より洗練されたマッチングアルゴリズムが必要になる場合があります</p></li></ul></li></ul><p><strong>8. クエリエンリッチメント</strong></p><ul><li><p>元のクエリを関連用語で拡張し、検索範囲を広げる手法。</p></li><li><p><em>プロセス：</em> </p><ol><li><p>元のクエリを分析する</p></li><li><p>同義語や意味的に関連するフレーズを生成する</p></li><li><p>クエリに以下の追加用語を追加します</p></li></ol></li><li><p><em>目的：</em> </p><ul><li><p>文書コーパス内の潜在的な一致の範囲を拡大する</p></li><li><p>特定の言語や専門用語を含むクエリの検索パフォーマンスを向上</p></li></ul></li><li><p><em>メリット：</em> </p><ul><li><p>元の検索語句と完全に一致しない関連文書を取得する可能性がある</p></li><li><p>クエリとドキュメント間の語彙の不一致を克服するのに役立ちます</p></li></ul></li><li><p><em>課題:</em> </p><ul><li><p>慎重に実装しないとクエリドリフトのリスクがある</p></li><li><p>検索プロセスにおける計算オーバーヘッドが増加する可能性がある</p></li></ul></li></ul><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">トップに戻る</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 14 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>