<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/cn/search-labs/author/han-xiang-choong</link>
    </image>
    <link>https://www.elastic.co/cn/search-labs/author/han-xiang-choong</link>
    <atom:link href="https://www.elastic.co/cn/search-labs/rss/author/han-xiang-choong.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[cn]]></language>
    <lastBuildDate>Wed, 23 Sep 2026 08:06:47 GMT</lastBuildDate>
  <item>
    <title><![CDATA[高级 RAG 技术第 2 部分：查询和测试]]></title>
    <description><![CDATA[讨论并实施可提高 RAG 性能的技术。第 2 部分（共 2 部分），重点是查询和测试高级 RAG 管道。]]></description>
    <content:encoded><![CDATA[<p><em>所有代码都可以 </em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em>在 Searchlabs 软件仓库的 advanced-rag-techniques 分支中</em></a>找到 <em>。</em></p><p>欢迎阅读我们关于高级 RAG 技术文章的第二部分！在<a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">本系列的第 1 部分</a>中，我们建立、讨论并实施了高级 RAG 管道的数据处理组件：</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="高级 RAG 管道" /><p>在这一部分，我们将继续查询和测试我们的实现。让我们直奔主题！</p><h3>目录</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#searching-and-retrieving,-generating-answers">搜索和检索，生成答案</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#enriching-queries-with-synonyms">用同义词丰富查询</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-hypothetical-document-embedding">HyDE（假设文档嵌入）</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hybrid-search">混合搜索</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#experiments">实验</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#summary-of-results">结果摘要</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-1-who-audits-elastic">测试 1：谁审核了 Elastic？</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-2--total-revenue-2023">测试 2：2023 年总收入</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-1">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-1">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-3-what-product-does-growth-primarily-depend-on-how-much">测试 3：增长主要依赖于什么产品？多少钱？</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-2">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-2">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-4-describe-employee-benefit-plan">测试 4：说明员工福利计划</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-3">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-3">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-5-which-companies-did-elastic-acquire">测试 5：Elastic 收购了哪些公司？</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-4">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-4">SimpleRAG</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#conclusion">结论</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#appendix">附录</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#prompts">提示</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#rag-question-answering-prompt">RAG 问题解答提示</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#elastic-query-generator-prompt">弹性查询生成器提示</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#potential-questions-generator-prompt">潜在问题生成器提示</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-generator-prompt">HyDE 生成器提示</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#sample-hybrid-search-query">混合搜索查询示例</a></p></li></ul></li></ul><h2>搜索和检索，生成答案</h2><p>让我们提出第一个问题，最好是主要在年度报告中找到的一些信息。怎么样？</p>Who audits Elastic?"
<p>现在，让我们运用一些技术来增强查询。</p><h3>用同义词丰富查询</h3><p>首先，让我们增强查询措辞的多样性，并将其转化为可轻松处理成 Elasticsearch 查询的形式。我们将借助 GPT-4o 将查询转换为 OR 子句列表。让我们来写下这个提示：</p>
ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<p>当应用到我们的查询时，GPT-4o 会生成基本查询和相关词汇的同义词。</p>'audits elastic OR 
elasticsearch audits OR 
elastic auditor OR 
elasticsearch auditor OR 
elastic audit firm OR 
elastic audit company OR 
elastic audit organization OR 
elastic audit service'
<p>在<code>ESQueryMaker</code> 类中，我定义了一个分割查询的函数：</p>def parse_or_query(self, query_text: str) -&gt; List[str]:
    # Split the query by 'OR' and strip whitespace from each term
    # This converts a string like "term1 OR term2 OR term3" into a list ["term1", "term2", "term3"]
    return [term.strip() for term in query_text.split(' OR ')]
<p>它的作用是将这串 OR 子句拆分成一个术语列表，使我们能够对关键文档字段进行多重匹配：</p>["original_text", 'keyphrases', 'potential_questions', 'entities']
<p>最后得出了这个疑问：</p> 'query': {
    'bool': {
        'must': [
            {
                'multi_match': {
                'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
                'fields': [
                    'original_text',
                'keyphrases',
                'potential_questions',
                'entities'
                ],
                'type': 'best_fields',
                'operator': 'or'
                }
            }
      ]
<p>这比原始查询涵盖的范围更广，有望降低因忘记同义词而错过搜索结果的风险。但我们可以做得更多。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">返回顶部</a></p><h3>HyDE（假设文档嵌入）</h3><p>让我们再次利用 GPT-4o 来实现<a href="https://arxiv.org/abs/2212.10496">HyDE</a>。</p><p>HyDE 的基本前提是生成一个假设文档--一种可能包含原始查询答案的文档。文件的真实性或准确性并不重要。有鉴于此，让我们写下下面的提示：</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<p>由于矢量搜索通常是通过余弦矢量相似性进行操作的，因此 HyDE 的前提是，我们可以通过文档与文档的匹配，而不是查询与文档的匹配，来获得更好的结果。</p><p>我们关心的是结构、流程和术语。事实性不强。GPT-4o 可以输出这样的 HyDE 文档：</p>'Elastic N.V., the parent company of Elastic, the organization known for developing Elasticsearch, is subject to audits to ensure financial accuracy, 
regulatory compliance, and the integrity of its financial statements. The auditing of Elastic N.V. is typically conducted by an external, 
independent auditing firm. This is common practice for publicly traded companies to provide stakeholders with assurance regarding the company\'s 
financial position and operations.\n\nThe primary external auditor for Elastic is the audit firm Ernst &amp; Young LLP (EY). Ernst &amp; Young is one of the 
four largest professional services networks in the world, commonly referred to as the "Big Four" audit firms. These firms handle a substantial number 
of audits for major corporations around the globe, ensuring adherence to generally accepted accounting principles (GAAP) and international financial 
reporting standards (IFRS).\n\nThe audit process conducted by EY involves several steps. Initially, the auditors perform a risk assessment to identify 
areas where misstatements due to error or fraud could occur. They then design audit procedures to test the accuracy and completeness of financial statements,
 which include examining financial transactions, assessing internal controls, and reviewing compliance with relevant laws and regulations. Upon completion of 
 the audit, Ernst &amp; Young issues an audit report, which includes the auditor’s opinion on whether the financial statements are free from material misstatement 
 and are presented fairly in accordance with the applicable financial reporting framework.\n\nIn addition to external audits by firms like Ernst &amp; Young, 
 Elastic may also be subject to internal audits. Internal audits are performed by the company’s own internal auditors to evaluate the effectiveness of internal 
 controls, risk management, and governance processes.\n\nOverall, the auditing process plays a crucial role in maintaining the transparency and reliability of 
 Elastic\'s financial information, providing confidence to investors, regulators, and other stakeholders.'
<p>它看起来非常可信，是我们希望索引的文档类型的理想候选者。我们将把它嵌入并用于混合搜索。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">返回顶部</a></p><h3>混合搜索</h3><p>这是我们搜索逻辑的核心。我们的词法搜索组件将是生成的 OR 子句字符串。我们的密集矢量组件将是嵌入式 HyDE 文档（又称搜索矢量）。我们使用 KNN 来有效识别与搜索向量最接近的几个候选文档。我们将词法搜索组件默认称为<em>TF-IDF 和 BM25 评分</em>。最后，将采用<a href="https://arxiv.org/abs/2407.01219">Wang 等</a>人推荐的 30/70 比例合并词性和密集向量得分。</p>def hybrid_vector_search(self, index_name: str, query_text: str, query_vector: List[float], 
                         text_fields: List[str], vector_field: str, 
                         num_candidates: int = 100, num_results: int = 10) -&gt; Dict:
    """
    Perform a hybrid search combining text-based and vector-based similarity.

    Args:
        index_name (str): The name of the Elasticsearch index to search.
        query_text (str): The text query string, which may contain 'OR' separated terms.
        query_vector (List[float]): The query vector for semantic similarity search.
        text_fields (List[str]): List of text fields to search in the index.
        vector_field (str): The name of the field containing document vectors.
        num_candidates (int): Number of candidates to consider in the initial KNN search.
        num_results (int): Number of final results to return.

    Returns:
        Dict: A tuple containing the Elasticsearch response and the search body used.
    """
    try:
        # Parse the query_text into a list of individual search terms
        # This splits terms separated by 'OR' and removes any leading/trailing whitespace
        query_terms = self.parse_or_query(query_text)

        # Construct the search body for Elasticsearch
        search_body = {
            # KNN search component for vector similarity
            "knn": {
                "field": vector_field,  # The field containing document vectors
                "query_vector": query_vector,  # The query vector to compare against
                "k": num_candidates,  # Number of nearest neighbors to retrieve
                "num_candidates": num_candidates  # Number of candidates to consider in the KNN search
            },
            "query": {
                "bool": {
                    # The 'must' clause ensures that matching documents must satisfy this condition
                    # Documents that don't match this clause are excluded from the results
                    "must": [
                        {
                            # Multi-match query to search across multiple text fields
                            "multi_match": {
                                "query": " ".join(query_terms),  # Join all query terms into a single space-separated string
                                "fields": text_fields,  # List of fields to search in
                                "type": "best_fields",  # Use the best matching field for scoring
                                "operator": "or"  # Match any of the terms (equivalent to the original OR query)
                            }
                        }
                    ],
                    # The 'should' clause boosts relevance but doesn't exclude documents
                    # It's used here to combine vector similarity with text relevance
                    "should": [
                        {
                            # Custom scoring using a script to combine vector and text scores
                            "script_score": {
                                "query": {"match_all": {}},  # Apply this scoring to all documents that matched the 'must' clause
                                "script": {
                                    # Script to combine vector similarity and text relevance
                                    "source": """
                                    # Calculate vector similarity (cosine similarity + 1)
                                    # Adding 1 ensures the score is always positive
                                    double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;
                                    # Get the text-based relevance score from the multi_match query
                                    double text_score = _score;
                                    # Combine scores: 70% vector similarity, 30% text relevance
                                    # This weighting can be adjusted based on the importance of semantic vs keyword matching
                                    return 0.7 * vector_score + 0.3 * text_score;
                                    """,
                                    # Parameters passed to the script
                                    "params": {
                                        "query_vector": query_vector,  # Query vector for similarity calculation
                                        "vector_field": vector_field  # Field containing document vectors
                                    }
                                }
                            }
                        }
                    ]
                }
            }
        }

        # Execute the search request against the Elasticsearch index
        response = self.conn.search(index=index_name, body=search_body, size=num_results)
        # Log the successful execution of the search for monitoring and debugging
        logger.info(f"Hybrid search executed on index: {index_name} with text query: {query_text}")
        # Return both the response and the search body (useful for debugging and result analysis)
        return response, search_body
    except Exception as e:
        # Log any errors that occur during the search process
        logger.error(f"Error executing hybrid search on index: {index_name}. Error: {e}")
        # Re-raise the exception for further handling in the calling code
        raise e
<p>最后，我们可以拼凑出一个 RAG 函数。我们的 RAG（从询问到答复）将遵循这一流程：</p><ol><li><p>将查询转换为 OR 子句。</p></li><li><p>生成 HyDE 文档并嵌入。</p></li><li><p>将二者作为混合搜索的输入。</p></li><li><p>检索前 N 个结果，将它们倒转，使最相关的得分是 LLM 上下文内存中"最近的" （反向打包） 反向打包示例：查询："Elasticsearch 查询优化技术" 检索文档（按相关性排序）：  LLM 上下文的反向顺序：  通过颠倒顺序，最相关的信息(1)会出现在上下文的最后，从而可能在生成答案时受到 LLM 的更多关注。</p><ol><li><p>"使用 bool 查询可有效组合多个搜索条件。"</p></li><li><p>"实施缓存策略，缩短查询响应时间。"</p></li><li><p>"优化索引映射，提高搜索性能。"</p></li><li><p>"优化索引映射，提高搜索性能。"</p></li><li><p>"实施缓存策略，缩短查询响应时间。"</p></li><li><p>"使用 bool 查询可有效组合多个搜索条件。"</p></li></ol></li><li><p>将上下文传递给 LLM 生成。</p></li></ol>def get_context(index_name, 
                match_query, 
                text_query, 
                fields, 
                num_candidates=100, 
                num_results=20, 
                text_fields=["original_text", 'keyphrases', 'potential_questions', 'entities'], 
                embedding_field="primary_embedding"):

    embedding=embedder.get_embeddings_from_text(text_query)

    results, search_body = es_query_maker.hybrid_vector_search(
        index_name=index_name,
        query_text=match_query,
        query_vector=embedding[0][0],
        text_fields=text_fields,
        vector_field=embedding_field,
        num_candidates=num_candidates,
        num_results=num_results
    )

    # Concatenates the text in each 'field' key of the search result objects into a single block of text.
    context_docs=['\n\n'.join([field+":\n\n"+j['_source'][field] for field in fields]) for j in results['hits']['hits']]

    # Reverse Packing to ensure that the highest ranking document is seen first by the LLM.
    context_docs.reverse()
    return context_docs, search_body

def retrieval_augmented_generation(query_text):
    match_query= gpt4o.generate_query(query_text)
    fields=['original_text']

    hyde_document=gpt4o.generate_HyDE(query_text)

    context, search_body=get_context(index_name, match_query, hyde_document, fields)

    answer= gpt4o.basic_qa(query=query_text, context=context)
    return answer, match_query, hyde_document, context, search_body

<p>让我们运行查询并得到答案：</p>According to the context, Elastic N.V. is audited by an independent registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of independent registered public accounting firm," which states:

"We have audited the accompanying consolidated balance sheets of Elastic N.V. [...] / s / pricewaterhouseco."
<p>不错。没错。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">返回顶部</a></p><h2>实验</h2><p>现在有一个重要问题需要回答。我们在这些实施中投入了如此多的精力和额外的复杂性，究竟得到了什么？</p><p>让我们来做个小小的比较。我们实施的 RAG 管道与基线混合搜索相比，没有任何增强功能。我们将进行一系列小测试，看看是否会发现任何实质性差异。我们将把刚刚实现的 RAG 称为 AdvancedRAG，把基本管道称为 SimpleRAG。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" alt="简单的 RAG 管道" /><h4>结果摘要</h4><p>本表总结了两种 RAG 管道的五次测试结果。我根据答案的细节和质量来判断每种方法的相对优劣，但这完全是主观判断。现将实际答案转载于下表，供您参考。说了这么多，让我们来看看他们的表现如何！</p><p>SimpleRAG 无法回答问题 1&amp; 5。AdvancedRAG 对问题 2、3 和 4 的回答也要详细得多。基于更多的细节，我认为 AdvancedRAG 的答案质量更高。</p><p>测试</p><p>问题</p><p>高级 RAG 性能</p><p>SimpleRAG 性能</p><p>AdvancedRAG 延迟</p><p>SimpleRAG 延迟</p><p>优胜者</p><p>1</p><p>谁审核 Elastic？</p><p>正确确定普华永道为审计员。</p><p>未能确定审计员。</p><p>11.6s</p><p>4.4s</p><p>AdvancedRAG</p><p>2</p><p>2023 年的总收入是多少？</p><p>提供了正确的收入数字。包括往年收入的补充情况。</p><p>提供了正确的收入数字。</p><p>13.3s</p><p>2.8s</p><p>AdvancedRAG</p><p>3</p><p>增长主要依靠什么产品？多少钱？</p><p>正确指出弹性云是关键驱动因素。包括总体收入情况&amp; 。</p><p>正确指出弹性云是关键驱动因素。</p><p>14.1s</p><p>12.8s</p><p>AdvancedRAG</p><p>4</p><p>说明员工福利计划</p><p>全面介绍了退休计划、医疗计划和其他福利。包括不同年份的具体捐款额。</p><p>提供了很好的福利概览，包括薪酬、退休计划、工作环境和 Elastic Cares 计划。</p><p>26.6s</p><p>11.6s</p><p>AdvancedRAG</p><p>5</p><p>Elastic 收购了哪些公司？</p><p>正确列出了报告中提到的近期收购（CmdWatch、Build Security 和 Optimyze）。提供了一些收购日期和收购价格。</p><p>未能从提供的上下文中检索到相关信息。</p><p>11.9s</p><p>2.7s</p><p>AdvancedRAG</p><h4>测试 1：谁审核了 Elastic？</h4><p>AdvancedRAG</p>According to the context, Elastic N.V. is audited by an independent 
registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of 
independent registered public accounting firm," which states:
"We have audited the accompanying consolidated balance sheets 
of Elastic N.V. [...] / s / pricewaterhouseco."
<p>SimpleRAG</p>According to the provided context, no specific auditing firm 
or company is mentioned as the auditor for Elastic. 
The context provided talks about various features, strengths, 
deployment options, and other details related to Elastic and 
its operations but does not specify which auditing firm is 
responsible for auditing Elastic's financial statements or 
internal controls.
<p><strong>摘要</strong>：SimpleRAG 没有将普华永道确定为审计机构</p><p>好吧，这其实挺让人惊讶的。这看起来像是 SimpleRAG 的搜索失败。没有检索到与审计有关的文件。让我们在下一个测试中降低难度。</p><h4>测试 2：2023 年总收入</h4><p>AdvancedRAG</p>### Total Revenue for 2023

According to the provided context, the total revenue for the 
year ended April 30, 2023, was **$1,068,989,000**. This 
information can be found several times in the provided text. 
Specifically, one part states:

"The following table sets forth our results of operations 
for the periods presented in dollars and as a percentage of 
our total revenue. ... year ended April 30, 2023, ... 
total revenue 1,068,989".

Additionally, it is mentioned here as well:

"For the years ended April 30, 2023, 2022, and 2021, 
representing year-over-year growth of 24% for the year 
ended April 30, 2023... total revenue $1,068,989".

Therefore, the total revenue for 2023 is $1.069 billion.
<p>SimpleRAG</p>According to the context provided, the total revenue for the fiscal year ended April 30, 2023, is:

$1,068,989,000 (or approximately $1.1 billion).
<p><strong>摘要</strong>：两个 RAG 都得到了正确答案：2023 年总收入为 1,068,989,000 美元</p><p>他们都在这里。看来，AdvancedRAG 可能获得了更多的文件？当然，答案会更加详细，并包含往年的信息。考虑到我们所做的改进，这是意料之中的，但现在下结论还为时过早。</p><p>让我们提高难度。</p><h4>测试 3：增长主要依赖于什么产品？多少钱？</h4><p>AdvancedRAG</p>### Product Growth Primarily Depends On:
According to the context, "our revenue is derived primarily from 
subscriptions of search, observability and security built into 
the elastic stack." The document mentions various strategies and 
factors contributing to the company’s growth which include investment 
in "elastic cloud," "elastic stack," and the conversion of free 
users to paid subscribers.

### Quantitative Data:
- **Revenue**: The document states, "Our total revenue was $1.1 billion, 
$862.4 million, and $608.5 million for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Subscription Revenue**: "Subscriptions accounted for 92%, 93%, 
and 93% of our total revenue for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Elastic Cloud Contribution**: "For the years ended April 30, 2023, 
2022, and 2021, Elastic Cloud contributed 40%, 35%, and 27% of our 
total revenue, respectively."
- **Customer Growth**: The number of customers increased from 
"over 15,000 as of April 30, 2021," to "approximately 20,200 as of 
April 30, 2023."

### Summary:
Product growth for the company primarily depends on the "Elastic Stack" 
(which includes search, observability, and security solutions) and 
particularly the "Elastic Cloud." The Elastic Cloud's growing contribution 
to revenue highlights its significance, which grew from 27% in 2021 to 40% in 2023.
<p>SimpleRAG</p>According to the provided context, growth primarily depends on **Elastic Cloud**, 
which is described as the company's "family of cloud-based offerings." 
The context also emphasizes the significance of Elastic Cloud in achieving 
the company's long-term growth potential and its role in increasing subscription revenue.

From the context:

- "Increasing adoption of Elastic Cloud... is an important growth opportunity 
for our business. Organizations are increasingly looking for hosted 
deployment alternatives with reduced administrative burdens."
- "We believe that offering Elastic Cloud is important for achieving 
our long-term growth potential, and we expect Elastic Cloud’s contribution 
to our subscription revenue to continue to increase."

In terms of the contribution to total revenue, here are the specific figures provided:

- Elastic Cloud contributed 40% of the total revenue for the year ended April 30, 2023.
- Elastic Cloud contributed 35% of the total revenue for the year ended April 30, 2022.
- Elastic Cloud contributed 27% of the total revenue for the year ended April 30, 2021.

This progression shows a clear upward trend in the importance and 
contribution of Elastic Cloud to the company's overall revenue.
<p><strong>摘要</strong>：两个 RAG 都正确地将弹性云确定为主要增长动力。不过，AdvancedRAG 包含更多细节，将订阅收入和客户增长考虑在内，并明确提及其他 Elastic 产品。</p><h4>测试 4：说明员工福利计划</h4><p>AdvancedRAG</p>### Employee Benefit Plan Overview

The provided context outlines various elements of the employee benefit plan at 
Elastic, often referred to as "Elasticians." Here are the key components 
described in the texts:

1. **401(k) Plan**:
   - **Coverage**: The 401(k) plan is available to substantially all U.S. 
   employees who meet minimum age and service requirements.
   - **Contributions**: Elastic makes contributions to the 401(k) plan up to 
   6% of the participating employee’s W-2 earnings and wages.
   - **Expenses**: For the fiscal years ended April 30, Elastic recorded 
   expenses of $17.9 million (2023), $15.2 million (2022), and $11.4 million (2021) 
   related to the 401(k) plan.
   - **Defined-Contribution Plans in Other Countries**: Elastic has 
   defined-contribution plans in various other countries and recorded respective 
   expenses of $9.4 million (2023), $7.2 million (2022), and $5.1 million (2021).

2. **Stock-Based Compensation**:
   - **Types of Awards**: Stock options, restricted stock units (RSUs), 
   and shares under the Employee Stock Purchase Plan (ESPP).
   - **Fair Value Measurement**: Fair value of these stock awards is 
   measured using models like Black-Scholes.
   - **Employee Stock Purchase Plan (2022 ESPP)**: 
     - Started in 2022, it allows employees to acquire ordinary 
     shares at a discount (85% of the market value at the beginning 
     or end of the offering period).
     - Offering periods are approximately six months long.

3. **Total Rewards Compensation**:
   - **Components**: Includes cash compensation as well as equity awards, 
   reflecting a comprehensive interest in the well-being of the employees.
   - **Programs**: Market competitive medical and dental programs, 
   mental health, and holistic well-being programs.
   - **Paid Time Off**: Features 16 weeks of paid leave for new 
   parents and life-planning benefits.

4. **Retirement and Income Protection Plans**:
   - **401(k) Match**: Elastic provides a dollar-for-dollar 
   match up to 6% of eligible earnings up to a plan-limit 
   maximum for U.S.-based employees.
   - **Similar Plans Outside the U.S.**: Competitive retirement 
   plans available for employees based outside the United States.

5. **Flexible Work Environment**: 
   - **Distributed Workforce**: Elastic operates with a global 
   distributed workforce model, allowing employees to work remotely 
   and providing a broad candidate pool.
   
6. **Learning and Development**:
   - **Programs**: Access to on-demand video-based learning, 
   leadership development programs, and other specific training 
   to develop managers and leaders.

7. **Employee Engagement and Surveys**:
   - **Mechanisms**: Annual employee engagement survey and 
   mid-year pulse survey.
   - **Objectives**: To monitor morale, develop action plans, 
   and maintain high engagement scores.

8. **Onboarding Program**:
   - **X-School**: An extensive new-hire orientation program to 
   help new employees meet peers and learn about the company's 
   products and solutions.

In summary, Elastic's employee benefit plan encompasses a broad 
spectrum of financial, health, well-being, and professional development 
benefits designed to support and enhance the lives of its employees. 
The plan is both competitive and comprehensive, aiming to attract,
retain, and engage its global workforce.
<p>SimpleRAG</p>The employee benefit plan at the Company includes several key 
components aimed at enhancing the overall well-being and 
professional development of employees, referred to as "Elasticians." 

1. **Compensation and Benefits:**
   - The Company provides market-competitive compensation 
   packages which typically include cash compensation and equity awards.
   - Comprehensive medical and dental programs are offered, 
   emphasizing mental health and holistic well-being.
   - Paid time off programs feature:
     - 16 weeks of paid leave for new parents.
     - Life-planning benefits.
     - Travel reimbursements for certain healthcare services.

2. **Retirement and Income Protection Plans:**
   - In the United States, employees are eligible for a 401k 
   plan with a dollar-for-dollar match by the Company up to 6% 
   of eligible earnings, subject to a plan limit.
   - Similar competitive retirement and income protection plans 
   are available for employees outside the United States.

3. **Work Environment:**
   - The Company offers a flexible work environment, allowing most 
   employees to work from anywhere as often as they would like.
   - Physical offices around the world are available for employees 
   who prefer to work on-site or need a space for face-to-face interactions.

4. **Community Involvement:**
   - The Elastic Cares Program supports charitable organizations 
   that matter to employees on both local and global levels.
   - The program includes:
     - Donation matching.
     - Providing the Company's technology for free to certain 
     nonprofit organizations.
     - Volunteer time off initiative.
<p><strong>摘要</strong>：AdvancedRAG 更深入、更详细地介绍了美国员工的 401K 计划，以及美国以外地区的缴费计划。报告还提到了 "健康与福利计划"，但没有提到 SimpleRAG 提到的 "Elastic Cares 计划"。</p><h4>测试 5：Elastic 收购了哪些公司？</h4><p>AdvancedRAG</p>Elastic acquired several companies over the years to enhance its 
technology and market presence. According to the provided context, 
the specific acquisitions include:

1. **CmdWatch Security Inc.**: Acquired on September 17, 2021. 
The total purchase consideration was $77.8 million.
2. **Build Security Ltd.**: Acquired on September 2, 2021. The 
combined purchase price for Build Security Ltd. and another acquisition was $57.2 million.
3. **Optimyze Cloud Inc.**: Acquired on November 1, 2021. Part 
of the combined purchase price mentioned above with Build Security Ltd.
4. **Endgame, Inc.**: Although the context mentions an amended 
and restated stock incentive plan, it indicates that Endgame, Inc. was 
associated with Elastic, suggesting it could have been acquired at some point.

These acquisitions are part of Elastic's strategy to 
enhance the technology underlying its security and 
observability offerings.
<p>SimpleRAG</p>The provided context does not specify any companies that Elastic has acquired. 
Therefore, based on the context, there is no information available about the companies acquired by Elastic.
<p><strong>摘要</strong>：SimpleRAG 无法检索到任何有关收购的相关信息，导致回答失败。AdvancedRAG 正确地列出了 CmdWatch、Build Security 和 Optimyze，它们是报告中列出的主要收购项目。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">返回顶部</a></p><h2>结论</h2><p>根据我们的测试，我们的先进技术似乎增加了所提供信息的范围和深度，有可能提高 RAG 答案的质量。</p><p>此外，可靠性也可能有所提高，因为 AdvancedRAG 可以正确回答<code>Which companies did Elastic acquire?</code> 和<code>Who audits Elastic</code> 等措辞含糊的问题，而 SimpleRAG 则不能。</p><p>不过，值得注意的是，在 5 个案例中的 3 个案例中，基本的 RAG 管道（包括混合搜索，但不包括其他技术）设法得出了能够捕捉到大部分关键信息的答案。</p><p>我们应该注意到，由于在数据准备和查询阶段加入了 LLM，AdvancedRAG 的延迟一般是 SimpleRAG 的 2-5 倍。这是一笔不小的费用，可能使 AdvancedRAG 只适用于优先考虑应答质量而不是延迟的情况。</p><p>在数据准备阶段，使用 Claude Haiku 或 GPT-4o-mini 等更小巧、更便宜的 LLM，就能减轻巨大的延迟成本。将高级模型留待生成答案时使用。</p><p>这与 Wang 等人的研究结果一致。结果表明，任何改进都是相对渐进的。简而言之，简单的基线 RAG 就能让您获得大部分体面的最终产品，而且成本更低，速度更快。对我来说，这是一个有趣的结论。对于速度和效率至关重要的使用案例，SimpleRAG 是明智的选择。对于需要榨取每一滴性能的使用案例，AdvancedRAG 中包含的技术可能会提供一条出路。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt56b7067a9d41d5a8/6a171119acf0886fb4be9c45/ea811706b6adc4731d90b925a9fefa0ac15901b4-1440x1060.jpg" alt="王家管道" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">返回顶部</a></p><h2>附录</h2><h3>提示</h3><h4>RAG 问题解答提示</h4><p>提示 LLM 根据查询和上下文生成答案。</p>BASIC_RAG_PROMPT = '''
You are an AI assistant tasked with answering questions based primarily on the provided context, while also drawing on your own knowledge when appropriate. Your role is to accurately and comprehensively respond to queries, prioritizing the information given in the context but supplementing it with your own understanding when beneficial. Follow these guidelines:

1. Carefully read and analyze the entire context provided.
2. Primarily focus on the information present in the context to formulate your answer.
3. If the context doesn't contain sufficient information to fully answer the query, state this clearly and then supplement with your own knowledge if possible.
4. Use your own knowledge to provide additional context, explanations, or examples that enhance the answer.
5. Clearly distinguish between information from the provided context and your own knowledge. Use phrases like "According to the context..." or "The provided information states..." for context-based information, and "Based on my knowledge..." or "Drawing from my understanding..." for your own knowledge.
6. Provide comprehensive answers that address the query specifically, balancing conciseness with thoroughness.
7. When using information from the context, cite or quote relevant parts using quotation marks.
8. Maintain objectivity and clearly identify any opinions or interpretations as such.
9. If the context contains conflicting information, acknowledge this and use your knowledge to provide clarity if possible.
10. Make reasonable inferences based on the context and your knowledge, but clearly identify these as inferences.
11. If asked about the source of information, distinguish between the provided context and your own knowledge base.
12. If the query is ambiguous, ask for clarification before attempting to answer.
13. Use your judgment to determine when additional information from your knowledge base would be helpful or necessary to provide a complete and accurate answer.

Remember, your goal is to provide accurate, context-based responses, supplemented by your own knowledge when it adds value to the answer. Always prioritize the provided context, but don't hesitate to enhance it with your broader understanding when appropriate. Clearly differentiate between the two sources of information in your response.

Context:
[The concatenated documents will be inserted here]

Query:
[The user's question will be inserted here]

Please provide your answer based on the above guidelines, the given context, and your own knowledge where appropriate, clearly distinguishing between the two:
'''
<h4>弹性查询生成器提示</h4><p>提示使用同义词丰富查询内容，并将其转换为 OR 格式。</p>ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<h4>潜在问题生成器提示</h4><p>提示生成潜在问题，丰富文件元数据。</p>RAG_QUESTION_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating questions for Retrieval-Augmented Generation (RAG) systems. Your task is to analyze a given document and create 10 diverse questions that would effectively test a RAG system's ability to retrieve and synthesize information from this document.

Guidelines:
1. Thoroughly analyze the entire document.
2. Generate exactly 10 questions that cover various aspects and levels of complexity within the document's content.
3. Create questions that specifically target:
   a. Key facts and information
   b. Main concepts and ideas
   c. Relationships between different parts of the content
   d. Potential applications or implications of the information
   e. Comparisons or contrasts within the document
4. Ensure questions require answers of varying lengths and complexity, from simple retrieval to more complex synthesis.
5. Include questions that might require combining information from different parts of the document.
6. Frame questions to test both literal comprehension and inferential understanding.
7. Avoid yes/no questions; focus on open-ended questions that promote comprehensive answers.
8. Consider including questions that might require additional context or knowledge to fully answer, to test the RAG system's ability to combine retrieved information with broader knowledge.
9. Number the questions from 1 to 10.
10. Output only the ten questions, without any additional text, explanations, or answers.

Document:
[The document content will be inserted here]

Generate 10 questions optimized for testing a RAG system based on this document:
'''
<h4>HyDE 生成器提示</h4><p>使用 HyDE 生成假设文档的提示</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<h3>混合搜索查询示例</h3>{'knn': {'field': 'primary_embedding',
  'query_vector': [0.4265527129173279,
   -0.1712949573993683,
   -0.042020395398139954,
   ...],
  'k': 100,
  'num_candidates': 100},
 'query': {'bool': {'must': [{'multi_match': {'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
      'fields': ['original_text',
       'keyphrases',
       'potential_questions',
       'entities'],
      'type': 'best_fields',
      'operator': 'or'}}],
   'should': [{'script_score': {'query': {'match_all': {}},
      'script': {'source': '\n                                        double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;\n                                        double text_score = _score;\n                                        return 0.7 * vector_score + 0.3 * text_score;\n                                        ',
       'params': {'query_vector': [0.4265527129173279,
         -0.1712949573993683,
         -0.042020395398139954,
        ...],
        'vector_field': 'primary_embedding'}}}}]}},
 'size': 10}
]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</guid>
    <category><![CDATA[向量数据库]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[高级 RAG 技术第 1 部分：数据处理]]></title>
    <description><![CDATA[讨论并实施可提高 RAG 性能的技术。第 1 部分（共 2 部分），重点介绍高级 RAG 管道的数据处理和摄取部分。]]></description>
    <content:encoded><![CDATA[<p><em>这是我们探索高级 RAG 技术的第一部分。 </em><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2"><em>点击此处查看第二部分！</em></a></p><p>最近发表的论文《<a href="https://arxiv.org/abs/2407.01219">在检索增强生成中寻找最佳实践</a>》对各种 RAG 增强技术的功效进行了实证评估，目的是为 RAG 找到一套最佳实践。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt671704ff06a4011d/6a170b3ea929cf2d19ae09d8/dafa7250e7c4ead4d9b4aed7c407509131929749-1440x572.png" alt="王建议的 RAG 管道" /><p>我们将实施其中一些建议的最佳实践，即旨在提高搜索质量的实践<strong>（句子分块、HyDE、反向打包）</strong>。</p><p>为简洁起见，我们将省略那些侧重于提高效率的技术<strong>（查询分类和摘要）</strong>。</p><p>我们还将实施一些未涉及但我个人认为有用且有趣的技术<strong>（元数据包含、复合多字段嵌入、查询丰富化）</strong>。</p><p>最后，我们将进行一个简短的测试，看看搜索结果和生成答案的质量与基线相比是否有所提高。让我们开始吧！</p><h2>RAG 概览</h2><p>RAG 的目的是通过检索外部知识库中的信息来丰富生成的答案，从而增强 LLM。通过提供特定领域的信息，LLM 可以快速适应训练数据范围之外的用例；比微调成本低得多，也更容易保持更新。</p><p>提高 RAG 质量的措施通常集中在两个方面：</p><ol><li><p>提高知识库的质量和清晰度。</p></li><li><p>提高搜索查询的覆盖面和针对性。</p></li></ol><p>这两项措施将实现提高法律硕士获得相关事实和信息的几率的目标，从而减少产生幻觉或利用自身知识的可能性--这些知识可能已经过时或不相关。</p><p>方法的多样性难以用几句话说清楚。为了更清楚地说明问题，让我们直接进入实施阶段。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="高级 RAG 管道" /><h3>目录</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#overview">概述</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">目录</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#set-up">设置</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#ingesting-processing-and-embedding-documents">摄取、处理和嵌入文件</a>  </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#data-ingestion">数据采集</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#sentence-level-token-wise-chunking">句子级标记分块</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#metadata-inclusion-and-generation">元数据的纳入和生成</a> </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#keyphrases-extracted-by-textrank">通过 TextRank 提取的关键词</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#potential-questions-generated-by-gpt-4o">GPT-4o 提出的潜在问题</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#entities-extracted-by-spacy">Spacy 提取的实体</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#composite-multi-field-embeddings">复合多场嵌入</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#indexing-to-elastic">索引至弹性</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#cat-break">猫休息</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#appendix">附录</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#definitions">定义</a></p></li></ul></li></ul><h2>设置</h2><p><em>所有代码均可 </em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em>在 Searchlabs 软件仓库中</em></a>找到 <em>。</em></p><p>先说第一件事。您需要以下材料</p><ol><li><p>弹性云部署</p></li><li><p>LLM 应用程序接口--我们在本笔记本中使用了 Azure OpenAI 上的 GPT-4o 部署</p></li><li><p>Python 3.12.4 或更高版本</p></li></ol><p>我们将运行<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/main.ipynb"> main.ipynb 笔记本 中的所有代码 。</a></p><p>继续 git 克隆该 repo，导航至 supporting-blog-content/Advanced-rag-techniques，然后运行以下命令：</p># Create a new virtual environment named 'rag_env'
python -m venv rag_env

# Activate the virtual environment (for Unix-based systems)
source rag_env/bin/activate

# (For Windows)
.\rag_env\Scripts\activate

# Install packages listed in requirements.txt
pip install -r requirements.txt
<p>完成后，创建一个<em>.env</em>文件，并填写以下字段（在<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/.env.example"><em>.env.example</em></a> 中引用）。感谢我的合著者 Claude-3.5 提出的有益意见。</p># Elastic Cloud: Found in the 'Deployment' page of your Elastic Cloud 
# console
ELASTIC_CLOUD_ENDPOINT=""
ELASTIC_CLOUD_ID=""

# Elastic Cloud: Created during deployment setup or in 'Security' 
# settings
ELASTIC_USERNAME=""
ELASTIC_PASSWORD=""

# Elastic Cloud: The name of the index you created in Kibana or via API
ELASTIC_INDEX_NAME=""

# Azure AI Studio: Found in 'Keys and Endpoint' section of your Azure 
# OpenAI resource
AZURE_OPENAI_KEY_1=""
AZURE_OPENAI_KEY_2=""
AZURE_OPENAI_REGION=""
AZURE_OPENAI_ENDPOINT=""

# Azure AI Studio: Found in 'Deployments' section of your Azure OpenAI 
# resource
AZURE_OPENAI_DEPLOYMENT_NAME=""

# Using BAAI/bge-small-en-v1.5 because I think it is a good balance of 
# resource efficiency and performance. 
HUGGINGFACE_EMBEDDING_MODEL="BAAI/bge-small-en-v1.5"
<p>接下来，我们将选择要摄取的文档，并将其放在文档文件夹中。在本文中，我们将使用<a href="https://s201.q4cdn.com/217177842/files/doc_downloads/OtherDocuments/2023/AnnualMeeting/Annual-Report-Fiscal-Year-2023.pdf"> Elastic N.V. 的《 2023 年年度报告》</a> 。这是一份相当具有挑战性的密集文件，非常适合对我们的 RAG 技术进行压力测试。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte292dc6030d496cc/6a170b40dc55de9b03e00dfc/e513b9d67adac43da794c25a5969b893127bbbe3-1440x395.jpg" alt="2023 年弹性年度报告" /><p>现在我们都准备好了，开始摄入。打开<em>main.ipynb</em>，执行前两个单元格以导入所有软件包并初始化所有服务。</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">返回顶部</a></p><h2>摄取、处理和嵌入文件</h2><h3>数据采集</h3><ul><li><p><em>个人感言：LlamaIndex 的便利性令我震惊。在还没有 LLM 和 LlamaIndex 的年代，录入各种格式的文档是一个痛苦的过程，需要从各处收集深奥的软件包。现在只需调用一个函数。狂野</em></p></li></ul><p><code>SimpleDirectoryReader</code> 将加载<code>directory_path.</code> 文件中的每个文档。对于<code>.pdf</code> 文件，它会返回一个文档对象列表，我将其转换为 Python 字典，因为我觉得它们更容易处理。</p># llamaindex_processor.py
from llama_index.core import SimpleDirectoryReader

class LlamaIndexProcessor:
   def __init__(self):
       pass 
   
   def load_documents(self, directory_path):
       ''' 
       Load all documents in directory
       '''
       reader = SimpleDirectoryReader(input_dir=directory_path)
       return reader.load_data()

# main.ipynb
llamaindex_processor=LlamaIndexProcessor()
documents=llamaindex_processor.load_documents('./documents/')
documents=[dict(doc_obj) for doc_obj in documents]
<p>每个字典都包含<code>text</code> 字段中的关键内容。它还包含有用的元数据，如页码、文件名、文件大小和类型。</p>{
  'id_': '5f76f0b3-22d8-49a8-9942-c2bbab14f63f',
  'metadata': {'page_label': '5',
   'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_path': '/Users/han/Desktop/Projects/truckasaurus/documents/Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_type': 'application/pdf',
   'file_size': 3724426,
   'creation_date': '2024-07-27',
   'last_modified_date': '2024-07-27'},
   'text': 'Table of Contents\nPage\nPART I\nItem 1. Business 3\n15 Item 1A. Risk Factors\nItem 1B. Unresolved Staff Comments 48\nItem 2. Properties 48\nItem 3. Legal Proceedings 48\nItem 4. Mine Safety Disclosures 48\nPART II\nItem 5. Market for Registrant's Common Equity, Related Stockholder Matters and Issuer Purchases of \nEquity Securities49\nItem 6. [Reserved] 49\nItem 7. Management's Discussion and Analysis of Financial Condition and Results of Operations 50\nItem 7A. Quantitative and Qualitative Disclosures About Market Risk 64\nItem 8. Financial Statements and Supplementary Data 66\nItem 9. Changes in and Disagreements With Accountants on Accounting and Financial Disclosure 100\n100\n101Item 9A. Controls and Procedures\nItem 9B. Other Information\nItem 9C. Disclosure Regarding Foreign Jurisdictions That Prevent Inspections 101\nPART III\n102\n102\n102\n102Item 10. Directors, Executive Officers and Corporate Governance\nItem 11. Executive Compensation\nItem 12. Security Ownership of Certain Beneficial Owners and Management, and Related Stockholder Matters  \nItem 13. Certain Relationships and Related Transactions, and Director Independence\nItem 14. Principal Accountant Fees and Services 102\nPART IV\n103\n105Item 15. Exhibits and Financial Statement Schedules  \nItem 16. Form 10-K Summary\nSignatures 106\ni',
   ...
}
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">返回顶部</a></p><h3>句子级标记分块</h3><p>首先要做的是将我们的文件缩减成标准长度的块状（以确保一致性和可管理性）。嵌入模型有独特的标记限制（可处理的最大输入尺寸）。标记是模型处理文本的基本单位。为防止信息丢失（内容截断或遗漏），我们应提供不超过这些限制的文本（将较长的文本分割成较小的片段）。</p><p>分块对性能有重大影响。在理想情况下，每个信息块都代表一个独立的信息片段，捕捉有关单个主题的上下文信息。分块方法包括字级分块（按字数分割文档）和语义分块（使用 LLM 识别逻辑断点）。</p><p>单词级的分块处理成本低、速度快、操作简单，但存在拆分句子从而破坏上下文的风险。语义分块的速度越来越慢，成本越来越高，尤其是在处理像116页的《弹性年度报告》这样的文档时。</p><p>让我们选择一种中间路线。句子级分块仍然简单，但比单词级分块能更有效地保留上下文，而且成本更低，速度更快。此外，我们还将采用一个滑动窗口来捕捉周围的一些上下文，并减轻分割段落的影响。</p># chunker.py 

import uuid
import re


class Chunker: 
    def __init__(self, tokenizer):
        self.tokenizer = tokenizer 
    
    def split_into_sentences(self, text):
        """Split text into sentences."""
        return re.split(r'(?&lt;=[.!?])\s+', text)
 
    def sentence_wise_tokenized_chunk_documents(self, documents, chunk_size=512, overlap=20, min_chunk_size=50):
        '''
        1. Split text into sentences.
        2. Tokenize using the provided tokenizer method.
        3. Build chunks up to the chunk_size limit.
        4. Create an overlap based on tokens - to preserve context.
        5. Only keep chunks that meet the minimum token size requirement.
        '''
        chunked_documents = []

        for doc in documents:
            sentences = self.split_into_sentences(doc['text'])
            tokens = []
            sentence_boundaries = [0]

            # Tokenize all sentences and keep track of sentence boundaries
            for sentence in sentences:
                sentence_tokens = self.tokenizer.encode(sentence, add_special_tokens=True)
                tokens.extend(sentence_tokens)
                sentence_boundaries.append(len(tokens))

            # Create chunks
            chunk_start = 0
            while chunk_start &lt; len(tokens):
                chunk_end = chunk_start + chunk_size

                # Find the last complete sentence that fits in the chunk
                sentence_end = next((i for i in sentence_boundaries if i &gt; chunk_end), len(tokens))
                chunk_end = min(chunk_end, sentence_end)

                # Create the chunk
                chunk_tokens = tokens[chunk_start:chunk_end]

                # Check if the chunk meets the minimum size requirement
                if len(chunk_tokens) &gt;= min_chunk_size:
                    # Create a new document object for this chunk
                    chunk_doc = {
                        'id_': str(uuid.uuid4()),
                        'chunk': chunk_tokens,
                        'original_text': self.tokenizer.decode(chunk_tokens),
                        'chunk_index': len(chunked_documents),
                        'parent_id': doc['id_'],
                        'chunk_token_count': len(chunk_tokens)
                    }

                    # Copy all other fields from the original document
                    for key, value in doc.items():
                        if key != 'text' and key not in chunk_doc:
                            chunk_doc[key] = value

                    chunked_documents.append(chunk_doc)

                # Move to the next chunk start, considering overlap
                chunk_start = max(chunk_start + chunk_size - overlap, chunk_end - overlap)

        return chunked_documents

# main.ipynb 
# Initialize Embedding Model
HUGGINGFACE_EMBEDDING_MODEL = os.environ.get('HUGGINGFACE_EMBEDDING_MODEL')
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

# Initialize Chunker
chunker=Chunker(embedder.tokenizer)
<p><code>Chunker</code> 类采用嵌入模型的标记化器对文本进行编码和解码。现在，我们将构建每块 512 个令牌的分块，其中有 20 个令牌重叠。为此，我们会将文本分割成句子，对这些句子进行标记化处理，然后将标记化处理后的句子添加到当前语块中，直到无法在不超出标记限制的情况下添加更多句子为止。</p><p>最后，将句子解码回原始文本进行嵌入，将其存储在名为<code>original_text</code> 的字段中。数据块存储在一个名为<code>chunk</code> 的字段中。为了减少噪音（又称无用文件），我们将丢弃长度小于 50 个 token 的文件。</p><p>让我们在文件上运行一下：</p>chunked_documents=chunker.sentence_wise_tokenized_chunk_documents(documents, chunk_size=512)
<p>然后得到类似这样的文本块：</p>print(chunked_documents[4]['original_text'])

[CLS] the aggregate market value of the ordinary shares held by non - affiliates of the registrant, 
based on the closing price of the shares of ordinary shares on the new york stock exchange on 
october 31, 2022 ( the last business day of the registrant 's second fiscal quarter ), was 
approximately $ 6. 1 billion. [SEP] [CLS] as of may 31, 2023, the registrant had 97, 390, 886 
ordinary shares, par value €0. 01 per share, outstanding. [SEP] [CLS] documents incorporated by 
reference portions of the registrant 's definitive proxy statement relating to the registrant 's 2
023 annual general meeting of shareholders are incorporated by reference into part iii of this annual 
...
...
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">返回顶部</a></p><h3>元数据的纳入和生成</h3><p>我们已将文件分块。现在是丰富数据的时候了。我想生成或提取额外的元数据。这些附加元数据可用于影响和提高搜索性能。</p><p>我们将定义一个<code>DocumentEnricher</code> 类，它的作用是接收文档列表（Python 字典）和处理器函数列表。这些函数将在文档的<code>original_text</code> 列中运行，并将其输出存储在新字段中。</p><p>首先，我们使用<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/nltk_processor.py">TextRank</a> 提取关键词。TextRank 是一种基于图的算法，它能根据词与词之间的关系对关键短语和句子的重要性进行排序，从而从文本中提取关键短语和句子。</p><p>接下来，我们将<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/llm.py">使用 GPT-4o 生成 potential_questions</a>。</p><p>最后，我们将使用<a href="https://spacy.io/"> Spacy</a> <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/entity_extractor.py">提取实体</a> 。</p><p>由于每项工作的代码都相当冗长和复杂，我就不在此赘述了。如果您感兴趣，这些文件已在下面的代码示例中标出。</p><p>让我们运行数据浓缩：</p># documentenricher.py
from tqdm import tqdm

class DocumentEnricher:

    def __init__(self):
        pass 

    def enrich_document(self, documents, processors, text_col='text'):
        for doc in tqdm(documents, desc="Enriching documents using processors: "+str(processors)): 
            for (processor, field) in processors: 
                metadata=processor(doc[text_col])
                if isinstance(metadata, list):
                    metadata='\n'.join(metadata)
                doc.update({field: metadata})
 
# main.ipynb
# Initialize processor classes 
nltkprocessor=NLTKProcessor() // nltk_processor.py
entity_extractor=EntityExtractor() // entity_extractor.py
gpt4o = LLMProcessor(model='gpt-4o') // llm.py

# Initialize LLM
documentenricher=DocumentEnricher()

# Create new fields in the documents - These are the outputs of the processor functions.
processors=[
    (nltkprocessor.textrank_phrases, "keyphrases"),
    (gpt4o.generate_questions, "potential_questions"),
    (entity_extractor.extract_entities, "entities")
    ]

# .enrich_document() will modify chunked_docs in place. 
# To view the results, we'll print chunked_docs in the next few cells!
documentenricher.enrich_document(chunked_docs, text_col='original_text', processors=processors)
<p>看看结果吧：</p><h4>通过 TextRank 提取的关键词</h4><p>这些关键短语是大块核心主题的替身。如果查询与网络安全有关，这块内容的得分就会提高。</p>print(chunked_documents[25]['keyphrases'])

'elastic agent stop', 'agent stop malware', 
'stop malware ransomware', 'malware ransomware environment', 
'ransomware environment wide', 'environment wide visibility', 
'wide visibility threat', 'visibility threat detection', 
'sep cl key', 'cl key feature'
<h4>GPT-4o 提出的潜在问题</h4><p>这些潜在问题可能与用户查询直接匹配，从而提高得分。我们会提示 GPT-4o 生成一些问题，这些问题可以用当前语块中的信息来回答。</p>print(chunked_documents[25]['potential_questions'])

1. What are the primary functions that Elastic Agent provides in terms of cybersecurity?
2. Describe how Logstash contributes to data management within an IT environment.
3. List and explain any key features of Logstash mentioned in the document.
4. How does Elastic Agent enhance environment-wide visibility in threat detection?
5. What capabilities does Logstash offer for handling data beyond simple collection?
6. In what ways does the document suggest that Elastic Agent stops malware and ransomware?
7. Can you identify any relationships between the functionalities of Elastic Agent and Logstash in an integrated environment?
8. What implications might the advanced threat detection capabilities of Elastic Agent have for organizational security policies?
9. Compare and contrast the roles of Elastic Agent and Logstash based on their described functions.
10. How might the centralized collection ability of Logstash support the threat detection capabilities of Elastic Agent?
<h4>Spacy 提取的实体</h4><p>这些实体的作用与关键词类似，但可以捕捉到组织和个人的名称，而关键词提取可能会遗漏这些名称。</p>print(chunked_documents[29]['entities'])

'appdynamics', 'apm data', 'azure sentinel', 
'microsoft', 'mcafee', 'broadcom', 'cisco', 
'dynatrace', 'coveo', 'lucidworks'
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">返回顶部</a></p><h3>复合多场嵌入</h3><p>现在，我们已经用更多的元数据丰富了我们的文档，我们可以利用这些信息创建更强大、更能感知上下文的嵌入。</p><p>让我们回顾一下目前的进程。我们在每份文档中都有四个关注领域。</p>{
    "chunk": "...",
    "keyphrases": "...", 
    "potential_questions": "...", 
    "entities": "..." 
}
<p>每个字段都代表了对文件背景的不同看法，可能突出了法律硕士应重点关注的关键领域。</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt84cb328fce6aae23/6a170b42964cea3e4408bbc4/aea1f513009a0c7c8545a79fad8f072a5bcae24c-1440x1067.jpg" alt="RAG 中的元数据丰富管道" /><p>我们的计划是嵌入每个字段，然后创建嵌入的加权和，即复合嵌入。</p><p>幸运的话，除了引入另一个可调整的超参数来控制搜索行为外，这种复合嵌入还能让系统变得更加了解上下文。</p><p>首先，让我们使用在 main.ipynb 笔记本开头导入的本地定义的嵌入模型，嵌入每个字段并就地更新每个文档。</p># EmbeddingModel defined in embedding_model.py
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

cols_to_embed=['keyphrases', 'potential_questions', 'entities']

embedding_cols=[]
for col in cols_to_embed:
    # Works on text input
    embedding_col=embedder.embed_documents_text_wise(chunked_documents, text_field=col)
    embedding_cols.append(embedding_col)
# Works on token input
embedding_col=embedder.embed_documents_token_wise(chunked_documents, token_field="chunk")
embedding_cols.append(embedding_col)
<p>每个嵌入函数都会返回嵌入的字段，即带有<code>_embedding</code> 后缀的原始输入字段。</p><p>现在我们来定义复合嵌入的权重：</p>embedding_cols=[
                'keyphrases_embedding',
                'potential_questions_embedding',
                'entities_embedding',
                'chunk_embedding']
combination_weights=[
                    0.1,
                    0.15,
                    0.05,
                    0.7
                ]
<p>通过权重，您可以根据用例和数据质量为每个组件分配优先级。直观地说，这些权重的大小取决于每个组件的语义值。由于大块文本本身的内容迄今为止最为丰富，我将其权重定为 70% 。由于实体最小，只是一个组织或个人名称列表，因此我将其权重定为 5% 。这些值的精确设置必须根据具体情况，根据经验来确定。</p><p>最后，让我们编写一个函数来应用权重，并创建我们的复合嵌入。为了节省空间，我们还将删除所有的组件嵌入。</p>from tqdm import tqdm 
def combine_embeddings(objects, embedding_cols, combination_weights, primary_embedding='primary_embedding'):
    # Ensure the number of weights matches the number of embedding columns
    assert len(embedding_cols) == len(combination_weights), "Number of embedding columns must match number of weights"
    
    # Normalize weights to sum to 1
    weights = np.array(combination_weights) / np.sum(combination_weights)
    
    for obj in tqdm(objects, desc="Combining embeddings"):
        # Initialize the combined embedding
        combined = np.zeros_like(obj[embedding_cols[0]])
        
        # Compute the weighted sum
        for col, weight in zip(embedding_cols, weights):
            combined += weight * np.array(obj[col])
        
        # Add the new combined embedding to the object
        obj.update({primary_embedding:combined.tolist()})
        
        # Remove the original embedding columns
        for col in embedding_cols:
            obj.pop(col, None)

combine_embeddings(chunked_documents, embedding_cols, combination_weights)
<p>至此，我们完成了文件处理工作。现在我们有了一个文档对象列表，看起来像这样：</p>{ 'id_': '7fe71686-5cd0-4831-9e79-998c6dbeae0c', 'chunk': [2312, 14613, ...], 'original_text': 'if an emerging growth company, indicate by check mark if the registrant has elected not to use the extended ...', 'chunk_index': 3, 'chunk_token_count': 399, 'metadata': {'page_label': '3', 'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf', ... 'keyphrases': 'sep cl unk\ncheck mark registrant\ncl unk indicate\nunk indicate check\nindicate check mark\nprincipal executive office\naccelerate filer unk\ncompany unk emerge\nunk emerge growth\nemerge growth company', 'potential_questions': '1. What are the different types of registrant statuses mentioned in the document?\n2. Under what section of the Sarbanes-Oxley Act must registrants file a report on the effectiveness of their internal ...', 'entities': 'the effe ctiveness of\nsection 13\nSEP\nUNK\nsection 21e\n1934\n1933\nu. s. c.\nsection 404\nsection 12\nal', 'primary_embedding': [-0.3946287803351879, -0.17586839850991964, ...] }
<h4>索引至弹性</h4><p>让我们将文档批量上传到 Elastic Search。为此，我很早就在<a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/elastic_helpers.py"><code>elastic_helpers.py</code></a> 中定义了一组 Elastic Helper 函数。这是一段非常冗长的代码，所以我们还是只看函数调用。</p><p><code>es_bulk_indexer.bulk_upload_documents</code> 利用 Elasticsearch 方便的动态映射，可以处理任何字典对象列表。</p># Initialize Elasticsearch
ELASTIC_CLOUD_ID = os.environ.get('ELASTIC_CLOUD_ID')
ELASTIC_USERNAME = os.environ.get('ELASTIC_USERNAME')
ELASTIC_PASSWORD = os.environ.get('ELASTIC_PASSWORD')
ELASTIC_CLOUD_AUTH = (ELASTIC_USERNAME, ELASTIC_PASSWORD)
es_bulk_indexer = ESBulkIndexer(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)
es_query_maker = ESQueryMaker(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)

# Define Index Name
index_name=os.environ.get('ELASTIC_INDEX_NAME')


# Create index and bulk upload 
index_exists = es_bulk_indexer.check_index_existence(index_name=index_name)
if not index_exists:
    logger.info(f"Creating new index: {index_name}")
    es_bulk_indexer.create_es_index(es_configuration=BASIC_CONFIG, index_name=index_name)

success_count = es_bulk_indexer.bulk_upload_documents(
    index_name=index_name, 
    documents=chunked_documents, 
    id_col='id_',
    batch_size=32
)
<p>前往 Kibana，确认所有文件都已编入索引。应该有 224 个。对于这么大的文件来说，还算不错！</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8efeface6effe01d/6a170b447d8d67652870e72a/1b3b07f6b98ceb65f6594ce4be83c5b0ed7e7cf9-1440x1380.jpg" alt="Kibana 索引" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">返回顶部</a></p><h2>猫休息</h2><p>我们休息一下吧，文章有点沉重，我知道。看看我的猫</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1db5595f71c12ff/6a170b450e2e49940241a0fe/baca4eb52b801b21ced97352cc55462f0a12d6b0-969x996.jpg" alt="汉族管道" /><p>真可爱帽子不见了，我半信半疑是她偷藏起来的：(</p><p>祝贺你们走到这一步 :)</p><p>请看<a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">第二部分</a>，了解我们对 RAG 管道的测试和评估！</p><h2>附录</h2><h3>定义</h3><p><strong>1.句子分块</strong></p><ul><li><p>RAG 系统中使用的一种预处理技术，用于将文本划分为更小的、有意义的单元。</p></li><li><p><em>过程：</em> </p><ol><li><p>输入：大段文本（如文档、段落）</p></li><li><p>输出：较小的文本片段（通常是句子或小句子组）</p></li></ol></li><li><p><em>目的是</em> </p><ul><li><p>创建细粒度、针对特定上下文的文本片段</p></li><li><p>允许更精确的索引和检索</p></li><li><p>提高 RAG 系统检索信息的相关性</p></li></ul></li><li><p><em>特点</em> </p><ul><li><p>分段具有语义意义</p></li><li><p>可独立索引和检索</p></li><li><p>通常保留一些上下文，以确保独立的可理解性</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>提高检索精度</p></li><li><p>使 RAG 管道的扩容更有针对性</p></li></ul></li></ul><p><strong>2.HyDE（假设文档嵌入）</strong></p><ul><li><p>在 RAG 系统中使用 LLM 生成用于查询扩展的假设文档的技术。</p></li><li><p><em>过程：</em>  </p><ol><li><p>向 LLM 输入查询</p></li><li><p>LLM 生成回答查询的假设文档</p></li><li><p>嵌入生成的文件</p></li><li><p>使用嵌入进行向量搜索</p></li></ol></li><li><p><em>主要区别</em> </p><ul><li><p>传统 RAG：将查询与文档匹配</p></li><li><p>HyDE：将文档匹配到文档</p></li></ul></li><li><p><em>目的是</em> </p><ul><li><p>提高检索性能，尤其是复杂或模糊查询的检索性能</p></li><li><p>捕捉比简短查询更丰富的语义上下文</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>利用 LLM 的知识扩展查询</p></li><li><p>有可能提高检索文件的相关性</p></li></ul></li><li><p><em>挑战：</em> </p><ul><li><p>需要额外的 LLM 推理，增加了延迟和成本</p></li><li><p>性能取决于生成的假设文件的质量</p></li></ul></li></ul><p><strong>3.反向包装</strong></p><ul><li><p>RAG 系统中使用的一种技术，用于在将搜索结果传递给 LLM 之前对其重新排序。</p></li><li><p><em>过程：</em> </p><ol><li><p>搜索引擎（如 Elasticsearch）按相关性降序返回文档。</p></li><li><p>顺序颠倒，将最相关的文件放在最后。</p></li></ol></li><li><p><em>目的是</em> </p><ul><li><p>利用 LLM 的新旧偏差，LLM 往往更关注其上下文中的最新信息。</p></li><li><p>确保最相关的信息"最新鲜的" 在 LLM 的上下文窗口中。</p></li></ul></li><li><p><em>举例说明：</em>原始顺序：[最相关、第二最相关、第三最相关、......] 倒序：[......，最重要的第三项，最重要的第二项，最相关的］</p></li></ul><p><strong>4.查询分类</strong></p><ul><li><p>通过确定查询是需要 RAG 还是可以直接由 LLM 回答来优化 RAG 系统效率的技术。</p></li><li><p><em>过程：</em> </p><ol><li><p>针对使用中的 LLM 开发定制数据集</p></li><li><p>训练专门的分类模型</p></li><li><p>使用模型对收到的查询进行分类</p></li></ol></li><li><p><em>目的是</em> </p><ul><li><p>避免不必要的 RAG 处理，提高系统效率</p></li><li><p>将查询引导至最合适的响应机制</p></li></ul></li><li><p><em>要求：</em> </p><ul><li><p>LLM 专用数据集和模型</p></li><li><p>不断改进以保持准确性</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>减少简单查询的计算开销</p></li><li><p>有可能缩短非 RAG 查询的响应时间</p></li></ul></li></ul><p><strong>5.总结</strong></p><ul><li><p>在 RAG 系统中压缩检索文档的技术。</p></li><li><p><em>过程：</em> </p><ol><li><p>检索相关文件</p></li><li><p>生成每份文件的简明摘要</p></li><li><p>在 RAG 管道中使用摘要而非完整文件</p></li></ol></li><li><p><em>目的是</em> </p><ul><li><p>关注基本信息，提高 RAG 性能</p></li><li><p>减少不相关内容的噪音和干扰</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>有可能提高 LLM 答复的相关性</p></li><li><p>允许在上下文限制内纳入更多文件</p></li></ul></li><li><p><em>挑战：</em> </p><ul><li><p>总结时有可能丢失重要细节</p></li><li><p>生成摘要的额外计算开销</p></li></ul></li></ul><p><strong>6.元数据的纳入</strong></p><ul><li><p>一种用额外的上下文信息来丰富文档的技术。</p></li><li><p><em>元数据类型：</em>  </p><ul><li><p>关键词</p></li><li><p>标题</p></li><li><p>日期</p></li><li><p>作者详细信息</p></li><li><p>简介</p></li></ul></li><li><p><em>目的是</em> </p><ul><li><p>增加 RAG 系统可用的背景信息</p></li><li><p>让法律硕士更清楚地了解文件内容和相关性</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>有可能提高检索的准确性</p></li><li><p>提高法律硕士评估文件实用性的能力</p></li></ul></li><li><p><em>实施：</em> </p><ul><li><p>可在文件预处理过程中完成</p></li><li><p>可能需要额外的数据提取或生成步骤</p></li></ul></li></ul><p><strong>7.复合多字段嵌入</strong></p><ul><li><p>RAG 系统的高级嵌入技术，可为不同的文档组件创建单独的嵌入。</p></li><li><p><em>过程：</em> </p><ol><li><p>确定相关字段（例如标题、关键词、简介、主要内容）</p></li><li><p>为每个字段生成单独的嵌入</p></li><li><p>合并或存储这些嵌入信息，以用于检索</p></li></ol></li><li><p><em>与标准方法的区别：</em> </p><ul><li><p>传统：对整个文档进行单一嵌入</p></li><li><p>复合：针对不同文档方面的多重嵌入</p></li></ul></li><li><p><em>目的是</em> </p><ul><li><p>创建更细致入微、更能感知上下文的文档表示法</p></li><li><p>在文件中获取更多来源的信息</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>有可能提高模糊或多方面查询的性能</p></li><li><p>允许在检索中更灵活地加权不同的文件内容</p></li></ul></li><li><p><em>挑战：</em> </p><ul><li><p>嵌入存储和检索流程的复杂性增加</p></li><li><p>可能需要更复杂的匹配算法</p></li></ul></li></ul><p><strong>8.丰富查询</strong></p><ul><li><p>一种用相关术语扩展原始查询以提高搜索覆盖率的技术。</p></li><li><p><em>过程：</em> </p><ol><li><p>分析原始查询</p></li><li><p>生成同义词和语义相关的短语</p></li><li><p>用这些附加术语来扩展查询</p></li></ol></li><li><p><em>目的是</em> </p><ul><li><p>增加文件语料库中潜在匹配的范围</p></li><li><p>提高使用特定或技术语言查询的检索性能</p></li></ul></li><li><p><em>优势：</em> </p><ul><li><p>可能检索到与原始查询条件不完全匹配的相关文档</p></li><li><p>有助于克服查询和文档之间的词汇不匹配问题</p></li></ul></li><li><p><em>挑战：</em> </p><ul><li><p>如果不认真执行，则有查询偏移的风险</p></li><li><p>可能会增加检索过程中的计算开销</p></li></ul></li></ul><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">返回顶部</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</guid>
    <category><![CDATA[向量数据库]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 14 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>