<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/kr/search-labs/author/han-xiang-choong</link>
    </image>
    <link>https://www.elastic.co/kr/search-labs/author/han-xiang-choong</link>
    <atom:link href="https://www.elastic.co/kr/search-labs/rss/author/han-xiang-choong.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[kr]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 12:39:10 GMT</lastBuildDate>
  <item>
    <title><![CDATA[고급 RAG 기술 2부: 쿼리 및 테스트]]></title>
    <description><![CDATA[RAG 성능을 향상시킬 수 있는 기술을 논의하고 구현합니다. 2부 2부에서는 고급 RAG 파이프라인 쿼리 및 테스트에 중점을 둡니다.]]></description>
    <content:encoded><![CDATA[<p><em>모든 코드는 </em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em>Searchlabs 리포지토리의 고급-걸레-기술 브랜치에서</em></a>찾을 수 있습니다<em>.</em></p><p>고급 RAG 기법에 대한 글 2부에 오신 것을 환영합니다! <a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">이 시리즈의 1부에서는</a> 고급 RAG 파이프라인의 데이터 처리 구성 요소를 설정하고, 논의하고, 구현하는 방법을 살펴봤습니다:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="고급 RAG 파이프라인" /><p>이 부분에서는 쿼리 및 구현 테스트를 진행하겠습니다. 바로 본론으로 들어가 보겠습니다!</p><h3>목차</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#searching-and-retrieving,-generating-answers">검색 및 검색, 답변 생성</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#enriching-queries-with-synonyms">동의어로 쿼리 강화하기</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-hypothetical-document-embedding">HyDE(가상의 문서 임베딩)</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hybrid-search">하이브리드 검색</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#experiments">실험</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#summary-of-results">결과 요약</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-1-who-audits-elastic">테스트 1: 누가 Elastic을 감사하나요?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-2--total-revenue-2023">테스트 2: 총 수익 2023년</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-1">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-1">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-3-what-product-does-growth-primarily-depend-on-how-much">테스트 3: 성장은 주로 어떤 제품에 의존하나요? 얼마예요?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-2">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-2">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-4-describe-employee-benefit-plan">테스트 4: 직원 복리후생 계획 설명</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-3">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-3">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-5-which-companies-did-elastic-acquire">테스트 5: Elastic은 어떤 회사를 인수했나요?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-4">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-4">SimpleRAG</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#conclusion">결론</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#appendix">부록</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#prompts">프롬프트</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#rag-question-answering-prompt">RAG 질문 답변 프롬프트</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#elastic-query-generator-prompt">Elastic 쿼리 생성기 프롬프트</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#potential-questions-generator-prompt">잠재적 질문 생성기 프롬프트</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-generator-prompt">HyDE 생성기 프롬프트</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#sample-hybrid-search-query">하이브리드 검색 쿼리 샘플</a></p></li></ul></li></ul><h2>검색 및 검색, 답변 생성</h2><p>첫 번째 질문은 주로 연례 보고서에서 찾을 수 있는 정보에 대해 물어보겠습니다. 어때요?</p>Who audits Elastic?"
<p>이제 몇 가지 기술을 적용하여 쿼리를 개선해 보겠습니다.</p><h3>동의어로 쿼리 강화하기</h3><p>먼저, 쿼리 문구의 다양성을 높이고 이를 Elasticsearch 쿼리로 쉽게 처리할 수 있는 형태로 바꿔보겠습니다. 쿼리를 OR 절의 목록으로 변환하기 위해 GPT-4o의 도움을 받겠습니다. 이 프롬프트를 작성해 보겠습니다:</p>
ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<p>쿼리에 적용하면 GPT-4o는 기본 쿼리의 동의어와 관련 어휘를 생성합니다.</p>'audits elastic OR 
elasticsearch audits OR 
elastic auditor OR 
elasticsearch auditor OR 
elastic audit firm OR 
elastic audit company OR 
elastic audit organization OR 
elastic audit service'
<p><code>ESQueryMaker</code> 클래스에서 쿼리를 분할하는 함수를 정의했습니다:</p>def parse_or_query(self, query_text: str) -&gt; List[str]:
    # Split the query by 'OR' and strip whitespace from each term
    # This converts a string like "term1 OR term2 OR term3" into a list ["term1", "term2", "term3"]
    return [term.strip() for term in query_text.split(' OR ')]
<p>이 OR 절의 문자열을 가져와 용어 목록으로 분할하여 주요 문서 필드에서 다중 일치를 수행할 수 있도록 하는 역할을 합니다:</p>["original_text", 'keyphrases', 'potential_questions', 'entities']
<p>마지막으로 이 쿼리로 마무리합니다:</p> 'query': {
    'bool': {
        'must': [
            {
                'multi_match': {
                'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
                'fields': [
                    'original_text',
                'keyphrases',
                'potential_questions',
                'entities'
                ],
                'type': 'best_fields',
                'operator': 'or'
                }
            }
      ]
<p>이렇게 하면 원래 쿼리보다 더 많은 기반을 포함하므로 동의어를 잊어버려 검색 결과가 누락되는 위험을 줄일 수 있습니다. 하지만 더 많은 일을 할 수 있습니다.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">맨 위로 돌아가기</a></p><h3>HyDE(가상의 문서 임베딩)</h3><p>이번에는 <a href="https://arxiv.org/abs/2212.10496">HyDE를</a> 구현하기 위해 GPT-4o를 다시 사용해 보겠습니다.</p><p>HyDE의 기본 전제는 가상의 문서, 즉 원래 쿼리에 대한 답변이 포함될 가능성이 있는 문서를 생성하는 것입니다. 문서의 사실 여부나 정확성은 문제가 되지 않습니다. 이를 염두에 두고 다음 프롬프트를 작성해 보겠습니다:</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<p>벡터 검색은 일반적으로 코사인 벡터 유사성을 기반으로 작동하므로, HyDE의 전제는 쿼리와 문서가 아닌 문서와 문서를 일치시킴으로써 더 나은 결과를 얻을 수 있다는 것입니다.</p><p>우리가 신경 쓰는 것은 구조, 흐름, 용어입니다. 사실과 다릅니다. GPT-4o는 다음과 같이 HyDE 문서를 출력합니다:</p>'Elastic N.V., the parent company of Elastic, the organization known for developing Elasticsearch, is subject to audits to ensure financial accuracy, 
regulatory compliance, and the integrity of its financial statements. The auditing of Elastic N.V. is typically conducted by an external, 
independent auditing firm. This is common practice for publicly traded companies to provide stakeholders with assurance regarding the company\'s 
financial position and operations.\n\nThe primary external auditor for Elastic is the audit firm Ernst &amp; Young LLP (EY). Ernst &amp; Young is one of the 
four largest professional services networks in the world, commonly referred to as the "Big Four" audit firms. These firms handle a substantial number 
of audits for major corporations around the globe, ensuring adherence to generally accepted accounting principles (GAAP) and international financial 
reporting standards (IFRS).\n\nThe audit process conducted by EY involves several steps. Initially, the auditors perform a risk assessment to identify 
areas where misstatements due to error or fraud could occur. They then design audit procedures to test the accuracy and completeness of financial statements,
 which include examining financial transactions, assessing internal controls, and reviewing compliance with relevant laws and regulations. Upon completion of 
 the audit, Ernst &amp; Young issues an audit report, which includes the auditor’s opinion on whether the financial statements are free from material misstatement 
 and are presented fairly in accordance with the applicable financial reporting framework.\n\nIn addition to external audits by firms like Ernst &amp; Young, 
 Elastic may also be subject to internal audits. Internal audits are performed by the company’s own internal auditors to evaluate the effectiveness of internal 
 controls, risk management, and governance processes.\n\nOverall, the auditing process plays a crucial role in maintaining the transparency and reliability of 
 Elastic\'s financial information, providing confidence to investors, regulators, and other stakeholders.'
<p>색인하려는 종류의 문서에 이상적인 후보처럼 꽤 그럴듯해 보입니다. 이를 임베드하여 하이브리드 검색에 사용하겠습니다.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">맨 위로 돌아가기</a></p><h3>하이브리드 검색</h3><p>이것이 바로 검색 로직의 핵심입니다. 어휘 검색 구성 요소는 생성된 OR 절 문자열이 됩니다. 고밀도 벡터 컴포넌트에는 HyDE 문서(일명 검색 벡터)가 내장됩니다. KNN을 사용하여 검색 벡터에 가장 가까운 여러 후보 문서를 효율적으로 식별합니다. 기본적으로 어휘 검색 컴포넌트를 <em>TF-IDF 및 BM25를 사용한 점수화라고</em> 부릅니다. 마지막으로, 어휘와 밀도 벡터 점수는 <a href="https://arxiv.org/abs/2407.01219">Wang 등이</a> 권장하는 30/70 비율을 사용하여 결합됩니다.</p>def hybrid_vector_search(self, index_name: str, query_text: str, query_vector: List[float], 
                         text_fields: List[str], vector_field: str, 
                         num_candidates: int = 100, num_results: int = 10) -&gt; Dict:
    """
    Perform a hybrid search combining text-based and vector-based similarity.

    Args:
        index_name (str): The name of the Elasticsearch index to search.
        query_text (str): The text query string, which may contain 'OR' separated terms.
        query_vector (List[float]): The query vector for semantic similarity search.
        text_fields (List[str]): List of text fields to search in the index.
        vector_field (str): The name of the field containing document vectors.
        num_candidates (int): Number of candidates to consider in the initial KNN search.
        num_results (int): Number of final results to return.

    Returns:
        Dict: A tuple containing the Elasticsearch response and the search body used.
    """
    try:
        # Parse the query_text into a list of individual search terms
        # This splits terms separated by 'OR' and removes any leading/trailing whitespace
        query_terms = self.parse_or_query(query_text)

        # Construct the search body for Elasticsearch
        search_body = {
            # KNN search component for vector similarity
            "knn": {
                "field": vector_field,  # The field containing document vectors
                "query_vector": query_vector,  # The query vector to compare against
                "k": num_candidates,  # Number of nearest neighbors to retrieve
                "num_candidates": num_candidates  # Number of candidates to consider in the KNN search
            },
            "query": {
                "bool": {
                    # The 'must' clause ensures that matching documents must satisfy this condition
                    # Documents that don't match this clause are excluded from the results
                    "must": [
                        {
                            # Multi-match query to search across multiple text fields
                            "multi_match": {
                                "query": " ".join(query_terms),  # Join all query terms into a single space-separated string
                                "fields": text_fields,  # List of fields to search in
                                "type": "best_fields",  # Use the best matching field for scoring
                                "operator": "or"  # Match any of the terms (equivalent to the original OR query)
                            }
                        }
                    ],
                    # The 'should' clause boosts relevance but doesn't exclude documents
                    # It's used here to combine vector similarity with text relevance
                    "should": [
                        {
                            # Custom scoring using a script to combine vector and text scores
                            "script_score": {
                                "query": {"match_all": {}},  # Apply this scoring to all documents that matched the 'must' clause
                                "script": {
                                    # Script to combine vector similarity and text relevance
                                    "source": """
                                    # Calculate vector similarity (cosine similarity + 1)
                                    # Adding 1 ensures the score is always positive
                                    double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;
                                    # Get the text-based relevance score from the multi_match query
                                    double text_score = _score;
                                    # Combine scores: 70% vector similarity, 30% text relevance
                                    # This weighting can be adjusted based on the importance of semantic vs keyword matching
                                    return 0.7 * vector_score + 0.3 * text_score;
                                    """,
                                    # Parameters passed to the script
                                    "params": {
                                        "query_vector": query_vector,  # Query vector for similarity calculation
                                        "vector_field": vector_field  # Field containing document vectors
                                    }
                                }
                            }
                        }
                    ]
                }
            }
        }

        # Execute the search request against the Elasticsearch index
        response = self.conn.search(index=index_name, body=search_body, size=num_results)
        # Log the successful execution of the search for monitoring and debugging
        logger.info(f"Hybrid search executed on index: {index_name} with text query: {query_text}")
        # Return both the response and the search body (useful for debugging and result analysis)
        return response, search_body
    except Exception as e:
        # Log any errors that occur during the search process
        logger.error(f"Error executing hybrid search on index: {index_name}. Error: {e}")
        # Re-raise the exception for further handling in the calling code
        raise e
<p>마지막으로 RAG 함수를 조합할 수 있습니다. 쿼리부터 답변까지 RAG는 이 흐름을 따릅니다:</p><ol><li><p>쿼리를 OR 절로 변환합니다.</p></li><li><p>HyDE 문서를 생성하고 임베드합니다.</p></li><li><p>둘 다 하이브리드 검색에 입력으로 전달합니다.</p></li><li><p>LLM의 컨텍스트 메모리(역 패킹) 역 패킹 예제에서 가장 관련성이 높은 점수가 "가장 최근의" 이 되도록 상위 n개의 결과를 검색하여 역 패킹합니다: 쿼리: "Elasticsearch 쿼리 최적화 기술" 검색된 문서(관련성 순으로 정렬):  LLM 컨텍스트에 대한 순서가 역순입니다:  순서를 반대로 하면 가장 관련성이 높은 정보(1)가 컨텍스트의 마지막에 표시되어 답변 생성 중에 LLM으로부터 더 많은 관심을 받을 수 있습니다.</p><ol><li><p>"부울 쿼리를 사용하여 여러 검색 기준을 효율적으로 결합할 수 있습니다."</p></li><li><p>"캐싱 전략을 구현하여 쿼리 응답 시간을 개선하세요."</p></li><li><p>"더 빠른 검색 성능을 위해 인덱스 매핑을 최적화하세요."</p></li><li><p>"더 빠른 검색 성능을 위해 인덱스 매핑을 최적화하세요."</p></li><li><p>"캐싱 전략을 구현하여 쿼리 응답 시간을 개선하세요."</p></li><li><p>"부울 쿼리를 사용하여 여러 검색 기준을 효율적으로 결합할 수 있습니다."</p></li></ol></li><li><p>생성할 컨텍스트를 LLM에 전달합니다.</p></li></ol>def get_context(index_name, 
                match_query, 
                text_query, 
                fields, 
                num_candidates=100, 
                num_results=20, 
                text_fields=["original_text", 'keyphrases', 'potential_questions', 'entities'], 
                embedding_field="primary_embedding"):

    embedding=embedder.get_embeddings_from_text(text_query)

    results, search_body = es_query_maker.hybrid_vector_search(
        index_name=index_name,
        query_text=match_query,
        query_vector=embedding[0][0],
        text_fields=text_fields,
        vector_field=embedding_field,
        num_candidates=num_candidates,
        num_results=num_results
    )

    # Concatenates the text in each 'field' key of the search result objects into a single block of text.
    context_docs=['\n\n'.join([field+":\n\n"+j['_source'][field] for field in fields]) for j in results['hits']['hits']]

    # Reverse Packing to ensure that the highest ranking document is seen first by the LLM.
    context_docs.reverse()
    return context_docs, search_body

def retrieval_augmented_generation(query_text):
    match_query= gpt4o.generate_query(query_text)
    fields=['original_text']

    hyde_document=gpt4o.generate_HyDE(query_text)

    context, search_body=get_context(index_name, match_query, hyde_document, fields)

    answer= gpt4o.basic_qa(query=query_text, context=context)
    return answer, match_query, hyde_document, context, search_body

<p>쿼리를 실행하여 답변을 확인해 보겠습니다:</p>According to the context, Elastic N.V. is audited by an independent registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of independent registered public accounting firm," which states:

"We have audited the accompanying consolidated balance sheets of Elastic N.V. [...] / s / pricewaterhouseco."
<p>멋지네요. 맞습니다.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">맨 위로 돌아가기</a></p><h2>실험</h2><p>지금 대답해야 할 중요한 질문이 있습니다. 이러한 구현에 많은 노력과 추가적인 복잡성을 투자하여 얻은 것은 무엇일까요?</p><p>간단한 비교를 해보겠습니다. 개선 사항 없이 기본 하이브리드 검색과 비교하여 구현한 RAG 파이프라인. 몇 가지 테스트를 실행하여 실질적인 차이를 발견할 수 있는지 확인해 보겠습니다. 방금 구현한 RAG를 AdvancedRAG라고 하고 기본 파이프라인을 SimpleRAG라고 하겠습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" alt="간단한 RAG 파이프라인" /><h4>결과 요약</h4><p>이 표에는 두 RAG 파이프라인에 대한 5가지 테스트 결과가 요약되어 있습니다. 답변의 세부 사항과 품질을 기준으로 각 방법의 상대적 우열을 판단했지만 이는 전적으로 주관적인 판단입니다. 이 표 아래에 실제 답변이 재현되어 있으므로 참고하시기 바랍니다. 그럼 이제 그 결과를 살펴보도록 하겠습니다!</p><p>SimpleRAG는 질문 1에 답할 수 없었습니다 &amp; 5. AdvancedRAG는 질문 2, 3, 4에 대해서도 훨씬 더 자세히 설명했습니다. 더 자세히 설명한 것을 바탕으로 저는 AdvancedRAG의 답변 품질이 더 좋다고 판단했습니다.</p><p>테스트</p><p>질문</p><p>고급RAG 성능</p><p>SimpleRAG 성능</p><p>AdvancedRAG 지연 시간</p><p>SimpleRAG 지연 시간</p><p>우승자</p><p>1</p><p>누가 Elastic을 감사하나요?</p><p>감사인으로 PwC를 올바르게 식별했습니다.</p><p>감사자를 식별하지 못했습니다.</p><p>11.6s</p><p>4.4s</p><p>AdvancedRAG</p><p>2</p><p>2023년 총 수익은 얼마였나요?</p><p>정확한 수익 수치를 제공했습니다. 전년도 수익에 대한 추가 컨텍스트가 포함되어 있습니다.</p><p>정확한 수익 수치를 제공했습니다.</p><p>13.3s</p><p>2.8s</p><p>AdvancedRAG</p><p>3</p><p>성장은 주로 어떤 제품에 의존하나요? 얼마예요?</p><p>핵심 동인으로 Elastic Cloud를 올바르게 식별했습니다. 전체 수익 컨텍스트 포함 &amp; 더 자세한 내용.</p><p>핵심 동인으로 Elastic Cloud를 올바르게 식별했습니다.</p><p>14.1s</p><p>12.8s</p><p>AdvancedRAG</p><p>4</p><p>직원 복리후생 계획 설명</p><p>퇴직 계획, 건강 프로그램 및 기타 혜택에 대한 포괄적인 설명을 제공합니다. 연도별 구체적인 기부 금액이 포함되어 있습니다.</p><p>보상, 퇴직 계획, 근무 환경, Elastic Cares 프로그램을 포함한 복리후생에 대한 좋은 개요를 제공했습니다.</p><p>26.6s</p><p>11.6s</p><p>AdvancedRAG</p><p>5</p><p>Elastic은 어떤 회사를 인수했나요?</p><p>보고서에 언급된 최근 인수를 올바르게 나열했습니다(CmdWatch, 빌드 시큐리티, 옵티마이즈). 일부 인수 날짜와 구매 가격을 제공했습니다.</p><p>제공된 컨텍스트에서 관련 정보를 검색하지 못했습니다.</p><p>11.9s</p><p>2.7s</p><p>AdvancedRAG</p><h4>테스트 1: 누가 Elastic을 감사하나요?</h4><p>AdvancedRAG</p>According to the context, Elastic N.V. is audited by an independent 
registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of 
independent registered public accounting firm," which states:
"We have audited the accompanying consolidated balance sheets 
of Elastic N.V. [...] / s / pricewaterhouseco."
<p>SimpleRAG</p>According to the provided context, no specific auditing firm 
or company is mentioned as the auditor for Elastic. 
The context provided talks about various features, strengths, 
deployment options, and other details related to Elastic and 
its operations but does not specify which auditing firm is 
responsible for auditing Elastic's financial statements or 
internal controls.
<p><strong>요약</strong>: SimpleRAG는 PWC를 감사인으로 지정하지 않았습니다.</p><p>사실 꽤 놀랍습니다. SimpleRAG 측의 검색 실패로 보입니다. 감사 관련 문서가 검색되지 않았습니다. 다음 테스트에서는 난이도를 조금 낮춰보겠습니다.</p><h4>테스트 2: 총 수익 2023년</h4><p>AdvancedRAG</p>### Total Revenue for 2023

According to the provided context, the total revenue for the 
year ended April 30, 2023, was **$1,068,989,000**. This 
information can be found several times in the provided text. 
Specifically, one part states:

"The following table sets forth our results of operations 
for the periods presented in dollars and as a percentage of 
our total revenue. ... year ended April 30, 2023, ... 
total revenue 1,068,989".

Additionally, it is mentioned here as well:

"For the years ended April 30, 2023, 2022, and 2021, 
representing year-over-year growth of 24% for the year 
ended April 30, 2023... total revenue $1,068,989".

Therefore, the total revenue for 2023 is $1.069 billion.
<p>SimpleRAG</p>According to the context provided, the total revenue for the fiscal year ended April 30, 2023, is:

$1,068,989,000 (or approximately $1.1 billion).
<p><strong>요약</strong>: 두 RAG 모두 2023년 총 수익 1,068,989,000달러라는 정답을 얻었습니다.</p><p>둘 다 바로 여기에 있었습니다. AdvancedRAG가 더 광범위한 문서를 확보한 것 같나요? 물론 답변은 더 상세하고 지난 몇 년간의 정보를 포함하고 있습니다. 개선 사항을 고려할 때 예상되는 부분이지만, 아직 확정하기에는 이르다고 할 수 있습니다.</p><p>난이도를 높여 보겠습니다.</p><h4>테스트 3: 성장은 주로 어떤 제품에 의존하나요? 얼마예요?</h4><p>AdvancedRAG</p>### Product Growth Primarily Depends On:
According to the context, "our revenue is derived primarily from 
subscriptions of search, observability and security built into 
the elastic stack." The document mentions various strategies and 
factors contributing to the company’s growth which include investment 
in "elastic cloud," "elastic stack," and the conversion of free 
users to paid subscribers.

### Quantitative Data:
- **Revenue**: The document states, "Our total revenue was $1.1 billion, 
$862.4 million, and $608.5 million for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Subscription Revenue**: "Subscriptions accounted for 92%, 93%, 
and 93% of our total revenue for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Elastic Cloud Contribution**: "For the years ended April 30, 2023, 
2022, and 2021, Elastic Cloud contributed 40%, 35%, and 27% of our 
total revenue, respectively."
- **Customer Growth**: The number of customers increased from 
"over 15,000 as of April 30, 2021," to "approximately 20,200 as of 
April 30, 2023."

### Summary:
Product growth for the company primarily depends on the "Elastic Stack" 
(which includes search, observability, and security solutions) and 
particularly the "Elastic Cloud." The Elastic Cloud's growing contribution 
to revenue highlights its significance, which grew from 27% in 2021 to 40% in 2023.
<p>SimpleRAG</p>According to the provided context, growth primarily depends on **Elastic Cloud**, 
which is described as the company's "family of cloud-based offerings." 
The context also emphasizes the significance of Elastic Cloud in achieving 
the company's long-term growth potential and its role in increasing subscription revenue.

From the context:

- "Increasing adoption of Elastic Cloud... is an important growth opportunity 
for our business. Organizations are increasingly looking for hosted 
deployment alternatives with reduced administrative burdens."
- "We believe that offering Elastic Cloud is important for achieving 
our long-term growth potential, and we expect Elastic Cloud’s contribution 
to our subscription revenue to continue to increase."

In terms of the contribution to total revenue, here are the specific figures provided:

- Elastic Cloud contributed 40% of the total revenue for the year ended April 30, 2023.
- Elastic Cloud contributed 35% of the total revenue for the year ended April 30, 2022.
- Elastic Cloud contributed 27% of the total revenue for the year ended April 30, 2021.

This progression shows a clear upward trend in the importance and 
contribution of Elastic Cloud to the company's overall revenue.
<p><strong>요약</strong>: 두 RAG 모두 Elastic Cloud를 핵심 성장 동력으로 정확하게 파악했습니다. 그러나 AdvancedRAG에는 구독 수익과 고객 성장을 고려한 더 자세한 내용이 포함되어 있으며, 다른 Elastic 제품에 대해서도 명시적으로 언급하고 있습니다.</p><h4>테스트 4: 직원 복리후생 계획 설명</h4><p>AdvancedRAG</p>### Employee Benefit Plan Overview

The provided context outlines various elements of the employee benefit plan at 
Elastic, often referred to as "Elasticians." Here are the key components 
described in the texts:

1. **401(k) Plan**:
   - **Coverage**: The 401(k) plan is available to substantially all U.S. 
   employees who meet minimum age and service requirements.
   - **Contributions**: Elastic makes contributions to the 401(k) plan up to 
   6% of the participating employee’s W-2 earnings and wages.
   - **Expenses**: For the fiscal years ended April 30, Elastic recorded 
   expenses of $17.9 million (2023), $15.2 million (2022), and $11.4 million (2021) 
   related to the 401(k) plan.
   - **Defined-Contribution Plans in Other Countries**: Elastic has 
   defined-contribution plans in various other countries and recorded respective 
   expenses of $9.4 million (2023), $7.2 million (2022), and $5.1 million (2021).

2. **Stock-Based Compensation**:
   - **Types of Awards**: Stock options, restricted stock units (RSUs), 
   and shares under the Employee Stock Purchase Plan (ESPP).
   - **Fair Value Measurement**: Fair value of these stock awards is 
   measured using models like Black-Scholes.
   - **Employee Stock Purchase Plan (2022 ESPP)**: 
     - Started in 2022, it allows employees to acquire ordinary 
     shares at a discount (85% of the market value at the beginning 
     or end of the offering period).
     - Offering periods are approximately six months long.

3. **Total Rewards Compensation**:
   - **Components**: Includes cash compensation as well as equity awards, 
   reflecting a comprehensive interest in the well-being of the employees.
   - **Programs**: Market competitive medical and dental programs, 
   mental health, and holistic well-being programs.
   - **Paid Time Off**: Features 16 weeks of paid leave for new 
   parents and life-planning benefits.

4. **Retirement and Income Protection Plans**:
   - **401(k) Match**: Elastic provides a dollar-for-dollar 
   match up to 6% of eligible earnings up to a plan-limit 
   maximum for U.S.-based employees.
   - **Similar Plans Outside the U.S.**: Competitive retirement 
   plans available for employees based outside the United States.

5. **Flexible Work Environment**: 
   - **Distributed Workforce**: Elastic operates with a global 
   distributed workforce model, allowing employees to work remotely 
   and providing a broad candidate pool.
   
6. **Learning and Development**:
   - **Programs**: Access to on-demand video-based learning, 
   leadership development programs, and other specific training 
   to develop managers and leaders.

7. **Employee Engagement and Surveys**:
   - **Mechanisms**: Annual employee engagement survey and 
   mid-year pulse survey.
   - **Objectives**: To monitor morale, develop action plans, 
   and maintain high engagement scores.

8. **Onboarding Program**:
   - **X-School**: An extensive new-hire orientation program to 
   help new employees meet peers and learn about the company's 
   products and solutions.

In summary, Elastic's employee benefit plan encompasses a broad 
spectrum of financial, health, well-being, and professional development 
benefits designed to support and enhance the lives of its employees. 
The plan is both competitive and comprehensive, aiming to attract,
retain, and engage its global workforce.
<p>SimpleRAG</p>The employee benefit plan at the Company includes several key 
components aimed at enhancing the overall well-being and 
professional development of employees, referred to as "Elasticians." 

1. **Compensation and Benefits:**
   - The Company provides market-competitive compensation 
   packages which typically include cash compensation and equity awards.
   - Comprehensive medical and dental programs are offered, 
   emphasizing mental health and holistic well-being.
   - Paid time off programs feature:
     - 16 weeks of paid leave for new parents.
     - Life-planning benefits.
     - Travel reimbursements for certain healthcare services.

2. **Retirement and Income Protection Plans:**
   - In the United States, employees are eligible for a 401k 
   plan with a dollar-for-dollar match by the Company up to 6% 
   of eligible earnings, subject to a plan limit.
   - Similar competitive retirement and income protection plans 
   are available for employees outside the United States.

3. **Work Environment:**
   - The Company offers a flexible work environment, allowing most 
   employees to work from anywhere as often as they would like.
   - Physical offices around the world are available for employees 
   who prefer to work on-site or need a space for face-to-face interactions.

4. **Community Involvement:**
   - The Elastic Cares Program supports charitable organizations 
   that matter to employees on both local and global levels.
   - The program includes:
     - Donation matching.
     - Providing the Company's technology for free to certain 
     nonprofit organizations.
     - Volunteer time off initiative.
<p><strong>요약</strong>: AdvancedRAG는 미국 거주 직원을 위한 401K 플랜을 언급하고 미국 외 지역의 기여금 플랜을 정의하는 등 훨씬 더 깊이 있고 자세하게 설명합니다. 또한 건강 및 웰빙 플랜에 대해서는 언급하고 있지만 SimpleRAG에서 언급하는 Elastic Cares 프로그램은 누락되어 있습니다.</p><h4>테스트 5: Elastic은 어떤 회사를 인수했나요?</h4><p>AdvancedRAG</p>Elastic acquired several companies over the years to enhance its 
technology and market presence. According to the provided context, 
the specific acquisitions include:

1. **CmdWatch Security Inc.**: Acquired on September 17, 2021. 
The total purchase consideration was $77.8 million.
2. **Build Security Ltd.**: Acquired on September 2, 2021. The 
combined purchase price for Build Security Ltd. and another acquisition was $57.2 million.
3. **Optimyze Cloud Inc.**: Acquired on November 1, 2021. Part 
of the combined purchase price mentioned above with Build Security Ltd.
4. **Endgame, Inc.**: Although the context mentions an amended 
and restated stock incentive plan, it indicates that Endgame, Inc. was 
associated with Elastic, suggesting it could have been acquired at some point.

These acquisitions are part of Elastic's strategy to 
enhance the technology underlying its security and 
observability offerings.
<p>SimpleRAG</p>The provided context does not specify any companies that Elastic has acquired. 
Therefore, based on the context, there is no information available about the companies acquired by Elastic.
<p><strong>요약</strong>: SimpleRAG가 인수에 대한 관련 정보를 검색하지 않아 답변이 실패했습니다. AdvancedRAG는 보고서에 나열된 주요 인수 기업인 CmdWatch, Build Security 및 Optimyze를 올바르게 나열합니다.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">맨 위로 돌아가기</a></p><h2>결론</h2><p>테스트 결과, 고급 기술은 제시되는 정보의 범위와 깊이를 증가시켜 잠재적으로 RAG 답변의 품질을 향상시키는 것으로 나타났습니다.</p><p>또한 <code>Which companies did Elastic acquire?</code>, <code>Who audits Elastic</code> 같은 모호한 단어로 된 질문은 AdvancedRAG에서는 정답을 맞혔지만 SimpleRAG에서는 그렇지 않았기 때문에 신뢰도가 개선될 수 있습니다.</p><p>그러나 5건 중 3건의 경우, 하이브리드 검색을 포함하지만 다른 기술은 사용하지 않는 기본 RAG 파이프라인이 대부분의 핵심 정보를 포착하는 답변을 생성했다는 점을 유념할 필요가 있습니다.</p><p>데이터 준비 및 쿼리 단계에 LLM이 통합되어 있기 때문에 AdvancedRAG의 지연 시간은 일반적으로 SimpleRAG보다 2~5배 더 길다는 점에 유의해야 합니다. 이는 상당한 비용이므로 지연 시간보다 응답 품질이 우선시되는 상황에서만 AdvancedRAG가 적합할 수 있습니다.</p><p>데이터 준비 단계에서 클로드 하이쿠나 GPT-4o-mini와 같은 더 작고 저렴한 LLM을 사용하면 상당한 지연 비용을 줄일 수 있습니다. 답변 생성을 위해 고급 모델을 저장합니다.</p><p>이는 Wang 등의 연구 결과와 일치합니다. 연구 결과에서 알 수 있듯이 개선 사항은 비교적 점진적으로 이루어졌습니다. 간단히 말해, 간단한 기본 RAG를 사용하면 저렴하고 빠르게 부팅할 수 있으면서도 괜찮은 최종 제품을 만들 수 있습니다. 저에게는 흥미로운 결론입니다. 속도와 효율성이 중요한 사용 사례의 경우 SimpleRAG가 현명한 선택입니다. 마지막 한 방울의 성능까지 끌어올려야 하는 사용 사례의 경우 AdvancedRAG에 통합된 기술이 한 가지 방법을 제시할 수 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt56b7067a9d41d5a8/6a171119acf0886fb4be9c45/ea811706b6adc4731d90b925a9fefa0ac15901b4-1440x1060.jpg" alt="왕 파이프라인" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">맨 위로 돌아가기</a></p><h2>부록</h2><h3>프롬프트</h3><h4>RAG 질문 답변 프롬프트</h4><p>쿼리 및 컨텍스트에 따라 LLM이 답변을 생성하도록 하기 위한 프롬프트입니다.</p>BASIC_RAG_PROMPT = '''
You are an AI assistant tasked with answering questions based primarily on the provided context, while also drawing on your own knowledge when appropriate. Your role is to accurately and comprehensively respond to queries, prioritizing the information given in the context but supplementing it with your own understanding when beneficial. Follow these guidelines:

1. Carefully read and analyze the entire context provided.
2. Primarily focus on the information present in the context to formulate your answer.
3. If the context doesn't contain sufficient information to fully answer the query, state this clearly and then supplement with your own knowledge if possible.
4. Use your own knowledge to provide additional context, explanations, or examples that enhance the answer.
5. Clearly distinguish between information from the provided context and your own knowledge. Use phrases like "According to the context..." or "The provided information states..." for context-based information, and "Based on my knowledge..." or "Drawing from my understanding..." for your own knowledge.
6. Provide comprehensive answers that address the query specifically, balancing conciseness with thoroughness.
7. When using information from the context, cite or quote relevant parts using quotation marks.
8. Maintain objectivity and clearly identify any opinions or interpretations as such.
9. If the context contains conflicting information, acknowledge this and use your knowledge to provide clarity if possible.
10. Make reasonable inferences based on the context and your knowledge, but clearly identify these as inferences.
11. If asked about the source of information, distinguish between the provided context and your own knowledge base.
12. If the query is ambiguous, ask for clarification before attempting to answer.
13. Use your judgment to determine when additional information from your knowledge base would be helpful or necessary to provide a complete and accurate answer.

Remember, your goal is to provide accurate, context-based responses, supplemented by your own knowledge when it adds value to the answer. Always prioritize the provided context, but don't hesitate to enhance it with your broader understanding when appropriate. Clearly differentiate between the two sources of information in your response.

Context:
[The concatenated documents will be inserted here]

Query:
[The user's question will be inserted here]

Please provide your answer based on the above guidelines, the given context, and your own knowledge where appropriate, clearly distinguishing between the two:
'''
<h4>Elastic 쿼리 생성기 프롬프트</h4><p>동의어로 쿼리를 보강하고 OR 형식으로 변환하는 프롬프트입니다.</p>ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<h4>잠재적 질문 생성기 프롬프트</h4><p>잠재적인 질문을 생성하고 문서 메타데이터를 보강하는 프롬프트입니다.</p>RAG_QUESTION_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating questions for Retrieval-Augmented Generation (RAG) systems. Your task is to analyze a given document and create 10 diverse questions that would effectively test a RAG system's ability to retrieve and synthesize information from this document.

Guidelines:
1. Thoroughly analyze the entire document.
2. Generate exactly 10 questions that cover various aspects and levels of complexity within the document's content.
3. Create questions that specifically target:
   a. Key facts and information
   b. Main concepts and ideas
   c. Relationships between different parts of the content
   d. Potential applications or implications of the information
   e. Comparisons or contrasts within the document
4. Ensure questions require answers of varying lengths and complexity, from simple retrieval to more complex synthesis.
5. Include questions that might require combining information from different parts of the document.
6. Frame questions to test both literal comprehension and inferential understanding.
7. Avoid yes/no questions; focus on open-ended questions that promote comprehensive answers.
8. Consider including questions that might require additional context or knowledge to fully answer, to test the RAG system's ability to combine retrieved information with broader knowledge.
9. Number the questions from 1 to 10.
10. Output only the ten questions, without any additional text, explanations, or answers.

Document:
[The document content will be inserted here]

Generate 10 questions optimized for testing a RAG system based on this document:
'''
<h4>HyDE 생성기 프롬프트</h4><p>HyDE를 사용하여 가상의 문서를 생성하는 프롬프트</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<h3>하이브리드 검색 쿼리 샘플</h3>{'knn': {'field': 'primary_embedding',
  'query_vector': [0.4265527129173279,
   -0.1712949573993683,
   -0.042020395398139954,
   ...],
  'k': 100,
  'num_candidates': 100},
 'query': {'bool': {'must': [{'multi_match': {'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
      'fields': ['original_text',
       'keyphrases',
       'potential_questions',
       'entities'],
      'type': 'best_fields',
      'operator': 'or'}}],
   'should': [{'script_score': {'query': {'match_all': {}},
      'script': {'source': '\n                                        double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;\n                                        double text_score = _score;\n                                        return 0.7 * vector_score + 0.3 * text_score;\n                                        ',
       'params': {'query_vector': [0.4265527129173279,
         -0.1712949573993683,
         -0.042020395398139954,
        ...],
        'vector_field': 'primary_embedding'}}}}]}},
 'size': 10}
]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[고급 RAG 기술 1부: 데이터 처리]]></title>
    <description><![CDATA[RAG 성능을 향상시킬 수 있는 기술을 논의하고 구현합니다. 2부 중 1부에서는 고급 RAG 파이프라인의 데이터 처리 및 수집 구성 요소에 초점을 맞춥니다.]]></description>
    <content:encoded><![CDATA[<p><em>고급 RAG 기법에 대한 탐험의 1부입니다. </em><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2"><em>2부를 보려면 여기를 클릭하세요!</em></a></p><p>최근 논문 ' <a href="https://arxiv.org/abs/2407.01219">검색 증강 세대의 모범 사례 찾기</a> '에서는 RAG의 모범 사례 집합으로 수렴하는 것을 목표로 다양한 RAG 향상 기술의 효과를 실증적으로 평가하고 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt671704ff06a4011d/6a170b3ea929cf2d19ae09d8/dafa7250e7c4ead4d9b4aed7c407509131929749-1440x572.png" alt="왕이 추천하는 RAG 파이프라인" /><p>검색 품질을 개선하기 위해 제안된 몇 가지 모범 사례, 즉 <strong>문장 청킹, HyDE, 리버스 패킹을</strong> 구현해 보겠습니다.</p><p>간결성을 위해 효율성 향상에 초점을 맞춘 기술 <strong>(쿼리 분류 및 요약)</strong>은 생략하겠습니다.</p><p>또한 다루지 않았지만 개인적으로 유용하고 흥미로웠던 몇 가지 기술 <strong>(메타데이터 포함, 복합 다중 필드 임베딩, 쿼리 강화)</strong>도 구현할 예정입니다.</p><p>마지막으로 간단한 테스트를 실행하여 검색 결과 및 생성된 답변의 품질이 기준선에 비해 개선되었는지 확인합니다. 시작해보자!</p><h2>RAG 개요</h2><p>RAG는 외부 지식 기반에서 정보를 검색하여 생성된 답변을 풍부하게 함으로써 LLM을 향상시키는 것을 목표로 합니다. 도메인별 정보를 제공함으로써 LLM은 학습 데이터의 범위를 벗어난 사용 사례에 빠르게 적응할 수 있으며, 미세 조정보다 훨씬 저렴하고 최신 상태로 유지하기 쉽습니다.</p><p>RAG의 품질을 개선하기 위한 조치는 일반적으로 두 가지 트랙에 중점을 둡니다:</p><ol><li><p>지식창고의 품질과 명확성을 향상시킵니다.</p></li><li><p>검색 쿼리의 범위와 구체성 개선.</p></li></ol><p>이 두 가지 조치를 통해 LLM이 관련 사실과 정보에 접근할 수 있는 확률을 높이고, 따라서 오래되었거나 관련성이 없는 자신의 지식에 의존하거나 착각할 가능성을 줄이려는 목표를 달성할 수 있습니다.</p><p>방법의 다양성은 몇 문장으로 명확히 설명하기 어렵습니다. 더 명확하게 설명하기 위해 바로 구현으로 넘어가겠습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="고급 RAG 파이프라인" /><h3>목차</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#overview">개요</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">목차</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#set-up">설정</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#ingesting-processing-and-embedding-documents">문서 수집, 처리 및 임베드하기</a>  </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#data-ingestion">데이터 수집</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#sentence-level-token-wise-chunking">문장 수준의 토큰 단위 청킹</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#metadata-inclusion-and-generation">메타데이터 포함 및 생성</a> </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#keyphrases-extracted-by-textrank">TextRank로 추출한 키프레이즈</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#potential-questions-generated-by-gpt-4o">GPT-4o에서 생성된 잠재적 질문</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#entities-extracted-by-spacy">Spacy에서 추출한 엔티티</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#composite-multi-field-embeddings">복합 멀티필드 임베딩</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#indexing-to-elastic">Elastic으로 인덱싱</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#cat-break">고양이 휴식</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#appendix">부록</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#definitions">정의</a></p></li></ul></li></ul><h2>설정</h2><p><em>모든 코드는 </em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em>Searchlabs 리포지토리에서</em></a>찾을 수 있습니다<em>.</em></p><p>먼저 해야 할 일이 있습니다. 다음이 필요합니다:</p><ol><li><p>Elastic Cloud 배포</p></li><li><p>LLM API - 이 노트북에서는 Azure OpenAI에서 GPT-4o 배포를 사용하고 있습니다.</p></li><li><p>Python 버전 3.12.4 이상</p></li></ol><p><a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/main.ipynb">main.ipynb 노트북에서</a>모든 코드를 실행할 것입니다.</p><p>계속해서 리포지토리를 git으로 복제하고 supporting-blog-content/advanced-rag-techniques로 이동한 다음 다음 명령을 실행합니다:</p># Create a new virtual environment named 'rag_env'
python -m venv rag_env

# Activate the virtual environment (for Unix-based systems)
source rag_env/bin/activate

# (For Windows)
.\rag_env\Scripts\activate

# Install packages listed in requirements.txt
pip install -r requirements.txt
<p>이 작업이 완료되면 <em>.env</em> 파일을 만듭니다. 파일을 열고 다음 필드를 채웁니다( <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/.env.example"><em>.env.example에서</em></a> 참조). 유용한 의견을 주신 공동 저자 Claude-3.5에게 감사의 마음을 전합니다.</p># Elastic Cloud: Found in the 'Deployment' page of your Elastic Cloud 
# console
ELASTIC_CLOUD_ENDPOINT=""
ELASTIC_CLOUD_ID=""

# Elastic Cloud: Created during deployment setup or in 'Security' 
# settings
ELASTIC_USERNAME=""
ELASTIC_PASSWORD=""

# Elastic Cloud: The name of the index you created in Kibana or via API
ELASTIC_INDEX_NAME=""

# Azure AI Studio: Found in 'Keys and Endpoint' section of your Azure 
# OpenAI resource
AZURE_OPENAI_KEY_1=""
AZURE_OPENAI_KEY_2=""
AZURE_OPENAI_REGION=""
AZURE_OPENAI_ENDPOINT=""

# Azure AI Studio: Found in 'Deployments' section of your Azure OpenAI 
# resource
AZURE_OPENAI_DEPLOYMENT_NAME=""

# Using BAAI/bge-small-en-v1.5 because I think it is a good balance of 
# resource efficiency and performance. 
HUGGINGFACE_EMBEDDING_MODEL="BAAI/bge-small-en-v1.5"
<p>다음으로 수집할 문서를 선택하고 문서 폴더에 넣습니다. 이 글에서는 <a href="https://s201.q4cdn.com/217177842/files/doc_downloads/OtherDocuments/2023/AnnualMeeting/Annual-Report-Fiscal-Year-2023.pdf">Elastic N.V. 연례 보고서 2023을</a> 사용하겠습니다. 꽤 도전적이고 밀도가 높은 문서로, RAG 기술을 스트레스 테스트하기에 적합합니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte292dc6030d496cc/6a170b40dc55de9b03e00dfc/e513b9d67adac43da794c25a5969b893127bbbe3-1440x395.jpg" alt="Elastic 연례 보고서 2023" /><p>이제 모든 준비가 완료되었으니 인제스트로 이동해 보겠습니다. <em>main.ipynb를</em> 열고 처음 두 셀을 실행하여 모든 패키지를 가져오고 모든 서비스를 초기화합니다.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">맨 위로 돌아가기</a></p><h2>문서 수집, 처리 및 임베드하기</h2><h3>데이터 수집</h3><ul><li><p><em>개인적 의견: 저는 라마인덱스의 편리함에 놀랐습니다. LLM과 라마인덱스가 나오기 전에는 다양한 형식의 문서를 수집하려면 여기저기서 난해한 패키지를 수집해야 하는 번거로운 과정이 필요했습니다. 이제 단일 함수 호출로 축소되었습니다. 야생.</em></p></li></ul><p><code>SimpleDirectoryReader</code> 은 <code>directory_path.</code> 파일에 있는 모든 문서를 로드합니다. <code>.pdf</code> 파일의 경우 문서 개체 목록을 반환하는데, 저는 작업하기 쉽다고 판단하여 Python 사전으로 변환합니다.</p># llamaindex_processor.py
from llama_index.core import SimpleDirectoryReader

class LlamaIndexProcessor:
   def __init__(self):
       pass 
   
   def load_documents(self, directory_path):
       ''' 
       Load all documents in directory
       '''
       reader = SimpleDirectoryReader(input_dir=directory_path)
       return reader.load_data()

# main.ipynb
llamaindex_processor=LlamaIndexProcessor()
documents=llamaindex_processor.load_documents('./documents/')
documents=[dict(doc_obj) for doc_obj in documents]
<p>각 사전에는 <code>text</code> 필드에 주요 콘텐츠가 포함되어 있습니다. 또한 페이지 번호, 파일 이름, 파일 크기 및 유형과 같은 유용한 메타데이터도 포함되어 있습니다.</p>{
  'id_': '5f76f0b3-22d8-49a8-9942-c2bbab14f63f',
  'metadata': {'page_label': '5',
   'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_path': '/Users/han/Desktop/Projects/truckasaurus/documents/Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_type': 'application/pdf',
   'file_size': 3724426,
   'creation_date': '2024-07-27',
   'last_modified_date': '2024-07-27'},
   'text': 'Table of Contents\nPage\nPART I\nItem 1. Business 3\n15 Item 1A. Risk Factors\nItem 1B. Unresolved Staff Comments 48\nItem 2. Properties 48\nItem 3. Legal Proceedings 48\nItem 4. Mine Safety Disclosures 48\nPART II\nItem 5. Market for Registrant's Common Equity, Related Stockholder Matters and Issuer Purchases of \nEquity Securities49\nItem 6. [Reserved] 49\nItem 7. Management's Discussion and Analysis of Financial Condition and Results of Operations 50\nItem 7A. Quantitative and Qualitative Disclosures About Market Risk 64\nItem 8. Financial Statements and Supplementary Data 66\nItem 9. Changes in and Disagreements With Accountants on Accounting and Financial Disclosure 100\n100\n101Item 9A. Controls and Procedures\nItem 9B. Other Information\nItem 9C. Disclosure Regarding Foreign Jurisdictions That Prevent Inspections 101\nPART III\n102\n102\n102\n102Item 10. Directors, Executive Officers and Corporate Governance\nItem 11. Executive Compensation\nItem 12. Security Ownership of Certain Beneficial Owners and Management, and Related Stockholder Matters  \nItem 13. Certain Relationships and Related Transactions, and Director Independence\nItem 14. Principal Accountant Fees and Services 102\nPART IV\n103\n105Item 15. Exhibits and Financial Statement Schedules  \nItem 16. Form 10-K Summary\nSignatures 106\ni',
   ...
}
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">맨 위로 돌아가기</a></p><h3>문장 수준의 토큰 단위 청킹</h3><p>가장 먼저 해야 할 일은 일관성과 관리 용이성을 위해 문서를 표준 길이의 덩어리로 줄이는 것입니다. 임베딩 모델에는 고유한 토큰 제한(처리할 수 있는 최대 입력 크기)이 있습니다. 토큰은 모델이 처리하는 텍스트의 기본 단위입니다. 정보 손실(내용 잘림 또는 누락)을 방지하기 위해 긴 텍스트를 더 작은 세그먼트로 분할하여 이러한 제한을 초과하지 않는 텍스트를 제공해야 합니다.</p><p>청크는 성능에 상당한 영향을 미칩니다. 이상적으로는 각 청크가 독립된 정보를 나타내며 단일 주제에 대한 맥락 정보를 캡처하는 것이 좋습니다. 청킹 방법에는 단어 수에 따라 문서를 분할하는 단어 수준 청킹과 LLM을 사용하여 논리적 중단점을 식별하는 시맨틱 청킹이 있습니다.</p><p>단어 단위 청크는 저렴하고 빠르며 쉽지만 문장이 분리되어 문맥이 깨질 위험이 있습니다. 시맨틱 청크는 특히 116페이지에 달하는 Elastic 연례 보고서와 같은 문서를 처리하는 경우 속도가 느려지고 비용이 많이 듭니다.</p><p>중간 접근 방식을 선택해 보겠습니다. 문장 수준 청킹은 여전히 간단하지만 단어 수준 청킹보다 훨씬 저렴하고 빠르면서 문맥을 더 효과적으로 보존할 수 있습니다. 또한 슬라이딩 창을 구현하여 주변 문맥을 일부 캡처하고 단락 분할의 영향을 완화할 것입니다.</p># chunker.py 

import uuid
import re


class Chunker: 
    def __init__(self, tokenizer):
        self.tokenizer = tokenizer 
    
    def split_into_sentences(self, text):
        """Split text into sentences."""
        return re.split(r'(?&lt;=[.!?])\s+', text)
 
    def sentence_wise_tokenized_chunk_documents(self, documents, chunk_size=512, overlap=20, min_chunk_size=50):
        '''
        1. Split text into sentences.
        2. Tokenize using the provided tokenizer method.
        3. Build chunks up to the chunk_size limit.
        4. Create an overlap based on tokens - to preserve context.
        5. Only keep chunks that meet the minimum token size requirement.
        '''
        chunked_documents = []

        for doc in documents:
            sentences = self.split_into_sentences(doc['text'])
            tokens = []
            sentence_boundaries = [0]

            # Tokenize all sentences and keep track of sentence boundaries
            for sentence in sentences:
                sentence_tokens = self.tokenizer.encode(sentence, add_special_tokens=True)
                tokens.extend(sentence_tokens)
                sentence_boundaries.append(len(tokens))

            # Create chunks
            chunk_start = 0
            while chunk_start &lt; len(tokens):
                chunk_end = chunk_start + chunk_size

                # Find the last complete sentence that fits in the chunk
                sentence_end = next((i for i in sentence_boundaries if i &gt; chunk_end), len(tokens))
                chunk_end = min(chunk_end, sentence_end)

                # Create the chunk
                chunk_tokens = tokens[chunk_start:chunk_end]

                # Check if the chunk meets the minimum size requirement
                if len(chunk_tokens) &gt;= min_chunk_size:
                    # Create a new document object for this chunk
                    chunk_doc = {
                        'id_': str(uuid.uuid4()),
                        'chunk': chunk_tokens,
                        'original_text': self.tokenizer.decode(chunk_tokens),
                        'chunk_index': len(chunked_documents),
                        'parent_id': doc['id_'],
                        'chunk_token_count': len(chunk_tokens)
                    }

                    # Copy all other fields from the original document
                    for key, value in doc.items():
                        if key != 'text' and key not in chunk_doc:
                            chunk_doc[key] = value

                    chunked_documents.append(chunk_doc)

                # Move to the next chunk start, considering overlap
                chunk_start = max(chunk_start + chunk_size - overlap, chunk_end - overlap)

        return chunked_documents

# main.ipynb 
# Initialize Embedding Model
HUGGINGFACE_EMBEDDING_MODEL = os.environ.get('HUGGINGFACE_EMBEDDING_MODEL')
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

# Initialize Chunker
chunker=Chunker(embedder.tokenizer)
<p><code>Chunker</code> 클래스는 임베딩 모델의 토큰라이저를 받아 텍스트를 인코딩하고 디코딩합니다. 이제 20개의 토큰이 겹치는 512개씩의 토큰 덩어리를 만들겠습니다. 이를 위해 텍스트를 문장으로 분할하고, 해당 문장을 토큰화한 다음 토큰 한도를 초과하지 않고 더 이상 추가할 수 없을 때까지 토큰화된 문장을 현재 청크에 추가합니다.</p><p>마지막으로 임베드할 원본 텍스트로 문장을 다시 디코딩하여 <code>original_text</code> 라는 필드에 저장합니다. 청크는 <code>chunk</code> 라는 필드에 저장됩니다. 노이즈(일명 쓸모없는 문서)를 줄이기 위해 길이가 50토큰보다 작은 문서는 모두 폐기합니다.</p><p>문서에서 실행해 보겠습니다:</p>chunked_documents=chunker.sentence_wise_tokenized_chunk_documents(documents, chunk_size=512)
<p>그리고 다음과 같은 텍스트 덩어리를 다시 가져옵니다:</p>print(chunked_documents[4]['original_text'])

[CLS] the aggregate market value of the ordinary shares held by non - affiliates of the registrant, 
based on the closing price of the shares of ordinary shares on the new york stock exchange on 
october 31, 2022 ( the last business day of the registrant 's second fiscal quarter ), was 
approximately $ 6. 1 billion. [SEP] [CLS] as of may 31, 2023, the registrant had 97, 390, 886 
ordinary shares, par value €0. 01 per share, outstanding. [SEP] [CLS] documents incorporated by 
reference portions of the registrant 's definitive proxy statement relating to the registrant 's 2
023 annual general meeting of shareholders are incorporated by reference into part iii of this annual 
...
...
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">맨 위로 돌아가기</a></p><h3>메타데이터 포함 및 생성</h3><p>문서를 정리했습니다. 이제 데이터를 보강할 차례입니다. 추가 메타데이터를 생성하거나 추출하고 싶습니다. 이 추가 메타데이터는 검색 성능에 영향을 미치고 향상시키는 데 사용할 수 있습니다.</p><p>문서 목록(파이썬 사전)과 프로세서 함수 목록을 받는 역할을 하는 <code>DocumentEnricher</code> 클래스를 정의하겠습니다. 이러한 함수는 문서의 <code>original_text</code> 열을 통해 실행되고 출력을 새 필드에 저장합니다.</p><p>먼저 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/nltk_processor.py">TextRank를</a> 사용하여 키프레이즈를 추출합니다. TextRank는 단어 간의 관계에 따라 중요도에 순위를 매겨 텍스트에서 핵심 구문과 문장을 추출하는 그래프 기반 알고리즘입니다.</p><p>다음으로 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/llm.py">GPT-4o를 사용하여 potential_questions를 생성하겠습니다</a>.</p><p>마지막으로 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/entity_extractor.py"></a> <a href="https://spacy.io/">Spacy를</a> 사용하여 엔티티를 추출합니다.</p><p>각 코드에 대한 설명은 상당히 길고 복잡하기 때문에 여기서는 반복하지 않겠습니다. 관심이 있으시다면 아래 코드 샘플에 파일이 표시되어 있습니다.</p><p>데이터 보강을 실행해 보겠습니다:</p># documentenricher.py
from tqdm import tqdm

class DocumentEnricher:

    def __init__(self):
        pass 

    def enrich_document(self, documents, processors, text_col='text'):
        for doc in tqdm(documents, desc="Enriching documents using processors: "+str(processors)): 
            for (processor, field) in processors: 
                metadata=processor(doc[text_col])
                if isinstance(metadata, list):
                    metadata='\n'.join(metadata)
                doc.update({field: metadata})
 
# main.ipynb
# Initialize processor classes 
nltkprocessor=NLTKProcessor() // nltk_processor.py
entity_extractor=EntityExtractor() // entity_extractor.py
gpt4o = LLMProcessor(model='gpt-4o') // llm.py

# Initialize LLM
documentenricher=DocumentEnricher()

# Create new fields in the documents - These are the outputs of the processor functions.
processors=[
    (nltkprocessor.textrank_phrases, "keyphrases"),
    (gpt4o.generate_questions, "potential_questions"),
    (entity_extractor.extract_entities, "entities")
    ]

# .enrich_document() will modify chunked_docs in place. 
# To view the results, we'll print chunked_docs in the next few cells!
documentenricher.enrich_document(chunked_docs, text_col='original_text', processors=processors)
<p>그리고 결과를 살펴보세요:</p><h4>TextRank로 추출한 키프레이즈</h4><p>이러한 키문구는 청크의 핵심 주제를 대신하는 역할을 합니다. 쿼리가 사이버 보안과 관련이 있는 경우 이 항목의 점수가 높아집니다.</p>print(chunked_documents[25]['keyphrases'])

'elastic agent stop', 'agent stop malware', 
'stop malware ransomware', 'malware ransomware environment', 
'ransomware environment wide', 'environment wide visibility', 
'wide visibility threat', 'visibility threat detection', 
'sep cl key', 'cl key feature'
<h4>GPT-4o에서 생성된 잠재적 질문</h4><p>이러한 잠재적 질문은 사용자 쿼리와 직접적으로 일치하여 점수를 높일 수 있습니다. 현재 청크에 있는 정보를 사용하여 답변할 수 있는 질문을 생성하라는 메시지를 GPT-4o에 표시합니다.</p>print(chunked_documents[25]['potential_questions'])

1. What are the primary functions that Elastic Agent provides in terms of cybersecurity?
2. Describe how Logstash contributes to data management within an IT environment.
3. List and explain any key features of Logstash mentioned in the document.
4. How does Elastic Agent enhance environment-wide visibility in threat detection?
5. What capabilities does Logstash offer for handling data beyond simple collection?
6. In what ways does the document suggest that Elastic Agent stops malware and ransomware?
7. Can you identify any relationships between the functionalities of Elastic Agent and Logstash in an integrated environment?
8. What implications might the advanced threat detection capabilities of Elastic Agent have for organizational security policies?
9. Compare and contrast the roles of Elastic Agent and Logstash based on their described functions.
10. How might the centralized collection ability of Logstash support the threat detection capabilities of Elastic Agent?
<h4>Spacy에서 추출한 엔티티</h4><p>이러한 엔티티는 키프레이즈와 유사한 용도로 사용되지만 키프레이즈 추출이 놓칠 수 있는 조직 및 개인 이름을 캡처합니다.</p>print(chunked_documents[29]['entities'])

'appdynamics', 'apm data', 'azure sentinel', 
'microsoft', 'mcafee', 'broadcom', 'cisco', 
'dynatrace', 'coveo', 'lucidworks'
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">맨 위로 돌아가기</a></p><h3>복합 멀티필드 임베딩</h3><p>이제 추가 메타데이터로 문서를 보강했으므로 이 정보를 활용하여 더욱 강력하고 컨텍스트를 인식하는 임베딩을 만들 수 있습니다.</p><p>프로세스의 현재 시점을 다시 한 번 살펴보겠습니다. 각 문서에는 네 가지 관심 분야가 있습니다.</p>{
    "chunk": "...",
    "keyphrases": "...", 
    "potential_questions": "...", 
    "entities": "..." 
}
<p>각 필드는 문서의 컨텍스트에 대한 다른 관점을 나타내며, 잠재적으로 LLM이 집중해야 할 핵심 영역을 강조할 수 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt84cb328fce6aae23/6a170b42964cea3e4408bbc4/aea1f513009a0c7c8545a79fad8f072a5bcae24c-1440x1067.jpg" alt="RAG의 메타데이터 강화 파이프라인" /><p>이러한 각 필드를 임베딩한 다음 복합 임베딩이라고 하는 임베딩의 가중치 합계를 생성하는 것이 계획입니다.</p><p>운이 좋으면 이 복합 임베딩을 통해 검색 동작을 제어하는 또 다른 조정 가능한 하이퍼파라미터를 도입하는 것 외에도 시스템이 컨텍스트를 더 잘 인식할 수 있게 됩니다.</p><p>먼저, main.ipynb 노트북의 시작 부분에서 가져온 로컬로 정의된 임베딩 모델을 사용해 각 필드를 임베드하고 각 문서를 제자리에 업데이트해 보겠습니다.</p># EmbeddingModel defined in embedding_model.py
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

cols_to_embed=['keyphrases', 'potential_questions', 'entities']

embedding_cols=[]
for col in cols_to_embed:
    # Works on text input
    embedding_col=embedder.embed_documents_text_wise(chunked_documents, text_field=col)
    embedding_cols.append(embedding_col)
# Works on token input
embedding_col=embedder.embed_documents_token_wise(chunked_documents, token_field="chunk")
embedding_cols.append(embedding_col)
<p>각 임베딩 함수는 <code>_embedding</code> 접두사가 붙은 원래 입력 필드인 임베딩의 필드를 반환합니다.</p><p>이제 컴포지트 임베딩의 가중치를 정의해 보겠습니다:</p>embedding_cols=[
                'keyphrases_embedding',
                'potential_questions_embedding',
                'entities_embedding',
                'chunk_embedding']
combination_weights=[
                    0.1,
                    0.15,
                    0.05,
                    0.7
                ]
<p>가중치를 사용하면 사용 사례와 데이터의 품질에 따라 각 구성 요소에 우선순위를 지정할 수 있습니다. 직관적으로 이러한 가중치의 크기는 각 구성 요소의 시맨틱 값에 따라 달라집니다. 청크 텍스트 자체가 가장 풍부하기 때문에 가중치를 70% 으로 지정합니다. 엔티티가 가장 작고 조직 또는 사람 이름의 목록에 불과하므로 가중치를 5% 로 할당합니다. 이러한 값에 대한 정확한 설정은 사용 사례별로 경험적으로 결정해야 합니다.</p><p>마지막으로 가중치를 적용하는 함수를 작성하고 복합 임베딩을 만들어 보겠습니다. 공간을 절약하기 위해 모든 컴포넌트 임베딩도 삭제할 것입니다.</p>from tqdm import tqdm 
def combine_embeddings(objects, embedding_cols, combination_weights, primary_embedding='primary_embedding'):
    # Ensure the number of weights matches the number of embedding columns
    assert len(embedding_cols) == len(combination_weights), "Number of embedding columns must match number of weights"
    
    # Normalize weights to sum to 1
    weights = np.array(combination_weights) / np.sum(combination_weights)
    
    for obj in tqdm(objects, desc="Combining embeddings"):
        # Initialize the combined embedding
        combined = np.zeros_like(obj[embedding_cols[0]])
        
        # Compute the weighted sum
        for col, weight in zip(embedding_cols, weights):
            combined += weight * np.array(obj[col])
        
        # Add the new combined embedding to the object
        obj.update({primary_embedding:combined.tolist()})
        
        # Remove the original embedding columns
        for col in embedding_cols:
            obj.pop(col, None)

combine_embeddings(chunked_documents, embedding_cols, combination_weights)
<p>이것으로 문서 처리를 완료했습니다. 이제 다음과 같은 문서 개체 목록이 생겼습니다:</p>{ 'id_': '7fe71686-5cd0-4831-9e79-998c6dbeae0c', 'chunk': [2312, 14613, ...], 'original_text': 'if an emerging growth company, indicate by check mark if the registrant has elected not to use the extended ...', 'chunk_index': 3, 'chunk_token_count': 399, 'metadata': {'page_label': '3', 'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf', ... 'keyphrases': 'sep cl unk\ncheck mark registrant\ncl unk indicate\nunk indicate check\nindicate check mark\nprincipal executive office\naccelerate filer unk\ncompany unk emerge\nunk emerge growth\nemerge growth company', 'potential_questions': '1. What are the different types of registrant statuses mentioned in the document?\n2. Under what section of the Sarbanes-Oxley Act must registrants file a report on the effectiveness of their internal ...', 'entities': 'the effe ctiveness of\nsection 13\nSEP\nUNK\nsection 21e\n1934\n1933\nu. s. c.\nsection 404\nsection 12\nal', 'primary_embedding': [-0.3946287803351879, -0.17586839850991964, ...] }
<h4>Elastic으로 인덱싱</h4><p>Elastic Search에 문서를 대량으로 업로드해 보겠습니다. 이를 위해 저는 오래 전에 <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/elastic_helpers.py"><code>elastic_helpers.py</code></a> 에서 Elastic Helper 함수 집합을 정의했습니다. 코드가 매우 길기 때문에 함수 호출을 살펴보는 데 집중하겠습니다.</p><p><code>es_bulk_indexer.bulk_upload_documents</code> 는 모든 사전 개체 목록에서 작동하며, Elasticsearch의 편리한 동적 매핑을 활용합니다.</p># Initialize Elasticsearch
ELASTIC_CLOUD_ID = os.environ.get('ELASTIC_CLOUD_ID')
ELASTIC_USERNAME = os.environ.get('ELASTIC_USERNAME')
ELASTIC_PASSWORD = os.environ.get('ELASTIC_PASSWORD')
ELASTIC_CLOUD_AUTH = (ELASTIC_USERNAME, ELASTIC_PASSWORD)
es_bulk_indexer = ESBulkIndexer(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)
es_query_maker = ESQueryMaker(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)

# Define Index Name
index_name=os.environ.get('ELASTIC_INDEX_NAME')


# Create index and bulk upload 
index_exists = es_bulk_indexer.check_index_existence(index_name=index_name)
if not index_exists:
    logger.info(f"Creating new index: {index_name}")
    es_bulk_indexer.create_es_index(es_configuration=BASIC_CONFIG, index_name=index_name)

success_count = es_bulk_indexer.bulk_upload_documents(
    index_name=index_name, 
    documents=chunked_documents, 
    id_col='id_',
    batch_size=32
)
<p>Kibana로 이동하여 모든 문서가 색인되었는지 확인합니다. 224개가 있어야 합니다. 이렇게 큰 문서치고는 나쁘지 않습니다!</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8efeface6effe01d/6a170b447d8d67652870e72a/1b3b07f6b98ceb65f6594ce4be83c5b0ed7e7cf9-1440x1380.jpg" alt="Kibana 색인" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">맨 위로 돌아가기</a></p><h2>고양이 휴식</h2><p>잠시만요, 기사가 좀 무겁네요, 알아요. 내 고양이를 확인하세요:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1db5595f71c12ff/6a170b450e2e49940241a0fe/baca4eb52b801b21ced97352cc55462f0a12d6b0-969x996.jpg" alt="한 파이프라인" /><p>사랑스러워요. 모자가 없어져서 어딘가에 훔쳐서 숨겨둔 것 같아요 :(</p><p>여기까지 오신 것을 축하드립니다 :)</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">2부에서는</a> RAG 파이프라인의 테스트 및 평가에 대해 알아보세요!</p><h2>부록</h2><h3>정의</h3><p><strong>1. 문장 청크</strong></p><ul><li><p>RAG 시스템에서 텍스트를 의미 있는 작은 단위로 나누기 위해 사용되는 전처리 기법입니다.</p></li><li><p><em>프로세스:</em> </p><ol><li><p>입력: 큰 텍스트 블록(예: 문서, 단락)</p></li><li><p>출력: 작은 텍스트 세그먼트(일반적으로 문장 또는 작은 문장 그룹)</p></li></ol></li><li><p><em>목적:</em> </p><ul><li><p>상황에 맞는 세분화된 텍스트 세그먼트 생성</p></li><li><p>보다 정확한 색인 및 검색 가능</p></li><li><p>RAG 시스템에서 검색된 정보의 관련성 향상</p></li></ul></li><li><p><em>특성:</em> </p><ul><li><p>세그먼트는 의미론적으로 의미가 있습니다.</p></li><li><p>독립적으로 색인 및 검색 가능</p></li><li><p>독립형 이해성을 보장하기 위해 일부 컨텍스트를 보존하는 경우가 많습니다.</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>검색 정밀도 향상</p></li><li><p>RAG 파이프라인에서 보다 집중적인 증강 지원</p></li></ul></li></ul><p><strong>2. HyDE(가상 문서 임베딩)</strong></p><ul><li><p>LLM을 사용하여 RAG 시스템에서 쿼리 확장을 위한 가상의 문서를 생성하는 기술입니다.</p></li><li><p><em>프로세스:</em>  </p><ol><li><p>LLM에 쿼리 입력</p></li><li><p>LLM은 쿼리에 대한 가상의 문서를 생성합니다.</p></li><li><p>생성된 문서 임베드</p></li><li><p>벡터 검색에 임베딩 사용</p></li></ol></li><li><p><em>주요 차이점:</em> </p><ul><li><p>기존 RAG: 쿼리를 문서와 일치시킵니다.</p></li><li><p>HyDE: 문서와 문서 매칭</p></li></ul></li><li><p><em>목적:</em> </p><ul><li><p>특히 복잡하거나 모호한 쿼리의 검색 성능 향상</p></li><li><p>짧은 쿼리보다 더 풍부한 의미론적 컨텍스트 캡처</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>LLM의 지식을 활용하여 쿼리를 확장합니다.</p></li><li><p>검색된 문서의 연관성을 잠재적으로 개선할 수 있습니다.</p></li></ul></li><li><p><em>도전 과제:</em> </p><ul><li><p>추가 LLM 추론이 필요하므로 지연 시간과 비용이 증가합니다.</p></li><li><p>성능은 생성된 가상 문서의 품질에 따라 달라집니다.</p></li></ul></li></ul><p><strong>3. 역포장</strong></p><ul><li><p>검색 결과를 LLM으로 전달하기 전에 순서를 바꾸기 위해 RAG 시스템에서 사용되는 기술입니다.</p></li><li><p><em>프로세스:</em> </p><ol><li><p>검색 엔진(예: Elasticsearch)은 관련성 내림차순으로 문서를 반환합니다.</p></li><li><p>순서가 뒤바뀌어 가장 관련성이 높은 문서가 마지막에 배치됩니다.</p></li></ol></li><li><p><em>목적:</em> </p><ul><li><p>해당 맥락에서 최신 정보에 더 집중하는 경향이 있는 LLM의 최신성 편향을 악용합니다.</p></li><li><p>LLM의 컨텍스트 창에서 가장 관련성 높은 정보가 "최신" 되도록 합니다.</p></li></ul></li><li><p><em>예시:</em> 원래 순서: [가장 관련성 높음, 두 번째 높음, 세 번째 높음, ...] 반전된 순서: [..., 세 번째로 많이, 두 번째로 많이, 가장 관련성 높음]</p></li></ul><p><strong>4. 쿼리 분류</strong></p><ul><li><p>쿼리에 RAG가 필요한지 아니면 LLM이 직접 답변할 수 있는지 판단하여 RAG 시스템 효율성을 최적화하는 기술입니다.</p></li><li><p><em>프로세스:</em> </p><ol><li><p>사용 중인 LLM에 맞는 사용자 지정 데이터 세트 개발</p></li><li><p>전문 분류 모델 훈련</p></li><li><p>모델을 사용하여 수신 쿼리 분류하기</p></li></ol></li><li><p><em>목적:</em> </p><ul><li><p>불필요한 RAG 처리를 방지하여 시스템 효율성 향상</p></li><li><p>가장 적절한 응답 메커니즘으로 직접 쿼리 보내기</p></li></ul></li><li><p><em>요구 사항:</em> </p><ul><li><p>LLM 전용 데이터 세트 및 모델</p></li><li><p>정확성 유지를 위한 지속적인 개선 사항</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>간단한 쿼리에 대한 계산 오버헤드 감소</p></li><li><p>RAG가 아닌 쿼리에 대한 응답 시간 개선 가능성</p></li></ul></li></ul><p><strong>5. 요약</strong></p><ul><li><p>RAG 시스템에서 검색된 문서를 압축하는 기술입니다.</p></li><li><p><em>프로세스:</em> </p><ol><li><p>관련 문서 검색</p></li><li><p>각 문서에 대한 간결한 요약 생성</p></li><li><p>RAG 파이프라인에서 전체 문서 대신 요약 사용</p></li></ol></li><li><p><em>목적:</em> </p><ul><li><p>필수 정보에 집중하여 RAG 성능 향상</p></li><li><p>관련성이 낮은 콘텐츠로 인한 노이즈 및 간섭 감소</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>잠재적으로 LLM 응답의 관련성 향상</p></li><li><p>컨텍스트 제한 내에서 더 많은 문서를 포함할 수 있습니다.</p></li></ul></li><li><p><em>도전 과제:</em> </p><ul><li><p>요약에서 중요한 세부 정보가 손실될 위험</p></li><li><p>요약 생성을 위한 추가 계산 오버헤드</p></li></ul></li></ul><p><strong>6. 메타데이터 포함</strong></p><ul><li><p>추가 컨텍스트 정보로 문서를 보강하는 기술입니다.</p></li><li><p><em>메타데이터의 유형:</em>  </p><ul><li><p>키문구</p></li><li><p>제목</p></li><li><p>날짜</p></li><li><p>저작자 세부 정보</p></li><li><p>블러브</p></li></ul></li><li><p><em>목적:</em> </p><ul><li><p>RAG 시스템에서 사용할 수 있는 컨텍스트 정보 증가</p></li><li><p>LLM에게 문서 콘텐츠와 관련성에 대한 보다 명확한 이해 제공</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>잠재적으로 검색 정확도 향상</p></li><li><p>문서 유용성을 평가하는 LLM의 능력 향상</p></li></ul></li><li><p><em>구현:</em> </p><ul><li><p>문서 전처리 중에 수행 가능</p></li><li><p>추가 데이터 추출 또는 생성 단계가 필요할 수 있습니다.</p></li></ul></li></ul><p><strong>7. 복합 멀티필드 임베딩</strong></p><ul><li><p>다양한 문서 구성 요소에 대해 별도의 임베딩을 생성하는 RAG 시스템용 고급 임베딩 기술입니다.</p></li><li><p><em>프로세스:</em> </p><ol><li><p>관련 필드 식별(예: 제목, 키문구, 광고 문구, 주요 콘텐츠)</p></li><li><p>각 필드에 대해 별도의 임베딩을 생성합니다.</p></li><li><p>검색에 사용할 수 있도록 이러한 임베딩을 결합하거나 저장하세요.</p></li></ol></li><li><p><em>표준 접근 방식과의 차이점:</em> </p><ul><li><p>기존: 전체 문서에 대한 단일 임베딩</p></li><li><p>합성: 다양한 문서 측면을 위한 다중 임베딩</p></li></ul></li><li><p><em>목적:</em> </p><ul><li><p>보다 미묘하고 컨텍스트를 인식하는 문서 표현 만들기</p></li><li><p>문서 내에서 더 다양한 소스의 정보를 캡처하세요.</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>모호하거나 다면적인 쿼리의 성능을 잠재적으로 개선합니다.</p></li><li><p>검색 시 다양한 문서 측면에 보다 유연한 가중치를 부여할 수 있습니다.</p></li></ul></li><li><p><em>도전 과제:</em> </p><ul><li><p>임베딩 스토리지 및 검색 프로세스의 복잡성 증가</p></li><li><p>보다 정교한 매칭 알고리즘이 필요할 수 있습니다.</p></li></ul></li></ul><p><strong>8. 쿼리 강화</strong></p><ul><li><p>검색 범위를 개선하기 위해 관련 용어로 원래 쿼리를 확장하는 기술입니다.</p></li><li><p><em>프로세스:</em> </p><ol><li><p>원본 쿼리 분석</p></li><li><p>동의어 및 의미적으로 연관된 구문 생성하기</p></li><li><p>다음 추가 용어를 사용하여 쿼리를 보강하세요.</p></li></ol></li><li><p><em>목적:</em> </p><ul><li><p>문서 말뭉치에서 잠재적인 일치 범위 늘리기</p></li><li><p>특정 언어 또는 기술 용어가 포함된 쿼리의 검색 성능 향상</p></li></ul></li><li><p><em>이점:</em> </p><ul><li><p>원래 쿼리 용어와 정확히 일치하지 않는 관련 문서를 검색할 수 있습니다.</p></li><li><p>쿼리와 문서 간의 어휘 불일치를 극복하는 데 도움이 될 수 있습니다.</p></li></ul></li><li><p><em>도전 과제:</em> </p><ul><li><p>신중하게 구현하지 않을 경우 쿼리 드리프트 위험</p></li><li><p>검색 프로세스에서 계산 오버헤드가 증가할 수 있습니다.</p></li></ul></li></ul><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">맨 위로 돌아가기</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 14 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>