<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Han Xiang Choong - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/es/search-labs/author/han-xiang-choong</link>
    </image>
    <link>https://www.elastic.co/es/search-labs/author/han-xiang-choong</link>
    <atom:link href="https://www.elastic.co/es/search-labs/rss/author/han-xiang-choong.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[es]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 10:21:29 GMT</lastBuildDate>
  <item>
    <title><![CDATA[Técnicas avanzadas de RAG parte 2: Consultas y pruebas]]></title>
    <description><![CDATA[Discutir e implementar técnicas que puedan aumentar el rendimiento de RAG. Parte 2 de 2, centrada en consultar y probar una pipeline avanzada de RAG.]]></description>
    <content:encoded><![CDATA[<p><em>Todo el código puede </em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em>encontrar en el repositorio Searchlabs, en la rama advanced-rag-techniques</em></a><em>.</em></p><p>¡Bienvenidos a la Parte 2 de nuestro artículo sobre Técnicas Avanzadas de RAG! En <a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1">la parte 1 de este serial</a>, establecimos, discutimos e implementamos los componentes de procesamiento de datos de la avanzada tubería RAG:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="Pipeline avanzado de RAG" /><p>En esta parte, vamos a proceder con la consulta y la prueba de nuestra implementación. ¡Vamos al grano!</p><h3>Índice</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#searching-and-retrieving,-generating-answers">Búsqueda y recuperación, generando respuestas</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#enriching-queries-with-synonyms">Enriqueciendo consultas con sinónimos</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-hypothetical-document-embedding">HyDE (Incrustación de Documentos Hipotéticos)</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hybrid-search">Búsqueda híbrida</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#experiments">Experimentos</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#summary-of-results">Resumen de resultados</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-1-who-audits-elastic">Prueba 1: ¿Quién audita a Elastic?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-2--total-revenue-2023">Test 2: ingresos totales 2023</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-1">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-1">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-3-what-product-does-growth-primarily-depend-on-how-much">Prueba 3: ¿De qué producto depende principalmente el crecimiento? ¿Cuánto?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-2">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-2">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-4-describe-employee-benefit-plan">Prueba 4: Describe el plan de beneficios para empleados</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-3">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-3">SimpleRAG</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#test-5-which-companies-did-elastic-acquire">Prueba 5: ¿Qué compañías adquirió Elastic?</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#advancedrag-4">AdvancedRAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#simplerag-4">SimpleRAG</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#conclusion">Conclusión</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#appendix">Apéndice</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#prompts">Prompts</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#rag-question-answering-prompt">Prompt de respuesta a preguntas RAG</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#elastic-query-generator-prompt">Prompt generador de consultas elástico</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#potential-questions-generator-prompt">Prompt generador de preguntas potenciales</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#hyde-generator-prompt">Prompt generador de HyDE</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#sample-hybrid-search-query">Consulta de búsqueda híbrida de ejemplo</a></p></li></ul></li></ul><h2>Búsqueda y recuperación, generando respuestas</h2><p>Hagamos nuestra primera pregunta, idealmente alguna información que se encuentre principalmente en el reporte anual. ¿Qué tal esto:</p>Who audits Elastic?"
<p>Ahora, apliquemos algunas de nuestras técnicas para mejorar la consulta.</p><h3>Enriqueciendo consultas con sinónimos</h3><p>Primero, mejoremos la diversidad de la redacción de la consulta y convirtiéramos en un formulario que pueda procesar fácilmente en una consulta de Elasticsearch. Aplicar la ayuda de GPT-4o para convertir la consulta en una lista de cláusulas OR. Vamos a escribir este prompt:</p>
ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<p>Cuando se aplica a nuestra consulta, GPT-4o genera sinónimos de la consulta base y el vocabulario relacionado.</p>'audits elastic OR 
elasticsearch audits OR 
elastic auditor OR 
elasticsearch auditor OR 
elastic audit firm OR 
elastic audit company OR 
elastic audit organization OR 
elastic audit service'
<p>En la clase <code>ESQueryMaker</code> , definí una función para dividir la consulta:</p>def parse_or_query(self, query_text: str) -&gt; List[str]:
    # Split the query by 'OR' and strip whitespace from each term
    # This converts a string like "term1 OR term2 OR term3" into a list ["term1", "term2", "term3"]
    return [term.strip() for term in query_text.split(' OR ')]
<p>Su función es tomar esta cadena de cláusulas OR y dividirlas en una lista de términos, permitiéndonos hacer una coincidencia múltiple en nuestros campos clave del documento:</p>["original_text", 'keyphrases', 'potential_questions', 'entities']
<p>Por fin terminando con esta pregunta:</p> 'query': {
    'bool': {
        'must': [
            {
                'multi_match': {
                'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
                'fields': [
                    'original_text',
                'keyphrases',
                'potential_questions',
                'entities'
                ],
                'type': 'best_fields',
                'operator': 'or'
                }
            }
      ]
<p>Esto cubre muchas más bases que la consulta original, con suerte reduciendo el riesgo de perder un resultado de búsqueda porque olvidamos un sinónimo. Pero podemos hacer más.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">Volver arriba</a></p><h3>HyDE (Incrustación de Documentos Hipotéticos)</h3><p>Vamos a reclutar de nuevo GPT-4o, esta vez para implementar <a href="https://arxiv.org/abs/2212.10496">HyDE</a>.</p><p>La premisa básica de HyDE es generar un documento hipotético: el tipo de documento que probablemente contenga la respuesta a la consulta original. La veracidad o exactitud del documento no es un problema. Con eso en mente, escribamos el siguiente prompt:</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<p>Dado que la búsqueda vectorial suele operar sobre similitud vectorial coseno, la premisa de HyDE es que podemos obtener mejores resultados emparejando documentos con documentos en lugar de consultas con documentos.</p><p>Lo que nos importa es la estructura, el flujo y la terminología. No tanto la factualidad. GPT-4o genera un documento HyDE así:</p>'Elastic N.V., the parent company of Elastic, the organization known for developing Elasticsearch, is subject to audits to ensure financial accuracy, 
regulatory compliance, and the integrity of its financial statements. The auditing of Elastic N.V. is typically conducted by an external, 
independent auditing firm. This is common practice for publicly traded companies to provide stakeholders with assurance regarding the company\'s 
financial position and operations.\n\nThe primary external auditor for Elastic is the audit firm Ernst &amp; Young LLP (EY). Ernst &amp; Young is one of the 
four largest professional services networks in the world, commonly referred to as the "Big Four" audit firms. These firms handle a substantial number 
of audits for major corporations around the globe, ensuring adherence to generally accepted accounting principles (GAAP) and international financial 
reporting standards (IFRS).\n\nThe audit process conducted by EY involves several steps. Initially, the auditors perform a risk assessment to identify 
areas where misstatements due to error or fraud could occur. They then design audit procedures to test the accuracy and completeness of financial statements,
 which include examining financial transactions, assessing internal controls, and reviewing compliance with relevant laws and regulations. Upon completion of 
 the audit, Ernst &amp; Young issues an audit report, which includes the auditor’s opinion on whether the financial statements are free from material misstatement 
 and are presented fairly in accordance with the applicable financial reporting framework.\n\nIn addition to external audits by firms like Ernst &amp; Young, 
 Elastic may also be subject to internal audits. Internal audits are performed by the company’s own internal auditors to evaluate the effectiveness of internal 
 controls, risk management, and governance processes.\n\nOverall, the auditing process plays a crucial role in maintaining the transparency and reliability of 
 Elastic\'s financial information, providing confidence to investors, regulators, and other stakeholders.'
<p>Parece bastante creíble, como el candidato ideal para los tipos de documentos que queremos indexar. Vamos a incrustar esto y usarlo para búsqueda híbrida.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">Volver arriba</a></p><h3>Búsqueda híbrida</h3><p>Esta es la base de nuestra lógica de búsqueda. Nuestro componente de búsqueda léxica serán las cadenas de cláusulas OR generadas. Nuestro componente vectorial denso será el documento HyDE embebido (también conocido como el vector de búsqueda). Empleamos KNN para identificar eficientemente varios documentos candidatos más cercanos a nuestro vector de búsqueda. Por defecto, llamamos a nuestro componente de búsqueda <em>léxico Puntaje con TF-IDF y BM25</em> . Finalmente, los puntajes léxicos y vectoriales densas se combinarán usando la proporción 30/70 recomendada por <a href="https://arxiv.org/abs/2407.01219">Wang et al</a>.</p>def hybrid_vector_search(self, index_name: str, query_text: str, query_vector: List[float], 
                         text_fields: List[str], vector_field: str, 
                         num_candidates: int = 100, num_results: int = 10) -&gt; Dict:
    """
    Perform a hybrid search combining text-based and vector-based similarity.

    Args:
        index_name (str): The name of the Elasticsearch index to search.
        query_text (str): The text query string, which may contain 'OR' separated terms.
        query_vector (List[float]): The query vector for semantic similarity search.
        text_fields (List[str]): List of text fields to search in the index.
        vector_field (str): The name of the field containing document vectors.
        num_candidates (int): Number of candidates to consider in the initial KNN search.
        num_results (int): Number of final results to return.

    Returns:
        Dict: A tuple containing the Elasticsearch response and the search body used.
    """
    try:
        # Parse the query_text into a list of individual search terms
        # This splits terms separated by 'OR' and removes any leading/trailing whitespace
        query_terms = self.parse_or_query(query_text)

        # Construct the search body for Elasticsearch
        search_body = {
            # KNN search component for vector similarity
            "knn": {
                "field": vector_field,  # The field containing document vectors
                "query_vector": query_vector,  # The query vector to compare against
                "k": num_candidates,  # Number of nearest neighbors to retrieve
                "num_candidates": num_candidates  # Number of candidates to consider in the KNN search
            },
            "query": {
                "bool": {
                    # The 'must' clause ensures that matching documents must satisfy this condition
                    # Documents that don't match this clause are excluded from the results
                    "must": [
                        {
                            # Multi-match query to search across multiple text fields
                            "multi_match": {
                                "query": " ".join(query_terms),  # Join all query terms into a single space-separated string
                                "fields": text_fields,  # List of fields to search in
                                "type": "best_fields",  # Use the best matching field for scoring
                                "operator": "or"  # Match any of the terms (equivalent to the original OR query)
                            }
                        }
                    ],
                    # The 'should' clause boosts relevance but doesn't exclude documents
                    # It's used here to combine vector similarity with text relevance
                    "should": [
                        {
                            # Custom scoring using a script to combine vector and text scores
                            "script_score": {
                                "query": {"match_all": {}},  # Apply this scoring to all documents that matched the 'must' clause
                                "script": {
                                    # Script to combine vector similarity and text relevance
                                    "source": """
                                    # Calculate vector similarity (cosine similarity + 1)
                                    # Adding 1 ensures the score is always positive
                                    double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;
                                    # Get the text-based relevance score from the multi_match query
                                    double text_score = _score;
                                    # Combine scores: 70% vector similarity, 30% text relevance
                                    # This weighting can be adjusted based on the importance of semantic vs keyword matching
                                    return 0.7 * vector_score + 0.3 * text_score;
                                    """,
                                    # Parameters passed to the script
                                    "params": {
                                        "query_vector": query_vector,  # Query vector for similarity calculation
                                        "vector_field": vector_field  # Field containing document vectors
                                    }
                                }
                            }
                        }
                    ]
                }
            }
        }

        # Execute the search request against the Elasticsearch index
        response = self.conn.search(index=index_name, body=search_body, size=num_results)
        # Log the successful execution of the search for monitoring and debugging
        logger.info(f"Hybrid search executed on index: {index_name} with text query: {query_text}")
        # Return both the response and the search body (useful for debugging and result analysis)
        return response, search_body
    except Exception as e:
        # Log any errors that occur during the search process
        logger.error(f"Error executing hybrid search on index: {index_name}. Error: {e}")
        # Re-raise the exception for further handling in the calling code
        raise e
<p>Finalmente, podemos reconstruir una función RAG. Nuestro RAG, desde la consulta hasta la respuesta, seguirá este flujo:</p><ol><li><p>Convertir la consulta en cláusulas OR.</p></li><li><p>Genera un documento HyDE e incrustalo.</p></li><li><p>Pásame ambos como entradas a la búsqueda híbrida.</p></li><li><p>Recuperar los resultados top-n, invertirlos para que el puntaje más relevante sea la "más reciente" en la memoria contextual del LLM (Empaquetado inverso) Ejemplo de empaquetado inverso: Consulta: "Técnicas de optimización de consultas Elasticsearch" Documentos recuperados (ordenados por relevancia): Orden invertido para el contexto del LLM: Al invertir el orden, la información más relevante (1) aparece al final en el contexto,  Potencialmente recibiendo más atención del LLM durante la generación de respuestas.</p><ol><li><p>"Emplea consultas bool para combinar múltiples criterios de búsqueda de forma eficiente."</p></li><li><p>"Implementar estrategias de caché para mejorar los tiempos de respuesta a las consultas."</p></li><li><p>"Optimizar los mapeos de índice para un rendimiento de búsqueda más rápido."</p></li><li><p>"Optimizar los mapeos de índice para un rendimiento de búsqueda más rápido."</p></li><li><p>"Implementar estrategias de caché para mejorar los tiempos de respuesta a las consultas."</p></li><li><p>"Emplea consultas bool para combinar múltiples criterios de búsqueda de forma eficiente."</p></li></ol></li><li><p>Pasa el contexto al LLM para que genere.</p></li></ol>def get_context(index_name, 
                match_query, 
                text_query, 
                fields, 
                num_candidates=100, 
                num_results=20, 
                text_fields=["original_text", 'keyphrases', 'potential_questions', 'entities'], 
                embedding_field="primary_embedding"):

    embedding=embedder.get_embeddings_from_text(text_query)

    results, search_body = es_query_maker.hybrid_vector_search(
        index_name=index_name,
        query_text=match_query,
        query_vector=embedding[0][0],
        text_fields=text_fields,
        vector_field=embedding_field,
        num_candidates=num_candidates,
        num_results=num_results
    )

    # Concatenates the text in each 'field' key of the search result objects into a single block of text.
    context_docs=['\n\n'.join([field+":\n\n"+j['_source'][field] for field in fields]) for j in results['hits']['hits']]

    # Reverse Packing to ensure that the highest ranking document is seen first by the LLM.
    context_docs.reverse()
    return context_docs, search_body

def retrieval_augmented_generation(query_text):
    match_query= gpt4o.generate_query(query_text)
    fields=['original_text']

    hyde_document=gpt4o.generate_HyDE(query_text)

    context, search_body=get_context(index_name, match_query, hyde_document, fields)

    answer= gpt4o.basic_qa(query=query_text, context=context)
    return answer, match_query, hyde_document, context, search_body

<p>Vamos a hacer nuestra consulta y obtener nuestra respuesta:</p>According to the context, Elastic N.V. is audited by an independent registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of independent registered public accounting firm," which states:

"We have audited the accompanying consolidated balance sheets of Elastic N.V. [...] / s / pricewaterhouseco."
<p>Muy bien. Así es.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">Volver arriba</a></p><h2>Experimentos</h2><p>Hay una pregunta importante que responder ahora. ¿Qué obtuvimos invirtiendo tanto esfuerzo y complejidad adicional en estas implementaciones?</p><p>Hagamos una pequeña comparación. La pipeline RAG que implementamos frente a la búsqueda híbrida base, sin ninguna de las mejoras que hicimos. Haremos un pequeño serial de pruebas para ver si notamos diferencias sustanciales. Nos referiremos al RAG que acabamos de implementar como AdvancedRAG, y al pipeline básico como SimpleRAG.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" alt="Simple RAG Pipeline" /><h4>Resumen de resultados</h4><p>Esta tabla resume los resultados de cinco pruebas de ambas tuberías RAG. Juzgué la superioridad relativa de cada método basándome en el detalle y la calidad de las respuestas, pero este es un juicio totalmente subjetivo. Las respuestas reales se reproducen a continuación de esta tabla para que lo consideres. Dicho esto, ¡echemos cómo les fue!</p><p>SimpleRAG no pudo responder a las preguntas 1 y 5. AdvancedRAG también profundizó mucho más en las preguntas 2, 3 y 4. Basándome en el mayor detalle, evalué mejor la calidad de las respuestas de AdvancedRAG.</p><p>Prueba</p><p>Pregunta</p><p>Rendimiento avanzado de RAG</p><p>Rendimiento de SimpleRAG</p><p>Latencia de AdvancedRAG</p><p>Latencia de SimpleRAG</p><p>Ganador</p><p>1</p><p>¿Quién audita Elastic?</p><p>Identificó correctamente a PwC como contralor.</p><p>No se identificó al contralor.</p><p>11,6</p><p>4,4</p><p>AdvancedRAG</p><p>2</p><p>¿Cuál fue el total de ingresos en 2023?</p><p>Proporcionó la cifra correcta de ingresos. Incluyó contexto adicional sobre ingresos de años anteriores.</p><p>Proporcionó la cifra correcta de ingresos.</p><p>13,3s</p><p>2,8</p><p>AdvancedRAG</p><p>3</p><p>¿De qué producto depende principalmente el crecimiento? ¿Cuánto?</p><p>Identificamos correctamente a Elastic Cloud como el motor clave. Incluyó el contexto general de ingresos y mayor detalle.</p><p>Identificamos correctamente a Elastic Cloud como el motor clave.</p><p>14.1</p><p>12,8</p><p>AdvancedRAG</p><p>4</p><p>Describe el plan de beneficios para empleados</p><p>Ofreció una descripción completa de los planes de jubilación, programas de salud y otros beneficios. Incluyó cantidades específicas de contribución para distintos años.</p><p>Ofreció una buena visión general de los beneficios, incluyendo compensación, planes de jubilación, entorno laboral y el programa Elastic Cares.</p><p>26,6</p><p>11,6</p><p>AdvancedRAG</p><p>5</p><p>¿Qué compañías adquirió Elastic?</p><p>Listé correctamente las adquisiciones recientes mencionadas en el reporte (CmdWatch, Build Security, Optimyze). Proporcioné algunas fechas de adquisición y precios de compra.</p><p>No se consiguió recuperar la información relevante del contexto proporcionado.</p><p>11,9</p><p>2,7</p><p>AdvancedRAG</p><h4>Prueba 1: ¿Quién audita a Elastic?</h4><p>AdvancedRAG</p>According to the context, Elastic N.V. is audited by an independent 
registered public accounting firm, PricewaterhouseCoopers (PwC). 
This information is found in the section titled "report of 
independent registered public accounting firm," which states:
"We have audited the accompanying consolidated balance sheets 
of Elastic N.V. [...] / s / pricewaterhouseco."
<p>SimpleRAG</p>According to the provided context, no specific auditing firm 
or company is mentioned as the auditor for Elastic. 
The context provided talks about various features, strengths, 
deployment options, and other details related to Elastic and 
its operations but does not specify which auditing firm is 
responsible for auditing Elastic's financial statements or 
internal controls.
<p><strong>Resumen</strong>: SimpleRAG no identificó a PWC como contralor</p><p>Vale, eso en realidad es bastante sorprendente. Eso parece un fallo de búsqueda por parte de SimpleRAG. No se recuperaron documentos relacionados con la auditoría. Bajemos un poco la dificultad con la siguiente prueba.</p><h4>Test 2: ingresos totales 2023</h4><p>AdvancedRAG</p>### Total Revenue for 2023

According to the provided context, the total revenue for the 
year ended April 30, 2023, was **$1,068,989,000**. This 
information can be found several times in the provided text. 
Specifically, one part states:

"The following table sets forth our results of operations 
for the periods presented in dollars and as a percentage of 
our total revenue. ... year ended April 30, 2023, ... 
total revenue 1,068,989".

Additionally, it is mentioned here as well:

"For the years ended April 30, 2023, 2022, and 2021, 
representing year-over-year growth of 24% for the year 
ended April 30, 2023... total revenue $1,068,989".

Therefore, the total revenue for 2023 is $1.069 billion.
<p>SimpleRAG</p>According to the context provided, the total revenue for the fiscal year ended April 30, 2023, is:

$1,068,989,000 (or approximately $1.1 billion).
<p><strong>Resumen</strong>: Ambos RAG obtuvieron la respuesta correcta: 1.068.989.000 dólares de ingresos totales en 2023</p><p>Ambos estaban justo aquí. ¿Parece que AdvancedRAG adquirió una gama más amplia de documentos? Sin duda, la respuesta es más detallada e incorpora información de años anteriores. Eso es de esperar dadas las mejoras que hicimos, pero es demasiado pronto para decidir.</p><p>Subamos la dificultad.</p><h4>Prueba 3: ¿De qué producto depende principalmente el crecimiento? ¿Cuánto?</h4><p>AdvancedRAG</p>### Product Growth Primarily Depends On:
According to the context, "our revenue is derived primarily from 
subscriptions of search, observability and security built into 
the elastic stack." The document mentions various strategies and 
factors contributing to the company’s growth which include investment 
in "elastic cloud," "elastic stack," and the conversion of free 
users to paid subscribers.

### Quantitative Data:
- **Revenue**: The document states, "Our total revenue was $1.1 billion, 
$862.4 million, and $608.5 million for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Subscription Revenue**: "Subscriptions accounted for 92%, 93%, 
and 93% of our total revenue for the years ended April 30, 2023, 
2022, and 2021, respectively."
- **Elastic Cloud Contribution**: "For the years ended April 30, 2023, 
2022, and 2021, Elastic Cloud contributed 40%, 35%, and 27% of our 
total revenue, respectively."
- **Customer Growth**: The number of customers increased from 
"over 15,000 as of April 30, 2021," to "approximately 20,200 as of 
April 30, 2023."

### Summary:
Product growth for the company primarily depends on the "Elastic Stack" 
(which includes search, observability, and security solutions) and 
particularly the "Elastic Cloud." The Elastic Cloud's growing contribution 
to revenue highlights its significance, which grew from 27% in 2021 to 40% in 2023.
<p>SimpleRAG</p>According to the provided context, growth primarily depends on **Elastic Cloud**, 
which is described as the company's "family of cloud-based offerings." 
The context also emphasizes the significance of Elastic Cloud in achieving 
the company's long-term growth potential and its role in increasing subscription revenue.

From the context:

- "Increasing adoption of Elastic Cloud... is an important growth opportunity 
for our business. Organizations are increasingly looking for hosted 
deployment alternatives with reduced administrative burdens."
- "We believe that offering Elastic Cloud is important for achieving 
our long-term growth potential, and we expect Elastic Cloud’s contribution 
to our subscription revenue to continue to increase."

In terms of the contribution to total revenue, here are the specific figures provided:

- Elastic Cloud contributed 40% of the total revenue for the year ended April 30, 2023.
- Elastic Cloud contributed 35% of the total revenue for the year ended April 30, 2022.
- Elastic Cloud contributed 27% of the total revenue for the year ended April 30, 2021.

This progression shows a clear upward trend in the importance and 
contribution of Elastic Cloud to the company's overall revenue.
<p><strong>Resumen</strong>: Ambos RAGs identificaron correctamente a Elastic Cloud como el principal motor de crecimiento. Sin embargo, AdvancedRAG incluye más detalles, teniendo en cuenta los ingresos por subscripción y el crecimiento de clientes, y menciona explícitamente otras ofertas de Elastic.</p><h4>Prueba 4: Describe el plan de beneficios para empleados</h4><p>AdvancedRAG</p>### Employee Benefit Plan Overview

The provided context outlines various elements of the employee benefit plan at 
Elastic, often referred to as "Elasticians." Here are the key components 
described in the texts:

1. **401(k) Plan**:
   - **Coverage**: The 401(k) plan is available to substantially all U.S. 
   employees who meet minimum age and service requirements.
   - **Contributions**: Elastic makes contributions to the 401(k) plan up to 
   6% of the participating employee’s W-2 earnings and wages.
   - **Expenses**: For the fiscal years ended April 30, Elastic recorded 
   expenses of $17.9 million (2023), $15.2 million (2022), and $11.4 million (2021) 
   related to the 401(k) plan.
   - **Defined-Contribution Plans in Other Countries**: Elastic has 
   defined-contribution plans in various other countries and recorded respective 
   expenses of $9.4 million (2023), $7.2 million (2022), and $5.1 million (2021).

2. **Stock-Based Compensation**:
   - **Types of Awards**: Stock options, restricted stock units (RSUs), 
   and shares under the Employee Stock Purchase Plan (ESPP).
   - **Fair Value Measurement**: Fair value of these stock awards is 
   measured using models like Black-Scholes.
   - **Employee Stock Purchase Plan (2022 ESPP)**: 
     - Started in 2022, it allows employees to acquire ordinary 
     shares at a discount (85% of the market value at the beginning 
     or end of the offering period).
     - Offering periods are approximately six months long.

3. **Total Rewards Compensation**:
   - **Components**: Includes cash compensation as well as equity awards, 
   reflecting a comprehensive interest in the well-being of the employees.
   - **Programs**: Market competitive medical and dental programs, 
   mental health, and holistic well-being programs.
   - **Paid Time Off**: Features 16 weeks of paid leave for new 
   parents and life-planning benefits.

4. **Retirement and Income Protection Plans**:
   - **401(k) Match**: Elastic provides a dollar-for-dollar 
   match up to 6% of eligible earnings up to a plan-limit 
   maximum for U.S.-based employees.
   - **Similar Plans Outside the U.S.**: Competitive retirement 
   plans available for employees based outside the United States.

5. **Flexible Work Environment**: 
   - **Distributed Workforce**: Elastic operates with a global 
   distributed workforce model, allowing employees to work remotely 
   and providing a broad candidate pool.
   
6. **Learning and Development**:
   - **Programs**: Access to on-demand video-based learning, 
   leadership development programs, and other specific training 
   to develop managers and leaders.

7. **Employee Engagement and Surveys**:
   - **Mechanisms**: Annual employee engagement survey and 
   mid-year pulse survey.
   - **Objectives**: To monitor morale, develop action plans, 
   and maintain high engagement scores.

8. **Onboarding Program**:
   - **X-School**: An extensive new-hire orientation program to 
   help new employees meet peers and learn about the company's 
   products and solutions.

In summary, Elastic's employee benefit plan encompasses a broad 
spectrum of financial, health, well-being, and professional development 
benefits designed to support and enhance the lives of its employees. 
The plan is both competitive and comprehensive, aiming to attract,
retain, and engage its global workforce.
<p>SimpleRAG</p>The employee benefit plan at the Company includes several key 
components aimed at enhancing the overall well-being and 
professional development of employees, referred to as "Elasticians." 

1. **Compensation and Benefits:**
   - The Company provides market-competitive compensation 
   packages which typically include cash compensation and equity awards.
   - Comprehensive medical and dental programs are offered, 
   emphasizing mental health and holistic well-being.
   - Paid time off programs feature:
     - 16 weeks of paid leave for new parents.
     - Life-planning benefits.
     - Travel reimbursements for certain healthcare services.

2. **Retirement and Income Protection Plans:**
   - In the United States, employees are eligible for a 401k 
   plan with a dollar-for-dollar match by the Company up to 6% 
   of eligible earnings, subject to a plan limit.
   - Similar competitive retirement and income protection plans 
   are available for employees outside the United States.

3. **Work Environment:**
   - The Company offers a flexible work environment, allowing most 
   employees to work from anywhere as often as they would like.
   - Physical offices around the world are available for employees 
   who prefer to work on-site or need a space for face-to-face interactions.

4. **Community Involvement:**
   - The Elastic Cares Program supports charitable organizations 
   that matter to employees on both local and global levels.
   - The program includes:
     - Donation matching.
     - Providing the Company's technology for free to certain 
     nonprofit organizations.
     - Volunteer time off initiative.
<p><strong>Resumen</strong>: AdvancedRAG entra en mucho más detalle y profundidad, mencionando el plan 401K para empleados con base en EE. UU., así como definiendo planes de contribución fuera de EE. UU. También menciona planes de salud y bienestar, pero no incluye el programa Elastic Cares, que menciona SimpleRAG.</p><h4>Prueba 5: ¿Qué compañías adquirió Elastic?</h4><p>AdvancedRAG</p>Elastic acquired several companies over the years to enhance its 
technology and market presence. According to the provided context, 
the specific acquisitions include:

1. **CmdWatch Security Inc.**: Acquired on September 17, 2021. 
The total purchase consideration was $77.8 million.
2. **Build Security Ltd.**: Acquired on September 2, 2021. The 
combined purchase price for Build Security Ltd. and another acquisition was $57.2 million.
3. **Optimyze Cloud Inc.**: Acquired on November 1, 2021. Part 
of the combined purchase price mentioned above with Build Security Ltd.
4. **Endgame, Inc.**: Although the context mentions an amended 
and restated stock incentive plan, it indicates that Endgame, Inc. was 
associated with Elastic, suggesting it could have been acquired at some point.

These acquisitions are part of Elastic's strategy to 
enhance the technology underlying its security and 
observability offerings.
<p>SimpleRAG</p>The provided context does not specify any companies that Elastic has acquired. 
Therefore, based on the context, there is no information available about the companies acquired by Elastic.
<p><strong>Resumen</strong>: SimpleRAG no recupera ninguna información relevante sobre adquisiciones, lo que lleva a una respuesta fallida. AdvancedRAG lista correctamente CmdWatch, Build Security y Optimyze, que fueron las adquisiciones clave listadas en el reporte.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">Volver arriba</a></p><h2>Conclusión</h2><p>Según nuestras pruebas, nuestras técnicas avanzadas parecen aumentar el rango y la profundidad de la información presentada, lo que podría mejorar la calidad de las respuestas RAG.</p><p>Además, puede haber mejoras en la fiabilidad, ya que preguntas formuladas de forma ambigua como <code>Which companies did Elastic acquire?</code> y <code>Who audits Elastic</code> fueron respondidas correctamente por AdvancedRAG pero no por SimpleRAG.</p><p>Sin embargo, conviene tener en cuenta que en 3 de cada 5 casos, la tubería básica de RAG, que incorpora Búsqueda Híbrida pero ninguna otra técnica, logró producir respuestas que capturaron la mayor parte de la información clave.</p><p>Cabe señalar que, debido a la incorporación de LLMs en las fases de preparación y consulta de datos, la latencia de AdvancedRAG suele ser entre 2 y 5 veces mayor que la de SimpleRAG. Este es un costo significativo que puede hacer que AdvancedRAG sea adecuado solo para situaciones donde la calidad de la respuesta se prioriza sobre la latencia.</p><p>Los importantes costos de latencia pueden aliviar usando un LLM más pequeño y barato como Claude Haiku o GPT-4o-mini en la fase de preparación de datos. Almacena los modelos avanzados para generar respuestas.</p><p>Esto está en línea con los hallazgos de Wang et al. Como muestran sus resultados, cualquier mejora realizada es relativamente incremental. En resumen, un RAG básico te lleva casi hasta un producto final decente, siendo además más barato y rápido. Para mí, es una conclusión interesante. Para casos de uso donde la velocidad y la eficiencia son clave, SimpleRAG es la opción sensata. Para casos de uso en los que hay que exprimir hasta la última gota de rendimiento, las técnicas incorporadas en AdvancedRAG pueden ofrecer una vía a seguir.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt56b7067a9d41d5a8/6a171119acf0886fb4be9c45/ea811706b6adc4731d90b925a9fefa0ac15901b4-1440x1060.jpg" alt="Oleoducto Wang" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2#table-of-contents">Volver arriba</a></p><h2>Apéndice</h2><h3>Prompts</h3><h4>Prompt de respuesta a preguntas RAG</h4><p>Prompt para que el LLM genere respuestas basadas en la consulta y el contexto.</p>BASIC_RAG_PROMPT = '''
You are an AI assistant tasked with answering questions based primarily on the provided context, while also drawing on your own knowledge when appropriate. Your role is to accurately and comprehensively respond to queries, prioritizing the information given in the context but supplementing it with your own understanding when beneficial. Follow these guidelines:

1. Carefully read and analyze the entire context provided.
2. Primarily focus on the information present in the context to formulate your answer.
3. If the context doesn't contain sufficient information to fully answer the query, state this clearly and then supplement with your own knowledge if possible.
4. Use your own knowledge to provide additional context, explanations, or examples that enhance the answer.
5. Clearly distinguish between information from the provided context and your own knowledge. Use phrases like "According to the context..." or "The provided information states..." for context-based information, and "Based on my knowledge..." or "Drawing from my understanding..." for your own knowledge.
6. Provide comprehensive answers that address the query specifically, balancing conciseness with thoroughness.
7. When using information from the context, cite or quote relevant parts using quotation marks.
8. Maintain objectivity and clearly identify any opinions or interpretations as such.
9. If the context contains conflicting information, acknowledge this and use your knowledge to provide clarity if possible.
10. Make reasonable inferences based on the context and your knowledge, but clearly identify these as inferences.
11. If asked about the source of information, distinguish between the provided context and your own knowledge base.
12. If the query is ambiguous, ask for clarification before attempting to answer.
13. Use your judgment to determine when additional information from your knowledge base would be helpful or necessary to provide a complete and accurate answer.

Remember, your goal is to provide accurate, context-based responses, supplemented by your own knowledge when it adds value to the answer. Always prioritize the provided context, but don't hesitate to enhance it with your broader understanding when appropriate. Clearly differentiate between the two sources of information in your response.

Context:
[The concatenated documents will be inserted here]

Query:
[The user's question will be inserted here]

Please provide your answer based on the above guidelines, the given context, and your own knowledge where appropriate, clearly distinguishing between the two:
'''
<h4>Prompt generador de consultas elástico</h4><p>Prompt para enriquecer consultas con sinónimos y convertirlas al formato OR.</p>ELASTIC_SEARCH_QUERY_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating Elasticsearch query strings. Your task is to create the most effective query string for the given user question. This query string will be used to search for relevant documents in an Elasticsearch index.

Guidelines:
1. Analyze the user's question carefully.
2. Generate ONLY a query string suitable for Elasticsearch's match query.
3. Focus on key terms and concepts from the question.
4. Include synonyms or related terms that might be in relevant documents.
5. Use simple Elasticsearch query string syntax if helpful (e.g., OR, AND).
6. Do not use advanced Elasticsearch features or syntax.
7. Do not include any explanations, comments, or additional text.
8. Provide only the query string, nothing else.

For the question "What is Clickthrough Data?", we would expect a response like:
clickthrough data OR click-through data OR click through rate OR CTR OR user clicks OR ad clicks OR search engine results OR web analytics

AND operator is not allowed. Use only OR.

User Question:
[The user's question will be inserted here]

Generate the Elasticsearch query string:
'''
<h4>Prompt generador de preguntas potenciales</h4><p>Prompt para generar posibles preguntas, enriquecer metadatos del documento.</p>RAG_QUESTION_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating questions for Retrieval-Augmented Generation (RAG) systems. Your task is to analyze a given document and create 10 diverse questions that would effectively test a RAG system's ability to retrieve and synthesize information from this document.

Guidelines:
1. Thoroughly analyze the entire document.
2. Generate exactly 10 questions that cover various aspects and levels of complexity within the document's content.
3. Create questions that specifically target:
   a. Key facts and information
   b. Main concepts and ideas
   c. Relationships between different parts of the content
   d. Potential applications or implications of the information
   e. Comparisons or contrasts within the document
4. Ensure questions require answers of varying lengths and complexity, from simple retrieval to more complex synthesis.
5. Include questions that might require combining information from different parts of the document.
6. Frame questions to test both literal comprehension and inferential understanding.
7. Avoid yes/no questions; focus on open-ended questions that promote comprehensive answers.
8. Consider including questions that might require additional context or knowledge to fully answer, to test the RAG system's ability to combine retrieved information with broader knowledge.
9. Number the questions from 1 to 10.
10. Output only the ten questions, without any additional text, explanations, or answers.

Document:
[The document content will be inserted here]

Generate 10 questions optimized for testing a RAG system based on this document:
'''
<h4>Prompt generador de HyDE</h4><p>Prompt para generar documentos hipotéticos usando HyDE</p>HYDE_DOCUMENT_GENERATOR_PROMPT = '''
You are an AI assistant specialized in generating hypothetical documents based on user queries. Your task is to create a detailed, factual document that would likely contain the answer to the user's question. This hypothetical document will be used to enhance the retrieval process in a Retrieval-Augmented Generation (RAG) system.

Guidelines:
1. Carefully analyze the user's query to understand the topic and the type of information being sought.
2. Generate a hypothetical document that:
   a. Is directly relevant to the query
   b. Contains factual information that would answer the query
   c. Includes additional context and related information
   d. Uses a formal, informative tone similar to an encyclopedia or textbook entry
3. Structure the document with clear paragraphs, covering different aspects of the topic.
4. Include specific details, examples, or data points that would be relevant to the query.
5. Aim for a document length of 200-300 words.
6. Do not use citations or references, as this is a hypothetical document.
7. Avoid using phrases like "In this document" or "This text discusses" - write as if it's a real, standalone document.
8. Do not mention or refer to the original query in the generated document.
9. Ensure the content is factual and objective, avoiding opinions or speculative information.
10. Output only the generated document, without any additional explanations or meta-text.

User Question:
[The user's question will be inserted here]

Generate a hypothetical document that would likely contain the answer to this query:
'''
<h3>Consulta de búsqueda híbrida de ejemplo</h3>{'knn': {'field': 'primary_embedding',
  'query_vector': [0.4265527129173279,
   -0.1712949573993683,
   -0.042020395398139954,
   ...],
  'k': 100,
  'num_candidates': 100},
 'query': {'bool': {'must': [{'multi_match': {'query': 'audits Elastic Elastic auditing Elastic audit process Elastic compliance Elastic security audit Elasticsearch auditing Elasticsearch compliance Elasticsearch security audit',
      'fields': ['original_text',
       'keyphrases',
       'potential_questions',
       'entities'],
      'type': 'best_fields',
      'operator': 'or'}}],
   'should': [{'script_score': {'query': {'match_all': {}},
      'script': {'source': '\n                                        double vector_score = cosineSimilarity(params.query_vector, params.vector_field) + 1.0;\n                                        double text_score = _score;\n                                        return 0.7 * vector_score + 0.3 * text_score;\n                                        ',
       'params': {'query_vector': [0.4265527129173279,
         -0.1712949573993683,
         -0.042020395398139954,
        ...],
        'vector_field': 'primary_embedding'}}}}]}},
 'size': 10}
]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2</guid>
    <category><![CDATA[Base de datos vectorial]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf605c8246989df32/6a1711178b73cbc61d18a11d/8da40067835ab8b4dc12fe52a51a6c26858ad32f-1440x1095.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Técnicas avanzadas de RAG parte 1: Procesamiento de datos]]></title>
    <description><![CDATA[Discutir e implementar técnicas que puedan aumentar el rendimiento de RAG. Parte 1 de 2, centrada en el componente de procesamiento e ingesta de datos de una tubería avanzada de RAG.]]></description>
    <content:encoded><![CDATA[<p><em>Esta es la Parte 1 de nuestra exploración sobre las Técnicas Avanzadas de RAG. </em><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2"><em>¡Haz clic aquí para la Parte 2!</em></a></p><p>El reciente artículo <a href="https://arxiv.org/abs/2407.01219">Searching for Best Practices in Retrieval-Augmented Generation</a> evalúa empíricamente la eficacia de diversas técnicas de mejora del RAG, con el objetivo de converger en un conjunto de mejores prácticas para el RAG.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt671704ff06a4011d/6a170b3ea929cf2d19ae09d8/dafa7250e7c4ead4d9b4aed7c407509131929749-1440x572.png" alt="Oleoducto RAG recomendado por Wang" /><p>Implementaremos algunas de estas mejores prácticas propuestas, concretamente aquellas que buscan mejorar la calidad de la búsqueda <strong>(fragmentación de oraciones, HyDE, empaquetado inverso).</strong></p><p>Por brevedad, omitiremos aquellas técnicas centradas en mejorar la eficiencia <strong>(Clasificación de Consultas y Resumen).</strong></p><p>También implementaremos algunas técnicas que no se trataron, pero que personalmente encuentro útiles e <strong>interesantes (Inclusión de Metadatos, Incrustaciones Compuestas Multi-Campo, Enriquecimiento de Consultas).</strong></p><p>Finalmente, realizaremos una breve prueba para ver si la calidad de nuestros resultados de búsqueda y respuestas generadas mejoró respecto a la línea base. ¡Vamos a ello!</p><h2>Resumen de RAG</h2><p>RAG tiene como objetivo mejorar los LLMs recuperando información de bases de conocimiento externas para enriquecer las respuestas generadas. Al proporcionar información específica de dominio, los LLM pueden adaptar rápidamente a casos de uso fuera del alcance de sus datos de entrenamiento; Significativamente más barato que el ajuste fino y más fácil de mantener actualizado.</p><p>Las medidas para mejorar la calidad de RAG suelen centrar en dos pistas:</p><ol><li><p>Mejorar la calidad y claridad de la base de conocimiento.</p></li><li><p>Mejorar la cobertura y especificidad de las consultas de búsqueda.</p></li></ol><p>Estas dos medidas lograrán el objetivo de mejorar las probabilidades de que el LLM tenga acceso a hechos e información relevantes, y así sea menos probable que alucine o se base en su propio conocimiento, que puede estar desactualizado o irrelevante.</p><p>La diversidad de métodos es difícil de aclarar en solo unas pocas frases. Vamos directamente a la implementación para aclarar las cosas.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" alt="Pipeline avanzado de RAG" /><h3>Índice</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#overview">Visión general</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Índice</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#set-up">Preparación</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#ingesting-processing-and-embedding-documents">Ingestión, procesamiento e incrustación de documentos</a>  </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#data-ingestion">Ingesta de datos</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#sentence-level-token-wise-chunking">Fragmentación a nivel de frase, por fichas</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#metadata-inclusion-and-generation">Inclusión y generación de metadatos</a> </p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#keyphrases-extracted-by-textrank">Frases clave extraídas por TextRank</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#potential-questions-generated-by-gpt-4o">Posibles preguntas generadas por GPT-4o</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#entities-extracted-by-spacy">Entidades extraídas por Spacy</a></p></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#composite-multi-field-embeddings">Incrustaciones compuestas multicampo</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#indexing-to-elastic">Indexación a Elastic</a></p></li></ul></li></ul></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#cat-break">Ruptura de gato</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#appendix">Apéndice</a></p><ul><li><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#definitions">Definiciones</a></p></li></ul></li></ul><h2>Preparación</h2><p><em>Todo el código puede </em><a href="https://github.com/elastic/elasticsearch-labs/tree/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques"><em>encontrar en el repositorio de Searchlabs</em></a><em>.</em></p><p>Primero lo primero. Necesitarás lo siguiente:</p><ol><li><p>Un despliegue de nube elástica</p></li><li><p>Una API LLM - Estamos usando un despliegue GPT-4o en Azure OpenAI en este cuaderno</p></li><li><p>Python Versión 3.12.4 o posterior</p></li></ol><p>Ejecutaremos todo el código desde <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/main.ipynb">el cuaderno main.ipynb.</a></p><p>Adelante, clona el repositorio por git, navega a supporting-blog-content/advanced-rag-techniques y luego ejecuta los siguientes comandos:</p># Create a new virtual environment named 'rag_env'
python -m venv rag_env

# Activate the virtual environment (for Unix-based systems)
source rag_env/bin/activate

# (For Windows)
.\rag_env\Scripts\activate

# Install packages listed in requirements.txt
pip install -r requirements.txt
<p>Una vez hecho esto, crea un <em>.env</em> y rellenar los siguientes campos (Referenciado en <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/.env.example"><em>.env.example</em></a>). Créditos a mi coautor, Claude-3.5, por los comentarios útiles.</p># Elastic Cloud: Found in the 'Deployment' page of your Elastic Cloud 
# console
ELASTIC_CLOUD_ENDPOINT=""
ELASTIC_CLOUD_ID=""

# Elastic Cloud: Created during deployment setup or in 'Security' 
# settings
ELASTIC_USERNAME=""
ELASTIC_PASSWORD=""

# Elastic Cloud: The name of the index you created in Kibana or via API
ELASTIC_INDEX_NAME=""

# Azure AI Studio: Found in 'Keys and Endpoint' section of your Azure 
# OpenAI resource
AZURE_OPENAI_KEY_1=""
AZURE_OPENAI_KEY_2=""
AZURE_OPENAI_REGION=""
AZURE_OPENAI_ENDPOINT=""

# Azure AI Studio: Found in 'Deployments' section of your Azure OpenAI 
# resource
AZURE_OPENAI_DEPLOYMENT_NAME=""

# Using BAAI/bge-small-en-v1.5 because I think it is a good balance of 
# resource efficiency and performance. 
HUGGINGFACE_EMBEDDING_MODEL="BAAI/bge-small-en-v1.5"
<p>A continuación, elegimos el documento a ingerir y lo colocaremos en la carpeta de documentos. Para este artículo, emplearemos el <a href="https://s201.q4cdn.com/217177842/files/doc_downloads/OtherDocuments/2023/AnnualMeeting/Annual-Report-Fiscal-Year-2023.pdf">Reporte Anual 2023 de Elastic N.V</a>. Es un documento bastante exigente y denso, perfecto para poner a prueba nuestras técnicas RAG.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte292dc6030d496cc/6a170b40dc55de9b03e00dfc/e513b9d67adac43da794c25a5969b893127bbbe3-1440x395.jpg" alt="Reporte Anual de Elastic 2023" /><p>Ahora que estamos listos, vamos a ingestión. Abre <em>main.ipynb</em> y ejecuta las dos primeras celdas para importar todos los paquetes e iniciar todos los servicios.</p><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Volver arriba</a></p><h2>Ingestión, procesamiento e incrustación de documentos</h2><h3>Ingesta de datos</h3><ul><li><p><em>Nota personal: Me sorprende la comodidad de LlamaIndex. En la antigüedad, antes de los LLMs y LlamaIndex, ingerir documentos de varios formatos era un proceso doloroso de recopilar paquetes esotéricos de todas partes. Ahora se reduce a una sola llamada de función. Salvaje.</em></p></li></ul><p>El <code>SimpleDirectoryReader</code> cargará todos los documentos del <code>directory_path.</code> Para <code>.pdf</code> archivos, devuelve una lista de objetos documento, que convierto a diccionarios de Python porque me resultan más fáciles de manejar.</p># llamaindex_processor.py
from llama_index.core import SimpleDirectoryReader

class LlamaIndexProcessor:
   def __init__(self):
       pass 
   
   def load_documents(self, directory_path):
       ''' 
       Load all documents in directory
       '''
       reader = SimpleDirectoryReader(input_dir=directory_path)
       return reader.load_data()

# main.ipynb
llamaindex_processor=LlamaIndexProcessor()
documents=llamaindex_processor.load_documents('./documents/')
documents=[dict(doc_obj) for doc_obj in documents]
<p>Cada diccionario contiene el contenido clave en el campo <code>text</code> . También contiene metadatos útiles como número de página, nombre de archivo, tamaño y tipo.</p>{
  'id_': '5f76f0b3-22d8-49a8-9942-c2bbab14f63f',
  'metadata': {'page_label': '5',
   'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_path': '/Users/han/Desktop/Projects/truckasaurus/documents/Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf',
   'file_type': 'application/pdf',
   'file_size': 3724426,
   'creation_date': '2024-07-27',
   'last_modified_date': '2024-07-27'},
   'text': 'Table of Contents\nPage\nPART I\nItem 1. Business 3\n15 Item 1A. Risk Factors\nItem 1B. Unresolved Staff Comments 48\nItem 2. Properties 48\nItem 3. Legal Proceedings 48\nItem 4. Mine Safety Disclosures 48\nPART II\nItem 5. Market for Registrant's Common Equity, Related Stockholder Matters and Issuer Purchases of \nEquity Securities49\nItem 6. [Reserved] 49\nItem 7. Management's Discussion and Analysis of Financial Condition and Results of Operations 50\nItem 7A. Quantitative and Qualitative Disclosures About Market Risk 64\nItem 8. Financial Statements and Supplementary Data 66\nItem 9. Changes in and Disagreements With Accountants on Accounting and Financial Disclosure 100\n100\n101Item 9A. Controls and Procedures\nItem 9B. Other Information\nItem 9C. Disclosure Regarding Foreign Jurisdictions That Prevent Inspections 101\nPART III\n102\n102\n102\n102Item 10. Directors, Executive Officers and Corporate Governance\nItem 11. Executive Compensation\nItem 12. Security Ownership of Certain Beneficial Owners and Management, and Related Stockholder Matters  \nItem 13. Certain Relationships and Related Transactions, and Director Independence\nItem 14. Principal Accountant Fees and Services 102\nPART IV\n103\n105Item 15. Exhibits and Financial Statement Schedules  \nItem 16. Form 10-K Summary\nSignatures 106\ni',
   ...
}
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Volver arriba</a></p><h3>Fragmentación a nivel de frase, por fichas</h3><p>Lo primero que hay que hacer es reducir nuestros documentos a fragmentos de una longitud estándar (para garantizar la coherencia y la manejabilidad). Los modelos de incrustación tienen límites únicos de tokens (tamaño máximo de entrada que pueden procesar). Los tokens son las unidades básicas de texto que procesan los modelos. Para evitar la pérdida de información (truncamiento u omisión de contenido), deberíamos proporcionar texto que no exceda esos límites (dividiendo textos más largos en segmentos más pequeños).</p><p>El chunking tiene un impacto significativo en el rendimiento. Idealmente, cada fragmento representaría una pieza de información autónoma, capturando información contextual sobre un único tema. Los métodos de fragmentación incluyen el fragmento a nivel de palabra, donde los documentos se dividen por el recuento de palabras, y el fragmento semántico, que emplea un LLM para identificar puntos de interrupción lógicos.</p><p>El fragmento a nivel de palabra es barato, rápido y sencillo, pero corre el riesgo de fragmentar las frases y así romper el contexto. El fragmento semántico se vuelve lento y caro, especialmente si se trata de documentos como el Reporte Anual de Elastic de 116 páginas.</p><p>Elijamos un enfoque intermedio. El fragmento a nivel de oración sigue siendo sencillo, pero puede preservar el contexto de forma más eficaz que el fragmento a nivel de palabra, siendo significativamente más barato y rápido. Además, implementaremos una ventana deslizante para capturar parte del contexto circundante y aliviar el impacto de dividir los párrafos.</p># chunker.py 

import uuid
import re


class Chunker: 
    def __init__(self, tokenizer):
        self.tokenizer = tokenizer 
    
    def split_into_sentences(self, text):
        """Split text into sentences."""
        return re.split(r'(?&lt;=[.!?])\s+', text)
 
    def sentence_wise_tokenized_chunk_documents(self, documents, chunk_size=512, overlap=20, min_chunk_size=50):
        '''
        1. Split text into sentences.
        2. Tokenize using the provided tokenizer method.
        3. Build chunks up to the chunk_size limit.
        4. Create an overlap based on tokens - to preserve context.
        5. Only keep chunks that meet the minimum token size requirement.
        '''
        chunked_documents = []

        for doc in documents:
            sentences = self.split_into_sentences(doc['text'])
            tokens = []
            sentence_boundaries = [0]

            # Tokenize all sentences and keep track of sentence boundaries
            for sentence in sentences:
                sentence_tokens = self.tokenizer.encode(sentence, add_special_tokens=True)
                tokens.extend(sentence_tokens)
                sentence_boundaries.append(len(tokens))

            # Create chunks
            chunk_start = 0
            while chunk_start &lt; len(tokens):
                chunk_end = chunk_start + chunk_size

                # Find the last complete sentence that fits in the chunk
                sentence_end = next((i for i in sentence_boundaries if i &gt; chunk_end), len(tokens))
                chunk_end = min(chunk_end, sentence_end)

                # Create the chunk
                chunk_tokens = tokens[chunk_start:chunk_end]

                # Check if the chunk meets the minimum size requirement
                if len(chunk_tokens) &gt;= min_chunk_size:
                    # Create a new document object for this chunk
                    chunk_doc = {
                        'id_': str(uuid.uuid4()),
                        'chunk': chunk_tokens,
                        'original_text': self.tokenizer.decode(chunk_tokens),
                        'chunk_index': len(chunked_documents),
                        'parent_id': doc['id_'],
                        'chunk_token_count': len(chunk_tokens)
                    }

                    # Copy all other fields from the original document
                    for key, value in doc.items():
                        if key != 'text' and key not in chunk_doc:
                            chunk_doc[key] = value

                    chunked_documents.append(chunk_doc)

                # Move to the next chunk start, considering overlap
                chunk_start = max(chunk_start + chunk_size - overlap, chunk_end - overlap)

        return chunked_documents

# main.ipynb 
# Initialize Embedding Model
HUGGINGFACE_EMBEDDING_MODEL = os.environ.get('HUGGINGFACE_EMBEDDING_MODEL')
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

# Initialize Chunker
chunker=Chunker(embedder.tokenizer)
<p>La clase <code>Chunker</code> incorpora el tokenizador del modelo de incrustación para codificar y decodificar texto. Ahora construiremos fragmentos de 512 tokens cada uno, con una superposición de 20 tokens. Para ello, dividiremos el texto en frases, tokenizaremos esas frases y luego agregaremos las frases tokenizadas a nuestro fragmento actual hasta que no podamos agregar más sin superar nuestro límite de tokens.</p><p>Finalmente, decodifica las frases de nuevo al texto original para incrustarlas, almacenándola en un campo llamado <code>original_text</code>. Los chunks se almacenan en un campo llamado <code>chunk</code>. Para reducir el ruido (es decir, documentos inútiles), descartaremos cualquier documento de menos de 50 tokens de longitud.</p><p>Vamos a repasarla por nuestros documentos:</p>chunked_documents=chunker.sentence_wise_tokenized_chunk_documents(documents, chunk_size=512)
<p>Y que me devolvan fragmentos de texto que se parezcan a esto:</p>print(chunked_documents[4]['original_text'])

[CLS] the aggregate market value of the ordinary shares held by non - affiliates of the registrant, 
based on the closing price of the shares of ordinary shares on the new york stock exchange on 
october 31, 2022 ( the last business day of the registrant 's second fiscal quarter ), was 
approximately $ 6. 1 billion. [SEP] [CLS] as of may 31, 2023, the registrant had 97, 390, 886 
ordinary shares, par value €0. 01 per share, outstanding. [SEP] [CLS] documents incorporated by 
reference portions of the registrant 's definitive proxy statement relating to the registrant 's 2
023 annual general meeting of shareholders are incorporated by reference into part iii of this annual 
...
...
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Volver arriba</a></p><h3>Inclusión y generación de metadatos</h3><p>Dividimos nuestros documentos. Ahora es el momento de enriquecer los datos. Quiero generar o extraer metadatos adicionales. Estos metadatos adicionales pueden emplear para influir y mejorar el rendimiento en las búsquedas.</p><p>Definiremos una clase <code>DocumentEnricher</code> , cuyo papel es incluir una lista de documentos (diccionarios de Python) y una lista de funciones del procesador. Estas funciones se ejecutarán sobre la columna <code>original_text</code> de los documentos y almacenarán sus salidas en nuevos campos.</p><p>Primero, extraemos las frases clave usando <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/nltk_processor.py">TextRank</a>. TextRank es un algoritmo basado en gráficos que extrae frases clave y oraciones del texto clasificando su importancia en función de las relaciones entre palabras.</p><p>A continuación, <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/llm.py">generaremos potential_questions usando GPT-4o</a>.</p><p>Finalmente, <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/entity_extractor.py">extraeremos entidades</a> usando <a href="https://spacy.io/">Spacy</a>.</p><p>Dado que el código de cada uno de estos es bastante extenso y complejo, me abstendré de reproducirlo aquí. Si te interesa, los archivos están marcados en los ejemplos de código que aparecen a continuación.</p><p>Vamos a ejecutar el enriquecimiento de datos:</p># documentenricher.py
from tqdm import tqdm

class DocumentEnricher:

    def __init__(self):
        pass 

    def enrich_document(self, documents, processors, text_col='text'):
        for doc in tqdm(documents, desc="Enriching documents using processors: "+str(processors)): 
            for (processor, field) in processors: 
                metadata=processor(doc[text_col])
                if isinstance(metadata, list):
                    metadata='\n'.join(metadata)
                doc.update({field: metadata})
 
# main.ipynb
# Initialize processor classes 
nltkprocessor=NLTKProcessor() // nltk_processor.py
entity_extractor=EntityExtractor() // entity_extractor.py
gpt4o = LLMProcessor(model='gpt-4o') // llm.py

# Initialize LLM
documentenricher=DocumentEnricher()

# Create new fields in the documents - These are the outputs of the processor functions.
processors=[
    (nltkprocessor.textrank_phrases, "keyphrases"),
    (gpt4o.generate_questions, "potential_questions"),
    (entity_extractor.extract_entities, "entities")
    ]

# .enrich_document() will modify chunked_docs in place. 
# To view the results, we'll print chunked_docs in the next few cells!
documentenricher.enrich_document(chunked_docs, text_col='original_text', processors=processors)
<p>Y echa un vistazo a los resultados:</p><h4>Frases clave extraídas por TextRank</h4><p>Estas frases clave son un sustituto de los temas centrales del fragmento. Si una consulta tiene que ver con ciberseguridad, el puntaje de este segmento se incrementará.</p>print(chunked_documents[25]['keyphrases'])

'elastic agent stop', 'agent stop malware', 
'stop malware ransomware', 'malware ransomware environment', 
'ransomware environment wide', 'environment wide visibility', 
'wide visibility threat', 'visibility threat detection', 
'sep cl key', 'cl key feature'
<h4>Posibles preguntas generadas por GPT-4o</h4><p>Estas posibles preguntas pueden coincidir directamente con las consultas de los usuarios, ofreciendo un aumento en el puntaje. Pedimos a GPT-4o que genere preguntas que pueden responder usando la información encontrada en el fragmento actual.</p>print(chunked_documents[25]['potential_questions'])

1. What are the primary functions that Elastic Agent provides in terms of cybersecurity?
2. Describe how Logstash contributes to data management within an IT environment.
3. List and explain any key features of Logstash mentioned in the document.
4. How does Elastic Agent enhance environment-wide visibility in threat detection?
5. What capabilities does Logstash offer for handling data beyond simple collection?
6. In what ways does the document suggest that Elastic Agent stops malware and ransomware?
7. Can you identify any relationships between the functionalities of Elastic Agent and Logstash in an integrated environment?
8. What implications might the advanced threat detection capabilities of Elastic Agent have for organizational security policies?
9. Compare and contrast the roles of Elastic Agent and Logstash based on their described functions.
10. How might the centralized collection ability of Logstash support the threat detection capabilities of Elastic Agent?
<h4>Entidades extraídas por Spacy</h4><p>Estas entidades cumplen un propósito similar al de las frases clave, pero capturan los nombres de organizaciones e individuos, que la extracción de frases clave puede pasar por alto.</p>print(chunked_documents[29]['entities'])

'appdynamics', 'apm data', 'azure sentinel', 
'microsoft', 'mcafee', 'broadcom', 'cisco', 
'dynatrace', 'coveo', 'lucidworks'
<p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Volver arriba</a></p><h3>Incrustaciones compuestas multicampo</h3><p>Ahora que enriquecimos nuestros documentos con metadatos adicionales, podemos aprovechar esta información para crear incrustaciones más robustas y conscientes del contexto.</p><p>Repasemos nuestro punto actual en el proceso. Tenemos cuatro campos de interés en cada documento.</p>{
    "chunk": "...",
    "keyphrases": "...", 
    "potential_questions": "...", 
    "entities": "..." 
}
<p>Cada campo representa una perspectiva diferente sobre el contexto del documento, lo que puede destacar un área clave en la que el LLM debe centrar.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt84cb328fce6aae23/6a170b42964cea3e4408bbc4/aea1f513009a0c7c8545a79fad8f072a5bcae24c-1440x1067.jpg" alt="Pipeline de Enriquecimiento de Metadatos en RAG" /><p>El plan es incrustar cada uno de estos campos y luego crear una suma ponderada de las incrustaciones, conocida como Incrustación Compuesta.</p><p>Con suerte, esta Incrustación Compuesta permitirá que el sistema sea más consciente del contexto, además de introducir otro hiperparámetro ajustable que controla el comportamiento de búsqueda.</p><p>Primero, embebamos cada campo y actualicemos cada documento en su lugar, usando nuestro modelo de incrustación definido localmente importado al inicio del cuaderno main.ipynb.</p># EmbeddingModel defined in embedding_model.py
embedder=EmbeddingModel(model_name=HUGGINGFACE_EMBEDDING_MODEL)

cols_to_embed=['keyphrases', 'potential_questions', 'entities']

embedding_cols=[]
for col in cols_to_embed:
    # Works on text input
    embedding_col=embedder.embed_documents_text_wise(chunked_documents, text_field=col)
    embedding_cols.append(embedding_col)
# Works on token input
embedding_col=embedder.embed_documents_token_wise(chunked_documents, token_field="chunk")
embedding_cols.append(embedding_col)
<p>Cada función de incrustación devuelve el campo de la incrustación, que es simplemente el campo de entrada original con un <code>_embedding</code> postfijo.</p><p>Ahora definamos las ponderaciones de nuestra incrustación compuesta:</p>embedding_cols=[
                'keyphrases_embedding',
                'potential_questions_embedding',
                'entities_embedding',
                'chunk_embedding']
combination_weights=[
                    0.1,
                    0.15,
                    0.05,
                    0.7
                ]
<p>Las ponderaciones te permiten asignar prioridades a cada componente, basándote en tu caso de uso y la calidad de tus datos. Intuitivamente, el tamaño de estos pesos depende del valor semántico de cada componente. Como el texto en fragmentos en sí es, con diferencia, el más rico, asigno un peso del 70%. Como las entidades son las más pequeñas, siendo solo una lista de nombres de organizaciones o personas, le asigno un peso del 5%. La configuración precisa de estos valores debe determinar empíricamente, caso de uso por caso.</p><p>Finalmente, escribamos una función para aplicar los pesos y creemos nuestra incrustación compuesta. También eliminaremos todas las incrustaciones de componentes para ahorrar espacio.</p>from tqdm import tqdm 
def combine_embeddings(objects, embedding_cols, combination_weights, primary_embedding='primary_embedding'):
    # Ensure the number of weights matches the number of embedding columns
    assert len(embedding_cols) == len(combination_weights), "Number of embedding columns must match number of weights"
    
    # Normalize weights to sum to 1
    weights = np.array(combination_weights) / np.sum(combination_weights)
    
    for obj in tqdm(objects, desc="Combining embeddings"):
        # Initialize the combined embedding
        combined = np.zeros_like(obj[embedding_cols[0]])
        
        # Compute the weighted sum
        for col, weight in zip(embedding_cols, weights):
            combined += weight * np.array(obj[col])
        
        # Add the new combined embedding to the object
        obj.update({primary_embedding:combined.tolist()})
        
        # Remove the original embedding columns
        for col in embedding_cols:
            obj.pop(col, None)

combine_embeddings(chunked_documents, embedding_cols, combination_weights)
<p>Con esto, completamos el procesamiento de los documentos. Ahora tenemos una lista de objetos documento que se ven así:</p>{ 'id_': '7fe71686-5cd0-4831-9e79-998c6dbeae0c', 'chunk': [2312, 14613, ...], 'original_text': 'if an emerging growth company, indicate by check mark if the registrant has elected not to use the extended ...', 'chunk_index': 3, 'chunk_token_count': 399, 'metadata': {'page_label': '3', 'file_name': 'Elastic_NV_Annual-Report-Fiscal-Year-2023.pdf', ... 'keyphrases': 'sep cl unk\ncheck mark registrant\ncl unk indicate\nunk indicate check\nindicate check mark\nprincipal executive office\naccelerate filer unk\ncompany unk emerge\nunk emerge growth\nemerge growth company', 'potential_questions': '1. What are the different types of registrant statuses mentioned in the document?\n2. Under what section of the Sarbanes-Oxley Act must registrants file a report on the effectiveness of their internal ...', 'entities': 'the effe ctiveness of\nsection 13\nSEP\nUNK\nsection 21e\n1934\n1933\nu. s. c.\nsection 404\nsection 12\nal', 'primary_embedding': [-0.3946287803351879, -0.17586839850991964, ...] }
<h4>Indexación a Elastic</h4><p>Subamos nuestros documentos en masa a Elastic Search. Para este propósito, hace mucho tiempo definí un conjunto de funciones auxiliares elásticas en <a href="https://github.com/elastic/elasticsearch-labs/blob/advanced-rag-techniques/supporting-blog-content/advanced-rag-techniques/elastic_helpers.py"><code>elastic_helpers.py</code></a>. Es un código muy largo, así que vamos a centrarnos en las llamadas a funciones.</p><p><code>es_bulk_indexer.bulk_upload_documents</code> funciona con cualquier lista de objetos de diccionario, aprovechando los convenientes mapeos dinámicos de Elasticsearch.</p># Initialize Elasticsearch
ELASTIC_CLOUD_ID = os.environ.get('ELASTIC_CLOUD_ID')
ELASTIC_USERNAME = os.environ.get('ELASTIC_USERNAME')
ELASTIC_PASSWORD = os.environ.get('ELASTIC_PASSWORD')
ELASTIC_CLOUD_AUTH = (ELASTIC_USERNAME, ELASTIC_PASSWORD)
es_bulk_indexer = ESBulkIndexer(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)
es_query_maker = ESQueryMaker(cloud_id=ELASTIC_CLOUD_ID, credentials=ELASTIC_CLOUD_AUTH)

# Define Index Name
index_name=os.environ.get('ELASTIC_INDEX_NAME')


# Create index and bulk upload 
index_exists = es_bulk_indexer.check_index_existence(index_name=index_name)
if not index_exists:
    logger.info(f"Creating new index: {index_name}")
    es_bulk_indexer.create_es_index(es_configuration=BASIC_CONFIG, index_name=index_name)

success_count = es_bulk_indexer.bulk_upload_documents(
    index_name=index_name, 
    documents=chunked_documents, 
    id_col='id_',
    batch_size=32
)
<p>Ve a Kibana y verifica que todos los documentos fueron indexados. Deberían ser 224. ¡No está mal para un documento tan extenso!</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8efeface6effe01d/6a170b447d8d67652870e72a/1b3b07f6b98ceb65f6594ce4be83c5b0ed7e7cf9-1440x1380.jpg" alt="Índice Kibana" /><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Volver arriba</a></p><h2>Ruptura de gato</h2><p>Vamos a hacer una pausa, el artículo es un poco pesado, lo sé. Mira a mi gato:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1db5595f71c12ff/6a170b450e2e49940241a0fe/baca4eb52b801b21ced97352cc55462f0a12d6b0-969x996.jpg" alt="Gasoducto Han" /><p>Adorable. El sombrero desapareció y sospecho que lo robó y escondió en algún sitio :(</p><p>¡Enhorabuena por llegar hasta aquí :)</p><p>¡Únete a mí en <a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-2">la Parte 2</a> para probar y evaluar nuestra cadena RAG!</p><h2>Apéndice</h2><h3>Definiciones</h3><p><strong>1. Fragmentación de frases</strong></p><ul><li><p>Una técnica de preprocesamiento empleada en sistemas RAG para dividir el texto en unidades más pequeñas y significativas.</p></li><li><p><em>Proceso:</em> </p><ol><li><p>Entrada: Gran bloque de texto (por ejemplo, documento, párrafo)</p></li><li><p>Salida: Segmentos de texto más pequeños (normalmente oraciones o pequeños grupos de oraciones)</p></li></ol></li><li><p><em>Propósito:</em> </p><ul><li><p>Crea segmentos de texto granulares y específicos de contexto</p></li><li><p>Permite una indexación y recuperación más precisas</p></li><li><p>Mejora la relevancia de la información recuperada en sistemas RAG</p></li></ul></li><li><p><em>Características:</em> </p><ul><li><p>Los segmentos tienen significado semántico</p></li><li><p>Puede ser indexado y recuperado de forma independiente</p></li><li><p>A menudo preserva cierto contexto para garantizar la comprensibilidad independiente</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Mejora la precisión de la recuperación</p></li><li><p>Permite una ampliación más enfocada en las canalizaciones RAG</p></li></ul></li></ul><p><strong>2. HyDE (Embedding de Documentos Hipotéticos)</strong></p><ul><li><p>Una técnica que emplea un LLM para generar un documento hipotético para la expansión de consultas en sistemas RAG.</p></li><li><p><em>Proceso:</em>  </p><ol><li><p>Consulta de entrada a un LLM</p></li><li><p>El LLM genera un documento hipotético que responde a la consulta</p></li><li><p>Incrustar el documento generado</p></li><li><p>Emplea la incrustación para búsqueda vectorial</p></li></ol></li><li><p><em>Diferencia clave:</em> </p><ul><li><p>RAG tradicional: Empareja la consulta con documentos</p></li><li><p>HyDE: Empareja documentos con documentos</p></li></ul></li><li><p><em>Propósito:</em> </p><ul><li><p>Mejorar el rendimiento de la recuperación, especialmente para consultas complejas o ambiguas</p></li><li><p>Captura un contexto semántico más rico que una consulta corta</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Aprovecha el conocimiento de LLM para ampliar las consultas</p></li><li><p>Puede mejorar potencialmente la relevancia de los documentos recuperados</p></li></ul></li><li><p><em>Desafíos:</em> </p><ul><li><p>Requiere inferencia adicional de LLM, aumentando la latencia y el costo</p></li><li><p>El rendimiento depende de la calidad del documento hipotético generado</p></li></ul></li></ul><p><strong>3. Empaquetado inverso</strong></p><ul><li><p>Una técnica empleada en sistemas RAG para reordenar los resultados de búsqueda antes de pasarlos al LLM.</p></li><li><p><em>Proceso:</em> </p><ol><li><p>El motor de búsqueda (por ejemplo, Elasticsearch) devuelve los documentos en orden descendente de relevancia.</p></li><li><p>El orden se invierte, colocando el documento más relevante al final.</p></li></ol></li><li><p><em>Propósito:</em> </p><ul><li><p>Aprovecha el sesgo de actualidad de los LLM, que tienden a centrar más en la información más reciente en su contexto.</p></li><li><p>Garantiza que la información más relevante esté "más fresca" en la ventana de contexto del LLM.</p></li></ul></li><li><p><em>Ejemplo:</em> Orden original: [Más relevante, Segundo más, Tercero más, ...] Orden invertido: [..., Tercero más, segundo más relevante]</p></li></ul><p><strong>4. Clasificación de consultas</strong></p><ul><li><p>Una técnica para optimizar la eficiencia del sistema RAG determinando si una consulta requiere RAG o puede ser respondida directamente por el LLM.</p></li><li><p><em>Proceso:</em> </p><ol><li><p>Desarrollar un conjunto de datos personalizado específico para el LLM en uso</p></li><li><p>Capacitar un modelo de clasificación especializado</p></li><li><p>Emplea el modelo para categorizar las consultas entrantes</p></li></ol></li><li><p><em>Propósito:</em> </p><ul><li><p>Mejorar la eficiencia del sistema evitando el procesamiento innecesario de RAG</p></li><li><p>Consulta directa al mecanismo de respuesta más adecuado</p></li></ul></li><li><p><em>Requisitos:</em> </p><ul><li><p>Conjunto de datos y modelo específicos de LLM</p></li><li><p>Refinamiento continuo para mantener la precisión</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Reduce la sobrecarga computacional para consultas simples</p></li><li><p>Potencialmente mejora el tiempo de respuesta para consultas que no son RAG</p></li></ul></li></ul><p><strong>5. Resumen</strong></p><ul><li><p>Una técnica para condensar documentos recuperados en sistemas RAG.</p></li><li><p><em>Proceso:</em> </p><ol><li><p>Recuperar documentos relevantes</p></li><li><p>Genera resúmenes concisos de cada documento</p></li><li><p>Emplea resúmenes en lugar de documentos completos en la tubería RAG</p></li></ol></li><li><p><em>Propósito:</em> </p><ul><li><p>Mejora el rendimiento de RAG centrándote en la información esencial</p></li><li><p>Reducir el ruido y las interferencias de contenido menos relevante</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Potencialmente mejora la relevancia de las respuestas de los LLM</p></li><li><p>Permite incluir más documentos dentro de los límites del contexto</p></li></ul></li><li><p><em>Desafíos:</em> </p><ul><li><p>Riesgo de perder detalles importantes en la resumen</p></li><li><p>Sobrecarga computacional adicional para la generación de resúmenes</p></li></ul></li></ul><p><strong>6. Inclusión de metadatos</strong></p><ul><li><p>Una técnica para enriquecer documentos con información contextual adicional.</p></li><li><p><em>Tipos de metadatos:</em>  </p><ul><li><p>Frases clave</p></li><li><p>Títulos</p></li><li><p>Fechas</p></li><li><p>Detalles de la autoría</p></li><li><p>Resumen</p></li></ul></li><li><p><em>Propósito:</em> </p><ul><li><p>Aumentar la información contextual disponible para el sistema RAG</p></li><li><p>Proporcionar a los LLMs una comprensión más clara del contenido y la relevancia del documento</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Potencialmente mejora la precisión de la recuperación</p></li><li><p>Mejora la capacidad del LLM para evaluar la utilidad de los documentos</p></li></ul></li><li><p><em>Implementación:</em> </p><ul><li><p>Se puede hacer durante el preprocesamiento de documentos</p></li><li><p>Puede requerir pasos adicionales de extracción de datos o generación</p></li></ul></li></ul><p><strong>7. Incrustaciones compuestas multicampo</strong></p><ul><li><p>Una técnica avanzada de incrustación para sistemas RAG que crea incrustaciones separadas para diferentes componentes del documento.</p></li><li><p><em>Proceso:</em> </p><ol><li><p>Identificar campos relevantes (por ejemplo, título, frases clave, resúmenes, contenido principal)</p></li><li><p>Genera incrustaciones separadas para cada campo</p></li><li><p>Combina o almacena estos embeddings para su uso en la recuperación</p></li></ol></li><li><p><em>Diferencia con el enfoque estándar:</em> </p><ul><li><p>Tradicional: Embedding único para todo el documento</p></li><li><p>Compuesto: Múltiples incrustaciones para diferentes aspectos del documento</p></li></ul></li><li><p><em>Propósito:</em> </p><ul><li><p>Crear representaciones documentales más matizadas y conscientes del contexto</p></li><li><p>Captura información de una mayor variedad de fuentes dentro de un documento</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Potencialmente mejora el rendimiento en consultas ambiguas o multifacéticas</p></li><li><p>Permite una ponderación más flexible de los diferentes aspectos del documento en la recuperación</p></li></ul></li><li><p><em>Desafíos:</em> </p><ul><li><p>Mayor complejidad en los procesos de incrustación, almacenamiento y recuperación</p></li><li><p>Puede requerir algoritmos de emparejamiento más sofisticados</p></li></ul></li></ul><p><strong>8. Enriquecimiento de consultas</strong></p><ul><li><p>Una técnica para ampliar la consulta original con términos relacionados para mejorar la cobertura de búsqueda.</p></li><li><p><em>Proceso:</em> </p><ol><li><p>Analizar la consulta original</p></li><li><p>Generar sinónimos y frases semánticamente relacionadas</p></li><li><p>Complementa la consulta con estos términos adicionales</p></li></ol></li><li><p><em>Propósito:</em> </p><ul><li><p>Ampliar el rango de posibles coincidencias en el corpus documental</p></li><li><p>Mejorar el rendimiento de recuperación para consultas con lenguaje específico o técnico</p></li></ul></li><li><p><em>Beneficios:</em> </p><ul><li><p>Potencialmente recupera documentos relevantes que no coinciden exactamente con los términos originales de la consulta</p></li><li><p>Puede ayudar a superar la discrepancia de vocabulario entre consultas y documentos</p></li></ul></li><li><p><em>Desafíos:</em> </p><ul><li><p>Riesgo de deriva de consulta si no se implementa cuidadosamente</p></li><li><p>Puede aumentar la sobrecarga computacional en el proceso de recuperación</p></li></ul></li></ul><p><a href="https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1#table-of-contents">Volver arriba</a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/advanced-rag-techniques-part-1</guid>
    <category><![CDATA[Base de datos vectorial]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Han Xiang Choong]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a4691874a19d8da/6a170b3f47d49c99f22d8a24/72b51ba2ae5e5977b56e5b915674753d6cfd0e56-1440x840.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 14 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>