<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Priscilla Parodi - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Priscilla Parodi - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/kr/search-labs/author/priscilla-parodi</link>
    </image>
    <link>https://www.elastic.co/kr/search-labs/author/priscilla-parodi</link>
    <atom:link href="https://www.elastic.co/kr/search-labs/rss/author/priscilla-parodi.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[kr]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 05:13:22 GMT</lastBuildDate>
  <item>
    <title><![CDATA[AI 표절: Elasticsearch를 통한 표절 탐지]]></title>
    <description><![CDATA[NLP 모델과 벡터 검색을 사용한 사용 사례를 중심으로 Elasticsearch를 사용해 AI 표절을 확인하는 방법을 알려드립니다.]]></description>
    <content:encoded><![CDATA[<p>표절은 콘텐츠의 일부 또는 전체를 복사하는 <strong>직접적</strong> 표절과 일부 단어나 문구를 변경하여 저자의 저작물을 다시 표현하는 <strong>의역적</strong> 표절로 나눌 수 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb6e139760e56ec98/6a171147dc55de0ad2e00edf/5d7073187fda829438aeec8d3a1194a5bea2ba57-1440x347.png" alt="" /><p>영감과 의역에는 차이가 있습니다. 콘텐츠를 읽고 영감을 얻은 다음 비슷한 결론에 도달하더라도 자신의 말로 아이디어를 탐구할 수 있습니다.</p><p>표절은 오랫동안 논의의 대상이 되어 왔지만, 콘텐츠의 제작과 게시가 가속화되면서 표절 문제는 계속 제기되고 있으며 지속적인 과제가 되고 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc2af1678cb04506b/6a171149cf4f251c2ab2d257/a0b9a98d729db09dae0a79315c001e6763c12704-1400x1016.png" alt="" /><p>이 문제는 표절 검사가 자주 이루어지는 서적, 학술 연구 또는 사법 문서에만 국한되지 않습니다. 또한 신문과 소셜 미디어까지 확장할 수 있습니다.</p><p>정보가 풍부하고 퍼블리싱에 쉽게 접근할 수 있는 상황에서 어떻게 하면 확장 가능한 수준에서 표절을 효과적으로 검사할 수 있을까요?</p><p>대학, 정부 기관 및 기업에서는 다양한 도구를 사용하지만, 간단한 <a href="https://www.elastic.co/search-labs/lexical-and-semantic-search-with-elasticsearch">어휘 검색을</a> 통해 직접적인 표절을 효과적으로 감지할 수 있지만, 가장 큰 문제는 <strong>의역된 콘텐츠를</strong>식별하는 데 있습니다.</p><h2>생성적 AI를 통한 표절 탐지</h2><p>제너레이티브 AI로 새로운 도전이 시작됩니다. AI가 생성한 콘텐츠를 복사할 경우 표절로 간주되나요?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8a1ed1b6fa3f1a56/6a17114b4a531bc40736aa69/9345b28d6d27c37469bc38e823c41780b4eabfe5-1440x875.png" alt="" /><p>예를 들어 <a href="https://openai.com/">OpenAI</a> <a href="https://openai.com/policies/terms-of-use">이용약관에는</a> OpenAI가 사용자를 위해 API로 생성한 콘텐츠에 대한 저작권을 주장하지 않는다고 명시되어 있습니다. 이 경우 생성 AI를 사용하는 개인은 생성된 콘텐츠를 인용 없이 원하는 대로 사용할 수 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbc2e08ab5b808e38/6a17114dab7f086cbadb9f93/e1f415f69a247666f02ddc81468944920c874cd7-968x814.png" alt="" /><p>그러나 효율성을 개선하기 위해 제너레이티브 AI를 사용하는 것에 대한 수용 여부는 여전히 논의의 여지가 있습니다.</p><p>표절 탐지에 기여하기 위해 OpenAI는 <a href="https://huggingface.co/roberta-base-openai-detector">탐지 모델을</a> 개발했지만 나중에 그 정확도가 충분히 높지 않다는 사실을 인정했습니다.</p><p><em>"이는 단독으로 탐지하기에는 정확도가 충분하지 않으며, 메타데이터 기반 접근 방식, 사람의 판단, 대중 교육과 함께 사용해야 더 효과적이라고 생각합니다."</em></p><p>하지만 더 많은 도구가 제공되면서 의역 및 인공지능 콘텐츠의 경우에도 표절을 감지할 수 있는 옵션이 늘어났습니다.</p><h2>Elasticsearch로 표절 탐지하기</h2><p>이 블로그에서는 메타데이터 검색을 넘어 자연어 처리(NLP) 모델과 벡터 검색의 사용 사례인 표절 탐지에 대해 한 가지 더 살펴보고자 합니다.</p><p>이는 NLP 관련 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/plagiarism-detection-with-elasticsearch/plagiarism_detection_es.ipynb">기사가</a> 포함된 <a href="https://www.sbert.net/">SentenceTransformers의</a> <a href="https://sbert.net/datasets/emnlp2016-2018.json">데이터 세트를</a> 활용하는 Python 예제를 통해 설명합니다. 이전에 Elasticsearch로 가져온 텍스트 <a href="https://huggingface.co/sentence-transformers/all-mpnet-base-v2">임베딩 모델로</a> 생성된 '초록' 임베딩을 고려하여 '의미론적 텍스트 유사성'을 수행하여 초록의 표절 여부를 확인합니다. 또한, AI가 생성한 콘텐츠인 AI 표절을 식별하기 위해 OpenAI에서 개발한 자연어 처리( <a href="https://huggingface.co/roberta-base-openai-detector">NLP) 모델도</a> Elasticsearch로 가져왔습니다.</p><p>다음 이미지에는 데이터 흐름이 나와 있습니다:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0464f88d3ed12070/6a17114fab7f084905db9f97/1ad89c98a2f42a497548ca3947749bad54ec1172-1440x880.png" alt="" /><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest.html">추론 프로세서가</a> 있는 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/inference-processor.html">수집</a> 파이프라인에서 'abstract' 단락은 768차원 벡터인 'abstract_vector.predicted_value'에 매핑됩니다.</p><p>매핑:</p>"abstract_vector.predicted_value": { # Inference results field
"type": "dense_vector", 
"dims": 768, # model embedding_size
"index": "true", 
"similarity": "dot_product" # When indexing vectors for approximate kNN search, you need to specify the similarity function for comparing the vectors.
<p>벡터 표현 간의 유사성은 '유사성' <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html#dense-vector-params">매개변수를</a> 사용하여 정의된 벡터 유사성 메트릭을 사용하여 측정합니다.</p><p><a href="https://en.wikipedia.org/wiki/Cosine_similarity">코사인은</a> 기본 유사성 지표로, '(1 + 코사인(쿼리, 벡터)) / 2'로 계산됩니다. / 2'. 원본 벡터를 보존해야 하고 미리 정규화할 수 없는 경우가 아니라면, 코사인 유사성을 수행하는 가장 효율적인 방법은 모든 벡터를 단위 길이로 정규화하는 것입니다. 이렇게 하면 검색 중에 추가 벡터 길이 계산을 수행하지 않고 대신 'dot_product'를 사용할 수 있습니다.</p><p>동일한 파이프라인에서 <a href="https://huggingface.co/roberta-base-openai-detector">텍스트 분류 모델을</a> 포함하는 또 다른 추론 프로세서는 콘텐츠가 사람이 작성한 '진짜' 콘텐츠인지, 아니면 AI가 작성한 '가짜' 콘텐츠인지 감지하여 각 문서에 'openai-detector.predicted_value'를 추가합니다.</p><p>수집 파이프라인:</p>client.ingest.put_pipeline( 
    id="plagiarism-checker-pipeline",
    processors = [
    {
      "inference": { #for ml models - to infer against the data that is being ingested in the pipeline
        "model_id": "roberta-base-openai-detector", #text classification model id
        "target_field": "openai-detector", # Target field for the inference results
        "field_map": { #Maps the document field names to the known field names of the model.
        "abstract": "text_field" # Field matching our configured trained model input. 
        }
      }
    },
    {
      "inference": {
        "model_id": "sentence-transformers__all-mpnet-base-v2", #text embedding model id
        "target_field": "abstract_vector", # Target field for the inference results
        "field_map": {
        "abstract": "text_field" # Field matching our configured trained model input. Typically for NLP models, the field name is text_field.
        }
      }
    }
    
  ]
)
<p>쿼리 시, 동일한 텍스트 임베딩 모델이 'query_vector_builder' 객체에서 쿼리 'model_text'의 벡터 표현을 생성하는 데도 사용됩니다.</p><p>k-근접 이웃(kNN) 검색은 유사성 메트릭으로 측정한 쿼리 벡터에 가장 가까운 k개의 벡터를 찾습니다.</p><p>각 문서의 _점수는 유사성에서 파생되며, 점수가 높을수록 순위가 높아집니다. 이는 문서가 의미론적으로 더 유사하다는 것을 의미합니다. &gt; 0.9점일 경우 '높은 유사성', &lt; 0.7점일 경우 '낮은 유사성', 그 외에는 '보통 유사성'으로 간주합니다. 사용 사례에 따라 표절로 인정되는 _점수 수준을 결정하기 위해 다양한 임계값을 유연하게 설정할 수 있습니다.</p><p>또한 텍스트 분류를 수행하여 텍스트 쿼리에서 AI가 생성한 요소도 확인합니다.</p><p>쿼리:</p>from elasticsearch import Elasticsearch
from elasticsearch.client import MlClient

#duplicated text - direct plagiarism test

model_text = 'Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at http://hucvl.github.io/recipeqa.'

response = client.search(index='plagiarism-checker', size=1,
    knn={
        "field": "abstract_vector.predicted_value",
        "k": 9,
        "num_candidates": 974,
        "query_vector_builder": { #The 'all-mpnet-base-v2' model is also employed to generate the vector representation of the query in a 'query_vector_builder' object.
            "text_embedding": {
                "model_id": "sentence-transformers__all-mpnet-base-v2",
                "model_text": model_text
            }
        }
    }
)

for hit in response['hits']['hits']:
    score = hit['_score']
    title = hit['_source']['title']
    abstract = hit['_source']['abstract']
    openai = hit['_source']['openai-detector']['predicted_value']
    url = hit['_source']['url']

    if score &gt; 0.9:
        print(f"\nHigh similarity detected! This might be plagiarism.")
        print(f"\nMost similar document: '{title}'\n\nAbstract: {abstract}\n\nurl: {url}\n\nScore:{score}\n\n")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

    elif score &lt; 0.7:
        print(f"\nLow similarity detected. This might not be plagiarism.")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

    else:
        print(f"\nModerate similarity detected.")
        print(f"\nMost similar document: '{title}'\n\nAbstract: {abstract}\n\nurl: {url}\n\nScore:{score}\n\n")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

ml_client = MlClient(client)

model_id = 'roberta-base-openai-detector' #open ai text classification model

document = [
    {
        "text_field": model_text
    }
]

ml_response = ml_client.infer_trained_model(model_id=model_id, docs=document)

predicted_value = ml_response['inference_results'][0]['predicted_value']

if predicted_value == 'Fake':
    print("\nNote: The text query you entered may have been generated by AI.\n")
<p>출력:</p>High similarity detected! This might be plagiarism.

Most similar document: 'RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes'

Abstract: Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at[ http://hucvl.github.io/recipeqa](http://hucvl.github.io/recipeqa).

url:[http://aclweb.org/anthology/D18-1166](http://aclweb.org/anthology/D18-1166)

Score:1.0
<p>이 예에서는 데이터 세트의 '추상' 값 중 하나를 텍스트 쿼리 'model_text'로 사용한 후 표절이 식별되었습니다. 유사도 점수는 1.0으로 높은 수준의 유사성, 즉 <strong>직접적인 표절을</strong> 나타냅니다. 벡터화된 쿼리와 문서는 예상대로 AI가 생성한 콘텐츠로 인식되지 않았습니다.</p><p>쿼리:</p>#similar text - paraphrase plagiarism test 

model_text = 'Comprehending and deducing information from culinary instructions represents a promising avenue for research aimed at empowering artificial intelligence to decipher step-by-step text. In this study, we present CuisineInquiry, a database for the multifaceted understanding of cooking guidelines. It encompasses a substantial number of informative recipes featuring various elements such as headings, explanations, and a matched assortment of visuals. Utilizing an extensive set of automatically crafted question-answer pairings, we formulate a series of tasks focusing on understanding and logic that necessitate a combined interpretation of visuals and written content. This involves capturing the sequential progression of events and extracting meaning from procedural expertise. Our initial findings suggest that CuisineInquiry is poised to function as a demanding experimental platform.'
<p>출력:</p>High similarity detected! This might be plagiarism.

Most similar document: 'RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes'

Abstract: Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at[ http://hucvl.github.io/recipeqa](http://hucvl.github.io/recipeqa).

url:[http://aclweb.org/anthology/D18-1166](http://aclweb.org/anthology/D18-1166)

Score:0.9302529

Note: The text query you entered may have been generated by AI.
<p>텍스트 쿼리 'model_text'를 유사한 단어의 반복을 최소화하면서 동일한 메시지를 전달하는 AI 생성 텍스트로 업데이트한 결과, 감지된 유사도는 여전히 높았지만 1.0이 아닌 0.9302529로 <strong>표절로</strong> 판정되었습니다. AI가 생성한 이 쿼리도 감지될 것으로 예상했습니다.</p><p>마지막으로, 이 문서 중 하나의 초록이 아닌 Elasticsearch에 대한 텍스트 쿼리 'model_text'를 고려한 결과, 탐지된 유사도는 0.68991005 으로 고려 임계값에 따라 유사도가 낮은 것으로 나타났습니다.</p><p>쿼리:</p>#different text - not a plagiarism

model_text = 'Elasticsearch provides near real-time search and analytics for all types of data.'
<p>출력:</p>Low similarity detected. This might not be plagiarism.
<p>AI가 생성한 텍스트 쿼리에서 표절이 정확하게 식별되었지만, 의역과 직접 복사한 콘텐츠의 경우 표절 탐지를 위해서는 다양한 측면을 고려해야 합니다.</p><p>AI가 생성한 콘텐츠 감지의 맥락에서 가치 있는 기여를 하는 모델을 살펴봤습니다. 그러나 독립형 탐지에는 내재된 한계가 있으므로 정확도를 높이기 위해 다른 방법을 통합해야 한다는 점을 인식하는 것이 중요합니다.</p><p>텍스트 임베딩 모델 선택에 따른 가변성도 고려해야 할 사항입니다. 각기 다른 데이터 세트로 학습된 모델에 따라 유사성 수준이 달라지며, 이는 생성된 텍스트 임베딩의 중요성을 강조합니다.</p><p>마지막으로, 이 예제에서는 문서의 초록을 사용했습니다. 그러나 표절 탐지는 대용량 문서와 관련된 경우가 많기 때문에 텍스트 길이 문제를 해결하는 것이 필수적입니다. 텍스트가 모델의 토큰 한도를 초과하는 경우가 많으므로 임베딩을 구축하기 전에 청크로 분할해야 하는 경우가 많습니다. 이를 처리하는 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.11/knn-search.html#nested-knn-search">실용적인 접근 방식은</a> dense_vector와 함께 중첩된 구조를 활용하는 것입니다.</p><h2>결론</h2><p>이 블로그에서는 특히 의역 및 AI 생성 콘텐츠에서 표절을 탐지하는 데 따르는 어려움과 이를 위해 시맨틱 텍스트 유사성 및 텍스트 분류를 사용하는 방법에 대해 설명했습니다.</p><p>이러한 방법을 결합하여 AI가 생성한 콘텐츠, 직접 표절 및 의역 표절을 성공적으로 식별한 표절 탐지 사례를 제공했습니다.</p><p>주요 목표는 탐지를 간소화하는 필터링 시스템을 구축하는 것이었지만, 검증을 위해서는 여전히 사람의 평가가 필수적이었습니다.</p><p>의미론적 텍스트 유사도 및 NLP에 대해 자세히 알아보려면 다음 링크도 확인해 보세요:</p><ul><li><p><a href="https://www.elastic.co/what-is/semantic-search">시맨틱 검색이란 무엇인가요?</a></p></li><li><p><a href="https://www.elastic.co/what-is/natural-language-processing">자연어 처리(NLP)란 무엇인가요?</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch">Elasticsearch를 사용한 어휘 및 시맨틱 검색</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">수집 파이프라인과 중첩된 벡터를 통해 대용량 문서를 청크 처리하면 구절 검색이 쉬워집니다.</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-plagiarism-checker-with-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-plagiarism-checker-with-elasticsearch</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt68a5bc2434a9b03b/6a1711510e2e49a09641a22a/83e05cd4f81799fbb7b7950ed87600e825ec81e9-1024x1024.png" length="0" type="image/png"/>
    <pubDate>Tue, 19 Dec 2023 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch를 사용한 어휘 및 시맨틱 검색]]></title>
    <description><![CDATA[이 블로그에서는 어휘 및 시맨틱 검색을 중심으로 Elasticsearch를 사용해 정보를 검색하는 다양한 접근 방식을 살펴보겠습니다.]]></description>
    <content:encoded><![CDATA[<p>검색은 검색어 또는 복합 검색어를 기반으로 가장 관련성이 높은 정보를 찾는 과정이며, 관련 검색 결과는 이러한 검색어와 가장 잘 일치하는 문서입니다. 검색과 관련된 여러 가지 과제와 방법이 있지만, <strong>질문에 대한 최상의 답변을 찾는다는</strong> 궁극적인 목표는 동일합니다.</p><p>이 목표를 고려하여 이 블로그 게시물에서는 텍스트 검색에 특히 초점을 맞춘 <strong>어휘 검색과 의미론적 검색을</strong>중심으로 Elasticsearch를 사용하여 정보를 검색하는 다양한 접근 방식을 살펴보겠습니다.</p><h2>필수 구성 요소</h2><p>이를 위해 이커머스 상품 정보를 시뮬레이션하기 위해 생성된 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/products-ecommerce.json">데이터 세트에</a> 대한 다양한 검색 시나리오를 보여주는 Python 예제를 제공합니다.</p><p>이 데이터 세트에는 각각 설명이 포함된 2,500개 이상의 제품이 포함되어 있습니다. 이러한 제품은 아래와 같이 76개의 개별 제품 카테고리로 분류되며, 각 카테고리에는 다양한 수의 제품이 포함되어 있습니다:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt34415151c00ced2b/6a17d8710b0bed6c9bdd342c/4104466050f3024b6bcaf382da2a702650f62227-1440x708.png" alt="" /><p><em>트리맵 시각화 - category.keyword(제품 카테고리)의 상위 22개 값</em></p><p>설정에는 다음이 필요합니다:</p><ul><li><p>Python 3.6 이상</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/client/python-api/current/index.html">Elastic Python 클라이언트</a></p></li><li><p>Elastic 8.8 이상 배포, 8GB 메모리 머신 러닝 노드 사용</p></li><li><p>Elastic에 사전 로드되어 배포에 설치 및 시작되는 <a href="https://www.elastic.co/guide/en/machine-learning/8.9/ml-nlp-elser.html">Elastic 학습된 Sparse EncodeR</a> 모델</p></li></ul><p>저희는 Elastic Cloud를 사용할 예정이며, <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">무료 체험판을 사용할 수</a> 있습니다.</p><p>이 블로그 게시물에 제공된 검색어 외에도 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/ecommerce_dense_sparse_project.ipynb">Python 노트북에서</a> 다음 프로세스를 안내해 드립니다:</p><ul><li><p>Python 클라이언트를 사용하여 Elastic 배포에 연결 설정하기</p></li><li><p>텍스트 임베딩 모델을 Elasticsearch 클러스터에 로드합니다.</p></li><li><p>피처 벡터와 고밀도 벡터를 인덱싱하기 위한 매핑으로 인덱스를 생성합니다.</p></li><li><p>텍스트 삽입 및 텍스트 확장을 위한 추론 프로세서가 포함된 수집 파이프라인 만들기</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0e95406ece8f08d9/6a17d8731d1b83d32e93e2f4/54a9a490a3b0cf1b9c2228bee8eddd3f566bd435-1418x1102.png" alt="" /><h2>어휘 검색 - 희소 검색</h2><p>텍스트 쿼리를 기반으로 Elasticsearch에서 문서의 관련성 순위를 매기는 고전적인 방식은 <strong>어휘 검색을 위한</strong> 희소 모델 <a href="https://en.wikipedia.org/wiki/Okapi_BM25">인 BM25 모델의</a> Lucene 구현을 사용합니다. 이 방법은 텍스트 검색의 전통적인 접근 방식을 따르며, 정확한 용어 일치 항목을 찾습니다.</p><p>이 검색을 가능하게 하기 위해 Elasticsearch는 텍스트 분석을 수행하여 <strong>텍스트 필드</strong> 데이터를 검색 가능한 형식으로 변환합니다.</p><p><strong>텍스트 분석</strong> 은 검색을 위해 관련 토큰을 추출하는 프로세스를 관리하는 일련의 규칙인<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analyzer-anatomy.html"> 분석기에</a> 의해 수행됩니다. 분석기에는<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenizers.html"> 토큰라이저가</a> 정확히 하나만 있어야 합니다. 토큰화 도구는 아래 예시와 같이 문자 스트림을 수신하여 개별 토큰(일반적으로 개별 단어)으로 분할합니다:</p><h3>어휘 검색을 위한 문자열 토큰화</h3>#Performs text analysis on a string and returns the resulting tokens.

# Define the text to be analyzed
text = "Comfortable furniture for a large balcony"

# Define the analyze request
request_body = {
  "analyzer": "standard",
  "text": text
}

# Perform the analyze request
response = client.indices.analyze(analyzer=request_body["analyzer"], text=request_body["text"])

# Extract and display the analyzed tokens
tokens = [token["token"] for token in response["tokens"]]
print("Analyzed Tokens:", tokens)
<p>출력</p>Analyzed Tokens: ['comfortable', 'furniture', 'for', 'a', 'large', 'balcony']
<p>이 예에서는 영어 문법 기반 토큰화를 제공하기 때문에 대부분의 사용 사례에서 잘 작동하는 기본 분석기인 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-standard-analyzer.html">표준</a> 분석기를 사용하고 있습니다. 토큰화를 통해 개별 조건에 따른 매칭이 가능하지만, 각 토큰은 여전히 문자 그대로 매칭됩니다.</p><p>검색 환경을 맞춤 설정하고 싶다면 다른<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-analyzers.html"> 기본 제공 분석기를</a> 선택할 수 있습니다. 예를 들어, <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-stop-analyzer.html">중지 분석기를</a> 사용하도록 코드를 업데이트하면 중지 단어 제거를 지원하여 문자가 아닌 모든 문자에서 텍스트를 토큰으로 분리할 수 있습니다.</p>...
# Define the analyze request
request_body = {
  "analyzer": "stop",
  "text": text
}
...
<p>출력</p>Analyzed Tokens: ['comfortable', 'furniture', 'large', 'balcony']
<p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html">기본</a> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-charfilters.html">제공</a> 분석기가 사용자의 요구 사항을 충족하지 못하는 경우, 0개 이상의 문자 필터,<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenizers.html">토큰화</a> 도구 및 0개 이상의 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenfilters.html">토큰 필터를 적절히 조합하여 사용자 지정</a> 분석기를 만들 수 있습니다.</p>"analyzer":  {

  "my_analyzer": {

    "type": "custom", #For custom analyzers, use a type of custom or omit the type parameter.

    "tokenizer": "standard", #Built-in or customized tokenizer

    "filter": ["lowercase", "synonym"] #Built-in or customized token filters
  }
}
<p>토큰화기와 토큰 필터를 결합한 위의 예에서 텍스트는 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-synonym-tokenfilter.html#:~:text=Elasticsearch%20will%20use%20the%20token,applied%20to%20the%20synonym%20entries.">동의어 토큰</a> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lowercase-tokenfilter.html">필터에</a> 의해 처리되기 전에 소문자 필터에 의해 소문자로 처리됩니다.</p><h2>어휘 일치</h2><p><a href="https://www.elastic.co/blog/practical-bm25-part-2-the-bm25-algorithm-and-its-variables">BM25는</a> 용어의 빈도와 중요도에 따라 주어진 검색 쿼리에 대한 문서의 관련성을 측정합니다.</p><p>아래 코드는 <em>"전자상거래 검색"</em> 인덱스의 "설명" 필드 값과 검색 쿼리를 고려하여 최대 2개의 문서를 <em>검색하는</em> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-match-query.html">일치</a> 쿼리를 수행합니다.<strong>"</strong><em><strong>넓은 발코니를 위한 편안한 가구 ".</strong></em></p><p>이 쿼리와 일치하는 문서로 간주되는 기준을 세분화하면 정확도를 높일 수 있습니다. 그러나 보다 구체적인 결과를 얻으려면 변형에 대한 허용 오차가 낮아지는 대가가 따릅니다.</p># BM25

response = client.search(size=2,
index="ecommerce-search",
query= {
  "match": {
    "description" : {  
      "query": "Comfortable furniture for a large balcony",
      "analyzer": "stop"
    }
  }
}
)

hits = response['hits']['hits']

if not hits:
  print("No matches found")

else:
  for hit in hits:
    score = hit['_score']
    product = hit['_source']['product']
    category = hit['_source']['category']
    description = hit['_source']['description']
    print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>출력</p>Score: 15.607948
Product: Barbie Dreamhouse
Category: Toys
Description: is a classic Barbie playset with multiple rooms, furniture, a large balcony, a pool, and accessories. It allows kids to create their dream Barbie world.

Score: 9.137739
Product: Comfortable Rocking Chair
Category: Indoor Furniture
Description: enjoy relaxing moments with this comfortable rocking chair. Its smooth motion and cushioned seat make it an ideal piece of furniture for unwinding.
<p>결과물을 분석한 결과, 가장 연관성이 높은 결과는 "<em>장난감 " 카테고리의</em>" 바비 드림하우스 " 제품이며,<em>설명에 "</em><em>가구</em>", " 대형" 및 <em>"발코니</em>" 라는 용어가 포함되어 있어 연관성이 높으며, 이 제품은 설명에 검색어와 일치하는 3개의 용어가 포함된 유일한 제품이며, 설명에 <em>"발코니"</em> 라는 용어가 포함된 유일한 제품이기도 합니다.</p><p>두 번째로 관련성이 높은 제품은 "<em>실내 가구</em>" 로 분류된<em>" 편안한 흔들</em>의자 " 이며 설명에 "<em>편안한</em>" 및 "<em>가구 " 라는</em>용어가 포함되어 있습니다. 데이터 세트에서 이 검색 쿼리의 2개 이상의 용어와 일치하는 제품은 3개뿐이며, 이 제품은 그 중 하나입니다.</p><p><em>"105개 제품의 설명에는 '</em> 편안함" '이, 4개 카테고리의 4개 제품 설명에는 ' <em>"가구"</em> '가 표시됩니다: <em>장난감</em>, <em>실내 가구, 실외 가구 및 '개 및 고양이 용품 &amp; 장난감'.</em></p><p>보시다시피, 검색어와 가장 연관성이 높은 제품은 장난감이고 두 번째로 연관성이 높은 제품은 실내 가구입니다. 이러한 문서가 일치하는 이유를 알 수 있도록 점수 계산에 대한 자세한 정보를 원한다면 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-explain.html"><em>explain</em></a> __query 매개 변수를 true로 설정하면 됩니다.</p><p>두 결과 모두 가장 관련성이 높은 결과이지만, 이 데이터 세트의 문서 수와 용어 발생을 모두 고려할 때 "<em>넓은 발코니에 어울리는 편안한 가구</em>" 쿼리의 의도는 장난감과 실내 가구를 제외한 실제 넓은 발코니에 어울리는 가구를 검색하는 것입니다.</p><p>어휘 검색은 비교적 <strong>간단하고 빠르지만</strong>, 사용자의 의도와 쿼리를 알지 못하면 가능한 모든 용어와 동의어를 알 수 없기 때문에 한계가 있습니다. 자연어 사용에서 흔히 볼 수 있는 현상은 <strong>어휘 불일치입니다</strong>. <a href="https://dl.acm.org/doi/abs/10.1145/32206.32212">연구에</a> 따르면, 같은 분야의 전문가들이 같은 사물의 이름을 다르게 지을 확률은 평균적으로 <strong>80% %에</strong> 달한다고 합니다.</p><p>이러한 한계로 인해 의미론적 지식을 통합하는 다른 채점 모델을 찾게 되었습니다. 자연어처럼 순차적인 입력 토큰을 처리하는 데 탁월한 트랜스포머 기반 모델은 문서와 쿼리의 수학적 표현을 모두 고려하여 검색의 기본 의미를 포착합니다. 이를 통해 문맥을 인식하는 조밀한 텍스트 벡터 표현이 가능해져 관련 콘텐츠를 찾는 정교한 방법인 <strong>시맨틱 검색을</strong> 강화할 수 있습니다.</p><h2>시맨틱 검색 - 고밀도 검색</h2><p>이러한 맥락에서 데이터를 의미 있는 벡터 값으로 변환한 후, 데이터 집합에서 쿼리 벡터와 가장 유사한 벡터 표현을 찾기 위해 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html">k-최근접 이웃(kNN)</a> 검색 알고리즘을 사용합니다. Elasticsearch는 kNN 검색을 위해 두 가지 방법, 즉 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#exact-knn">정확한 무차별 kNN과</a> ANN이라고도 하는 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#approximate-knn">대략적인 kNN을</a> 지원합니다.</p><p>무차별 대입 방식은 정확한 결과를 보장하지만 대규모 데이터 세트에는 잘 확장되지 않습니다. 근사 kNN은 성능 향상을 위해 정확도를 일부 희생하여 가장 가까운 이웃을 효율적으로 찾습니다.</p><p>kNN 검색과 고밀도 벡터 인덱스에 대한 Lucene의 지원을 통해 Elasticsearch는 다양한 <a href="http://ann-benchmarks.com/">앤 벤치마크 데이터 세트에서</a> 강력한 검색 성능을 보여주는 계층적 탐색 가능한 작은 세계(HNSW) 알고리즘을 활용할 수 있습니다. 아래 예제 코드를 사용하여 Python에서 대략적인 kNN 검색을 수행할 수 있습니다.</p><h3>대략적인 kNN을 사용한 시맨틱 검색</h3># KNN - approximate kNN

response = client.search(index='ecommerce-search', size=2,
knn={
  "field": "description_vector.predicted_value",
  "k": 50, # Number of nearest neighbors to return as top hits.
#The optimal value of k is dependent on the data. It can vary in different scenarios.

  "num_candidates": 500, # Number of nearest neighbor candidates to consider per shard.

#Increasing num_candidates tends to improve the accuracy of the final k results.

  "query_vector_builder": { # Object indicating how to build a query_vector. kNN search enables you to perform semantic search by using a previously deployed text embedding model, the steps for this process are demonstrated in the Python notebook.
    "text_embedding": { 
      "model_id": "sentence-transformers__all-mpnet-base-v2", # Text embedding model id
      "model_text": "Comfortable furniture for a large balcony" # Query
    }
  }
}
)

for hit in response['hits']['hits']:
        
  score = hit['_score']
  product = hit['_source']['product']
  category = hit['_source']['category']
  description = hit['_source']['description']
  print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>이 코드 블록은 제품 데이터 세트의 " description" 필드가 포함된 것을 고려하여<em>Elasticsearch의 kNN을</em><em>사용하여 " 큰 발코니를 위한 편안한 가구 "</em> 의 벡터화된 쿼리(query_vector_build)와 유사한 설명을 가진 최대 두 개의 제품을 반환합니다.</p><p>제품 임베딩은 이전에 추론 프로세서가 포함된 수집 파이프라인에서 생성되었습니다. <em>"</em><a href="https://huggingface.co/sentence-transformers/all-mpnet-base-v2"><em>all-mpnet-base-v2</em></a><em>"</em> 텍스트 임베딩 모델을 포함하는 추론 프로세서를 사용하여 파이프라인에서 수집되는 데이터에 대해 추론하는 방식으로 생성되었습니다.</p><p>이 모델은 다음을 사용하여 사전 학습된 모델의 평가를 기반으로 선택되었습니다. <em>"</em><a href="https://github.com/UKPLab/sentence-transformers/blob/master/docs/package_reference/sentence_transformer/evaluation.md"><em>sentence_transformers.evaluation</em></a><em>"</em> 훈련 중 모델을 평가하는 데 다양한 클래스가 사용됩니다. "all-mpnet-base-v2" 모델은 <a href="https://www.sbert.net/docs/pretrained_models.html">문장-변환 순위에서</a> 가장 우수한 평균 성능을 보였으며, <a href="https://huggingface.co/spaces/mteb/leaderboard">대용량 텍스트 임베딩 벤치마크(MTEB)</a> 리더보드에서도 유리한 위치를 확보했습니다. 이 모델은<a href="https://huggingface.co/microsoft/mpnet-base"> Microsoft/MPnet-Base</a> 모델을 사전 학습하고 1B 문장 쌍 데이터 세트에서 미세 조정하여 문장을 768차원의 고밀도 벡터 공간에 매핑합니다.</p><p>또는 도메인별 데이터에 맞게 세밀하게 조정된 다른 모델을 사용할 수도 있습니다.</p><p>출력</p>Score: 0.79207325
Product: Patio Sofa Set with Ottoman
Category: Outdoor Furniture
Description: is a versatile and comfortable patio sofa set, including a sofa, ottoman, and coffee table, great for outdoor lounging.

Score: 0.7836937
Product: Patio Sofa Set with Canopy
Category: Outdoor Furniture
Description: is a luxurious and comfortable patio sofa set with a canopy, providing shade and style for outdoor lounging.
<p><em>출력은 선택한 모델,</em> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search-filter-example"><em>필터</em></a> <em>및</em> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#tune-approximate-knn-for-speed-accuracy"><em>대략적인 kNN 튜닝에</em></a>따라 달라질 수 있습니다<em>.</em></p><p><em>검색어에 " outdoor</em>" 라는 단어가 명시적으로 언급되지 않았음에도 불구하고 kNN 검색 결과는 모두 "<em>아웃도어</em>가구 " 카테고리에 속하며, 이는 문맥에서 의미 이해의 중요성을 강조합니다.</p><p>고밀도 벡터 검색은 몇 가지 장점이 있습니다:</p><ul><li><p>시맨틱 검색 활성화</p></li><li><p>대규모 데이터 세트를 처리할 수 있는 확장성</p></li><li><p>다양한 데이터 유형을 처리할 수 있는 유연성</p></li></ul><p>하지만 <strong>밀도 높은 벡터 검색에는 고유한 문제도</strong> 있습니다:</p><ul><li><p>사용 사례에 적합한 임베딩 모델 선택하기</p></li><li><p>모델을 선택한 후에는 도메인별 데이터 세트에서 성능을 최적화하기 위해 모델을 미세 조정해야 할 수 있으며, 이 과정에는 도메인 전문가의 참여가 필요합니다.</p></li><li><p>또한 고차원 벡터를 인덱싱하는 데는 계산 비용이 많이 들 수 있습니다.</p></li></ul><h2>시맨틱 검색 - 학습된 희소 검색</h2><p>시맨틱 검색을 수행하는 또 다른 방법인 학습된 희소 검색에 대해 알아보겠습니다.</p><p>스파스 모델로서, 수십 년에 걸친 최적화의 혜택을 누리고 있는 Elasticsearch의 Lucene 기반 반전 인덱스를 활용합니다. 그러나 이 접근 방식은 단순히 BM25와 같은 어휘 채점 기능을 사용하여 동의어를 추가하는 것 이상의 의미를 갖습니다. 대신, 더 심층적인 언어 규모 지식을 사용하여 학습된 연관성을 통합하여 관련성을 최적화합니다.</p><p>검색 쿼리를 확장하여 원래 쿼리에 없는 관련 용어를 포함시킴으로써, 아래 예에서 볼 수 있듯이 <a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-elser.html">Elastic 학습형 스파스 인코더는</a> <strong>스파스 벡터 임베딩을 개선합니다</strong>.</p><h3>Elastic 학습형 스파스 인코더를 사용한 스파스 벡터 검색</h3># Elastic Learned Sparse Encoder

response = client.search(index='ecommerce-search', size=2,
query={
  "text_expansion": {
    "ml.tokens": {
      "model_id":"elser_model",
      "model_text":"Comfortable furniture for a large balcony"                
    }
  }
}
)

for hit in response['hits']['hits']:

  score = hit['_score']
  product = hit['_source']['product']
  category = hit['_source']['category']
  description = hit['_source']['description']
  print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>출력</p>Score: 14.405318
Product: Garden Lounge Set with Side Table
Category: Garden Furniture
Description: is a comfortable and stylish garden lounge set, including a sofa, chairs, and a side table for outdoor relaxation.

Score: 14.281318
Product: Rattan Patio Conversation Set
Category: Outdoor Furniture
Description: is a stylish and comfortable outdoor furniture set, including a sofa, two chairs, and a coffee table, all made of durable rattan material.
<p>이 경우 결과에는 "<em>야외 가구</em>" 와 매우 유사한 제품을 제공하는 "<em>정원 가구 " 카테고리가 포함됩니다.</em></p><p>분석하여 "ml.tokens", 학습된 희소 검색이 생성한 토큰이 포함된 "rank_features" 필드를 보면, 생성된 다양한 토큰 중 검색 쿼리의 일부가 아니지만 "<em>relax</em>" (편안한), "<em>sofa</em>" (가구), "<em>outdoor</em>" (발코니)와 같이 의미상 여전히 연관성이 있는 용어가 있다는 것을 알 수 있습니다.</p><p>아래 이미지는 용어 확장이 있는 경우와 없는 경우 모두 쿼리와 함께 이러한 용어 중 일부를 강조 표시합니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7e869985347b19a7/6a17d875e31791dc2c2d56a1/dd86607fce6137d843a3ec002390eaa988b432f9-1440x502.png" alt="" /><p>이 모델은 문맥 인식 검색을 제공하고 어휘 불일치 문제를 완화하는 동시에 해석 가능한 결과를 제공하는 데 도움이 됩니다. 도메인별 재학습이 적용되지 않은 경우에도 고밀도 벡터 모델을 능가하는 성능을 발휘할 수 있습니다.</p><h2>하이브리드 검색: 어휘 검색과 시맨틱 검색을 결합하여 관련성 높은 결과 제공</h2><p>검색에 관한 한 만능 솔루션은 존재하지 않습니다. 이러한 검색 방법에는 각각 장점도 있지만 문제점도 있습니다. 사용 사례에 따라 가장 적합한 옵션이 변경될 수 있습니다. 종종 검색 방법 간에 상호 보완적으로 최상의 결과를 얻을 수 있습니다. 따라서 관련성을 높이기 위해 각 방법의 강점을 결합하는 방법을 살펴볼 것입니다.</p><p><strong>하이브리드 검색을</strong> 구현하는 방법에는 선형 조합, 각 점수에 가중치를 부여하는 방법, 가중치를 지정할 필요가 없는 상호 순위 융합(RRF) 등 여러 가지가 있습니다.</p><h3>Elasticsearch: 어휘 검색과 시맨틱 검색의 두 가지 장점을 모두 갖춘 최고의 솔루션</h3># BM25 + Elastic Learned Sparse Encoder (Linear Combination)

response = client.search(index='ecommerce-search', size=2,

query= {
  "bool": {
    "should": [
    {
      "match": {
        "description" : {  
          "query": "A dining table and comfortable chairs for a large balcony",
          "boost": 1
        }
      }
    },                   
    {
      "text_expansion": {
        "ml.tokens": {
          "model_id": "elser_model",
          "model_text": "A dining table and comfortable chairs for a large balcony",
          "boost": 1
        }
      }
     }
    ]
  }
}
)

# The boost value is 1 for the text expansion and match query. This means that the relevance score of the results of these queries are not boosted. You can specify a boost value to give a weight to each score in the sum. The scores will be calculated as: score = boost value * match_score + boost value * text_expansion_score

for hit in response['hits']['hits']:

  score = hit['_score']
  product = hit['_source']['product']
  category = hit['_source']['category']
  description = hit['_source']['description']
  print(f"\nScore: {score}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p>이 코드에서는 "<em>큰 발코니를 위한 식탁과 편안한 의자</em>" 라는 값을 갖는 두 개의 쿼리를 사용하여 하이브리드 검색을 수행했습니다. "<em>가구</em>" 을 검색어로 사용하는 대신 찾고 있는 내용을 지정하고 있으며, 두 검색 모두 동일한 필드 값인 "설명" 을 고려하고 있습니다. 순위는 BM25와 ELSER 점수에 동일한 가중치를 부여한 선형 조합으로 결정됩니다.</p><p>출력</p>Score: 31.628141
Product: Garden Dining Set with Swivel Rockers
Category: Garden Furniture
Description: is a functional and comfortable garden dining set, including a table and chairs with swivel rockers for easy movement.

Score: 31.334227
Product: Garden Dining Set with Swivel Chairs
Category: Garden Furniture
Description: is a functional and comfortable garden dining set, including a table and chairs with swivel seats for convenience.
<p>아래 코드에서는 쿼리에 동일한 값을 사용하되, 상호 순위 융합 방법을 사용하여 BM25(쿼리 파라미터)와 kNN(knn 파라미터)의 점수를 결합하여 문서를 결합하고 순위를 매깁니다.</p># BM25 + KNN (RRF)

response = client.search(index='ecommerce-search', size=2,
query={
  "bool": {
    "should": [
    {
      "match": {
        "description": {
        "query": "A dining table and comfortable chairs for a large balcony"
        }
      }
    }
    ]
  }
},
knn={
  "field": "description_vector.predicted_value",
  "k": 50,
  "num_candidates": 500,
  "query_vector_builder": {
    "text_embedding": {
      "model_id": "sentence-transformers__all-mpnet-base-v2",
      "model_text": "A dining table and comfortable chairs for a large balcony"
    }
  }
},
rank={
  "rrf": { # Reciprocal rank fusion
    "window_size": 50, # This value determines the size of the individual result sets per query.
    "rank_constant": 20 # This value determines how much influence documents in individual result sets per query have over the final ranked result set.
  }
}
)

for hit in response['hits']['hits']:
        
  rank = hit['_rank']
  category = hit['_source']['category']
  product = hit['_source']['product']
  description = hit['_source']['description']
  print(f"\nRank: {rank}\nProduct: {product}\nCategory: {category}\nDescription: {description}\n")
<p><em>RRF 기능은 기술 미리보기 중입니다. 구문은 GA 이전에 변경될 가능성이 높습니다.</em></p><p>출력</p>Rank: 1
Product: Patio Dining Set with Bench
Category: Outdoor Furniture
Description: is a spacious and functional patio dining set, including a dining table, chairs, and a bench for additional seating.

Rank: 2
Product: Garden Dining Set with Swivel Chairs
Category: Garden Furniture
Description: is a functional and comfortable garden dining set, including a table and chairs with swivel seats for convenience.
<p>여기에서도 다양한 필드와 값을 사용할 수 있으며, 이러한 예제 중 일부는 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/ecommerce_dense_sparse_project.ipynb">Python 노트북에서</a> 사용할 수 있습니다.</p><p>보시다시피, Elasticsearch를 사용하면 기존의 어휘 검색과 벡터 검색(희소 또는 고밀도)의 두 가지 장점을 모두 활용하여 목표에 <strong>도달하고 질문에 대한 최상의 답을 찾을</strong>수 있습니다 .</p><p>여기에 언급된 접근 방식에 대해 계속 알아보고 싶다면 다음 블로그가 유용할 수 있습니다:</p><ul><li><p><a href="https://www.elastic.co/blog/improving-information-retrieval-elastic-stack-hybrid">Elastic Stack에서 정보 검색 개선: 하이브리드 검색</a></p></li><li><p><a href="https://www.elastic.co/blog/vector-search-elasticsearch-rationale">Elasticsearch의 벡터 검색: 설계의 근거</a></p></li><li><p><a href="https://www.elastic.co/blog/lexical-ai-powered-search-elastic-vector-database">Elastic의 벡터 데이터베이스로 어휘 및 AI 기반 검색을 최대한 활용하는 방법</a></p></li><li><p><a href="https://www.elastic.co/blog/may-2023-launch-sparse-encoder-ai-model">Elastic 학습형 스파스 인코더를 소개합니다: 시맨틱 검색을 위한 Elastic의 AI 모델</a></p></li><li><p><a href="https://www.elastic.co/blog/may-2023-launch-information-retrieval-elasticsearch-ai-model">Elastic Stack에서 정보 검색 개선: 새로운 검색 모델인 Elastic 학습형 스파스 인코더를 소개합니다.</a></p></li></ul><p>Elasticsearch는 벡터 검색을 구축하는 데 필요한 모든 도구와 함께 벡터 데이터베이스를 제공합니다:</p><ul><li><p>Elasticsearch <a href="https://www.elastic.co/elasticsearch/vector-database">벡터 데이터베이스</a></p></li><li><p>Elastic의 <a href="https://www.elastic.co/enterprise-search/vector-search">벡터 검색</a> 사용 사례</p></li></ul><h2>결론</h2><p>이 블로그 게시물에서는 특히 텍스트, 어휘 및 의미 검색에 초점을 맞춰 Elasticsearch를 사용해 정보를 검색하는 다양한 접근 방식을 살펴봤습니다. 이를 보여드리기 위해 이커머스 제품 정보가 포함된 데이터 세트를 사용하여 다양한 검색 시나리오를 보여주는 Python 예제를 제공했습니다.</p><p>BM25로 기존 어휘 검색을 검토하고 어휘 불일치 등의 장점과 문제점에 대해 논의했습니다. 우리는 이 문제를 극복하기 위해 시맨틱 지식을 통합하는 것이 중요하다고 강조했습니다. 또한 시맨틱 검색을 가능하게 하는 고밀도 벡터 검색에 대해 논의하고, 고차원 벡터를 색인화할 때의 계산 비용 등 이 검색 방법과 관련된 문제점에 대해서도 다뤘습니다.</p><p>반면에 스파스 벡터는 압축률이 매우 높다고 언급했습니다. 따라서 원래 쿼리에 없는 관련 용어를 포함하도록 검색 쿼리를 확장하는 Elastic의 학습된 스파스 인코더에 대해 설명했습니다.</p><p>검색에 있어 만능 솔루션은 존재하지 않습니다. 각 검색 방법에는 장단점이 있습니다. 따라서 하이브리드 검색의 개념에 대해서도 논의했습니다.</p><p>보시다시피, Elasticsearch를 사용하면 기존의 어휘 검색과 벡터 검색이라는 두 가지 장점을 모두 누릴 수 있습니다!</p><p>시작할 준비가 되셨나요? 사용 가능한 <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/lexical-and-semantic-search-with-elasticsearch/ecommerce_dense_sparse_project.ipynb">Python 노트북을</a> 확인하고 <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">Elastic Cloud 무료 체험판을</a> 시작하세요.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <category><![CDATA[Python]]></category>
    <category><![CDATA[쿼리 언어]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbd7e961f594005e7/6a17d80f033c8d981f6bb009/d240bfef29e9d432069059b312dd044eb76eec6c-1440x840.png" length="0" type="image/png"/>
    <pubDate>Tue, 03 Oct 2023 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch에서 NLP와 벡터 검색으로 챗봇 기능 강화하기]]></title>
    <description><![CDATA[벡터 검색과 NLP가 어떻게 챗봇 기능을 향상시키는지 살펴보고 Elasticsearch가 그 과정을 어떻게 촉진하는지 알아보세요.]]></description>
    <content:encoded><![CDATA[<p>대화형 인터페이스는 한동안 사용되어 왔으며 고객 서비스, 정보 검색, 작업 자동화 등 다양한 작업을 지원하는 수단으로 점점 더 인기를 얻고 있습니다. 일반적으로 음성 어시스턴트나 메시징 앱을 통해 액세스하는 이러한 인터페이스는 사용자가 보다 효율적으로 쿼리를 해결할 수 있도록 사람의 대화를 시뮬레이션합니다.</p><p>기술이 발전함에 따라 챗봇은 사용자에게 개인화된 경험을 제공하면서 더 복잡한 작업을 신속하게 처리하는 데 사용됩니다. 자연어 처리(NLP)를 통해 챗봇은 사용자의 언어를 처리하고 메시지의 의도를 파악하여 관련 정보를 추출할 수 있습니다. 예를 들어, 명명된 개체 인식은 텍스트의 주요 정보를 일련의 카테고리로 분류하여 추출합니다. 감정 분석은 감정 어조를 식별하고, 질문은 쿼리에 대한 '답변'을 식별합니다. NLP의 목표는 알고리즘이 인간의 언어를 처리하고 대량의 텍스트에서 관련 구절 찾기, 텍스트 요약, 새롭고 독창적인 콘텐츠 생성 등 기존에는 인간만이 할 수 있었던 작업을 수행할 수 있도록 하는 것입니다.</p><p>이러한 고급 NLP 기능은 <a href="https://www.elastic.co/what-is/vector-search">벡터 검색이라는</a> 기술을 기반으로 합니다. Elastic은 벡터 검색을 기본적으로 지원하여 정확하고 근사한 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search">kNN(k-최근접 이웃) 검색을</a> 수행하고, NLP를 지원하여 사용자 정의 또는 <a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-model-ref.html#ml-nlp-model-ref">타사 모델을</a> Elasticsearch에서 직접 사용할 수 있습니다.</p><p>이 블로그 게시물에서는 벡터 검색과 NLP가 어떻게 챗봇 기능을 향상시키는지 살펴보고 Elasticsearch가 그 과정을 어떻게 촉진하는지 보여드리겠습니다. 벡터 검색에 대한 간략한 개요부터 시작하겠습니다.</p><h2>벡터 검색</h2><p>인간은 문자로 된 언어의 의미와 문맥을 이해할 수 있지만, 기계는 그렇지 못합니다. 벡터가 필요한 이유입니다. 기계는 텍스트를 벡터 표현(텍스트의 의미를 수치로 표현한 것)으로 변환함으로써 이러한 한계를 극복할 수 있습니다. 기존 검색과 비교했을 때, 벡터는 빈도에 기반한 키워드와 어휘 검색에 의존하는 대신 숫자 값에 대해 정의된 연산을 사용하여 텍스트 데이터를 처리할 수 있습니다.</p><p>이를 통해 벡터 검색은 쿼리 벡터가 주어진 유사성을 나타내기 위해 "임베딩 공간" 의 거리를 사용하여 유사한 개념이나 컨텍스트를 공유하는 데이터를 찾을 수 있습니다. 데이터가 유사하면 해당 벡터도 비슷해집니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt53a615a8ba0ac931/6a17d795fbc5f8b257491910/08542abf8108aace288745b1aca8579b476ddc1b-1440x618.png" alt="" /><p>벡터 검색은 NLP 애플리케이션뿐만 아니라 이미지 및 동영상 처리 등 비정형 데이터가 관련된 다양한 영역에서도 활용되고 있습니다.</p><p>챗봇 플로우에서는 사용자의 쿼리에 대해 여러 가지 접근 방식이 있을 수 있으며, 그 결과 더 나은 사용자 경험을 위해 정보 검색을 개선할 수 있는 다양한 방법이 있습니다. 각 대안에는 고유한 장점과 단점이 있으므로 사용 가능한 데이터와 리소스, 교육 시간(해당되는 경우) 및 예상 정확도를 고려하는 것이 중요합니다. 다음 섹션에서는 질문 답변 NLP 모델에 대한 이러한 측면을 다룹니다.</p><h2>질문 답변</h2><p>질문 답변(QA) 모델은 자연어로 묻는 질문에 답변하도록 설계된 NLP 모델의 한 유형입니다. 사용자가 여러 리소스에서 답변을 추론해야 하는 질문이 있는데 문서에 이미 존재하는 목표 답변이 없는 경우, 제너레이티브 QA 모델이 유용할 수 있습니다. 그러나 이러한 모델은 계산 비용이 많이 들고 도메인 관련 학습에 많은 양의 데이터가 필요하기 때문에 도메인 외의 질문을 처리하는 데 특히 유용할 수 있지만 일부 상황에서는 실용성이 떨어질 수 있습니다.</p><p>반면에 사용자가 특정 주제에 대해 질문이 있고 실제 답변이 문서에 있는 경우 추출형 QA 모델을 사용할 수 있습니다. 이러한 모델은 소스 문서에서 직접 답변을 추출하여 투명하고 검증 가능한 결과를 제공하므로 간단하고 효율적인 방식으로 질문에 답하고자 하는 기업이나 조직에 보다 실용적인 옵션이 됩니다.</p><p>아래 예는 <a href="https://huggingface.co/deepset/minilm-uncased-squad2">Hugging Face에서 사용할 수</a> 있고 Elasticsearch에 배포된 사전 학습된 추출 QA 모델을 사용하여 주어진 컨텍스트에서 답을 추출하는 방법을 보여줍니다:</p>POST _ml/trained_models/deepset__minilm-uncased-squad2/deployment/_infer
{
    "docs": [{"text_field": "Canvas is a data visualization and presentation application within Kibana. With Canvas, live data can be pulled directly from Elasticsearch and combined with colors, images, text, and other customized options to create dynamic, multi-page displays."}],
    "inference_config": {"question_answering": {"question": "What is Kibana Canvas?"}}
}


{
  "predicted_value": "a data visualization and presentation application",
  "start_offset": 10,
  "end_offset": 59,
  "prediction_probability": 0.28304219431376443
}
<p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-deploy-models.html">학습된 모델을 배포합니다.</a></p><p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-ner-example.html#ex-ner-ingest">추론 수집 파이프라인에 모델을 추가합니다.</a></p><p>사용자 쿼리를 처리하고 정보를 검색하는 다양한 방법이 있으며, 비정형 데이터를 다룰 때는 여러 언어 모델과 데이터 소스를 사용하는 것이 효과적인 대안이 될 수 있습니다. 이를 설명하기 위해 선택한 문서에서 추출한 데이터를 고려한 답변으로 쿼리에 응답하는 데 사용되는 챗봇의 데이터 처리 예시가 있습니다.</p><h2>챗봇 데이터 처리: NLP 및 벡터 검색</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta418a9c54bb16cf9/6a17d7975772624b371bca43/c2d1a2f110e937b1d3e5df0d5caac3c906c98fb0-1440x748.png" alt="" /><p>위와 같이 챗봇의 데이터 처리는 크게 세 부분으로 나눌 수 있습니다:</p><ul><li><p><strong>벡터 처리:</strong> 이 파트에서는 문서를 벡터 표현으로 변환합니다.</p></li><li><p><strong>사용자 입력 처리:</strong> 사용자 쿼리에서 관련 정보를 추출하고 시맨틱 검색 및 하이브리드 검색을 수행하는 부분입니다.</p></li><li><p><strong>최적화:</strong> 이 부분에는 모니터링이 포함되며 챗봇의 안정성, 최적의 성능 및 우수한 사용자 경험을 보장하는 데 매우 중요합니다.</p></li></ul><h2>벡터 처리</h2><p><strong>처리</strong> 부분의 첫 번째 단계는 각 문서의 구성 요소를 결정한 다음 각 요소를 벡터 표현으로 변환하는 것입니다. 이러한 표현은 다양한 데이터 형식에 대해 생성할 수 있습니다.</p><p>사전 학습된 모델과 라이브러리를 포함하여 임베딩을 계산하는 데 사용할 수 있는 다양한 방법이 있습니다.</p><p>이러한 표현에 대한 검색 및 검색의 효과는 기존 데이터와 사용된 방법의 품질 및 관련성에 따라 달라진다는 점에 유의하세요.</p><p>벡터가 계산되면 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html">dense_vector</a> 필드 유형으로 Elasticsearch에 저장됩니다.</p>PUT &lt;target&gt;
{
  "mappings": {
    "properties": {
      "doc_part_vector": {
        "type": "dense_vector",
        "dims": 3
      },
      "doc_part" : {
        "type" : "keyword"
      }
    }
  }
}
<h2>챗봇 사용자 입력 처리</h2><p><strong>사용자</strong> 입장에서는 질문을 받은 후 질문을 진행하기 전에 가능한 모든 정보를 추출하는 것이 유용합니다. 이는 사용자의 의도를 파악하는 데 도움이 되며, 이 경우 이를 지원하기 위해 <a href="https://huggingface.co/dslim/bert-base-NER">NER(Named Entity Recognition) 모델을</a> 사용하고 있습니다. NER은 명명된 엔티티를 식별하고 미리 정의된 엔티티 카테고리로 분류하는 프로세스입니다.</p>POST _ml/trained_models/dslim__bert-base-ner/deployment/_infer
{
  "docs": { "text_field": "How many people work for Elastic?"}
}


{
  "predicted_value": "How many people work for [Elastic](ORG&amp;Elastic)?",
  "entities": [
    {
      "entity": "Elastic",
      "class_name": "ORG",
      "class_probability": 0.4993975435876747,
      "start_pos": 25,
      "end_pos": 32
    }
  ]
}
<p>필수 단계는 아니지만 구조화된 데이터나 위 또는 다른 NLP 모델 결과를 사용하여 사용자의 쿼리를 분류하면 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search-filter-example">필터를</a> 사용하여 kNN 검색을 제한할 수 있습니다. 이렇게 하면 처리해야 하는 데이터의 양을 줄여 성능과 정확도를 향상시킬 수 있습니다.</p>    "filter": {
      "term": {
        "org": "Elastic"
      }
    }
<h2>시맨틱 검색 및 하이브리드 검색</h2><p>프롬프트는 사용자 쿼리에서 시작되고 챗봇은 가변성과 모호성이 있는 인간의 언어를 처리해야 하므로 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#semantic-search">시맨틱 검색이</a> 매우 적합합니다. Elasticsearch에서는 쿼리 문자열과 <a href="https://huggingface.co/sentence-transformers/msmarco-MiniLM-L-12-v3">임베딩 모델의</a> ID를 쿼리 벡터 빌더 객체에 전달하여 한 단계로 시맨틱 검색을 수행할 수 있습니다. 그러면 쿼리를 벡터화하고 kNN <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-search.html">검색을</a> 수행하여 쿼리의 의미에 가장 가까운 상위 k개의 일치 항목을 검색합니다:</p>POST /&lt;target&gt;/_search
{
  "knn": {
    "field": "doc_part_vector",
    "k": 5,
    "num_candidates": 20,
    "query_vector_builder": {
      "text_embedding": {
        "model_id": "&lt;text-embedding-model-id&gt;",
        "model_text": "&lt;query_string&gt;"
      }
    }
  }
 }
<p><a href="https://www.elastic.co/guide/en/machine-learning/8.7/ml-nlp-text-emb-vector-search-example.html">엔드투엔드 예제: 텍스트 임베딩 모델을 배포하고 이를 시맨틱 검색에 사용하는 방법.</a> Elasticsearch는 <strong>희소 모델인</strong> Okapi BM25의 Lucene 구현을 사용해 텍스트 쿼리의 관련성 순위를 매기고, <strong>고밀도 모델은</strong> <strong>시맨틱 검색에</strong> 사용합니다. 텍스트 쿼리에서 <strong> 얻은,</strong> <strong>벡터</strong> 일치와 <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#_combine_approximate_knn_with_other_features"><strong>일치의 강점을</strong></a> <strong>모두 결합하려면 하이브리드 검색을 수행할</strong> 수 있습니다:</p>POST &lt;target&gt;/_search
{
  "query": {
          "match": {
            "content": {
              "query": "&lt;query_string&gt;"
            }
        }
  },
  "knn": {
    "field": "doc_part_vector",
    "query_vector_builder": {
      "text_embedding": {
    "model_id": "&lt;text-embedding-model-id&gt;",
     "model_text": "&lt;query_string&gt;"
      }
    },
    "filter": {
      "term": {
        "org": "Elastic"
      }
    }
  }
}
<h3>희소 모델과 고밀도 모델을 결합하면 최상의 결과를 얻을 수 있는 경우가 많습니다.</h3><p>희소 모델은 일반적으로 짧은 쿼리와 특정 용어에서 더 나은 성능을 발휘하는 반면, 고밀도 모델은 컨텍스트와 연관성을 활용합니다. 이러한 방법이 서로 어떻게 비교되고 보완되는지 자세히 알아보려면 검색을 위해 특별히 훈련된 두 가지 고밀도 모델과 BM25를 비교하여 벤치마킹해 보세요.</p><p>가장 관련성이 높은 결과는 일반적으로 사용자에게 제공된 첫 번째 답변일 수 있으며,_score는 반환된 문서의 <strong>관련성을</strong> 결정하는 데 사용되는 숫자입니다.</p><h2>챗봇 최적화</h2><p>챗봇의 사용자 경험, 성능 및 안정성을 개선하기 위해 하이브리드 스코어링을 적용하는 것 외에도 다음과 같은 접근 방식을 통합할 수 있습니다: <strong>감정 분석:</strong> 대화가 전개되는 동안 사용자의 댓글과 반응을 파악하기 위해 <a href="https://huggingface.co/distilbert-base-uncased-finetuned-sst-2-english">감성 분석 모델을</a> 통합할 수 있습니다:</p>POST _ml/trained_models/distilbert-base-uncased-finetuned-sst-2-english/deployment/_infer
{
  "docs": { "text_field": "That was not my question!"}
}


{
  "predicted_value": "NEGATIVE",
  "prediction_probability": 0.980080439016437
}
<p><a href="https://www.elastic.co/blog/chatgpt-elasticsearch-openai-meets-private-data"><strong>GPT의 기능</strong></a> <strong>:</strong> 전반적인 경험을 향상시키기 위한 대안으로, Elasticsearch의 검색 정확도와 OpenAI의 GPT 질문 답변 기능을 결합하여 <a href="https://platform.openai.com/docs/guides/chat">채팅 완료 API를</a> 활용하여 이러한 상위 k 문서를 컨텍스트로 고려하여 사용자 모델이 생성한 응답으로 돌아갈 수 있습니다. <em>프롬프트: "이 질문에 답하세요 &lt;user_question&gt; 이 문서만 사용 &lt;top_search_result&gt;"</em></p><p><strong>관찰 가능성:</strong> 챗봇의 성능을 보장하는 것은 매우 중요하며, 이를 위해서는 모니터링이 필수적인 요소입니다. 챗봇 상호작용을 캡처하는 로그 외에도 응답 시간, 지연 시간 및 기타 관련 챗봇 메트릭을 추적하는 것이 중요합니다. 이를 통해 패턴과 추세를 파악하고 이상 징후를 탐지할 수 있으며<a href="https://www.elastic.co/blog/monitor-openai-api-gpt-models-opentelemetry-elastic">, Elastic Observability</a> 도구를 사용하면 이러한 정보를 수집하고 분석할 수 있습니다.</p><h2>요약</h2><p>이 블로그 게시물에서는 NLP와 벡터 검색이 무엇인지 살펴보고 문서의 벡터 표현에서 추출한 데이터를 고려하여 사용자 쿼리에 응답하는 데 사용되는 챗봇의 예를 살펴봅니다.</p><p>앞서 살펴본 바와 같이 챗봇은 NLP와 벡터 검색을 사용하여 정형화된 타깃 데이터를 넘어서는 복잡한 작업을 수행할 수 있습니다. 여기에는 여러 데이터 소스 및 형식을 컨텍스트로 사용하여 특정 제품 또는 비즈니스 관련 쿼리에 대한 추천 및 답변을 제공하는 동시에 개인화된 사용자 경험을 제공하는 것도 포함됩니다.</p><p>사용 사례는 고객의 문의를 지원하는 고객 서비스 제공부터 단계별 안내, 권장 사항 제안 또는 작업 자동화를 통해 개발자의 문의를 돕는 것까지 다양합니다. 목표와 기존 데이터에 따라 다른 모델과 방법을 활용하여 더 나은 결과를 달성하고 전반적인 사용자 경험을 개선할 수도 있습니다.</p><p>다음은 유용할 수 있는 주제에 대한 몇 가지 링크입니다:</p><ol><li><p><a href="https://www.elastic.co/blog/how-to-deploy-natural-language-processing-nlp-getting-started">자연어 처리(NLP)를 배포하는 방법: 시작하기</a></p></li><li><p><a href="https://www.elastic.co/blog/overview-image-similarity-search-in-elastic">Elasticsearch의 이미지 유사도 검색 개요</a></p></li><li><p><a href="https://www.elastic.co/blog/chatgpt-elasticsearch-openai-meets-private-data">ChatGPT와 Elasticsearch: OpenAI와 개인 데이터의 만남</a></p></li><li><p><a href="https://www.elastic.co/blog/monitor-openai-api-gpt-models-opentelemetry-elastic">OpenTelemetry와 Elastic을 통한 OpenAI API 및 GPT 모델 모니터링</a></p></li><li><p><a href="https://www.elastic.co/blog/why-technology-leaders-need-vector-search">IT 리더가 검색 환경을 개선하기 위해 벡터 검색이 필요한 5가지 이유</a></p></li></ol><p>Elasticsearch에 NLP와 기본 벡터 검색을 통합하면 속도, 확장성, 검색 기능을 활용하여 정형 또는 비정형 데이터에 관계없이 대량의 데이터를 처리할 수 있는 매우 효율적이고 효과적인 챗봇을 만들 수 있습니다.</p><p>시작할 준비가 되셨나요? <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">Elastic Cloud 무료 체험판을</a> 시작하세요.</p><p><em>이 블로그 게시물에서는 해당 소유자가 소유하고 운영하는 타사 생성 AI 도구를 사용했거나 참조했을 수 있습니다. Elastic은 타사 도구에 대한 어떠한 통제권도 없으며 해당 도구의 콘텐츠, 운영 또는 사용에 대해 어떠한 책임이나 의무도 지지 않으며 그러한 도구의 사용으로 인해 발생할 수 있는 손실이나 손해에 대해서도 책임을 지지 않습니다. 개인 정보, 민감한 정보 또는 기밀 정보가 포함된 AI 도구를 사용할 때는 주의를 기울여 주세요. 제출하는 모든 데이터는 AI 학습 또는 기타 목적으로 사용될 수 있습니다. 회원님이 제공한 정보가 안전하게 보호되거나 기밀로 유지된다는 보장은 없습니다. 사용하기 전에 생성 AI 도구의 개인정보 보호 관행과 이용 약관을 숙지해야 합니다.</em></p><p><em>Elastic, Elasticsearch 및 관련 상표는 미국 및 기타 국가에서 Elasticsearch N.V.의 상표, 로고 또는 등록 상표입니다. 기타 모든 회사 및 제품명은 해당 소유자의 상표, 로고 또는 등록 상표입니다.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/enhancing-chatbot-capabilities-with-nlp-and-vector-search-in-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/enhancing-chatbot-capabilities-with-nlp-and-vector-search-in-elasticsearch</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c545fc80b6d79d6/6a170214839dfad776dcfd6f/d968e646240cd3ef7c79b5124d562a5f951d812b-1440x840.png" length="0" type="image/png"/>
    <pubDate>Wed, 21 Jun 2023 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>