<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Quynh Nguyen - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Quynh Nguyen - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/kr/search-labs/author/quynh-nguyen</link>
    </image>
    <link>https://www.elastic.co/kr/search-labs/author/quynh-nguyen</link>
    <atom:link href="https://www.elastic.co/kr/search-labs/rss/author/quynh-nguyen.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[kr]]></language>
    <lastBuildDate>Tue, 29 Sep 2026 00:42:01 GMT</lastBuildDate>
  <item>
    <title><![CDATA[하이브리드 검색 재랭킹을 통한 다국어 임베딩 모델 관련성 향상]]></title>
    <description><![CDATA[Elasticsearch에서 Cohere의 재랭커와 하이브리드 검색을 사용해 E5 다국어 임베딩 모델 검색 결과의 정확도를 개선하는 방법을 알아보세요.]]></description>
    <content:encoded><![CDATA[<h2>소개</h2><p><a href="https://www.elastic.co/search-labs/blog/multilingual-embedding-model-deployment-elasticsearch">이 시리즈의 마지막 파트에서는</a> Elastic의 사전 학습된 E5 모델(그리고 Hugging Face의 다른 다국어 텍스트 임베딩 모델)을 배포하는 과정을 살펴보고, Elasticsearch와 Kibana를 사용해 텍스트 데이터에서 고밀도 벡터 임베딩을 생성하는 방법에 대해 알아보았습니다. 이 블로그에서는 이러한 임베딩의 결과를 살펴보고 다국어 모델을 활용할 때 얻을 수 있는 중요한 이점을 강조합니다.</p><p>이제 색인 <code>coco_multilingual</code> 을 만들었으므로 검색을 수행하면 참조할 수 있도록 'en' 필드가 있는 여러 언어로 된 문서가 표시됩니다:</p># GET coco_multilingual/_search
    {
       "_index": "coco_multilingual",
       "_id": "WAiXQJYBgf6odR9bLohZ",
       "_score": 1,
       "_source": {
         "description": "Ein Parkmeßgerät auf einer Straße mit Autos",
         "en": "A row of parked cars sitting next to parking meters.",
         "language": "de",
         "vector_description": {...}
       }
     },
     . . .<h2>영어로 검색 수행하기</h2><p>영어로 검색을 수행해보고 얼마나 잘 검색되는지 확인해 보겠습니다:</p>GET coco_multi/_search
{
"size": 10,
"_source": [
  "description", "language", "en"
],
"knn": {
  "field": "vector_description.predicted_value",
  "k": 10,
  "num_candidates": 100,
  "query_vector_builder": {
    "text_embedding": {
      "model_id": ".multilingual-e5-small_linux-x86_64_search",
      "model_text": "query: kitty"
    }
  }
}
}{
       "_index": "coco_multi",
       "_id": "JQiXQJYBgf6odR9b6Yz0",
       "_score": 0.9334303,
       "_source": {
         "description": "Eine Katze, die in einem kleinen, gepackten Koffer sitzt.",
         "en": "A brown and white cat is in a suitcase.",
         "language": "de"
       }
     },
      {
       "_index": "coco_multi",
       "_id": "3AiXQJYBgf6odR9bFod6",
       "_score": 0.9281012,
       "_source": {
         "description": "Una bambina che tiene un gattino vicino a una recinzione blu.",
         "en": "A little girl holding a kitten next to a blue fence.",
         "language": "it"
       }
     },
     . . .<p>이 쿼리는 놀라울 정도로 단순해 보이지만, 내부적으로는 모든 언어의 모든 문서에서 'kitty'라는 단어가 포함된 숫자를 검색하고 있습니다. 그리고 벡터 검색을 수행하기 때문에 'kitty'와 관련이 있을 수 있는 모든 단어를 의미론적으로 검색할 수 있습니다: "고양이", "새끼 고양이", "고양이", "가토"(이탈리아어), "메오"(베트남어), 고양이(한국어), 猫(중국어) 등이 있습니다. 그 결과, 검색어가 영어로 되어 있어도 다른 모든 언어로 된 콘텐츠도 검색할 수 있습니다. 예를 들어, 고양이(<code>ying on something</code> )를 검색하면 이탈리아어, 네덜란드어 또는 베트남어로 된 문서도 표시됩니다. 효율성에 대해 이야기해 보세요!</p><h2>다른 언어로 된 콘텐츠 검색 수행하기</h2>GET coco_multi/_search
{  
 "size": 100,
 "_source": [
   "description", "language", "en"
 ],
 "knn": {
   "field": "vector_description.predicted_value",
   "k": 50,
   "num_candidates": 1000,
   "query_vector_builder": {
     "text_embedding": {
       "model_id": ".multilingual-e5-small_linux-x86_64_search",
       "model_text": "query: kitty lying on something"
     }
   }
 }
}{
 "description": "A black kitten lays on her side beside remote controls.",
 "en": "A black kitten lays on her side beside remote controls.",
 "language": "en"
},
{
 "description": "un gattino sdraiato su un letto accanto ad alcuni telefoni ",
 "en": "A black kitten lays on her side beside remote controls.",
 "language": "it"
},
{
 "description": "eine Katze legt sich auf ein ausgestopftes Tier",
 "en": "a cat lays down on a stuffed animal",
 "language": "de"
},
{
 "description": "Một chú mèo con màu đen nằm nghiêng bên cạnh điều khiển từ xa.",
 "en": "A black kitten lays on her side beside remote controls.",
 "language": "vi"
}
. . .<p>마찬가지로 한국어로 '고양이'로 키워드 검색을 수행하면 의미 있는 결과를 얻을 수 있습니다. 여기서 놀라운 점은 이 색인에는 한국어로 된 문서가 하나도 없다는 것입니다!</p>GET coco_multi/_search
{
 "size": 100,
 "_source": [
   "description", "language", "en"
 ],
 "knn": {
   "field": "vector_description.predicted_value",
   "k": 50,
   "num_candidates": 1000,
   "query_vector_builder": {
     "text_embedding": {
       "model_id": ".multilingual-e5-small_linux-x86_64_search",
       "model_text": "query: 고양이"
     }
   }
 }
} {
       {
         "description": "eine Katze legt sich auf ein ausgestopftes Tier",
         "en": "a cat lays down on a stuffed animal",
         "language": "de"
       }
     },
     {
       {
         "description": "Một con chó và con mèo đang ngủ với nhau trên một chiếc ghế dài màu cam.",
         "en": "A dog and cat lying  together on an orange couch. ",
         "language": "vi"
       }
     },<p>임베딩 모델은 공유 의미 공간에서 의미를 나타내므로 색인된 캡션과 다른 언어로 쿼리해도 관련 이미지를 검색할 수 있습니다.</p><h2>하이브리드 검색 및 재랭킹으로 관련성 높은 검색 결과 얻기</h2><p>예상대로 관련 결과가 나타나서 기쁘게 생각합니다. 하지만 이커머스나 가장 적합한 상위 5~10개의 결과로 범위를 좁혀야 하는 RAG 애플리케이션과 같은 실제 환경에서는 재랭크 모델을 사용하여 가장 관련성이 높은 결과의 우선 순위를 지정할 수 있습니다.</p><p>여기서 베트남어로 "고양이는 무슨 색인가요?"라고 묻는 쿼리를 수행하면 많은 결과가 나오지만 상위 1, 2위가 가장 관련성이 높지 않을 수 있습니다.</p>GET coco_multi/_search
{
 "size": 20,
 "_source": [
   "description",
   "language",
   "en"
 ],
 "knn": {
   "field": "vector_description.predicted_value",
   "k": 20,
   "num_candidates": 1000,
   "query_vector_builder": {
     "text_embedding": {
       "model_id": ".multilingual-e5-small_linux-x86_64_search",
       "model_text": "query: con mèo màu gì?"
     }
   }
 }
}<p>결과에는 모두 고양이 또는 어떤 형태의 색상이 언급되어 있습니다:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt979f5944b1708042/6a17ef76420229ff6829f6aa/33e1e887dbbdd1066cfedc7375f5e3b46538529e-859x847.png" alt="" /><p>이제 개선해 봅시다! <a href="https://cohere.com/blog/rerank-3pt5">Cohere의</a>다국어 재랭크 모델을 통합하여 질문에 해당하는 추론을 개선해 보겠습니다.</p>PUT _inference/rerank/cohere_rerank
{
 "service": "cohere",
 "service_settings": {
   "api_key": "your_api_key",
   "model_id": "rerank-v3.5"
 },
 "task_settings": {
   "top_n": 10,
   "return_documents": true
 }
}


GET coco_multi/_search
{
"size": 10,
"_source": [
  "description",
  "language",
  "en"
],
"retriever": {
  "text_similarity_reranker": {
    "retriever": {
      "rrf": {
        "retrievers": [
          {
            "knn": {
              "field": "vector_description.predicted_value",
              "k": 50,
              "num_candidates": 100,
              "query_vector_builder": {
                "text_embedding": {
                  "model_id": ".multilingual-e5-small_linux-x86_64_search",
                  "model_text": "query: con mèo màu gì?" // English: What color is the cat?
                }
              }
            }
          }
        ],
        "rank_window_size": 100,
        "rank_constant": 0
      }
    },
    "field": "description",
    "inference_id": "cohere_rerank",
    "inference_text": "con mèo màu gì?"
  }
}
} {
       "_index": "coco_multi",
       "_id": "rQiYQJYBgf6odR9bBYyH",
       "_score": 1.5501487,
       "_source": {
         "description": "Hai cái điện thoại được đặt trên một cái chăn cạnh một con mèo con màu đen.",
         "en": "A black kitten lays on her side beside remote controls.",
         "language": "vi"
       }
     },
     {
       "_index": "coco_multi",
       "_id": "swiXQJYBgf6odR9b04uf",
       "_score": 1.5427427,
       "_source": {
         "description": "Một con mèo sọc nâu nhìn vào máy quay.", // Real translation: A brown striped cat looks at the camera 
         "en": "This cat is sitting on a porch near a tire.",
         "language": "vi"
       }
     },<p>이제 최고의 결과를 통해 저희 애플리케이션은 새끼 고양이의 색이 검은색 또는 줄무늬가 있는 갈색이라고 자신 있게 대답할 수 있습니다. 여기서 더욱 흥미로운 점은 벡터 검색이 실제로 원본 데이터 세트의 영어 캡션에서 누락된 부분을 찾아냈다는 점입니다. 참조 영어 번역에서 갈색 줄무늬 고양이를 놓쳤음에도 불구하고 이를 찾아낼 수 있습니다. 이것이 바로 벡터 검색의 힘입니다.</p><h2>결론</h2><p>이 블로그에서는 다국어 임베딩 모델의 유용성과 Elasticsearch를 활용하여 모델을 통합하여 임베딩을 생성하고 하이브리드 검색 및 재랭커로 관련성과 정확도를 효과적으로 개선하는 방법을 살펴봤습니다. 원하는 언어와 데이터 세트에 <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/nlp/ml-nlp-e5">대해 즉시 사용 가능한 E5 모델을 사용하여 자체 클라우드</a> <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">클러스터를</a> 생성하여 다국어 의미론적 검색을 사용해 볼 수 있습니다.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/multilingual-embedding-model-hybrid-search-reranking</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/multilingual-embedding-model-hybrid-search-reranking</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <category><![CDATA[운영]]></category>
    <dc:creator><![CDATA[Quynh Nguyen]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf625e3f63fcd9f54/6a17ef7796142a61f8eb1bcd/d341b04acecc8eeec321f5404e1643447ecc8526-720x420.png" length="0" type="image/png"/>
    <pubDate>Mon, 03 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch에서 다국어 임베딩 모델 배포하기]]></title>
    <description><![CDATA[Elasticsearch에서 벡터 검색 및 언어 간 검색을 위한 e5 다국어 임베딩 모델을 배포하는 방법을 알아보세요.]]></description>
    <content:encoded><![CDATA[<h2>소개</h2><p>전 세계 사용자가 있는 세계에서는 다국어 정보 검색(CLIR)이 매우 중요합니다. CLIR은 검색을 단일 언어로 제한하는 대신 <em>모든</em> 언어로 정보를 찾을 수 있도록 하여 사용자 경험을 개선하고 운영을 간소화합니다. 이커머스 고객이 자신의 언어로 상품을 검색하면 데이터를 미리 현지화할 필요 없이 적절한 결과가 표시되는 글로벌 시장을 상상해 보세요. 또는 학술 연구자들이 뉘앙스와 복잡성이 있는 논문을 모국어로 검색할 수 있으며, 출처가 다른 언어로 되어 있어도 검색할 수 있습니다.</p><p>다국어 텍스트 임베딩 모델을 사용하면 바로 그렇게 할 수 있습니다. 임베딩은 텍스트의 의미를 숫자 벡터로 표현하는 방법입니다. 이 벡터는 비슷한 의미를 가진 텍스트가 고차원 공간에서 서로 가깝게 위치하도록 설계되었습니다. 특히 다국어 텍스트 임베딩 모델은 여러 언어에서 동일한 의미를 가진 단어와 구문을 유사한 벡터 공간에 매핑하도록 설계되었습니다.</p><p>오픈 소스 다국어 E5와 같은 모델은 대개 대조 학습과 같은 기술을 사용하여 방대한 양의 텍스트 데이터를 학습합니다. 이 접근 방식에서 모델은 유사한 의미를 가진 텍스트 쌍(양의 쌍)과 서로 다른 의미를 가진 텍스트 쌍(음의 쌍)을 구별하는 방법을 학습합니다. 모델은 양성 쌍 간의 유사도는 최대화하고 음성 쌍 간의 유사도는 최소화하도록 생성하는 벡터를 조정하도록 학습됩니다. 다국어 모델의 경우 이 학습 데이터에는 서로 다른 언어로 된 텍스트 쌍이 포함되어 있어 모델이 여러 언어에 대한 공유 표현 공간을 학습할 수 있습니다. 이렇게 생성된 임베딩은 쿼리의 언어에 관계없이 텍스트 임베딩 간의 유사성을 사용하여 관련 문서를 찾는 교차 언어 검색을 비롯한 다양한 NLP 작업에 사용할 수 있습니다.</p><h2>다국어 벡터 검색의 이점</h2><ul><li><p><strong>뉘앙스</strong>: 벡터 검색은 키워드 매칭을 넘어 의미론적 의미를 포착하는 데 탁월합니다. 이는 언어의 맥락과 미묘한 차이를 이해해야 하는 작업에 매우 중요합니다.</p></li><li><p><strong>언어 간 이해</strong>: 쿼리와 문서가 서로 다른 어휘를 사용하는 경우에도 여러 언어에서 효과적으로 정보를 검색할 수 있습니다.</p></li><li><p><strong>연관성</strong>: 쿼리와 문서 간의 개념적 유사성에 초점을 맞춰 보다 관련성 높은 결과를 제공합니다.</p></li></ul><p>예를 들어, 여러 국가에서 소셜 미디어가 정치 담론에 미치는 영향( ")을 연구하는 학계 연구자(" )가 있다고 가정해 보겠습니다. 벡터 검색을 사용하면 "l'impatto dei social media sul discorso politico" (이탈리아어) 또는 "ảnh hưởng của mạng xã hội đối với diễn ngôn chính trị" (베트남어) 같은 검색어를 입력하면 관련 문서를 영문으로 찾을 수 있습니다, 스페인어 또는 기타 색인된 언어로 된 관련 문서를 찾아보세요. 벡터 검색은 정확한 키워드가 포함된 논문뿐만 아니라 소셜 미디어가 정치에 미치는 영향에 대한 <em>개념을</em> 논의하는 논문도 찾아내기 때문입니다. 이를 통해 연구의 폭과 깊이를 크게 향상시킬 수 있습니다.</p><h2>시작하기</h2><p>기본으로 제공되는 E5 모델을 사용하여 Elasticsearch를 사용하여 CLIR을 설정하는 방법은 다음과 같습니다. 여러 언어로 된 이미지 캡션이 포함된 <a href="https://huggingface.co/datasets/romrawinjp/multilingual-coco">오픈 소스 다국어 COCO 데이터셋을</a> 사용하여 두 가지 유형의 검색을 시각화해 보겠습니다:</p><ol><li><p>하나의 영어 데이터 세트에서 다른 언어로 된 쿼리 및 검색어, 그리고</p></li><li><p>여러 언어로 된 문서가 포함된 데이터 집합을 기반으로 여러 언어로 쿼리할 수 있습니다.</p></li></ol><p>그런 다음 하이브리드 검색과 재랭크의 힘을 활용하여 검색 결과를 더욱 개선할 것입니다.</p><h2>필수 구성 요소</h2><ul><li><p>Python 3.6+</p></li><li><p>Elasticsearch 8+</p></li><li><p>Elasticsearch Python 클라이언트: pip 설치 elasticsearch</p></li></ul><h2>데이터 세트</h2><p><a href="https://huggingface.co/datasets/romrawinjp/multilingual-coco">COCO 데이터 세트</a> 는 대규모 캡션 데이터 세트입니다. 데이터 세트의 각 이미지는 여러 언어로 캡션되어 있으며, 언어별로 여러 번역본을 사용할 수 있습니다. 데모용으로 각 번역을 개별 문서로 색인화하여 참조할 수 있도록 가장 먼저 제공되는 영어 번역과 함께 보여드리겠습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc7e508e9a7dffe8/6a17f3e2b1e113249579f394/d4f0632529c71a22fbdecf21c9f4f0bb64b8e69c-1600x567.png" alt="" /><h3>1단계: 다국어 COCO 데이터 세트 다운로드</h3><p>블로그를 단순화하고 쉽게 따라갈 수 있도록 여기서는 간단한 API 호출을 통해 restval의 처음 100개 행을 로컬 JSON 파일에 로드합니다. 또는 허깅페이스의 라이브러리 데이터셋을 사용하여 전체 데이터셋 또는 데이터셋의 하위 집합을 로드할 수도 있습니다.</p>import requests
import json
import os
### Download multilingual coco dataset into a json file (for easy viewing)
### Here we are retrieving first 100 rows for this example
### Alternatively, you can use `datasets` library from Hugging Face
url = "https://datasets-server.huggingface.co/rows?dataset=romrawinjp%2Fmultilingual-coco&amp;config=default&amp;split=restval&amp;offset=0&amp;length=100"
response = requests.get(url)


if response.status_code == 200:
   data = response.json()
   output_file = "multilingual_coco_sample.json" 
   ### Loading the downloaded content into a json file locally
   with open(output_file, "w", encoding="utf-8") as f:
       json.dump(data, f, indent=4, ensure_ascii=False)
   print(f"Data successfully downloaded and saved to {output_file}")
else:
   print(f"Failed to download data: {response.status_code}")
   print(response.text)<p>데이터가 JSON 파일에 성공적으로 로드되면 다음과 비슷한 내용이 표시됩니다:</p><p><code>Data successfully downloaded and saved to multilingual_coco_sample.json</code></p><h3>2단계: (Elasticsearch를 시작하고) Elasticsearch에서 데이터 색인하기</h3><p>a) 로컬 Elasticsearch 서버를 시작합니다.</p><p>b) Elasticsearch 클라이언트를 시작합니다.</p>from elasticsearch import Elasticsearch
from getpass import getpass


# Initialize Elasticsearch client
es = Elasticsearch(getpass("Host: "), api_key=getpass("API Key: "))


index_name = "coco"


# Create the index if it doesn't exist
if not es.indices.exists(index=index_name):
   es.indices.create(index=index_name, body=mapping)<p>c) 인덱스 데이터</p># Load the JSON data
with open('./multilingual_coco_sample.json', 'r') as f:
   data = json.load(f)


rows = data["rows"]
# List of languages to process
languages = ["en", "es", "de", "it", "vi", "th"]


# For each image, we will process each individual caption as its own document
bulk_data = []
for data in rows:
   row = data["row"]
   image = row.get("image")
   image_url = image["src"]


   # Process each language
   for lang in languages:
       # Skip if language not present in this row
       if lang not in row:
           continue


       # Get all descriptions for this language
 # along with first available English caption for reference
       descriptions = row[lang]
       first_eng_caption = row["en"][0]


       # Prepare bulk indexing data
       for description in descriptions:
           if description == "":
               continue
           # Add index operation
           bulk_data.append(
               {"index": {"_index": index_name}}
           )
           # Add document
           bulk_data.append({
               "language": lang,
               "description": description,
               "en": first_eng_caption,
               "image_url": image_url,
           })


# Perform bulk indexing
if bulk_data:
   try:
       response = es.bulk(operations=bulk_data)
       if response["errors"]:
           print("Some documents failed to index")
       else:
           print(f"Successfully bulk indexed {len(bulk_data)} documents")
   except Exception as e:
       print(f"Error during bulk indexing: {str(e)}")


print("Indexing complete!")<p>데이터가 색인되면 다음과 비슷한 내용이 표시됩니다:</p><p><code>Successfully bulk indexed 4840 documents</code></p><p><code>Indexing complete!</code></p><h3>3단계: E5 학습된 모델 배포하기</h3><p>Kibana에서 스택 관리 &gt; <strong>학습된 모델</strong> 페이지로 이동하고 .multilingual-e5-small_linux-x86_64에 대한 <strong>배포를</strong> 클릭합니다. 옵션을 선택합니다. 이 E5 모델은 Linux-x86_64에 최적화된 소규모 다국어 버전으로, 즉시 사용할 수 있습니다. '배포'를 클릭하면 배포 설정 또는 vCPU 구성을 조정할 수 있는 화면이 표시됩니다. 간단하게 하기 위해 사용량에 따라 배포를 자동으로 확장하는 적응형 리소스를 선택한 기본 옵션을 사용하겠습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfbc09867063a8e6f/6a17f3e3148009a295b4889d/95cd8f352425d1db2d04b00c3c88d1e71d1ef19a-1600x440.png" alt="" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt264a2016341e9b6f/6a17f3f0e8fbce18de3a1aa7/1599d99949dda8267acc58f400a403a3af5373ef-1600x655.png" alt="" /><p>선택적으로 다른 텍스트 임베딩 모델을 사용하려는 경우 사용할 수 있습니다. 예를 들어, BGE-M3를 사용하려면 <a href="https://www.elastic.co/docs/reference/elasticsearch/clients/eland/machine-learning#ml-nlp-pytorch">Elastic의 Eland Python 클라이언트를</a> 사용하여 HuggingFace에서 모델을 가져올 수 있습니다.</p>export MODEL_ID="bge-m3"
export HUB_MODEL_ID="BAAI/bge-m3"
export CLOUD_ID={{CLOUD_ID}}
export ES_API_KEY={{API_KEY}}
docker run -it --rm docker.elastic.co/eland/eland \
eland_import_hub_model --cloud-id $CLOUD_ID --es-api-key $ES_API_KEY --hub-model-id $HUB_MODEL_ID --es-model-id $MODEL_ID --task-type text_embedding --start<p>그런 다음 학습된 모델 페이지로 이동하여 가져온 모델을 원하는 구성으로 배포합니다.</p><h3>4단계: 배포된 모델을 사용하여 원본 데이터에 대한 임베딩을 벡터화하거나 생성합니다.</h3><p>임베딩을 생성하려면 먼저 텍스트를 가져와 추론 텍스트 임베딩 모델을 통해 실행할 수 있는 수집 파이프라인을 만들어야 합니다. 이 작업은 Kibana의 사용자 인터페이스 또는 Elasticsearch의 API를 통해 수행할 수 있습니다.</p><p><strong>Kibana 인터페이스를 통해 이 작업을 수행하려면</strong>, 학습된 모델을 배포한 후 <strong>테스트 </strong>버튼을 클릭합니다. 이렇게 하면 생성된 임베딩을 테스트하고 미리 볼 수 있습니다. <code>coco</code>인덱스에 대한 새 데이터 보기를 만들고, 데이터 보기를 새로 만든 코코 데이터 보기로 설정하고, 필드를 임베딩을 생성할 필드이므로 <code>description</code> 로 설정합니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt24a2b9a9a5111bbc/6a17f3f13e9e452c7fba15d5/cfe189e13dc118d325e7fb90bdace0c912e29f51-1088x1600.png" alt="" /><p>잘 작동합니다! 이제 수집 파이프라인을 생성하고 원본 문서를 재색인하고 파이프라인을 통과하여 임베딩이 포함된 새 인덱스를 생성할 수 있습니다. <strong>파이프라인 생</strong>성을 클릭하면 임베딩을 만드는 데 필요한 프로세서가 자동으로 채워지는 파이프라인 생성 프로세스를 안내합니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte39c5ad52702103d/6a17f3f3e9ea87dd05a9c734/1e043c1c3279b66fbdf19c06b41e76e613043998-1600x1126.png" alt="" /><p>마법사는 데이터를 수집하고 처리하는 동안 장애를 처리하는 데 필요한 프로세서를 자동으로 채울 수도 있습니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt63e002a79d177cf4/6a17f3f596142a3e91eb1c48/8804d31b4f869078e3b2245040bbb0ab1720a94a-1600x1084.png" alt="" /><p>이제 수집 파이프라인을 만들어 보겠습니다. 파이프라인의 이름을 <code>coco_e5</code> 으로 지정합니다. 파이프라인이 성공적으로 생성되면, 마법사에서 원래 색인된 데이터를 새 색인으로 재색인하여 임베딩을 생성하는 데 즉시 파이프라인을 사용할 수 있습니다. <strong>색인 재생성을 </strong>클릭하여 프로세스를 시작합니다.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt39243d9ad1779fdf/6a17f3f696142a13eaeb1c4c/e34b1b18f5b24420d4581fe4d657c569926c2023-1600x1126.png" alt="" /><h2>보다 복잡한 구성의 경우, Elasticsearch API를 사용할 수 있습니다.</h2><p>일부 모델의 경우 모델 학습 방식에 따라 임베딩을 생성하기 전에 실제 입력에 특정 텍스트를 미리 추가하거나 추가해야 할 수 있으며, 그렇지 않으면 성능이 저하될 수 있습니다.</p><p>예를 들어, e5의 경우 모델은 입력 텍스트가 "passage: {content of passage}". 이를 위해 수집 파이프라인을 활용해 보겠습니다: 새로운 수집 파이프라인 <strong>벡터화_descriptions를</strong> 생성하겠습니다. 이 파이프라인에서는 임시 <code>temp_desc</code> 필드를 새로 만들고, "passage: "를 <code>description</code> 텍스트에 추가하고, 모델에서 <code>temp_desc</code> 을 실행하여 텍스트 임베딩을 생성한 다음 <code>temp_desc</code> 을 삭제합니다.</p>PUT _ingest/pipeline/vectorize_descriptions
{
"description": "Pipeline to run the descriptions text_field through our inference text embedding model",
"processors": [
 {
   "set": {
     "field": "temp_desc",
     "value": "passage: {{description}}"
   }
 },
 {
   "inference": {     
"field_map": {
       "temp_desc": "text_field"
     },
     "model_id": ".multilingual-e5-small_linux-x86_64_search",
     "target_field": "vector_description"
   }
 },
 {
   "remove": {
     "field": "temp_desc"
   }
 }
]
}<p>또한 생성된 벡터에 어떤 <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/dense-vector#dense-vector-quantization">양자화 유형을</a> 사용할지 지정할 수도 있습니다. 기본적으로 Elasticsearch는 <code>int8_hnsw</code> 을 사용하지만 여기서는 각 차원을 단일 비트 정밀도로 축소하는 <a href="https://www.elastic.co/search-labs/blog/better-binary-quantization-lucene-elasticsearch">Better Binary Quantization</a> (또는 <code>bqq_hnsw</code>)을 사용하려고 합니다. 이렇게 하면 메모리 사용량이 96% (또는 32배)로 줄어드는 대신 정확도는 더 높아집니다. 이 정량화 유형을 선택한 이유는 나중에 정확도 손실을 개선하기 위해 리랭커를 사용할 것이라는 것을 알고 있기 때문입니다.</p><p>이를 위해 <strong>coco_multi라는</strong> 새 인덱스를 생성하고 매핑을 지정합니다. 여기서 마법은 <strong>벡터_설명</strong>필드에 있으며, 여기서 인덱스_옵션의유형을 <strong>bbq_hnsw로</strong> 지정합니다.</p>PUT coco_multi
{
 "mappings": {
   "properties": {
     "description": {
       "type": "text"
     },
     "en": {
       "type": "text"
     },
     "image_url": {
       "type": "keyword"
     },
     "language": {
       "type": "keyword"
     },
     "vector_description.predicted_value": {
       "type": "dense_vector",
       "dims": 384,
       "index": "true",
       "similarity": "cosine",
       "index_options": {
         "type": "bbq_hnsw" 
       }
     }
   }
 }
}<p>이제 설명 필드를 '벡터화'하거나 임베딩을 생성하는 수집 파이프라인을 사용하여 원본 문서를 새 인덱스로 재색인할 수 있습니다.</p>POST _reindex?wait_for_completion=false
{
 "source": {
   "index": "coco"
 },
 "dest": {
   "index": "coco_multilingual",
   "pipeline": "vectorize_descriptions"
 }
}<p>여기까지입니다! 우리는 Elasticsearch와 Kibana로 다국어 모델을 성공적으로 배포했으며, Kibana 사용자 인터페이스 또는 Elasticsearch API를 통해 Elastic으로 데이터로 벡터 임베딩을 생성하는 방법을 단계별로 배웠습니다. 이 시리즈의 두 번째 파트에서는 다국어 모델 사용의 결과와 뉘앙스에 대해 살펴봅니다. 그 동안 자체 <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">클라우드 클러스터를 생성하여</a> 원하는 언어와 데이터 세트에서 <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/nlp/ml-nlp-e5">즉시 사용 가능한 E5 모델을 사용하여 다국어 시맨틱 검색을</a> 사용해 볼 수 있습니다.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/multilingual-embedding-model-deployment-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/multilingual-embedding-model-deployment-elasticsearch</guid>
    <category><![CDATA[벡터 데이터베이스]]></category>
    <category><![CDATA[운영]]></category>
    <dc:creator><![CDATA[Quynh Nguyen]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt59254226694f93a6/6a17f3f81480098988b488a1/8f2aa7bebb6b2f701e274ba7282273f9ab4abed6-720x432.png" length="0" type="image/png"/>
    <pubDate>Wed, 22 Oct 2025 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>