<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Priscilla Parodi - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Priscilla Parodi - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/author/priscilla-parodi</link>
    </image>
    <link>https://www.elastic.co/search-labs/author/priscilla-parodi</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/author/priscilla-parodi.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 13:13:23 GMT</lastBuildDate>
  <item>
    <title><![CDATA[AI plagiarism: Plagiarism detection with Elasticsearch]]></title>
    <description><![CDATA[Here's how to check for AI plagiarism using Elasticsearch, focusing on use cases with NLP models and Vector Search.]]></description>
    <content:encoded><![CDATA[<p>Plagiarism can be <strong>direct</strong>, involving the copying of parts or the entire content, or <strong>paraphrased</strong>, where the author's work is rephrased by changing some words or phrases.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb6e139760e56ec98/6a171147dc55de0ad2e00edf/5d7073187fda829438aeec8d3a1194a5bea2ba57-1440x347.png" alt="" /><p>There is a distinction between inspiration and paraphrasing. It is possible to read a content, get inspired, and then explore the idea with your own words, even if you come to a similar conclusion.</p><p>While plagiarism has been a topic of discussion for a long time, the accelerated production and publication of content have kept it relevant and posed an ongoing challenge.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc2af1678cb04506b/6a171149cf4f251c2ab2d257/a0b9a98d729db09dae0a79315c001e6763c12704-1400x1016.png" alt="" /><p>This challenge isn't limited to books, academic research, or judicial documents, where plagiarism checks are frequently conducted. It can also extend to newspapers and even social media.</p><p>With the abundance of information and easy access to publishing, how can plagiarism be effectively checked on a scalable level?</p><p>Universities, government entities, and companies employ diverse tools, but while a straightforward <a href="https://www.elastic.co/search-labs/lexical-and-semantic-search-with-elasticsearch">lexical search</a> can effectively detect direct plagiarism, the primary challenge lies in identifying <strong>paraphrased content.</strong></p><h2>Plagiarism detection with Generative AI</h2><p>A new challenge emerges with Generative AI. Is content generated by AI considered plagiarism when copied?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8a1ed1b6fa3f1a56/6a17114b4a531bc40736aa69/9345b28d6d27c37469bc38e823c41780b4eabfe5-1440x875.png" alt="" /><p>The <a href="https://openai.com/">OpenAI</a> <a href="https://openai.com/policies/terms-of-use">terms of use</a>, for example, specify that OpenAI will not claim copyright over content generated by the API for users. In this case, individuals using their Generative AI can use the generated content as they prefer without citation.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbc2e08ab5b808e38/6a17114dab7f086cbadb9f93/e1f415f69a247666f02ddc81468944920c874cd7-968x814.png" alt="" /><p>However, the acceptance of using Generative AI to improve efficiency remains a topic of discussion.</p><p>In an effort to contribute to plagiarism detection, OpenAI developed a <a href="https://huggingface.co/roberta-base-openai-detector">detection model</a> but later acknowledged that its accuracy is not sufficiently high.</p><p><em>"We believe this is not high enough accuracy for standalone detection and needs to be paired with metadata-based approaches, human judgment, and public education to be more effective."</em></p><p>The challenge persists; however, with the availability of more tools, there are now increased options for detecting plagiarism, even in cases of paraphrased and AI content.</p><h2>Detecting plagiarism with Elasticsearch</h2><p>Recognizing this, in this blog we are exploring one more use case with Natural Language Processing (NLP) models and Vector Search, plagiarism detection, beyond metadata searches.</p><p>This is demonstrated with <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/plagiarism-detection-with-elasticsearch/plagiarism_detection_es.ipynb">Python examples</a>, where we utilize a <a href="https://sbert.net/datasets/emnlp2016-2018.json">dataset</a> from <a href="https://www.sbert.net/">SentenceTransformers</a> containing NLP-related articles. We check the abstracts for plagiarism by performing 'semantic textual similarity' considering 'abstract' embeddings generated with a <a href="https://huggingface.co/sentence-transformers/all-mpnet-base-v2">text embedding model</a> previously imported into Elasticsearch. Additionally, to identify AI-generated content — AI plagiarism, an <a href="https://huggingface.co/roberta-base-openai-detector">NLP model</a> developed by OpenAI was also imported into Elasticsearch.</p><p>The following image illustrates the data flow:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0464f88d3ed12070/6a17114fab7f084905db9f97/1ad89c98a2f42a497548ca3947749bad54ec1172-1440x880.png" alt="" /><p>During the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest.html">ingest pipeline</a> with an <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/inference-processor.html">inference processor</a>, the 'abstract' paragraph is mapped to a 768-dimensional vector, the 'abstract_vector.predicted_value'.</p><p>Mapping:</p>"abstract_vector.predicted_value": { # Inference results field
"type": "dense_vector", 
"dims": 768, # model embedding_size
"index": "true", 
"similarity": "dot_product" # When indexing vectors for approximate kNN search, you need to specify the similarity function for comparing the vectors.
<p>The similarity between vector representations is measured using a vector similarity metric, defined using the 'similarity' <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html#dense-vector-params">parameter</a>.</p><p><a href="https://en.wikipedia.org/wiki/Cosine_similarity">Cosine</a> is the default similarity metric, computed as '(1 + cosine(query, vector)) / 2'. Unless you need to preserve the original vectors and cannot normalize them in advance, the most efficient way to perform cosine similarity is to normalize all vectors to unit length. This helps avoid performing extra vector length computations during the search, instead use 'dot_product'.</p><p>In this same pipeline, another inference processor containing the <a href="https://huggingface.co/roberta-base-openai-detector">text classification model</a> detects whether the content is 'Real' probably written by humans, or 'Fake' probably written by AI, adding the 'openai-detector.predicted_value' to each document.</p><p>Ingest Pipeline:</p>client.ingest.put_pipeline( 
    id="plagiarism-checker-pipeline",
    processors = [
    {
      "inference": { #for ml models - to infer against the data that is being ingested in the pipeline
        "model_id": "roberta-base-openai-detector", #text classification model id
        "target_field": "openai-detector", # Target field for the inference results
        "field_map": { #Maps the document field names to the known field names of the model.
        "abstract": "text_field" # Field matching our configured trained model input. 
        }
      }
    },
    {
      "inference": {
        "model_id": "sentence-transformers__all-mpnet-base-v2", #text embedding model id
        "target_field": "abstract_vector", # Target field for the inference results
        "field_map": {
        "abstract": "text_field" # Field matching our configured trained model input. Typically for NLP models, the field name is text_field.
        }
      }
    }
    
  ]
)
<p>At query time, the same text embedding model is also employed to generate the vector representation of the query 'model_text' in a 'query_vector_builder' object.</p><p>A k-nearest neighbor (kNN) search finds the k nearest vector to the query vector measured by the similarity metric.</p><p>The _score of each document is derived from the similarity, ensuring that a larger score corresponds to a higher ranking. This means that the document is more similar semantically. As a result, we are printing three possibilities: if score &gt; 0.9, we are considering 'high similarity'; if &lt; 0.7, 'low similarity’, otherwise, 'moderate similarity’. You have the flexibility to set different threshold values to determine what level of _score qualifies as plagiarism or not, based on your use case.</p><p>Additionally, text classification is performed to also check for AI-generated elements in the text query.</p><p>Query:</p>from elasticsearch import Elasticsearch
from elasticsearch.client import MlClient

#duplicated text - direct plagiarism test

model_text = 'Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at http://hucvl.github.io/recipeqa.'

response = client.search(index='plagiarism-checker', size=1,
    knn={
        "field": "abstract_vector.predicted_value",
        "k": 9,
        "num_candidates": 974,
        "query_vector_builder": { #The 'all-mpnet-base-v2' model is also employed to generate the vector representation of the query in a 'query_vector_builder' object.
            "text_embedding": {
                "model_id": "sentence-transformers__all-mpnet-base-v2",
                "model_text": model_text
            }
        }
    }
)

for hit in response['hits']['hits']:
    score = hit['_score']
    title = hit['_source']['title']
    abstract = hit['_source']['abstract']
    openai = hit['_source']['openai-detector']['predicted_value']
    url = hit['_source']['url']

    if score &gt; 0.9:
        print(f"\nHigh similarity detected! This might be plagiarism.")
        print(f"\nMost similar document: '{title}'\n\nAbstract: {abstract}\n\nurl: {url}\n\nScore:{score}\n\n")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

    elif score &lt; 0.7:
        print(f"\nLow similarity detected. This might not be plagiarism.")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

    else:
        print(f"\nModerate similarity detected.")
        print(f"\nMost similar document: '{title}'\n\nAbstract: {abstract}\n\nurl: {url}\n\nScore:{score}\n\n")

        if openai == 'Fake':
            print("This document may have been created by AI.\n")

ml_client = MlClient(client)

model_id = 'roberta-base-openai-detector' #open ai text classification model

document = [
    {
        "text_field": model_text
    }
]

ml_response = ml_client.infer_trained_model(model_id=model_id, docs=document)

predicted_value = ml_response['inference_results'][0]['predicted_value']

if predicted_value == 'Fake':
    print("\nNote: The text query you entered may have been generated by AI.\n")
<p>Output:</p>High similarity detected! This might be plagiarism.

Most similar document: 'RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes'

Abstract: Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at[ http://hucvl.github.io/recipeqa](http://hucvl.github.io/recipeqa).

url:[http://aclweb.org/anthology/D18-1166](http://aclweb.org/anthology/D18-1166)

Score:1.0
<p>In this example, after utilizing one of the 'abstract' values from our dataset as the text query 'model_text', plagiarism was identified. The similarity score is 1.0, indicating a high level of similarity — <strong>direct plagiarism</strong>. The vectorized query and document were not recognized as AI-generated content, which was expected.</p><p>Query:</p>#similar text - paraphrase plagiarism test 

model_text = 'Comprehending and deducing information from culinary instructions represents a promising avenue for research aimed at empowering artificial intelligence to decipher step-by-step text. In this study, we present CuisineInquiry, a database for the multifaceted understanding of cooking guidelines. It encompasses a substantial number of informative recipes featuring various elements such as headings, explanations, and a matched assortment of visuals. Utilizing an extensive set of automatically crafted question-answer pairings, we formulate a series of tasks focusing on understanding and logic that necessitate a combined interpretation of visuals and written content. This involves capturing the sequential progression of events and extracting meaning from procedural expertise. Our initial findings suggest that CuisineInquiry is poised to function as a demanding experimental platform.'
<p>Output:</p>High similarity detected! This might be plagiarism.

Most similar document: 'RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes'

Abstract: Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images. With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge. Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems. The data and leaderboard are available at[ http://hucvl.github.io/recipeqa](http://hucvl.github.io/recipeqa).

url:[http://aclweb.org/anthology/D18-1166](http://aclweb.org/anthology/D18-1166)

Score:0.9302529

Note: The text query you entered may have been generated by AI.
<p>By updating the text query 'model_text' with an AI-generated text that conveys the same message while minimizing the repetition of similar words, the detected similarity was still high, but the score was 0.9302529 instead of 1.0 — <strong>paraphrase plagiarism</strong>. It was also expected that this query, which was generated by AI, would be detected.</p><p>Lastly, considering the text query 'model_text' as a text about Elasticsearch, which is not an abstract of one of these documents, the detected similarity was 0.68991005, indicating low similarity according to the considered threshold values.</p><p>Query:</p>#different text - not a plagiarism

model_text = 'Elasticsearch provides near real-time search and analytics for all types of data.'
<p>Output:</p>Low similarity detected. This might not be plagiarism.
<p>Although plagiarism was accurately identified in the text query generated by AI, as well as in cases of paraphrasing and direct copied content, navigating the landscape of plagiarism detection involves acknowledging various aspects.</p><p>In the context of AI-generated content detection, we explored a model that makes a valuable contribution. However, it is crucial to recognize the inherent limitations in standalone detection, necessitating the incorporation of other methods to boost the accuracy.</p><p>The variability introduced by the choice of text embedding models is another consideration. Different models, trained with distinct datasets, result in varying levels of similarity, highlighting the importance of the text embeddings generated.</p><p>Lastly, in these examples, we used the document's abstract. However, plagiarism detection often involves large documents, making it essential to address the challenge of text length. It is common for the text to exceed a model's token limit, requiring segmentation into chunks before building embeddings. A <a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.11/knn-search.html#nested-knn-search">practical approach</a> to handling this involves utilizing nested structures with dense_vector.</p><h2>Conclusion</h2><p>In this blog, we discussed the challenges of detecting plagiarism, particularly in paraphrased and AI-generated content, and how semantic textual similarity and text classification can be used for this purpose.</p><p>By combining these methods, we provided an example of plagiarism detection where we successfully identified AI-generated content, direct and paraphrased plagiarism.</p><p>The primary goal was to establish a filtering system that simplifies detection but human assessment remains essential for validation.</p><p>If you are interested in learning more about semantic textual similarity and NLP, we encourage you to also check out these links:</p><ul><li><p><a href="https://www.elastic.co/what-is/semantic-search">What is semantic search?</a></p></li><li><p><a href="https://www.elastic.co/what-is/natural-language-processing">What is natural language processing (NLP)?</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/lexical-and-semantic-search-with-elasticsearch">Lexical and Semantic Search with Elasticsearch</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines">Chunking Large Documents via Ingest pipelines plus nested vectors equals easy passage search</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-plagiarism-checker-with-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-plagiarism-checker-with-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt68a5bc2434a9b03b/6a1711510e2e49a09641a22a/83e05cd4f81799fbb7b7950ed87600e825ec81e9-1024x1024.png" length="0" type="image/png"/>
    <pubDate>Tue, 19 Dec 2023 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Enhancing chatbot capabilities with NLP and vector search in Elasticsearch]]></title>
    <description><![CDATA[Explore how vector search and NLP work to enhance chatbot capabilities and see how Elasticsearch facilitates the process.]]></description>
    <content:encoded><![CDATA[<p>Conversational interfaces have been around for a while and are becoming increasingly popular as a means of assisting with various tasks, such as customer service, information retrieval, and task automation. Typically accessed through voice assistants or messaging apps, these interfaces simulate human conversation in order to help users resolve their queries more efficiently.</p><p>As technology advances, chatbots are used to handle more complex tasks — and quickly — while still providing a personalized experience for users. Natural language processing (NLP) enables chatbots to process the user's language, identifies the intent behind their message, and extracts relevant information from it. For example, Named Entity Recognition extracts key information in a text by classifying them into a set of categories. Sentiment Analysis identifies the emotional tone, and Question Answering the “answer” to a query. The goal of NLP is to enable algorithms to process human language and perform tasks that historically only humans were capable of, such as finding relevant passages among large amounts of text, summarizing text, and generating new, original content.</p><p>These advanced NLP capabilities are built upon a technology known as <a href="https://www.elastic.co/what-is/vector-search">vector search</a>. Elastic has native support for vector search, performing exact and approximate <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search">k-nearest neighbor (kNN) search</a>, and for NLP, enabling the use of custom or <a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-model-ref.html#ml-nlp-model-ref">third-party models</a> directly in Elasticsearch.</p><p>In this blog post, we will explore how vector search and NLP work to enhance chatbot capabilities and demonstrate how Elasticsearch facilitates the process. Let's begin with a brief overview of vector search.</p><h2>Vector search</h2><p>Although humans can comprehend the meaning and context of written language, machines cannot do the same. This is where vectors come in. By converting text into vector representations (numerical representations of the meaning of the text), machines can overcome this limitation. Compared to a traditional search, instead of relying on keywords and lexical search based on frequencies, vectors enable the process of text data using operations defined for numerical values.</p><p>This allows vector search to locate data that shares similar concepts or contexts by using distances in the "embedding space" to represent similarity given a query vector. When the data is similar, the corresponding vectors will be alike.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt53a615a8ba0ac931/6a17d795fbc5f8b257491910/08542abf8108aace288745b1aca8579b476ddc1b-1440x618.png" alt="" /><p>Vector search is not only utilized in NLP applications, but it’s also used in various other domains where unstructured data is involved, including image and video processing.</p><p>In a chatbot flow, there can be several approaches to users' queries, and as a result, there are different ways to improve information retrieval for a better user experience. Since each alternative has its own set of advantages and possible disadvantages, it is essential to take into account the available data and resources, as well as the training time (when applicable) and expected accuracy. In the following section, we will cover these aspects for question-answering NLP models.</p><h2>Question-answering</h2><p>A question-answering (QA) model is a type of NLP model that is designed to answer questions asked in natural language. When users have questions that require inferring answers from multiple resources, without a pre-existing target answer available in the documents, generative QA models can be useful. However, these models can be computationally expensive and require large amounts of data for domain related training, which may make them less practical in some situations, even though this method can be particularly valuable to handle out-of-domain questions.</p><p>On the other hand, when users have questions on a specific topic, and the actual answer is present in the document, extractive QA models can be used. These models directly extract the answer from the source document, providing transparent and verifiable results, making them a more practical option for businesses or organizations that want to provide a simple and efficient way of answering questions.</p><p>The example below demonstrates the use of a pre-trained extractive QA model, <a href="https://huggingface.co/deepset/minilm-uncased-squad2">available on Hugging Face</a> and deployed into Elasticsearch, to extract answers from a given context:</p>POST _ml/trained_models/deepset__minilm-uncased-squad2/deployment/_infer
{
    "docs": [{"text_field": "Canvas is a data visualization and presentation application within Kibana. With Canvas, live data can be pulled directly from Elasticsearch and combined with colors, images, text, and other customized options to create dynamic, multi-page displays."}],
    "inference_config": {"question_answering": {"question": "What is Kibana Canvas?"}}
}


{
  "predicted_value": "a data visualization and presentation application",
  "start_offset": 10,
  "end_offset": 59,
  "prediction_probability": 0.28304219431376443
}
<p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-deploy-models.html">Deploy trained models.</a></p><p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-ner-example.html#ex-ner-ingest">Add a model to an inference ingest pipeline.</a></p><p>There are various ways to handle user queries and retrieve information, and using multiple language models and data sources can be an effective alternative when dealing with unstructured data. To illustrate this, we have an example of the data processing of a chatbot employed to respond to queries with answers considering data extracted from selected documents.</p><h2>Chatbot data processing: NLP and vector search</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta418a9c54bb16cf9/6a17d7975772624b371bca43/c2d1a2f110e937b1d3e5df0d5caac3c906c98fb0-1440x748.png" alt="" /><p>As shown above, the data processing for our chatbot can be divided into three parts:</p><ul><li><p><strong>Vector processing:</strong> This part converts documents into vector representations.</p></li><li><p><strong>User input processing:</strong> This part extracts relevant information from the user query and performs semantic search and hybrid retrieval.</p></li><li><p><strong>Optimization:</strong> This part includes monitoring and is crucial for ensuring the chatbot's reliability, optimal performance, and great user experience.</p></li></ul><h2>Vector processing</h2><p>For the <strong>processing</strong> part, the first step is to determine component parts of each document to then convert each element to a vector representation; these representations can be created for a wide range of data formats.</p><p>There are various methods that can be used to compute embeddings, including pre-trained models and libraries.</p><p>It's important to note that the effectiveness of search and retrieval on these representations depends on the existing data and the quality and relevance of the method used.</p><p>As the vectors are computed, they are stored in Elasticsearch with a <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/dense-vector.html">dense_vector</a> field type.</p>PUT &lt;target&gt;
{
  "mappings": {
    "properties": {
      "doc_part_vector": {
        "type": "dense_vector",
        "dims": 3
      },
      "doc_part" : {
        "type" : "keyword"
      }
    }
  }
}
<h2>Chatbot user input processing</h2><p>For the <strong>user</strong> part, after receiving a question, it's useful to extract all possible information from it before proceeding. This helps to understand the user's intention, and in this case, we are using a <a href="https://huggingface.co/dslim/bert-base-NER">Named Entity Recognition model (NER)</a> to assist with that. NER is the process of identifying and classifying named entities into predefined entity categories.</p>POST _ml/trained_models/dslim__bert-base-ner/deployment/_infer
{
  "docs": { "text_field": "How many people work for Elastic?"}
}


{
  "predicted_value": "How many people work for [Elastic](ORG&amp;Elastic)?",
  "entities": [
    {
      "entity": "Elastic",
      "class_name": "ORG",
      "class_probability": 0.4993975435876747,
      "start_pos": 25,
      "end_pos": 32
    }
  ]
}
<p>Although not a necessary step, by using structured data or the above or another NLP model result to categorize the user's query, we can restrict the kNN search using a <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-search-filter-example">filter</a>. This helps to improve performance and accuracy by reducing the amount of data that needs to be processed.</p>    "filter": {
      "term": {
        "org": "Elastic"
      }
    }
<h2>Semantic search and hybrid retrieval</h2><p>Since the prompt originates from user queries and the chatbot needs to process human language with its variability and ambiguity, <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#semantic-search">semantic search</a> is a great fit. In Elasticsearch, you can perform semantic search in a single step by passing the query string and the ID of the <a href="https://huggingface.co/sentence-transformers/msmarco-MiniLM-L-12-v3">embedding model</a> into a query_vector_builder object. This will vectorize the query and perform kNN <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-search.html">search</a> to retrieve top k matches that are closest in meaning to the query:</p>POST /&lt;target&gt;/_search
{
  "knn": {
    "field": "doc_part_vector",
    "k": 5,
    "num_candidates": 20,
    "query_vector_builder": {
      "text_embedding": {
        "model_id": "&lt;text-embedding-model-id&gt;",
        "model_text": "&lt;query_string&gt;"
      }
    }
  }
 }
<p><a href="https://www.elastic.co/guide/en/machine-learning/8.7/ml-nlp-text-emb-vector-search-example.html">End-to-end example: How to deploy a text embedding model and use it for semantic search.</a> Elasticsearch uses the Lucene implementation of the Okapi BM25, a <strong>sparse model</strong> , to rank text queries for relevance, while <strong>dense models</strong> are used for <strong>semantic search</strong>. To <strong>combine</strong> the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#_combine_approximate_knn_with_other_features"><strong>strengths of both</strong></a> <strong>,</strong> vector matches and matches obtained from the text query, you can perform a <strong>hybrid retrieval</strong> :</p>POST &lt;target&gt;/_search
{
  "query": {
          "match": {
            "content": {
              "query": "&lt;query_string&gt;"
            }
        }
  },
  "knn": {
    "field": "doc_part_vector",
    "query_vector_builder": {
      "text_embedding": {
    "model_id": "&lt;text-embedding-model-id&gt;",
     "model_text": "&lt;query_string&gt;"
      }
    },
    "filter": {
      "term": {
        "org": "Elastic"
      }
    }
  }
}
<h3>Combining both sparse and dense models often yields the best results</h3><p>Sparse models generally perform better on short queries and specific terminologies, while dense models leverage context and associations. If you want to learn more about how these methods compare and complement each other, <a href="https://www.elastic.co/blog/improving-information-retrieval-elastic-stack-benchmarking-passage-retrieval">here</a> we benchmark BM25 against two dense models that have been specifically trained for retrieval.</p><p>The most relevant result can usually be the first answer given to the user, the<a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-search.html#search-api-response-body-score">_score</a> is a number used to determine the <strong>relevance</strong> of the returned document.</p><h2>Chatbot optimization</h2><p>To help improve the user experience, performance, and reliability of your chatbot, in addition to applying hybrid scoring, you can incorporate the following approaches: <strong>Sentiment Analysis:</strong> To provide awareness of user comments and reactions as the dialog unfolds, you can incorporate a <a href="https://huggingface.co/distilbert-base-uncased-finetuned-sst-2-english">sentiment analysis model</a>:</p>POST _ml/trained_models/distilbert-base-uncased-finetuned-sst-2-english/deployment/_infer
{
  "docs": { "text_field": "That was not my question!"}
}


{
  "predicted_value": "NEGATIVE",
  "prediction_probability": 0.980080439016437
}
<p><a href="https://www.elastic.co/blog/chatgpt-elasticsearch-openai-meets-private-data"><strong>GPT's capabilities</strong></a> <strong>:</strong> As an alternative to enhance the overall experience, you can combine Elasticsearch's search relevance with OpenAI's GPT question-answering capabilities, utilizing the <a href="https://platform.openai.com/docs/guides/chat">Chat Completion API</a> to return to the user model-generated responses considering these top k documents as a context. <em>Prompt: "answer this question &lt;user_question&gt; using only this document &lt;top_search_result&gt;"</em></p><p><strong>Observability:</strong> Ensuring the performance of any chatbot is crucial, and monitoring is an essential component in achieving this. In addition to logs that capture chatbot interactions, it's important to track response time, latency, and other relevant chatbot metrics. By doing so, you can identify patterns, trends, and even detect anomalies.<a href="https://www.elastic.co/blog/monitor-openai-api-gpt-models-opentelemetry-elastic">Elastic Observability</a> tools enable you to collect and analyze this information.</p><h2>Summary</h2><p>This blog post covers what NLP and vector search are and delves into an example of a chatbot employed to respond to user queries by considering data extracted from the vector representation of documents.</p><p>As demonstrated, using NLP and vector search, chatbots are capable of performing complex tasks that go beyond structured, targeted data. This includes making recommendations and answering specific product or business-related queries using multiple data sources and formats as context, while also providing a personalized user experience.</p><p>Use cases range from providing customer service by assisting customers with their inquiries to helping developers with their queries, by providing step-by-step guidance, suggesting recommendations, or even automating tasks. Depending on the goal and existing data, other models and methods can also be utilized to achieve even better results and improve the overall user experience.</p><p>Here are some links on the topic that may be useful:</p><ol><li><p><a href="https://www.elastic.co/blog/how-to-deploy-natural-language-processing-nlp-getting-started">How to deploy natural language processing (NLP): Getting started</a></p></li><li><p><a href="https://www.elastic.co/blog/overview-image-similarity-search-in-elastic">Overview of image similarity search in Elasticsearch</a></p></li><li><p><a href="https://www.elastic.co/blog/chatgpt-elasticsearch-openai-meets-private-data">ChatGPT and Elasticsearch: OpenAI meets private data</a></p></li><li><p><a href="https://www.elastic.co/blog/monitor-openai-api-gpt-models-opentelemetry-elastic">Monitor OpenAI API and GPT models with OpenTelemetry and Elastic</a></p></li><li><p><a href="https://www.elastic.co/blog/why-technology-leaders-need-vector-search">5 reasons IT leaders need vector search to improve search experiences</a></p></li></ol><p>By incorporating NLP and native vector search in Elasticsearch, you can take advantage of its speed, scalability, and search capabilities to create highly efficient and effective chatbots capable of handling large amounts of data, whether structured or unstructured.</p><p>Ready to get started? Begin a <a href="https://cloud.elastic.co/registration?onboarding_token=vectorsearch&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">free trial of Elastic Cloud</a>.</p><p><em>In this blog post, we may have used or we may refer to third party generative AI tools, which are owned and operated by their respective owners. Elastic does not have any control over the third party tools and we have no responsibility or liability for their content, operation or use, nor for any loss or damage that may arise from your use of such tools. Please exercise caution when using AI tools with personal, sensitive or confidential information. Any data you submit may be used for AI training or other purposes. There is no guarantee that information you provide will be kept secure or confidential. You should familiarize yourself with the privacy practices and terms of use of any generative AI tools prior to use.</em></p><p><em>Elastic, Elasticsearch and associated marks are trademarks, logos or registered trademarks of Elasticsearch N.V. in the United States and other countries. All other company and product names are trademarks, logos or registered trademarks of their respective owners.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/enhancing-chatbot-capabilities-with-nlp-and-vector-search-in-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/enhancing-chatbot-capabilities-with-nlp-and-vector-search-in-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Priscilla Parodi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c545fc80b6d79d6/6a170214839dfad776dcfd6f/d968e646240cd3ef7c79b5124d562a5f951d812b-1440x840.png" length="0" type="image/png"/>
    <pubDate>Wed, 21 Jun 2023 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>