チャット履歴とフォローアップの質問

上記のプロセスは、ユーザーが 1 つの質問しかできない場合に適しています。しかし、このアプリケーションでは追加の質問も許可されており、これによりいくつかの追加の複雑さが生じます。たとえば、新しい質問を LLM に送信するときに追加のコンテキストとして含めることができるように、以前の質問と回答をすべて保存する必要があります。

このアプリケーションのチャット履歴は、Elasticsearch と Langchain の統合の一部である別のクラスであるElasticsearchChatMessageHistoryクラスを通じて管理されます。関連する質問と回答の各グループは、使用されたセッション ID への参照とともに Elasticsearch インデックスに書き込まれます。

def get_elasticsearch_chat_message_history(index, session_id):
    return ElasticsearchChatMessageHistory(
        es_connection=elasticsearch_client, index=index, session_id=session_id
    )

INDEX_CHAT_HISTORY = os.getenv(
    "ES_INDEX_CHAT_HISTORY", "workplace-app-docs-chat-history"
)

chat_history = get_elasticsearch_chat_message_history(
    INDEX_CHAT_HISTORY, session_id
)

前のセクションで気づいたかもしれませんが、LLM からの応答はチャンク単位でクライアントにストリーム配信されますが、完全な応答でanswer変数が生成されます。これは、各やり取りの後に、応答と質問を履歴に追加できるようにするためです。

chat_history.add_user_message(question)
chat_history.add_ai_message(answer)

クライアントがリクエスト URL のクエリ文字列でsession_id引数を送信する場合、その質問は同じセッションでの以前の質問のコンテキストで行われたものと見なされます。

このアプリケーションがフォローアップの質問に対して採用しているアプローチは、LLM を使用して会話全体を要約した要約された質問を作成し、それを検索フェーズで使用することです。この目的は、潜在的に大量の質問と回答の履歴に対してベクトル検索を実行することを避けることです。このタスクを実行するロジックは次のとおりです。

if len(chat_history.messages) > 0:
    # create a condensed question
    condense_question_prompt = render_template(
        'condense_question_prompt.txt', question=question,
        chat_history=chat_history.messages)
    condensed_question = get_llm().invoke(condense_question_prompt).content
else:
    condensed_question = question

docs = store.as_retriever().invoke(condensed_question)

これは主な質問の処理方法と多くの類似点がありますが、この場合は LLM のストリーミング インターフェースを使用する必要がないため、代わりにinvoke()メソッドが使用されます。

質問を要約するには、ファイル **api/templates/condense_question_prompt.txt` に保存されている別のプロンプトが使用されます。

Given the following conversation and a follow up question, rephrase the follow up question to be a standalone question, in its original language.

Chat history:
{% for dialogue_turn in chat_history -%}
{% if dialogue_turn.type == 'human' %}Question: {{ dialogue_turn.content }}{% elif dialogue_turn.type == 'ai' %}Response: {{ dialogue_turn.content }}{% endif %}
{% endfor -%}
Follow Up Question: {{ question }}
Standalone question:

このプロンプトには、セッションのすべての質問と応答に加えて、最後に新しいフォローアップの質問が表示されます。LLM は、すべての情報を要約した簡略化された質問を提供するように指示されています。

LLM が生成フェーズで可能な限り多くのコンテキストを取得できるように、取得されたドキュメントとフォローアップの質問とともに、会話の完全な履歴がメインプロンプトに追加されます。以下は、サンプル アプリケーションで使用されるプロンプトの最終バージョンです。

Use the following passages and chat history to answer the user's question. 
Each passage has a NAME which is the title of the document. After your answer, leave a blank line and then give the source name of the passages you answered from. Put them in a comma separated list, prefixed with SOURCES:.

Example:

Question: What is the meaning of life?
Response:
The meaning of life is 42.

SOURCES: Hitchhiker's Guide to the Galaxy

If you don't know the answer, just say that you don't know, don't try to make up an answer.

----

{% for doc in docs -%}
---
NAME: {{ doc.metadata.name }}
PASSAGE:
{{ doc.page_content }}
---

{% endfor -%}
----
Chat history:
{% for dialogue_turn in chat_history -%}
{% if dialogue_turn.type == 'human' %}Question: {{ dialogue_turn.content }}{% elif dialogue_turn.type == 'ai' %}Response: {{ dialogue_turn.content }}{% endif %}
{% endfor -%}

Question: {{ question }}
Response:

要約された質問の使用方法は、ニーズに合わせて調整できることに注意してください。一部のアプリケーションでは、生成フェーズでも要約された質問を送信すると、トークン数も減ってより効果的であることがわかります。あるいは、要約された質問をまったく使用せず、常にチャット履歴全体を送信すると、より良い結果が得られるかもしれません。これで、このアプリケーションがどのように動作するかを十分に理解し、さまざまなプロンプトを試して、自分のユースケースに最適なものを見つけることができると思います。

最先端の検索体験を構築する準備はできましたか?

十分に高度な検索は 1 人の努力だけでは実現できません。Elasticsearch は、データ サイエンティスト、ML オペレーター、エンジニアなど、あなたと同じように検索に情熱を傾ける多くの人々によって支えられています。ぜひつながり、協力して、希望する結果が得られる魔法の検索エクスペリエンスを構築しましょう。

はじめましょう