<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Agentic AI - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Agentic AI - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/blog/category/agentic-ai</link>
    </image>
    <link>https://www.elastic.co/search-labs/blog/category/agentic-ai</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/category/agentic-ai.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 12:07:25 GMT</lastBuildDate>
  <item>
    <title><![CDATA[Ask Elastic Agent Builder why it's slow: Natural-language trace analysis]]></title>
    <description><![CDATA[Four agent performance questions your Agent Builder traces can answer, covering token spend by model, tool error rates, slow conversation turns, and recent prompts. The ES|QL for each is here, including the type cast SUM() needs.]]></description>
    <content:encoded><![CDATA[<p>Ask Elastic Agent Builder how many tokens your agents burned today, and it writes the Elasticsearch Query Language (ES|QL) and runs it against your OpenTelemetry (OTel) trace data. Then it answers in the chat UI. The same holds for your other agent performance questions: Which tool fails most often? Which conversation turns are slowest? What have users actually been asking? Elastic Agent Builder is designed for this purpose. It allows users to ship agents grounded in your data in minutes. </p><p>If tracing is on, the <code>agent-builder-traces</code> <a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview">skill</a> is already loaded on every agent in your space and there’s nothing to install. The <a href="https://www.elastic.co/search-labs/blog/opentelemetry-tracing-agent-builder">first post</a> in our series covered enabling OTel tracing and the out-of-the-box (OOTB) dashboards, along with threshold alerts. This post is about asking.</p><h2>What is the agent-builder-traces skill?</h2><p><code>agent-builder-traces</code> is a built-in Agent Builder skill. It takes a natural-language question, generates an ES|QL query from it, executes that query against your trace index, and returns a plain-language summary. The index it targets is <code>traces-agent_builder.otel-&lt;space-id&gt;</code>, where <code>&lt;space-id&gt;</code> is the Kibana space that you’re working in. Using the exact space-scoped index pattern, rather than a wildcard, keeps data from other spaces out of your results.</p><p>Under the hood, the skill uses a single inline tool: <code>agent-builder-traces.generate_esql</code>. Instead of calling this tool directly, you ask a question and the agent calls the tool on your behalf, passing your question as the prompt for ES|QL generation. The tool resolves the current space's trace index, builds a query using the default model, executes it against Elasticsearch, and returns the result.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt556675b25cdd19b9/6ab4ad2b69958c862adc8c79/unnamed.png" alt="Elastic Agent Builder chat calling the agent-builder-traces skill to answer a token usage question in plain English" /><h2>Which privacy settings control what your traces capture?</h2><p>Several fields containing sensitive information are off by default. You can enable them in <strong>Gen AI Settings</strong> (<strong>Stack Management</strong> &gt; <strong>AI Assistants</strong>), under the "Agent Builder Traces" section. Expand the <strong>Advanced privacy</strong> settings to find them.</p><p><strong>Setting key</strong></p><p><strong>What it captures</strong></p><p><code>agentBuilder:tracing:includeUserPrompts</code></p><p>User messages</p><p><code>agentBuilder:tracing:includeLlmResponses</code></p><p>Large language model (LLM) response text</p><p><code>agentBuilder:tracing:includeToolDetails</code></p><p>Tool call arguments and results</p><p><code>agentBuilder:tracing:includeSystemPrompt</code></p><p>The agent's system prompt</p><p><code>agentBuilder:tracing:includeRealNames</code></p><p>Real tool, agent, and conversation names (hashed when off)</p><p><code>agentBuilder:tracing:includeRealIds</code></p><p>Real conversation and workflow IDs (hashed when off)</p><p><code>agentBuilder:tracing:includeUserData</code></p><p>Real user IDs and usernames (hashed when off)</p><h3>Why is my trace analysis query returning empty rows?</h3><p>If the skill returns empty rows or tells you that a field is unavailable, check two things: whether the relevant setting is enabled in your configuration, and whether your query time window overlaps with any recorded spans. The skill reports exactly what the query returned; it doesn’t fabricate content when fields are empty..</p><h2>What agent performance questions can you ask?</h2><p>Examples of questions that you can ask include:</p><h3>How many tokens have my agents used, by model?</h3><p><strong>Ask:</strong> <em>How many input and output tokens have my agents used in the last 24 hours, broken down by model?</em></p><p><strong>ES|QL:</strong> </p><p>The <code>TO_LONG()</code> cast around each token field is required. These fields can surface as mixed integer and long types across index generations, and ES|QL’s <code>SUM()</code> needs an explicit numeric conversion before it can aggregate them. If you write your own queries against this index and hit unexpected type errors, this is usually why.</p><h3>Which tools are failing most often?</h3><p><strong>Ask:</strong> <em>What is the error rate for each tool over the last 7 days?</em></p><p><strong>ES|QL: </strong></p><p>The tool error rate query aggregates across spans named <code>‘execute_tool’</code> and compares <code>status.code == "Error"</code> counts to total counts per tool name. The result shows which tools are failing most often. </p><h3>Which conversation turns are slowest?</h3><p><strong>Ask:</strong> <em>Show me the 10 slowest conversation turns in the last hour.</em></p><p><strong>ES|QL: </strong></p><p>The skill queries spans where <code>span.name LIKE "invoke_agent *"</code> and <code>attributes.elastic.inference.span.kind == "CHAIN"</code>. These correspond to individual conversation turns, from the moment a user sends a message to when the agent returns a response. Durations are stored in nanoseconds, so the skill converts to seconds before sorting.</p><h3>What have users been asking my agents?</h3><p><strong>Ask:</strong> <em>What questions have users been asking in the last 30 minutes?</em></p><p><strong>ES|QL:</strong> </p><p>The recent-prompts query returns data only when <code>agentBuilder:tracing:includeUserPrompts</code> is set to <code>true</code>. When the setting is off, the <code>attributes.gen_ai.input.messages</code> field is empty and the skill will suggest that you check your<strong> Include User Prompts</strong> privacy setting.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt207b044228f7d38b/6ab4adecf2c7a075d08b4ba0/unnamed.png" alt="Agent Builder trace analysis showing empty input messages because the Include User Prompts privacy setting is off" /><h2>When should you use Discover or Kibana Lens instead?</h2><p>The skill targets one index pattern: <code>traces-agent_builder.otel-&lt;space-id&gt;</code>. Two scenarios fall outside that boundary.</p><h3>Can the skill query my own application indices?</h3><p>If you want to run ad hoc questions against data that you've indexed yourself, use a general data exploration skill or write ES|QL directly in Discover. The <code>agent-builder-traces</code> skill won’t query outside the Agent Builder traces index.</p><h3>Can the skill build or edit dashboards?</h3><p>The skill focuses on ad hoc queries and generating summary text. If your goal is to build a permanent visualization from your trace data, use Lens or the <code>dashboard-management</code> skill instead. We covered that specific process in depth in the first post in our series.</p><p>The skill sits alongside two other ways of watching your agents.</p><p><strong>Tool</strong></p><p><strong>Best for</strong></p><p><strong>You use it when</strong></p><p><code>agent-builder-traces</code> skill</p><p>Ad hoc agent performance questions</p><p>You notice something and want an answer without leaving the chat</p><p>OOTB dashboards</p><p>Ongoing visibility</p><p>You want trends over time without asking anything</p><p>Threshold alerts</p><p>Automated monitoring</p><p>You define a condition once and get notified when it's breached</p><h2>How to build evaluation pipelines from agent trace data</h2><p>Spans in <code>traces-agent_builder.otel-&lt;space-id&gt;</code> capture real execution data from every conversation turn and LLM call running in your space. That record is the raw material for evaluation pipelines: automated checks on response quality and latency budgets per agent type, along with regression detection when you update a system prompt. The next post in our series covers how to build those eval loops from trace data that Agent Builder is already collecting.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-performance-trace-analysis-agent-builder</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-performance-trace-analysis-agent-builder</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Operations]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Meghan Murphy]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc61c395d7859ef8/6ab4acdee2d12a17d37d48c9/unnamed.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Agentic workflows in Elasticsearch: pause an AI agent for human approval, resume 72 hours later]]></title>
    <description><![CDATA[Build AI agent orchestration where the workflow waits for a human approval and then executes the fix on its own, with nothing extra to provision and the whole decision trail queryable in Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p>An AI agent receives a question and processes it within seconds. The agent responds before the session expires. That model works well for question and answer or code generation. And it works well for point-in-time analysis. But what happens when the agent needs to wait for a human to approve something and that human is in a meeting or dealing with another incident? The session expires and context is lost. The work starts over from scratch.</p><p></p><p>Elasticsearch Workflows solves this with a persistent execution state. A workflow can pause at a human approval gate for days and resume exactly where it left off. Every decision is persisted in Elasticsearch as searchable data.</p><p></p><p>In this article, we’ll build this in practice with a real scenario: Documents that failed during ingestion get stuck in a data stream’s failure store. </p><p></p><p>In this scenario, every step is reproducible. By the end, you’ll have an end-to-end workflow triggered by alerts, with pre-execution approval, automatic remediation, post-execution verification, and a rejection-and-revision path.</p><h2>What you’ll learn</h2><p>We’ll cover how to:</p><ul><li><p>Use the <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/failure-store">failure store</a> to capture and remediate documents that failed ingestion.</p></li><li><p>Build a long-running workflow with structured and binary approval gates using <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/flow-control-steps"><code>waitForInput</code></a> and <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/wait-for-approval"><code>waitForApproval</code></a>.</p></li><li><p>Invoke Elastic Agent Builder AI agents automatically after human approval.</p></li><li><p>Trigger workflows automatically from alerting rules.</p></li><li><p>Handle approval and rejection paths (real human-in-the-loop, not just a checkbox).</p></li></ul><h2>What’s a long-running AI agent?</h2><p>The difference between a session-bound agent and a long-running agent is the nature of the work rather than the speed of the large language model (LLM).</p><p></p><p>A session-bound agent lives within a session: It receives a question and processes it. Then it responds to the question. If the process dies, the context dies with it. </p><p></p><p>Operational processes, like data remediation or infrastructure provisioning, or like change review, are different. The processing itself is fast. What takes time are human decisions at unpredictable moments. The person who needs to approve a reindex might be handling an incident or in a different time zone. Or they may simply be at lunch. Long-running agents handle these breaks in workflow by assembling separate infrastructure: a database to store execution state, a server for the orchestrator, a UI for approvals, and integrations for notifications. </p><p></p><p>Elasticsearch Workflows runs where the data already lives. The workflow's long-term execution state is persisted in Elasticsearch, surviving restarts and waiting days between steps. AI agents come from Agent Builder and maintain a session only for the duration of each step, such as when the step ends or when the session ends. The workflow holds the context that spans across steps and days, and alerts come from the same rules already monitoring your data. The approval UI is in Kibana, and there’s no additional infrastructure to provision.</p><p></p><p>The workflow's internal execution state is managed by Kibana. The <code>remediation-runs</code> index is the audit trail that the workflow itself writes and that you can query like any other Elasticsearch index.</p><h2>What’s the Elasticsearch failure store?</h2><p>When Elasticsearch receives a document it cannot index, it has two options: Reject the document with an error, or store it somewhere safe for later analysis. The <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/failure-store">failure store</a> is that second option.</p><p></p><p>Imagine a <code>logs-demo-app</code> data stream with the <code>price</code> field mapped as <code>float</code>. If the source application sends <code>"price": "N/A"</code> instead of a number, Elasticsearch cannot index the document. With the failure store enabled, it’s redirected to a dedicated index within the data stream itself rather than losing that document.</p><p></p><p>Documents in the failure store preserve the original content and include information about the error, including the exception type, the message, the pipeline and processor that failed, and even the stack trace. You can query the failure store using the <code>data_stream::failures</code> syntax. </p><p></p><p>A typical document looks like this:</p>{
  "@timestamp": "2026-07-16T14:45:02.111Z",
  "document": {
    "id": "AZ9rY09teEMlWkReFa9e",
    "index": "logs-demo-app",
    "source": {
      "user_id": "u-test-1",
      "price": "INVALID",
      "message": "Order failed"
    }
  },
  "error": {
    "type": "document_parsing_exception",
    "message": "failed to parse field [price] of type [float]"
  }
}<p>The <code>document.source</code> field contains the original document that failed, and the <code>error</code> field describes what went wrong. This is exactly the structure that the <code>failure-analyst</code> agent will read to diagnose the problem.</p><p>But the documents remain stuck there. Someone needs to diagnose the problem and fix the pipeline or mapping. They also need to reindex the documents back into the data stream. This is exactly the kind of work that combines AI automation with human approval, and it’s what we’re going to build.</p><h2>How the AI agent orchestration works, from alert to resolution</h2><p>The workflow connects four building blocks: an alerting rule, two AI agents, human approval gates, and an audit index. The diagram below shows how these components interact.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt50eed826ed4fd509/6ab0ff9b3a38313fbe723342/agentic-workflow-diagram.png" alt="Agentic workflow diagram: alerting rule, failure-analyst diagnosis, Gate 1, automated execution, and Gate 2 verify" /><p>The alerting rule monitors <code>logs-demo-app::failures</code> every minute. When it finds documents written within the previous five minutes, it starts a workflow execution.</p><p>The workflow first runs <code>read_failures</code> to retrieve the failed documents. It then passes them to the <code>failure-analyst</code> agent through the <code>diagnose</code> step. The agent inspects the destination mapping and identifies the root cause. It then produces a structured remediation plan without making any changes.</p><p>The workflow records this plan in the <code>remediation-runs</code> index with the status <code>awaiting_fix_approval</code> and pauses at Gate 1.</p><p>Gate 1 uses <code>waitForInput</code>, a structured form that collects the reviewer's decision and optional notes. While the workflow waits, its state remains persisted in Elasticsearch. No agent session or polling loop needs to remain active, which means that the gate can wait for hours or days without continuously consuming compute resources.</p><p>If the reviewer approves the plan, the workflow invokes the <code>remediation-executor</code> agent through the <code>execute_fix</code> step. The agent calls the <code>execute-failure-store-fix</code> skill and creates the ingest pipeline. It also performs the bounded reindex and returns a detailed execution report.</p><p>Gate 2 appears only after the execution finishes. It uses <code>waitForApproval</code> to present the agent's report and asks the reviewer to choose between <strong>Yes, mark as resolved </strong>and<strong> No, escalate</strong>. The workflow then records the final outcome in <code>remediation-runs</code> and includes the complete report in <code>agent_report</code>.</p><p>If the reviewer rejects the initial plan at Gate 1, the workflow sends the diagnosis and the reviewer's feedback back to the <code>failure-analyst</code> agent, which then produces a revised plan that’s presented at Gate 1b. If approved, the revised plan follows the same automatic execution and verification path. If the plan is rejected again, the workflow records the case as <code>fix_rejected</code> and ends without applying any changes.</p><p>The <code>remediation-runs</code> index records the diagnosis, revision requests, revised plans, execution reports, and terminal outcomes as searchable audit documents. The shared workflow <code>execution_id</code> correlates records belonging to the same remediation run.</p><h3>Production consideration: Exception-based verification</h3><p>The diagram above shows the deliberately conservative implementation used in this tutorial: Every completed remediation pauses at Gate 2, including executions that appear fully successful. This makes the human-verification mechanism explicit and may be appropriate for low-volume or high-risk environments, but it can create review fatigue at scale. A production extension would add a deterministic verification step after execution. When the expected documents are present in the destination and the reindex reports no failures or policy-defined errors, the workflow can record resolved automatically. Only partial, failed, or ambiguous outcomes should pause at Gate 2 for site reliability engineering (SRE) review and possible escalation.</p><p>The following sections walk through each part: the data stream setup, the agents and skills, the alerting rule, and the workflow.</p><h2>Prerequisites</h2><p>To follow this tutorial, you’ll need:</p><ul><li><p>An Elastic Cloud Serverless project or an Elastic Stack 9.4+ deployment.</p></li><li><p>Agent Builder (generally available [GA] on Elasticsearch projects, enabled by default).</p></li><li><p>The following environment variables configured:</p></li></ul>export ES_URL="https://&lt;your-project&gt;.es.&lt;region&gt;.gcp.elastic.cloud:443"
export ES_API_KEY="&lt;your-api-key&gt;"
export KIBANA_ENDPOINT="https://&lt;your-project&gt;.kb.&lt;region&gt;.gcp.elastic.cloud"<h2>Setting up a data stream with the failure store enabled</h2><p>All code in this tutorial, including the workflow YAML, setup scripts, and agent instructions, is available in this <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/ai-agent-orchestration-human-approval-workflow">repository</a>. </p><p>
Start by getting the code from the repository:</p># Clone the repository without checking out all files
git clone --filter=blob:none --sparse https://github.com/elastic/elasticsearch-labs.git
cd elasticsearch-labs

# Check out only this companion folder
git sparse-checkout set supporting-blog-content/ai-agent-orchestration-human-approval-workflow
cd supporting-blog-content/ai-agent-orchestration-human-approval-workflow<p>Then run the setup script:</p>./scripts/01-setup-failure-store.sh<p>The script prepares the test environment: It removes previous test resources, creates a failure-store-enabled index template, ingests three valid baseline documents, initializes the failure store, and creates the <code>remediation-runs</code> audit index. The <code>price</code> field is mapped as <code>float</code> with <code>ignore_malformed: false</code> ; nonnumeric values generate a parsing error and get redirected to the failure store instead of being silently dropped. The data stream itself is created automatically when the first document is indexed.</p><p>The failure store is enabled at the index template level, not on the data stream directly. The template includes a <strong><code>data_stream_options</code></strong> block that tells Elasticsearch to activate the failure store for any data stream created from that template:</p>"data_stream_options": {
  "failure_store": {
    "enabled": true
  }
}<p>The data stream is created automatically upon the first ingestion request, inheriting the failure store configuration from the index template. From that point on, any document that cannot be indexed is redirected to the failure store instead of being rejected with an indexing error.</p><p>The script also ingests three valid documents with numeric <code>price</code> values. These documents establish the healthy baseline: The data stream exists and contains valid data before any ingestion failures occur.</p><p>After the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/ai-agent-orchestration-human-approval-workflow/scripts/01-setup-failure-store.sh"><code>01-setup-failure-store.sh</code></a> script completes, the environment is ready:</p>logs-demo-app           → 3 valid documents
logs-demo-app::failures → 0 documents (empty, index initialized)
remediation-runs        → created with explicit keyword mappings<p>The invalid documents that trigger the alerting rule and start the workflow are ingested later, when you run <code>02-trigger-test.sh</code> in the testing section. This sequence shows a realistic failure scenario: The data stream is operating normally with valid data, and then an upstream change causes malformed documents to arrive.</p><h3>Importing the agents and skills into Agent Builder</h3><p>Run the restore script to import the two agents and two skills, along with the workflow into your Kibana instance:</p>pip install requests python-dotenv
python3 scripts/restore.py<p>Expected output:</p>Restoring skills
  created  failure-store-remediation-planner
  created  execute-failure-store-fix
  Skills: created=2, updated=0

Restoring agents
  created  failure-analyst
  created  remediation-executor
  Agents: created=2, updated=0

Restoring workflow
  imported failure_store_remediation

Restore completed successfully<p>Running the script again updates existing components without creating duplicates.</p><h2>The AI agents that diagnose and execute the fix</h2><p>After running the <code>restore</code> script, two agents will be available in Agent Builder: one to diagnose failures and another to apply approved fixes. In Kibana, go to <strong>Agent Builder </strong>and verify that both agents and their associated skills were created correctly.</p><h3><code>failure-analyst</code>, the read-only diagnosis agent</h3><p>This agent reads failed documents from the failure store and inspects the destination mapping. It produces a structured remediation plan that includes an ingest pipeline definition and bounded replay instructions. The plan also includes a risk assessment. The agent doesn’t execute the proposed remediation.</p><p></p><p>The agent instructions establish its objective and read-only boundary. Key instructions include:</p>## Description

Analyzes failed documents from the failure store, identifies root causes, and proposes remediation pipelines.

## Instructions

# Role

You are a cautious Elasticsearch data quality analyst specializing in data
stream failure stores.

# Goal

Analyze a batch of related failure-store documents and produce a safe,
evidence-based remediation plan for human review.

Do not execute any changes.

# Required approach

- Identify the common root cause and the affected document pattern.
- Propose the smallest bounded remediation that addresses the demonstrated failure.
- Do not propose an executable remediation when required context is missing.
- Do not create pipelines, change mappings, update templates, run reindex,
  or perform any write operation.

...<p>Its associated <code>failure-store-remediation-planner</code> skill provides the specialized rules used to construct a safe replay. For example:</p>...

- A remediation pipeline that replays indexing or mapping failures from a failure store must use `recover_failure_document` as its first processor. 
- For a small human-reviewed batch, filter the reindex source using the exact failure-store document `_id` values supplied in the input. 
- Mapping changes, template changes, failure-store deletion, and upstream application changes must be listed only as manual follow-up recommendations.

...<p>The response contract is also constrained so that the workflow can consume the plan reliably:</p>Return exactly one JSON object and no additional explanatory text.
Do not use Markdown code fences, preambles, summaries, or attachments.<p>Together, the agent instructions and skill allow <code>failure-analyst</code> to produce an actionable but bounded remediation plan while leaving execution under the control of the workflow and its human approval gate.</p><h3><code>remediation-executor</code>, the agent that applies the approved fix</h3><p>The workflow invokes this agent after the remediation plan has been approved at Gate 1. Unlike the first agent, <code>remediation-executor</code> is allowed to perform the write operations required to apply the approved remediation.</p><p></p><p>Its agent instructions intentionally keep this role focused on immediate execution:</p>## Description
Executes Elasticsearch operations: creates pipelines and runs reindex.

## Instructions
You execute Elasticsearch remediation operations.

When asked to run a remediation, invoke the skill
"execute-failure-store-fix" directly and immediately.

Do not ask for confirmation. Do not ask clarifying questions.

After the skill completes, write a plain-text summary as your final message.
Include: whether the pipeline was created, how many documents were matched,
how many were successfully reindexed, and how many failed. This message is
captured by the workflow and shown to the human reviewer.<p>The associated <code>execute-failure-store-fix</code> skill defines how the approved plan is retrieved and executed. Key instructions include:</p>## Name

execute-failure-store-fix

## Description

1. The prompt will include a "Workflow execution ID". Use it to search the
   remediation-runs index for the approved plan for this specific execution.

2. Determine which field contains the remediation plan:
   - If status is "awaiting_fix_approval", read the "diagnosis" field.
   - If status is "awaiting_fix_approval_v2", read the
     "revised_diagnosis" field.

3. From that JSON object, extract:
   - remediation_pipeline.pipeline_id
   - remediation_pipeline.pipeline_definition
   - remediation_pipeline.reindex_request
   - affected_docs.data_stream

4. Create the ingest pipeline in Elasticsearch using the extracted
   pipeline_definition.

5. Execute the reindex_request exactly as specified in the approved plan.
   Do not substitute or hardcode the destination — use
   affected_docs.data_stream from the approved plan.

6. Report how many documents were matched, successfully reindexed,
   and failed..<p>The skill searches by both <code>execution_id</code> and approval status. This ensures that the executor retrieves the plan approved for the current workflow execution rather than an unrelated remediation run. The status determines whether it uses the original <code>diagnosis</code> or the <code>revised_diagnosis</code> produced during the revision path.</p><p></p><p>This separation gives the two agents distinct responsibilities: <code>failure-analyst</code> investigates the failure and proposes a bounded plan without making changes, while <code>remediation-executor</code> applies the selected plan only after the workflow has completed the required human approval step.</p><p></p><p>The setup scripts import the complete agent instructions and skill definitions. The excerpts above highlight the instructions most relevant to understanding the responsibilities and behavior of each agent. The complete definitions are available in <code>backup/save_agents.json</code> and <code>backup/save_skills.json</code> in the repository.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0bff70653cc57d12/6ab10598f2c7a0e6b88b38df/failure-analyst-remediation-executor.png" alt="Agent Builder in Kibana showing the Failure Analyst and Remediation Executor AI agents after import" /><p>After verifying that the agents and their skills have been imported correctly, the next step is to configure the alerting rule that detects new documents in the failure store and starts the remediation workflow.</p><h2>How do I trigger a workflow from an Elasticsearch alert</h2><p>The remediation process starts with an Elasticsearch alerting rule that monitors the failure store and starts the workflow when new failures are detected. Because the <code>restore.py</code> script has already imported and enabled <code>failure_store_remediation</code>, you can connect the workflow while creating the rule.</p><h3>Creating the alerting rule that watches the failure store</h3><p>The rule checks the failure store every minute and creates an alert when it finds documents written within the configured five-minute time window.</p><ol><li><p>Go to <strong>Management → Rules → Create rule</strong>.</p></li><li><p>Select <strong>Elasticsearch query</strong>.</p></li><li><p>Configure:</p></li><ol><ul><li><p><strong>Query type</strong>: <code>ES|QL</code>.</p></li><li><p><strong>Query</strong>:</p><p></p></li><li><p>Select time field: <code>@timestamp</code>.</p></li><li><p>Select alert group: <strong>Create an alert if matches are found.</strong></p></li><li><p>Time window: <strong><code>5 minutes</code></strong>.</p></li><li><p>Check every: <strong><code>1 minute</code></strong>.</p></li></ul></ol><li><p>In the <strong>Actions</strong> section, click <strong>Add action → Workflows</strong>.</p></li><li><p>Select the <code>failure_store_remediation</code> workflow.</p></li><li><p>Under <strong>Run workflow</strong> for, select <strong>New alerts</strong>.</p></li><li><p>Set <strong>Action frequency</strong> to <strong>Run per alert</strong>.</p></li><li><p>Save and enable the rule.</p></li></ol><p>Whenever the rule creates a new alert, it starts one <code>failure_store_remediation</code> workflow execution for that alert. The failure store is still empty at this stage; the invalid documents are ingested later by <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/ai-agent-orchestration-human-approval-workflow/scripts/02-trigger-test.sh"><code>02-trigger-test.sh</code></a>.</p><p></p><p>With the alert trigger configured, let’s examine the workflow that performs the diagnosis, approval, remediation, and verification steps.</p><h2>Building the agentic workflow, step by step</h2><p>At this point, the workflow has already been created in Kibana by the <code>restore.py</code> script executed earlier. To understand how the workflow operates, we’ll examine it incrementally, using relevant YAML excerpts to explain the role of each step in the remediation process. The complete workflow definition is available in the tutorial repository.</p><h3>Declaring the alert trigger and the 7-day timeout</h3><p>The workflow declares an alert trigger that allows it to be initiated by the alerting rule created earlier.</p>version: "1"
name: failure_store_remediation
description: &gt;
  Long-running workflow that remediates failed documents in a data stream's
  failure store. Triggered by an alerting rule when failures are detected.
  An AI agent diagnoses the root cause and proposes a fix. After human
  approval, the remediation-executor agent runs the fix automatically,
  then pauses for verification.
enabled: true
tags: [failure-store, remediation, agentic]
settings:
  timeout: 7d

triggers:
  - type: alert<p>Two details are important here. First, the <code>timeout: 7d</code> defines that the workflow can remain active for up to seven days. We set this explicitly because the workflow will pause at human approval gates that may take hours or days to respond.</p><p></p><p>Second, declaring <code>type: alert</code> makes the workflow available to alerting. The workflow action configured in the rule establishes the connection between the alert and <code>failure_store_remediation</code>.</p><h3>Reading failed documents from the failure store</h3><p>The workflow begins by retrieving failed documents using the <code>::failures</code> syntax:</p>steps:
  - name: read_failures
    type: elasticsearch.search
    with:
      index: "logs-demo-app::failures"
query:
	    range:
		"@timestamp":
		gte: "now-5m"
		lte: "now"
size: 50
sort:
        - "@timestamp": "desc"<p>The range query limits the search to failure-store documents created during the same five-minute window monitored by the alerting rule. Within that window, <code>size: 50</code> bounds the batch to the 50 most recent matching documents. For larger volumes, use <code>search_after</code> or pagination.</p><h3>How the AI agent diagnoses the root cause</h3><p>The next step sends the failed documents to the <code>failure-analyst</code> agent:</p>  - name: diagnose
    type: ai.agent
    agent-id: failure-analyst
    timeout: 300s
with:
      message: |
        Analyze these failed documents from the "logs-demo-app" failure store.

        Failed documents:
        {{ steps.read_failures.output.hits.hits | json }}

        Identify the root cause, classify the failure type, propose a concrete
        fix with ingest pipeline processors, generate a remediation pipeline to
        reindex the failed documents, and assess the risk.

        YOUR RESPONSE MUST BE A SINGLE JSON OBJECT ONLY.
  Do not write any explanation, preamble, markdown, or commentary.
  Start your response with { and end with }.
  Use exactly these top-level keys: root_cause, failure_type, affected_docs,
  proposed_fix, remediation_pipeline, risk_assessment.<p>The <code>{{ steps.read_failures.output.hits.hits | json }}</code> template injects the documents retrieved in the previous step into the agent's prompt. The agent analyzes the error patterns and identifies that the <code>price</code> field received string values where floating-point values were expected. It then proposes a remediation pipeline.</p><p>The agent's output arrives in <code>steps.diagnose.output.message</code> as a single JSON object. The exact diagnosis, proposed remediation, pipeline definition, and reindex request may vary between executions because they’re generated by the AI agent from the failures it receives. In our tests, the agent consistently identified the root cause and generated working remediation pipelines.</p><p>Despite this natural variability, the response contract remains fixed so that the workflow can persist the plan and the executor can reliably extract its executable components. The response always contains the six required top-level keys: <code>root_cause</code>, <code>failure_type</code>, <code>affected_docs</code>, <code>proposed_fix</code>, <code>remediation_pipeline</code>, and <code>risk_assessment</code>.</p><h4>Recording the diagnosis for audit</h4><p>Before requesting human approval, the workflow writes the diagnosis and proposed remediation plan to the <code>remediation-runs</code> index. This creates a durable audit record that can later be retrieved by the <code>remediation-executor</code> if the plan is approved:</p>- name: record_diagnosis
  type: elasticsearch.index
  with:
    index: remediation-runs
    document:
      "@timestamp": "{{ now | date_to_xmlschema }}"
      execution_id: "{{ execution.id }}"
      data_stream: "logs-demo-app"
      triggered_by_rule: "{{ event.rule.name }}"
      failure_count: "{{ steps.read_failures.output.hits.hits | size }}"
      failure_count_total: "{{ steps.read_failures.output.hits.total.value }}"
      diagnosis: "{{ steps.diagnose.output.message }}"
      status: awaiting_fix_approval<p>The <code>execution_id</code> correlates the diagnosis with the current workflow execution. The record stores both the number of failure documents loaded into the bounded batch and the total number of failures matching the query. After human approval, the executor uses the <code>execution_id</code> together with the approval status to retrieve the correct remediation plan.</p><h3>How do human approval gates work in Elasticsearch Workflows?</h3><p>The workflow uses two kinds of pause steps for different kinds of human decisions. Gate 1 uses <code>waitForInput</code> to collect a structured approval decision and reviewer feedback. Gate 2 uses <code>waitForApproval</code> because verification requires a binary choice: Mark the case as resolved, or escalate it.</p><p>
</p><p><strong>Gate 1</strong></p><p><strong>Gate 1b</strong></p><p><strong>Gate 2</strong></p><p>Step name</p><p><code>gate_fix</code></p><p><code>gate_fix_revised</code></p><p>gate_verify (approved path) / gate_verify_v2 (revised path)</p><p>Step type</p><p><code>waitForInput</code></p><p><code>waitForInput</code></p><p><code>waitForApproval</code></p><p>Timeout</p><p>72h</p><p>72h</p><p>72h</p><p>Decision shape</p><p>Approve or reject, plus notes</p><p>Approve or reject, plus notes</p><p>Yes, mark as resolved / No, escalate</p><p>Presents</p><p>AI diagnosis and proposed plan</p><p>Revised plan</p><p>remediation-executor report</p><p>Appears</p><p>Always</p><p>Only after Gate 1 rejection</p><p>Only after execution completes</p><p></p><p>Gate 1 is defined like this:</p>- name: gate_fix
  type: waitForInput
  timeout: 72h
  with:
    message: |
      GATE 1/2 - Fix Approval

      Data stream: logs-demo-app
      Failed documents: {{ steps.read_failures.output.hits.total.value }}

      The AI agent has diagnosed the failures. Review the full diagnosis
      in the 'diagnose' step output above, then approve or reject.
    schema:
      type: object
      properties:
        approved:
          type: boolean
          title: "Approve fix"
          default: true
        notes:
          type: string
          title: "Feedback (required if rejecting)"<p>When the workflow reaches this step, execution pauses, and its state is persisted in Elasticsearch. No agent session or polling loop needs to remain active while the gate waits. Every gate accepts a response for up to 72 hours, subject to the workflow's overall seven-day timeout.</p><p></p><p>The schema collects an approval decision and an optional <code>notes</code> field. When the proposal is rejected, the workflow passes the notes to the <code>failure-analyst</code> agent as context for revising the remediation plan. The testing section shows the approval and rejection payloads at the point where each is submitted.</p><h3>What happens when the reviewer approves or rejects the plan</h3><p>When Gate 1 is submitted, <code>route_fix</code> resumes the workflow without rerunning the completed steps. The simplified excerpt below shows the first transition in each branch. (The complete executable YAML is available in the tutorial repository.)</p>- name: route_fix
  type: if
  condition: "steps.gate_fix.output.response.approved: true"
  steps:
    - name: execute_fix
      type: ai.agent
      agent-id: remediation-executor
      timeout: 300s
      with:
        message: |
          Run remediation for the "logs-demo-app" failure store.

          Workflow execution ID: {{ execution.id }}

          The fix has been approved by a human reviewer. Execute the skill
          'execute-failure-store-fix' immediately without asking for confirmation.<p>Approval sends the current workflow <code>execution_id</code> to the <code>remediation-executor</code>. The executor uses this identifier together with the approval status to retrieve the remediation plan associated with this specific execution before applying it and continuing to Gate 2.</p><p>If Gate 1 is rejected, the workflow stores the reviewer’s feedback and asks the <code>failure-analyst</code> agent to produce a revised proposal. Gate 1b (<code>gate_fix_revised</code>) then presents that proposal for another human decision. When the revised plan is approved, <code>execute_fix_revised</code> sends the same workflow execution ID to the executor:</p>- name: execute_fix_revised
  type: ai.agent
  agent-id: remediation-executor
  timeout: 300s
  with:
    message: |
      Run the revised remediation for the "logs-demo-app" failure store.

      Workflow execution ID: {{ execution.id }}

      The revised fix has been approved by a human reviewer. Execute the skill
      'execute-failure-store-fix' immediately without asking for confirmation.<p>The executor uses the approval status to determine which plan belongs to the current path: the original <code>diagnosis</code> associated with <code>awaiting_fix_approval</code> or the <code>revised_diagnosis</code> associated with <code>awaiting_fix_approval_v2</code>. Both paths then continue to Gate 2. If the revised proposal is rejected at Gate 1b, the workflow records <code>fix_rejected</code> and ends without applying any changes.</p><h3>Verifying the result at Gate 2 before closing the case</h3><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf5f2a8d89684da90/6ab10a9c44675f24756f7a95/gate-2-human-approval.png" alt="Gate 2 human approval in Elasticsearch Workflows: remediation-executor report with resolve or escalate buttons" /><p>Gate 2 is a verification checkpoint. By the time it appears, the <code>remediation-executor</code> agent has already created the ingest pipeline and run the reindex. It has also returned a detailed report. The gate embeds the executor's report, <code>steps.execute_fix.output.message</code> on the approved path, or <code>steps.execute_fix_revised.output.message</code> when the revised plan was executed, and asks the reviewer to select either <strong>Yes (mark as resolved) </strong>or<strong> No (escalate)</strong>.</p><p>Gate 2 uses <code>waitForApproval</code> because the workflow only needs a binary verification decision. The operation details, including the pipeline ID, documents matched and reindexed, version conflicts, failures, and execution errors, come from the <code>remediation-executor</code> agent’s report rather than from fields entered manually by the reviewer.</p><p>In our validated execution, the agent created the <code>logs-demo-app-price-remediation</code> pipeline and matched and reindexed all five approved documents. It reported zero version conflicts and zero failures. The <code>logs-demo-app</code> data stream increased from three to eight documents, while the five original records remained in <code>logs-demo-app::failures</code>. This is expected because reindex copies documents; it doesn’t remove them from the failure store.</p><p>If the reviewer selects <strong>Yes, mark as resolved</strong>, <code>route_verify</code> writes a new document to <code>remediation-runs</code> with status <code>resolved</code> and stores the complete agent report in <code>agent_report</code>. If the reviewer selects <strong>No, escalate</strong>; the workflow writes an <code>escalated</code> record with the same report. The revised path mirrors this through <code>route_verify_v2</code>, producing the same statuses and the same <code>agent_report</code> field.</p><p>Gate 2 exists because a technically completed reindex can still produce partial or unexpected results. Escalating instead of automatically marking the remediation as resolved keeps the audit trail honest. Whether to keep this second gate depends on your environment and risk tolerance.</p><h2>The complete workflow</h2><p>After reviewing the workflow, remember that the complete code is available in the repository in the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/ai-agent-orchestration-human-approval-workflow/workflow/failure-store-remediation.yaml"><code>failure-store-remediation.yaml</code></a> file.</p><p>The workflow can finish in five ways, depending on the decisions made at Gate 1, Gate 1b, and Gate 2. The diagram below shows how the original and revised remediation paths lead to <code>resolved</code>, <code>escalated</code>, or <code>fix_rejected</code>.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt250701b4468e5ced/6ab10bb2b29b5da685c6ad09/five-completion-paths.png" alt="Five completion paths in the failure store remediation workflow: resolved, escalated and fix_rejected outcomes" /><p>Every terminal outcome is appended to the <code>remediation-runs</code> audit history. Across the workflow paths, the audit records preserve the diagnosis, reviewer feedback when a revision is requested, the executor report, and the terminal outcome.</p><p>With the rule and its workflow action already configured, the end-to-end flow is ready to test.</p><h2>Testing the workflow end to end</h2><p>With the alerting rule enabled, run the trigger script to ingest five documents with invalid <code>price</code> values and current timestamps:</p>./scripts/02-trigger-test.sh<p>The rule searches the previous five minutes, so it detects the new <code>failure-store</code> documents during its next evaluation. Go to <strong>Workflows → Executions</strong>, and open the new execution. Confirm that <code>read_failures</code>, <code>diagnose</code>, and <code>record_diagnosis</code> have completed and that <code>gate_fix</code> is waiting for input.</p><h3>Approving the fix at Gate 1</h3><p>At Gate 1, approve the proposed remediation by submitting:</p>{  "approved": true }<p>The workflow resumes at <code>route_fix</code> and invokes the <code>remediation-executor</code> agent through <code>execute_fix</code>. It pauses at Gate 2 with the execution report.</p><p>The following clip shows Gate 1 being approved and the workflow resuming automatically at <code>route_fix</code> and <code>execute_fix</code>.</p><p></p><p></p><p>Before completing Gate 2, verify that the five remediated documents were added to the data stream:</p>GET logs-demo-app/_count<p>Expected result:</p>{
  "count": 8
}<p>The data stream started with three valid documents, so a successful replay increases the total to eight. The original records remain in the failure store because reindex copies documents to the destination rather than removing the source records.</p><p>Complete Gate 2 by selecting <strong>Yes, mark as resolved</strong> or <strong>No, escalate</strong>. The workflow stores the selected outcome and the execution report in <code>remediation-runs</code>.</p><h3>Rejecting the plan and reviewing a revised one</h3><p>To test the rejection path, reset the environment and start another execution:</p>./scripts/03-reset-full-test-environment.sh --apply
./scripts/01-setup-failure-store.sh
./scripts/02-trigger-test.sh<p>At Gate 1, reject the proposal with specific feedback:</p>{
  "approved": false,
  "notes": "The proposed remediation should distinguish numeric strings from non-numeric strings. Convert numeric strings to a float, preserve non-numeric values in price_raw, and remove price only when conversion fails."
}<p>Although <code>notes</code> is optional in the current schema, provide specific feedback whenever you reject a proposal. The workflow passes this value to the <code>failure-analyst</code> agent when requesting a revised plan.</p><p>The workflow records the feedback and asks the agent to revise the proposal. It pauses at Gate 1b. To approve the revised plan, submit:</p>{
  "approved": true
}<p>To reject it definitively, submit:</p>{
  "approved": false,
  "notes": "Explain why the revised proposal should not be executed."
}<p>Approving the revision sends it through the same execution and Gate 2 verification path described above. Rejecting it records <code>fix_rejected</code> and ends the workflow without applying the remediation.</p><h2>Conclusion</h2><p>The pattern built in this tutorial combines three elements: persistent workflow state, specialized AI agents, and human control at the points where judgment matters. The <code>failure-analyst</code> agent diagnoses the problem and proposes a bounded remediation. Gate 1 pauses before any change is made, giving the reviewer control over execution. After approval, the <code>remediation-executor</code> agent applies the fix automatically, and Gate 2 pauses again so the result can be verified before the case is marked as resolved or escalated.</p><p>Persistent execution state solves more than the problem of expiring agent sessions. The diagnosis, reviewer feedback, execution report, and final decision are stored as searchable data in Elasticsearch and correlated by the workflow <code>execution_id</code>. This correlation allows the executor to retrieve the plan approved for the current workflow run, even when multiple remediation cases exist in <code>remediation-runs</code>. It also makes it possible to analyze how many remediations were approved on the first attempt and which failure types are most frequently rejected, along with how long approval gates remain open and which cases require escalation.</p><p>This workflow pattern also applies to index promotions, mapping changes, enrichment pipeline validation, infrastructure operations, and other processes where automation and human judgment must coexist: Diagnose the problem, pause for approval, execute the approved action automatically, pause for verification, and record the final outcome.</p><p></p><h2>Resources</h2><ul><li><p><a href="https://www.elastic.co/docs/explore-analyze/workflows">Elasticsearch Workflows documentation</a></p></li><li><p><a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Agent Builder documentation</a></p></li><li><p><a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/failure-store">Failure store documentation</a></p></li><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/ai-agent-orchestration-human-approval-workflow">Tutorial repository</a></p></li></ul><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-agent-orchestration-human-approval-workflow</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-agent-orchestration-human-approval-workflow</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Index Data]]></category>
    <category><![CDATA[Operations]]></category>
    <dc:creator><![CDATA[Alex Salgado]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c030f52a5d4283d/6ab0f53597bf87491eb22cf8/ai-agent-orchestration.png" length="0" type="image/png"/>
    <pubDate>Mon, 21 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[You and your AI agent shouldn't be using curl: Introducing the Elastic CLI and Agent Skills]]></title>
    <description><![CDATA[Elastic CLI reaches every Elasticsearch, Kibana and Cloud API from one command, and it's what Elastic Agent Skills run on. Input is validated against a JSON Schema before anything leaves your machine, and API keys stay in your OS keychain.]]></description>
    <content:encoded><![CDATA[<p><a href="https://github.com/elastic/cli">The Elastic CLI</a> gives you one command for every public Elastic API: Elasticsearch, Kibana, and Elastic Cloud's control plane, including Serverless projects. Learn <code>elastic es search</code>, and you already know how <code>elastic kb data-views list</code> behaves. It's built to be driven by an AI coding agent as easily as by you, so every command takes JSON in and out and validates input against a JSON Schema before sending anything. Plus, it exits with a code that an agent can branch on. Administrators control which commands run at all, and API keys go to your OS keychain, never into a large language model (LLM) transcript. Our Agent Skills now run on it. The command line interface (CLI) is in technical preview today.</p><p><a href="https://cloud.elastic.co/serverless-registration">Start a free Elastic Cloud Serverless trial</a> or <a href="https://cloud.elastic.co/login">log in to Elastic Cloud</a> to follow along, and install the CLI via <a href="https://www.npmjs.com/">npm</a>: </p>npm install -g @elastic/cli<h2>Designing the Elastic CLI for people and AI agents</h2><p>A useful side effect of building flexible tools for developers is that they’re more useful to AI agents, too. It’s also the tool that our<a href="https://github.com/elastic/agent-skills"> Agent Skills</a> now use to get things done, closing <a href="https://www.elastic.co/search-labs/blog/agent-skills-elastic">the loop that we opened in March 2026</a>, when we said that a CLI for agent workflows was coming.</p><p>The CLI gives every public API across Elasticsearch, Kibana, and Elastic Cloud one shape: the same flags, input and output conventions, authentication method, and failure mode. Consistency is the ergonomic feature; everything else is built on it.</p><p>Agents need the same thing, only stricter. An agent won’t know how to craft a valid CLI command or notice if a tool “feels” wrong; it needs output it can parse and input it can validate before sending, along with failures it can branch on. Agents are now a first-class interface to Elastic, alongside people, whether they live on the platform or in your editor and terminal.<a href="https://github.com/elastic/agent-skills"> Agent Skills</a>, and now the CLI, are how we serve the second kind, so those needs are built into the core of the CLI rather than tacked on.</p><h3>JSON input and output for every command</h3><p>Agents love structured text, and almost every Elastic API already speaks JSON, so first-class JSON support was a hard requirement. Developers who are quick with a <u><code>jq</code></u> query will be equally satisfied.</p><ul><li><p><strong>JSON output:</strong> Any command, like <code>elastic version</code> and <code>elastic es indices delete ...</code>, supports <code>--json</code>, which prints JSON-parseable output to stdout and nothing else. Failed commands print <code>{"error": {"code": "...", "message": "..."}}</code> to stderr.</p></li><li><p><strong>JSON input:</strong> Every command that takes input accepts JSON on stdin or via <code>--input-file</code>. Every top-level key in that JSON also works as a CLI argument, and inline arguments take precedence, so you can keep a big request body in a file and tweak a value or two per invocation.</p></li><li><p><strong>JSON Schema as </strong><strong><code>-</code></strong><strong><code>-help</code></strong><strong> output:</strong> Pass <code>--help --json</code> to any command, and it prints a valid JSON Schema for its input, which also feeds nicely into codegen tools. <code>elastic cli-schema</code> prints the whole command tree.</p></li></ul><h3>Exit codes that an AI agent can branch on</h3><p>Agents loop on exit codes as much as they do over stdout. All failure modes are distinguishable from success, even if stdout and stderr are never read.</p><h3>Safety rails: Keychain storage, allow lists, and validation</h3><p>No model uses tools perfectly 100% of the time, so an agent-friendly CLI should provide safety rails wherever possible.</p><ul><li><p><strong>Contexts and secret storage:</strong> Connection details live in named contexts in <code>~/.elasticrc.yml</code>, <code>kubectl</code>-style; switch with <code>--use-context</code>. API commands never take credentials as flags. <code>elastic config context add</code> writes API keys to your OS keychain (macOS, Linux, Windows) and leaves a <code>$(keychain:...)</code> reference in the YAML; <code>$(env:...)</code>, <code>$(cmd:...)</code>, and <code>$(file:...)</code> work, too. Creating a Serverless project with <code>--save-as</code> writes its credentials straight to the keychain and never prints them, so nothing leaks into logs or LLM transcripts.</p></li><li><p><strong>Allowlists/blocklists:</strong> A <code>commands.allowed</code> (or <code>commands.blocked</code>) list in the config file, globally or per context, ensures that only the commands an administrator wants are runnable.</p></li></ul>commands:
   allowed:
     - version
     - stack.es.search
     - stack.es.esql.*<ul><li><p><strong>Validation:</strong> Every command has a JSON Schema, so inputs are validated before any request is sent. Add <code>--dry-run</code> to any command that takes input, and it validates and exits without sending anything.</p></li><li><p><strong>Confirmation:</strong> Destructive commands prompt in a terminal. In a noninteractive session, where agents live, they refuse to run without <code>--yes</code> and say so in a structured error.</p></li><li><p><strong>Sanitization:</strong> Index, field, and pipeline names have length limits and forbidden characters. <code>elastic sanitize index-name '&lt;value&gt;'</code> (and <code>field-name</code>, <code>pipeline-name</code>, …) prints a version stripped of anything invalid.</p></li></ul><h3>Keeping API responses inside an agent's context window</h3><p>Elastic APIs return a lot of data, and an agent’s context window is finite. Three controls help keep unnecessary text out of the context window:</p><ul><li><p><strong>Field masks:</strong> <code>--output-fields</code> takes a comma-separated list, with dot notation for nested fields.</p></li></ul>elastic es info --output-fields 'name,version.number'
 # {
 #   "name": "serverless",
 #   "version": { "number": "9.5.0" }
 # }<ul><li><p><strong>String templates:</strong> For total control, <code>--output-template</code> takes a <a href="https://mustache.github.io/">mustache</a>-style template.</p></li></ul>elastic es info --output-template 'ES version: {{ version.number }}'
# ES version: 9.5.0<ul><li><strong>Command profiles:</strong> <code>--command-profile</code> serverless (or <code>default_profile: serverless</code> in your config) hides Elastic Cloud Hosted commands and the Elasticsearch namespaces that don’t exist on Serverless. That means less to scroll past and less for an agent to guess wrong. It’s the profile we recommend for agents.</li></ul><p><strong>Control</strong></p><p><strong>What it does</strong></p><p><strong>Syntax</strong></p><p><strong>When to use</strong></p><p>Field mask</p><p>Returns only the fields you name, using dot notation for nested fields</p><p><code>--output-fields 'name,version.number'</code></p><p>You want valid JSON back, just less of it. This is the default choice for agents parsing structured output.</p><p>String template</p><p>Renders the response through a mustache-style template</p><p><code>--output-template 'ES version: {{ version.number }}'</code></p><p>You need one value in a specific shape, for a shell variable, a log line, or a prompt.</p><p>Command profile</p><p>Hides commands and namespaces that don't apply to your deployment</p><p><code>--command-profile serverless</code>or <code>default_profile: serverless</code></p><p>You want a smaller command surface so an agent has less to scroll past and less to guess wrong. This is recommended for agents.</p><p></p><h2>Helpers for bulk ingest, scroll search, and msearch</h2><p>Some of Elasticsearch’s most popular APIs have a learning curve, so elastic es helpers wraps them:</p><ul><li><p><code>scroll-search</code>: Stream a large result set as NDJSON with paging handled for you.</p></li><li><p><code>bulk-ingest</code>: Ingest from a file, a directory, or stdin (NDJSON, JSON arrays, or CSV) with streaming, batching, concurrency, and retries.</p></li><li><p><code>msearch</code>: Send multiple searches in one request.</p></li><li><p><code>watch</code>: Print new documents from an index to stdout as they’re indexed. This is great for piping into logging tools.</p></li></ul><p><code>elastic es</code> and <code>elastic kb</code> are aliases for <code>elastic stack elasticsearch</code> and <code>elastic stack kibana</code>. If we don’t ship a command you need, <code>elastic extension create</code> scaffolds one for you.</p><h2>Searching Elastic docs from the terminal</h2><p>If you or your agents don’t know which API to use, elastic docs search (plus docs read and docs ask) searches Elastic’s documentation from the terminal, returning Markdown or <code>--json</code>. These are experimental. You’ll see a warning until you pass <code>--accept-experimental</code>, so explore, but don’t script against them yet.</p><h2>Shell completion for Bash, Zsh, and Fish</h2><p>Autocomplete hooks are available for Bash, Zsh, and Fish, and they always respect your <code>commands.allowed</code> or <code>commands.blocked</code> policy.</p><h2>How Elastic Agent Skills use the CLI</h2><p><a href="https://github.com/elastic/agent-skills">Agent Skills</a> teach an AI coding agent how an Elastic expert approaches a job; for example, which cluster health field is the verdict or how to stage a reindex so it doesn’t fall over. They capture process and judgment but not transport. A skill that embeds <a href="https://curl.se/">curl</a> with an auth header has hard-coded a hostname, key, and runtime, and it breaks when any of those change.</p><p>So our skills now use a <em>universal</em> format that runs unchanged in any runtime that can execute the <code>elastic</code> CLI, including Claude Code, Codex, Cursor, and GitHub Copilot. The body refers to operations in HTTP shorthand (<code>GET /_cluster/health</code>, <code>POST /_query</code>), and an operations table at the end binds each to a CLI command. That table is the only place transport appears:</p><p>HTTP API (shorthand)</p><p><code>elastic</code> CLI command</p><p><code>GET /{index}/_mapping</code></p><p><code>elastic es indices get-mapping --index '&lt;index&gt;'</code></p><p><code>POST /_query</code></p><p><code>elastic es esql query --format tsv --query "&lt;esql&gt;"</code></p><p><code>POST cloud:/api/v1/serverless/projects/elasticsearch</code></p><p><code>elastic cloud serverless projects search create --input-file &lt;json&gt; --wait --save-as &lt;ctx&gt;</code></p><p>Every universal skill also inherits a blunt preamble; that is, use the CLI, don’t guess credentials, don’t call the HTTP API directly, and never ask the user to paste an API key into the chat.</p><p>The two halves need each other. The skill supplies the expertise that the model doesn’t have, and the CLI supplies a way to act on it that’s validated, credential-safe, and scoped by your allowlist. Tell your agent to <em>spin up a Serverless project and load products.csv into it</em>, and the provisioning skill creates it with <code>--save-as</code>. The ingest skill dry-runs a mapping and loads with <code>elastic es bulk</code>, and the Elasticsearch Query Language (ES|QL) skill writes a query that parses on the first try. Every step returns JSON, exits non-zero on failure, and can only do what your policy allows.</p><p>Skills for Elastic Cloud onboarding and provisioning, Elastic Workflows, and Kubernetes investigation are available today. Skills for Elasticsearch query, ingest, reindex, and index design, plus Kibana dashboards and alerting, are close behind.</p><h2>What’s in the Elastic CLI technical preview, and what’s next</h2><p>The preview covers all public Elasticsearch Serverless, Kibana Serverless, and Elastic Cloud APIs. Hosted-only 9.x Elasticsearch API coverage is nearly 100%, and hosted-only 9.x Kibana APIs will be added soon.</p><p>We’re actively planning more developer experience work, including broader coverage for all supported stack releases, more helpers for common workflows, more skills in the public catalog, and loading the same skills into agents that run on the Elastic platform itself. What shapes that list is hearing how you and your agents use the CLI. Tell us what’s awkward and what’s missing, along with what you’d automate next.</p><h2>Install the Elastic CLI and Agent Skills</h2><p>The Elastic CLI is available now on npm (Node.js 22+). Install it and the skills together:</p>npm install -g @elastic/cli # or: npx -y @elastic/cli --help
npx skills add elastic/agent-skills<p>Then, add a context and check it:</p>elastic config context add prod --es-url https://&lt;project&gt;.es.us-east-1.aws.elastic.cloud --es-api-key &lt;KEY&gt;
 elastic status<p>Even without a project, you can <a href="https://cloud.elastic.co/serverless-registration">start a free Serverless trial</a> in about a minute, with no credit card. If you already have a project, <a href="https://cloud.elastic.co/login">log in</a> and create API keys for Elastic Cloud and your Elasticsearch clusters. Before pointing an agent at anything real, start with a trial project, a read-only key, and a scoped commands.allowed list. Be sure to take five minutes to read the <a href="https://github.com/elastic/agent-skills#security-considerations">security notes</a> in the skills repo.</p><p>Replace all those curl commands in your Bash scripts, and add some usage instructions to your AGENTS.md. Then let your agent’s skills work efficiently and accurately with our APIs. Let us know what you think, and don’t hesitate to<a href="https://github.com/elastic/cli/issues"> open an issue</a> if you find a bug or if your use case isn’t well supported. Your feedback directly shapes what we build next.</p><h2>Elastic CLI and Agent Skills resources</h2><ul><li><p><a href="https://github.com/elastic/cli">Elastic CLI on GitHub</a> and<a href="https://github.com/elastic/cli/tree/main/docs/cli"> CLI documentation</a></p></li><li><p><a href="https://github.com/elastic/agent-skills">Elastic Agent Skills on GitHub</a></p></li><li><p><a href="https://agentskills.io">agentskills.io specification</a></p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/deploy/elastic-cloud/serverless">Elastic Cloud Serverless documentation</a></p></li><li><p><a href="https://github.com/elastic/cli/issues">Report a CLI issue</a> ·<a href="https://github.com/elastic/agent-skills/issues"> Report a skills issue</a> ·<a href="https://discuss.elastic.co/"> Discuss</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-cli-ai-agents</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-cli-ai-agents</guid>
    <category><![CDATA[Developer Experience]]></category>
    <category><![CDATA[AI Tools ]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Josh Mock,Matt Ryan]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt111f4783cff3ef01/6aa7bad035eddc3a1a11d192/image1.png" length="0" type="image/png"/>
    <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Trust, but benchmark: How we let an AI agent optimize Elasticsearch]]></title>
    <description><![CDATA[We share how we built a harness that automatically identifies and implements optimizations in the Elasticsearch codebase.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch executes a diverse set of workloads, including sustained heavy index building and real-time search and analytics. Delivering excellent performance across the board requires going broad in coverage while simultaneously diving deep enough into the codebase to understand optimization opportunities for each workload. Traditionally, human attention has been the bottleneck in this process; there simply aren't enough engineering hours to scrutinize every hot code path looking for inefficiency across a large and evolving surface area.</p><p>However, with the rapid progression of coding agents, performance optimization has become a task we can tackle semiautomatically. Unlike many software engineering challenges, optimizing code offers a cheap and objective verifier. If you ask an AI model to make code faster, there’s a hard number at the end telling you exactly what happened, backed by profiling tools that explain why. This makes performance a perfect candidate for automation, provided you can actually trust the numbers.</p><p>If you simply point a coding agent at a benchmark, you typically get low signal-to-noise: wins that fall inside the variance of the environment, or variations caused by thermal throttling rather than better code. To capture optimizations that actually benefit Elasticsearch users, we had to bridge the gap between "checkable in principle" and "checked in practice." We built a highly trustworthy measurement loop: a <a href="https://en.wikipedia.org/wiki/Agent_harness">harness</a> that assumes the agent will be wrong a good fraction of the time but reliably catches and proves it when it’s right and then helps guide it where to look next.</p><p>Once the machinery is in place, the results speak for themselves. By letting this harness loose on the codebase, we've already begun uncovering meaningful wins across the stack. In part 2 of this post, we’ll dive into some examples it has found so far, including string conversion inefficiencies in Elasticsearch Query Language (ES|QL), an improvement to our NEON vector dot product implementation, and an upgrade opportunity for the gzip library we were using. In this part, we’ll take a look at the design choices we made and how they relate to the broader topic of effective harness development.</p><h2>The AI code optimization pipeline architecture</h2><p>The first step in any software engineering problem is to identify the correct high-level components. We made an architectural choice that turned out to be very helpful for this problem: separate understanding where opportunities exist from the loop making code changes. The agent starts with a real workload but only uses it to mine information about where to seek performance improvements. At this stage, it’s instructed to go broad and consider a range of performance-related signals. Once it has found and classified the hot spots, the agent reads the context of the code around them to understand the optimization opportunities. We use a separate task to condense the ranked list of hot spots into artifacts that a loop can iterate against in minutes: a microbenchmark that we prove exercises the hot path in its real operating regime. Finally, we use a proposer-verifier loop to actually make changes to the codebase to improve performance on the benchmark. This hands off to validation to assess the impact on real workloads at the end. Our CLI (<code>atune</code>) supplies the tools this process needs, and the rest is largely automated by a set of task-specific instructions.</p><p>For context, our high-level architecture looks like the following. Pink boxes are the humans, and teal boxes are the agent. There are three task types, one skeleton loop, and one referee.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0e561ef2d3b61421/6aa3e83fd1556f0fcc7269f8/unnamed.png" alt="AI code optimization pipeline: exploration to benchmark, human approval, exploitation, validation and PR review" /><p><em>An exploration task profiles a real workload and produces ranked opportunities; a human promotes one into an exploitation task, which iterates against an approved microbenchmark and commits each accepted experiment; a validation run on the real workload guards the result before a human reviews and opens the PR. Where no benchmark covers the hot path, a benchmark task authors one and a human approves it into a registry. A performance atlas informs every task and accumulates what each one learns.</em></p><h2>Why performance optimization suits autonomous agents</h2><p>Three properties make a task ideally suited for autonomous work, and it's worth being explicit about them because they provide a checklist you can use to evaluate automation candidates. You want:</p><ol><li><p>An objective verdict so that the agent can be held to something other than its own opinion.</p></li><li><p>A dense guiding signal so that it knows where to look next instead of guessing.</p></li><li><p>A bounded blast radius so that being wrong is affordable.</p></li></ol><p>Performance gives you all three. Benchmarks provide the verdict and profilers provide the gradient, while a rejected patch costs you wall-clock time rather than correctness. The change gets reverted, and the reason gets recorded. Life goes on. Regarding the first two, we've come to think the gradient matters more than the verdict. ″This got faster″ is binary, whereas a profile hints at what to try next. An agent can generate its next hypothesis conditioned on a rich guiding signal, the richer the better, rather than grinding through a list.</p><h2>Signals are what the agent gets to see</h2><p>A useful mental model is that the CLI is the agent's sensory apparatus. That changes how you design each command. Rather than exposing a capability, you design it to return a clear and concise answer to a question about the task at hand, and, when relevant, an explanation the model can reason over, instead of raw data it has to parse and interpret. These are the signals that the agent acts on, and we ended up with the following for our harness:</p><p><strong>Signal</strong></p><p><strong>Question it answers</strong></p><p>Facet-decomposed macro profile</p><p>Where in the code does real workload time go, per query type?</p><p>Allocation and lock sampling, in the same capture</p><p>Is the cost cycles, garbage, or contention?</p><p>Cost-composition classification</p><p>Is this in scope compute, other product code, GC, JIT tax, or parked threads?</p><p>Input-shape instrumentation</p><p>What does the workload actually feed this code?</p><p>Statistical verdict</p><p>Did this change help, at this measured noise floor?</p><p>Allocation-rate comparison</p><p>Did the new code end up allocating more?</p><p>Interpreted disassembly</p><p>Why did that result happen?</p><p>End-to-end A/B guard, with differential profile attribution</p><p>Did anything appear to break, and was it us?</p><p>Environment check</p><p>Is this machine even fit to measure right now?</p><p>Upstream duplicate search</p><p>Has somebody already reported or fixed this?</p><p>Four of these signals are worth dwelling on, because in each case the tool encodes a judgment that the agent would otherwise have had to keep making by hand.</p><p>Facet-decomposition is the clearest example. Blended CPU shares hide breadth: if you profile a mixed query workload, the grouping hash map insert and the percentiles sketch update can both show up as single-digit percentages of the total run time and look comparable. They aren't comparable at all, because the hash map insert is paid for by nearly every aggregation query, while the sketch update is only paid for when somebody asks for percentiles. So the profiler runs each named query facet as its own race, and every opportunity the agent records carries a breadth field (universal, broad, or narrow) and gets ranked by headroom × tractability × breadth. A universal 3% beats a narrow 10%. Putting the ranking function in the tool prevents it from having to be rediscovered on every run.</p><p>The cost-composition classification works as a router. Rather than handing the agent a flat top-N list of frame names, it buckets every sampled stack into ″in scope compute,″ ″other product compute,″ ″GC,″ ″JIT and safepoint overhead,″ ″off CPU waiting,″ and ″parked threads.″ Each of those buckets implies a different kind of investigation. If GC is above about 15%, the real target is allocation rate, and the CPU top frames will actively mislead you, because they show where objects were collected rather than where they were created. If JIT and safepoint overhead is above about 25%, you're looking at a ceiling rather than an opportunity, since no in-scope code change will move it. If threads are parked and core utilization is low, this is a concurrency problem and CPU flame graphs are the wrong instrument entirely. We wrote that mapping into the playbook as a table, so the model (even a cheap one) reads a profile the way that an experienced engineer would, rather than reaching straight for the top frame. The classification is a prior for forming a hypothesis, though, not a substitute for evidence, so the agent still has to cite specific frames when it proposes an experiment.</p><p>Interpreted disassembly is a tool we hadn't originally provided, but it most definitely earns its place. It helps answer <em>why</em>, the question that unblocks the next hypothesis. Flame graphs tell you where the time goes; they rarely tell you why a change made things worse. So <code>atune asm</code> runs the benchmark briefly with the JIT told to print the assembly for one hot method, captures both sides of the working-tree diff, reduces each to the final C2 compilation, normalizes the addresses, and diffs them. The diff alone would likely still be 4,000 lines of aarch64, so on top of it sits an interpretation layer: a per-mnemonic delta, a net instruction count, the compilation tier that was actually captured, and a vectorization signal that counts vector register references on each side and raises a warning when they're eliminated, halved, or narrowed from <a href="https://en.wikipedia.org/wiki/Advanced_Vector_Extensions">Advanced Vector Extensions</a> (AVX) to <a href="https://en.wikipedia.org/wiki/Streaming_SIMD_Extensions">Streaming SIMD Extensions</a> (SSE) width. The playbook then maps mnemonic patterns to causes:</p><p><strong>Pattern in the diff</strong></p><p><strong>Likely cause</strong></p><p><code>b.eq</code><code>/</code><code>b.ne</code> up, <code>csel</code> down</p><p>New unpredictable branches</p><p>Clusters of <code>str</code><code>/</code><code>ldr</code> against the stack pointer</p><p>The compiler ran out of registers</p><p>NEON loads replaced by scalar compares</p><p>The vector path degraded</p><p>In one experiment, the agent fused two <a href="https://en.wikipedia.org/wiki/Single_instruction,_multiple_data">SIMD</a> mask extractions into one, and the benchmark regressed by 26%. The vectorization warning explained it in about 10 seconds. Without that tool, the agent has a dead end and no working model of the machine; with it, it has a corrected model and several new ideas.</p><p>The fourth signal is less a single tool than a habit; the instruments check themselves. Core utilization is derived two independent ways, from sample density and from process sampling, so the two can be compared. The classification is rejected if the unclassified bucket exceeds a budget, on the grounds that a breakdown which can't account for its own samples shouldn't be reasoned over. The disassembly capture warns when the compilation it caught isn't the steady-state one. The upstream duplicate search is restricted to read-only commands, and that restriction is enforced by a test that greps the source, so no future edit can quietly reintroduce the ability to file anything. Each of these exists because a tool that can be confidently wrong is worse than a tool that is merely absent.</p><p>One small piece of design is worth highlighting as a specific instance of good return practice. <code>atune compare</code> returns 0 for improved, 1 for no change, 2 for regressed, and 3 for error, and the loop branches on this code. That means no parsing and no ambiguity about what the verdict was. Plus, no tokens are spent interpreting prose.</p><p>If you take one thing away from our CLI design, it’s a broader design principle. In a general setting, the interesting thing isn't the individual signals we found useful to understand performance; it's that the CLI is capturing and packaging the judgment of an experienced performance engineer into tools that return answers rather than raw data. Structurally imposing good judgment about the problem an agent is tasked with improves outcomes. The right CLI is as much part of that story as the instructions. Furthermore, tokens are saved by tools that return decisions and digests rather than data. That means a comparison verdict instead of raw JMH output; a triage summary sitting on top of a 4,000-line disassembly diff.</p><h2>From 20-second probes to hours of validation</h2><p>Building a verifier that you can afford to consult is a separate problem from building one that you can trust. This covers the affordability part. Or, if you like aphorisms, real workloads are where truth lives and where iteration goes to die. A macro profile takes 45 to 60 minutes, and an end-to-end validation run takes hours. But a microbenchmark takes minutes. That's why the exploration-then-exploitation split works; you go broad on the real workload once and then hand off to a microbenchmark that you can iterate against in minutes.</p><p>The catch is that the handoff is only sound if the microbenchmark exercises the hot path in a realistic operating regime; that is, with the right cardinality and right data distribution. If you get that wrong, your fast loop spins fast but in the wrong direction. Until we finalized the handoff procedure, we saw cases where the agent accepted changes on a benchmark whose key distributions happened to flatter it, and only the end-to-end run caught the problem.</p><h3>Handing off from exploration to exploitation</h3><p>The handoff between the two phases is structured rather than informal. An exploration task's primary deliverable is a set of opportunity records, and each one carries the scope paths that an exploitation task would be allowed to edit, a headroom estimate, a classification (constant factor, structural,</p><p>allocation, or concurrency), and the benchmark it would gate on (or an explicit "no benchmark coverage" flag, if none exists). Each also carries a narrow test pattern so that the correctness gate stays cheap. The test pattern field exists because of a specific incident; an exploration task omitted it, and the resulting exploitation task ran a very heavy test suite on every experiment. The fix was to change the upstream artifact rather than add an instruction downstream. This is a pattern we use repeatedly; make and record decisions as early as possible rather than re-derive them each time.</p><h3>Validating a new benchmark before it can gate anything</h3><p>Where benchmark coverage is genuinely missing, a dedicated benchmark task authors one, and that new benchmark has to pass validity checks before anything can rely on it. Two of them are mechanical: what fraction of the benchmark’s hot self-time comes from frames that actually appear in the production profile and whether the parameters fall inside the input shapes we measured. The third is a checklist that the agent has to attest item by item. It exists to avoid the JIT getting a simpler world than production.</p><ol><li><p>Inputs have to be reshuffled rather than fixed or sorted, so the branch predictor doesn’t get too good.</p></li><li><p>Results have to be consumed, or dead-code elimination deletes the thing that you meant to measure.</p></li><li><p>Inputs must not be compile-time constants, or they get folded away.</p></li><li><p>Call sites have to see roughly the product mix of types, because a monomorphic call site inlines, whereas a megamorphic one doesn’t.</p></li></ol><p>A human then approves it into a hash-pinned registry. Until that happens, it's inert, because task setup refuses any task citing an unapproved benchmark. We think of that approval as the strongest gate in the system, and it's deliberately placed. An approved benchmark can decide accept or reject in every future task, so it's the one place where we ask for a human signature on an artifact rather than on a decision.</p><h3>The validation ladder</h3><p>Underneath all of this sits a hierarchy of feedback mechanisms with the property that each rung is cheaper and weaker than the one below it, and the cheap tiers are for rejection only.</p><p><strong>Tier</strong></p><p><strong>Cost</strong></p><p><strong>Role</strong></p><p>probe</p><p>~20–60 s</p><p>Directionally right? Can never accept</p><p>codegen capture</p><p>~4 min</p><p>Why did that happen?</p><p>screen</p><p>~5–15 min</p><p>Cheap statistical filter</p><p>confirm</p><p>~20–60 min</p><p>The accept decision</p><p>end-to-end</p><p>hours</p><p>Regression guard, advisory</p><p>The asymmetry between accept and reject is doing real work here. A probe is a single paired fork, whereas the accept predicate requires a full confirm run with a matching calibration record. Agents are very good at telling believable stories; indeed, they're trained on many tasks judged by both LLMs and by humans, so being convincing is actively rewarded. We don't want an agent to be able to promote a cheap signal into a decision by being persuasive about it.</p><p>For this task benchmark, wall-clock is a real cost consideration. A confirm run might take an hour. So one has to weigh carefully all the costs involved when choosing the setup. A model that lands one hypothesis in four typically beats a cheaper one landing one in 10 by a margin on end-to-end metrics. The usual instinct to down-spec the model on a long-running loop is exactly backward here. When your loop has an uncertain outcome and significant costs beyond the tokens it consumes, you may well find yourself in the same situation.</p><h2>How do you know a performance improvement is real?</h2><p><em>Is it faster?</em> is a statistical question. So we made the accept predicate code rather than judgment and put it somewhere the agent can't bypass. There are four ideas in the accept decision, and in each case, the alternative we rejected is as informative as the choice we made.</p><h3>Forks are the statistical unit</h3><p>Each <a href="https://github.com/openjdk/jmh">JMH</a> fork collapses to its mean, and verdicts come from an exact <a href="https://en.wikipedia.org/wiki/Mann%E2%80%93Whitney_U_test">two-sided Mann-Whitney U test</a> at α = 0.05 plus a seeded bootstrap confidence interval over three to five fork means per side. The reason is that iterations within a fork share JIT and heap state and are therefore autocorrelated, so treating them as independent samples manufactures significance out of nothing. We rejected comparing single-run scores, which is pure noise, and iteration-level <a href="https://en.wikipedia.org/wiki/Student%27s_t-test">t-tests</a>, which can be confidently wrong. Seeding the bootstrap means that a rerun reproduces the verdict exactly because the agent needs to be able to tell the difference between a result that changed and a result that was never stable.</p><h3>Pair candidate and baseline in time</h3><p>The screen (three forks, optionally over a subset of parameters) exists only to kill bad hypotheses in maybe 10 minutes instead of an hour. The accept decision itself comes from a confirm run that measures candidate and baseline back to back using a stash-flip. In an unstable environment, thermal and background drift only cancels if both sides ran under the same conditions, so an hours-old baseline is really a different experiment.</p><h3>The noise floor is measured, not assumed</h3><p>The minimum effect size that a task will accept has to clear the A/A-calibrated coefficient of variation for that specific benchmark on that specific machine, and both the confirm run and the comparison refuse to proceed without a matching calibration record. A 1–2% improvement on a laptop is indistinguishable from noise, and because the floor varies by benchmark and by JDK, any global constant you pick will be too loose somewhere and too tight somewhere else.</p><h3>The accept rule is composite and deliberately conservative</h3><p>A parameter combination counts as improved only if p &lt; α, the effect clears the calibrated floor, and the confidence interval excludes zero. Overall acceptance then requires that nothing regressed (not the primary benchmarks and not the guards) and that at least one primary combination improved. The headline figure is the <a href="https://en.wikipedia.org/wiki/Geometric_mean">geometric mean</a> of the per-combination speedup ratios, which is always positive and composes across experiments, so a task's cumulative improvement is a meaningful number rather than a sum of incomparable percentages.</p><h3>What must not get slower</h3><p>Guards deserve a note of their own, because they answer a different question to the primary benchmarks; not <em>Did this get faster?</em> but <em>What must not get slower while it does?</em> The task definition lists them separately for that reason, and they're typically the operations and the input regimes that the change isn't aimed at. If you're optimizing insert throughput on a hash table, iteration is a guard and so is a collision-heavy key distribution. A change that improves the common case by weakening the hash function will look good on uniformly distributed keys while catastrophically degrading more adversarial inputs. We know that because it happened; the collision distribution was missing from the matrix that accepted one of our early experiments, and the task now carries a comment telling future readers never to drop it again for a hash-quality-sensitive scope. While guards cost wall-clock on every confirm, they also surface edge-case regressions, and omitting them can be much more costly in the long run.</p><p>Runtime isn't the only thing worth guarding, either. The confirm run also captures normalized allocation rate on both sides and flags any change that buys speed with more than about 15% extra garbage. That one is advisory rather than blocking, because sometimes the trade is the right one. However, it's the kind of regression a purely time-based accept rule would happily wave through but might raise a red flag to an experienced performance engineer with better understanding of the calling context.</p><h3>The end-to-end gate is one-sided and default open</h3><p>The end-to-end gate is a different statistical problem: small n, high noise, and a very strong prior that we should accept based on our microbenchmark results. Our first design treated accept and reject on an equal footing, and it produced multiple clearly spurious rejections, so the redesign is one-sided and default open. An operation is flagged only when the median regression exceeds the threshold and every candidate repetition is slower than every base repetition. That full-separation criterion is the nonparametric one-sided test at this sample size, and it's robust to the single outlier repetition that would occasionally fool us otherwise. A flag then also has to be corroborated against a differential CPU flame graph, where only a rise in the task's own in-scope CPU share counts as real. Near misses get reported for transparency but don't trigger triage, and nothing is ever auto-rejected; a flag is a request for human attention rather than a verdict. We have a final backstop which is the large suite of performance tests that we already run against Elasticsearch on a daily basis.</p><p>The one-sided gate lesson generalizes beyond benchmarking. A noisy gate should be one-sided and default open where there is strong prior reason to accept. A symmetric threshold on a noisy signal doesn't just cost you real wins, it teaches the loop to distrust its own instruments, and that’s a much more expensive failure.</p><h2>Exploration and exploitation need different permissions</h2><p>Exploration and exploitation might look like two phases of one activity, but they have different inputs (a macro workload versus a pinned scope) and different outputs (ranked opportunities versus commits). They also have different failure modes, which means they want different permissions. We made the split a first-class property of a task, which lets us enforce it; an exploration task literally cannot commit. The baselining, comparison, checkpointing, and validation CLI all refuse exploration tasks, benchmarking allows probes only, and every probe diff is always reverted. A broad, speculative survey is safe because nothing it does can edit the code.</p><p>A nice ancillary benefit is prompt focus. Each type reads one playbook, in full, with the others explicitly not loaded. If you try to write a single document covering both "find where the headroom is" and "land a validated win inside this scope", you get something that does neither well, because the instructions for good exploration (follow the profile, widen the net, a broad survey is preferred) are close to the opposite of the instructions for good exploitation (one hypothesis at a time, minimal diff, never widen scope).</p><h2>AI agent memory: Journals, knowledge bases, and postmortems</h2><p>Sessions are ephemeral, but what you can learn from them isn't, so the harness accumulates three durable assets, plus one disposable view derived from them. These are a journal of the code changes we’ve tried, a knowledge base of how the code performs, postmortems of when the harness failed and, because sessions can be stopped and resumed, a session summary. What makes them work together is a clean ownership rule about which kind of fact goes where.</p><p>The journal records what we tried and measured. It's append-only, one file per task, and one record per experiment, and it's written before the code is edited. Rejections carry a forward-looking note in the form "do not retry X because Y", which is probably the highest value line, because it's what stops the next session re-deriving a dead end. Records also carry the environment and the driving model, which means that hypothesis hit rates are comparable across models.</p><p>The knowledge base records how the code works and how it performs. It's an indexed collection of per-area summaries, each stamped with the commit it was written against, and every playbook ends with an upkeep step that appends whatever durable facts the run turned up. Because the performance characteristics of the JDK also change from time to time, for example, a new <a href="https://download.java.net/java/early_access/loom/docs/api/jdk.incubator.vector/jdk/incubator/vector/Vector.html">Vector API</a> might implement vector masking more efficiently on AArch64, findings from profile data are also tagged with the JVM version they apply to.</p><p>The postmortems record mistakes that the agent has made in the past, and they're indexed by symptom rather than by date. The question a session actually has when a number looks wrong is <em>Have I seen this shape of wrongness before?</em>, and a chronological list doesn't answer it. So the rows read like "validation fails on operations structurally unrelated to your diff", or "screen reports no matching combinations".</p><p>Keeping the journal and the knowledge base distinct sounds pedantic but isn't because without the rule, both of them turn into a diary that has a tendency to bloat the context window or miss critical information in context.</p><p>The disposable state is a session handoff, and our advice is to never write it by hand and to avoid asking a model for a session summary, if possible. Our harness regenerates a one-page digest mechanically from the journal, so a resumed session doesn’t spend time and tokens reconstructing where the task had got to. Because it’s derived rather than authored, it can’t drift from the record in the way hand-maintained content does. That’s also why it isn’t part of the audit trail; because it’s cheap to regenerate it from the journal.</p><p>Mechanical session handoffs point to two lessons, and they turn out to be the same one. Every time a person appears in the loop, it’s a source of friction and an opportunity for error. And every time you reach for a model, ask whether code can do the same job. This seems like an odd thing to advocate in a project whose primary premise is delegating to a model, but the habit is easy to fall into once you have one to hand. Judgment is expensive, wherever it happens to sit, so the person and model both have to earn their place on merit.</p><p>Two further things are critical for durable agent memory. The first is that knowledge rots, so you have to lint it. Elasticsearch's main branch moves daily, which means the knowledge base's file citations go stale, so one linter checks them. Another checks that every command, flag, and path cited in the agent-facing docs actually exists, that every postmortem is linked from the index, and that every task type has a playbook, and it runs as part of the test suite. Treating prose written for an agent as a testable artifact is the reason it stays true. The second is that loading discipline is half of memory. The instruction is to load the index and then the one or two summaries matching the task's scope; never to bulk-load the rest. Memory you can't afford to read isn't memory.</p><h2>Coding agent guardrails: Containment, scope, and stop conditions</h2><p>Elasticsearch is millions of lines of code, and an unscoped "make it faster" run against a codebase that size is unreviewable and unfalsifiable. It’s also expensive. The converse is that a verifier only protects what it can see, so we have to apply the same boundary to what it can change. The outcome is a task that fixes its goal, allowed paths, benchmarks, thresholds, and stop conditions before any code is edited, and none of those are things the agent may change during a run.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2bae81c11c37e29b/6aa3e867d1556f62a67269fc/unnamed.png" alt="Coding agent guardrails: human owns scope and thresholds, agent owns judgement, atune CLI enforces mechanically" /><p><em>The human owns the safety envelope, defining scope, thresholds, and stop conditions before any code is edited, and owning every handoff that crosses a trust boundary. The agent session owns judgment: reading profiles, forming one hypothesis at a time, editing inside scope. The atune CLI owns mechanical enforcement: scope checks, correctness gates, statistics, calibration, stop conditions, environment checks, and the signals. Persisted state is the record: a worktree pinned at a base commit, an append-only journal, and generated reports.</em></p><p>Containment happens in two layers. The first layer is coarse and task-agnostic; it’s a single static permissions file lets the session edit the per-task worktree (which includes a local branch of Elasticsearch, all journal entries, and CLI artifacts) and the knowledge base, and it denies the pristine Elasticsearch clone, the task definitions, the harness config, and the audit trail. The second is precise and per-task: a git-level scope check that runs over tracked and untracked files before any build, so an out-of-scope edit is rejected before it can be benchmarked or committed.</p><p>Two containment layers have a nice corollary. Widening the first layer to the whole worktree costs nothing, because nothing out of scope survives the second one. Coarse containment plus precise scope beats trying to make a single mechanism do both jobs, which is what we tried first and which left us with a permissions file that needed editing for every new task.</p><p>Stop conditions are mechanical. There's a maximum number of experiments and of consecutive rejections, along with a cumulative improvement target, and proposing a new experiment is refused once one of them fires. A human can override; the agent can't. That asymmetry is what makes the gates independent of which model happens to be driving.</p><p>Tests are add-only. New test files ship with the checkpoint, and modifying or deleting an existing test is blocked by the scope check. This closes the single most tempting shortcut in the entire problem space by construction rather than by instruction, which seems like the right way to handle any shortcut you'd otherwise have to keep asking an agent not to take.</p><p>The harness lives outside the code it optimizes. The Elasticsearch clone is a separate, gitignored directory, and each task gets a worktree pinned at a base commit. The payoffs compound; the subject repo stays pristine and upstream mergeable with nothing related to the harness leaking into a PR, and the audit trail is versioned independently of a codebase that moves daily. Also, tasks are isolated from each other and from any developer checkout. Targeting a newer Elasticsearch means creating a new task rather than repointing an old one, because that task's numbers are tied to its base. It also means that the harness is retargetable in principle, since the Elasticsearch-specific parts are configuration, knowledge base, and benchmark registry rather than architecture.</p><h3>The decisions the agent never makes</h3><p>The rule we settled on is that the agent runs the loops and a human owns every step that crosses a trust boundary. That is creating work, blessing a measurement instrument, publishing a branch, or acting outside the repo. None of those are in the agent's allowlist, and each has a reason worth stating:</p><ul><li><p>Humans have to sign off on the task because the thing being constrained can't set its own limits.</p></li><li><p>Calibrating the noise floor needs to be done once per benchmark, per machine, and by default is measured rather than assumed. However, it's rather expensive and we allow a human to override if they know the environment well.</p></li><li><p>Deciding what to optimize next by promoting an opportunity is a judgment and a new scope.</p></li><li><p>Since benchmarks go on to gate other tasks, we consider reviewing this artifact part of the correctness safety net.</p></li><li><p>We leave outward-facing actions, such as pushing a branch or filing an issue, to a human until we're confident in the process.</p></li><li><p>We allow actions to be forced, but the override has to sit outside the thing being overridden.</p></li></ul><p>How the human actions get surfaced in the workflow matters. The generated report and the session handoff both print the human actions currently due, at the moment they become due, rather than leaving them to be inferred from the playbook.</p><p>We're deliberately not taking a position on how permanent the manual processes are. The right amount of supervision for a new technology is an empirical question, and we'd rather measure it than argue about it. We’ve started with a relatively high degree of supervision because that's the cheap direction in which to be wrong (a gate you never needed is easier to remove than a regression you shipped) and because the harness makes the question answerable. Every gate is a named, logged transition, so over time we can see which of them ever changed an outcome and which only ever cost friction. In summary, measure first, and then refine.</p><h2>Building the harness is the same kind of loop</h2><p>A lot of the harness design didn't fall out of an initial design document. The signal set, the ranking function, the shape of a task, and the exact wording of a playbook rule each came from watching a run go wrong. If there's one piece of advice here that generalizes, it's to use the thing before it's ready and to instrument your own disappointment.</p><p>The clearest example is a rule we now call <em>distrust surprising results</em>. A validation run reported that every operation had regressed, the worst of them by 14.8%. It was wrong twice over. A target operation pattern had overmatched a completely different code path, and a stale output directory from an earlier run was being read alongside the new one. Offered a coherent story, the agent took it and reverted a change that was actually good.</p><p>What went into the playbook after that incident is not "be careful." It's a three-step check to run before acting on a surprising verdict:</p><ol><li><p>Trace the code path, and confirm that the thing which moved can even reach your diff.</p></li><li><p>Read the raw per-repetition data rather than the summary, and recompute one headline number by hand.</p></li><li><p>Compare the report's shape against a known good run, because a structurally different report implicates the pipeline rather than the code.</p></li></ol><p>Alongside that, there’s another important rule, which is if the agent concludes that the harness is buggy, it must <em>not</em> fix it mid-run, because a mid-run harness change makes every result in that run incomparable. It should journal the evidence and stop.</p><p>Improving the harness is itself a loop worth describing. Asking the model to review its own transcripts and the harness documents, and to propose the rule itself, usually works well. It's good at spotting where its own instructions were ambiguous, in a way that's hard to reproduce by rereading the instructions yourself. What makes that output useful is having somewhere for it to land: a terse rule in the playbook, the narrative in a dated postmortem, a symptom keyed index row, and a linter that keeps the citations honest.</p><p>Restraint turns out to be part of the same discipline. The design document carries an explicit list of extension points that we've deliberately not built, because the need for them is still speculative. That's the same "don't guess, wait for evidence" rule we impose on the optimization loop, applied to ourselves.</p><h2>How this applies beyond performance optimization</h2><p>A few of these themes aren't specific to performance work or to Elasticsearch.</p><p>Verifiable work is the current frontier. The same insight drives <a href="https://arxiv.org/pdf/2411.15124">reinforcement learning with verifiable rewards</a>, and it’s what <a href="https://arxiv.org/pdf/2506.13131">AlphaEvolve</a> is built around. The tasks agents are consistently good at are the ones that come with a cheap oracle (tests, compilers, benchmarks), and the interesting move isn't finding more such domains but manufacturing oracles for domains that lack them. Performance is an instructive case precisely because the oracle is, in some senses, obvious, and yet building it still took most of the engineering.</p><p>The referee pattern generalizes, too. Separating a fallible optimizer from mechanical enforcement is the same shape as sandboxed execution and policy engines: judgment in the model, invariants in code. The practical consequence is that the system's safety doesn't depend on which model drives it. A weaker model wastes benchmark time, but it can't corrupt the code or accept a bogus win.</p><p><a href="https://en.wikipedia.org/wiki/Goodhart%27s_law">Goodhart</a> is a standing adversary for any optimization task that uses an agent, and it has a <a href="https://arxiv.org/pdf/2209.13085">formal treatment worth reading</a>. An agent optimizes <em>exactly</em> what you tell it to, so the two-tier benchmark structure, add-only tests, benchmark approval registry, adversarial guard distributions, and a marker file that makes timing an instrumented build mechanically impossible are all one design theme wearing different clothes.</p><p>Tools are context engineering. The <a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">consensus is drifting away from "expose everything"</a> and toward a few well-shaped <a href="https://www.anthropic.com/engineering/writing-tools-for-agents">tools that return digests</a>: the principles of progressive disclosure, self-documenting interfaces, and verdicts rather than payloads.</p><p>Memory is becoming architecture. A rules file, per-task playbooks, a durable knowledge base, and an append-only journal form a hierarchy with different lifetimes, owners, and loading rules, and the hard part is eviction and staleness rather than storage; hence, the linters.</p><p>Finally, human-in-the-loop is a dial rather than a switch, so where it should sit is something to measure per domain rather than assert.</p><h2>What's in part 2 of this post</h2><p>The harness design is a set of hypotheses about what autonomous performance work needs, and the harness was built so that we could test them. Part 2 is that test: the first four PRs we raised using it, their gains, how many hypotheses it took to get each one, which gates actually caught something, and where the harness got in its own way. That includes the changes which didn't survive end-to-end validation, since as usual, the rejections are as informative.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-code-optimization-elasticsearch-agent-harness</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-code-optimization-elasticsearch-agent-harness</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Thomas Veasey,Chris Hegarty]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf9efb1bf91e4491b/6aa3e56cd909f868f36d69a6/unnamed.png" length="0" type="image/png"/>
    <pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Know your facts: How Elasticsearch AI Indices let agents skip the reading and keep the answer]]></title>
    <description><![CDATA[A technical walkthrough of precomputing facts into an Elasticsearch AI Index, so agents answer from a single ES|QL query instead of reading whole documents, with fewer tokens and lower latency.]]></description>
    <content:encoded><![CDATA[<p>Pulling whole documents into an agent's context to answer one question is expensive, and the cost compounds with every miss. In this walkthrough, we precompute the facts instead. A Kibana workflow distills each document into a fact-level Knowledge Indicator (KI), stored in an Elasticsearch AI Index and retrieved with a single Elasticsearch Query Language (ES|QL) query. On the same question, an agent answering from KIs reached the same grounded answer using fewer tokens and lower latency than reading raw documents, without loading a single full document into context. These facts are precomputed once and then stored for use by future agents when they encounter similar queries. This is Part 2 of our series on building context with AI indices; <a href="https://www.elastic.co/search-labs/blog/ai-index-building-context-agents">Part 1</a> covered routing agents to the right index.</p><p>Managing context depends on good retrieval. Rather than have agents rediscover the same content for every question, burning tokens by retracing similar steps over and over again, Elastic’s agentic AI capabilities enable us to precompute these details and store them in a structured, searchable form, and they let agents load that context directly. We call this precomputed unit of context a Knowledge Indicator.</p><p>The default agentic retrieval augmented generation (RAG) pattern does the opposite. It retrieves whole documents and dumps them into the model's context at query time, paying for that retrieval in tokens and latency on every single question. Precomputing the answer as a KI moves that cost out of the hot path and does it once.</p><h2>How it works: AI Index, Kibana Workflows, and the query-ki skill</h2><p>Building context through AI indices has three main parts: the AI Index (a special Elasticsearch index where KIs live), Kibana Workflows to create your KIs, and a <code>query-ki</code> skill to help agents directly query KIs using ES|QL: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt201b0bf84c5f5002/6a8fef16ecdaa77015050aa9/unnamed.png" alt="AI Index architecture: Kibana Workflows write Knowledge Indicators, agents read them via the query-ki ES|QL skill" /><p>This blog post is similar to Part 1 in that we’re using the same core building blocks. But in this post, we’re demonstrating a very different use case. Instead of precomputing index metadata, we’re distilling specific <em>facts</em> from our indexed documents that may be used to directly answer agents’ questions without subsequent searches. We've also provided a <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/precomputed-context-technical-walkthrough-part-2/index-facts-kis.ipynb">notebook</a>, if you'd like to create the same KIs yourself, end to end, as you go through these examples. </p><h3>Prerequisites: Elasticsearch Serverless and an LLM API key</h3><p>This tutorial assumes you have:</p><ol><li><p>An Elasticsearch Serverless project. You can <a href="https://cloud.elastic.co/registration?onboarding_token=search&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">sign up for a trial</a> if you don't have one.</p></li><li><p>An API key to access your Elasticsearch project.</p></li><li><p>An OpenAI-compatible large language model (LLM) API key, to access AI indices via Deep Agents scripts.</p></li></ol><h2>Load the BrowseComp-Plus sample corpus into Elasticsearch</h2><p>First, we’ll need some sources. Sources can be data that already exists in your Elasticsearch indices or external data accessed via connectors or ES|QL data sources. 
For this blog, we’ll create an index, <code>browsecomp-plus</code>, to hold our example data, with the following mappings:</p>{
  "browsecomp-plus": {
    "mappings": {
      "_meta": {
        "description": "BrowseComp-Plus corpus: ~100k human-verified web documents (news articles, Wikipedia entries, institutional pages) used as a reasoning-intensive browsing/QA retrieval benchmark. BM25-only index."
      },
      "properties": {
        "docid": {
          "type": "keyword",
          "meta": {
            "description": "Stable corpus document id."
          }
        },
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document text: title, date, and body content."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document title (from the document's front matter)."
          }
        },
        "url": {
          "type": "keyword",
          "meta": {
            "description": "Source URL the document was crawled from."
          }
        }
      }
    }
  }
}<p>and populate it with a small sample of BrowseComp-Plus data via the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk"><code>_bulk</code> API</a>. You can use the supporting <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/precomputed-context-technical-walkthrough-part-2/index-facts-kis.ipynb">notebook</a> to load a sample of this data in your project. </p><h2>Create the AI Index that stores your KIs</h2><p>Just like in Part 1, the first step is to create an AI Index:</p>PUT ai-index-idx-my-corpus<p>This is preconfigured with the same required mappings as we listed out in Part 1. We perform hybrid search here using <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_text</code></a> out of the box.</p><h2>How agents retrieve KIs using ES|QL</h2><p>A KI is a document in the AI Index. What makes KIs useful is <em>retrieval</em>, or querying the AI Index to find the right content. This query is packaged within a small, portable skill that’s harness-agnostic and can be run in any agent harness. </p><p>Here’s a sample <code>query-ki</code> skill:</p> ---
name: query-ki
description: &gt;-
  Retrieve Knowledge Indicators (precomputed context) from the Elasticsearch AI
  Index before answering. Use it to find which index to search (routing profiles)
  or to look up precomputed facts without reading source documents. Trigger on any question that depends on specific facts, names, dates, or on choosing a data source.
allowed-tools: esql_query
---

# Retrieving Knowledge Indicators

Knowledge Indicators (KIs) live in Elasticsearch indices named <code>ai-index-*</code>.
Retrieve them by calling the <code>esql_query</code> tool with the query below. Substitute
the user's question for <code>&lt;query&gt;</code>, and <code>corpus_entry</code> as the <code>&lt;ki_type&gt;</code> for facts.

```esql
FROM ai-index-idx-* METADATA _id, _index, _score
| WHERE type == "&lt;ki_type&gt;"
| FORK
    (WHERE MATCH(content, "&lt;query&gt;") OR MATCH(description, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
    (WHERE MATCH(content.semantic, "&lt;query&gt;") OR MATCH(description.semantic, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
| FUSE
| SORT _score DESC
| KEEP title, content, description, tags
| LIMIT 5
```

Ground your answer in what the query returns, and cite the KI titles you used. If
nothing relevant comes back, say so rather than guessing.<p>Save this as<code>skills/query-ki/SKILL.md</code>.</p><p>Here’s what this skill is doing: </p><ul><li><p>We’re defining <code>corpus_entry</code> as our KI use case.</p></li><li><p>We’re performing a hybrid ES|QL search on our AI indices, filtering by the appropriate <code>type</code>, using reciprocal rank fusion (RRF) as the default method to fuse results.</p></li><li><p>The KI results will directly ground the agent’s answer when determining what facts are relevant to the users’ query.</p></li></ul><p>When we say that AI indices and KIs are <em>harness-agnostic</em>, it’s because the skill is just instructions plus a query. It will work in Elastic Agent Builder, a Kibana workflow agent, Claude Code, or any other harness. We’ll be using Deep Agents for examples of how to query it outside the Kibana ecosystem. Since an AI Index is, at its core, an Elasticsearch index, you can also explore your data directly. </p><h2>Precompute facts as KIs for agentic RAG</h2><p>In this example, we extract actual facts so agents can retrieve an answer without consuming a full document. We generate one fact-based KI per selected document, though the actual number and structure of KIs you generate are completely customizable.</p><p>We'll use a sample of the <a href="https://github.com/texttron/BrowseComp-Plus">BrowseComp-Plus</a> corpus, indexed into a <code>browsecomp-plus</code> index, with <code>docid</code>, <code>url</code>, <code>title</code>, and <code>text</code> fields.</p><h3>Baseline: Retrieving whole documents with RRF</h3><p>As a baseline, here's a simple <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion">RRF</a> query:</p>POST /_query?format=txt
{
  "query": """
    FROM browsecomp-plus METADATA _score, _id, _index
    | FORK
        (WHERE match(title, "What was the actress who played Torvi from Vikings also known for?") | SORT _score DESC | LIMIT 100)
        (WHERE match(text,  "What was the actress who played Torvi from Vikings also known for?") | SORT _score DESC | LIMIT 100)
    | FUSE // uses RRF by default
    | SORT _score DESC
    | KEEP _id, title, text
    | LIMIT 10
  """
}<p>This drops several hundred words of raw body text into the model's context. It may work, but it's expensive, and the cost compounds with every miss.</p><h3>Build the Kibana workflow</h3><p>The workflow below reads a batch of documents with a single ES|QL query and writes one fact-level KI per document into the AI Index. Each iteration runs two steps: <code>generate_ki</code> distills a raw document into a structured KI, and <code>sink_ki</code> writes it to the AI Index keyed on <code>docid</code> so reruns are idempotent.</p><p>Copy and paste the following YAML into the <a href="https://www.elastic.co/docs/explore-analyze/workflows">Elastic Workflows</a> editor:</p>version: '1'
name: browsecomp-plus-doc-ki
description: Query the BrowseComp-Plus corpus with ES|QL, generate a KI per doc with an AI agent, and bulk-write each into the AI Index as a corpus_entry.
enabled: true
tags:
  - precomputed-context
  - browsecomp-plus
triggers:
  - type: manual
steps:
  - name: query_corpus
    type: elasticsearch.esql.query
    with:
      # WHERE drops empty bodies and restricts to the curated KI_DOCIDS -- the
      # specific documents this example's question depends on -- so the workflow
      # generates only a handful of KIs instead of one per corpus document.
      # SUBSTRING keeps the prompt bounded (a full body would blow the context window).
      # Column order drives the foreach.item[N] indices:
      #   item[0]=docid  item[1]=title  item[2]=url  item[3]=text
      query: &gt;
        FROM browsecomp-plus
        | WHERE text IS NOT NULL AND docid IN ("11589", "50639", "64501", "41758", "57766", "84983", "82008")
        | KEEP docid, title, url, text
        | EVAL text = SUBSTRING(text, 1, 12000)

  - name: loop_corpus_docs
    type: foreach
    foreach: '{{ steps.query_corpus.output.values }}'
    steps:
      # Turn the raw doc into a retrieval-optimized Knowledge Indicator.
      - name: generate_ki
        type: ai.agent
        timeout: 300s
        with:
          message: &gt;
            You are a knowledge engineer building a Knowledge Indicator (KI)
            for an enterprise document-retrieval corpus. A KI is a compact,
            high-signal record that a hybrid (BM25 + semantic) search engine
            and an AI agent use to FIND and JUDGE the source document without
            reading it in full.

            Read the document below and extract a faithful, richly structured KI.
            Follow these rules strictly:
            - Be 100% grounded: never state anything not supported by the text.
            - Prefer concrete, named specifics (people, organizations, products,
              dates, places, figures) over vague phrasing.
            - Write for retrieval, not prose flourish. No marketing language.
            - If a field cannot be determined from the text, return an empty
              string or empty array rather than guessing.

            Document ID: {{ foreach.item[0] }}
            Original Title: {{ foreach.item[1] }}
            Source URL: {{ foreach.item[2] }}
            Document Body:
            {{ foreach.item[3] }}
          schema:
            type: object
            properties:
              title:
                type: string
                description: A concise, specific, human-readable title (&lt;= 12 words).
              summary:
                type: string
                description: A dense 3-5 sentence factual summary capturing the document's main claims, named entities, and conclusions. PRIMARY semantic search surface.
              answers_questions:
                type: array
                items:
                  type: string
                description: 2-5 natural-language questions this document can authoritatively answer.
              key_entities:
                type: array
                items:
                  type: string
                description: 3-10 salient named entities (people, organizations, products, places, dates) explicitly mentioned in the text.
              topics:
                type: array
                items:
                  type: string
                description: 3-8 short topic/category labels.
              tagline:
                type: string
                description: A single ultra-short phrase (&lt;= 6 words) as a quick-reference label.
            required:
              - title
              - summary
              - answers_questions
              - key_entities
              - topics

      # Direct bulk write to the AI Index. The explicit <code>index</code> action row sets
      # _id = docid so re-runs upsert in place (idempotent). <code>index:</code> in <code>with</code>
      # supplies the default target index for the bulk request.
      - name: sink_ki
        type: elasticsearch.bulk
        with:
          index: ai-index-idx-my-corpus
          operations:
            - index:
                _id: '{{ foreach.item[0] }}'
            - '@timestamp': '{{ execution.startedAt | date: "%Y-%m-%dT%H:%M:%S.%LZ" }}'
              type: corpus_entry
              title: '{{ foreach.item[1] | default: steps.generate_ki.output.structured_output.title }}'
              tags:
                - browsecomp-plus
              references:
                uri: '{{ foreach.item[2] }}'
              attributes:
                docid: '{{ foreach.item[0] }}'
                url: '{{ foreach.item[2] }}'
                source_index: browsecomp-plus
                tagline: '{{ steps.generate_ki.output.structured_output.tagline }}'
                topics: '{{ steps.generate_ki.output.structured_output.topics | json }}'
                answers_questions: '{{ steps.generate_ki.output.structured_output.answers_questions | json }}'
                key_entities: '{{ steps.generate_ki.output.structured_output.key_entities | json }}'
              content: &gt;
                === SOURCE / PROVENANCE ===
                Backing Elasticsearch index: browsecomp-plus
                Document ID (docid): {{ foreach.item[0] }}
                Source URL: {{ foreach.item[2] }}
                Retrieve the full original document with ES|QL:
                FROM browsecomp-plus | WHERE docid == "{{ foreach.item[0] }}"
                === KNOWLEDGE INDICATOR ===
                {{ steps.generate_ki.output.structured_output.summary }}
                Questions this document answers: {{ steps.generate_ki.output.structured_output.answers_questions | join: " | " }}
                Key entities: {{ steps.generate_ki.output.structured_output.key_entities | join: ", " }}
              description: &gt;
                {{ steps.generate_ki.output.structured_output.tagline }}.
                Topics: {{ steps.generate_ki.output.structured_output.topics | join: ", " }}.
                Entities: {{ steps.generate_ki.output.structured_output.key_entities | join: ", " }}.<p>Here’s what this workflow is doing: </p><ul><li><p><code>query_corpus</code> runs an ES|QL query against the <code>browsecomp-plus</code> index, applying some rules, like dropping documents with empty bodies and trimming each body to 12,000 chars so the agent prompt stays inside the context window.</p></li><ul><li><p>Note: In this example, we’re cherry-picking some concrete KI IDs, because generating KIs for every document in the index would take a long time, and we want this exercise to be short for those following along.</p></li></ul><li><p><code>loop_corpus_docs</code> iterates over every returned document, running the following two steps per document: </p></li><ul><li><p><code>generate_ki</code> reads the document and calls an LLM to emit a strictly grounded, structured KI.</p></li><li><p><code>sink_ki</code> bulk-writes each KI into the AI Index (<code>ai-index-idx-my-corpus</code>) as a KI of type <code>corpus_entry</code>. It forces <code>_id</code> to be the same as the document’s <code>docid</code> so rerunning the workflow is idempotent.</p></li></ul></ul><p>To summarize, this workflow turns each raw corpus document into a compact, searchable metadata record that agents can find and judge without reading the full source into the context window.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5a37c99585270959/6a8ff23da1b20b401c8728c7/unnamed.png" alt="Kibana Workflow browsecomp-plus-doc-ki: query_corpus, generate_ki and sink_ki write a corpus_entry KI to the AI Index" /><p>This workflow is used for example purposes, and the same <code>foreach</code> caveat as in Part 1 applies. For scale, use <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/composition"><code>workflow.executeAsync</code></a> or native parallel support. The <a href="https://www.elastic.co/docs/explore-analyze/workflows/reference/cheat-sheet">cheat sheet</a> is useful for optimizing Workflows. There could also be cost and efficiency gains in production by using <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/ai-steps#ai-prompt"><code>ai.prompt</code></a> or by choosing different models with which to create KIs. </p><h3>Inspect the KIs in your AI Index</h3><p>Once the workflow runs, you can query the AI Index to browse what was written:</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt626b3ba4dc3a52c0/6a8ff268971ef9107f537cb5/unnamed.png" alt="ES|QL query in Kibana Discover returning five corpus_entry Knowledge Indicators from an Elasticsearch AI Index" /><p>Here’s an example of what one of the KI documents looks like: </p>{
  "_index": "ai-index-idx-my-corpus",
  "_id": "57766",
  "_version": 1,
  "_seq_no": 0,
  "_primary_term": 1,
  "found": true,
  "_source": {
    "@timestamp": "2026-08-05T20:35:39.034Z",
    "type": "corpus_entry",
    "title": "Vikings (TV series) - Wikipedia",
    "tags": [
      "browsecomp-plus"
    ],
    "references": {
      "uri": "https://en.wikipedia.org/wiki/Vikings_%28TV_series%29"
    },
    "attributes": {
      "docid": "57766",
      "url": "https://en.wikipedia.org/wiki/Vikings_%28TV_series%29",
      "source_index": "browsecomp-plus",
      "tagline": "Ragnar Lothbrok's rise and legacy",
      "topics": """["Historical drama television","Viking Age","Norse mythology and sagas","Canadian-Irish co-production","Television cast and production","Medieval Scandinavia"]""",
      "answers_questions": """["When did the Vikings TV series premiere and on which network?","Who created and wrote the Vikings TV series?","Where was the Vikings TV series filmed?","Who are the main cast members of Vikings?","What historical and literary sources inspired the Vikings TV series?"]""",
      "key_entities": """["Michael Hirst","Travis Fimmel","Katheryn Winnick","History Channel","Amazon Prime Video","Ashford Studios","County Wicklow, Ireland","Vikings: Valhalla","Ragnar Lodbrok","Wardruna"]"""
    },
    "content": """=== SOURCE / PROVENANCE === Backing Elasticsearch index: browsecomp-plus Document ID (docid): 57766 Source URL: https://en.wikipedia.org/wiki/Vikings_%28TV_series%29 Retrieve the full original document with ES|QL: FROM browsecomp-plus | WHERE docid == "57766" === KNOWLEDGE INDICATOR === Vikings is a historical drama television series created and written by Michael Hirst, co-produced between Canada and Ireland, that premiered on the History Channel on March 3, 2013, and concluded on March 3, 2021, after 6 seasons and 89 episodes. The series is inspired by the sagas of legendary Norse hero Ragnar Lodbrok — drawing on 13th-century texts Ragnars saga Loðbrókar and Ragnarssona þáttr, as well as Saxo Grammaticus' Gesta Danorum — and follows Ragnar's rise from farmer to Scandinavian king, then the exploits of his sons across England, Scandinavia, Kievan Rus', the Mediterranean, and North America. Principal cast includes Travis Fimmel as Ragnar Lothbrok, Katheryn Winnick as Lagertha, Gustaf Skarsgård as Floki, and Alexander Ludwig as Bjorn Ironside, among many others. The series was filmed entirely in Ireland at Ashford Studios and County Wicklow, with additional location shoots in Iceland, Morocco, Norway, and Canada; the first season budget was US$40 million. A sequel series, Vikings: Valhalla, premiered on Netflix on February 25, 2022. Questions this document answers: When did the Vikings TV series premiere and on which network? | Who created and wrote the Vikings TV series? | Where was the Vikings TV series filmed? | Who are the main cast members of Vikings? | What historical and literary sources inspired the Vikings TV series? Key entities: Michael Hirst, Travis Fimmel, Katheryn Winnick, History Channel, Amazon Prime Video, Ashford Studios, County Wicklow, Ireland, Vikings: Valhalla, Ragnar Lodbrok, Wardruna
""",
    "description": """Ragnar Lothbrok's rise and legacy. Topics: Historical drama television, Viking Age, Norse mythology and sagas, Canadian-Irish co-production, Television cast and production, Medieval Scandinavia. Entities: Michael Hirst, Travis Fimmel, Katheryn Winnick, History Channel, Amazon Prime Video, Ashford Studios, County Wicklow, Ireland, Vikings: Valhalla, Ragnar Lodbrok, Wardruna.
"""
  }
}<h3>Query KIs from LangChain Deep Agents</h3><p>We’ll use <a href="https://docs.langchain.com/oss/python/deepagents/overview">LangChain Deep Agents</a> with an OpenAI-compatible key to show that AI indices and KIs will work with any agent harness, inside and outside of Kibana’s Agent Builder ecosystem. </p><p>First, let’s create <code>facts_baseline_agent.py</code> to measure our baseline before applying KIs: </p># Example question: What was the actress who played Torvi from Vikings also known for?
import os
import sys
import time
from elasticsearch import Elasticsearch

from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM browsecomp-plus | LIMIT 5'.
               Full-text search syntax: WHERE MATCH(field, "value") — not field MATCH "value".
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


@tool
def get_mapping(index: str) -&gt; dict:
    """Return the field mapping for an Elasticsearch index or pattern."""
    return es.indices.get_mapping(index=index).body


baseline_agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query, get_mapping],  # no query-ki skill
    system_prompt=(
        "You are a research assistant answering questions about a document corpus "
        "stored in the Elasticsearch index <code>browsecomp-plus</code> (fields: docid, url, "
        "title, text). You have NOT memorized the corpus. Answer by querying the raw "
        "index directly with ES|QL via the esql_query tool. "
        "Full-text search syntax: WHERE MATCH(field, \"value\") — never use field MATCH \"value\". "
        "Use get_mapping if you are unsure of field names. Ground your answer strictly "
        "in the rows returned, and cite the docid or url you used."
    ),
)

start = time.perf_counter()
result = baseline_agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>When I ran this baseline query, <em>What was the actress who played Torvi from Vikings also known for?</em>, it output the following: </p>% python facts_baseline_agent.py "What was the actress who played Torvi from Vikings also known for?"

--- Tool calls ---
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Torvi Vikings actress") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Georgia Hirst") | LIMIT 10'}
  [read_file] {'file_path': '/large_tool_results/toolu_bdrk_01SVEwHdtc9uE1zdfgCwchb5', 'offset': 0, 'limit': 50}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Torvi Vikings TV") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Vikings series cast actress") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Michael Hirst Vikings") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Torvi") | LIMIT 20'}
  [read_file] {'file_path': '/large_tool_results/toolu_bdrk_01PS3ryzML5BstqMXnV4c65N', 'offset': 0, 'limit': 100}
Total: 8

--- Usage ---
Tokens: 386187 (input 384940, output 1247)
Latency: 44.86s

--- Answer ---
Perfect! I found the information. Based on the document from docid 11589 (Georgia Hirst's Wikipedia page), I now have the answer to the question.

**Georgia Hirst**, the actress who played Torvi in the History Channel series Vikings (2014-2020), is also known for:

1. **Her role as Becky in the zombie horror movie "Ravers" (2018)** - This film had its world premiere at the 2018 London FrightFest Film Festival.

2. **Being the daughter of Michael Hirst**, the creator and writer of Vikings. This is a notable connection that's frequently mentioned about her.

3. **Her role in the short film "Dungeness" (2018)**, directed by Stephen Gates, where she was the sole actor.

4. **Her role as Grace in the interactive rom-com "Five Dates" (2020)**.

5. **Her advocacy work** - In 2017, she made an appearance on Sky News urging young women to get smear tests after being diagnosed with precancerous cells at age 22.

6. **Her philanthropy** - She joined The Celebs in 2021 to record a cover of The Beatles' "Let It Be" in support of the Mind charity.

So while she's primarily known for her role as Torvi in Vikings (where she appeared in 63 episodes across seasons 2-6), she's also recognized for her work in horror films like "Ravers," her health advocacy, and being part of the Hirst family that created the show.<p>(Note: Deep Agents automatically adds the <code>read_file</code> tool to handle paginated tool results, which is why it shows up in the output.) </p><p>Next, let’s create an agent that knows how to use our <code>query-ki</code> skill, <code>facts_ki_agent.py</code>: </p># Example question: What was the actress who played Torvi from Vikings also known for?
import os
import sys
import time
from elasticsearch import Elasticsearch
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent
from deepagents.backends.filesystem import FilesystemBackend

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM ai-index-idx-* | LIMIT 5'.
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


# FilesystemBackend loads skills from disk, relative to root_dir.
backend = FilesystemBackend(root_dir=".", virtual_mode=False)

agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query],
    skills=["skills"],
    backend=backend,
    system_prompt=(
        "You are a research assistant answering questions about a document corpus. "
        "You have NOT memorized the corpus. When a question depends on specific facts, "
        "names, dates, or events, use the query-ki skill to retrieve Knowledge "
        "Indicators before answering. Ground your answer strictly in what it returns, "
        "and cite the KI titles you used."
    ),
)

start = time.perf_counter()
result = agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>This agent will query the KI indices to generate the answer, rather than pulling every document into context. </p><p>When I ran these results against the same query, here was the output: </p>% python facts_ki_agent.py "What was the actress who played Torvi from Vikings also known for?"

--- Tool calls ---
  [read_file] {'file_path': './query-ki/SKILL.md', 'limit': 1000}
  [esql_query] {'query': 'FROM ai-index-idx-* METADATA _id, _index, _score\n| WHERE type == "corpus_entry"\n| FORK\n    (WHERE MATCH(co
Total: 2

--- Usage ---
Tokens: 27625 (input 27037, output 588)
Latency: 15.22s

--- Answer ---
Based on the retrieved information, **Georgia Hirst** played Torvi in Vikings. She is also known for:

1. **Ravers (2018)** - A zombie horror film where she played the character Becky. This was her debut film and premiered at the London FrightFest Film Festival.

2. **Being the daughter of Michael Hirst** - The creator and writer of Vikings itself, making her connection to the show a notable family affair. Her older half-sister Maude Hirst also appeared in Vikings as Helga.

3. **Cervical cancer awareness advocacy** - She has publicly advocated for cervical cancer screening after being diagnosed with precancerous cells at age 22 and successfully recovering through treatment.

4. **Charity work** - In 2021, she participated in a celebrity cover of The Beatles' "Let It Be" in support of the Mind charity (a mental health organization), alongside Anne Hegerty, Ivan Kaye, Eunice Olumide, and Shona McGarty.

**Sources cited:** "Georgia Hirst" and "Georgia Hirst - Wikipedia" Knowledge Indicators from the AI Index.<h2>How much can precomputing facts reduce agent token usage?</h2><p>Both agents had similar conclusions, but they took far different paths to get there: </p><p>The same question and the same grounded answer result in 93% fewer tokens and two tool calls instead of eight, when answering from KIs.</p><p>
</p><p>Baseline (No AI Index)</p><p>With AI Index</p><p>Total tool calls</p><p>8</p><p>2</p><p><code>read_file</code> calls</p><p>2</p><p>1</p><p><code>esql_query</code> calls</p><p>6, all against the <code>browsecomp-plus</code> index</p><p>1, from <code>ai-index-idx-*</code></p><p>Tokens consumed</p><p>386,187</p><p>27,625</p><p>Latency</p><p>44.86s</p><p>15.22s</p><p>Answer</p><p>Grounded, correct</p><p>Grounded, correct</p><p>Exact tool call counts, latency, and answers will vary between runs and using different agents. </p><p>Both agents produced solid, grounded answers. The difference is cost. Querying KIs from the AI Index cut token use by 93% and cut latency by roughly two thirds. Here’s how both paths went, side by side:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt52fac74907997882/6a8ff2a049d4293b02a4fd64/unnamed.png" alt="Agentic RAG tool calls: 8 calls and 386,187 tokens without Knowledge Indicators, 2 calls and 27,625 tokens with them" /><p>That was in-depth, but it shows what AI indices and Workflows do together: the same answer, at a fraction of the tokens.</p><h2>Build precomputed context in Elasticsearch Serverless</h2><p>This walkthrough shows how to generate more sophisticated KIs based on documented facts and query them for knowledge retrieval use cases using Elasticsearch primitives. </p><p>Managing context is critical in agentic search systems. And at its core, context is a retrieval problem. AI indices help you manage context within the Elastic Stack. Try it out in Serverless, and let us know what you think in our <a href="https://discuss.elastic.co/top?period=monthly">Discuss forums</a> or the <code>#stack-kibana</code> channel in our <a href="https://elasticstack.slack.com/signup#/domain-signup">Community Slack</a>.</p><p>We’d also love to hear from you about what use cases you’d like to solve using AI indices.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agentic-rag-precomputed-facts-ai-index</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agentic-rag-precomputed-facts-ai-index</guid>
    <category><![CDATA[AI Tools ]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Kathleen DeRusso,Matt Nowzari ,Apostolos Matsagkas,Peter Pišljar]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3e1939e01169bb08/6a8fedbec8ced9f736055f59/1.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Let the big model think, let the small model work: Splitting LLM costs in Elastic Workflows]]></title>
    <description><![CDATA[Build an Elastic workflow that sends a data sample to a large model to propose classification labels. A human signs off, then a smaller model applies them across the full corpus.]]></description>
    <content:encoded><![CDATA[<p>Split the expensive part of large language model (LLM) classification from the cheap part. This article builds an <a href="https://www.elastic.co/docs/explore-analyze/workflows">Elastic workflow</a> where Claude Sonnet reads a stratified sample of NASA pilot incident reports and proposes classification labels based on what it finds. A human reviews the schema and signs off, and then <a href="https://mistral.ai/news/mistral-small-3-1/">Mistral Small 3.1</a> applies the labels across the full corpus. The routing is YAML, the results land in Elasticsearch as structured data, and the pattern works wherever you have free text that needs labeling.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0ceb64be1f3ad4e5/6a87ef8373f743fa6848688d/image2.png" alt="Example NASA ASRS pilot incident report showing a free-text narrative describing a near-miss at an uncontrolled airfield, the type of document classified by the LLM pipeline" /><p><a href="https://asrs.arc.nasa.gov/">NASA Aviation Safety Reporting System (ASRS)</a> reports describe unusual events during flights, such as missed altitudes, confusing clearances, runway issues, or mechanical problems. Each report already has an official category, like altitude deviation, course deviation, or ground encounter. In this article, we ask a different question: <em>What does this report reveal about the pilot who wrote it?</em> The idea is to ask a model to infer a schema grounded on the data to classify the report based on criteria that help us figure out information about the report writers. Then ask a second model to apply the labels.</p><p><em><strong>You can find the full workflow definitions and helper scripts </strong></em><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/larger-llms-task-planning-smaller-llms-execution"><em><strong>here</strong></em></a><em><strong>.</strong></em></p><h2>What you need to run this LLM pipeline</h2><ul><li><p>Elastic Stack 9.4+ or Elastic Cloud Serverless. Elastic Workflows has been generally available (GA) since 9.4.</p></li><li><p>Elastic Agent Builder enabled in your deployment.</p></li><li><p>A <a href="https://www.elastic.co/docs/reference/kibana/connectors-kibana/ai-connector">Kibana generative AI (GenAI) connector</a> pointing at Claude Sonnet (or an equivalent reasoning model). This is the planner.</p></li><li><p>A <a href="https://console.mistral.ai/api-keys">Mistral API key</a>. We’ll use it to register an Elasticsearch inference endpoint.</p></li><li><p>Python 3.10+ with <code>elasticsearch&gt;=9.0</code> and <code>pandas</code>. Used by the dataset loader.</p></li></ul><h2>How two-tier LLM orchestration works</h2><p>The workflow has two jobs: Decide what labels should exist, and then apply those labels to every report.</p><p><strong>The first job is open-ended.</strong> A large model reads a varied sample of reports and proposes a small schema of categorical fields. A field is one way to describe the writer, such as <code>attribution_style</code> or <code>procedure_orientation</code>. Each field has a few allowed values, such as <code>self_critical</code>, <code>system_attributing</code>, or <code>balanced</code>.</p><p><strong>The second job is repeatable.</strong> After a human approves the schema, a smaller model reads each report and chooses one value for each field.</p><p>We use Elastic Workflows because the steps are known ahead of time: sample reports, propose labels, wait for approval, classify every document, and store the results. Writing those steps in YAML makes the process reproducible, observable, and cheaper to rerun.</p><h3><strong>Why split LLM work across two model tiers?</strong></h3><p>A small model could handle classification, but schema discovery is a different shape of problem. It requires reading a diverse sample, spotting latent patterns, and proposing complex structures. In practice, smaller models over-anchor on surface keywords and produce redundant or nonexclusive fields.</p><p>Classification is simpler, the schema exists, the values are enumerated, and the task is to pick one per field. A smaller model handles this reliably and at a fraction of the cost, since it runs once per document across the entire corpus.</p><p><em>Large</em> and <em>small</em> here mean reasoning capability. In this article, Claude Sonnet plays the planner and Mistral Small 3.1 plays the executor.</p><h2>Classifying NASA pilot reports with a two-tier LLM pipeline</h2><p>We’ll use the NASA ASRS database, which collects voluntary, anonymous incident reports from pilots, controllers, and mechanics. The dataset is public, and the reports are written as free-text narratives.</p><p>What we want to ask is:</p><p><em>What does this report reveal about the pilot who wrote it?</em></p><p>The planner reads a varied sample of reports and decides which distinctions are meaningful based on how the reports are actually written.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6b7f0b9d35ef59e3/6a87efa5b6895193cca26e8e/image4.png" alt="Elastic Workflow pipeline diagram showing schema discovery by a large language model, human approval via waitForInput, and classification by a small language model storing results in Elasticsearch" /><p><strong>Step</strong></p><p><strong>Role</strong></p><p><strong>Model tier</strong></p><p><code>sample</code></p><p>Pull a diverse subset of reports from the corpus.</p><p>(no LLM)</p><p><code>discover</code></p><p>Read the question and the sample, propose a schema of fields with enum values.</p><p><strong>Large</strong></p><p><code>approve</code></p><p>Human reviews the proposed schema and approves or edits it.</p><p>(Human via <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/wait-for-input"><code>waitForInput</code></a>)</p><p><code>apply</code></p><p>Iterate over the corpus, assign one value per field to each report.</p><p><strong>Small</strong></p><p><code>store</code></p><p>Write the schema and the per-document field values to Elasticsearch.</p><p>(No LLM)</p><h2>Registering Mistral and Claude as Elasticsearch inference endpoints</h2><p>The small model will be registered as an Elasticsearch <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/infer-service-mistral.html">inference endpoint</a> using the native <code>mistral</code> service integration. </p>INFERENCE_ID = "mistral-small-extractor"

es.inference.put(
    task_type="chat_completion",
    inference_id=INFERENCE_ID,
    inference_config={
        "service": "mistral",
        "service_settings": {
            "api_key": MISTRAL_API_KEY,
            "model": "mistral-small-latest",
            # 6 RPM is conservative for the Mistral free tier to avoid 429s.
            "rate_limit": {"requests_per_minute": 6},
        },
    },
)<p>The alias <code>mistral-small-latest</code> resolves to <a href="https://mistral.ai/news/mistral-small-3-1">Mistral Small 3.1</a>. It has a 128k context window and supports JSON-mode output.</p><p>The large model will be an <a href="https://www.elastic.co/docs/reference/kibana/connectors-kibana/ai-connector">AI connector</a> pointing at Claude Sonnet. The Agent Builder UI walks you through creating the connector. Take a note of the connector ID since we’ll reference it from the workflow.</p><h2>Indexing NASA ASRS incident reports into Elasticsearch</h2><p>The ASRS dataset is indexed with keyword mappings for aggregation fields and text mappings for the narratives the models will read.</p><p>Download the ASRS CSV (the database publishes quarterly extracts at the <a href="https://asrs.arc.nasa.gov/search/database.html">ASRS Database Online</a> page), and index it. The mappings are:</p>{
  "properties": {
    "acn":          { "type": "keyword" },
    "flight_phase": { "type": "keyword" },
    "anomaly":      { "type": "keyword" },
    "synopsis":     { "type": "text" },
    "narrative":    { "type": "text" }
  }
}<p>The mapping types follow how each field is used. <code>flight_phase</code> and <code>anomaly</code> are mapped as <code>keyword</code> because we’ll run terms aggregations on them to build the sample, and aggregations need exact, non-analyzed values. <code>narrative</code> and <code>synopsis</code> are mapped as <code>text</code> because they hold free-form prose that the models will read. The companion notebook has the full loader script that reads the CSV and bulk-indexes the documents.</p><h2>Building a stratified sample for the planning LLM</h2><p>The YAML snippets in this and the following sections are steps of the workflow definition that the notebook registers via the Workflows API. The first two steps generate a representative sample: They aggregate by flight phase and by anomaly and pull a few documents per bucket with <code>top_hits</code>.</p>- name: by_phase
  type: elasticsearch.request
  with:
    method: POST
    path: "/incident_reports/_search"
    body:
      size: 0
      aggs:
        per_phase:
          terms:
            field: flight_phase
            size: 8
          aggs:
            sampled_docs:
              top_hits:
                size: 5
                _source: ["acn", "synopsis", "narrative"]

- name: by_anomaly
  type: elasticsearch.request
  with:
    method: POST
    path: "/incident_reports/_search"
    body:
      size: 0
      aggs:
        per_anomaly:
          terms:
            field: anomaly
            size: 8
          aggs:
            sampled_docs:
              top_hits:
                size: 3
                _source: ["acn", "synopsis", "narrative"]<h2>How the large LLM discovers a classification schema from the data</h2><p>The prompt needs both the question and the sample. A question alone may produce generic labels disconnected from the corpus, and a sample alone produces descriptive clusters that ignore the angle of the question. </p><p>When both are present and the output is structured, the model produces labels that are grounded in the data and oriented to the task: a schema of categorical fields, each with two to four mutually exclusive value options backed by evidence from the sample.</p><p>Here’s the planner step from the workflow:</p>- name: discover
  type: ai.prompt
  connector-id: "claude-sonnet"
  with:
    systemPrompt: |
      You design categorical schemas for use by downstream classifiers. A
      schema is a small set of fields, each with a few mutually exclusive
      values. Every value you propose must be grounded in evidence from the
      provided sample and must serve the stated question. You do not invent
      values that are not supported by at least two documents in the sample.
      You do not propose fields that a reasonable analyst could have written
      without reading the documents.
    prompt: |
      Question:
      ${{ inputs.goal }}

      Sample documents stratified by flight phase:
      ${{ steps.by_phase.output.aggregations.per_phase.buckets | json }}

      Sample documents stratified by anomaly type:
      ${{ steps.by_anomaly.output.aggregations.per_anomaly.buckets | json }}

      Propose between 2 and 4 categorical fields that:
      - serve the question (you can explain how)
      - depend on patterns visible in the sample (you can cite document IDs)
      - would not be obvious to someone who has not read the sample

      For each field, return: name (snake_case), definition, why_useful,
      and values (2 to 4 mutually exclusive options).

      For each value, return: value (snake_case) and definition.
    schema:
      type: object
      properties:
        fields:
          type: array
          minItems: 2
          maxItems: 4
          items:
            type: object
            required: [name, definition, why_useful, values]
            properties:
              name: { type: string }
              definition: { type: string }
              why_useful: { type: string }
              values:
                type: array
                minItems: 2
                maxItems: 4
                items:
                  type: object
                  required: [value, definition]
                  properties:
                    value: { type: string }
                    definition: { type: string }
    temperature: 0.3<p>The structured output schema enforces the shape of the response:</p><p> </p><ul><li><p><code>name</code>: Identifier for the categorical field.</p></li><li><p><code>definition</code>: What this field measures, in one sentence.</p></li><li><p><code>why_useful</code>: How this field serves the question; this also helps the downstream classifier understand the intent.</p></li><li><p><code>values</code>: Two to four mutually exclusive options. Each has a <code>value</code> and a <code>definition</code>.</p></li></ul><p>Here’s an example of the produced schema. We can see how the writer is being classified and the reasons why the model decided to create the category. <code>definition</code>and <code>why_useful</code> fields are used by the second model to classify the documents.</p>{
  "fields": [
    {
      "name": "attribution_style",
      "definition": "How the reporter frames responsibility for what happened.",
      "why_useful": "Surfaces reporting culture independent of the technical event. Useful for training and safety-management programmes that want to distinguish reporter style from incident type.",
      "values": [
        {
          "value": "self_critical",
          "definition": "Assigns the cause primarily to their own action, even when external factors clearly contributed."
        },
        {
          "value": "system_attributing",
          "definition": "Frames the cause as external: ATC, equipment, weather, or organisational factors."
        },
        {
          "value": "balanced",
          "definition": "Distributes responsibility across self and system without emphasising either."
        }
       ]
    },
    {
      "name": "procedure_orientation",
      "definition": "How the reporter relates to written procedure.",
      "why_useful": "Distinguishes pilots who frame events through SOPs from those who frame them through personal judgment.",
      "values": [
        // procedure_first, experience_first (same structure as above)
      ]
    }
  ]
}<h2>Human-in-the-loop schema approval with waitForInput</h2><p>The proposed schema is now passed to a person for approval. Elastic Workflows has a <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/wait-for-input"><code>waitForInput</code></a> step that pauses the workflow with a schema, exposes a form, and resumes when the input is submitted.</p><p><code>waitForInput</code> has no timeout of its own, so if nobody responds, the execution <a href="https://www.elastic.co/docs/explore-analyze/workflows/authoring-techniques/human-in-the-loop#what-happens-while-the-workflow-is-paused">waits indefinitely</a>. To put a limit on that, set a workflow-level <code>settings.timeout</code>; if it elapses before the reviewer submits the form, the execution is canceled.</p>- name: human_gate
  type: waitForInput
  with:
    message: "Review and edit the proposed schema. The approved fields will be applied across the full corpus."
    schema:
      type: object
      required: [approved_fields]
      properties:
        approved_fields:
          type: array
          items:
            type: object
            properties:
              name: { type: string }
              definition: { type: string }
              values:
                type: array
                items:
                  type: object
                  properties:
                    value: { type: string }
        notes:
          type: string<p>When the workflow reaches this step, the execution pauses and the Kibana UI shows an "Action is required" badge. Clicking <strong>Provide action</strong> opens a form where the reviewer can paste or edit the schema JSON. Since <code>waitForInput</code> cannot be prepopulated from a previous step, the code polls the <em>discover</em> step output and prints a paste-ready JSON block that can be copied directly into this form.</p>discover = step_output(execution_id, "discover")  # polls until the step completes

# Strip <code>why_useful</code> (not part of the human_gate form) and wrap in the shape
# expected by the waitForInput form so this is paste-ready.
approved_fields = [
    {
        "name": field["name"],
        "definition": field["definition"],
        "values": [
            {"value": v["value"], "definition": v["definition"]}
            for v in field["values"]
        ],
    }
    for field in discover["content"]["fields"]
]

print(json.dumps({"approved_fields": approved_fields, "notes": ""}, indent=2))<p>The <code>step_output</code> helper (in the notebook) polls the execution via <code>GET /api/workflows/executions/{id}</code> until the <em>discover</em> step completes and then returns its output.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf3771a69b51089e7/6a87efeca8b3236eb5cc01f5/image1.png" alt="Kibana execution view showing an Elastic Workflow paused at the waitForInput step with the Provide action button highlighted for human-in-the-loop schema approval" /><p>Code JSON output pasted on Kibana:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt85d91c93043ad0d1/6a87f00d6ea6da5cfe0083e0/image5.png" alt="Kibana Provide action modal displaying the approved classification schema JSON with pilot experience level fields, where a reviewer edits the schema before the workflow resumes" /><p>The reviewer can keep useful fields, rewrite unclear ones, merge overlapping values, and add notes. After approval, the workflow resumes and sends the final schema to the executor step.</p><p>For a new corpus, keep this human gate in place. Once the schema is stable, you can auto-approve and only fall back to review when it’s worth it: Route just the low-confidence extractions to a person, or compare a new discovery run against the schema stored in the <code>schemas</code> index and trigger review only when fields or values change beyond a threshold.</p><h2>Classifying the full corpus with a smaller LLM</h2><p>By the time the workflow reaches this step, the open-ended part of the job is over. From here, the small model takes over and classifies each report against the approved schema.</p>- name: fetch_corpus
  type: elasticsearch.request
  with:
    method: POST
    path: "/incident_reports/_search"
    body:
      size: 100
      _source: ["acn", "narrative"]
      query:
        match_all: {}

- name: classify_all
  type: foreach
  foreach: "${{ steps.fetch_corpus.output.hits.hits }}"
  iteration-on-failure:
    retry:
      max-attempts: 5
      delay: "3s"
    fallback:
      - name: notify_failure
        type: slack_api.postMessage
        connector-id: "team-alerts"
        with:
          channelNames:
            - "#pipeline-alerts"
          text: "Classification failed for ACN ${{ foreach.item._source.acn }} after all retries."
    continue: true
  steps:
    - name: classify
      type: ai.agent
      inference-id: "mistral-small-extractor"
      timeout: "120s"
      with:
        message: |
          You will classify the following report against a fixed schema.
          For each field in the schema, assign exactly one of its value
          options, or null if none of the values clearly applies. Include
          the short quote that supports the assignment and a confidence
          score between 0 and 1. Set review_required to true if any field
          returned null or any confidence is below 0.5.

          Schema:
          ${{ steps.human_gate.output.approved_fields | json }}

          Report:
          ${{ foreach.item._source.narrative }}
        schema:
          type: object
          properties:
            field_values:
              type: object
              additionalProperties: true
            review_required: { type: boolean }
    - name: write_extraction
      type: elasticsearch.index
      with:
        index: extractions
        document:
          acn: "${{ foreach.item._source.acn }}"
          field_values: "${{ steps.classify.output.structured_output.field_values }}"
          review_required: "${{ steps.classify.output.structured_output.review_required }}"<p><em>Note: The classification step uses </em><em><code>ai.agent</code></em><em> instead of </em><em><code>ai.prompt</code></em><em> because </em><a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/ai-steps#step-types"><em><code>ai.agent</code></em><em> accepts an </em><em><code>inference-id</code></em></a><em>, which lets it call the Elasticsearch </em><em><code>_inference</code></em><em> endpoint directly, while </em><em><code>ai.prompt</code></em><em> only accepts a </em><em><code>connector-id</code></em><em>.</em></p><p>The <code>fetch_corpus</code> step is the third <code>elasticsearch.request</code> in the workflow, so it’s worth saying why we read from the index again. The first two (<code>by_phase</code> and <code>by_anomaly</code>) only pulled a small stratified sample for the planner to reason over, not the data to label. Now that the schema is approved, <code>fetch_corpus</code> pulls the documents we actually want to classify. We cap it at 100 with <code>match_all</code> to keep the demo fast; this is where you would page through the full corpus.</p><p>For every field, it returns a value (or null), a confidence, and a short quote. Setting <code>additionalProperties: true</code> in the JSON schema lets the step return one entry per field without the workflow having to know the field names ahead of time. A stored extraction looks like this:</p>{
  "acn": "2238341",
  "field_values": {
    "attribution_style": {
      "value": "self_critical",
      "confidence": 0.82,
      "quote": "I should have caught the altitude bust earlier"
    },
    "procedure_orientation": {
      "value": "procedure_first",
      "confidence": 0.44,
      "quote": "we ran the QRH before doing anything else"
    }
  },
  "review_required": true
}<p>Here, <code>review_required</code> is <code>true</code> because <code>procedure_orientation</code> came back at <code>0.44</code> confidence, below our <code>0.5</code> threshold, which is the signal a confidence-based quality gate would act on.</p><p>The <code>fetch_corpus</code> step pulls the documents to classify. The <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/foreach"><code>foreach</code></a>step iterates over them sequentially, and <code>iteration-on-failure</code> handles the errors: <code>retry</code> covers transient API errors from the inference endpoint, and, if all attempts fail, the <code>fallback</code> step posts to <a href="https://www.elastic.co/docs/reference/kibana/connectors-kibana/slack-action-type#slack-workflow-examples">Slack</a> so the failure doesn’t pass silently. (An <a href="https://www.elastic.co/docs/reference/kibana/connectors-kibana/email-action-type">email</a> connector works the same way.) <code>continue: true</code> then lets the loop move on to the next document instead of failing the whole run. </p><p><em>For production-scale corpora, consider using </em><a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/composition#workflow-executeasync"><em><code>executeAsync</code></em></a><em>, which is the fan-out version of execute.</em></p><h2>Writing schemas and extractions back to Elasticsearch</h2><p>The workflow produces two things: the approved schema and the per-document field values. The <code>store_schema</code> step runs right after the human gate, before the classification step fans out:</p>- name: store_schema
  type: elasticsearch.index
  with:
    index: schemas
    document:
      question: "${{ inputs.goal }}"
      approved_fields: "${{ steps.human_gate.output.approved_fields }}"
      reviewer_notes: "${{ steps.human_gate.output.notes }}"<p>Each extraction is written inside the <code>foreach</code> loop, so results are persisted as they’re produced rather than batched at the end.</p><p>The <code>schemas</code> index holds one document per discovery run (question, approved fields, reviewer notes). The <code>extractions</code> index holds one document per report per schema version. </p><h2>What this two-tier LLM orchestration pattern gives you</h2><p>We built one Elastic workflow that pulls a stratified sample from an incident report index, sends it with a question to a large reasoning model to generate a classification schema based on the data and a user-defined angle, pauses for human approval, and then iterates over the full corpus with a small Mistral model that assigns one value per field. </p><p>The approved schema and per-document field values are written back to Elasticsearch as structured data.</p><p>The point of the exercise is that two different shapes of work, schema discovery, and schema application can use two different model tiers and that a workflow lets you write the routing decision down.</p><h2>Next steps for your own LLM pipeline</h2><ul><li><p>Try it on a corpus of your own. The pattern doesn’t care whether the input is incident reports, customer feedback, weekly status updates, or property listings.</p></li><li><p>Promote the <code>foreach</code> step to <code>workflow.executeAsync</code> once you’re comfortable for parallel fan-out at scale.</p></li><li><p>Schedule the rediscovery workflow on a cron trigger so you can discover different schema variations based on the data that comes in.</p></li><li><p>Read the <a href="https://www.elastic.co/docs/explore-analyze/workflows">Elastic Workflows documentation</a> for the full step catalog.</p></li></ul><h3><strong>Related reading</strong></h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/build-ai-agents-elastic-inference-service">Build AI agents with Elastic Inference Service</a> (EIS) covers the broader multi-model wiring pattern via EIS, complementary to the Workflows-orchestrated split shown here.</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/ai-agentic-workflows-elastic-ai-agent-builder">How to build AI agentic workflows with Elasticsearch</a> is a higher-level survey of how Agent Builder and Workflows fit together.</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/langextract-elasticsearch-tutorial-usage-example">LangExtract and Elasticsearch tutorial</a> explores a different extraction pattern using a hand-authored schema; useful for contrast with the discover-then-apply approach above.</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/llm-orchestration-elastic-workflows</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/llm-orchestration-elastic-workflows</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Integrations]]></category>
    <category><![CDATA[AI Tools ]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdf42ed268fb6f953/6a87ef5c386ac3fab0adf4e2/image3.png" length="0" type="image/png"/>
    <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Ask the source: Scaling code search to a billion lines with Elasticsearch and Elastic Agent Builder]]></title>
    <description><![CDATA[Sourcerer matches Claude Code and Codex on code retrieval quality and searches up to thousands of times faster than grep. Every answer links back to the exact files and lines across repos and versions.]]></description>
    <content:encoded><![CDATA[<p>Code is the source of truth for its own behavior; it’s always authoritative and never outdated. Definitive answers live at a specific commit in a particular repository, but enterprise deployments depend on many versioned projects working together. Understanding how it all works is a major code search effort.</p><p><em>Does App A v1.2.3 support Feature X? Is it compatible with App B v9.8.7 when running on Kubernetes? Will I need more JVM heap space?</em></p><p>These are the kinds of questions that our field teams handle constantly. Answering them is harder than it looks. Documentation offers context, but it's an abstraction that can't anticipate every possible question. When we hit one it doesn't cover, our options are to interrupt an engineer who should be developing code or to hunt through that code ourselves. Often we don't have the time or expertise to navigate that much of it.</p><p>Coding agents do this well for a single repository on your laptop. Wouldn't it be great if we could scale that to our entire code estate? As a field engineer, I wanted that capability to serve my customers: agentic code intelligence across every project, dependency, platform, and version that we support. And I wanted it grounded in linked citations and always available to everyone as a service.</p><p>So I built it with <a href="https://www.elastic.co/elasticsearch">Elasticsearch</a> and <a href="https://www.elastic.co/elasticsearch/agent-builder">Elastic Agent Builder</a> and packaged it into a command line interface (CLI). I released it under an Apache 2.0 license and called it <a href="https://github.com/elastic/sourcerer">Sourcerer</a>. This blog post reports multiple performance benchmarks of Sourcerer as a code research agent and walks through the design and rationale of its implementation.</p><h2>Sourcerer</h2><p><a href="https://github.com/elastic/sourcerer">Sourcerer</a> explores code like a frontier coding agent, searching across many versioned repositories as fast as it would in a single repository, and it generates answers with linked citations that establish trust.</p><p>Sourcerer consists of:</p><ol><li><p>A set of configuration files for tools, skills, and agents in Agent Builder.</p></li><li><p>A set of index templates to store and search code from Git commit snapshots.</p></li><li><p>A CLI to install those assets and index and prune commit snapshots from remote Git repositories.</p></li></ol><p>At Elastic, we're using Sourcerer to support our customers with verifiable information about our software directly from the source. Our internal deployment has indexed over a billion lines of code from our own public and private repositories. It also includes our core dependencies, such as Apache Lucene and OpenJDK, along with our common integrations, like Kubernetes and OpenTelemetry. Our solution architects, customer architects, consulting architects, and support engineers no longer have to hunt for answers in documentation or reach out to our engineers who should be building software rather than supporting it.</p><h2>Code search benchmarks</h2><h3>Agentic code retrieval</h3><p><a href="https://arxiv.org/abs/2606.07297">SWE-Explore</a> is a new benchmark, published on June 5, 2026, by Zhang et al., that evaluates "how well coding agents explore, localize, and rank repository context." It appears to be the only benchmark that specifically tests agentic code retrieval quality. I ran the benchmark with Sourcerer to see how it performs and compares to the other coding agents from the original paper, and again with Claude Code to measure and compare its token usage and task durations with Sourcerer's.</p><h4>Retrieval scores</h4><p>Sourcerer performs as well as frontier coding agents on relevance metrics for code retrieval (see Figure 1). The composite retrieval score is the arithmetic mean of all retrieval metrics weighed by their Pearson correlations (<em>r</em>) as reported in the paper; I did this to rank the agents by a measurement of overall retrieval quality. Sourcerer trailed Claude Code by 0.002 and surpassed Codex by 0.022 on a 0.0–1.0 scale, which should be interpreted as a statistical tie, given that the results vary slightly on each run due to the indeterminism of large language models (LLMs).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt12808f4ed4f16659/6a7deb00cbb9ed3b2b53f944/image2.png" alt="" /><p></p><p>Generally, all the coding agents, including Sourcerer, performed well on precision metrics and suboptimally on recall metrics, although recall metrics had lower Pearson correlations and thus less importance. Table 1 shows Sourcerer's retrieval scores alongside the scores of the other agents tested in the original paper (page 8, table 6). "SignalReg" is the inverse of what the authors called "NoiseReg" (that is, 1 – NoiseReg); I did this to keep that metric consistent with the other metrics whose ranges imply that higher is better.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd9283fe7e6209465/6a7deb0ec2602368e0b38642/image4.png" alt="" /><h4>Token usage and task duration</h4><p>Zhang et al. didn’t publish metrics for token usage or task durations, so I ran the benchmark again to capture those metrics for Claude Code. Given the high token cost of running the benchmark with any agent that uses an LLM, I opted to test only one coding agent, and Claude Code was the one that I expected most people would find useful as a comparison.</p><p>Compared to Claude Code, Sourcerer used ~13.9% more tokens to complete all 848 benchmark tasks. Sourcerer used 156,379,977 input tokens and 1,612,175 output tokens, while Claude Code used 137,629,435 input tokens and 1,137,176 output tokens. Sourcerer took ~8.9% longer to complete all 848 benchmark tasks. Sourcerer took 40,717 seconds, and Claude Code took 37,460 seconds.</p><p>I view these results on token usage and task duration as an acceptable modest tax in exchange for efficiently searching across multiple repositories and versions. That said, there’s room to explore optimizations to Sourcerer's tools, skills, and system prompt, the harness of Agent Builder, or the search engine of Elasticsearch and Lucene.</p><h4>Single-repo vs. multi-repo scope</h4><p>Critically, the SWE-Explore benchmarks only measure the retrieval scores, token usage, and task duration of agents searching within the boundaries of a single commit snapshot of a repository for any given task. This is the typical search space of a development coding agent. Sourcerer's intended scope is much broader, covering many commit snapshots of many repositories. The benchmark on "search speed and scalability," covered next in this report, shows Sourcerer's unique advantage when searching across many repositories.</p><h4>Retrieval benchmark methodology</h4><p><a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-a.-swe-explore-benchmark-configuration">Appendix A</a> explains the configuration of these benchmarks in detail.</p><p>Sourcerer searched all indexed representations of the <a href="https://huggingface.co/datasets/SWE-Explore-Bench/SWE-Explore-Bench">SWE-Explore-Bench dataset</a> (see <a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-a.-swe-explore-benchmark-configuration">Appendix A</a>). Both benchmark runs used the same GPT-5.4 model that was used in the original paper. Sourcerer communicated with GPT-5.4 through <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service (EIS)</a>, while Claude Code communicated with GPT-5.4 through a shim proxy to be compatible with the OpenAI API.</p><p>I made a best effort to ensure that Sourcerer's benchmark task prompt was like-for-like with Claude Code's (see <a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-a.-swe-explore-benchmark-configuration">Appendix A</a>). Both agents received identical instructions for their roles and tasks, along with output formats, and differed only in their brief harness-specific instructions. I instructed Sourcerer not to use its repo discovery skill and instead gave it explicit repo filtering instructions, ensuring that it was on a level playing field with Claude Code, which already receives the resolved directories. An alternative could have been to leave Sourcerer's repo discovery skill active while instructing Claude Code to find the repository in a filesystem that has all the repositories for the benchmark. I left Sourcerer's system prompts and skills, in addition to its tools, unmodified from their defaults, given that we're comparing the two harnesses overall, and much of which in Claude Code is closed source and not visible or controllable anyway.</p><h3>Code search speed and scalability</h3><p>Coding agents tend to use external tools to match substrings or regular expressions as a first line of retrieval. LLMs are trained to use shell commands, like <code>ls</code> and <code>grep</code>, when exploring code on a filesystem. Claude Code's own built-in <a href="https://code.claude.com/docs/en/tools-reference#grep-tool-behavior">Grep</a> tool invokes <a href="https://github.com/BurntSushi/ripgrep"><code>ripgrep</code></a>. This is the behavior I wanted to reproduce in Elasticsearch.</p><p>Elasticsearch has a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field type that can scale regular expression matching to billions of documents. Sourcerer mimics the inputs and outputs of <code>grep</code> in Elasticsearch using <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.grep.yml"><code>sourcerer.code.grep</code></a>, an Elasticsearch Query Language (ES|QL) tool that performs an<a href="https://www.elastic.co/docs/reference/query-languages/sql/sql-like-rlike-operators"><code>RLIKE</code></a> query on a<a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field of an index where each document has the contents of a single line of code. Sourcerer also provides<a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.search.yml"><code>sourcerer.code.search</code></a>, which performs a BM25-ranked <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/search-functions/match"><code>MATCH</code></a> query against an analyzed text field, for relevance-ranked discovery rather than exact substring retrieval.</p><p>I benchmarked the speed of both approaches against the speed of <code>ripgrep</code> and <code>grep</code> on a filesystem using two different corpus sizes and pattern rarities. Corpus sizes included one commit (7,215,509 lines of code) and 52 commits (200,325,684 lines of code) from the<a href="https://github.com/elastic/elasticsearch"> elastic/elasticsearch</a> repository. The single commit covers the release tag for v9.4.3. The 52 commits cover the latest patch release tag for every major and minor release from v6.0.1 to v9.4.3. Sourcerer searched the corpus as indexed in Elasticsearch, while <code>ripgrep</code>  and <code>grep</code> searched the corpus as stored on a filesystem, reflecting their respective use cases. The regular expression patterns included one that appears rarely among the commits (DiskBBQ) and one that appears commonly among the commits (XContentType).</p><h4>sourcerer.code.grep</h4><p>The first search I benchmarked was a rare pattern for DiskBBQ that appears only in some commits:</p><p><code>.*[dD][iI][sS][kK][-_]?[bB][bB][qQ].*</code></p><p>Search latency (in seconds) spanning a single commit (605 matches found from 7,215,509 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>0.069s</p><p>0.124s</p><p>0.167s</p><p>0.020s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>0.027s</p><p>0.029s</p><p>0.046s</p><p>0.005s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>0.788s</p><p>0.800s</p><p>0.809s</p><p>0.005s</p><p>~6.5x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>0.081s</p><p>0.088s</p><p>0.109s</p><p>0.009s</p><p>~3.0x slower</p><p><code>grep</code></p><p>Cold</p><p>3.565s</p><p>3.580s</p><p>3.821s</p><p>0.055s</p><p>~28.9x slower</p><p><code>grep</code></p><p>Warm</p><p>0.825s</p><p>0.827s</p><p>0.833s</p><p>0.002s</p><p>~28.5x slower</p><p>Search latency (in seconds) spanning 52 commits (1,041 matches found from 200,325,684 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>0.159s</p><p>0.164s</p><p>0.270s</p><p>0.028s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>0.027s</p><p>0.031s</p><p>0.053s</p><p>0.006s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>22.356s</p><p>22.364s</p><p>22.459s</p><p>0.026s</p><p>~136.4x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>16.017s</p><p>16.123s</p><p>16.297s</p><p>0.058s</p><p>~520.1x slower</p><p><code>grep</code></p><p>Cold</p><p>101.962s</p><p>102.386s</p><p>104.281s</p><p>0.752s</p><p>~624.3x slower</p><p><code>grep</code></p><p>Warm</p><p>73.082s</p><p>73.507s</p><p>74.915s</p><p>0.536s</p><p>~2,371.2x slower</p><p>Table 2. p0/p50/p100/stdev retrieval speeds of <code>sourcerer.code.grep</code>, <code>ripgrep</code> , and <code>grep</code>, under cold and warm caches, at two corpus scopes (20 runs per method per cache state; three warmup runs discarded before each warm-cache measurement). All percentiles computed via linear interpolation. Ratios are computed against <code>sourcerer.code.grep</code>'s p50 at the matching cache state. See the “Methodology” section for cache definitions and query/command syntax.</p><p>The <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field type indexes trigrams of each value and uses them as a filter to narrow the candidate set for a regular expression before verifying full matches. For a selective pattern like this one (605 matches out of 7.2 million lines, 1041 out of 200 million) this approaches sublinear time complexity relative to corpus size. <code>ripgrep</code>  and <code>grep</code> both perform scans with linear time complexity, with <code>ripgrep</code>  using multithreading and single instruction, multiple data–accelerated (SIMD-accelerated) literal prefiltering to speed up searches, but neither has a mechanism to skip the vast majority of a corpus the way that an indexed trigram search can.</p><p>This shows up starkly in how each approach scales. Going from the single-commit corpus to the all-commits corpus is a 27.8x increase in line count. <code>sourcerer.code.grep</code> warm-cache time barely moves from 0.029 seconds to 0.031 seconds. ripgrep's warm-cache time goes from 0.089 seconds to 16.1 seconds, a 183x change; and grep's goes from 0.83 seconds to 1.2 minutes, an 89x change. Both filesystem tools scale worse than linearly with corpus size on this hardware, while the indexed approach is nearly flat.</p><p>The second search I benchmarked was a common pattern for XContentType that appears in all commits:</p><p><code>.*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*</code></p><p>Search latency (in seconds) spanning a single commit (7,999 matches found from 7,215,509 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>0.158s</p><p>0.172s</p><p>0.224s</p><p>0.015s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>0.055s</p><p>0.057s</p><p>0.140s</p><p>0.023s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>0.794s</p><p>0.804s</p><p>0.824s</p><p>0.007s</p><p>~4.7x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>0.093s</p><p>0.095s</p><p>0.096s</p><p>0.001s</p><p>~1.7x slower</p><p><code>grep</code></p><p>Cold</p><p>3.385s</p><p>3.399s</p><p>3.421s</p><p>0.010s</p><p>~19.8x slower</p><p><code>grep</code></p><p>Warm</p><p>0.681s</p><p>0.683s</p><p>0.686s</p><p>0.001s</p><p>~12.1x slower</p><p>Search latency (in seconds) spanning 52 commits (290,662 matches found from 200,325,684 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>1.587s</p><p>1.644s</p><p>1.756s</p><p>0.047s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>1.480s</p><p>1.541s</p><p>1.627s</p><p>0.040s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>22.426s</p><p>22.440s</p><p>22.548s</p><p>0.028s</p><p>~13.6x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>16.021s</p><p>16.276s</p><p>16.498s</p><p>0.119s</p><p>~10.6x slower</p><p><code>grep</code></p><p>Cold</p><p>97.699s</p><p>97.922s</p><p>99.150s</p><p>0.404s</p><p>~59.6x slower</p><p><code>grep</code></p><p>Warm</p><p>68.883s</p><p>69.184s</p><p>70.374s</p><p>0.360s</p><p>~44.9x slower</p><p>Table 3. p0/p50/p100/stdev retrieval speeds for the pattern <code>.*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*</code> under the same conditions as Table 2.</p><p>The relative search latencies of <code>sourcerer.code.grep</code> compared to <code>grep</code> shows why the pattern's match count matters as much as the corpus size:</p><p><strong>Corpus</strong></p><p><strong>Corpus size</strong></p><p><strong>Cache</strong></p><p><strong>Speed of sourcerer.code.grep with a rare pattern (DiskBBQ)</strong></p><p><strong>Speed of sourcerer.code.grep with a common pattern (XContentType)</strong></p><p>1 commit</p><p>7,215,509 lines</p><p>Cold</p><p>~28.9x faster</p><p>~19.8x faster</p><p>1 commit</p><p>7,215,509 lines</p><p>Warm</p><p>~28.5x faster</p><p>~12.1x faster</p><p>52 commits</p><p>200,325,684 lines</p><p>Cold</p><p>~624.3x faster</p><p>~59.6x faster</p><p>52 commits</p><p>200,325,684 lines</p><p>Warm</p><p>~2,371.2x faster</p><p>~44.9x faster</p><p>Table 4. This table shows how much faster <code>sourcerer.code.grep</code> was compared to <code>grep</code> when searching across two different corpus sizes and two different pattern rarities.</p><p>The pattern with far more matches shows a dramatically smaller Elasticsearch advantage, most strikingly at all-commits scope, where the advantage drops from 2,371x to 45x. The reason is visible in the absolute numbers: <code>sourcerer.code.grep</code>'s warm-cache time at all-commits scope jumps from 0.031 seconds (DiskBBQ) to 1.541 seconds (XContentType), a 49.7x increase for a 279x increase in match count, while <code>grep</code>'s warm-cache time barely changes (73.5 seconds to 69.2 seconds, effectively flat, since it scans the same number of bytes regardless of how many of them match). The <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field's trigram index has sublinear-in-corpus-size behavior that comes specifically from narrowing the candidate set before verification. Once a pattern matches hundreds of thousands of lines, the bottleneck shifts from narrowing candidates to collecting and serializing all of them, a cost that scales with match count rather than corpus size. <code>sourcerer.code.grep</code> still wins by a wide margin even in this less favorable case, but the margin depends heavily on how selective the search is, not just how large the corpus is.</p><h4>sourcerer.code.search</h4><p>Sourcerer's other retrieval tool, <code>sourcerer.code.search</code>, performs a BM25-ranked <code>MATCH</code> query rather than an exact-substring regex match. Its speed on both patterns is included below for reference.</p><p><strong>Pattern</strong></p><p><strong>Corpus scope</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>Matches found</strong></p><p>DiskBBQ</p><p>One commit</p><p>0.013s</p><p>0.017s</p><p>0.018s</p><p>0.002s</p><p>288</p><p>DiskBBQ</p><p>52 commits</p><p>0.014s</p><p>0.020s</p><p>0.046s</p><p>0.007s</p><p>493</p><p>XContentType</p><p>One commit</p><p>0.062s</p><p>0.063s</p><p>0.154s</p><p>0.020s</p><p>7,177</p><p>XContentType</p><p>52 commits</p><p>2.203s</p><p>2.359s</p><p>2.611s</p><p>0.125s</p><p>249,737</p><p>Table 5. p0/p50/p100/stdev retrieval speeds of <code>sourcerer.code.search</code>, warm cache only, for both patterns at both corpus scopes (20 runs per row; cold-cache figures omitted; see Methodology). Match counts are <code>sourcerer.code.search</code>'s own, not the regex-based methods' BM25 matches on tokens rather than substrings, so these aren’t directly comparable to Tables 2–4.</p><h4>What drives the speed advantage</h4><p><code>sourcerer.code.grep</code> outperformed <code>ripgrep</code> and <code>grep</code> at every corpus scale and pattern rarity tested, along with every cache state tested. But it wasn’t by a fixed margin. The advantage ranged from ~3–30x on a common pattern to over 2,300x on a rare pattern, because indexed and brute-force search respond to different things. The trigram narrowing of the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field does less work as a pattern gets more selective, while <code>grep</code> and <code>ripgrep</code> do the same amount of work regardless of how much of the corpus happens to match. <code>sourcerer.code.search</code> is a third option for the cases where the exact string isn't known at all.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcc34bd5e704c0112/6a7deb29b8c2e6ffc9be51a1/image1.png" alt="" /><p></p><p>Figure 2: Sourcerer's <code>sourcerer.code.grep</code> tool outperformed <code>ripgrep</code> and <code>grep</code> in every combination of corpus size and pattern rarity benchmarked, sometimes by multiple orders of magnitude. Its outperformance was strongest when searching rare patterns in large corpora and weakest when searching common patterns in small corpora.</p><p>These search speed benchmarks show what <em>scalable</em> means in practice for an enterprise code search agent. It's being able to search the histories of any number of repositories at interactive speeds, just like how a developer coding agent searches the working state of a single repository on a filesystem.</p><h4>Speed benchmark methodology</h4><p><a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-b.-code-search-speed-and-scalability-benchmark-configuration">Appendix B</a> explains the configuration of this benchmark. Note that while the Elasticsearch deployment had two data nodes each with the same specs as the virtual machine used for <code>ripgrep</code> and <code>grep</code>, the benchmark consisted of one index with one primary shard and one replica shard. Each search ran on a single shard, which means that <code>sourcerer.code.grep</code> had the same amount of vCPUs and memory available per search as <code>ripgrep</code> and <code>grep</code>, despite having twice as much capacity across the overall deployment. Having two data nodes actually incurred a slight latency <em>penalty</em> compared to just one data node. For brevity, I've omitted results from the benchmark with one data node. We don't recommend single-node deployments in production, so it's worth including the realistic latency overhead that comes with a multi-node deployment in this benchmark.</p><p>I compared all retrieval methods under both cold and warm caches, 20 cold runs and 20 warm runs per method. The definitions of cold and warm caches weren’t like-for-like between the Elasticsearch and filesystem benchmarks. For <code>ripgrep</code> and <code>grep</code>, a <em>cold cache</em> meant performing a full-page cache drop (<code>sync; echo 3 &gt; /proc/sys/vm/drop_caches</code>) immediately before each cold run; and a <em>warm cache</em> meant three discarded warmup executions immediately followed by the 20 measured runs, with no cache drops in between. For ES|QL, a <em>cold cache</em> meant calling <code>POST /_cache/clear</code> before each run. This only clears Elasticsearch's internal caches, not the OS page caches of the data nodes, which can't be cleared by hand on Elastic Cloud Hosted (ECH). A <em>warm cache</em> for ES|QL meant three discarded warmup queries immediately before the 20 measured runs, to help ensure that the caches on both data nodes would be warm. <code>sourcerer.code.search</code>'s cold-cache figures are omitted from Table 5 for the same reason discussed elsewhere in this post: Its cold-cache measurements came back statistically indistinguishable from its own warm-cache measurements, evidence that the OS-level cache-clearing limitation affects it more than it affects <code>sourcerer.code.grep</code>, which showed a consistent, physically sensible cold/warm gap throughout.</p><h3>Code search indexing throughput</h3><p>I didn't conduct a formal benchmark of indexing throughout. I'll share my general observations instead.</p><p>Typically, I see a sustained indexing throughput of 20K–25K lines per second on data nodes that each have ~16GiB RAM and ~8 vCPU on c4a-highcpu instances on Google Cloud Platform (GCP). That includes writing to a primary shard and its replica shard. I've seen throughput as high as ~60K lines per second on <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a> with <a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-tier-autoscaling">Search Power</a> set to "Performant."</p><p>An engineering team at Elastic compared Sourcerer's indexing throughput to semantic code search implementations that used either sparse vector generation (<a href="https://www.elastic.co/docs/explore-analyze/machine-learning/nlp/ml-nlp-elser">Elastic Learned Sparse EncodeR [ELSER])</a> or dense vector embedding generation (<a href="https://jina.ai/models/jina-embeddings-v5-text-small/">Jina</a>). Sourcerer indexed ~25x faster than <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/nlp/ml-nlp-elser">.elser-2-elastic</a> and ~15x faster than <a href="https://jina.ai/models/jina-embeddings-v5-text-small/">.jina-embeddings-v5-text-small</a>, while retrieval quality was similar among all of them. More concretely, what took ~6 hours to index with ELSER took ~14 minutes to index with Sourcerer.</p><h2>Code search solution design and rationale</h2><p>The remainder of this blog post explains the rationale for my design decisions of Sourcerer, giving expert insights for practitioners of Elasticsearch and generative AI (GenAI).</p><h3>Goals</h3><p>Ultimately, we want an agent that answers questions about deployed software and its supporting infrastructure by searching the primary sources of truth (the code itself) and generating verifiable responses that cite those sources so they can be trusted. Inspired by <a href="https://arxiv.org/abs/2605.15184">the success of coding agents with grep</a>, my main functional goal for Sourcerer was to reproduce the search behavior of a coding agent and generate responses with citations, all using Agent Builder. My nonfunctional goals were to keep it fast and scalable, as well as accurate, when searching across many versioned repositories, while maintaining acceptable costs and ease of use. Of these goals, reproducing the search behaviors of coding agents would be the most consequential, as it would dictate the access pattern, index design, and query design, plus their effects on nonfunctional goals.</p><h3>Access pattern</h3><p>From a human perspective, the intended access pattern is simple: We expect to ask natural language questions about software and receive plain language answers grounded in the source of truth. From the perspective of the agent handling those questions, the intended access pattern is to reproduce the search behavior of coding agents to find what it needs. The LLMs used by coding agents are heavily trained to explore code with shell commands, like<code>ls</code>or <code>find</code>, <code>grep</code> or <code>ripgrep</code>, <code>cat</code>, <code>head</code>, <code>tail</code>, and so on. <a href="https://code.claude.com/docs/en/tools-reference">Claude Code's built-in tools</a>, such as <a href="https://code.claude.com/docs/en/tools-reference#glob-tool-behavior">Glob</a> and <a href="https://code.claude.com/docs/en/tools-reference#grep-tool-behavior">Grep</a>, provide similar functions.</p><p>I chose to go with the grain of how models are trained. So my intended access pattern for Sourcerer was to reproduce the names, inputs, and outputs of shell commands as <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools/esql-tools">ES|QL tools in Agent Builder</a>. That way, the Agent Builder harness would allow an LLM to use its trained intuition to achieve similar results as a frontier harness, like Claude Code or Codex, without having to fill the LLM's limited context window with instructions for using a different search interface. This was the main context engineering problem to solve with Sourcerer. Solving it would enable faster searches that scale across many repositories at once, allowing an agent to answer questions about deployments in which many different versioned software projects work together.</p><h3>Index design</h3><p>I decided on three index templates to fulfill this access pattern:</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/index_templates/sourcerer-v2-refs.json"><code>sourcerer-refs</code></a>: Each document indexes the high-level metadata for a single Git reference or <em>ref</em> identified by its unique commit hash, which can have a tag name or branch name associated with it. A ref represents an entire snapshot of a repository at a point in time. This is a small index. The agent mainly uses this to discover the repositories and snapshots that are available to search.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/index_templates/sourcerer-v2-files.json"><code>sourcerer-files</code></a>: Each document indexes the metadata for a single file of a given ref. This is a larger index. The agent mainly uses this to navigate files and directories using <code>ls</code> semantics.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/index_templates/sourcerer-v2-lines.json"><code>sourcerer-lines</code></a>: Each document indexes the contents of a single, numbered line of code for a given file of a given ref. Yes, every line of code becomes a document. This is the largest index (but perhaps not as large as you might expect). The agent mainly uses this to search and view code using <code>grep</code> and <code>cat</code> semantics.</p></li></ul><p>The design is almost entirely denormalized. Each index has the same namespacing fields for fast, joinless filtering. Each indexed ref stores all files and lines from its commit snapshot, rather than storing diffs and reconstructing them at search time or storing unique files and lines with a mutable array of ref names associated with each. These choices trade duplicative storage (the cheapest compute resource) for faster searches and less segment merging pressure.</p><h4>Namespacing</h4><p>All three indices use four fields to namespace the ref, file, or line of code:</p><ul><li><p><code>git.host</code>: A Git hosting provider (for example, github, gitlab).</p></li><li><p><code>git.org</code>: An account name (for example, elastic).</p></li><li><p><code>git.repo</code>: A Git repository (for example, elasticsearch, kibana).</p></li><li><p><code>git.commit</code>: A commit hash, stored as the full 40-character SHA-1 digest for integrity.</p></li></ul><p>The document <code>_id</code> hashes for files and lines are also namespaced by <code>{git.host}</code>, <code>{git.org}</code>, <code>{git.repo}</code>, and <code>{git.commit}</code> to allow for idempotent indexing. That means you can safely rerun an indexing job without duplicating any documents.</p><p>Likewise, the index names are namespaced with the same semantics, using tildes (<code>~</code>) as a reliable separator since it's a disallowed character in Git repository names and organization names:</p><ul><li><p><code>sourcerer-v*-files~{git.host}~{git.org}~{git.repo}</code></p></li><li><p><code>sourcerer-v*-lines~{git.host}~{git.org}~{git.repo}</code></p></li></ul><p>This namespace convention has many benefits:</p><ul><li><p>Agents can quickly narrow the search space for refs and files, along with lines of code, by these common scoping fields, keeping searches fast and focused.</p></li><li><p>The semantics reflect common permission boundaries. You can reproduce the access policies of your Git hosting provider by implementing your choice of index-level security and/or document-level security based on host or organization or based on repository.</p></li><li><p>You can instantly delete indices for a whole repository or organization, or for a host, without an expensive <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-delete-by-query"><code>_delete_by_query</code></a>.</p></li><li><p>The index names are future proofed for different levels of granularity. Sourcerer might eventually allow indexing code by <code>{git.host}</code>, <code>{git.host}~{git.org}</code>, or <code>{git.host}~{git.org}~{git.repo}~{git.commit}</code> for selective shard sizing optimizations. The query syntax would be unaffected because they target index aliases (<code>sourcerer-files</code> and <code>sourcerer-lines</code>), not individual indices.</p></li></ul><h4>Settings</h4><p>Three index settings help to optimize storage costs and search speed:</p><ul><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/sorting">Index sorting</a> gives faster searches and better compression at the cost of reduced indexing throughput. Each index sorts documents on disk by <code>git.host</code>, <code>git.org</code>, <code>git.repo</code>, <code>git.commit</code>. File and line documents are further sorted by <code>file.path</code>, and line documents are further sorted by <code>line.number</code>.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-source-field#synthetic-source">Synthetic <code>_source</code></a> discards <code>_source</code> and instead reconstructs it as needed when reindexing. None of the queries access <code>_source</code>, which makes it dead weight. Enabling this setting reclaims ~50% storage space in the files and lines indices. While it requires an Enterprise license, enabling it on a non-licensed deployment won’t prevent the index from being created; instead the setting will be ignored.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/index-modules#index-codec"><code>best_compression</code></a> is a fallback for Elastic deployments that lack an Enterprise license to use synthetic <code>_source</code>. It provides decent compression for <code>_source</code> (~11% storage savings by my observations) in exchange for a modest tax on indexing throughput (~15% slower), while search speeds are essentially unaffected because the queries don't fetch <code>_source</code>.</p></li></ul><p>I use <a href="https://www.elastic.co/docs/manage-data/data-store/aliases">index aliases</a> to support zero-downtime upgrades when reindexing to a new schema.</p><h4>Mappings</h4><p>The indices mainly use <code>keyword</code> fields. They facilitate efficient filtering with basic wildcard support and aggregations, along with optimal storage usage and indexing throughput.</p><p>The <code>line.content</code> field is indexed both as a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field and as a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/text"><code>text</code></a> field with <a href="https://www.elastic.co/docs/manage-data/data-store/text-analysis">tokenization</a> and <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/similarity">similarity</a> settings tuned for code search. This gives agents the option to search code using familiar and effective <code>grep</code>-like regular expressions on the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field or using the inverted index of the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/text"><code>text</code></a> field to return only the highest ranking matched lines to reduce token usage and increase search speed. Both options are remarkably fast and scalable, taking milliseconds to finish in most cases.</p><h4>Shards</h4><p>While shards are becoming less relevant with the rise of stateless platforms like <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a>, I want the solution to accommodate all deployment modes of Elasticsearch. So I've given attention to the effect of the index design on the number and size of primary shards. Based on my observations, I expect that most users won’t need to give much attention to shards.</p><p>By default, Elasticsearch enforces a <a href="https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/size-shards#shard-count-per-node-recommendation">soft limit of 1,000 shards per data node</a> (including replicas). That means this solution will hit a soft limit of just under 250 repositories indexed per data node, because each repository is written to a files index and a lines index, each with one primary shard and one replica. Additionally, there’s a conventional best practice of limiting shard sizes to ~50GB, which affects how many refs you can index per repository. Both of these limits can be pushed a bit. But they reveal that this solution design really is optimized for its intended use case of searching the commit snapshots of supported, deployed software. You wouldn't use Sourcerer to index every repository on the Internet, and you shouldn't use it to index every ephemeral development branch. Plus, you should decide how many refs are worth retaining for each repository.</p><p>Repository-level granularity of indices appears to strike the right balance of shard counts and shard sizes. For reference, Kibana is one of the largest repositories on GitHub (<a href="https://stacey-gammon.github.io/repo-stats/">source</a>). I observed its shard size to be a manageable 60GB–75GB when retaining only the latest patch release for every major and minor version release from v6.0.0 to v9.5.0. That's great coverage for the Elastic deployments we see in the wild. If that's one of the largest repositories out there, you can expect just about any other repository to fit in a single shard as long as you have a reasonable retention policy, which I discuss in the next section (“Pruning”).</p><h4>Pruning</h4><p>By default, Sourcerer retains everything you index. Pruning lets you delete old refs to prevent unbounded growth. You can define ref retention policies based on the age of refs and the number of refs indexed in the repo. You can also define these policies based on the number of semantic versions indexed in the repo at any level of granularity for major, minor, patch, build, and prerelease versions. Some common configurations are to retain only the latest commit of the default branch or the most recent patch release tag for every major and minor version release tag.</p><h3>Elastic Agent Builder tools</h3><p>With the index design in place, we can review the tools that query those indices.</p><p><a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools/esql-tools">ES|QL tools in Agent Builder</a> are parameterized ES|QL queries with descriptions to guide the agent's use of them. Sourcerer has tools for several purposes: repo discovery, file discovery, code search, and code display. These tools reproduce the names, inputs, and outputs of shell commands that coding agents prefer to use when exploring code, making them intuitive enough for the LLM to use with minimal instructions passed into its context window.</p><h4>Repo discovery</h4><p>These are typically the first tools that the agent calls. Unlike most coding agents, which search within a single repository on a filesystem, Sourcerer is aware that its search space likely has multiple repositories and versions, and so its first step is to decide which repos and refs to scope its searches to.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.repos.list.yml"><code>sourcerer.repos.list</code></a>: Lists the repos that are available to search.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.repos.search.yml"><code>sourcerer.repos.search</code></a>: Lists the repos whose file contents best match a given query.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.refs.list.yml"><code>sourcerer.refs.list</code></a>: Lists the repos and refs that are available to search.</p></li></ul><h4>File discovery</h4><p>All file discovery tools support glob matching (<code>*</code> and <code>**</code>) on file paths.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.ls.yml"><code>sourcerer.files.ls</code></a>: Lists files and directories that match a given pattern.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.tree.yml"><code>sourcerer.files.tree</code></a>: Lists files and directories that match a given pattern in a tree-like format.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.wc.yml"><code>sourcerer.files.wc</code></a>: Counts lines, words, characters, bytes, and longest lines for each matching file.</p></li></ul><h4>Code search</h4><p>All file code search tools support glob matching (<code>*</code> and <code>**</code>) on file paths.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.grep.yml"><code>sourcerer.code.grep</code></a>: Searches lines of code using <a href="https://www.elastic.co/docs/reference/query-languages/sql/sql-like-rlike-operators"><code>RLIKE</code></a> on a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field for rapid execution of regular expressions.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.search.yml"><code>sourcerer.code.search</code></a>: Searches lines of code using <a href="https://www.elastic.co/docs/reference/query-languages/sql/sql-functions-search#sql-functions-search-match"><code>MATCH</code></a> on a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/text"><code>text</code></a> field that has been tuned for code search.</p></li></ul><h4>Code retrieval</h4><p>All file code retrieval tools concatenate the desired lines of any matching file and return them as a single, contiguous block of code in <code>grep -n</code> format, which is a format preferred by coding agents, including Claude Code's built-in <a href="https://code.claude.com/docs/en/tools-reference#grep-tool-behavior">Grep</a> tool. This lets the agent see a faithful representation of file contents with line-level attribution for precise citations, without requiring the agent to reconstruct the contents or infer line numbers through reasoning.</p><p>To illustrate, here's how the <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.head.yml"><code>sourcerer.files.head</code></a> tool formats the first five lines of <a href="https://raw.githubusercontent.com/elastic/kibana/refs/tags/v9.5.0/README.md">Kibana's <code>README.md</code></a> file, reconstructed from five documents from the lines index:</p>1:# Kibana
2:
3:Kibana is the open source interface to query, analyze, visualize, and manage your data stored in Elasticsearch.
4:
5:- [Getting Started](#getting-started)<p>All file code retrieval tools support glob matching (<code>*</code> and <code>**</code>) on file paths. Agents typically use these tools to display the contents of a single file.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.cat.yml"><code>sourcerer.files.cat</code></a>: Concatenates and displays all lines for each matching file.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.head.yml"><code>sourcerer.files.head</code></a>: Concatenates and displays the first <code>n</code> lines for each matching file.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.tail.yml"><code>sourcerer.files.tail</code></a>: Concatenates and displays the last <code>n</code> lines for each matching file.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.read_lines.yml"><code>sourcerer.files.read_lines</code></a>: Concatenates and displays the range of lines between two given line numbers for each matching file.</p></li></ul><h3>Agent Builder skills</h3><p>With the tools implemented, we can review the skills that guide the agent's proper use of them. Agents typically invoke the follow skills in this order:</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/repo-discovery/SKILL.md"><code>sourcerer-repo-discovery</code></a>: Guides the agent in discovering and selecting the repositories that are available to search for a given prompt. While this is typically the first skill an agent invokes, the agent might return to it when tracing dependencies from other repositories or when answering questions that span multiple repositories or versions.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/ref-resolution/SKILL.md"><code>sourcerer-ref-resolution</code></a>: Guides the agent in resolving the names of tags or branches to their unique, immutable commit hashes. This lets the agent reliably filter its searches to a single commit snapshot.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/code-search/SKILL.md"><code>sourcerer-code-search</code></a>: Guides the agent in exploring code, with basic best practices on when and how to use the available tools for maximum efficiency.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/code-citations/SKILL.md"><code>sourcerer-code-citations</code></a>: Guides the agent in citing files, directories, lines of code, and ranges of lines of code. Sourcerer auto-generates more specific citation skills for each major Git hosting provider, so that its citation links conform to the URL formats of each respective host.</p></li></ul><h3>Agent system prompt</h3><p>The final packaging of the agent comes with a <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_agents/sourcerer.yml">system prompt</a> that succinctly describes the agent's role and its high-level instructions. The configure file that defines the system prompt also defines the tools and skills that are made available to the agent, so that it can only execute what we permit it to execute.</p><h3>Sourcerer CLI</h3><p>The <a href="https://github.com/elastic/sourcerer">Sourcerer CLI</a> assists with setup and indexing, along with pruning to keep operations simple. Configuration is managed through a <a href="https://github.com/elastic/sourcerer/blob/main/specs/sourcerer-yml.md"><code>sourcerer.yml</code></a> configuration file.</p><ul><li><p><code>sourcerer setup</code>: Idempotently loads the index templates and Agent Builder configurations, in addition to Kibana dashboards. This is typically a one-time operation and takes a few seconds.</p></li><li><p><code>sourcerer index</code>: Checks for new refs that match patterns defined in <a href="https://github.com/elastic/sourcerer/blob/main/specs/sourcerer-yml.md"><code>sourcerer.yml</code></a> and then idempotently indexes them, skipping any refs that have already been indexed or that qualify for pruning based on retention policies. It calls <code>git</code> to clone repos and to list remote refs, as well as to  check out refs.</p></li><li><p><code>sourcerer prune</code>: Checks the retention policies for any refs that qualify for pruning and then deletes them from all three indices using <code>_delete_by_query</code>.</p></li></ul><p>You can easily schedule indexing and pruning using external schedulers, like cron. For our internal use at Elastic, we maintain <a href="https://github.com/elastic/sourcerer/blob/main/specs/sourcerer-yml.md"><code>sourcerer.yml</code></a> files in a private Git repository and schedule indexing and pruning with <a href="https://github.com/features/actions">GitHub Actions</a>. I prefer to index code frequently, while pruning outside of normal working hours to prevent agents from suddenly losing context in the middle of a conversation.</p><h3>Feature license summary</h3><p>For transparency, here’s a summary of the license levels for all non–open source software (non-OSS) features referenced in this solution design.</p><p>Subscription features:</p><ul><li><p><a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Agent Builder</a> (optional; you can query the indices from a different harness using the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/">Elasticsearch API</a>).</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-source-field#synthetic-source">Synthetic <code>_source</code></a> (optional).</p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/controlling-access-at-document-field-level#document-level-security">Document-level security</a> (optional).</p></li></ul><p>Free features, proprietary to Elastic (not OSS as defined by the Open Source Initiative [OSI]):</p><ul><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a>.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code> field</a>.</p></li></ul><p>A subscription gives you the magic of "everything just works" with Agent Builder, along with resource optimizations, finer security permissions, and platform support. Without a subscription, you can still index and prune code with the Sourcerer CLI and search the code using the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/">Elasticsearch API</a>.</p><h2>Conclusion</h2><p>Sourcerer demonstrates that Elasticsearch + Agent Builder is an exceptional solution for agentic enterprise code intelligence, as evidenced in many ways:</p><ul><li><p><strong>Accurate:</strong> Sourcerer's agentic code retrieval quality is at parity with frontier coding agents as demonstrated by its performance on the academic benchmark SWE-Explore.</p></li><li><p><strong>Fast:</strong> Sourcerer's query speed ranges from milliseconds in a single repository to multiple seconds across a billion lines of code. Indexing is also much faster than could be achieved with vector embedding generations.</p></li><li><p><strong>Scalable:</strong> Elasticsearch sharding enables horizontal scaling, making it possible to search the current and historical states of an entire enterprise software estate. Alternatively, the stateless architecture of <a href="https://www.elastic.co/cloud/serverless">Serverless</a> naturally scales without having to plan shards.</p></li><li><p><strong>Resilient:</strong> Elasticsearch replication enables high availability to keep the agent operational 24/7. Likewise, the stateless architecture of <a href="https://www.elastic.co/cloud/serverless">Serverless</a> naturally provides high availability.</p></li><li><p><strong>Polyglot:</strong> Sourcerer's indexing and retrieval methods are completely language-agnostic and tolerant of malformed code.</p></li><li><p><strong>Secure:</strong> Elasticsearch <a href="https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/controlling-access-at-document-field-level">document-level security</a> enforces access policies for humans and agents at the organization, repository, and commit levels.</p></li><li><p><strong>Efficient:</strong> Sourcerer demonstrates an efficient use of storage, memory, compute, and token consumption for its intended use case.</p></li><li><p><strong>Manageable:</strong> The Sourcerer CLI, and its use of native <code>git</code> commands and Elastic REST APIs, makes it easy to get started with and operate, as well as schedule.</p></li><li><p><strong>Universal:</strong> Organizations from all industries build software, and Git is the source of truth for ~85% of them (<a href="https://fosspost.org/git-market-share-statistics/">source</a>). Sourcerer has value to all of these organizations. </p></li></ul><p>Sourcerer is currently less than two months old, and the benchmark results suggest that there’s room to improve recall and token efficiency, along with task duration. I'll continue to work on this project in the near future.</p><h2>Try it yourself</h2><p>Sourcerer depends on Elasticsearch and Kibana. You can get started with those in a couple ways:</p><ul><li><p><a href="https://www.elastic.co/cloud/cloud-trial-overview">Elastic Cloud</a> (includes an Enterprise trial).</p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">Local setup with Docker</a> (includes an Enterprise trial).</p></li></ul><p>Then you can get started with <a href="https://github.com/elastic/sourcerer">Sourcerer</a>.</p><p>I built this with <a href="https://www.elastic.co/elasticsearch/agent-builder">Agent Builder</a>. What will you build?</p><h2>References</h2><p>Sen, S. (2026). <em>Is Grep All You Need? How Agent Harnesses Reshape Agentic Search</em>. arXiv.<a href="https://arxiv.org/abs/2605.15184"> https://arxiv.org/abs/2605.15184</a></p><p>Zhang, S. (2026). <em>SWE-Explore: Benchmarking How Coding Agents Explore Repositories</em>. arXiv.<a href="https://arxiv.org/abs/2606.07297"> https://arxiv.org/abs/2606.07297</a></p><h2>Appendices</h2><h3>Appendix A. SWE-Explore benchmark configuration</h3><p>Benchmark environment:</p><ul><li><p>Sourcerer version: v1.0.0 (commit hash: <a href="https://github.com/elastic/sourcerer/tree/26d2e84e3f9e4c1e1598a48532eba9a739465c3e">26d2e84e3f9e4c1e1598a48532eba9a739465c3e</a>)</p></li><li><p>Elastic Cloud Hosted (ECH):</p></li><ul><li><p>Region: GCP - Los Angeles (us-west2)</p></li><li><p>CPU Optimized hardware (c4a-highcpu)</p></li><li><p>3x Elasticsearch data nodes (each with 16GiB RAM, 8 vCPUs)</p></li><li><p>2x Kibana instances (each with 2GiB RAM)</p></li><li><p>Elastic stack version: v9.5.0 (commit hash: <a href="https://github.com/elastic/elasticsearch/tree/dbedef007f580447413782705acb9afec41f945b">dbedef007f580447413782705acb9afec41f945b</a>)</p></li></ul><li><p>Claude Code version: 2.1.202</p></li><li><p>LLM: GPT-5.4</p></li></ul><p>Indices as reported by <code>GET /_cat/indices</code>:</p>index                                          pri rep docs.count docs.deleted store.size
sourcerer-v1-files~ansible~ansible               1   1     234423            0     18.5mb
sourcerer-v1-files~apache~druid                  1   1      46720            0      4.1mb
sourcerer-v1-files~apache~lucene                 1   1      47932            0      4.2mb
sourcerer-v1-files~astral-sh~ruff                1   1      49021            0      9.1mb
sourcerer-v1-files~astropy~astropy               1   1      38994            0      6.1mb
sourcerer-v1-files~axios~axios                   1   1        339            0       89kb
sourcerer-v1-files~babel~babel                   1   1      24798            0      4.9mb
sourcerer-v1-files~briannesbitt~carbon           1   1      14054            0      1.2mb
sourcerer-v1-files~burntsushi~ripgrep            1   1        408            0    232.4kb
sourcerer-v1-files~caddyserver~caddy             1   1       2162            0    533.9kb
sourcerer-v1-files~django~django                 1   1    1329708            0     91.7mb
sourcerer-v1-files~element-hq~element-web        1   1      29426            0      6.1mb
sourcerer-v1-files~facebook~docusaurus           1   1       9288            0        2mb
sourcerer-v1-files~faker-ruby~faker              1   1       1130            0    207.2kb
sourcerer-v1-files~fastlane~fastlane             1   1      13663            0      1.7mb
sourcerer-v1-files~flipt-io~flipt                1   1      10518            0    873.8kb
sourcerer-v1-files~fluent~fluentd                1   1       3340            0    744.4kb
sourcerer-v1-files~fmtlib~fmt                    1   1       1243            0    133.7kb
sourcerer-v1-files~future-architect~vuls         1   1       3952            0      828kb
sourcerer-v1-files~gin-gonic~gin                 1   1        217            0     76.6kb
sourcerer-v1-files~gohugoio~hugo                 1   1      11545            0      1.2mb
sourcerer-v1-files~google~gson                   1   1       1432            0    349.4kb
sourcerer-v1-files~gravitational~teleport        1   1     100317            0     16.3mb
sourcerer-v1-files~hashicorp~terraform           1   1      13511            0      1.5mb
sourcerer-v1-files~immutable-js~immutable-js     1   1        458            0      120kb
sourcerer-v1-files~internetarchive~openlibrary   1   1      35448            0      5.9mb
sourcerer-v1-files~javaparser~javaparser         1   1       5154            0      1.3mb
sourcerer-v1-files~jekyll~jekyll                 1   1        720            0    272.9kb
sourcerer-v1-files~jordansissel~fpm              1   1        167            0     98.3kb
sourcerer-v1-files~jqlang~jq                     1   1       1172            0    291.4kb
sourcerer-v1-files~laravel~framework             1   1      29021            0      4.9mb
sourcerer-v1-files~matplotlib~matplotlib         1   1     133013            0     19.9mb
sourcerer-v1-files~micropython~micropython       1   1      15631            0        3mb
sourcerer-v1-files~mrdoob~three.js               1   1       9858            0      1.1mb
sourcerer-v1-files~mwaskom~seaborn               1   1        633            0      181kb
sourcerer-v1-files~navidrome~navidrome           1   1      10301            0      2.1mb
sourcerer-v1-files~nlohmann~json                 1   1       1090            0    281.9kb
sourcerer-v1-files~nodebb~nodebb                 1   1     104805            0     14.5mb
sourcerer-v1-files~nushell~nushell               1   1       8990            0      1.7mb
sourcerer-v1-files~pallets~flask                 1   1        251            0    120.1kb
sourcerer-v1-files~php-cs-fixer~php-cs-fixer     1   1       8300            0        1mb
sourcerer-v1-files~phpoffice~phpspreadsheet      1   1      17140            0      3.6mb
sourcerer-v1-files~preactjs~preact               1   1       3446            0    772.5kb
sourcerer-v1-files~projectlombok~lombok          1   1      24153            0        4mb
sourcerer-v1-files~prometheus~prometheus         1   1       3416            0    644.9kb
sourcerer-v1-files~protonmail~webclients         1   1      84927            0     14.9mb
sourcerer-v1-files~psf~requests                  1   1        928            0    329.4kb
sourcerer-v1-files~pydata~xarray                 1   1       5114            0    910.6kb
sourcerer-v1-files~pylint-dev~pylint             1   1      24945            0      3.8mb
sourcerer-v1-files~pytest-dev~pytest             1   1       8880            0      1.3mb
sourcerer-v1-files~qutebrowser~qutebrowser       1   1      29453            0      4.6mb
sourcerer-v1-files~reactivex~rxjava              1   1       1959            0    523.7kb
sourcerer-v1-files~redis~redis                   1   1      12646            0        2mb
sourcerer-v1-files~rubocop~rubocop               1   1      15349            0      3.3mb
sourcerer-v1-files~scikit-learn~scikit-learn     1   1      39030            0      6.2mb
sourcerer-v1-files~sharkdp~bat                   1   1       2526            0    789.3kb
sourcerer-v1-files~sphinx-doc~sphinx             1   1      55638            0      8.8mb
sourcerer-v1-files~sympy~sympy                   1   1     112572            0     16.7mb
sourcerer-v1-files~tokio-rs~axum                 1   1       1474            0    365.9kb
sourcerer-v1-files~tokio-rs~tokio                1   1       5175            0        1mb
sourcerer-v1-files~tutao~tutanota                1   1       9267            0      1.6mb
sourcerer-v1-files~uutils~coreutils              1   1       3226            0    806.6kb
sourcerer-v1-files~valkey-io~valkey              1   1       6681            0      1.4mb
sourcerer-v1-files~vuejs~core                    1   1       2785            0      750kb
sourcerer-v1-lines~ansible~ansible               1   1   23067027            0      6.9gb
sourcerer-v1-lines~apache~druid                  1   1   10230691            0      1.4gb
sourcerer-v1-lines~apache~lucene                 1   1    9857695            0      1.5gb
sourcerer-v1-lines~astral-sh~ruff                1   1    5895776            0    797.9mb
sourcerer-v1-lines~astropy~astropy               1   1   16855521            0      5.1gb
sourcerer-v1-lines~axios~axios                   1   1     104915            0     17.3mb
sourcerer-v1-lines~babel~babel                   1   1     661054            0     96.4mb
sourcerer-v1-lines~briannesbitt~carbon           1   1    2256429            0      339mb
sourcerer-v1-lines~burntsushi~ripgrep            1   1     121406            0     20.5mb
sourcerer-v1-lines~caddyserver~caddy             1   1     423391            0     60.4mb
sourcerer-v1-lines~django~django                 1   1  187595831            0     26.4gb
sourcerer-v1-lines~element-hq~element-web        1   1    5435563            0        1gb
sourcerer-v1-lines~facebook~docusaurus           1   1    1098502            0      173mb
sourcerer-v1-lines~faker-ruby~faker              1   1     281240            0     72.8mb
sourcerer-v1-lines~fastlane~fastlane             1   1    3774359            0    515.3mb
sourcerer-v1-lines~flipt-io~flipt                1   1    2318144            0    339.1mb
sourcerer-v1-lines~fluent~fluentd                1   1     616236            0     87.9mb
sourcerer-v1-lines~fmtlib~fmt                    1   1     525144            0     78.7mb
sourcerer-v1-lines~future-architect~vuls         1   1    1340951            0    410.3mb
sourcerer-v1-lines~gin-gonic~gin                 1   1      42186            0      6.2mb
sourcerer-v1-lines~gohugoio~hugo                 1   1    1323287            0    415.2mb
sourcerer-v1-lines~google~gson                   1   1     269864            0     82.1mb
sourcerer-v1-lines~gravitational~teleport        1   1   35961497            0      5.3gb
sourcerer-v1-lines~hashicorp~terraform           1   1    1913651            0    284.3mb
sourcerer-v1-lines~immutable-js~immutable-js     1   1     132082            0     41.1mb
sourcerer-v1-lines~internetarchive~openlibrary   1   1    6505234            0    903.1mb
sourcerer-v1-lines~javaparser~javaparser         1   1     756387            0    233.7mb
sourcerer-v1-lines~jekyll~jekyll                 1   1      57072            0      9.2mb
sourcerer-v1-lines~jordansissel~fpm              1   1      32285            0      8.8mb
sourcerer-v1-lines~jqlang~jq                     1   1     373367            0     56.1mb
sourcerer-v1-lines~laravel~framework             1   1    4680025            0      1.2gb
sourcerer-v1-lines~matplotlib~matplotlib         1   1   24426468            0      3.6gb
sourcerer-v1-lines~micropython~micropython       1   1    2104199            0      320mb
sourcerer-v1-lines~mrdoob~three.js               1   1    5519778            0   1014.6mb
sourcerer-v1-lines~mwaskom~seaborn               1   1     219103            0       35mb
sourcerer-v1-lines~navidrome~navidrome           1   1    1293741            0    405.1mb
sourcerer-v1-lines~nlohmann~json                 1   1     174859            0     53.5mb
sourcerer-v1-lines~nodebb~nodebb                 1   1    6621106            0      2.2gb
sourcerer-v1-lines~nushell~nushell               1   1    1424029            0    405.3mb
sourcerer-v1-lines~pallets~flask                 1   1      34538            0     11.1mb
sourcerer-v1-lines~php-cs-fixer~php-cs-fixer     1   1    1272280            0    357.6mb
sourcerer-v1-lines~phpoffice~phpspreadsheet      1   1    1999435            0      291mb
sourcerer-v1-lines~preactjs~preact               1   1     975966            0    274.3mb
sourcerer-v1-lines~projectlombok~lombok          1   1    1480367            0    470.7mb
sourcerer-v1-lines~prometheus~prometheus         1   1    1132908            0    484.6mb
sourcerer-v1-lines~protonmail~webclients         1   1   22247802            0      3.2gb
sourcerer-v1-lines~psf~requests                  1   1     337782            0    144.1mb
sourcerer-v1-lines~pydata~xarray                 1   1    2413091            0    714.7mb
sourcerer-v1-lines~pylint-dev~pylint             1   1    1188057            0    351.7mb
sourcerer-v1-lines~pytest-dev~pytest             1   1    1882734            0    532.2mb
sourcerer-v1-lines~qutebrowser~qutebrowser       1   1    6734755            0        1gb
sourcerer-v1-lines~reactivex~rxjava              1   1     486774            0    135.3mb
sourcerer-v1-lines~redis~redis                   1   1    3551629            0        1gb
sourcerer-v1-lines~rubocop~rubocop               1   1    2899319            0    763.3mb
sourcerer-v1-lines~scikit-learn~scikit-learn     1   1   10940081            0      3.2gb
sourcerer-v1-lines~sharkdp~bat                   1   1     317845            0      118mb
sourcerer-v1-lines~sphinx-doc~sphinx             1   1   15102452            0      3.9gb
sourcerer-v1-lines~sympy~sympy                   1   1   45538891            0     13.8gb
sourcerer-v1-lines~tokio-rs~axum                 1   1     146486            0     40.6mb
sourcerer-v1-lines~tokio-rs~tokio                1   1    1068294            0    293.9mb
sourcerer-v1-lines~tutao~tutanota                1   1    2932455            0    985.3mb
sourcerer-v1-lines~uutils~coreutils              1   1     478330            0    127.7mb
sourcerer-v1-lines~valkey-io~valkey              1   1    1876977            0    575.5mb
sourcerer-v1-lines~vuejs~core                    1   1     684268            0    190.9mb
sourcerer-v1-refs                                1   1        847            0    457.5kb<p>Benchmark task prompt:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3b4aec652d09215e/6a816866f6ab8740b5d16c31/Screenshot_2026-08-16_at_10.35.51_a.m..png" alt="" /><p>Table 6. This table compares <a href="http://google.com/url?q=https://github.com/Qiushao-E/SWE-Explore-Bench/blob/main/explorers/claude_code.py%23L30-L47&amp;sa=D&amp;source=docs&amp;ust=1786869254292102&amp;usg=AOvVaw3XAbMzUTtNbrrEmcvVpLY9">Claude Code's task prompt </a>used in the original paper and <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/commands/benchmark/swe_explore_bench/explorer.py#L145-L177">Sourcerer's task prompt</a> used in this benchmark run. The highlights indicate where Sourcerer's task prompt differed from Claude Code's: red indicates an instruction from Claude Code's task prompt that wasn't used by Sourcerer's, and green indicates an instruction that was unique to Sourcerer's task prompt.</p><h3>Appendix B. Code search speed and scalability benchmark configuration</h3><p>Benchmark environment:</p><ul><li><p>Sourcerer version: v2.0.0 (commit hash: <a href="https://github.com/elastic/sourcerer/tree/dcc373b3bf7b24168e00ed7b172aea7227f705f9">dcc373b3bf7b24168e00ed7b172aea7227f705f9</a>)</p></li><li><p>Elastic Cloud Hosted (ECH) for ES|QL:</p></li></ul><ul><li><p>Region: GCP - Los Angeles (us-west2)</p></li><li><p>CPU Optimized hardware (c4a-highcpu)</p></li><li><p>2x Elasticsearch data nodes (each with 16GiB RAM, 8 vCPUs)</p></li><li><p>1x Kibana instances (2GiB RAM)</p></li><li><p>Elastic stack version: v9.5.0 (commit hash: <a href="https://github.com/elastic/elasticsearch/tree/8d4246a64bc255212407b1b313fe402391299c88">8d4246a64bc255212407b1b313fe402391299c88</a>)</p></li></ul><ul><li><p>Virtual machine for <code>ripgrep</code> and <code>grep</code>:</p></li><ul><li><p>Region: GCP - Los Angeles (us-west2-a)</p></li><li><p>CPU Optimized hardware (c4a-highcpu-8-lssd)</p></li><li><p>16GiB RAM</p></li><li><p>Image: projects/ubuntu-os-cloud/global/images/ubuntu-2404-noble-arm64-v20260717</p></li><li><p>Provisioned IOPS: 3300</p></li><li><p>Provisioned throughput: 215</p></li></ul></ul><p>Indices as reported by <code>GET /_cat/indices</code>:</p>index                                           pri rep docs.count docs.deleted store.size
sourcerer-v2-lines~github~elastic~elasticsearch   1   0  200387235            0     44.8gb<p>Cluster settings adjusted to avoid truncating the returned match count on frequent patterns:</p>PUT /_cluster/settings
{
  "persistent": {
    "esql.query.result_truncation_max_size": 1000000
  }
}<p>Regular expression used by <code>ripgrep</code> and <code>grep</code>:</p><ul><li><p>DiskBBQ: <code>.*[dD][iI][sS][kK][-_]?[bB][bB][qQ].*</code></p></li><li><p>XContentType: <code>.*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*</code></p></li></ul><p>ES|QL query syntax used for <code>sourcerer.code.grep</code>:</p><p>ES|QL query syntax used for <code>sourcerer.code.search</code>:</p><p>ES|QL query parameters used in all tests:</p><ul><li><p><code>git_host="github"</code></p></li><li><p><code>git_org="elastic"</code></p></li><li><p><code>git_repo="elasticsearch"</code></p></li><li><p><code>n=1000000</code></p></li></ul><p>ES|QL query parameters used for specific tests:</p><ul><li><p>Corpus with single commit: <code>git_commit="45f6a06b1b441b41fe711059b8720013173e7c89"</code></p></li><li><p>Corpus with all commits: <code>git_commit="*"</code></p></li><li><p><code>sourcerer.code.grep</code> pattern for rare pattern (DiskBBQ): <code>regex=".*[dD][iI][sS][kK][-_]?[bB][bB][qQ].*"</code></p></li><li><p><code>sourcerer.code.grep</code> pattern for common pattern (XContentType): <code>regex=".*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*"</code></p></li><li><p><code>sourcerer.code.search</code> pattern for common pattern (DiskBBQ): <code>q="DiskBBQ"</code></p></li><li><p><code>sourcerer.code.search</code> pattern for common pattern (XContentType): <code>q="XContentType"</code></p></li></ul><p><code>n</code>  was set high enough to recover the true, uncapped match count for each pattern (605 and 1,041 matches for DiskBBQ;  7,999 and 290,662 matches for XContentType), confirmed in each case by comparing <code>hits_returned</code> against <code>ripgrep</code>'s and <code>grep</code>'s own match counts on the same underlying data.</p><p>Command syntax used for <code>ripgrep</code> :</p><p><code>rg -n --no-ignore -j 6</code></p><p>Command syntax used for <code>grep</code>:</p><p><code>grep -E -r -n --binary-files=without-match --exclude-dir=.git</code> </p><p>Directories of each cloned repository snapshot and the approximate total sizes of their nonbinary files tracked by Git, which is the search space of <code>ripgrep</code> and <code>grep</code>:</p>54M     elastic-elasticsearch-v6.0.1
55M     elastic-elasticsearch-v6.1.4
56M     elastic-elasticsearch-v6.2.4
79M     elastic-elasticsearch-v6.3.2
83M     elastic-elasticsearch-v6.4.3
89M     elastic-elasticsearch-v6.5.4
95M     elastic-elasticsearch-v6.6.2
99M     elastic-elasticsearch-v6.7.2
100M    elastic-elasticsearch-v6.8.23
99M     elastic-elasticsearch-v7.0.1
99M     elastic-elasticsearch-v7.1.1
135M    elastic-elasticsearch-v7.10.2
137M    elastic-elasticsearch-v7.11.2
140M    elastic-elasticsearch-v7.12.1
144M    elastic-elasticsearch-v7.13.4
147M    elastic-elasticsearch-v7.14.2
149M    elastic-elasticsearch-v7.15.2
154M    elastic-elasticsearch-v7.16.3
157M    elastic-elasticsearch-v7.17.29
103M    elastic-elasticsearch-v7.2.1
105M    elastic-elasticsearch-v7.3.2
109M    elastic-elasticsearch-v7.4.2
113M    elastic-elasticsearch-v7.5.2
118M    elastic-elasticsearch-v7.6.2
124M    elastic-elasticsearch-v7.7.1
127M    elastic-elasticsearch-v7.8.1
132M    elastic-elasticsearch-v7.9.3
149M    elastic-elasticsearch-v8.0.1
152M    elastic-elasticsearch-v8.1.3
173M    elastic-elasticsearch-v8.10.4
182M    elastic-elasticsearch-v8.11.4
187M    elastic-elasticsearch-v8.12.2
191M    elastic-elasticsearch-v8.13.4
196M    elastic-elasticsearch-v8.14.3
205M    elastic-elasticsearch-v8.15.5
230M    elastic-elasticsearch-v8.16.6
233M    elastic-elasticsearch-v8.17.10
241M    elastic-elasticsearch-v8.18.8
250M    elastic-elasticsearch-v8.19.18
149M    elastic-elasticsearch-v8.2.3
153M    elastic-elasticsearch-v8.3.3
155M    elastic-elasticsearch-v8.4.3
159M    elastic-elasticsearch-v8.5.3
161M    elastic-elasticsearch-v8.6.2
164M    elastic-elasticsearch-v8.7.1
168M    elastic-elasticsearch-v8.8.2
170M    elastic-elasticsearch-v8.9.2
231M    elastic-elasticsearch-v9.0.8
244M    elastic-elasticsearch-v9.1.10
254M    elastic-elasticsearch-v9.2.8
267M    elastic-elasticsearch-v9.3.7
305M    elastic-elasticsearch-v9.4.3]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Dave Moore]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf9ea6549b4ea8708/6a7deaae0da673fe0c57c6ea/image3.png" length="0" type="image/png"/>
    <pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Your AI agent doesn't need your API key: OAuth 2.1 for Elasticsearch MCP server authentication]]></title>
    <description><![CDATA[OAuth 2.1 lets you connect AI agents to the Elasticsearch MCP server with a browser sign-in instead of an API key. Your agent gets a short-lived token tied to your permissions that you can revoke any time.]]></description>
    <content:encoded><![CDATA[<p>Claude Desktop, Cursor or any MCP host can now connect to your Elasticsearch data with a one-time browser sign-in. OAuth 2.1 is now GA for the <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/mcp-server">Agent Builder MCP server</a> in Elastic Cloud Serverless. Instead of pasting an API key into a config file, your agent gets a short-lived token tied to your permissions. Every connection can be audited individually and revoked without affecting anything else on your account. Refresh tokens roll for 30 days, so you rarely need to sign in again. Org owners can see exactly who authorized which agents, and Agent Builder is the first Elastic surface using this model, with the rest of the Elastic API to follow.</p><h2>Why OAuth is better than API keys for MCP server authentication</h2><p>With OAuth, tokens are short-lived credentials that expire on their own, so a leaked token is a narrowing window rather than a standing grant. Every token traces back to an explicit consent: which user, which client, which time.</p><p>In contrast, an API key is a long-lived credential. It lives in a configuration file on the machine that runs the agent, working for whoever uses it until someone rotates or deletes it. That model is manageable for a CI pipeline you wrote and deployed for your team. It gets uncomfortable when the key is held by an AI agent that assembles its own requests, retrieves untrusted content that may contain injected instructions, and sometimes passes context to sub-agents.</p><p>The failure mode is familiar from every credential-leak postmortem: the key ends up somewhere it shouldn't (a log file, a prompt), and from that moment anyone who has the key can access everything its creator could. The audit trail doesn't help much: API key logs tell you that a key with a given name did something, but not who authorized the client that used it or when.</p><p><strong>Attribute</strong></p><p><strong>OAuth 2.1</strong></p><p><strong>API Key</strong></p><p>Credential lifetime</p><p>Short-lived, auto-refreshing</p><p>Long-lived until manually rotated</p><p>Audit trail</p><p>User, client and timestamp per connection</p><p>Key name only</p><p>Revocation</p><p>Per connection, immediate</p><p>Requires key rotation</p><p>Blast radius</p><p>Single client connection</p><p>Everything the key creator can access</p><h2>How to set up OAuth 2.1 for the Elasticsearch MCP server</h2><p>Setup is a one-time step per project:</p><p>1. In your Elastic Cloud Serverless project, open Agent Builder → Tools library → MCP clients → Create MCP client (OAuth). This gives you a client ID and the MCP server URL.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt083d8f85c74de3e0/6a7db08bb8c2e6c02dbe5027/image8.png" alt="Agent Builder MCP clients page showing the Client ID and MCP server URL needed for OAuth setup in Elastic Cloud Serverless" /><p></p><p>2. Add the server to your MCP host. In Cursor or Claude Desktop, the configuration file entry looks like this:</p>{
  "mcpServers": {
        "kibana-mcp": {
          "command": "npx",
          "args": [
            "mcp-remote",    "https://&lt;your-project&gt;.kb.&lt;region&gt;.aws.elastic.cloud/api/agent_builder/mcp",
            "--static-oauth-client-info",
            "{\"client_id\":\"MYCLIENTID111\"}"
          ]
     }
   }
}<p>3. The first time the agent calls a tool, the host opens your browser on an Elastic Cloud consent screen. You sign in with your normal Elastic Cloud credentials and see what the agent is requesting access to.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte83792ca688a17df/6a7db0cac8b7ac2b925221e8/image1.png" alt="Elastic Cloud OAuth consent screen where a user authorizes an MCP client to access Agent Builder in a Serverless project" /><p>4. Click Authorize. This creates an <a href="https://www.elastic.co/docs/deploy-manage/app-connections">application connection</a> between you, your machine, and your project, and Elastic Cloud issues the agent a short-lived access token.</p><p>From there, your agent has access to your project’s data. The mechanics are invisible. When the access token expires, the host refreshes it using a refresh token with a 30-day rolling expiry, so you are not re-authenticating every hour. If you stop trusting a connection, open the application connections page in Elastic Cloud and revoke the connection. Both tokens die immediately, and nothing else about your account changes.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9ce697d344c24ebd/6a7db12a2888394e8707eaaa/image4.png" alt="Application connections page in Elastic Cloud showing an active OAuth MCP client connection with option to revoke access" /><p>The agent acts with the permissions of the user who consented. Calls to <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools">Agent Builder tools</a>, such as ES|QL queries, Workflows, and Streams, are evaluated based on your role. If you want an agent with less access than your own, authorize the connection as a user with a tighter role, or keep using a scoped API key for that workload. </p><h2>How to manage and revoke MCP server connections in Elastic Cloud</h2><p>Org owners can list every active application connection across an organization or a specific project: which MCP hosts were authorized, by whom, and when. Revocation is per connection, so cutting off one misbehaving agent does not disturb the authorizing user's account or any other integration. There is no shared credential to rotate and no blast radius beyond the one client.</p><p>This is the practical difference for teams. An OAuth app connection records who consented, to which client, and when. An API key log shows that <code>agent-key-3</code> queried an index, but not who authorized that agent. </p><h2>What's next for OAuth authentication across the Elastic API</h2><p>Agent Builder is the first Elastic surface behind OAuth, and the same authorization model will carry to the rest of the Elastic API surface, including Elasticsearch and Elastic Cloud management, as we work toward a single Elastic MCP endpoint.</p><p>To try it now, open Agent Builder in an Elastic Cloud Serverless project and create an OAuth client, or start with the <a href="https://www.elastic.co/docs/deploy-manage/app-connections/oauth-clients">documentation</a>. If you don't have a project yet, you can <a href="https://cloud.elastic.co/registration">start a free trial</a>.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-mcp-server-oauth-authentication</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-mcp-server-oauth-authentication</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Elastic Cloud Serverless]]></category>
    <dc:creator><![CDATA[Alex Chalkias]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt08bd254292cdfc15/6a7dafffef5bef01744fa2be/image7.png" length="0" type="image/png"/>
    <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Building context in Elasticsearch: how AI Indices power smarter agents using fewer tokens]]></title>
    <description><![CDATA[Store AI agent context in an AI Index and power smarter agents using fewer tokens. Step-by-step walkthrough with ES|QL and Kibana Workflows included.]]></description>
    <content:encoded><![CDATA[<p>Agents burn tokens exploring your data before they answer anything, inspecting mappings, sampling documents, probing which index to use. Elasticsearch AI Indices let you precompute that work once and store it as a Knowledge Indicator (KI): a structured, searchable record agents retrieve directly instead of rediscovering from scratch. This walkthrough shows you how to build the full pipeline: create an AI Index, generate routing KIs with a Kibana Workflow, and wire them to any agent harness via a portable ES|QL skill. We've also provided a <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/building-context-technical-walkthrough-part-1">notebook</a> if you'd like to run it yourself end to end as you go through the examples in this blog. This is Part 1 in a blog series providing a technical walkthrough to managing your context through KIs and AI indices. </p><p>While AI indices will be included in future Stack releases, today we recommend using Serverless.</p><h2>How it works: AI Index, Kibana Workflows, and the query-ki skill</h2><p>Building context in this walkthrough has three moving parts:</p><ol><li><p>An <strong>AI Index</strong>, where KIs live. It's a regular Elasticsearch index or data stream with a specific naming convention triggering component templates to configure the right mappings automatically.</p></li><li><p><strong>Kibana Workflows</strong>, which read from your data sources, run an LLM to structure content into KIs, and write those KIs into the AI Index.</p></li><li><p>A <strong><code>query-ki</code></strong><strong> skill</strong>, a skill that queries KIs directly from the AI Index using ES|QL, and that a chat agent can call as a tool.</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfb84f57dc630f26b/6a7b37268bcd80ea30262453/image4.png" alt="Architecture: Kibana Workflows write KIs to an AI Index, agents read AI agent context via a query-ki skill" /><p></p><h3>Prerequisites</h3><p>This tutorial assumes you have:</p><ol><li><p>An Elasticsearch Serverless project. You can <a href="https://cloud.elastic.co/registration?onboarding_token=search&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">sign up for a trial</a> if you don't have one.</p></li><li><p>An API key to access your Elasticsearch project. </p></li></ol><h2>Create sample indices for agent routing</h2><p>First, we’ll need some sources. Sources can be data that already exists in your Elasticsearch indices, or external data accessed via connectors or ES|QL data sources. </p><p>For this blog, we’ll create some indices with example data. We’ll start with an example using three datasets: <a href="https://huggingface.co/datasets/BeIR/fiqa">BEIR/fiqa</a> (financial), <a href="https://huggingface.co/datasets/BeIR/nfcorpus">beir-nfcorpus</a> (biomedical/nutrition), and <a href="https://huggingface.co/datasets/BeIR/scifact">beir-scifact</a> (scientific fact-checking). Each index is populated with its own <code>_meta.description</code>. </p><p>Here are the mappings we define for these indices: </p>{
  "beir-fiqa": {
    "mappings": {
      "_meta": {
        "description": "FiQA: financial question answering corpus from StackExchange Finance community posts and web crawls. Covers investments, banking, taxes, and market analysis. BM25-only index."
      },
      "properties": {
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document body text."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document or article title."
          }
        }
      }
    }
  }
}


{
  "beir-nfcorpus": {
    "mappings": {
      "_meta": {
        "description": "NFCorpus: biomedical information retrieval corpus from NutritionFacts.org. Contains nutrition science and medical research documents on diet, disease, and health interventions. BM25-only index."
      },
      "properties": {
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document body text."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document or article title."
          }
        }
      }
    }
  }
}


{
  "beir-scifact": {
    "mappings": {
      "_meta": {
        "description": "SciFact: scientific fact-checking corpus of biomedical research abstracts used to verify factual claims in peer-reviewed literature. BM25-only index."
      },
      "properties": {
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document body text."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document or article title."
          }
        }
      }
    }
  }
}<p>Then, using the above convenience scripts, load a handful of documents into each index with the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk"><code>_bulk</code></a> API.</p><p>Now imagine an agent with a question and the indices we’ve just created. The agent has no idea which one is relevant at the start. Without pre-computed context, it either performs exploratory lookups (mappings, test searches) to figure out which source to use, or searches all three and hopes the merged results contain something useful. Either approach costs tokens, and if you multiply that inefficiency across every query an agent makes, it adds up.</p><h2>Create your AI Index</h2><p>Before generating any KIs, you need an index to store them. We call this an <strong>AI Index</strong>.</p><p>The naming convention is what triggers automatic configuration. Any index whose name starts with <code>ai-index-idx-</code> is a regular index; <code>ai-index-ds-</code> is a data stream. You’ll want to choose data streams for observability use cases, time series data, and when recency is important. Conversely, standard indices are a good choice for static data that will exist for a long while, where recency is not as much of a concern, and may need to occasionally be updated on demand. This naming convention is required for AI indices. </p><p>When Elasticsearch sees the <code>ai-index-</code> prefixes, it automatically applies component templates that configure the right mappings and settings.</p><p>Creating an AI Index is a single call:</p>PUT ai-index-idx-my-corpus<p>To see exactly what the component templates applied, inspect the mappings:</p>GET ai-index-idx-my-corpus/_mapping<p>The response shows the fields every AI Index gets out of the box:</p>{
  "ai-index-idx-my-corpus": {
    "mappings": {
      "properties": {
        "@timestamp": {
          "type": "date"
        },
        "attributes": {
          "type": "flattened"
        },
        "content": {
          "type": "text",
          "fields": {
            "semantic": {
              "type": "semantic_text",
              "inference_id": ".jina-embeddings-v5-text-small"
            }
          }
        },
        "description": {
          "type": "text",
          "fields": {
            "semantic": {
              "type": "semantic_text",
              "inference_id": ".jina-embeddings-v5-text-small"
            }
          }
        },
        "references": {
          "properties": {
            "uri": {
              "type": "keyword"
            }
          }
        },
        "tags": {
          "type": "keyword"
        },
        "title": {
          "type": "text",
          "fields": {
            "semantic": {
              "type": "semantic_text",
              "inference_id": ".jina-embeddings-v5-text-small"
            }
          }
        },
        "type": {
          "type": "keyword"
        }
      }
    }
  }
}<p><code>title</code>, <code>description</code>, and <code>content</code> are each a <code>text</code> field with a <code>.semantic</code> sub-field of type <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">semantic_text</a>, supporting hybrid retrieval.</p><p>Data stream indices (<code>ai-index-ds-*</code>) additionally carry a default 90-day data retention policy. This blog uses a standard index (<code>ai-index-idx-*</code>).</p><h2>Index Metadata as a Knowledge Indicator</h2><p>The target use case for this example is how the <code>query-index-metadata-ki</code> skill can route an agent to the correct Elasticsearch index, even when index or field names are vague. This reduces mistakes from choosing the wrong index or formulating queries based on incomplete schema exploration.</p><p>Since we're creating KIs for our own indices, we can give the LLM a head start: annotate index mappings with human-written <code>_meta.description</code> content. The workflow generates better KIs with more context to work from.</p><p>To address this, we'll manually create a <a href="https://www.elastic.co/docs/reference/kibana">Kibana Workflow</a> that profiles each index and writes routing KIs into the AI Index. The workflow chains four steps:</p><p></p><p><strong>Step</strong></p><p><strong>Type</strong></p><p><strong>What it does</strong></p><p><code>get_mapping</code></p><p><code>elasticsearch.request</code></p><p>Read the mapping, including <code>_meta.description</code> and per-field descriptions.</p><p><code>sample_docs</code></p><p><code>elasticsearch.search</code></p><p>Pull a few real documents so the profile reflects actual value shapes.</p><p><code>profile_index</code></p><p><code>ai.agent</code></p><p>Generate a structured index profile as structured output.</p><p><code>sink_index_ki</code></p><p><code>elasticsearch.bulk</code></p><p>Write the profile into the AI Index as a KI.</p><p>Paste the following YAML into the <a href="https://www.elastic.co/docs/explore-analyze/workflows">Workflows</a> editor:</p>version: '1'
name: beir-index-profile-ki
description: Profile an index into an index-selection Knowledge Indicator.
enabled: true
tags:
  - context-management
  - index-selection

triggers:
  - type: manual

consts:
  indices:
    - beir-fiqa
    - beir-nfcorpus
    - beir-scifact

steps:
  - name: loop_indices
    type: foreach
    foreach: '{{ consts.indices | json }}'
    iteration-on-failure:
      continue: true
    steps:
      - name: get_mapping
        type: elasticsearch.request
        with:
          method: GET
          path: '/{{ foreach.item }}/_mapping'

      - name: sample_docs
        type: elasticsearch.search
        with:
          index: '{{ foreach.item }}'
          size: 3
          query:
            match_all: {}

      - name: profile_index
        type: ai.agent
        timeout: 120s
        with:
          message: &gt;
            You are a data steward building an INDEX PROFILE for an enterprise
            data catalog. Downstream, an AI agent uses these profiles to decide
            WHICH Elasticsearch index to query for a given user question -- this
            is an index-SELECTION aid, not a place to answer the question itself.

            You are given (a) the index name, (b) its Elasticsearch mapping
            including human-written descriptions in `_meta.description` and each
            field's `meta.description`, and (c) a few sample documents. Produce a
            faithful, decision-useful profile. Rules:
            - Ground everything in the provided mapping + samples. Never invent
              fields, values, or purpose. If unknown, use an empty string/array.
            - Optimize for routing: make it obvious what kinds of questions this
              index can authoritatively answer, and what it canNOT.
            - Prefer concrete field names and real example values from the
              samples over vague phrasing.
            - For joins, surface shared keys (e.g. *_id fields) that link this
              index to sibling indices, since cross-index questions hinge on them.

            Index name: {{ foreach.item }}

            Elasticsearch mapping (JSON):
            {{ steps.get_mapping.output | json }}

            Sample documents (JSON):
            {{ steps.sample_docs.output.hits.hits | map: '_source' | json }}
          schema:
            type: object
            properties:
              display_name:
                type: string
                description: A concise human-readable name for what this index represents (&lt;= 8 words).
              purpose:
                type: string
                description: 2-4 sentences describing what this index stores and its role. PRIMARY semantic surface for matching a question to this index.
              answers_questions:
                type: array
                items:
                  type: string
                description: 3-7 representative natural-language questions this index can authoritatively answer.
              does_not_contain:
                type: array
                items:
                  type: string
                description: 1-4 things a searcher might wrongly expect here but that live elsewhere, to prevent mis-routing.
              key_fields:
                type: array
                items:
                  type: string
                description: 3-10 of the most query-relevant fields as "field_name - what it is".
              when_to_use:
                type: string
                description: A single crisp routing heuristic - when should an agent pick THIS index? (&lt;= 30 words).
              example_esql:
                type: string
                description: One realistic, runnable ES|QL query against this index answering one of answers_questions.
            required:
              - display_name
              - purpose
              - answers_questions
              - key_fields
              - when_to_use

      - name: sink_index_ki
        type: elasticsearch.request
        with:
          method: PUT
          path: '/ai-index-idx-my-corpus/_doc/{{ foreach.item | url_encode }}'
          body:
            '@timestamp': '{{ "now" | date: "%Y-%m-%dT%H:%M:%S.%LZ" }}'
            type: index_metadata_entry
            title: '{{ steps.profile_index.output.structured_output.display_name | default: foreach.item }}'
            tags:
              - index-profile
              - '{{ foreach.item }}'
            attributes:
              display_name: '{{ steps.profile_index.output.structured_output.display_name }}'
              purpose: '{{ steps.profile_index.output.structured_output.purpose }}'
              when_to_use: '{{ steps.profile_index.output.structured_output.when_to_use }}'
              answers_questions: '{{ steps.profile_index.output.structured_output.answers_questions | json }}'
              does_not_contain: '{{ steps.profile_index.output.structured_output.does_not_contain | json }}'
              key_fields: '{{ steps.profile_index.output.structured_output.key_fields | json }}'
              example_esql: '{{ steps.profile_index.output.structured_output.example_esql }}'
              source_index: '{{ foreach.item }}'
            content: &gt;
              === SOURCE / PROVENANCE ===
              This is an INDEX PROFILE for routing/index-selection.
              Backing Elasticsearch index: {{ foreach.item }}
              Inspect it directly with ES|QL:
              FROM {{ foreach.item }} | LIMIT 10
              === WHAT THIS INDEX IS ===
              {{ steps.profile_index.output.structured_output.purpose }}
              Questions this index can answer: {{ steps.profile_index.output.structured_output.answers_questions | join: " | " }}
              When to use this index: {{ steps.profile_index.output.structured_output.when_to_use }}
              Example query:
              {{ steps.profile_index.output.structured_output.example_esql }}
            description: &gt;
              Index profile: {{ steps.profile_index.output.structured_output.display_name }}.
              Does NOT contain: {{ steps.profile_index.output.structured_output.does_not_contain | join: "; " }}.
              Key fields: {{ steps.profile_index.output.structured_output.key_fields | join: "; " }}.<p>Let's walk through what this workflow does. We loop over three specified indices with a <code>foreach</code> loop. For each:</p><ol><li><p><code>get_mapping</code> fetches the Elasticsearch index mappings, including any <code>_meta.description</code> annotations we added earlier.</p></li><li><p><code>sample_docs</code> pulls 3 real documents. Concrete examples give the LLM much better signal than schema alone.</p></li><li><p><code>profile_index</code> calls <code>ai.agent</code> with the index name, mappings, and sample documents. The LLM returns structured output describing the index's purpose, key fields, and an example ES|QL query showing how to use it.</p></li><li><p><code>sink_index_ki</code> writes the result into the AI Index as a KI of type <code>index_metadata_entry</code>, keyed on the index name so re-runs are idempotent.</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfe75c3b235a08332/6a7b387e28883919bf07de96/image1.png" alt="Kibana Workflow generating Knowledge Indicators: get_mapping, sample_docs, profile_index, and sink to AI Index" /><p>A few things to point out: </p><ul><li><p>This workflow hard-codes a specific set of indices. In practice, you could derive the list from an index pattern or a dynamic source. </p></li><li><p>The <code>foreach</code> loop also runs iterations sequentially, which is fine for this guide but slow in production because each iteration involves an LLM call. For scale, use <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/composition">workflow.executeAsync</a> or native parallel support. The <a href="https://www.elastic.co/docs/explore-analyze/workflows/reference/cheat-sheet">cheat sheet</a> has tips on both.</p></li><li><p>In the <code>profile_index</code> step, the agent prompt is the special sauce. This is what shapes the accuracy and usefulness of the KIs. </p></li><li><p>Using <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/ai-steps#ai-prompt"><code>ai.prompt</code></a> can improve workflow efficiency (and cost) if you don’t need to load other tools. </p></li><li><p>Cost can be controlled in multiple ways. Richer prompts and structured output often result in higher token utilization, and of course the model you choose significantly impacts total costs. The <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service</a> (EIS) can be a great playground to test different models against the <code>profile_index</code>’s <code>ai.agent</code> step to compare how different models stack up against each other when generating KIs. </p></li></ul><h3>Query your AI Index to verify Knowledge Indicators</h3><p>Once the beir-index-profile-ki workflow runs, query the AI Index directly in the Discover tab using the following ES|QL query to confirm what got written:</p><p></p><p>This will result in the following output: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7a497f27ccbb343c/6a7b38b81f7b5adf6678fdf1/image2.png" alt="ES|QL query results from an AI Index showing three index metadata Knowledge Indicators in Kibana Discover" /><h3>Build a portable skill to retrieve AI agent context</h3><p>Retrieval is a critical component in an AI index. A KI is a document in the AI Index, and finding one is a single ES|QL query. We package that query as a small, portable skill so any agent can call it, regardless of the harness it runs in.</p><p>We write the skill as a SKILL.md: a YAML header with a name and description, followed by markdown instructions. This is the same Agent Skills format that many harnesses, including Claude Code, LangChain's Deep Agents, and others, load directly. </p><p>The harness reads the header content up front, and only pulls in the full instructions when a question matches the description. The one thing the skill asks of the harness is a way to run ES|QL against Elasticsearch.</p><p>Here is a sample <code>query-index-metadata-ki</code> skill: </p>---
name: query-index-metadata-ki
description: &gt;-
  Retrieve Knowledge Indicators (pre-computed context) from the Elasticsearch AI
  Index before answering. Use it to find which index to search (routing profiles).
  Trigger on any question that depends on choosing a data source.
allowed-tools: esql_query
---

# Retrieving Knowledge Indicators

Knowledge Indicators (KIs) live in Elasticsearch indices named `ai-index-*`.
Retrieve them by calling the `esql_query` tool with the query below. Substitute
the user's question for `&lt;query&gt;`, and `index_metadata_entry` as the `&lt;ki_type&gt;` for routing profiles.

```esql
FROM ai-index-idx-* METADATA _id, _index, _score
| WHERE type == "&lt;ki_type&gt;"
| FORK
    (WHERE MATCH(content, "&lt;query&gt;") OR MATCH(description, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
    (WHERE MATCH(content.semantic, "&lt;query&gt;") OR MATCH(description.semantic, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
| FUSE
| SORT _score DESC
| KEEP title, content, description, tags
| LIMIT 5
```

Ground your answer in what the query returns, and cite the KI titles you used. If
nothing relevant comes back, say so rather than guessing.<p>Let’s break down what the skill is doing: </p><ul><li><p>We’re defining <code>index-metadata-entry</code> as a KI type/use case.</p></li><li><p>We’re performing a hybrid ES|QL search on our AI indices, filtering by the appropriate <code>type</code> using RRF as the default method to fuse results.</p></li><li><p>The KI results will directly ground the agent’s answer when determining what indices are relevant to the query.</p></li></ul><p>Because the skill is just instructions plus a query, it travels wherever your agent does. You can point the same file at a Kibana Workflow agent, Claude Code, LangChain Deep Agents, or any other harness without changing a line of it.</p><h3>Connect your AI Index to an agent harness </h3><p>We want to demonstrate how you can use AI indices to query your data with any harness. For these examples, we’ll use LangChain Deep Agents and an OpenAI-compatible key, but any other agent harness can be easily substituted in, including Elastic Agent Builder.</p><p>First, let’s create a baseline to see how an agent will perform without using KIs:</p># Example question: Is there scientific evidence that vitamin D supplementation prevents cancer?
import os
import sys
import time
from elasticsearch import Elasticsearch
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM beir-fiqa | LIMIT 5'.
               Full-text search syntax: WHERE MATCH(field, "value") — not field MATCH "value".
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


@tool
def get_mapping(index: str) -&gt; dict:
    """Return the field mapping for an Elasticsearch index or pattern."""
    return es.indices.get_mapping(index=index).body


baseline_agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query, get_mapping],
    system_prompt=(
        "You are a research assistant with access to three Elasticsearch indices: "
        "beir-fiqa, beir-nfcorpus, and beir-scifact. "
        "You do NOT know which index is relevant for a given question. "
        "Use get_mapping to inspect an index's description and fields, "
        "then query the most relevant one with esql_query. "
        "Ground your answer strictly in what the queries return."
    ),
)

start = time.perf_counter()
result = baseline_agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>Here’s a modified example that could run the same agent, but now with the ability to search AI indices to return KIs: </p># Example question: Is there scientific evidence that vitamin D supplementation prevents cancer?
import os
import sys
import time
from elasticsearch import Elasticsearch
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent
from deepagents.backends.filesystem import FilesystemBackend

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM beir-fiqa | LIMIT 5'.
               Full-text search syntax: WHERE MATCH(field, "value") — not field MATCH "value".
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


backend = FilesystemBackend(root_dir=".", virtual_mode=False)

agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query],
    skills=["skills"],
    backend=backend,
    system_prompt=(
        "You are a research assistant with access to several Elasticsearch indices. "
        "You do NOT know which index is relevant for a given question. "
        "Before searching, always use the query-ki skill with type 'index_metadata_entry' "
        "to retrieve the routing profile for the right index, then query that index directly. "
        "Ground your answer strictly in what the queries return and cite the KI you used for routing."
    ),
)

start = time.perf_counter()
result = agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>This agent will always query the KI indices to get the answer.</p><h2>How much do Knowledge Indicators reduce agent token usage?</h2><p>Since we’re using agents, the results of these scripts are non-deterministic. However, when I ran these results against the query <code>Is there scientific evidence that vitamin D supplementation prevents cancer?</code>, both agents led to the same conclusion, but they took different paths to get there: </p><p>
</p><p>Baseline (No AI Index)</p><p>With AI Index</p><p>Total tool calls</p><p>12</p><p>8</p><p><code>read_file</code> calls</p><p>0</p><p>2</p><p><code>get_mapping</code> calls</p><p>3</p><p>0</p><p><code>esql_query</code> calls</p><p>9</p><p>6 </p><p>Total indices queried</p><p>2 (bounced between <code>beir-scifact</code> and <code>beir-nfcorpus</code>)</p><p>1 (<code>beir-nfcorpus</code>)</p><p>Tokens consumed</p><p>167,763</p><p>92,711</p><p>Latency</p><p>39.58s</p><p>36.15s</p><p>Answer</p><p>Grounded, correct</p><p>Grounded, correct</p><p>The KI answers were both grounded and correct, but an interesting datapoint is the fact that the overall tool usage and token utilization was smaller when using KIs (latency was roughly equivalent). Here’s how both paths went, side by side: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcce26759ded49327/6a7b3983288839683507dea0/image5.png" alt="Agent comparison: 12 tool calls without KIs vs 8 with KIs, 45% fewer tokens, same grounded answer" /><h2>Run the full AI Index pipeline in Serverless</h2><p>This walkthrough offered a deep dive into the build-it-yourself version of AI indices and KIs. In production, you wouldn't hand-write these workflows; a setup agent would generate them, and a feedback loop would refine KIs from the agent's own traces. But the primitives are exactly what you just used: extract KIs with a workflow, store them in an AI Index, and retrieve them with a skill.</p><p>Managing context is key to a relevant and efficient agentic search system, and AI indices are a way to manage this context with the full power of the Elastic stack. Try it out in Serverless and let us know what you think in our <a href="https://discuss.elastic.co/top?period=monthly">Discuss forums</a> or the <code>#stack-kibana</code> channel in our <a href="https://elasticstack.slack.com/signup#/domain-signup">Community Slack</a>! </p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-index-building-context-agents</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-index-building-context-agents</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Kathleen DeRusso,Matt Nowzari ,Apostolos Matsagkas,Peter Pisljar]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt82d8e9495cde2e30/6a7b37020c5aa95ac5f88c47/image3.png" length="0" type="image/png"/>
    <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Faster, cheaper support investigations with precomputed context]]></title>
    <description><![CDATA[Precomputed context cut input tokens by 58% and latency by 40% in Elastic’s support agent, making support investigations more efficient by reducing repeated retrieval.]]></description>
    <content:encoded><![CDATA[<p>When a support engineer is assigned a case, the questions can sound straightforward: What happened, what evidence supports the root cause, and what should happen next? The answers may be spread across a case record, a long conversation feed, linked engineering issues and comments, knowledge articles, and related cases. Precomputed context gathers and organizes evidence from those related records before the agent receives a question. The agent can then begin with the case relationships already identified, rather than reconstructing them during each response. In our evaluation, this approach resulted in lower input token use and latency, without a statistically significant reduction in factuality.</p><p>Before adding precomputed context, the Support team at Elastic used an agent for case investigation, root cause analysis (RCA), triage, and related work. To answer a question, the agent had to discover relevant indices, inspect schemas, issue several queries, reconcile conflicting updates, and assemble a response. Questions ranged from narrow state lookups to multisource investigations:</p><ul><li><p>What is the current status and priority of support case 1234567?</p></li><li><p>After client ABC’s cluster migrated certificates, it returned Secure Sockets Layer (SSL) handshake and certificate_unknown errors. What caused the failure, and how should it be fixed? Which knowledge base articles are relevant?</p></li><li><p>Client ABC’s cluster stopped processing indexing requests and returned 'rejected execution of primary operation' errors. What was the root cause, and how was service restored?</p></li></ul><p>The next question about the same case or about a related case often required much of the same orientation and synthesis work. That repetition consumed input tokens and added latency. It also increased the chance of invalid queries or wrong source selection.</p><p>Earlier work on <a href="https://www.elastic.co/search-labs/blog/pre-computed-context-llm-agent-costs">precomputed context</a> provided a starting point. It described extracting useful context ahead of time and storing it as structured Knowledge Indicators (KIs). Then it explained how an agent retrieves those compact records before scanning raw documents.</p><p>For this support use case, we wanted to understand what context was worth preparing in advance and how the choice of context affected response efficiency and reliability. We also wanted to know what retrieval rules were needed when an agent could use several KI types. This article follows that design process and reports the findings.</p><h2><strong>Starting with index profiles</strong></h2><p>We began with index profile KIs. A workflow read each support index mapping and sampled a few documents. It generated a compact profile describing the index’s purpose, the questions it could answer, important fields, exclusions, verified joins, time coverage, example Elasticsearch Query Language (ES|QL) queries, and more. A representative KI document with selected fields looked like this:</p>{
  "_id": "index-profile-support-cases",
  "_source": {
    "type": "ki",
    "title": "Index profile: Support cases",
    "origin": { "uri": "ki://support-cases" },
    "tags": ["ki-kind:index-profile", "index-selection", "support-data"],
    "content": [
      "BACKING_INDEX: support-cases",
      "PURPOSE: Primary metadata and the customer-reported problem for enterprise support cases (status, product, severity, reported symptom).",
      "QUESTIONS_ANSWERED: What is the status of a case? | Which cases affect a given product or version? | What symptom was reported? | Which high-severity cases are still open?",
      "WHEN_TO_USE: Start here to find or filter cases by status, product, or severity before pulling the conversation thread or root-cause detail from other indices.",
      "KEY_FIELDS: case_number, subject, status, product, severity, created_at, resolved_at",
      "EXAMPLE_QUERY: FROM support-cases | WHERE product == \"elasticsearch\" AND status == \"open\" | KEEP case_number, subject, severity, created_at | LIMIT 20"
    ],
    "description": "Tells the agent which index holds support-case metadata and how to query it, so it can choose the right data source and narrow to relevant cases before deeper analysis."
  }
}<p>The profiles helped the agent choose a source and form a query. They reduced orientation work, especially when index or field names were unclear, but they didn’t remove the main cost of a case investigation. After selecting the indices, the agent still had to gather the case record and reconstruct the conversation. They also still had to follow links to engineering and knowledge sources and then reconcile the evidence. Routing was useful, but the repeated work was synthesis.</p><h2><strong>Changing the unit of context</strong></h2><p>That observation led to case-level KIs. The workflow materializes one evidence-aware snapshot per eligible case. It gathers the case record, recent material conversation events, linked engineering issues and comments, and linked knowledge articles. The resulting KI document contains a stable case identifier and source references. It also contains tags and freshness metadata.</p><p>The design became more detailed as we worked through support specific nuances. Support conversations contain provisional theories, later corrections, administrative status changes, and occasionally separate incidents inside one case. A single fluent summary can flatten those distinctions. The case schema therefore records incident phases, a hypothesis ledger with confirmed, inferred, contradicted, and unresolved states, root cause status and confidence, technical outcome, reusable learning, and more.</p><p>Since large language models (LLMs) can hallucinate some of these details, we added deterministic checks after generation. For example:</p><ul><li><p>An unresolved root cause must be empty and have low confidence. </p></li><li><p>An inferred root cause cannot have high confidence. </p></li><li><p>A case marked closed in a customer relationship management (CRM) system is evaluated separately from whether the technical issue was resolved. </p></li></ul><p>These checks repair structural contradictions without requiring another model call to reinterpret the evidence. A representative KI document with selected fields looked like this:</p>{
  "_id": "1234567",
  "_source": {
    "type": "ki",
    "title": "Case 1234567: Cluster writes blocked after disk flood-stage watermark breach",
    "origin": { "uri": "case://1234567" },
    "tags": [
      "ki-kind:case-intelligence",
      "support-data",
      "cohort:closed",
      "technical-outcome:resolved",
      "rca-confidence:high"
    ],
    "references": [
      { "uri": "case-number://1234567" },
      { "uri": "https://github.com/elastic/elasticsearch/issues/00000" }
    ],
    "content": [
      "CASE_NUMBER: 1234567",
      "STATUS: Closed",
      "PRIORITY: High",
      "SUMMARY: A production cluster stopped accepting writes after a data node crossed the disk flood-stage watermark, which put all indices into read-only mode; freeing disk and clearing the block restored writes.",
      "PROBLEM_AND_IMPACT: Indexing failed cluster-wide with 'FORBIDDEN/12/index read-only'; the customer's ingest pipeline was stalled for 1 hour.",
      "PRODUCTS_AND_COMPONENTS: Elasticsearch, disk-based shard allocation.",
      "INCIDENT_PHASES: Phase 1: disk usage crossed the 95% flood-stage watermark, indices auto-set to read-only. || Phase 2: disk freed and read-only block cleared, writes resumed.",
      "HYPOTHESIS_LEDGER: Flood-stage watermark breach auto-applied a read-only block (confirmed); node stats showed disk at 96% and cluster logs recorded the flood-stage event.",
      "ROOT_CAUSE: A data node exceeded the flood-stage disk watermark, so Elasticsearch automatically applied a cluster-wide read-only block to protect the nodes.",
      "ROOT_CAUSE_STATUS: confirmed",
      "ROOT_CAUSE_CONFIDENCE: high",
      "RESOLUTION_OR_CURRENT_STATE: Freed disk space (removed stale snapshots/indices), cleared the read-only block, and verified writes resumed; recommended more headroom plus watermark alerting.",
      "TECHNICAL_OUTCOME: resolved",
      "REUSABLE_LEARNING: When every index goes read-only at once, check disk watermarks first; the flood-stage block is applied automatically but must be cleared manually after space is freed."
    ],
    "description": "Distilled root cause of a single case: symptom, confirmed root cause with high confidence, resolution, and a reusable lesson so the agent can explain the fix and find precedent for similar cases."
  }
}<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt996a89c947c17745/6a7addd247f9e476c0c2bbe5/image3.png" alt="" /><h2><strong>Generating KIs only when they’re useful</strong></h2><p>After a few iterations, we realized that generating a KI for every case adds cost and noise. Many cases have little usable evidence. In some, there’s only a short intake message; in others, there’s no substantive feed or acknowledgement from a support engineer. So we added gates and filters to the workflow. For example:</p><ul><li><p>A case proceeds when it has linked engineering evidence or at least two material conversational events, including evidence of support participation.</p></li><li><p>The workflow compares the stored KI watermark with the newest timestamp across the case and its material feed. An unchanged case reuses its existing KI instead of regenerating it every time.</p></li><li><p>The workflow limits corpus selection by focusing generation on recent cases or those otherwise likely to be queried, especially when the source corpus is large.</p></li><li><p>The workflow also refreshes the case KI when linked sources can change independently, rather than relying on an incomplete watermark.</p></li></ul><p>Deferred cases remain available through raw lookup and can be reconsidered after new activity arrives. Provenance is part of the stored KI: References point back to linked sources, and the output records what was missing or unverifiable. The KI shortens routine investigation, while raw records remain available for current status, complete history, attachments, and evidence outside the generated snapshot.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf80a0849cde2fe70/6a7added0f118a90462d82fa/image2.png" alt="" /><h2><strong>The retrieval skill is part of the system</strong></h2><p>While workflows generate KIs, the agent requires retrieval guidance to use them effectively. To address this, we developed a retrieval skill which provides ES|QL templates for querying KIs stored in the index and guidance on when to use each retrieval method. For instance, a known case number is searched lexically because it’s an exact identifier. Questions about similar cases or precedents use hybrid lexical and semantic retrieval, while questions about official knowledge, engineering issues, comments, or other non-case entities begin with an index profile and then query the selected raw source.</p><p>The skill also assigns a clear role to each KI family. An index profile provides routing and field guidance, while a case KI provides a bounded snapshot of the evidence relevant to one case. If no case KI is available, the agent treats that as a cache miss and follows the index profile guidance back to the underlying data, rather than assuming that the available context is complete.</p><p>That distinction matters during a typical case lookup. For a question such as <em>What is the current priority of case 1234567?</em>, the agent retrieves the case KI by case <code>_id</code> to understand the investigation and its supporting evidence. It then checks the live case record for values that can change, including status and priority, along with updated_at. If the live record is newer than the KI’s last update, the live value takes precedence. The KI remains useful for durable evidence, while the source record remains the reference for current state. The next eligible refresh rebuilds the KI using timestamps from the case and its material feed, in addition to linked sources.</p><p>This separation was shaped by an evaluation failure. In an earlier version, the agent used the case KI to plan retrieval but issued a raw source query using an inferred schema instead of the fields supplied by the index profile. The query failed, causing retries and additional model work. Later versions of the skill made the precedence rules, source roles, field guidance, and recovery steps for mapping errors explicit.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt14369daa9b806298/6a7ade06288839337f07dced/image1.png" alt="" /><h2><strong>What we observed </strong></h2><p>We evaluated the agent using 20 distinct questions, with approximately 10 for each workflow: case investigation/postmortem summarization and shorter technical support Q&amp;A across several source types. We used three trials per question to measure variation from run to run. The examples in this post were tested on an Elastic Cloud Hosted deployment running Elasticsearch 9.4.2.</p><p>Different agent configurations used the same answering model and comparable tool conditions. They varied only in whether the agent had access to KIs and, if so, which ones: index-profile KIs, case-level KIs, or both. We measured factuality against expected answers, input and output tokens, latency, and tool execution failures. A separate LLM judge scored factuality.</p><p>Case-level KIs produced the clearest signal of operational efficiency in both workflows. Relative to the raw index baseline, observed input token use and latency changed as follows:
</p><p><strong>Workflow</strong></p><p><strong>Input tokens</strong></p><p><strong>Latency</strong></p><p>Case investigation</p><p>About 58% lower</p><p>About 40% lower</p><p>Technical support Q&amp;A</p><p>About 43% lower</p><p>About 17% lower</p><p>We noticed that index profiles shortened source orientation but left consolidation to the agent, whereas case-level KIs supplied a compact evidence bundle aligned with the requested output. That led to fewer raw queries and reduced the opportunity for schema- and partition-related tool errors.</p><p>Factuality varied by workflow, but we didn’t observe a statistically significant reduction in this evaluation.</p><p>A note on interpreting these results: The evaluation was performed on a small dataset, which is common and useful early in the agent development lifecycle when large labeled evaluation datasets aren’t yet available. We therefore treat the results as directional signals for comparing variants. This mirrors strategies outlined in engineering blogs by <a href="https://www.elastic.co/search-labs/blog/ai-agent-evaluation-elastic">Elastic</a> and other companies, such as <a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents">Anthropic</a>: Start with a small number of tasks drawn from real failures, and then expand the suite as effects become smaller and the product matures.</p><h2><strong>Why case-level KIs fit support work</strong></h2><p>The case-level strategy matched the shape of support investigation work in several ways:</p><ul><li><p>A case number is a stable retrieval anchor.</p></li><li><p>The expensive operation is repeated consolidation across the same source relationships.</p></li><li><p>Much of the evidence used for explanation is durable, while volatile states can be verified separately.</p></li><li><p>The KI schema mirrors the type of questions an engineer asks during case summary and RCA work.</p></li><li><p>Maturity and freshness controls limit generation to cases with enough evidence and likely reuse.</p></li></ul><p>Index profiles still have value. They help with source discovery, schema orientation, non-case questions, and fallback. For case investigation, they leave the cross-source reconstruction inside the response path. Case-level KIs remove part of that recurring work, which is the main reason we expected lower token use and latency in this domain.</p><h2><strong>The combined strategy exposed a composition problem</strong></h2><p>While evaluating the responses, we made a counterintuitive observation: Providing both index profiles and case KIs didn’t improve on using case KIs alone. Inspection of the trace showed that the agent understood some of the case context but lacked a dependable rule for composing the two KI families. It used a case-level KI for planning but then ignored the index profile’s schema guidance when querying raw data, leading to failures and suboptimal responses.</p><p>This resulted in a practical lesson for improving guidance in the retrieval skill: Context sources need roles and precedence. The agent must know which source can support an answer, which source only routes to evidence, when verification is required, and what to do after a cache miss or mapping error. Two individually useful context types can create additional work when those contracts are implicit.</p><h2><strong>What we learned about precomputed context</strong></h2><p>In this support workflow, the most useful unit of context was a case and the evidence connected to it: the case record, conversation, linked engineering work, and relevant knowledge. Preparing that evidence ahead of time meant that the agent didn’t have to reconstruct the same relationships for every response. Compared with the raw index baseline, the case-level approach was associated with lower observed input token use and latency, while factuality didn’t show a statistically significant reduction.</p><p>Because case state can change, the system still checks source data for values, such as priority and case status, and falls back to raw data when the available evidence doesn’t justify a case-level KI. Next, we plan to test how well this design holds with support data from external systems, such as Salesforce, and to identify any adjustments needed.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/reduce-token-usage-precomputed-context</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/reduce-token-usage-precomputed-context</guid>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Abhimanyu Anand]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte973a631a35f846f/6a7add9cc8b7acfe04521420/image4.png" length="0" type="image/png"/>
    <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Your agents have been keeping receipts: turning Elastic Agent Builder's built-in OTel traces into token cost dashboards in Kibana]]></title>
    <description><![CDATA[Your Agent Builder agents already log every LLM call as an OTel trace, and that agent tracing data can power token cost dashboards and budget alerts before one runaway conversation quietly wrecks your month.]]></description>
    <content:encoded><![CDATA[<p>Every Elastic Agent Builder conversation already generates a full OpenTelemetry trace. LLM calls, tool executions, token counts, all logged by default into Elasticsearch data streams you can query with ES|QL. Most teams don't look at this data until something breaks, which means they're sitting on usage trends, latency bottlenecks, and cost signals they could have caught earlier. This post covers how to build token cost dashboards in Kibana, set alerts that fire when a conversation blows past 256,000 tokens, and use the waterfall timeline to see exactly where your agent spent its time.</p><h2>What is an Agent Builder OTel trace and what does it capture?</h2><p>When your agent runs, Agent Builder records everything that happened as an <a href="https://opentelemetry.io/docs/concepts/signals/traces/">OpenTelemetry (OTel) trace</a>. Think of a trace as a receipt for a single conversation turn. Every LLM request, tool call, and agent action is recorded as an individual span in Elasticsearch, which is a unit of work or operation. When opted in, additional details like user prompts, LLM responses, tool outputs, and conversation IDs are captured as structured span attributes on the chat span. All of this is scoped to your Kibana space.</p><h2>How to enable agent tracing and privacy controls in Kibana</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc85524772be260c0/6a6a33e610787d6661a77fa1/a0200d33905eee1439500fc3b7bd0c6a74af1fc8-1999x1266.png" alt="Agent Builder Traces settings in Kibana showing tracing toggle and advanced privacy controls for OTel trace data" /><p>To begin capturing trace data, ensure the following toggles under <strong>Agent Traces</strong> in Gen AI Settings are active within your environment:</p><ul><li><p><strong><code>agentBuilder:tracing:enabled</code></strong> — This gen AI setting manages the collection of traces and is enabled by default.</p></li></ul><p>Advanced privacy controls, located under the default tracing toggle, also let you collect message content. While prompts and tool outputs are masked by default, you may choose to enable them to support more robust traces:</p><ul><li><p><strong><code>agentBuilder:tracing:includeUserPrompts</code></strong></p></li><li><p><strong><code>agentBuilder:tracing:includeLlmResponses</code></strong></p></li><li><p><strong><code>agentBuilder:tracing:includeToolDetails</code></strong></p></li><li><p><strong><code>agentBuilder:tracing:includeSystemPrompt</code></strong></p></li><li><p><strong><code>agentBuilder:tracing:includeRealNames</code></strong><strong>:</strong> Retains real agent/tool names instead of anonymizing to custom</p></li><li><p><strong><code>agentBuilder:tracing:includeRealIds</code></strong>: Retains the actual conversation identifiers instead of the default hashed versions. This means trace data collects original IDs, which can link traces to specific user sessions (PII).</p></li></ul><p>Only enable these if you understand what data your agents handle and have appropriate data governance in place.</p><h2>How Agent Builder stores OTel trace data in Elasticsearch</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltadcdb2eeb64bb453/6a6a33e515fc5c4fa99e4935/0b22d6473e530801a7433a662a77051ecbc22814-1999x1436.png" alt="Waterfall view of an Agent Builder OTel trace showing span hierarchy, LLM call durations, and tool executions" /><p>The Agent Builder utilizes OpenTelemetry semantic conventions. This results in a structured hierarchy of spans that provides a granular view of the agent's internal logic:</p><p>Span</p><p>Type</p><p>What it captures</p><p>`invoke_agent &lt;name&gt;`</p><p>CHAIN</p><p>Full turn lifecycle, from user input to final reply</p><p>`invoke_agent &lt;name&gt;`</p><p>AGENT</p><p>Single agent execution: reasoning, tool calls, reply</p><p>`chat &lt;model&gt;`</p><p>LLM</p><p>One LLM request: model, latency, token counts</p><p>`execute_tool &lt;toolName&gt;`</p><p>TOOL</p><p>Tool invocation: arguments, duration, result</p><p>Trace data is written to a dedicated data stream per Kibana space, keeping conversation data cleanly isolated. To query your traces in Discover, target the index for your space directly:</p><p>For the default space, that’s <code>traces-agent_builder.otel-default</code>. If the advanced privacy controls are turned on, then those span attributes will also be shipped to the traces data stream with the original spans. This index lets you query the raw message content to see what's actually being said in conversations. It is best practice to avoid using wildcards to prevent mixing data from unrelated spaces.</p><p>Agent Builder ships with a built-in skill called <code>agent-builder-traces</code>, installed automatically when<code>agentBuilder:tracing:enabled</code> is on. You can use it to ask questions directly about your trace data, making it easy to explore agent behavior without writing ES|QL from scratch.</p><h2>How to debug agent behaviour with the OTel trace waterfall view</h2><p>The trace waterfall shows every step of an Agent Builder session as a timeline. To open it, navigate to the specific turn in the conversation UI and select the trace icon.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0c19c6699996667b/6a6a33e704e0aa0bbe5528f9/29984c0b4f0c290170ceb48e01e404bdc25db92d-602x158.png" alt="Trace icon in the Agent Builder conversation UI used to open the OTel trace waterfall view" /><p>This launches a waterfall timeline breaking down every step of your agent's execution. At the top level, you'll see the <code>invoke_agent</code> parent span with the full end-to-end duration of your agent run. Nested beneath it are chat spans, each representing a single LLM request and showing exactly how long the model took to respond. Alongside those are <code>execute_tool</code> spans, one per tool call, where you can see which tool was called, what arguments it received, and how long it ran. This allows you to trace the exact sequence your agent followed, pinpoint timing bottlenecks, and see where errors occurred.</p><h2>How to build token cost dashboards from trace data</h2><p>Discover gives you raw trace data, but most teams want answers to operational questions like "how many tokens did we burn today?", "which tool is called most often?", and "how many unique users interacted with the agent this week?". These would require a dashboard built directly against the trace data.</p><p>There is an Elastic-managed out-of-the-box dashboard called <em>[Elastic] Agent Builder Overview</em> that can be installed by clicking in the top-right corner of the <em>Agent Traces</em> section within GenAI Settings.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt21bd488325991b1f/6a6a33e8d57c1dc19ac13eec/805391d2c18083589b4e64f450e2bc33ba32be9a-1999x117.avif" alt="" /><p>It contains basic details spanning Token Usage and Cost, Conversation Volume and Latency, Agent Execution, and Tool Call Frequency and Errors.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5775187fb11ae16e/6a6a33e955755b13022bd23e/70b96d6ff2e96358831a1bf8a08aeecec8c41b0e-1999x1115.png" alt="Elastic Agent Builder Overview dashboard showing token usage, cost metrics, and LLM request counts" /><p>However, a custom dashboard may be more efficient. If you want something more tailored, build <a href="https://www.elastic.co/docs/explore-analyze/visualize/lens">Lens</a> panels directly against the OTel trace index. Some useful panels to start with:</p><ol><li><p>Most active conversations by token spend</p></li></ol><p>create a horizontal bar chart against <code>traces-agent_builder.otel-&lt;space-id&gt;</code>. Set the x-axis to <code>gen_ai.conversation.id</code> and sort descending and limit to the top 10. Set the y-axis to a sum of <code>gen_ai.usage.input_tokens</code>plus<code>gen_ai.usage.output_tokens</code>. Input the formula as: </p><p>Conversations with the most LLM round-trips</p><p>Option to create this visualization with a simple ES|QL query that would look like:</p><p>There is also a <code>dashboard-management</code> skill that can be used to help create traces visualizations using natural language.</p><h2>How to set token cost alerts for Agent Builder conversations</h2><p>Token consumption is the most direct cost lever for LLM-based agents. A single runaway conversation can blow through your monthly budget before anyone notices.</p><p>Elastic alerting lets you define a threshold rule directly against the trace data. Navigate toObservability &gt; Alerts &gt; Manage Rules &gt; Create Ruleand selectElasticsearch queryas the rule type.</p><p>A rule that fires when any single conversation exceeds 256,000 tokens looks like this as an ES|QL rule:</p><p>Set the schedule to run every 15 minutes and configure the action to send a Slack notification or open a PagerDuty incident. The <code>gen_ai.conversation.id</code> value in the alert payload gives you the exact conversation to inspect.</p><h2>What’s coming next for Agent Builder observability</h2><p>Agent traces give you visibility that goes far beyond debugging. Once you've built dashboards and configured alerts against Agent Builder trace data, you have a live pulse on how your agents are behaving in production. If you haven't already, spin up Agent Builder in your Kibana space, make sure tracing is enabled, and run a few conversations. Check Discover, pull up the waterfall view, and see what your agent is actually doing under the hood.</p><p>This is the first in a series of posts on Agent Builder observability. Coming up, we'll go deeper on using the<code>agent-builder-traces</code> skill to query your data conversationally, building custom evaluation pipelines from trace data, and using traces to feed conversation history back into your agents. Your agents have been keeping secrets. It's time to make them talk.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/opentelemetry-tracing-agent-builder</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/opentelemetry-tracing-agent-builder</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Operations]]></category>
    <dc:creator><![CDATA[Meghan Murphy,Pablo Neves Machado]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte3979255ddfc7f45/6a17e25ffaa913812f93c7cb/92c517a2e7b36122a18feee317a0215981b62b6b-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[One prompt, a complete workflow: Elastic's AI agent writes your automation for you]]></title>
    <description><![CDATA[Elastic Workflows takes a plain-text prompt and generates YAML you can inspect, version and run against your Elasticsearch data. Now GA, with human-in-the-loop workflows in Slack, parallel execution, and 10 new connectors.]]></description>
    <content:encoded><![CDATA[<h2>One prompt, a complete workflow: Elastic's AI agent writes your automation for you</h2><p>Elastic Workflows now writes its own YAML. You type what you want automated in plain language, the Elastic AI Agent generates a complete workflow against a typed schema, and nothing runs until you've read it. YAML is why this works: it gives the model a constrained, well-typed target, so what comes back is actual building blocks you can edit and run.</p><p><strong>What is new in Elastic Workflows 9.5:</strong></p><ul><li><p><strong>Natural language authoring</strong> is GA and on by default: describe an automation, and the Elastic AI Agent writes the workflow, you review and run it.</p></li><li><p><strong>Versioning</strong> with diff and one-click rollback is GA.</p></li><li><p>Three new experimental previews (behind an advanced setting): a visual mode that renders a workflow as a graph, human-in-the-loop steps that reach people in Slack for input or approval, and parallel execution.</p></li><li><p>More to build on: new connectors, event triggers that react to Cases activity, token metering for AI steps, and a queue strategy for concurrency.</p></li></ul><p>Workflows is the automation engine built into the Elastic platform. It reached general availability in 9.4, enabled by default and running against your Elasticsearch data with the connectors and access controls you already have. This post walks through what 9.5 adds.</p><h2>Why YAML makes AI workflow automation work</h2><p>YAML is the authoring language for Elastic Workflows because it's declarative, version-controllable, diffable, and portable across environments. It reads the same in a pull request as it does in the editor.</p><p>It was also a bet. Large language models (LLMs) are very good at generating structured, well-typed content, and a workflow language is close to an ideal target for that. Ask for prose and a model can wander. Ask for a workflow against a typed schema, with named step types and validated inputs, and there is a right shape for the answer.</p><p>In 9.5 that bet pays off, and it is GA. Inside the workflow editor, you write what you want in plain language:</p><p>When a detection alert fires for a host, pull the last 24 hours of related logs, ask the AI step to summarize what happened, and post the summary to the on-call Slack channel.</p><p>The Elastic AI Agent generates the workflow: the trigger, the Elasticsearch query, the AI summarize step, the Slack step, wired together with the right inputs and outputs. You get inspectable, editable YAML back. Nothing runs until you read it, adjust it, and decide to run it. You can also point it at a workflow you already have and describe the change you want, and it edits in place.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt024bee0a140cf88c/6a6a33d32e4eb74ad5d64368/7cd5a718081ce86a246df7fda74d17a2e8b7613f-1999x1124.png" alt="AI workflow builder generating YAML from a natural language prompt in the Elastic Workflows editor with AI Agent panel" /><p>The output is worth reading because there is a rich, well-typed language underneath it. A short prompt expands into real building blocks:</p><ul><li><p><code>foreach</code> and <code>while</code> loops, with guardrails that stop runaway execution.</p></li><li><p><code>switch</code> for clean multi-way branching.</p></li><li><p>data steps like <code>data.filter</code> and <code>data.aggregate</code> for in-flight transforms.</p></li><li><p><code>on-failure</code> handling on every step, so you can retry, continue, or abort.</p></li><li><p><code>workflow.execute</code>, so one workflow can call another and you assemble new automation from pieces you have already tested.</p></li></ul><p>Natural language gets you the first draft fast; the language underneath is what makes that draft real.</p><h2>Workflow versioning with diff and one-click rollback</h2><p>Versioning is GA in 9.5. Every workflow now has version history: every change is tracked and diffable, and you can roll back to any prior version in one click. You see who changed what and when, compare any two versions side by side, and undo a bad edit without reconstructing it by hand. This is the change-control foundation teams asked for before they would run automation against production systems.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9901bbd63fbe380f/6a6a33dc99442c082ddf1d88/63eecf4cdb28702d03752e0dacd99c7ff9b8f285-1920x1080.gif" alt="Toggling between the Elastic Workflows YAML editor and the visual graph view of a workflow's steps and branches" /><p>Versioning pairs with the production controls already in place: granular role-based access control (RBAC) over who creates, edits, runs, and views workflows, every management action written to the security audit log, and import/export that moves workflows between environments with their connector references intact.</p><p>Versioning in the product is the near end of a longer arc. A workflow is a declarative YAML definition, plain text with a well-defined schema, which means it already fits the tooling built for code: it can be diffed, reviewed, and version-controlled. Where we are headed is full, bidirectional integration with the version control systems you already use, so that a workflow could live in your repository, move through review, and deploy the same way the rest of your software does. That is coming, and the same bet that made natural language authoring work, a declarative and well-typed language, is what will let you manage workflows as code.</p><h2>Visual mode, human-in-the-loop workflows, and parallel execution</h2><p>Three of the newest 9.5 additions ship in Experimental. To try them, turn on <strong>Elastic Workflows: Experimental Features</strong> in <strong>Stack Management → Advanced Settings</strong> (it requires a page reload). Here is what each one does.</p><h3>Visual workflow editor: see the logic as a graph</h3><p>You can now switch a workflow between the YAML editor and a visual mode that renders the workflow as a graph. The graph lays out your steps, branches, and flow control, so you can see the logic and the paths a run can take at a glance, alongside the YAML. It is read-only in 9.5: you still author in YAML, and the graph stays in sync as you edit. It is the first step toward a full drag-and-drop builder, which is coming next.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4c3f9054539e5440/6a6a33e3c40efb720bd3094c/4f225a895d29df4d5af136167024f1403adfbc27-1920x1080.gif" alt=" Elastic Workflows YAML editor showing a security workflow with foreach loops, conditional logic and on-failure handling" /><h3>Human-in-the-loop workflows with approval steps in Slack</h3><p>Workflows could already pause for a person in 9.4 with <code>waitForInput</code>, which presents a schema-defined form and lets the response drive what happens next. 9.5 adds <code>waitForApproval</code> for the most common case of all, a binary approve or reject with labels you choose, and it extends both steps to their first external surface: Slack. Both pause the run until someone answers, with a timeout so a workflow never hangs forever. Not every decision should be fully automated, and this is how you put the human on exactly the steps that need one.</p><p><code>waitForInput</code> is the flexible one: define a schema for the input you want back, and the response comes through typed. Reach for it when the choice is more than yes or no. Here, an Observability workflow has caught a service-level objective (SLO) burn-rate alert on <code>payment-service</code> and asks the on-call engineer which mitigation to run:</p>- name: ask_sre
  type: waitForInput
  with:
    message: "payment-service is burning its error budget. Which mitigation should we run?"
    schema:
      type: object
      properties:
        mitigation:
          type: string
          enum: [restart, scale_up, monitor]
        reason:
          type: string
      required: [mitigation]
    channels:
      slack_api:
        connector-id: my-slack-connector
        channels: ["sre-oncall"]<p><code>waitForApproval</code> is the new one in 9.5, for a straight approve or reject. Here a security workflow has decided a host should be contained, but gates that destructive action on a human before it runs:</p>- name: request_containment_approval
  type: waitForApproval
  timeout: 24h
  with:
    message: &gt;
      Isolate {{ event.alerts[0].host.name }}? This will cut the host off from
      the network until it is manually released.
    approveLabel: Isolate host
    rejectLabel: Leave connected
    channels:
      slack_api:
        connector-id: my-slack-connector
        channels: ["soc-response"]
- name: act_on_decision
  type: switch
  expression: "{{ steps.request_containment_approval.output.response.approved }}"
  cases:
    - match: "true"
      steps:
        - name: isolate_host
          # ... run the containment action
    - match: "false"
      steps:
        - name: keep_monitoring
          # ... skip containment, keep watching<p>The new piece is the <code>channels</code> block. A wait step can now deliver to where people already are, and in 9.5 that means Slack: the workflow posts the request to a Slack channel, the person responds from Slack, and the workflow resumes with their answer. No one has to be sitting in Kibana for the automation to move. Slack is the first external surface, and richer delivery experiences across more channels are on the way.</p><h3>Parallel execution: run independent workflow steps at once</h3><p>By default a workflow runs one step after another, which is what you want when each step depends on the last. But plenty of work does not: enriching an alert from three sources, checking a file against several reputation services, investigating a handful of leads. Run those sequentially and the workflow is only as fast as the sum of its parts, when it could be as fast as the slowest one. The new <code>parallel</code> step lets you run independent work at the same time. It works two ways.</p><p>The first is when you know the work ahead of time. You define a fixed set of tasks, and they run at the same time, so you gather all the results in one step instead of waiting for each in turn. Enriching a security alert from two sources at once is the classic case:</p>- name: enrich
  type: parallel
  branches:
    - name: virustotal
      steps:
        - name: scan_hash
          type: virustotal.scanFileHash
          # ... pass the alert's file hash
    - name: ip_reputation
      steps:
        - name: check_ip
          type: abuseipdb.checkIp
          # ... pass the alert's source IP<p>The second is when you do not know the work ahead of time. You give the step a list, and it runs the same work once per item, all at the same time, up to a concurrency limit you set. Root cause analysis is a good example. An earlier AI step generates a set of hypotheses for why a service is degrading, and you do not know in advance how many there will be or what they are. Rather than investigate them one after another, you pass the list into a parallel step, and it investigates every hypothesis at once:</p>- name: investigate_hypotheses
  type: parallel
  foreach: "{{ steps.generate_hypotheses.output.hypotheses }}"
  concurrency:
    max: 5
  steps:
    - name: investigate_hypothesis
      type: ai.agent
      # runs once per hypothesis, up to 5 at a time
      # the agent gathers evidence for {{ foreach.item }} and scores it<p>You control how many run at once with <strong>concurrency</strong>, and the engine caps both the concurrency and the total number of parallel tasks so a workflow cannot spawn unbounded work. All the results are available to the next step, so the workflow runs the parallel work, then continues once every task finishes.</p><h2>More in Elastic Workflows: connectors, triggers, token metering and concurrency</h2><p>Beyond the headline features, 9.5 widens what a workflow can reach and react to.</p><h3>New connectors: BigQuery, Snowflake, HubSpot, Cortex XSOAR and more</h3><p>The connector catalog keeps growing, with native connectors added in 9.5 for:</p><ul><li><p>BigQuery</p></li><li><p>Snowflake</p></li><li><p>Box</p></li><li><p>Dropbox</p></li><li><p>OneDrive</p></li><li><p>Outlook</p></li><li><p>Azure Blob</p></li><li><p>Google Cloud Functions</p></li><li><p>HubSpot</p></li><li><p>Cortex XSOAR connector for security automation</p></li></ul><p>More are on the way, and when there is not a dedicated connector for the system you need, the <code>http</code> step is the escape hatch: it can securely call any API endpoint, with credentials supplied by a connector rather than written into the YAML.</p><h3>Event-driven triggers for Elastic Cases</h3><p>A workflow starts from a trigger, and 9.5 widens what a workflow can respond to. Cases now emit events a workflow can subscribe to:</p><ul><li><p>A case is created.</p></li><li><p>A case is updated.</p></li><li><p>Its status changes.</p></li><li><p>A comment is added.</p></li><li><p>An attachment is added.</p></li></ul><p>So a workflow can run the moment a case opens, to enrich it, tag it, or notify the right channel, or when its status flips to a state you care about, rather than polling for changes. Alert-triggered workflows also receive richer rule context now, including the rule's tags, type, and parameters, so the workflow has more to work with before it acts.</p><h3>Token usage and cost tracking for AI workflow steps</h3><p>Workflows can call AI steps: <code>ai.prompt</code> for a freeform prompt, <code>ai.classify</code> to sort something into categories, <code>ai.agent</code> to hand a task to an Agent Builder agent. In an automation that runs thousands of times a day, those calls add up. 9.5 now reports token usage for every AI step, input, output, cached, and total, both per step and for the whole run. You can see exactly what the AI in a workflow consumes, track it over time, and tune a prompt or a model choice with the numbers in front of you.</p><h3>Workflow concurrency: cancel, drop or queue</h3><p>A workflow's concurrency setting decides what happens when a new execution starts before the last one finishes. 9.5 adds a third strategy, so you can pick the behavior that fits the workflow:</p><p>Strategy</p><p>Use it when</p><p>Cancel-in-progress</p><p>Only the latest execution matters, like recomputing a current state</p><p>Drop</p><p>An execution already in flight covers the situation and extras are redundant</p><p>Queue (new in 9.5)</p><p>Every execution matters and order does, so they line up and run one after another, with a queue size and time-to-live you control</p><p>Queue is what you reach for when executions touch the same resource or should not overlap. Audit logging also covers more of the lifecycle in 9.5, including restoring a workflow from its version history.</p><h2>Get started with Elastic Workflows</h2><p>The fastest way to see this is to describe something you want automated. Open the workflow editor in 9.5, type it in plain language, and read the YAML that comes back. Natural language authoring is on by default. To try the visual mode, human-in-the-loop steps, and parallel execution, turn on <strong>Elastic Workflows: Experimental Features</strong> in <strong>Stack Management → Advanced Settings</strong>.</p><p>The theme across 9.5 is a shorter path from idea to running automation. You describe what you want and AI drafts it, you see it as a graph and version it as you go, you pause it for a person in Slack when a step needs judgment, and you run independent work in parallel. For the full details, see the <a href="https://www.elastic.co/docs/explore-analyze/workflows">Workflows documentation</a>.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-workflow-automation-natural-language</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-workflow-automation-natural-language</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Integrations]]></category>
    <dc:creator><![CDATA[Tinsae Erkailo,Shahar Glazner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltda47d75430c4fa7c/6a17f5cae3179149242d5963/d5d04bbcfc3925f48f3487ea4c7e0dd2205316d0-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How Elasticsearch detects multiple change points in time series with 0.99 recall]]></title>
    <description><![CDATA[ES|QL's CHANGE_POINT command finds structural shifts, variance changes and spikes in any metric in ~1ms, without tuning anything per series.]]></description>
    <content:encoded><![CDATA[<p>The current generation of agentic models are remarkably good system troubleshooters. Given a hypothesis and the means to test it, they can reason about a failing service much the way a seasoned SRE does: form a theory, look for corroborating evidence, discard it when the data disagrees, and narrow in on a root cause. Their main limitation is not their ability to reason but their reach: they can only investigate what their tools let them see.</p><p>This is where Elasticsearch earns its place in the stack. It is already where a great deal of operational telemetry lives (logs, metrics, traces, events), and it exposes them through an expressive query and aggregation layer. That makes it a natural tool for an agent debugging a live system: it can slice by attribute, aggregate over time, correlate across signals, and drill from a symptom down to the documents that produced it.</p><p>We've been building an agentic layer on top of Elasticsearch that continuously monitors a system and root-causes issues as they arise. A recurring primitive in that workflow is time series event analysis. For example, given an error rate, a p99 latency, a queue depth and a throughput counter, tell me whether something happened, what it was and when. A transient spike in errors, a regime change in latency, and a step up in CPU usage are the signatures of the underlying fault, and they're typically what an agent examines first as it forms and tests hypotheses.</p><p>Elasticsearch has shipped a single-change-point aggregation for some time. It answers "did this series change?" with one verdict and the most significant change it found. That's a good fit for a dashboard, but less so for an agent, which often wants to interrogate a long window and enumerate everything of interest in it: the error spike at 02:14, the latency regime shift at 02:30, the throughput dip while the pod was being rescheduled. So we upgraded the capability to detect and report multiple events of multiple kinds in a single series. At the same time, we took the opportunity to further harden it to work reliably against whatever telemetry the agent points it at. This post describes how it works.</p><h2>Why single change point detection isn't enough for agents</h2><p>Concretely, we want a single entry point that takes a numeric time series and returns a small list of interesting events, each with a type, a location, a significance, and some key characteristics. We care about three classes of event, because they map onto three different kinds of underlying fault:</p><p>Event type</p><p>What changes</p><p>Detection channel</p><p>Example fault</p><p>Structural change</p><p>Level (step) or slope (trend) shifts to a new sustained regime</p><p>Value channel</p><p>Config push doubles baseline latency; memory leak turns a flat curve into a ramp</p><p>Distribution change</p><p>Noise level (variance) shifts while the mean holds steady</p><p>Dispersion channel</p><p>Service responds erratically at the same average latency</p><p>Point anomaly</p><p>Isolated spike or dip against a stable background</p><p>Value channel (pulse detector)</p><p>Single burst of errors; one-minute throughput drop during GC pause</p><p>The hard part is not detecting any one of these on clean, well-behaved data. The hard part is doing it on arbitrary telemetry without per-series tuning. The agent does not know in advance whether the series it is examining is near-constant, smoothly drifting, <a href="https://en.wikipedia.org/wiki/Homoscedasticity_and_heteroscedasticity">heteroscedastic</a> (quiet in places and noisy in others), sparsely populated, or has a magnitude of . It is not scalable to hand-pick parameters for every series it needs to analyze. Whatever we build has to be robust to all of that while maintaining excellent recall and precision. If it fails to detect important events it runs the risk of missing key corroborating evidence for a working hypothesis. Conversely, an analysis tool that reports an event for every minor fluctuation will pollute the context the agent reasons over.</p><p>The design goals, in priority order, are: correct on diverse data out of the box, parsimonious (report only what matters), and cheap enough to run interactively across many series.</p><h2>How PELT and BIC power change point detection</h2><p>Change-point detection is a well studied field. The classical offline formulation searches for the segmentation of a series that minimizes a penalized cost: a per-segment goodness-of-fit term plus a penalty for each added break to stop the optimizer from putting a boundary between every pair of points. Solved naively, this is combinatorial, but PELT (<a href="https://arxiv.org/pdf/1101.1438">Pruned Exact Linear Time, Killick et al.</a>) finds the optimal partition in roughly linear time by using a dynamic program to prune candidate boundaries that can never be optimal. On the labeling side, comparing nested models by an information criterion such as the Bayesian Information Criterion (BIC) gives a principled, scale-aware way to decide whether a candidate break is really important and what sort of change it constitutes.</p><p>These are good building blocks, and we use them. However, the textbook recipe assumes more than telemetry gives you. It typically assumes a single change type (a mean shift), a known and stationary noise level, and reasonably benign numerics. Real telemetry violates all three: variance changes matter as much as mean changes, the noise level is unknown, often heavy-tailed and changing, and the data spans extreme magnitudes and degenerate cases, such as perfectly constant segments, that wreck an ill-conditioned polynomial fit or a fixed-variance cost. Most of the engineering I describe below is about closing that gap.</p><h2>Splitting one time series into three detection channels</h2><p>Rather than trying to find one detector that does everything, we run three focused detectors and then merge their findings. Two of the three are the same structural detector applied to two different views of the data, or channels.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blted8fd73801bac283/6a6a33ed15fc5c4ade9e493d/02fa7892061a4343a7720c7269c8df82ef67603f-1508x1132.png" alt="Elasticsearch time series split into value and dispersion channels detecting step changes, spikes and variance shifts" /><p>Two detectors run on the value channel: a structural detector that flags step and trend changes, and a pulse detector that identifies point spikes and dips. The dispersion channel, a windowed measure of spread, is fed to a second copy of the structural detector, where a variance change shows up as a level shift and is relabeled a distribution change. The two channels are complementary by construction: a step is a level shift the value channel flags but only has a single large first difference the dispersion channel ignores, while a variance change is invisible to the value channel yet shows up clearly in the dispersion channel. A thin orchestration layer then merges and de-duplicates the streams.</p><p>Keeping the three concerns separate makes each one tractable. A mean-shift detector and a variance-shift detector pull in opposite directions if you try to fuse them; a point-anomaly detector and a regime detector need opposite robustness settings. Separated, each can be tuned to its job.</p><h3>PELT with a scale-free cost</h3><p>For the structural channel, we fit each candidate segment with a low-order polynomial (constant or linear) and score it with the profiled-variance Gaussian cost. If a segment  of length  has residual sum of squares , the cost is</p><p>This is the negative log-likelihood of a Gaussian segment after profiling out the variance, i.e., after substituting the maximum-likelihood estimate  back into the log-likelihood. The total objective PELT minimizes is the sum of segment costs plus a per-break penalty,</p><p>Here,  is the BIC complexity term (number of parameters times  where  is the number of values in the time series) and the scale factor lets us trade sensitivity against parsimony in one place.</p><p>Using the profiled cost rather than a cost against a fixed global variance is a deliberate and important choice. The global noise level of telemetry is unreliable: on a smoothly varying series the natural estimate (the spread of first differences) can collapse toward zero, and a fixed-variance cost then treats every wiggle as enormously significant and over-segments. The profiled cost depends only on the ratio , so it is invariant to the absolute scale and immune to that failure. A small floor on  keeps the logarithm finite, so a zero-residual segment is not rewarded without bound.</p><h3>Keeping the fit stable using robust local weighting</h3><p>Before PELT runs, we robustly down-weight points so that an excursion does not create spurious breaks or drag a segment boundary onto itself. Crucially, these weights enter into the weighted residual moments, so they shape not just the segment fit but its residual variance, and so the segment costs themselves. Each point gets a Cauchy weight, , of its residual  from a rolling-median baseline against a robust scale : points near the local median keep full weight and points far from it are progressively discounted. A nice side effect of measuring excursions from the median on a window centered on each point is that clean structural breaks do not get down-weighted at all, because the majority of values land on the same side of the break as the point whose residual is being computed. So a sustained regime keeps full weight, but a lone spike does not.</p><p>There is a nice justification for scoring a weighted Gaussian cost when what we really want is robustness to a heavy tail. The Cauchy weight is exactly the <a href="https://en.wikipedia.org/wiki/Iteratively_reweighted_least_squares">iteratively reweighted least squares</a> (IRLS) weight of its loss: , so the weighted normal equations  are identical to the Cauchy M-estimator's estimating equations . A weighted-mean (or weighted-line) fit at those weights is therefore a <a href="https://en.wikipedia.org/wiki/M-estimator">Cauchy M-estimate</a>, not a Gaussian one. The cost we actually evaluate inherits the same properties. Because  is concave in , its tangent at the current residual lies above it. This gives a pointwise bound  with  the weight at the tangent point; summing, the weighted residual sum of squares  is a tangent upper bound on the total Cauchy loss, touching it in both value and gradient at the weights' anchor point. Minimizing the weighted RSS is thus one step of a <a href="https://en.wikipedia.org/wiki/MM_algorithm">majorize–minimize scheme</a> that provably decreases the true Cauchy objective, and the profiled-variance cost we feed the BIC is that majorizer standing in for the Cauchy deviance. The only approximation is that we anchor the weights once, at the rolling-median baseline, rather than iterating IRLS to its fixed point; this is exact for the inliers that sit near the baseline, and correct for gross outliers, whose vanishing weight removes them from the cost wherever the bound is loosest.</p><p>Finally, the trick that makes this work on heteroscedastic data is that the residual is judged primarily against a <em>local</em> robust scale, not a <em>global</em> one: the MAD of residuals in the same sliding window. On a series that is quiet in one stretch and noisy in another, this means a spike in the quiet stretch that is multiple local sigmas is correctly suppressed. The local MAD can collapse on quiet stretches, so we use a backstop that is a fraction of a global composite of robust scales and a floor related to the quantization error for discrete series and numerical precision otherwise.</p><h3>From candidates to labeled events using BIC verification</h3><p>PELT gives a globally optimal penalized segmentation, so we take its boundaries as candidates and verify each one. For a candidate at index  we look at the window of length  spanning to its nearest neighboring candidates and compare a no-change null against step and trend alternatives by BIC,</p><p>where  counts the fitted parameters (the same parameter count as in the PELT penalty above). We map the BIC gain of an alternative over the null to a significance via , to turn a threshold into a decision boundary). We treat this  as a significance score for ranking and thresholding, not as a calibrated tail probability.</p><p>At this stage we allow higher-degree models to avoid splitting smoothly varying trends. These are problematic in PELT itself because it considers short segments, which they overfit. We keep the most parsimonious alternative that clears the significance threshold and survives a persistence check. The persistence check re-scores with the weights immediately around the candidate muted, and if the evidence collapses, the "change" was driven by a few extreme points – an excursion, not a regime change – and we reject it. The polynomial order is applied symmetrically to the null and to each side of the split, so the alternative is always the same model class merely split at the candidate, and therefore strictly more flexible. The whole process can be thought of as Bayesian model selection with a preference for the null.</p><p>When no candidate survives, we still say something useful: we report the series as "stationary" (best no-change model is a constant) or "non-stationary" (best model has a slope), with the trend direction. For an agent, "this series is cleanly trending up over the window" is itself a finding.</p><h3>Detecting distribution changes with a dispersion channel</h3><p>A variance change is indirectly visible to the mean channel – worse, the robust weighting there actively mutes the excursions that signal it. So we detect it on a separate dispersion channel and reuse the exact same structural detector, because on this channel a variance change is just an ordinary level change.</p><p>The channel is built from one sample per non-overlapping window. Within a window, we take the <a href="https://en.wikipedia.org/wiki/Interquartile_range">inter-quartile range</a> of the first differences, rescaled to a standard-deviation equivalent (), and pass it through :</p><p>Then . Three choices matter here. First-differencing cancels level and slope, so a mean step contributes a single large difference rather than inflating the whole window, and a steady ramp produces a flat channel. Non-overlapping windows keep the samples independent; overlapping windows <a href="https://en.wikipedia.org/wiki/Autocorrelation">autocorrelates</a> the channel and makes the segmenter over-detect. And the IQR is used rather than the median (which is too robust and will miss a window that is 40% noisy then flatlines) or the raw standard deviation (which is not robust enough since one spike's two large differences inflate the window). Because the dispersion channel is a fraction of the original length, the verifier there is restricted to a lower-order null so a genuine low-high-low variance bump is not absorbed.</p><p>The functional form  is worth dwelling on, because each part earns its place. The log makes the channel respond to ratios of noise level rather than absolute differences. Variance changes in telemetry are typically multiplicative: a regime is "twice as noisy". On a raw-scale channel, a doubling shows up as an enormous absolute jump at a high baseline and a negligible one at a low baseline, so an additive step-cost detector would find variance changes trivially in loud series and miss them in quiet ones. Under a log, a factor- change in scale is the same offset  wherever it occurs, which is exactly the additive-step behavior the structural detector is built for. The " is a soft floor. A bare  diverges to  as the scale goes to zero, which is precisely what happens on a near constant stretch, and would manufacture a huge spurious step at the first noisy window after it. Conversely,  is finite and smooth at zero, behaves linearly () while the noise is small, and recovers the multiplicative  behavior once the noise is appreciable. This gives graceful degradation instead of a singularity, and with no tuned epsilon to pick. Note that squaring the scale (using a variance instead) only doubles the dynamic range; it makes no difference to the detector either way, since  differs only by a constant the threshold absorbs.</p><h3>Detecting point anomalies as excursions from a local baseline</h3><p>Spikes and dips are detected as point excursions from the local rolling-median baseline. Working from the local residual rather than raw values means level structure is removed and smooth curvature is tracked; even for time series that change significantly the detector is sensitive to significant local deviations.</p><p>The pipeline is a generous proposer followed by a strict gate:</p><ol><li><p>Propose every point whose residual exceeds a threshold number of robust sigmas. The scale is the larger of the global first-difference noise (which stays meaningful on smooth data where most residuals are exactly zero) and a composite of robust scales of the residuals (which inflates once a frequent large-residual population appears).</p></li><li><p>Merge adjacent same-sign candidates into excursions, dropping any that span a full minimum segment: that is a regime, and is owned by the structural channels.</p></li><li><p>Rank and cap the excursions by peak <a href="https://en.wikipedia.org/wiki/Standard_score">z-score</a>, keeping the top , so a pathological series cannot drown the output.</p></li><li><p>Gate using one shared null: build a <a href="https://en.wikipedia.org/wiki/Kernel_density_estimation">Gaussian KDE</a> from the series with all of the retained excursions removed, and keep an excursion only if its peak's Bonferroni-corrected upper-/lower-tail probability under that null clears the threshold.</p></li></ol><p>Removing all the tested excursions from the single null at once is a key trick. The leave-one-out alternative – score each excursion against a null containing the others – lets the largest spike and dip mask everything else. Removing them together means several genuinely distinct excursions are each judged against the remainder and all survive, while a recurring population is still rejected.</p><p>The proposer and the gate ask deliberately different questions, and that distinction drives two further choices. The proposer works on residuals from the rolling median since it wants recall, and a residual is what tells you a point stands out from its local neighborhood. The gate is value-based: it asks "is this magnitude one we see at other times in the series?", so a spike to a level that recurs elsewhere — such as periodic batch jobs — is suppressed even though it is a large local residual. Those are the right semantics for an agent, but they expose a heteroscedasticity problem, because telemetry noise is almost always a function of magnitude. Periodic spikes can sit orders of magnitude above the background, and a single KDE bandwidth fitted to the whole value range is then far too narrow up in the high tail. So ordinary large values come back as significant, a steady source of false positives.</p><p>The fix is a <a href="https://en.wikipedia.org/wiki/Variance-stabilizing_transformation">variance-stabilizing transform</a>. We run the value gate in  space, where  is a robust measure of the spread of the background.  is linear for  and logarithmic for , which turns a multiplicative (magnitude-dependent) spread into a roughly constant one, so a single bandwidth is valid across orders of magnitude. It is also odd and finite at zero, so exact zeros and sign changes (dips below a small baseline) need no special handling, unlike a bare log. Crucially, it is monotone and so does not change what is tested (for any monotone function , ) so it only fixes the estimate of that tail probability.</p><p>One subtlety closes the loop. The KDE null and the kernel bandwidth are taken from different scales, on purpose. The null is the stabilized background values: so it models any mode in the data, which is what makes a recurring large magnitude unsurprising. However, the bandwidth is taken from the stabilized residuals, not the stabilized values, because a genuine level change makes the value distribution bimodal, and a bandwidth computed from that bimodal spread would balloon, masking a real spike sitting on top of a shifted regime. The residual removes the step, so the bandwidth always reflects within-regime noise and the gate stays sensitive to a deviation that is extreme relative to its own neighborhood if it is also outside the envelope for the series as a whole.</p><h3>Merging structural, distribution, and point anomaly events</h3><p>Finally, an orchestration layer merges the structural, distribution, and point-anomaly event streams. Structural and distribution events that mark the same regime boundary are de-duplicated to the more significant one (a boundary that shifts both level and spread is one event, not two). Pulses are a separate stream added after de-duplication, because a spike that lands on a structural boundary is a real, separate finding and must not be suppressed. Everything is then mapped back from the internal value-array index space to source-bucket indices.</p><h2>Handling extreme magnitudes and edge cases</h2><p>What makes this usable as an unattended tool is a collection of defensive choices for the cases that break naive implementations:</p><ul><li><p>Variance computed as  loses all precision at large magnitudes: a constant series at  can manufacture phantom change points purely from floating-point error. We center every PELT input by a constant offset first; the polynomial RSS is invariant to that shift in exact arithmetic, but the working magnitudes drop from  to .</p></li><li><p>Using raw indices as the regressor results in poor condition polynomial fits: the largest moment is , which is about  for a cubic over a 2000-point window. This trips the SVD singularity guard and silently degrades the fit. Mapping  affinely onto  leaves the fit identical (RSS is invariant under reparametrization) but every moment becomes .</p></li><li><p>Scale-free cost, as described, means the segmentation cost doesn't depend on tuning a noise estimate.</p></li><li><p>Down-weighting wants a primarily <em>local</em> scale: suppress whatever is anomalous in its own neighbourhood. It uses the maximum of MAD and a small global floor on the differences from the rolling median. Conversely, spike/dip detection wants a <em>global</em> scale: we care about global outliers. It uses a composite of robust scales of all differences. Using the wrong one in either place produces characteristic failures – irrelevant spikes in a quiet segment, or locally large excursions creating spurious breaks – and maintaining separate channels allows us to pick appropriately.</p></li><li><p>Using one p-value threshold, <a href="https://en.wikipedia.org/wiki/Bonferroni_correction">Bonferroni-corrected</a> by the number of candidates, applied consistently across all three detectors, means that "how surprised should I be" is consistent everywhere.</p></li></ul><p>The recurring theme is that the difference between a detector that works in a notebook and one that works on a firehose of production time series is mainly in handling the edge cases gracefully.</p><h2>Performance: ~1ms per series on a single core</h2><p>Detection cost is dominated by PELT. Its segment cost is a profiled-variance linear fit, which we evaluate in constant time from prefix-summed weighted moments rather than maintaining a regression per candidate boundary, so a single segment cost is a handful of array reads and a 2×2 solve. Cost grows a little faster than linearly with series length since PELT's pruned candidate set does not stay constant on noisy data. Therefore, to bound the worst case on very long series, we downsample ahead of detection: above a cap (2000 samples), the series is collapsed into macro-buckets, keeping two samples per bucket: the median and the largest local deviation. This is inspired by the <a href="https://www.vldb.org/pvldb/vol7/p797-jugel.pdf">M4 downsampling scheme</a>, but because we need only the median and the largest excursion for structural-change and outlier detection, respectively, we can then afford to double the bucket resolution. The downsampled series carries its original bucket indices, so every reported event still maps back to a real source bucket; below the cap it is a no-op. The whole analysis is a single pass over the (possibly downsampled) series with no per-series configuration. This lets the agent call it freely across many signals.</p><p>In absolute terms, this means the analysis is comfortably interactive. On a single core, post-warmup, a typical series of 140–350 buckets is analysed in about 1 ms (≈220,000 buckets/s), and a series long enough to hit the downsample cap (say 5,000 buckets, collapsed to 2,000) takes about 40 ms, which is the effective worst case per call. For an agent issuing a handful of these calls per investigation, and parallelising across the many series in a `BY` query, the latency is negligible.</p><h2>Evaluation on synthetic and production telemetry</h2><h3>Synthetic benchmark</h3><p>We evaluate first on a synthetic generator with known ground truth, because it lets us measure the things that matter precisely. The generator produces random time series with diverse behaviors, which we group into three families. The <strong>positive</strong> family injects known events of each type: clean and noisy step changes (up and down, SNR around 10, at several positions), trend onsets and ramps (including a flat–ramp–flat sequence with two boundaries), variance changes with a constant mean (single steps and a low–high–low bump), and isolated spikes and dips. The <strong>null</strong> family ideally produces no event: stationary noise, perfectly constant series, smooth quadratic drift and clean ramps, and periodic signals. (A variance change with a constant mean is not null: it is a distribution change, an abrupt step on the dispersion channel, and so it belongs in the positive family above. The related null requirement, that such a change does not surface on the value channel as a step, is checked separately.) The <strong>adversarial</strong> family stresses the robustness machinery: a perfectly flat series at a magnitude of  (for which a naive variance arithmetic manufactures phantom breaks here through catastrophic cancellation) and, conversely, a genuine spike or step in a noisy baseline as high as  (which must still be found and located, the high baseline notwithstanding); a spike on top of a step; a within-regime spike after a 100 level jump; a recurring train of equal peaks (a population, not individual spikes); wide sustained excursions (a structural change, not a spike or dip); and fuzzed random level-shift series. On these series we track recall per event type, precision and the false-positive rate on the null family, localization error, parsimony (events per series and adherence to the count limit), and invariance under constant offsets, rescaling and extreme magnitudes.</p><p>The table below summarizes what the suite tests.</p><p>Event family</p><p>Representative scenarios</p><p>Required outcome</p><p>Localization tolerance</p><p>Step</p><p>clean / noisy, up / down, single and multiple, several positions</p><p>detected</p><p>≤ 4–8 buckets</p><p>Trend</p><p>slope change and flat–ramp–flat, clean and in noise</p><p>detected</p><p>≤ 8–12 buckets</p><p>Distribution</p><p>variance step (mean constant), low–high–low bump</p><p>detected</p><p>≤ 1 dispersion window</p><p>Spike / dip</p><p>isolated, multiple distinct, on heavy-tailed and high-magnitude series, within-regime after a step</p><p>detected, capped at max(5, 2% of n)</p><p>≤ 2 buckets</p><p>Null series</p><p>stationary noise, constant, smooth drift / ramp, periodic</p><p>no event reported</p><p>n/a</p><p>Invariance / robustness</p><p>constant offset, rescaling, 10^5–10^9 magnitudes, spike-on-step, recurring population, wide excursion</p><p>result unchanged / no spurious event</p><p>n/a</p><p>Running this over 400 series per scenario, just over half a million buckets in total, and scoring with the same two categories we use for the real-data evaluation below (any regime change versus point spikes and dips) gives:</p><p>Event type</p><p>Recall</p><p>Precision</p><p>Median localization error</p><p>Structural change</p><p>0.994</p><p>0.664</p><p>0</p><p>Spike / dip</p><p>0.847</p><p>0.746</p><p>0</p><p>Two things stand out. When an event is detected, it is placed essentially exactly: the median localization error is zero buckets for both categories; and the point-wise accuracy, 0.998, is directly comparable to the 0.995 we report on real telemetry below: the overwhelming majority of buckets are correctly left unmarked.</p><p>The precision figures are lower than on real data, and understandably so. A sixth of this population is a <em>hostile</em> null family (periodic signals, smooth drift, clean ramps) chosen precisely because they tempt a detector into a spurious break, and every false alarm on them counts against precision. The resulting false-positive rate is 0.5% per bucket, with 15% of null series carrying at least one spurious event. We treat this as the conservative end of the range: on the real-telemetry mix below, where the null series are less adversarial, precision rises to 0.85 (structural) and 0.90 (spikes/dips). Finally, offset and scale invariance holds on all 400 series: the same events, to within a few buckets, whether the series is shifted by a constant or rescaled by up to three orders of magnitude.</p><p>That raw precision also misses how the result is consumed. The agent reads events most-significant-first, so a false positive only does harm if it outranks a genuine one, and by and large it does not. The median p-value of a true positive is about , against about  for a false positive: the genuine events are typically overwhelmingly more significant. Concretely, if we keep only the top  events per series by significance (with  the true count) precision rises to 0.94, and a randomly chosen true positive is more significant than a randomly chosen false positive 93% of the time. It is not a perfectly clean separation: a strong periodicity or a sharp curve genuinely can contain a significant-looking break, which is why that figure is 0.94 rather than 1. However, the ranking is reliable enough that an agent reading from the top, or applying a stricter significance cut-off, sees the real events first and the false alarms as a lower-significance tail. This is also why exposing the detector's significance to the agent (see <a href="https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql#whats-next-for-es|ql-time-series-analysis">What's next</a>) matters more than squeezing the raw precision higher.</p><h3>Real cloud telemetry</h3><p>To evaluate on real data, we scraped around 300 metrics from our production cloud environment. These cover HTTP status-code counts, failed memory allocations, memory usage, network usage, page faults, CPU usage, and throttling metrics, measured both per instance and aggregated across the fleet as a whole. Their values range over more than 12 orders of magnitude, and they display a variety of behaviors including ramps, periodicity, step changes, distribution changes, and trend changes. The figure below shows a sample of series together with the detections we make on them.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte0bc6fbc138cca1c/6a6a33ee065b162280702003/2c30fd3e2f97e32b5b47bf66a923f73b08ec015d-2510x1284.png" alt="Grid of 12 production cloud telemetry series with detected structural changes, distribution changes and anomalies marked by Elasticsearch's change point detector" /><p>To get a sense of the accuracy on this set data, we labeled a subset of 150 most interesting time series, marking the visually clearest features in each. This labeling is not necessarily optimized for our target use case, where false negatives are typically more problematic than false positives: an agent consumes these results as part of a broader investigation and can pull additional information to corroborate them. Even so, we find excellent agreement with human judgment on these series. Since each series comprises between 140 and 350 points, the point-wise accuracy, at 0.995, is extremely high: the great majority of points are correctly identified as neither a change, a spike, nor a dip. But the more telling metrics are recall and precision on the human-labeled points, shown in the table below. The human labels did not attempt to categorize each change, so we break the results down only into "structural changes" and "spikes / dips" — the same two categories, and the same point-wise accuracy and recall/precision metrics, as the synthetic benchmark above.</p><p>Event type</p><p>Recall</p><p>Precision</p><p>Structural change</p><p>0.94</p><p>0.89</p><p>Spike / dip</p><p>0.97</p><p>0.92</p><p>It is worth covering exactly why we get disagreements. These largely fall into three categories: spikes and dips in context, isolated breaks, and small-magnitude breaks in stable series. We deliberately do not try to detect spikes and dips that are unusual only in their immediate context; that is, not globally unusual but visually significant relative to an inferred periodicity in the data, for example. Trying to account for these without fully modeling the seasonality in the data hurt precision more than it helped recall, and we have a separate persistent anomaly-detection process that builds more complete models of baseline behavior over time. Isolated breaks are an artifact we decided to live with: PELT's cost function tends to isolate a change point with a few values intermediate between the two regimes, because absorbing it into either neighboring span inflates that span's cost. Humans are good at judging such situations visually and assign a single change point. Finally, small-magnitude changes are simply not visually obvious. We detect them deliberately and regard this as a strength of a quantitative approach, since they are often the early precursors of an incident whose later, larger effects drown them out.</p><h2>How the agent uses change point results in ES|QL</h2><p>To the agent, all of this is one tool call: it points it at a collection of time series and gets back a typed, located, ranked list of events. That list is small by construction, which means it drops cleanly into the model's context without crowding out everything else it is reasoning about. Because the result contains multiple events, a single call over a window can hand the agent the whole local story – "error spike at 02:14, latency regime change at 02:30, throughput dip at 02:31" – and let it correlate across signals to a root cause.</p><p>Operationally, we expose this through both the ES|QL <code>CHANGE_POINT</code> <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/change-point">command</a> and the <code>change_point</code> <a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-change-point-aggregation">aggregation</a>. Currently, we have not extended the <code>change_point</code> aggregation to return multiple change points since it breaks backwards compatibility of the output schema. It just returns the most significant event. We don't have the same restriction for ES|QL since it returns change points annotated onto the table rows to which they apply. We do plan to revisit the output schema for both ES|QL and the aggregation in a later version. We'd like to migrate to optionally returning significance in log-space, which doesn't underflow, and including a short verbal description of each change, which we expect to help agents when seeing just the change points themselves.</p><p>ES|QL is Elasticsearch's piped query language, and <code>CHANGE_POINT</code> runs the detector as one stage in a pipeline. Its <code>BY</code> clause enables it to analyze many series at once (one per group) so the agent can, in a single query, segment every service's latency or every host's error rate side by side rather than issuing a call per series. The actual leverage, compared to the <code>change_point</code> aggregation, is composability: the events come back as ordinary rows in the pipeline, so the agent then has the entire ES|QL language to manipulate them downstream. It can filter to a window, join change points against deploy markers, count events per service, rank by significance, feed the survivors into a further aggregation, and so on.</p><p>For example, suppose the agent wants to find out whether any servers have recently seen a sudden CPU spike or a prolonged step change in CPU usage over the last 12 hours, and whether that might point to a load-balancing issue. It could use the following query:</p><p>Here it's using a <code>STATS ... BY host.pod</code> to see how the detected events cluster across other dimensions of the data, such as the Kubernetes pod, and so judge whether they share a common cause.</p><h2>What's next for ES|QL time series analysis</h2><p>As far as detecting events of interest in time series, the foundation is in place: a single, robust, parsimonious tool that turns a raw telemetry series into the short list of events that actually matter, which is exactly the kind of reach an agentic SRE needs. Going forward, we plan to explore the best mechanism for feeding the detector's uncertainty to the agent, so that a borderline event can be flagged as "worth a second look" rather than silently included or dropped. Also, this is the first of several analytical tools we plan to build into the ES|QL query language to enable agents to triage and RCA issues more effectively; so stay tuned for further updates.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[ML Research]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Thomas Veasey]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt097721c03648a84e/6a6a33ef0a222b4c70877f32/8f6b95800c65fe389d3e8d8281e8e8dc351f734d-992x342.png" length="0" type="image/png"/>
    <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[AI shopping agents: Why context comes before the query]]></title>
    <description><![CDATA[AI shopping agents that guess at your vocabulary make expensive mistakes. Pre-computed catalog context stops the guessing before the first tool call.]]></description>
    <content:encoded><![CDATA[<p>The race is on for retailers to match the evolving expectations of their customers to provide significantly richer online shopping experiences. Customers want to go beyond searching for products; they want interactive, personalized, and proactive guidance powered by AI. The challenge is that, although large language models (LLMs) can be extremely powerful, how do you construct a system that acts like a knowledgeable employee of your store in a fast, accurate, and cost-effective way? We’ll discuss the challenges of building these AI shopping assistants as well as emerging context engineering approaches to optimize how they work.</p><p>AI shopping agents fail not because the model is wrong but because the agent arrives at every conversation not knowing your catalog, vocabulary, or business rules. It has to discover all of this context through tool calls, and that discovery is the cost. Precomputing a structured context layer from signals you already hold (vocabulary, policies, user profiles, session behavior) cuts the exploratory work the agent does before it can answer and makes its behavior governed and predictable.</p><p>In the analogous document-retrieval case, precomputing context reduced input tokens by up to 75% on a controlled benchmark; we expect comparable savings in ecommerce because the exploration pattern is the same, though the exact figure will vary by catalog and query mix. The signals are richer in ecommerce than in almost any other domain, and most retailers already have them. The question is whether or not these signals are assembled in a form that the agent can use before it starts reasoning. Elastic’s broad mix of search capabilities, semantic, hybrid, keyword, filtering, and aggregations make it a compelling choice for not only building your core retrieval tools but also as the vital context engine.</p><h2><strong>Why AI shopping agents fail in production</strong></h2><p>Retailers investing in AI shopping assistants are discovering an uncomfortable gap between what the demos promise and what the first few months in production deliver.</p><p>The assistant takes four seconds to respond. It confidently recommends a product in a size range that doesn't exist for that item. It tells a returning customer about a coat style they bought 18 months ago and returned. It filters by a category name that doesn't match the internal taxonomy and returns zero results. The customer gives up and abandons the chat altogether.</p><p>These aren't model failures. The frontier models powering these agents are capable of extraordinary reasoning when they have the right information in front of them. The problem is that the agent arrives at the conversation knowing nothing about the retailer's catalog, customer, or business rules that govern what should and shouldn't be recommended. It has to learn all of this through the conversation itself, making exploratory tool calls to discover what departments exist, what filter values are valid, and what the brand's policies are on certain product types. Every one of those discovery calls costs latency and tokens before the agent has said a single useful thing to the customer.</p><p>Latency matters in ecommerce in a way it doesn't in many other contexts. Shoppers expect response times measured in seconds, not the minutes that enterprise knowledge-base agents routinely take. It’s well established that slower responses reduce engagement and conversion in online retail. An AI shopping assistant that thinks visibly for five seconds before answering a question about gift ideas isn't a feature; it's friction.</p><p>The fix isn't a faster model or a bigger context window. <strong>The agent's latency and cost problem is a context problem.</strong> This can be solved by carefully computing context before the agent call, not during it.</p><h3><strong>What an AI shopping agent knows before it searches</strong></h3><p>Imagine you’re the agent. You have a search tool connected to the catalog, and this query arrives:</p><p><code>"an outfit for an autumn wedding"</code></p><p>What do you actually do with that?</p><p>Start with what you don't know. An outfit for a man or a woman? Is "autumn" a color, a season, a style of fabric, or just when the wedding happens? And who is this shopper, someone who has bought from you for years, or a stranger? Do they buy expensive designer brands or do they always hunt out a bargain on sale? You have none of these answers. So you do what an agent does when it's working blind: You ask loads of follow-up questions, guess, or fire off a series of exploratory searches to find out what departments and filter values even exist, watching the seconds tick by before you've offered the customer anything at all.</p><p>Hold onto that feeling of working blind. The rest of this article is about what changes when the agent is handed the answers first.</p><h2><strong>Why ecommerce is different from general RAG</strong></h2><p>Most of the published work on reducing agent costs focuses on document retrieval: question-answering over corpora of articles, reports, or knowledge-base entries. A recent experiment from the Elastic team (<a href="https://www.elastic.co/search-labs/blog/pre-computed-context-llm-agent-costs">Cutting agent costs with pre-computed context</a>) demonstrated that pre-extracting structured facts from documents before the agent call reduced input token consumption by up to 75% and improved answer accuracy from 60% to 92% on a hard factual benchmark. That improvement came in stages, with the largest jump driven by feeding the agent's own wrong answers back into the extraction step rather than by precomputing context alone, which is a distinction we'll return to when we discuss governance.</p><p>Ecommerce applies the same principle to a fundamentally different structure. A product catalog isn't a document corpus. It's a highly structured index of items with strict field semantics, a domain-specific vocabulary of brand names, color codes, and category hierarchies, and a layer of business rules that override pure relevance in specific situations.</p><p>The failure modes that result are distinct from document retrieval augmented generation (RAG) failures:</p><ul><li><p><strong>Vocabulary mismatch:</strong> A customer asks for a "navy jumper." The agent constructs a filter against a field where the canonical value is NAVY and the category is stored as Knitwear &amp; Jumpers. Without a vocabulary mapping, the agent either guesses colors and categories and gets it wrong, or makes multiple exploratory calls to discover what values exist before it can filter correctly.</p></li><li><p><strong>Hallucinated filter values:</strong> Without knowing which filter dimensions are valid for a given query, agents can construct queries against fields that don't exist or with values that return zero results. A filter like category: knitwear looks reasonable; <code>masterCategoryNames: "Knitwear &amp; Jumpers"</code> is what the index actually contains. If the agent doesn’t know, it either hallucinates or has to do a separate tool call to find out, causing another LLM loop, which costs time and tokens.</p></li><li><p><strong>Context-free personalization:</strong> The same query from two different customers, one who typically shops in the premium range and dresses for formal occasions, and one who buys primarily casualwear under £40, should return different results. Without profile context, the agent treats every query identically, which is worse than a well-tuned keyword search because it creates the impression of a personal assistant while delivering generic answers.</p></li></ul><h2><strong>The context layer: What signals it needs</strong></h2><p>The reason ecommerce is particularly well suited to precomputed context is that retailers already hold an unusually rich set of signals. The challenge isn't data availability; it's assembly.</p><p>The signals fall into two groups. Two of them, the catalog vocabulary and the business policies, are the genuinely original work and the heart of this approach. The rest, live facet state, user profiles, and session history, are valuable but closer to table stakes, signals that most teams already understand how to fetch. Here's what most mid-to-large retailers have and what each signal prevents.</p><p>Signal</p><p>What it prevents</p><p>Effort to build</p><p>Catalog vocabulary</p><p>Vocabulary mismatch and hallucinated filter values; the agent guessing at colors, categories, or brand names instead of resolving them to canonical field values</p><p>One-time engineering effort (full-catalog aggregation); incremental to maintain as new categories and brands are added</p><p>Business policies</p><p>Recommendations that ignore legal or trading requirements, for example, missing age verification on alcohol queries or routing that misses a gluten-free range</p><p>Human-authored and governed, not automated; ongoing review as new policy types are added</p><p>Live facet state</p><p>Recommending filters that return zero results or out-of-stock options for the current query</p><p>Runs in parallel with vocabulary and policy lookups; leans on existing catalog and retrieval infrastructure</p><p>User profile</p><p>Making a returning customer restate sizes, budget, or brand preferences they've already given</p><p>Fastest signal to retrieve, a single document lookup by user ID</p><p>Session and purchase history</p><p>Re-recommending an item the customer already dismissed or bought and returned</p><p>Most aspirational layer; depends on customer relationship management (CRM) and analytics integration, best added once the agent is already live</p><p>The reason we say <em>semantic metadata</em> and not just <em>metadata</em> is that we’re trying to match the semantic (meaning) of the intent rather than the exact words. If a user is searching for “teal,” we should be able to understand that this is a color and which colors exist in our products that are semantically similar to teal, even if none of them are actually teal. So, if we search for “teal,” we might want to return:</p><p><code>ProductColours = “aquamarine, turquoise”</code></p><p>Hopefully, you can see how this semantic metadata is bridging the gap between user intent and the agent’s knowledge of the products.</p><h3><strong>Catalog vocabulary: The foundation of context engineering</strong></h3><p>Catalog vocabulary is the layer that does the most work, and it's the one most worth getting right first.</p><p>A vocabulary index maps natural language to the exact field values and category paths used in the product index. It answers questions like: <em>What does "navy" map to?</em> <em>Which categories fall under "knitwear"?</em> <em>Is "Autograph" a brand or a range?</em> <em>What's the correct spelling of a competitor brand the agent might encounter in a query?</em></p><p>What makes this more than a synonym list is how it's queried. The interesting thing a shopper says is rarely an exact field value. They say "something cozy for fall," not <code>colour: NAVY</code> and <code>masterCategoryNames: "Knitwear &amp; Jumpers"</code>. So the vocabulary index needs to resolve fuzzy, natural language intent into precise, exact-match filters, and that requires both kinds of matching at once: semantic search to understand that "cozy" leans toward knitwear and fleece, and exact keyword matching to pin the result to the canonical values the product index actually stores. A metadata index that supports both on the same documents is, in effect, a translation layer between how customers talk and how the catalog is structured.</p><p>This is also where the index earns the description "semantic metadata layer" rather than "lookup table." Each entry is a small natural language description of a facet value or schema concept, so the agent can match against meaning and then read back the exact filter to use. For a typical fashion retailer, this covers hundreds of color values, brand aliases, category synonyms, and size-range conventions. Building the semantic metadata layer from a full-catalog aggregation is a one-time engineering effort; maintaining it is incremental as new categories and brands are added. Without it, an agent encountering an unfamiliar term must either guess or make exploratory tool calls to discover what's there.</p><h3><strong>Business policies in the context layer</strong></h3><p>The second original layer is policy. Some queries carry implicit business requirements that pure relevance cannot handle. A query for "wine gift for a friend" should trigger an age-verification reminder in markets where it's legally required. A query for "gluten-free food gift" should route away from general confectionery toward the specific gluten-free range. A query mentioning "wedding guest outfit" in spring should apply different weighting than the same query in November.</p><p>These are policies, and most retail search teams already write them. They just call them boost rules, merchandising overlays, or synonym configurations. The difference in an agentic context is that instead of being applied silently as query modifications, they're surfaced as readable hints the agent can use when deciding how to frame its answer and which products to surface. The agent doesn't have to infer your trading rules from the catalog; it's handed them.</p><p>The critical point: These policies encode business intent, not just relevance. A policy that routes alcohol queries through an age-appropriate flow isn't a retrieval optimization; it's a trading requirement. That's why this layer must be human-authored and governed, not generated automatically from traffic patterns, a point we return to in the governance section.</p><p>Together, vocabulary and policy are what make the agent behave like it understands your business rather than just your data. The remaining three signals sharpen the experience, but they're more familiar engineering.</p><h3><strong>Facet, profile, and session signals in the context layer</strong></h3><ul><li><p><strong>Live facet state:</strong> Before the agent recommends filters, it should know which filters are available and how many results each returns for this specific query. An agent that suggests "filter by size 8" without knowing that size 8 is out of stock for this query undermines the customer's trust immediately. A facet state query against the product index, run in parallel with the vocabulary and policy lookups, returns the counts, ranges, and available values specific to the current query.</p></li><li><p><strong>User profile:</strong> A persistent profile (sizes, color preferences, budget range, brand affinities) lets a returning customer skip restating what they've already told you. It's typically the fastest signal to retrieve, a single document lookup by user ID.</p></li><li><p><strong>Session and purchase history:</strong> Within a session, the agent should know what the customer has already seen, dismissed, or added to their basket, so it doesn't re-recommend a dismissed item or repeat itself. Longer-term purchase history extends this, and signals like return history are richer still, but using them well depends on data most retailers hold in systems that aren't yet wired into their search path. This is the most aspirational layer and the one best approached last, once the earlier layers are delivering value. There must also be balance when building the initial context not to overinflate the size, which will slow down the first reply and increase token costs. There is, therefore, a careful balance to strike between providing information like purchase history in the initial context or providing it as a tool for the agent to use during conversation, but users will expect that, if they’re logged in, the agent should know what they’ve purchased. The precise optimal context is likely to be specific to each implementation and customer experience and will require careful testing.</p></li></ul><h2><strong>How context engineering works before the LLM call</strong></h2><p>The pattern that makes this work is simple to describe and moderately involved to implement:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt017dd46a26cf7bc6/6a6119f6f81792b87a07f605/e814160fc7e599aa3cd1603be366cc726114948f-1024x559.png" alt="Diagram of context engineering for AI shopping agents: a context layer resolves vocabulary, policy, facets, profile and session signals in parallel before the LLM call" /><p>Without this context build, the agent makes the same discoveries on its own, but through LLM-driven tool calls that each cost a full inference round trip. As a rough rule of thumb, an exploratory tool (for example, GetFilterValues) call tends to land somewhere in the region of 600ms to 1 second of latency in practice; an agent that discovers the vocabulary, checks policy hints, and retrieves facet state through three separate tool calls before it even begins answering can, therefore, add two to three seconds to the response time, and that's before it performs the actual product search. These are order-of-magnitude estimates, not benchmarked figures, and the real numbers depend heavily on the model, the network path, and how the tools are implemented.</p><p>Replacing those discovery calls with a parallel fetch (multiple context-building queries can be done in parallel) that runs before the LLM is invoked removes most of that cost. The context build takes roughly the same wall time as a single LLM tool call, but it replaces three or four of them. The LLM would either need to repeat failed searches or do its own context-building tool calls sequentially to get enough context to be successful. Precomputing context is also deterministic rather than subject to the model's tool-selection choices, which gives the business more control to fine-tune the experience.</p><p>The token reduction follows the same logic: Each exploratory tool call returns raw data the model must process. A preassembled context summary replaces that raw data with structured facts the model can consume in one pass. At the scale of usage possible in public-facing retail websites, this token cost saving can be significant. This work on document search (<a href="https://www.elastic.co/search-labs/blog/pre-computed-context-llm-agent-costs">Cutting agent costs with pre-computed context</a>) showed up to 75% input token reduction on a controlled benchmark. That benchmark was document retrieval rather than ecommerce, and its authors are explicit that the multiplier isn't a fixed number you can expect everywhere. We expect the direction to hold in ecommerce, because the exploration pattern is the same; it just runs against a structured catalog rather than a document corpus. The magnitude is something each team should measure against its own traffic.</p><h3><strong>The AI shopping agent with full context</strong></h3><p>Remember the query that left you guessing: an outfit for an autumn wedding. Run it again, but this time, before you have to think, you’re handed a precomputed context as a short brief:</p><ul><li><p>This shopper’s name is Sarah, female, age 32, and she buys womenswear, size 12.</p></li><li><p>She typically buys your mid-tier ranges.</p></li><li><p>"Autumn" here matches these specific color labels: “rust”, “burgundy”, “forest green”, “camel”.</p></li><li><p>Matching departments: “Womenswear”, “Menswear”.</p></li><li><p>Matching tags: "occasion dresses”, “trouser suits”, “wedding”.</p></li><li><p>She already has a burgundy bag in her basket.</p></li><li><p>House rule for wedding-guest looks: Complete the outfit. Show hats and accessories, not just the dress.</p></li></ul><p>Suddenly, you’re not guessing; you’re styling for this customer. And the interesting part is how those facts combine rather than just stack. The season proposes a whole autumn palette; the burgundy bag already in her basket narrows that palette to the few tones that coordinate with it; the house rule tells you to finish the look with a matching fascinator rather than stopping at the dress. Together, these facts let you answer like someone who knows both this customer and this shop, in a single pass, with nothing invented.</p><p>That short brief is exactly what the signal stack produces: the shopper's profile, the resolved vocabulary, the live basket, and the business policy. These are assembled in parallel and placed in front of the agent before its first move, so the "<em>What do I even do with this?</em>" problem never has to be solved one expensive tool call at a time.</p><h2><strong>Governing the context layer without automated drift</strong></h2><p>One difference between the approach described here and automated knowledge extraction systems is worth addressing directly: In ecommerce, the context index cannot self-update without human review.</p><p>The policies that govern how an agent responds to gift queries, alcohol queries, or queries from customers in certain age brackets aren't just relevance configurations; they're trading decisions with potential legal and brand implications. An automated system that generates new policies from traffic patterns, without review, is a compliance risk before it's a technical asset.</p><p>This is actually the right constraint for most retail organizations, and it aligns with how search teams already work. Merchandisers write boost rules. Search teams maintain synonym configurations. Content teams approve what language appears in automated recommendations. The context policy layer is the same kind of governed configuration; it just serves a different consumer, namely, the agent's reasoning step rather than the query pipeline.</p><p>It's worth noting where this differs from the document-retrieval work referenced earlier. In that experiment, the biggest accuracy gain came from an automated feedback loop that fed the agent's wrong answers straight back into the extractor. That works well for factual question-answering, where "right" and "wrong" are unambiguous. In ecommerce, the equivalent signals still surface automatically, but a human decides what to do with them, because the changes carry trading and compliance weight. The loop is the same shape; the publication step has a person in it.</p><p>The governance loop that works in practice has two tiers:</p><ul><li><p><strong>Automatic signal surfacing:</strong> Zero-result queries, repeated reformulations on the same topic, and sessions that end without a purchase after an agent interaction are all signals that something in the context layer is missing or wrong. These surface automatically as candidates for improving the experience, for example: a vocabulary term that didn't resolve or a policy that didn't fire on a query type it should have covered. To do this, you need a thorough log of conversations, including the reasoning and tool call trace in a platform like Elastic. This allows you to analyze the performance of the agent using both structured tools, for example, percentage increase in thumbs-down conversations and semantically. You could also run an automated review of conversations about "gifts" around December to characterize the thumbs-up/down ratio across an AB test of two agents.</p></li><li><p><strong>Human authorship and review:</strong> The search or merchandising team reviews candidates and authors the appropriate vocabulary entry or policy. Policies go through approval before publication. This typically mirrors the workflow that already exists for synonym changes or boost rule modifications; the tooling is the only new part.</p></li></ul><h2><strong>How to implement context engineering in phases</strong></h2><ul><li><p><strong>Phase 1: Vocabulary layer</strong> (highest leverage, bounded engineering task).</p></li><li><p><strong>Phase 2: Facet state and initial policies</strong> (leans on the same catalog and retrieval primitives).</p></li><li><p><strong>Phase 3: User profiles and session signals</strong> (requires CRM and analytics integration; best added when the agent is active).</p></li><li><p><strong>Phase 4: Governed feedback loop</strong> (shifts to organizational alignment; surfaces gaps for merchandising teams).</p></li></ul><p>The full signal stack described above doesn't need to be built at once, and the order isn’t arbitrary. The highest-leverage starting point is also the lowest in implementation complexity: the vocabulary layer. A semantic metadata index built from a full-catalog aggregation (canonical color values, brand aliases, category paths, field names) is a bounded engineering task, and an agent that can resolve "navy jumper" to color: NAVY, masterCategoryNames: "Knitwear &amp; Jumpers" before its first tool call is materially better than one that discovers this through trial and error. If you build nothing else, build this.</p><p>Facet state and the first set of policies can follow close behind, often in parallel, because they lean on the same catalog and the same retrieval primitives. The later layers, such as user profiles, session signals, and the governed feedback loop, are where the work shifts from search engineering to organizational alignment. To implement CRM system integration, merchandising workflow changes, and the analytics needed to surface gaps can take a significant amount of work. Those layers are more valuable once the agent is already in regular use and generating the traffic signals that make the governed loop worth running. The important feature is that each layer stands on its own, so a retailer gets real value from phase one without committing to phase six.</p><h2><strong>What infrastructure does a context layer need?</strong></h2><p>Precomputing context at the depth described here places specific requirements on the underlying platform. It's worth being explicit about these, because the temptation in early agent builds is to reach for the simplest available tool for each capability.</p><ul><li><p><strong>Semantic search</strong> to match natural language queries against the vocabulary index and surface the right canonical values. Fuzzy keyword matching alone won't resolve ambiguity between similar brand names or color terms.</p></li><li><p><strong>A percolator</strong> to implement the policy layer. A percolator reverses the usual search direction: Instead of matching a query against stored documents, it stores the queries and matches an incoming piece of text (here, the customer's message) against them. That's exactly what policy matching needs, because each policy is essentially a saved pattern that says "When a query looks like this, surface this hint."</p></li><li><p><strong>Real-time aggregations</strong> over the full product catalog to produce accurate facet state at query time. Precomputed facet snapshots go stale quickly in active catalogs; query-time aggregations are the more reliable source.</p></li><li><p><strong>Document retrieval by key</strong> for user profiles: fast, single-document lookups by user ID that must complete within the context build window.</p></li><li><p><strong>Structured and semantic logging and analytics</strong> over query traces and agent interactions, which are the raw material for the governed loop's automatic signal surfacing.</p></li></ul><p>These are standard capabilities of a mature search and analytics platform, rather than six separate systems, and Elasticsearch provides all of them in one place. That matters less as a procurement point than as an architectural one: When semantic matching, percolation, aggregations, profile lookups, and analytics all run against the same catalog in the same cluster, the context layer stays consistent with the search layer by construction. Splitting these across a separate vector store and a separate analytics platform is a legitimate choice, but it adds operational surface and introduces a consistency problem between two systems that are reasoning about the same products. The context infrastructure is simplest to run when it lives where the product data already lives and can be updated without a separate extract, transform, load (ETL) step.</p><h2><strong>Conclusion: Context engineering is a search team's job to own</strong></h2><p>The retailers who run effective AI shopping experiences at scale aren't the ones with the largest models or the most generous token budgets. They're the ones who have done the work to make their catalog, vocabulary, and tpolicies legible to an agent before it starts reasoning.</p><p>The good news is that most of the work is already done. The vocabulary is implicit in the catalog. The policies exist as merchandising rules and compliance guidelines. The user profiles are in the CRM. The session signals are in the analytics stream. The gap isn't data; it's the assembly layer that turns those signals into a structured context the agent can consume before it starts reasoning.</p><p>The search team already owns the vocabulary, policies, and merchandising workflows. The context layer is the right home for work the search team is already doing, in a form that serves the agent as well as the query pipeline. And, because it grows every time a gap is found and filled, it behaves less like a setup cost and more like an asset that compounds.</p><p>To begin building a context layer in Elastic, you can start a <a href="https://www.elastic.co/cloud?utm_campaign=G-TXT-EMEA-UK+CA-Core-EN-Lead_Gen-CloudTrials-BR&amp;utm_content=Brand-Cloud&amp;utm_source=google&amp;utm_medium=cpc&amp;device=c&amp;utm_term=elastic%20cloud%20trial&amp;utm_id=701610000005lJVAAY&amp;gad_source=1&amp;gad_campaignid=22979576770&amp;gbraid=0AAAAADrDgoJnVYpNJwbmfVxoTcZSmr4S8&amp;gclid=CjwKCAjwx7LSBhB3EiwAjcodxPrvWRgCciehjf-6cu_sOb7FxbwDEJiS8Dpl95oQo7D2J61zXLJrgRoCtgQQAvD_BwE">trial of Elastic Cloud</a> or <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">run locally</a>. You should become familiar with configuring <a href="https://www.elastic.co/docs/solutions/search/semantic-search">semantic search</a>, and if you’re interested in how to build, store, and match search policies at query time, you will enjoy <a href="https://www.elastic.co/search-labs/blog/elasticsearch-percolator-search-governance">this blog</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-shopping-agents-context-engineering</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-shopping-agents-context-engineering</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Hybrid Search]]></category>
    <category><![CDATA[Relevance]]></category>
    <dc:creator><![CDATA[Matthew Adams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt360319cdeac21a74/6ab3e9caf3b1262ef3650cdb/diagram-customer-query-query-input-context-build.webp" length="0" type="image/webp"/>
    <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[137,000 people, zero human decisions: agentic disaster response with Elasticsearch]]></title>
    <description><![CDATA[Find out how a Kibana detection rule, a workflow and an AI agent automatically relocated 137,000 military personnel across seven installations when a hurricane hit, no dispatcher required.]]></description>
    <content:encoded><![CDATA[<p>Elastic just coordinated the automated evacuation of 137,000 military personnel across seven installations, with no human in the loop. A Category 4 hurricane hits the Hampton Roads coastline. Elasticsearch's geospatial enrichment identifies every facility in the impact zone at index time. A Kibana detection rule fires. A workflow starts an AI agent conversation. The agent reasons through capacity, distance and branch compatibility, then dispatches 16 evacuation and intake notifications in a single pass. From raw GDACS event to coordinated action, automatically.</p><p>Every year, natural disasters force emergency managers, military commanders, and public safety officials to make high-stakes decisions in compressed time frames. These decisions traditionally rely on phone trees, spreadsheets, and institutional knowledge spread across dozens of people. The coordination overhead alone costs critical time.</p><p>This post demonstrates how Elastic can power a responsive, agentic coordination system for disaster response that detects a threat, reasons through the logistics, and takes action automatically. To make it concrete, we built a simulation: a fictitious Category 4 hurricane threatening the Hampton Roads coastline triggers the automated relocation of over 137,000 personnel across seven military installations.</p><p><strong>Disclaimer:</strong> <strong>This is an entirely fictitious scenario built for demonstration purposes. </strong>Hurricane ELARA-26 doesn’t exist. Installation locations are based on real, publicly available geographic data (the U.S. Department of Defense [DoD] Military Installations, Ranges, and Training Areas [MIRTA] dataset), but all operational data, like personnel counts, housing capacity, assets, contact emails, and mission profiles, are completely fabricated. Nothing in this demo reflects actual military readiness, capability, or operational procedures.</p><h2>Why automated disaster response requires geospatial and agentic coordination</h2><p>When a natural disaster threatens critical infrastructure, the coordination challenge is immediate:</p><ul><li><p>Which facilities are in the impact zone?</p></li><li><p>How many personnel need to move?</p></li><li><p>Where can they go, and do those facilities have capacity?</p></li><li><p>Who needs to be notified right now?</p></li></ul><p>These questions don't wait. Neither should the answers.</p><h2>Deploy the pipeline: prerequisites and setup</h2><p>Follow the instructions <a href="https://github.com/tehbooom/elastic_natural_disaster/blob/main/README.md">here in the example repo</a> to deploy a local Elastic cluster with Elastic Inference Service (EIS) via <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/connect-self-managed-cluster-to-eis#set-up-eis-with-cloud-connect">Cloud Connect</a>.</p><h2>How the Elasticsearch agentic disaster response pipeline works</h2><p>The pipeline has seven layers that work together end to end:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf09bfae87ab35bec/6a4693ef31bdbbe3ef8b33ae/61814cddea0409162fb057c2113e0a496c105238-1999x275.png" alt="Pipeline flowchart Alt text: Horizontal flowchart with seven labeled boxes connected by arrows: GDACS feed, ingest pipeline, enrich (geo_shape), detection rule, workflow, AI agent, and email." /><ol><li><p><strong>Data ingestion:</strong> Global Disaster Alert and Coordination System (GDACS) disaster events sent to Elasticsearch</p></li><li><p>I<strong>ngest pipeline</strong>: GeoJSON is ingested and normalized to Elastic Common Schema (ECS).</p></li><li><p><strong>Geospatial enrichment:</strong> The event's affected area polygon is matched against indexed military installation boundaries.</p></li><li><p><strong>Alerting:</strong> A Kibana detection rule fires when a disaster intersects any installation.</p></li><li><p><strong>Workflow automation:</strong> The alert triggers a Kibana workflow that starts an AI agent conversation.</p></li><li><p><strong>AI reasoning:</strong> The agent reasons through affected facilities, their assets, and nearest supporting facilities to determine relocation of all assets and personnel.</p></li><li><p><strong>Email notifications</strong>: The agent dispatches emails to all recipients for incoming and outgoing personnel and or assets.</p></li></ol><p>Let's walk through each layer.</p><h2>Step 1: Indexing military installations with geo boundaries</h2><p>The foundation is the DoD MIRTA dataset from <a href="https://source.coop/seerai/hifld/military-installations-ranges-and-training-areas-mirta-dod-sites---boundaries">source.coop/seerai/hifld</a>. This dataset provides a <code>geo_shape</code> of type <code>Point</code> for each installation; centroid coordinates rather than full boundary polygons.</p><p>Each installation document in the <code>mitra-facilities</code> index is enriched with operational profile data (all fictitious), beyond what MIRTA provides:</p>{
  "entity_name": "Naval Station Norfolk",
  "branch_of_service": "Navy",
  "mission_function_type": "fleet_support",
  "personnel_count": 50000,
  "housing_capacity": 55000,
  "temporary_housing_capacity": 10000,
  "logistics_capabilities": ["fuel", "airlift", "sealift", "medical"],
  "available_assets": [
    { "type": "helicopters", "count": 24 },
    { "type": "transport_vehicles", "count": 150 }
  ],
  "contact_email": "norfolk.ops@navy.mil.gov.fake",
  "operational_status": "act",
  "is_joint_base": false,
  "entity_geo_location": { "type": "polygon", "coordinates": [...] }
}<p>This rich index is what enables the AI agent to make intelligent allocation decisions; not just "here are nearby bases," but "here are bases with available housing capacity, compatible mission types, and the logistics to receive incoming assets."</p><h2>Step 2: Ingesting and normalizing GDACS events</h2><p>GDACS publishes real-time GeoJSON for earthquakes, tropical cyclones, floods, wildfires, volcanoes, and droughts. We ingest this feed into a data stream (<code>logs-gdacs.events-*</code>) using a custom ingest pipeline that normalizes the raw GeoJSON to ECS fields.</p><p>The GDACS ingest pipeline does several things worth noting:</p><p><strong>Geometry extraction:</strong> The centroid is stored as a <code>geo_point</code> for map display, and the impact polygon is stored as a <code>geo_shape</code> in <code>gdacs.affected_area</code>, which is the field used for intersection queries later.</p><p><strong>Severity normalization:</strong> Each disaster type has a different severity scale. A tropical cyclone is measured in km/h wind speed; an earthquake in Richter magnitude. The pipeline maps all of them to a normalized 0–100 score:</p>// Painless snippet from the ingest pipeline
if (type == 'TC') {
  norm = Math.min(100.0, Math.max(0.0, (val - 40.0) / 2.6));
} else if (type == 'EQ') {
  norm = Math.min(100.0, Math.max(0.0, (val - 4.0) * 20.0));
}<p>The normalized severity score then maps to a <code>severity_level</code> label (<code>low</code>, <code>medium</code>, <code>high</code>, <code>critical</code>) used for alert severity mapping in the detection rule.</p><p><strong>ECS alignment:</strong> <code>event.kind: alert, event.category: threat</code>, timestamps mapped to <code>event.start/event.end</code>, and a stable fingerprint-based <code>_id</code> for deduplication.</p><h2>Step 3: Geospatial enrichment: Finding affected facilities at index time</h2><p>Elasticsearch's geo_match enrich policy matches the disaster polygon against every installation boundary at index time, with no query-time join required. Instead of querying at search time, we use an <strong>enrich processor</strong> in the ingest pipeline to match the disaster's impact polygon against every installation boundary <em>as the document is indexed</em>.</p><p>The enrich policy is a <code>geo_match</code> policy:</p>{
  "geo_match": {
    "indices": "mitra-facilities",
    "match_field": "entity_geo_location",
    "enrich_fields": [
      "entity_name",
      "entity_type",
      "entity_station_number",
      "entity_geo_city_name",
      "entity_geo_region_name"
    ]
  }
}<p>The processor runs at the end of the ingest pipeline:</p>{
  "enrich": {
    "policy_name": "facilities-geo",
    "field": "gdacs.affected_area",
    "target_field": "affected_facilities",
    "shape_relation": "INTERSECTS",
    "max_matches": 128
  }
}<p><code>INTERSECTS</code> catches any installation whose boundary touches or overlaps the disaster polygon and even partial intersections. The result is that every GDACS event document is stored with an <code>affected_facilities</code> nested array that tells us exactly which installations are in the impact zone. No join query needed.</p><h2>Step 4: Detection rule: Alerting on facility impact</h2><p>A Kibana detection rule watches the <code>logs-gdacs.events-*</code> data stream and fires when a GDACS event has been enriched with at least one affected facility:</p>Query: affected_facilities: { entity_name: * }<p>The rule runs on an hourly schedule (covering a <code>now-1h</code> to <code>now</code> window) and uses dynamic severity mapping; the <code>gdacs.severity_level</code> field computed by the ingest pipeline drives the alert severity automatically.</p><p>Alert severity also drives the risk score via field mapping:</p>"risk_score_mapping": [
  {
    "field": "gdacs.normalized_severity",
    "operator": "equals",
    "value": ""
  }
]<p>When the rule fires, it passes the full alert context, including the enriched <code>affected_facilities</code> array with installation names, types, and locations, downstream to a Kibana workflow.</p><h2>Step 5: Workflow automation: Bridging alert to agent</h2><p>Kibana Workflows handle the handoff from detection to response. The natural disaster response workflow is triggered by the alert:</p>triggers:
  - type: alert
steps:
  - name: start_convo
    type: kibana.request
    with:
      method: "POST"
      path: "/api/agent_builder/converse"
      body:
        agent_id: "mitra.response"
        input: "New Natural Disaster Alert: {{ event.alerts | json }}"<p>The entire alert payload (disaster type, severity, affected area, and the list of impacted installations) is forwarded to the AI agent as its initial context. The agent takes it from there.</p><h2>Step 6: The AI agent: From data to coordinated action</h2><p>The <code>mitra.response</code> agent takes the full alert payload and, in a single agentic loop, assesses scope, finds receiving facilities, allocates personnel and dispatches evacuation and intake notifications, all without human intervention.</p><p>The agent has two tools available:</p><ul><li><p><strong><code>mitra.nearest_facility</code></strong>queries the <code>mitra-facilities</code> index using a geo_shape query, sorted by distance from a given coordinate, returning up to 50 nearby active installations with available capacity.</p></li><li><p><strong><code>mitra.send_email</code></strong> iterates over a JSON array of facility objects and dispatches formatted evacuation or receiving notifications.</p></li></ul><p>The agent's instruction set defines a clear workflow:</p><ol><li><p><strong>Assess the situation.</strong> Parse the alert, identify affected facilities, and determine disaster scope.</p></li><li><p><strong>Inventory what needs to move.</strong> Personnel counts, critical assets, housing requirements per facility.</p></li><li><p><strong>Find destination facilities.</strong> Call <code>mitra.nearest_facility</code> for each affected installation, filtering out facilities still in the danger zone.</p></li><li><p><strong>Make allocation decisions.</strong> Reason through single versus multi-facility solutions, branch compatibility, housing capacity, asset support.</p></li><li><p><strong>Send coordination emails.</strong> Dispatch evacuation orders to source facilities and intake notifications to receiving facilities.</p></li><li><p><strong>Produce a summary report. </strong>Produces a short summary of all affected facilities, total personnel, assets moved, destination facilities, and any concerns to the chat for review.</p></li></ol><p>The agent's allocation logic follows real-world constraints: Don't exceed housing capacity, prefer same-branch relocations when possible, use joint bases for multi-branch overflow, and prioritize distance to minimize transit time.</p><h3>The nearest facility tool</h3><p>The underlying workflow query uses <code>geo_shape</code> with a circle filter and <code>_geo_distance</code> sorting:</p>"query": {
  "bool": {
    "filter": [
      {
        "geo_shape": {
          "entity_geo_location": {
            "shape": {
              "type": "circle",
              "coordinates": [{{ inputs.lon }}, {{ inputs.lat }}],
              "radius": "5000km"
            },
            "relation": "intersects"
          }
        }
      },
      { "term": { "operational_status.keyword": "act" } }
    ]
  }
},
"sort": [
  {
    "_geo_distance": {
      "entity_geo_point": { "lat": {{ inputs.lat }}, "lon": {{ inputs.lon }} },
      "order": "asc",
      "unit": "km"
    }
  }
],
"script_fields": {
  "available_capacity": {
    "script": {
      "source": "Math.max(0, doc['housing_capacity'].value - doc['personnel_count'].value)"
    }
  }
}<p>Available capacity is computed at query time via a script field which calculates housing capacity minus current personnel count. The agent uses this to allocate personnel across destinations without exceeding limits.</p><h2>Hurricane ELARA-26: agentic coordination of 137,000 personnel, end to end</h2><p>Hurricane ELARA-26 is a Category 4 storm (213 km/h maximum winds) projected to make landfall in the Hampton Roads area of Virginia. When the GDACS event is ingested, the affected area polygon intersects seven major military installations in the region. The detection rule fires. The workflow kicks off an agent conversation.</p><p>Within a single agentic loop, the agent:</p><ul><li><p>Identified seven facilities in the impact zone, with a combined 137,372 personnel.</p></li><li><p>Called <code>mitra.nearest_facility</code> to find receiving facilities outside the storm track.</p></li><li><p>Distributed personnel across nine receiving facilities based on available housing capacity and distance.</p></li><li><p>Generated and dispatched evacuation orders to all seven affected installations.</p></li><li><p>Generated and dispatched intake notifications to all nine receiving facilities.</p></li><li><p>Produced a full coordination summary, similar to below:</p></li></ul><p><strong>Facilities evacuated:</strong></p><p>Facility</p><p>Personnel</p><p>Naval Station Norfolk</p><p>50,000</p><p>Joint Expeditionary Base Little Creek-Fort Story</p><p>18,000</p><p>Naval Air Station Oceana</p><p>15,355</p><p>Naval Air Station Oceana Dam Neck Annex</p><p>17,509</p><p>NG State Military Reservation Camp Pendleton</p><p>9,707</p><p>Joint Base Langley-Eustis</p><p>15,000</p><p>Naval Weapons Station Yorktown</p><p>11,801</p><p><strong>Receiving facilities:</strong></p><p>Facility</p><p>Distance</p><p>Incoming personnel</p><p>Fort Gregg-Adams</p><p>97 km</p><p>~40,000</p><p>Marine Corps Base Quantico</p><p>148 km</p><p>~30,000</p><p>Naval Support Facility Indian Head</p><p>151 km</p><p>~30,000</p><p>Joint Base Andrews</p><p>180 km</p><p>~30,000</p><p>Naval Air Station Patuxent River</p><p>141 km</p><p>~10,000</p><p>NG MTA Camp Butner</p><p>174 km</p><p>~5,000</p><p>NG Bethany Beach Training Site</p><p>209 km</p><p>~4,707</p><p>Rivanna Station</p><p>140 km</p><p>~7,500</p><p>Def Gen Supply Center</p><p>22 km</p><p>~6,000</p><p>Assets relocated include transport vehicles, helicopters, patrol boats, medical units, engineering vehicles, generators, water trailers, shelter kits, and communication systems.</p><h3>Automated email notifications</h3><p>Once the agent finalized its allocation plan, it called <code>mitra.send_email</code> and dispatched 16 emails in a single pass; that is, evacuation orders to all seven affected installations and intake notifications to all nine receiving facilities. Each message included destination facility, incoming personnel count, assets to move, and a coordination contact. What would have taken hours of phone trees completed automatically the moment the agent finished reasoning.</p><h3>Extending agentic disaster response with RAG and policy grounding</h3><p>This demo is purely from structured data, like capacity numbers, distances, and operational status. Elastic's semantic search and retrieval augmented generation (RAG) capabilities can make the agent significantly smarter, with two additions:</p><p><strong>Historical response retrieval:</strong> Index past after-action reports, Federal Emergency Management Agency (FEMA) incident summaries, and disaster response records as vector embeddings. When a new event fires, the agent can semantically retrieve how similar events were handled, informing allocation decisions with institutional knowledge rather than capacity math alone.</p><p><strong>Policy and doctrine grounding:</strong> Index DoD emergency management directives, installation continuity of operations plans, and commander guidance. The agent can retrieve and cite the actual policies governing a response, ensuring every decision is grounded in doctrine rather than inference.</p><p>Both follow the same Elastic-native approach:; An inference pipeline generates embeddings at index time, and a semantic search tool is exposed to the agent. The coordination pipeline stays the same. The agent just gets smarter.</p><h2>Why Elasticsearch is the right platform for agentic public sector response</h2><p>This isn’t a chatbot. It isn’t a dashboard. It's a responsive agentic workflow system, one that detected a threat, reasoned through a complex logistics problem, and coordinated the relocation of 137,000 people without a human in the loop. That kind of outcome is only possible because every capability it depends on lives in a single, unified platform.</p><p>Elasticsearch's geospatial support (<code>geo_point</code>, <code>geo_shape</code>, enrichment policies, and distance-based sorting) handles the spatial reasoning that makes intersection detection and facility lookup possible at scale. Semantic search and vector embeddings ground agents in truth, ensuring AI reasoning is based on what's actually in your data rather than hallucinated assumptions. Kibana's detection engine, Workflows, Agent Builder, and Agent Builder tools wire it all together into a pipeline that goes from raw event to coordinated action with no external glue code required.</p><p>No other platform brings this together the way Elastic does. The combination of real-time indexing, geospatial precision, semantic retrieval, and agentic orchestration, all in one stack, with enterprise-grade security and observability built in, is what separates Elastic from tools that do one of these things well but require you to stitch the rest together yourself.</p><h2>Agentic geospatial response for emergency management, fire, law enforcement and public health</h2><p>The same architecture applies wherever people, facilities, and real-time events intersect. The specific data changes. The pipeline doesn't.</p><p><strong>Emergency management:</strong> FEMA and state offices of emergency management can map shelter locations, staging areas, and vulnerable populations against incoming National Weather Service (NWS) severe weather polygons, triggering automated resource pre-positioning before a storm makes landfall.</p><p><strong>Fire and emergency medical services:</strong> Fire departments can overlay unit locations and response zones against wildfire perimeters or structure fire clusters, automatically routing mutual aid requests to the nearest available units with the right equipment.</p><p><strong>Law enforcement:</strong> Agencies can correlate active incident locations with school zones, critical infrastructure, and officer positions, triggering geo-aware lockdown notifications or resource dispatch without waiting for manual triage.</p><p><strong>Public school safety:</strong> School districts can monitor real-time threat feeds against campus boundaries. When a threat intersects a school's perimeter, an agent can immediately notify administration, initiate lockdown communications, and coordinate law enforcement response, all before a dispatcher picks up a phone.</p><p><strong>Public health:</strong> Health departments can match disease surveillance data or environmental hazard zones against clinic locations, population density layers, and supply depot inventories to route resources where they're needed most.</p><p>Sector</p><p>Use case</p><p>Elastic capability</p><p>Emergency management</p><p>Match shelter locations against NWS severe weather polygons</p><p>geo_shape enrichment + Kibana Workflows</p><p>Fire and EMS</p><p>Overlay unit locations against wildfire perimeters</p><p>geospatial routing + nearest-facility query</p><p>Law enforcement</p><p>Correlate incidents with school zones and officer positions</p><p>geo-aware alert rules + agent dispatch</p><p>Public school safety</p><p>Monitor threat feeds against campus perimeters</p><p>detection rules + automated notification</p><p>Public health</p><p>Match hazard zones against clinic locations and supply depots</p><p>semantic search + geospatial enrichment</p><p>The data is different in every scenario. The underlying pattern of ingest, enrich at index time, detect intersection, trigger agentic response, and act is all the same. Elastic gives public sector organizations the platform to build it once and adapt it everywhere.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-agentic-disaster-response</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-agentic-disaster-response</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Kibana]]></category>
    <dc:creator><![CDATA[Alec Carpenter]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt969cad2694920de4/6a4693f37746672ad42675b5/cb292a501835472598dee30bef25c77afc54db6c-720x420.png" length="0" type="image/png"/>
    <pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Creating an Elasticsearch MCP server with TypeScript]]></title>
    <description><![CDATA[Learn how to create an Elasticsearch MCP server with TypeScript and Claude Desktop.]]></description>
    <content:encoded><![CDATA[<p>When working with large knowledge bases in Elasticsearch, finding information is only half the battle. Engineers often need to synthesize results from multiple documents, generate summaries, and trace answers back to their sources. Model Context Protocol (MCP) provides a standardized way to connect Elasticsearch with large language model–powered (LLM-powered) applications to accomplish this. While Elastic offers official solutions, like Elastic Agent Builder (which includes an <a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">MCP endpoint</a> among its features), building a custom MCP server gives you full control over search logic, result formatting, and how retrieved content is passed to an LLM for synthesis, summaries, and citations.</p><p>In this article, we’ll explore the benefits of building a custom Elasticsearch MCP server and show how to create one in TypeScript that connects Elasticsearch to LLM-powered applications.</p><h2>Why build a custom Elasticsearch MCP server?</h2><p>Elastic provides some alternatives for <a href="https://www.elastic.co/docs/solutions/search/mcp">MCP servers</a>:</p><ul><li><p><a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">Elastic Agent Builder MCP server for Elasticsearch 9.2+</a></p></li><li><p><a href="https://github.com/elastic/mcp-server-elasticsearch?tab=readme-ov-file#elasticsearch-mcp-server">Elasticsearch MCP server for older versions (Python)</a></p></li></ul><p>If you need more control over how your MCP server interacts with Elasticsearch, building your own custom server gives you the flexibility to tailor it exactly to your needs. For example, Agent Builder's MCP endpoint is limited to Elasticsearch Query Language (ES|QL) queries, while a custom server allows you to use the full Query DSL. You also gain control over how results are formatted before being passed to the LLM and can integrate additional processing steps, like the OpenAI-powered summarization we'll implement in this tutorial.</p><p>By the end of this article, you’ll have an MCP server in TypeScript that searches for information stored in an Elasticsearch index, summarizes it, and provides citations. We'll use Elasticsearch for retrieval, OpenAI's <code>gpt-4o-mini</code> model to summarize and generate citations, and Claude Desktop as the MCP client and UI to take in user queries and give responses. The end result is an internal knowledge assistant that helps engineers discover and synthesize best practices across their organization’s technical docs.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltad9133cb083ad352/6a170c19b0367d411e72bd5b/ec5771a874cf9740d4cac6888622cbe8cd6aede7-1999x1133.png" alt="Creating an Elastic MCP server with TypeScript and Claude Desktop." /><h2>Prerequisites:</h2><ul><li><p>Node.js 20 +</p></li><li><p>Elasticsearch</p></li><li><p>OpenAI API key</p></li><li><p>Claude Desktop</p></li></ul><h3>What is MCP?</h3><p><a href="https://www.elastic.co/what-is/mcp">MCP</a> is an open standard, created by <a href="https://www.anthropic.com/news/model-context-protocol">Anthropic</a>, that provides secure, bidirectional connections between LLMs and external systems, like Elasticsearch. You can read more about the current state of MCP in <a href="https://www.elastic.co/search-labs/blog/mcp-current-state">this article</a>.</p><p>The MCP landscape is <a href="https://www.elastic.co/search-labs/blog/mcp-current-state#mcp-project-updates:-transport,-elicitation,-and-structured-tooling">evolving every day</a>, with servers available for a wide range of use cases. On top of that, it’s easy to build your own custom MCP server, as we’ll show in this article.</p><h3>MCP clients</h3><p>There’s a long <a href="https://modelcontextprotocol.io/clients">list of available MCP clients</a>, each with its own characteristics and limitations. For simplicity and popularity, we’ll use <a href="https://claude.ai/download">Claude Desktop</a> as our MCP client. It will serve as the chat interface where users can ask questions in natural language, and it will automatically invoke the tools exposed by our MCP server to search documents and generate summaries.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt06fd7a02042094e1/6a170c1b14b2700024e3c651/66eb0b11473347b6cf2d85718251eeac38d6249d-1999x1491.png" alt="Claude 4.5 Sonnet page, with the note, &quot;Coffee and Claude time? How can I help you today?&quot;" /><h2>Creating an Elasticsearch MCP server</h2><p>Using the <a href="https://github.com/modelcontextprotocol/typescript-sdk">TypeScript SDK</a>, we can easily create a server that understands how to query our Elasticsearch data based on a user query input.</p><p>Here are the steps in this article to integrate the Elasticsearch MCP server with the Claude Desktop client:</p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#configure-mcp-server-for-elasticsearch">Configure MCP server for Elasticsearch.</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#load-the-mcp-server-into-claude-desktop">Load the MCP server into Claude Desktop.</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude#test-it-out">Test it out.</a></p></li></ol><h3>Configure MCP server for Elasticsearch</h3><p>To begin, let's initialize a node application:</p>npm init -y<p>This will create a <code>package.json</code> file, and with it, we can start installing the necessary dependencies for this application.</p>npm install @elastic/elasticsearch @modelcontextprotocol/sdk openai zod &amp;&amp; npm install --save-dev ts-node @types/node typescript<ul><li><p><strong>@elastic/elasticsearch</strong> will give us access to the Elasticsearch Node.js library.</p></li><li><p><strong>@modelcontextprotocol/sdk</strong> provides the core tools to create and manage an MCP server, register tools, and handle communication with MCP clients.</p></li><li><p><strong>openai</strong> allows interaction with OpenAI models to generate summaries or natural language responses.</p></li><li><p><a href="https://zod.dev/"><strong>zod</strong></a>helps define and validate structured schemas for input and output data in each tool.</p></li></ul><p><code>ts-node</code>, <code>@types/node</code>, and <code>typescript</code> will be used during development to type the code and compile the scripts.</p><h4>Set up the dataset</h4><p>To provide the data that Claude Desktop can query using our MCP server, we’ll use a mock <a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/dataset.json">internal knowledge base dataset</a>. Here’s what a document from this dataset will look like:</p>{
    "id": 5,
    "title": "Logging Standards for Microservices",
    "content": "Consistent logging across microservices helps with debugging and tracing. Use structured JSON logs and include request IDs and timestamps. Avoid logging sensitive information. Centralize logs in Elasticsearch or a similar system. Configure log rotation to prevent storage issues and ensure logs are searchable for at least 30 days.",
    "tags": ["logging", "microservices", "standards"]
}<p>To ingest the data, we prepared a script that creates an index in Elasticsearch and loads the dataset into it. You can find it <a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/setup.ts">here</a>.</p><h4>MCP server</h4><p>Create a file named <a href="https://github.com/Delacrobix/typescript-elasticsearch-mcp/blob/main/index.ts"><code>index.ts</code></a> and add the following code to import the dependencies and handle environment variables:</p>// index.ts
import { z } from "zod";
import { Client } from "@elastic/elasticsearch";
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import OpenAI from "openai";

const ELASTICSEARCH_ENDPOINT =
  process.env.ELASTICSEARCH_ENDPOINT ?? "http://localhost:9200";
const ELASTICSEARCH_API_KEY = process.env.ELASTICSEARCH_API_KEY ?? "";
const OPENAI_API_KEY = process.env.OPENAI_API_KEY ?? "";
const INDEX = "documents";<p>Also, let’s initialize the clients to handle the Elasticsearch and OpenAI calls:</p>const openai = new OpenAI({
  apiKey: OPENAI_API_KEY,
});

const _client = new Client({
  node: ELASTICSEARCH_ENDPOINT,
  auth: {
    apiKey: ELASTICSEARCH_API_KEY,
  },
});<p>To make our implementation more robust and ensure structured input and output, we'll define schemas using <a href="https://zod.dev/"><code>zod</code></a>. This allows us to validate data at runtime, catch errors early, and make the tool responses easier to process programmatically:</p>const DocumentSchema = z.object({
  id: z.number(),
  title: z.string(),
  content: z.string(),
  tags: z.array(z.string()),
});

const SearchResultSchema = z.object({
  id: z.number(),
  title: z.string(),
  content: z.string(),
  tags: z.array(z.string()),
  score: z.number(),
});

type Document = z.infer&lt;typeof DocumentSchema&gt;;
type SearchResult = z.infer&lt;typeof SearchResultSchema&gt;;<p>Learn more about structured outputs <a href="https://www.elastic.co/search-labs/blog/structured-outputs-elasticsearch-guide">here</a>.</p><p>Now let’s initialize the MCP server:</p>const server = new McpServer({
  name: "Elasticsearch RAG MCP",
  description:
    "A RAG server using Elasticsearch. Provides tools for document search, result summarization, and source citation.",
  version: "1.0.0",
});<h4>Defining the MCP tools</h4><p>With everything configured, we can start writing the tools that will be exposed by our MCP server. This server exposes two tools:</p><ul><li><p><strong><code>search_docs</code></strong><strong>: </strong>Searches for documents in Elasticsearch using full-text search.</p></li><li><p><strong><code>summarize_and_cite</code></strong><strong>:</strong> Summarizes and synthesizes information from previously retrieved documents to answer a user question. This tool also adds citations referencing the source documents.</p></li></ul><p>Together, these tools form a simple “retrieve-then-summarize” workflow, where one tool fetches relevant documents and the other uses those documents to generate a summarized, cited response.</p><h4>Tool response format</h4><p>Each tool can accept arbitrary input parameters, but it must respond with the following structure:</p><ul><li><p><strong>Content:</strong> This is the response of the tool in an unstructured format. This field is usually used to return text, images, audio, links, or embeddings. For this application, it will be used to return formatted text with the information generated by the tools.</p></li><li><p><strong>structuredContent: </strong>This is an optional return used to provide the results of each tool in a structured format. This is useful for programmatic purposes. Although it isn't used in this MCP server, it can be useful if you want to develop other tools or process the results programmatically.</p></li></ul><p>With that structure in mind, let’s dive into each tool in detail.</p><h4>Search_docs tool</h4><p>This tool performs a <a href="https://www.elastic.co/docs/solutions/search/full-text">full-text search</a> in the Elasticsearch index to retrieve the most relevant documents based on the user query. It highlights key matches and provides a quick overview with relevance scores.</p>server.registerTool(
  "search_docs",
  {
    title: "Search Documents",
    description:
      "Search for documents in Elasticsearch using full-text search. Returns the most relevant documents with their content, title, tags, and relevance score.",
    inputSchema: {
      query: z
        .string()
        .describe("The search query terms to find relevant documents"),
      max_results: z
        .number()
        .optional()
        .default(5)
        .describe("Maximum number of results to return"),
    },
    outputSchema: {
      results: z.array(SearchResultSchema),
      total: z.number(),
    },
  },
  async ({ query, max_results }) =&gt; {
    if (!query) {
      return {
        content: [
          {
            type: "text",
            text: "Query parameter is required",
          },
        ],
        isError: true,
      };
    }

    try {
      const response = await _client.search({
        index: INDEX,
        size: max_results,
        query: {
          bool: {
            must: [
              {
                multi_match: {
                  query: query,
                  fields: ["title^2", "content", "tags"],
                  fuzziness: "AUTO",
                },
              },
            ],
            should: [
              {
                match_phrase: {
                  title: {
                    query: query,
                    boost: 2,
                  },
                },
              },
            ],
          },
        },
        highlight: {
          fields: {
            title: {},
            content: {},
          },
        },
      });

      const results: SearchResult[] = response.hits.hits.map((hit: any) =&gt; {
        const source = hit._source as Document;

        return {
          id: source.id,
          title: source.title,
          content: source.content,
          tags: source.tags,
          score: hit._score ?? 0,
        };
      });

      const contentText = results
        .map(
          (r, i) =&gt;
            `[${i + 1}] ${r.title} (score: ${r.score.toFixed(
              2,
            )})\n${r.content.substring(0, 200)}...`,
        )
        .join("\n\n");

      const totalHits =
        typeof response.hits.total === "number"
          ? response.hits.total
          : (response.hits.total?.value ?? 0);

      return {
        content: [
          {
            type: "text",
            text: `Found ${results.length} relevant documents:\n\n${contentText}`,
          },
        ],
        structuredContent: {
          results: results,
          total: totalHits,
        },
      };
    } catch (error: any) {
      console.log("Error during search:", error);

      return {
        content: [
          {
            type: "text",
            text: `Error searching documents: ${error.message}`,
          },
        ],
        isError: true,
      };
    }
  }
);<p><em>We configure </em><a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-fuzzy-query"><em><code>fuzziness</code></em></a><em><code>: “AUTO”</code></em><em> to have a variable typo tolerance based on the length of the token that’s being analyzed. We also set </em><em><code>title^2</code></em><em> to increase the score of the documents where the match happens on the title field.</em></p><h4>summarize_and_cite tool</h4><p>This tool generates a summary based on documents retrieved in the previous search. It uses OpenAI’s <code>gpt-4o-mini</code> model to synthesize the most relevant information to answer the user’s question, providing responses derived directly from the search results. In addition to the summary, it also returns citation metadata for the source documents used.</p>server.registerTool(
  "summarize_and_cite",
  {
    title: "Summarize and Cite",
    description:
      "Summarize the provided search results to answer a question and return citation metadata for the sources used.",
    inputSchema: {
      results: z
        .array(SearchResultSchema)
        .describe("Array of search results from search_docs"),
      question: z.string().describe("The question to answer"),
      max_length: z
        .number()
        .optional()
        .default(500)
        .describe("Maximum length of the summary in characters"),
      max_docs: z
        .number()
        .optional()
        .default(5)
        .describe("Maximum number of documents to include in the context"),
    },
    outputSchema: {
      summary: z.string(),
      sources_used: z.number(),
      citations: z.array(
        z.object({
          id: z.number(),
          title: z.string(),
          tags: z.array(z.string()),
          relevance_score: z.number(),
        })
      ),
    },
  },
  async ({ results, question, max_length, max_docs }) =&gt; {
    if (!results || results.length === 0 || !question) {
      return {
        content: [
          {
            type: "text",
            text: "Both results and question parameters are required, and results must not be empty",
          },
        ],
        isError: true,
      };
    }

    try {
      const used = results.slice(0, max_docs);

      const context = used
        .map(
          (r: SearchResult, i: number) =&gt;
            `[Document ${i + 1}: ${r.title}]\\n${r.content}`
        )
        .join("\n\n---\n\n");

      // Generate summary with OpenAI
      const completion = await openai.chat.completions.create({
        model: "gpt-4o-mini",
        messages: [
          {
            role: "system",
            content:
              "You are a helpful assistant that answers questions based on provided documents. Synthesize information from the documents to answer the user's question accurately and concisely. If the documents don't contain relevant information, say so.",
          },
          {
            role: "user",
            content: `Question: ${question}\\n\\nRelevant Documents:\\n${context}`,
          },
        ],
        max_tokens: Math.min(Math.ceil(max_length / 4), 1000),
        temperature: 0.3,
      });

      const summaryText =
        completion.choices[0]?.message?.content ?? "No summary generated.";

      const citations = used.map((r: SearchResult) =&gt; ({
        id: r.id,
        title: r.title,
        tags: r.tags,
        relevance_score: r.score,
      }));

      const citationText = citations
        .map(
          (c: any, i: number) =&gt;
            `[${i + 1}] ID: ${c.id}, Title: "${c.title}", Tags: ${c.tags.join(
              ", ",
            )}, Score: ${c.relevance_score.toFixed(2)}`,
        )
        .join("\n");

      const combinedText = `Summary:\\n\\n${summaryText}\\n\\nSources used (${citations.length}):\\n\\n${citationText}`;

      return {
        content: [
          {
            type: "text",
            text: combinedText,
          },
        ],
        structuredContent: {
          summary: summaryText,
          sources_used: citations.length,
          citations: citations,
        },
      };
    } catch (error: any) {
      return {
        content: [
          {
            type: "text",
            text: `Error generating summary and citations: ${error.message}`,
          },
        ],
        isError: true,
      };
    }
  }
);<p>Finally, we need to start the server using <a href="https://github.com/modelcontextprotocol/typescript-sdk?tab=readme-ov-file#stdio">stdio</a>. This means the MCP client will communicate with our server by reading and writing to its standard input and output streams. stdio is the simplest transport option and works well for local MCP servers launched as subprocesses by the client. Add the following code at the end of the file:</p>const transport = new StdioServerTransport();
server.connect(transport);<p>Now compile the project using the following command:</p>npx tsc index.ts --target ES2022 --module node16 --moduleResolution node16 --outDir ./dist --strict --esModuleInterop<p>This will create a <code>dist</code> folder, and inside it, an <code>index.js</code> file.</p><h3>Load the MCP server into Claude Desktop</h3><p>Follow <a href="https://modelcontextprotocol.io/docs/develop/connect-local-servers">this guide</a> to configure the MCP server with Claude Desktop. In the Claude configuration file, we need to set the following values:</p>{
  "mcpServers": {
    "elasticsearch-rag-mcp": {
      "command": "node",
      "args": [   "/Users/user-name/app-dir/dist/index.js"
      ],
      "env": {
        "ELASTICSEARCH_ENDPOINT": "your-endpoint-here",
        "ELASTICSEARCH_API_KEY": "your-api-key-here",
        "OPENAI_API_KEY": "your-openai-key-here"
      }
    }
  }
}<p>The <code>args</code> value should point to the compiled file in the <code>dist</code> folder. You also need to set the environment variables in the configuration file with the exact same names defined in the code.</p><h3>Test it out</h3><p>Before executing each tool, click on <strong>Search and Tools</strong> to make sure that the tools are enabled. Here you can also enable or disable each one:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt395a7337021f9820/6a170c1c67045bb74d45c228/172981c2a54adabc70d5819013c3007670935605-1999x1002.png" alt="Claude 4.5 Sonnet page, with the note, &quot;Good afternoon, Jeff. How can I help you today?&quot;" /><p>Finally, let’s test the MCP server from the Claude Desktop chat and start asking questions:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf4ac458dc0206271/6a170c1e66c4f91328f8c072/03654c0f8c53c714f801fba8b25747071179209b-1999x1353.png" alt="User search request in Claude Desktop chat for documents about authentication methods and role-based access control, along with Claude's responses." /><p>For the question “<strong>Search for documents about authentication methods and role-based access control</strong>”, the <code>search_docs</code> tool is executed and returns the following results:</p>Most Relevant Documents:
Access Control and Role Management (highest relevance) - This document covers role-based access control (RBAC) principles, including ensuring users only have necessary permissions, regular auditing of user roles, revoking inactive accounts, and implementing just-in-time access for sensitive operations.
User Authentication with OAuth 2.0 - This document explains OAuth 2.0 authentication, which enables secure delegated access without credential sharing. It covers configuring identity providers, token management with limited scope and lifetime, and secure storage of refresh tokens.
Container Security Guidelines - While primarily about container security, this document touches on access control aspects like running containers as non-root users and avoiding embedded credentials.
Incident Response Playbook - This mentions role assignment during incidents (incident commander, communications lead, etc.), which relates to access control in emergency scenarios.
Logging Standards for Microservices - This document includes guidance on avoiding logging sensitive information, which is relevant to authentication security.<p>The response is, “Great! I found 5 relevant documents about authentication methods and role-based access control. Here's what was found:”</p><p>The tool call returns the source documents as part of its response payload, which are later used to generate citations.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbaf48a9468770ce2/6a170c21509168bffee1bb14/25ff4c7e9563d99752f95540dafdc7fd211a66e3-800x530.gif" alt="Claude 4.5 Sonnet page, with scrolling responses that include the five relevant documents about authentication methods and role-based access control." /><p>It’s also possible to chain multiple tools in a single interaction. In this case, Claude Desktop analyzes the user's question and determines that it needs to first call <code>search_docs</code> to retrieve relevant documents and to then pass those results to <code>summarize_and_cite</code> to generate the final answer, all without requiring separate prompts from the user:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta46ff45274e64192/6a170c230c4857a91501aac1/e6a8a46acb4236e77058f18bcd2f0737b5882c05-1999x1101.png" alt="Claude Desktop chat, with the note, &quot;Jeff returns,&quot; plus a new user question, &quot;What are the main recommendations to improve authentication and access control across our systems? Include references.&quot;" /><p>In this case, for the query “<strong>What are the main recommendations to improve authentication and access control across our systems? Include references.</strong>”, we obtained the following results:</p>Based on the documentation, here are the main recommendations to improve authentication and access control across your systems:

Key Recommendations

1. Implement Role-Based Access Control (RBAC)
Ensure users have only the permissions necessary for their job functions. This principle of least privilege minimizes security risks by limiting access to what's truly needed. [1]

2. Regular Access Audits
Conduct regular audits of user roles and promptly revoke access for inactive accounts. This helps prevent unauthorized access from dormant or compromised accounts. [1]

3. Just-in-Time (JIT) Access
Implement just-in-time access for sensitive operations, where permissions are granted only when needed and for a limited duration. This reduces the attack surface by minimizing standing privileges. [1]

4. OAuth 2.0 for Secure Authentication
Use OAuth 2.0 to enable secure delegated access without sharing user credentials. This protocol allows third-party applications to access resources on behalf of users without exposing passwords. [2]

5. Token Security and Management
Configure your identity provider to issue tokens with limited scope and lifetime

Store refresh tokens securely
Validate access tokens consistently to maintain security [2]
References

Access Control and Role Management (Tags: security, access-control)
User Authentication with OAuth 2.0 (Tags: authentication, oauth)
These recommendations work together to create a defense-in-depth approach, where multiple security layers protect your systems from unauthorized access.<p>As in the previous step, we can see the response from each tool for this question:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8f633c518e708a99/6a170c25ab7f082991db9ed6/cb606d356b2f7d5e4878a5eff71bc881869ac0ee-800x585.gif" alt="Claude Desktop chat page, with scrolling text that includes the response from each tool for the question, “What are the main recommendations to improve authentication and access control across our systems? Include references.”" /><p><em>Note: If a submenu appears asking whether you approve the use of each tool, select </em><em><strong>Always allow</strong></em><em> or </em><em><strong>Allow once</strong></em><em>.</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6627ee0bff1862df/6a170c266f7f040f6f91488c/aea942ba9b0037526ea215bec65690f1a5c3099c-1522x250.png" alt="Claude Desktop &quot;Always allow&quot; and &quot;Allow once&quot; options for a user to choose from." /><h2>Conclusion</h2><p>MCP servers represent a significant step toward standardizing LLM tools for both local and remote applications. Though full compatibility is still in the works, we’re moving fast in that direction.</p><p>In this article, we learned how to build a custom MCP server in TypeScript that connects Elasticsearch to LLM-powered applications. Our server exposes two tools: <code>search_docs</code> for retrieving relevant documents using Query DSL; and <code>summarize_and_cite</code> for generating summaries with citations via OpenAI models and Claude Desktop as client UI.</p><p>The future of compatibility between different client and server providers looks promising. Next steps include adding more functionalities and flexibility to your agent. There’s a practical <a href="https://www.elastic.co/search-labs/blog/llm-functions-elasticsearch-intelligent-query">article</a> on how you can parameterize your queries using search templates to gain precision and flexibility.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-mcp-server-typescript-claude</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Integrations]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5600198cb47666a5/6a170c28509168ce3ae1bb18/0bb24c05fff391f42070c2883182ea6fe9cb9680-1280x720.png" length="0" type="image/png"/>
    <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Using Elasticsearch Inference API along with Hugging Face models]]></title>
    <description><![CDATA[Learn how to connect Elasticsearch to Hugging Face models using inference endpoints, and build a multilingual blog recommendation system with semantic search and chat completions.]]></description>
    <content:encoded><![CDATA[<p>In recent updates, Elasticsearch introduced a native integration to connect to models hosted on the <a href="https://endpoints.huggingface.co/">Hugging Face Inference Service</a>. In this post, we’ll explore how to configure this integration and perform inference through simple API calls using a large language model (LLM). We’ll use <a href="https://huggingface.co/HuggingFaceTB/SmolLM3-3B">SmolLM3-3B</a>, a lightweight general-purpose model with a good balance between resource usage and answer quality.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9094997548bd70f8/6a170d6a839dfa0ad6dcff54/7ddadf1976421a860a7d62087239adb9150d808b-1999x1388.png" alt="Scatter plot showing several small language models plotted by model size (in billions of parameters) on the x‑axis and win rate (percentage) on the y‑axis. SmolLM3‑3B appears near the top of the efficiency trend, with a higher win rate than other models of similar size." /><h2>Prerequisites</h2><ul><li><p><strong>Elasticsearch 9.3 or Elastic Cloud Serverless: </strong>You can create a cloud deployment following <a href="https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud">these instructions</a>, or you can use the <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart#local-dev-quick-start"><code>start-local</code></a> quickstart instead.</p></li><li><p><strong>Python 3.12: </strong>Download Python <a href="https://www.python.org/">here</a>.</p></li><li><p><strong>Hugging Face </strong><a href="https://huggingface.co/docs/hub/en/security-tokens">access token</a>.</p></li></ul><h2>Chat completions using a Hugging Face inference endpoint</h2><p>First, we’ll build a practical example that connects Elasticsearch to a Hugging Face <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put">inference endpoint</a> to generate AI-powered recommendations from a collection of blog posts. For the app knowledge base, we’ll use a dataset of company blog articles, which contains valuable but often hard-to-navigate information.</p><p>With this endpoint, <a href="https://www.elastic.co/docs/solutions/search/semantic-search">semantic search</a> retrieves the most relevant articles for a given query, and a Hugging Face LLM generates short, contextual recommendations based on those results.</p><p>Let’s take a look at a high-level overview of the information flow we’re going to build:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf217b7b7db4e1e6c/6a170d6ca929cf8022ae0a3b/1dfbc2323438feaaa42e13ab242dd1f7166f74aa-1200x676.png" alt="Flow diagram showing an Elasticsearch index feeding semantic search results into an inference endpoint, which returns article recommendations." /><p>In this article, we’ll test <strong>SmolLM3-3B </strong>capacity tocombine its compact size with strong multilingual reasoning and tool-calling capabilities. Based on a search query, we’ll send all the matching content (in English and Spanish) to the LLM to generate a list of recommended articles with a custom-made description based on the search query and results.</p><p>Here’s what the UI of an article site with an AI recommendations generation system could look like.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt20e69b9a06fecd65/6a170d6e839dfa6f97dcff58/8d3b86b212f28ff279f2da67a33e6134039f0e4e-1999x949.png" alt="UI of an article site with an AI recommendations generation system, listing three examples, with text in English and titles in either English or Spanish." /><p>You can find the full implementation of this application in the linked <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/notebook.ipynb">notebook</a>.</p><h3>Configuring Elasticsearch inference endpoints</h3><p>To use the Elasticsearch <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put-hugging-face">Hugging Face inference endpoint</a>, we need two important elements: a Hugging Face API key and a running Hugging Face endpoint URL. It should look like this:</p>PUT _inference/chat_completions/hugging-face-smollm3-3b
{
    "service": "hugging_face",
    "service_settings": {
        "api_key": "hugging-face-access-token", 
        "url": "url-endpoint" 
    }
}<p>The Hugging Face inference endpoint in Elasticsearch supports different task types: <code>text_embedding</code>, <code>completion</code>, <code>chat_completion</code>, and <code>rerank</code>. In this blog post, we use <code>chat_completion</code> because we need the model to generate conversational recommendations based on the search results and a system prompt.This endpoint allows us to perform chat completions directly from Elasticsearch in a simple way using the Elasticsearch API:</p>POST _inference/chat_completion/hugging-face-smollm3-3b/_stream
{
  "messages": [
      { "role": "user", "content": "&lt;user prompt&gt;" }
  ]
}<p>This will serve as the core of the application, receiving the prompt and the search results that will pass through the model. With the theory covered, let’s start implementing the application.</p><h4>Setting up ​​inference endpoint on Hugging Face</h4><p>To deploy the Hugging Face model, we’re going to use <a href="https://huggingface.co/inference-endpoints/dedicated">Hugging Face one-click deployments</a>, an easy and fast service for deploying model endpoints. Keep in mind that this is a paid service, and using it may incur additional costs. This step will create the model instance that will be used to generate the recommendations of the articles.</p><p>You can pick a model from the one-click catalog:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta7bdfa43d6766324/6a170d6fb339d59e5476a039/b816e9fba1fe172687bf58f5143fb1f838c1077f-549x331.png" alt="Interface view of a model catalog filtered to “smoll3,” showing one model named “smollm3‑3b” with text generation, vLLM, GPU 1× Nvidia L4, and a listed price of $0.8, plus a note suggesting extending the search to all Hugging Face models." /><p>Let’s pick the <strong>SmolLM3-3B</strong> model:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdb0a2e6ffd7deb20/6a170d710c48574b7401aafc/610d3aba0429f3666c2df3616d513eb6a4397c0c-502x478.png" alt="Interface for creating an endpoint for the SmolLM3‑3B model, showing the model name, a &quot;verified by Hugging Face&quot; note, an endpoint name field, a cost of $0.80 per hour per running replica, a cURL option, and a &quot;Create Endpoint&quot; button." /><p>From here, grab the Hugging Face endpoint URL:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt25714021711ed6ff/6a170d72c1e8a54853f88336/025094ddb2cfbd1f0f216a5ec4e119b0f4fa2c42-646x328.png" alt="Dashboard view of a Hugging Face inference endpoint named “smollm3‑3b‑pnz,” showing a green Running status, one active replica, zero requests in the last hour, navigation tabs, and the displayed endpoint URL." /><p>As mentioned in the Elasticsearch <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-inference-put-hugging-face">Hugging Face inference endpoints documentation</a>, text generation requires a model that’s compatible with the OpenAI API. For that reason, we need to append the <code>/v1/chat/completions</code> subpath to the Hugging Face endpoint URL. The final result will look like this:</p>https://j2g31h0futopfkli.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions<p>With this in place, we can start coding in a Python notebook.</p><h4>Generating Hugging Face API key</h4><p>Create a <a href="https://huggingface.co/join">Hugging Face account</a>, and obtain an API token by following <a href="https://huggingface.co/docs/hub/en/security-tokens#user-access-tokens">these instructions</a>. You can choose between three token types: <em>fine-grained</em> (recommended for production, as it provides access only to specific resources); <em>read</em> (for read-only access); or <em>write</em> (for read and write access). For this tutorial, a read token is sufficient, since we only need to call the inference endpoint. Save this key for the next step.</p><h4>Setting up Elasticsearch inference endpoint</h4><p>First, let’s declare an Elasticsearch Python client:</p>os.environ["ELASTICSEARCH_API_KEY"] = "your-elasticsearch-api-key"
os.environ["ELASTICSEARCH_URL"] = "https://xxxx.us-central1.gcp.cloud.es.io:443"

es_client = Elasticsearch(
    os.environ["ELASTICSEARCH_URL"], api_key=os.environ["ELASTICSEARCH_API_KEY"]
)<p>Next, let’s create an Elasticsearch inference endpoint that uses the Hugging Face model. This endpoint will allow us to generate responses based on the blog posts and the prompt passed to the model.</p>INFERENCE_ENDPOINT_ID = "smollm3-3b-pnz"

os.environ["HUGGING_FACE_INFERENCE_ENDPOINT_URL"] = (
 "https://j2g31h0futopfkli.us-east-1.aws.endpoints.huggingface.cloud/v1/chat/completions"
)
os.environ["HUGGING_FACE_API_KEY"] = "hf_xxxxx"

resp = es_client.inference.put(
        task_type="chat_completion",
        inference_id=INFERENCE_ENDPOINT_ID,
        body={
            "service": "hugging_face",
            "service_settings": {
                "api_key": os.environ["HUGGING_FACE_API_KEY"],
                "url": os.environ["HUGGING_FACE_INFERENCE_ENDPOINT_URL"],
            },
        },
    )<h3>Dataset</h3><p>The dataset contains the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/dataset.json">blog posts</a> that will be queried, representing a multilingual content set used throughout the workflow:</p>// Articles dataset document example: 
{
    "id": "6",
    "title": "Complete guide to the new API: Endpoints and examples",
    "author": "Tomas Hernandez",
    "date": "2025-11-06",
    "category": "tutorial",
    "content": "This guide describes in detail all endpoints of the new API v2. It includes code examples in Python, JavaScript, and cURL for each endpoint. We cover authentication, resource creation, queries, updates, and deletion. We also explain error handling, rate limiting, and best practices. Complete documentation is available on our developer portal."
  }<h4>Elasticsearch mappings</h4><p>With the dataset defined, we need to create a data schema that properly fits the blog post structure. The following <a href="https://www.elastic.co/docs/manage-data/data-store/mapping">index mappings</a> will be used to store the data in Elasticsearch:</p>INDEX_NAME = "blog-posts"

mapping = {
    "mappings": {
        "properties": {
            "id": {"type": "keyword"},
            "title": {
                "type": "object",
                "properties": {
                    "original": {
                        "type": "text",
                        "copy_to": "semantic_field",
                        "fields": {"keyword": {"type": "keyword"}},
                    },
                    "translated_title": {
                        "type": "text",
                        "fields": {"keyword": {"type": "keyword"}},
                    },
                },
            },
            "author": {"type": "keyword", "copy_to": "semantic_field"},
            "category": {"type": "keyword", "copy_to": "semantic_field"},
            "content": {"type": "text", "copy_to": "semantic_field"},
            "date": {"type": "date"},
            "semantic_field": {"type": "semantic_text"},
        }
    }
}


es_client.indices.create(index=INDEX_NAME, body=mapping)<p>Here, we can see more clearly how the data is structured. We’ll use semantic search to retrieve results based on natural language, along with the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/copy-to"><code>copy_to</code></a> property to copy the field contents into the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_text</code></a> field. Additionally, the <code>title</code> field contains two subfields: the <code>original</code> subfield stores the title in either English or Spanish, depending on the original language of the article; and the <code>translated_title</code> subfield is present only for Spanish articles and contains the English translation of the original title.</p><h3>Ingesting data</h3><p>The following code snippet ingests the blog posts dataset into Elasticsearch using the <a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript/bulk_examples">bulk API</a>:</p>def build_data(json_file, index_name):
    with open(json_file, "r") as f:
        data = json.load(f)

    for doc in data:
        action = {"_index": index_name, "_source": doc}
        yield action


try:
    success, failed = helpers.bulk(
        es_client,
        build_data("dataset.json", INDEX_NAME),
    )
    print(f"{success} documents indexed successfully")

    if failed:
        print(f"Errors: {failed}")
except Exception as e:
    print(f"Error: {str(e)}")<p>Now that we have the articles ingested into Elasticsearch, we need to create a function capable of searching against the <code>semantic_text</code> field:</p>def perform_semantic_search(query_text, index_name=INDEX_NAME, size=5):
    try:
        query = {
            "query": {
                "match": {
                    "semantic_field": {
                        "query": query_text,
                    }
                }
            },
            "size": size,
        }

        response = es_client.search(index=index_name, body=query)
        hits = response["hits"]["hits"]

        return hits
    except Exception as e:
        print(f"Semantic search error: {str(e)}")
        return []<p>We also need a function that calls the inference endpoint. In this case, we’ll call the endpoint using the <strong><code>chat_completion</code></strong>task type to get streaming responses:</p>def stream_chat_completion(messages: list, inference_id: str = INFERENCE_ENDPOINT_ID):
    url = f"{ELASTICSEARCH_URL}/_inference/chat_completion/{inference_id}/_stream"
    payload = {"messages": messages}
    headers = {
        "Authorization": f"ApiKey {ELASTICSEARCH_API_KEY}",
        "Content-Type": "application/json",
    }

    try:
        response = requests.post(url, json=payload, headers=headers, stream=True)
        response.raise_for_status()

        for line in response.iter_lines(decode_unicode=True):
            if line:
                line = line.strip()

                if line.startswith("event:"):
                    continue

                if line.startswith("data: "):
                    data_content = line[6:]

                    if not data_content.strip() or data_content.strip() == "[DONE]":
                        continue

                    try:
                        chunk_data = json.loads(data_content)

                        if "choices" in chunk_data and len(chunk_data["choices"]) &gt; 0:
                            choice = chunk_data["choices"][0]
                            if "delta" in choice and "content" in choice["delta"]:
                                content = choice["delta"]["content"]
                                if content:
                                    yield content

                    except json.JSONDecodeError as json_err:
                        print(f"\nJSON decode error: {json_err}")
                        print(f"Problematic data: {data_content}")
                        continue

    except requests.exceptions.RequestException as e:
        yield f"Error: {str(e)}"<p>Now we can write a function that calls the semantic search function, along with the <code>chat_completions</code> inference endpoint and the recommendations endpoint, to generate the data that will be allocated in the cards:</p>def recommend_articles(search_query, index_name=INDEX_NAME, max_articles=5):
    print(f"\n{'='*80}")
    print(f"🔍 Search Query: {search_query}")
    print(f"{'='*80}\n")

    articles = perform_semantic_search(search_query, index_name, size=max_articles)

    if not articles:
        print("❌ No relevant articles found.")
        return None, None

    print(f"✅ Found {len(articles)} relevant articles\n")

    # Build context with found articles
    context = "Available blog articles:\n\n"
    for i, article in enumerate(articles, 1):
        source = article.get("_source", article)
        context += f"Article {i}:\n"
        context += f"- Title: {source.get('title', 'N/A')}\n"
        context += f"- Author: {source.get('author', 'N/A')}\n"
        context += f"- Category: {source.get('category', 'N/A')}\n"
        context += f"- Date: {source.get('date', 'N/A')}\n"
        context += f"- Content: {source.get('content', 'N/A')}\n\n"

    system_prompt = """You are an expert content curator that recommends blog articles.

    Write recommendations in a conversational style starting with phrases like:
    - "If you're interested in [topic], this article..."
    - "This post complements your search with..."
    - "For those looking into [topic], this article provides..."


    FORMAT REQUIREMENTS:
    - Return ONLY a JSON array
    - Each element must have EXACTLY these three fields: "article_number", "title", "recommendation"
    - If the original title is in spanish, use the "translated_title" subfield in the "title" field

    Keep each recommendation concise (2-3 sentences max) and focused on VALUE to the reader.

    EXAMPLE OF CORRECT FORMAT:
    [
        {"article_number": 1, "title": "Article title in english", "recommendation": "If you are interested in [topic], this article provides..."},
        {"article_number": 2, "title": "Article title in english", "recommendation": " for those looking into [topic], this article provides..."}
    ]

    Return ONLY the JSON array following this exact structure."""

    user_prompt = f"""Search query: "{search_query}"

    Generate recommendations for the following articles: {context}
    """

    messages = [
        {"role": "system", "content": "/no_think"},
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_prompt},
    ]

    # LLM generation
    print(f"{'='*80}")
    print("🤖 Generating personalized recommendations...\n")

    full_response = ""

    for chunk in stream_chat_completion(messages):
        print(chunk, end="", flush=True)
        full_response += chunk

    return context, articles, full_response<p>Finally, we need to extract the information and format it to be printed:</p>def display_recommendation_cards(articles, recommendations_text):
    print("\n" + "=" * 100)
    print("📇 RECOMMENDED ARTICLES".center(100))
    print("=" * 100 + "\n")

    # Parse JSON recommendations - clean tags and extract JSON
    recommendations_list = []
    try:

        # Clean up &lt;think&gt; tags
        cleaned_text = re.sub(
            r"&lt;think&gt;.*?&lt;/think&gt;", "", recommendations_text, flags=re.DOTALL
        )
        # Remove markdown code blocks ( ... ``` or ``` ... ```)
        cleaned_text = re.sub(r"```(?:json)?", "", cleaned_text)
        cleaned_text = cleaned_text.strip()

        parsed = json.loads(cleaned_text)

        # Extract recommendations from list format
        for item in parsed:
            article_number = item.get("article_number")
            title = item.get("title", "")
            rec_text = item.get("recommendation", "")

            if article_number and rec_text:
                recommendations_list.append(
                    {
                        "article_number": article_number,
                        "title": title,
                        "recommendation": rec_text,
                    }
                )
    except json.JSONDecodeError as e:
        print(f"⚠️  Could not parse recommendations as JSON: {e}")
        return

    for i, article in enumerate(articles, 1):
        source = article.get("_source", article)

        # Card border
        print("┌" + "─" * 98 + "┐")

        # Find recommendation and title for this article number
        recommendation = None
        title = None
        for rec in recommendations_list:
            if rec.get("article_number") == i:
                recommendation = rec.get("recommendation")
                title = rec.get("title")
                break

        # Print title
        title_lines = textwrap.wrap(f"📌 {title}", width=94)
        for line in title_lines:
            print(f"│  {line}".ljust(99) + "│")

        # Card border
        print("├" + "─" * 98 + "┤")

        # Print recommendation
        if recommendation:
            recommendation_lines = textwrap.wrap(recommendation, width=94)
            for line in recommendation_lines:
                print(f"│  {line}".ljust(99) + "│")

        # Card bottom
        print("└" + "─" * 98 + "┘")<p>Let’s test this by asking a question about the security blog posts:</p>search_query = "Security and vulnerabilities"

context, articles, recommendations = recommend_articles(search_query)

print("\nElasticsearch context:\n", context)

# Display visual cards
display_recommendation_cards(articles, recommendations)<p>Here we can see the cards in the console generated by the workflow:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4aa221a08a51aeb3/6a170d7460084be1413c45d6/730d35212594bb3db30447c3ea7e2a92857287b7-1999x1515.png" alt="Section titled “Recommended Articles” showing five boxed article summaries, including topics on an authentication system vulnerability, migration risks, REST API v2 performance and authentication improvements, notification system changes, and a complete guide to the new API." /><p>You can see the full results, including all hits and the LLM response, in <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-inference-api-and-hugging-face/results.md">this file</a>.</p><p>We’re asking for articles related to: “Security and vulnerabilities.” This question is used as the search query against the documents stored in Elasticsearch. The retrieved results are then passed to the model, which generates recommendations based on their content. As we can see, the model did a great job generating engaging short text that can motivate the reader to click on it.</p><h2>Conclusion</h2><p>This example shows how Elasticsearch and Hugging Face can be combined to create a fast and efficient centralized system for AI applications. This approach reduces manual effort and provides flexibility, thanks to Hugging Face’s extensive model catalog. Using SmolLM3-3B, in particular, shows how compact, multilingual models can still deliver meaningful reasoning and content generation when paired with semantic search. Together, these tools offer a scalable and effective foundation for building intelligent content analysis and multilingual applications.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/hugging-face-elasticsearch-inference-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/hugging-face-elasticsearch-inference-api</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Integrations]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f961af4cb26ec97/6a170d767d8d6790c770e790/1417d6ff033712206c9bd4bcc22074ee3437ce96-1999x1125.png" length="0" type="image/png"/>
    <pubDate>Mon, 23 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[SearchClaw: Bring Elasticsearch to OpenClaw with composable skills]]></title>
    <description><![CDATA[Give your local AI agent access to Elasticsearch data using OpenClaw, composable skills, and agents, no custom code required.]]></description>
    <content:encoded><![CDATA[<p>In recent weeks, <a href="https://openclaw.ai/">OpenClaw</a> has been appearing frequently in AI community discussions, particularly among developers interested in agents, automation, and local runtimes. The project gained traction quickly, which naturally raised a technical question:</p><p><em>What real problem does it solve for engineers?</em></p><p><strong>OpenClaw</strong> is a self-hosted gateway for AI agents: a single runtime that coordinates execution, treats agents as isolated processes, and uses skills (structured instructions in markdown files) as the unit of integration. Conceptually, this isn’t entirely different from what we already do with command line interfaces (CLIs) and scripts, but it’s now formalized around agent-driven workflows.</p><p>This led to a practical exploration within the Elastic Stack:</p><p><em>If we treat OpenClaw as an orchestration runtime, how does it behave when Elasticsearch is the back end? And how straightforward is integration using OpenClaw skills?</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8eb543674b064eee/6a170e562b835f4412f4b2bc/ec61e65f54b96b83975b52b2d88305170001d9bd-1999x1445.png" alt="Chart showing GitHub star history, from 2020 to 2026, for four different open‑source automation and AI‑agent frameworks: OpenClaw, LangChain, CrewAI, and n8n-io." /><p>Let's build an integration using composable skills.</p><h2><strong>Solution architecture</strong></h2><p>In this tutorial, we’ll teach OpenClaw how to access and query Elasticsearch data through a custom read-only skill, and we’ll then demonstrate how it composes multiple skills together; for example, combining Elasticsearch queries with real-time weather data to generate dynamic reports.</p><p>Before diving into the hands-on steps, let’s look at what we’re building. The solution is composed of three integrated layers that work together through OpenClaw orchestration.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3b365123e76e74dd/6a170e587d8d67349c70e7d2/ca8dc124a7410ba036ddf887eee011c42125cdf3-1270x680.png" alt="SearchClaw (OpenClaw and Elasticsearch) solution architecture, with the OpenClaw Gateway Runtime as the central hub. It loads skills for context and then interacts directly with each back end." /><h3>Layer 1: Storage and search (Elasticsearch)</h3><p>The data layer runs on Elasticsearch via <a href="https://github.com/elastic/start-local"><code>start-local</code></a>, a single command that spins up Elasticsearch and Kibana locally with Docker.</p><p>Two sample indices demonstrate different use cases:</p><ul><li><p><strong><code>fresh_produce</code></strong><strong>:</strong> 10 products with semantic search (ecommerce scenario)</p></li><li><p><strong><code>app-logs-synthetic</code></strong><strong>:</strong> 30 log entries across four services (observability scenario)</p></li></ul><p>The same read-only skill works with both indices without any reconfiguration; the agent inspects the mapping and adapts its queries accordingly.</p><h3>Layer 2: Orchestration (OpenClaw Gateway)</h3><p>The gateway receives natural language requests and loads the Elasticsearch skill, and the large language model (LLM) decides which queries to construct. The skill is a pure <strong><code>SKILL.md</code></strong> with reference docs, meaning that its operations require no custom code.</p><p>To understand how the gateway organizes this, two core OpenClaw concepts are worth knowing:</p><ul><li><p><strong>Agents:</strong> Independent AI instances, each with its own configuration, workspace, and set of skills. You can run multiple agents for different purposes.</p></li><li><p><strong>Workspace:</strong> A folder that defines an agent’s context:<strong><code>AGENTS.md</code></strong> (the agent’s permanent briefing), <strong><code>.env</code></strong>(credentials), and a <strong><code>skills/</code></strong> directory. Think of it as the agent’s working environment.</p></li></ul><h3>Layer 3: Skills (composable capabilities)</h3><p>Skills are structured instructions in markdown files (<code>SKILL.md</code>) that teach the agent how to use specific tools or APIs. They can be global (available to all agents), workspace-specific, or bundled with OpenClaw. The agent selectively loads only the skills relevant to each request.</p><p>This tutorial uses two skills:</p><ul><li><p><strong><code>Elasticsearch-openclaw</code></strong><strong> (custom, built for this tutorial):</strong> A read-only skill that teaches the agent how to search, filter, aggregate, and explore Elasticsearch indices using curl.</p></li><li><p><strong><code>Weather</code></strong><strong> (community skill, used for composition demo):</strong> A skill that fetches current weather conditions from external APIs.</p></li></ul><p>Later in the tutorial, we'll demonstrate how OpenClaw composes both skills in a single request, querying Elasticsearch products based on real-time weather data without any custom integration code.</p><h4>Read-only by design</h4><p>The <code>elasticsearch-openclaw</code> skill is <strong>read-only by design</strong>. It provides patterns for searching, filtering, and aggregating data, but it never writes, updates, or deletes. This minimizes the security footprint when giving AI agents access to your Elasticsearch cluster.</p><p>Even if the agent environment is compromised, your data remains safe from modification or deletion. This is enforced through:</p><ul><li><p><strong>Skill design:</strong> No write operation patterns in <code>SKILL.md</code> or reference files.</p></li><li><p><strong>API key permissions:</strong> The tutorial uses a read-only API key with only <code>read</code> and <code>view_index_metadata</code> privileges.</p></li><li><p><strong>Agent instructions:</strong> <code>AGENTS.md</code> explicitly states "You can SEARCH, FILTER, and AGGREGATE data, but you can NEVER write, update, or delete."</p></li></ul><p>This security-first approach is why infrastructure setup (index creation, data loading) must be done manually; by design, the agent cannot do it for you.</p><h2><strong>Prerequisites</strong></h2><p>To follow this tutorial, you’ll need:</p><p><strong>Software and tools:</strong></p><ul><li><p>Docker Desktop installed and running (Docker Engine with Compose V2).</p></li><li><p>Elasticsearch running locally via <code>start-local</code>. (We’ll set this up in the next section.)</p></li><li><p>Jina API key (free): <a href="https://jina.ai/embeddings">https://jina.ai/embeddings</a>.</p></li><li><p>OpenClaw installed: <a href="https://openclaw.ai">https://openclaw.ai</a>.</p></li></ul><h3><strong>Setting up the environment</strong></h3><p>Start by cloning the starter project, which contains the skill, workspace configuration, and Dev Tools scripts:</p>git clone https://github.com/salgado/elasticsearch-openclaw-start-blog
cd elasticsearch-openclaw-start-blog<p>The repository contains:</p>elasticsearch-openclaw-start-blog/
├── devtools_fresh_produce.md         ← Creates fresh_produce index (10 products)
├── devtools_app_logs_synthetic.md    ← Creates app-logs-synthetic index (30 logs)
└── openclaw-workspace-elastic-blog/
    ├── AGENTS.md                      ← Agent briefing
    ├── .env.example                   ← Credentials template<p><em><strong>Note:</strong></em><em> The </em><em><code>devtools*.md</code></em><em> files contain Kibana Dev Tools commands formatted as reference documentation.</em></p><h4>Installing OpenClaw</h4><p>OpenClaw is a self-hosted gateway. This means you maintain full control over execution and data, but you need to prepare your local environment or server.</p><p>I installed OpenClaw on a separate machine, which is why I included the disclaimer below.</p><p><strong>** Security and responsibility disclaimer **</strong></p><p>Since OpenClaw is an early-stage, rapidly evolving open-source project, the community has raised important discussions about potential security vulnerabilities, especially around token handling and third-party script execution.</p><p><strong>Deployment recommendations:</strong></p><ul><li><p><strong>Isolated environments:</strong> If you’re not an advanced infrastructure security user, we recommend installing OpenClaw strictly in isolated, controlled environments (such as a dedicated virtual machine [VM], a rootless Docker container, or a test machine).</p></li><li><p><strong>Do not use in production:</strong> Avoid running the gateway on servers containing sensitive data or with unrestricted access to your corporate network until the project reaches a more stable, audited version.</p></li><li><p><strong>Least privilege:</strong> We reinforce the need to use Elasticsearch API keys with restricted permissions (read-only) to mitigate risks, in case the environment is compromised.</p></li><li><p><strong>Network segmentation:</strong> Both Elasticsearch and OpenClaw bind to <code>localhost</code> by default. Keep it that way, unless you have a specific reason to expose them.</p></li><li><p><strong>Credential rotation:</strong> Rotate API keys periodically. OpenClaw stores credentials locally, so treat the machine’s security as the perimeter.</p></li><li><p><strong>Audit logging:</strong> Enable Elasticsearch audit logging to track all API calls made by OpenClaw. This creates a full trail of what the agent accessed and when.</p></li><li><p><strong>Keep the installation up to date.</strong></p></li></ul><p>For a deeper analysis of the security architecture and deployment options, consult the <a href="https://docs.openclaw.ai">official OpenClaw documentation</a>.</p><h4>Runtime installation</h4><p>OpenClaw manages daemons and skill isolation via CLI. Since it’s a recent project that has undergone naming changes, we recommend strictly following the <a href="https://docs.openclaw.ai/install">official documentation</a> to ensure installation compatibility.</p># Global gateway installation
curl -fsSL https://openclaw.ai/install.sh | bash<h2><strong>Preparing the Elasticsearch back end</strong></h2><p>Before connecting any agent runtime, we need a working Elasticsearch environment with data to query and a secure, <strong>read-only access layer</strong>. In the next two sections, we’ll spin up Elasticsearch locally using <code>start-local</code>, create an index with <code>semantic_text</code> and Jina v5 embeddings, load sample data, validate that semantic search works, and generate a read-only API key. Once this foundation is in place, the Elasticsearch side is complete and we can focus entirely on teaching the agent how to use it.</p><h3>Part 1: Setting up Elasticsearch locally</h3><p>Start a local Elasticsearch and Kibana instance with a single command:</p>curl -fsSL https://elastic.co/start-local | sh<p>Once complete: Elasticsearch at <code>http://localhost:9200</code>, Kibana at <code>http://localhost:5601</code>, and credentials in <code>elastic-start-local/.env</code>.</p><h3>Part 2: Configuring the index in Kibana Dev Tools</h3><p>Open <code>http://localhost:5601</code> → Dev Tools and run <code>devtools_fresh_produce.md</code> in order.</p><ul><li><p><strong>Step 1:</strong> Replace <code>YOUR_JINA_API_KEY</code> with your actual Jina API key (free).</p></li><li><p><strong>Step 2:</strong> Save the encoded field immediately; it cannot be retrieved later.</p></li></ul><p>The key commands in the Dev Tools file are:</p><p><strong>Create the Jina inference endpoint:</strong></p>PUT _inference/text_embedding/jina-embeddings-v5
{
  "service": "jinaai",
  "service_settings": {
    "api_key": "YOUR_JINA_API_KEY",
    "model_id": "jina-embeddings-v5-text-small"
  }
}<p><strong>Create the index with </strong><strong><code>semantic_text</code></strong><strong>:</strong></p>PUT /fresh_produce
{
  "mappings": {
    "properties": {
      "name": {
        "type": "text",
        "fields": { "keyword": { "type": "keyword" } }
      },
      "description": { "type": "text" },
      "category": { "type": "keyword" },
      "price": { "type": "float" },
      "stock_kg": { "type": "float" },
      "on_sale": { "type": "boolean" },
      "image_url": { "type": "keyword" },
      "semantic_content": {
        "type": "semantic_text",
        "inference_id": "jina-embeddings-v5"
      }
    }
  }
}<p>The <code>semantic_text</code> field type handles embedding generation automatically at index time.</p><p><strong>Index sample products</strong> using the bulk API (see <code>devtools_fresh_produce.md</code> for the full dataset of 10 products).</p><p><strong>Validate semantic search:</strong></p>GET /fresh_produce/_search
{
  "query": {
    "semantic": {
      "field": "semantic_content",
      "query": "healthy colorful meals"
    }
  },
  "size": 3,
  "_source": ["name", "description", "category"]
}<p>The semantic query type handles inference on the query side automatically; no need to specify model IDs or embedding details.</p><p><strong>Create a read-only API key:</strong></p>POST /_security/api_key
{
  "name": "openclaw-readonly",
  "role_descriptors": {
    "reader": {
      "cluster": ["monitor"],
      "indices": [
        {
          "names": ["fresh_produce", "app-logs-synthetic"],
          "privileges": ["read", "view_index_metadata"]
        }
      ]
    }
  }
}<p>Save the encoded value from the response. This is your API key for the OpenClaw configuration.</p><h2>Connecting to OpenClaw</h2><p>With the Elasticsearch back end ready, we can now wire it into OpenClaw. Several Elasticsearch integrations already exist in the ecosystem, from <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/mcp-server">Elastic’s own Model Context Protocol (MCP) server</a> to community-built MCP servers. However, most of these offer full CRUD access or are designed for different agent runtimes. Given that the technology is still in its early stages and security remains a primary concern, I chose to build a dedicated skill, simple, read-only, and purpose-built for OpenClaw. This approach ensures that the agent can search, filter, and aggregate data but never modify it, keeping the blast radius minimal even if the environment is compromised.</p><p>In the next sections, we’ll configure credentials, install the skill, create a dedicated agent, and explore how the workspace ties everything together.</p><h3>Install the skill and create the agent</h3><h4>Step 1: Configure credentials</h4><p>From the cloned repository, configure the credentials by copying the environment template and filling in your Elasticsearch URL and the read-only API key:</p>cp openclaw-workspace-elastic-blog/.env.example 
openclaw-workspace-elastic-blog/.env<p>Edit the .env file with these two values:</p>ELASTICSEARCH_URL: http://localhost:9200 (from start-local)
ELASTICSEARCH_API_KEY: The encoded value from the read-only API key you created in Part 2 (the POST /_security/api_key response)<p>Example .env file:</p>ELASTICSEARCH_URL=http://localhost:9200
ELASTICSEARCH_API_KEY=VnVaRmxLSDRCQxxxxxxxxbGVfa2V5<h4>Step 2: Install the skill from ClawHub</h4><p><a href="https://clawhub.ai/">ClawHub</a> is OpenClaw's public skill registry. Think of it as npm for AI agent skills. At the time of this writing, ClawHub hosts over 3,200 skills, covering everything from Slack and GitHub integrations to Internet of Things (IoT) device automation. For this tutorial, we created <code>elasticsearch-openclaw</code>, a custom skill focused on read-only queries using <code>semantic_text</code>, aggregations, and observability on Elasticsearch 9.x. It’s published on ClawHub so you can install it directly. As a best practice, only install skills from trusted sources with known provenance; as with any package manager, review the content before granting access to your agent.</p><p>The <code>elasticsearch-openclaw</code> skill is published on ClawHub.</p><p><strong>Recommended:</strong> Open the OpenClaw Web UI (http://127.0.0.1:18789/) and ask:</p>Install the elasticsearch-openclaw skill from https://clawhub.ai/salgado/elasticsearch-openclaw<p>OpenClaw will:</p><ul><li><p>Fetch the skill from ClawHub.</p></li><li><p>Install it in the appropriate directory.</p></li><li><p>Confirm when ready to use.</p></li></ul><h4>Step 3: Create the agent</h4><p>Do this by registering a dedicated agent with its own workspace, and then restart the gateway to load the new configuration:</p>openclaw agents add elasticsearch-agent \
  --workspace ~/path/to/elasticsearch-openclaw-start-blog/openclaw-workspace-elastic-blog \
  --non-interactive

openclaw gateway restart<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltffac27a4e3fd8fe9/6a170e5a964ceac34508bc5d/abc051a513b0cc7dff4a7f02493d51e220c72ad4-1999x1095.png" alt="OpenClaw web chat screen, with the focus on &quot;Find products, in my Elasticsearch, that would be good for a fresh salad.&quot;" /><h3>Understanding the workspace</h3><p>Now that the agent is running, let’s look at what makes it tick.</p><h4><code>AGENTS.md</code></h4><p>The <code>AGENTS.md</code> file is the agent’s permanent briefing. It defines who the agent is, what it can do, and how it should behave. For our Elasticsearch agent, this file instructs the agent about the available indices, the read-only constraint, and the preferred query patterns.</p><h4>Skills: When they make a difference</h4><p>Without skill</p><p>With `elasticsearch-openclaw` skill</p><p>Agent has no knowledge of Elasticsearch query syntax.</p><p>Agent knows semantic, full-text, filtered, and aggregation patterns.</p><p>Agent might attempt write operations.</p><p>Agent is instructed to never write, update, or delete.</p><p>Agent guesses field names and types.</p><p>Agent inspects mappings first and then constructs appropriate queries.</p><p>Generic curl commands with trial and error.</p><p>Structured query templates with best practices for Elasticsearch 9.x.</p><h2><strong>Exploring with the agent</strong></h2><p>With the Elasticsearch back end configured and the OpenClaw agent connected, it’s time to see what the agent can actually do. In the next sections, we’ll test natural language queries, explore observability data, and compose multiple skills together.</p><h3><strong>Testing in OpenClaw</strong></h3><p>Open the OpenClaw web UI, and try some natural language queries. The agent will inspect the index mapping, choose the appropriate query type, and return results.</p><p>Type:</p>“Find products that would be good for a healthy summer salad.”<p>Result:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a0015d9a2d3abfd/6a170e5b1949f754dce7aad9/d5b4bbe71ad56af5462bccc1475bd10d5233abd9-1011x557.png" alt="OpenClaw web chat page with &quot;Semantic search working&quot; message, along with a list of salad ingredients." /><p>Others ideas to explore:</p><ul><li><p><strong>Index exploration:</strong> &gt; “What indices do I have in Elasticsearch? Show me the fields of <code>fresh_produce</code>.”</p></li><li><p><strong>Filtered search:</strong> &gt; “Show me all products on sale under $15.”</p></li><li><p><strong>Aggregations:</strong> &gt; “What’s the average price by category?”</p></li></ul><h3>Observability</h3><p>To demonstrate that the skill works beyond a single use case, the repository includes a second index: <code>app-logs-synthetic</code>, with 30 synthetic log entries across four fictional services, created from <code>devtools_app_logs_synthetic.md</code>.</p><h4>Setting up the log data</h4><p>Since the skill is read-only, you need to populate the index first. The <code>devtools_app_logs_synthetic.md</code> file contains <strong>five commands</strong> (three for setup and two for verification):</p><ul><li><p><strong><code>Create ingest pipeline</code></strong><strong>:</strong> Adds @timestamp to log entries automatically.</p></li><li><p><strong><code>Create index mapping</code></strong><strong>:</strong> Defines the <code>app-logs-synthetic</code> structure (classic fields only, no <code>semantic_text</code>).</p></li><li><p><strong><code>Bulk insert logs</code></strong><strong>:</strong> Loads 30 synthetic log entries across four services.</p></li><li><p><strong><code>Count query</code></strong><strong>:</strong> Verify 30 documents were indexed.</p></li><li><p><strong><code>Sample search</code></strong><strong>:</strong> Quick test to confirm that data is queryable.</p></li></ul><h4>How to run:</h4><ol><li><p>Open Kibana Dev Tools: http://localhost:5601 → Dev Tools.</p></li><li><p>Copy each numbered block from the .md file.</p></li><li><p>Paste into the Dev Tools console.</p></li><li><p>Press <em><strong>Ctrl/Cmd+Enter</strong></em> to execute.</p></li><li><p>Wait for a successful response before continuing to the next block.</p></li></ol><p>This creates the <code>app-logs-synthetic</code> index with sample data ready for querying.</p><p>Try this query in the OpenClaw web UI:</p>Show me the distribution of HTTP status codes across all services.<p>Result:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8600731ee46ee0f8/6a170e5d961e69607fc4cfae/d35fc1c0ea6d647f1c85163eb0ab8e268c6c4f89-1002x565.png" alt="OpenClaw web chat with &quot;the full picture across 30 logs,&quot; listing &quot;ok,&quot; &quot;bad requests,&quot; &quot;server errors,&quot; and more." /><p>Other ideas to explore:</p><ul><li><p>“How many 500 errors do I have in <code>app-logs-synthetic</code>? Which services are failing?”</p></li><li><p>“Which endpoints have the slowest response times?”</p></li><li><p>“What happened with the <code>payment-service</code> in the last 24 hours?”</p></li></ul><p>This is the same skill, same agent, same setup, just pointed at different data. The agent inspects the new index mapping, adapts its queries, and returns relevant results without any reconfiguration.</p><h2><strong>Composing skills in action</strong></h2><p>This is where composable skills truly shine. Start by asking the agent:</p>Install the weather skill.<p>OpenClaw will search for the weather skill, automatically attempt the installation, and guide you through the process. Just follow the on-screen instructions; no new API key is required for the weather skill. Afterward, try this:</p>“Find the products on sale in the fresh_produce index that match today’s weather in São Paulo. Generate a nice HTML report with product cards using the image_url field from each document, price, description, and stock. Save it to ~/Desktop/report.html and open it in the browser.”<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb362465ddf7556fe/6a170e5f509168515ce1bb8a/14fa4303bb2f1eb19530d8844f09c99948b3c752-1965x1079.png" alt="SearchClaw results for &quot;products on sale that match today's weather,&quot; including images of watermelon and avocado." /><p>In a single request, the agent chains multiple skills: the <strong>weather skill</strong> to check current conditions, the <strong>Elasticsearch skill </strong>to run a hybrid search on products that match the context, and its built-in file and browser tools to generate an HTML report and open it. No custom integration code, no glue scripts, just skills composed by the LLM at runtime.</p><p>This is what makes OpenClaw different from a traditional automation framework. You don’t preprogram the workflow. You describe the outcome, and the agent figures out the composition.</p><h2><strong>Conclusion</strong></h2><p>SearchClaw started as a simple experiment and ended up demonstrating what composable, LLM-driven integration looks like in practice. The key takeaway is not the individual tools (all are familiar) but the approach. Instead of writing a specific application with hardcoded queries, we gave the agent capabilities and let it compose solutions dynamically. This is what makes OpenClaw native: composable, LLM-driven, and local-first.</p><p>As with any early-stage project, OpenClaw should be used thoughtfully, especially regarding security and environment isolation. The read-only skill approach demonstrated here is one way to limit risk while still unlocking the value of your Elasticsearch data.</p><p>The full code is available in the repository and can serve as a starting point for your own integrations: <a href="https://github.com/salgado/elasticsearch-openclaw-start-blog">https://github.com/salgado/elasticsearch-openclaw-start-blog</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/openclaw-elasticsearch-ai-agents</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/openclaw-elasticsearch-ai-agents</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Integrations]]></category>
    <dc:creator><![CDATA[Alex Salgado]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltee8bba6ac4830abd/6a170e60cdacbf48277d2a92/ce3248c3cb7a352e3fdafef4ac8116ab998ab4f4-1950x1137.png" length="0" type="image/png"/>
    <pubDate>Tue, 10 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Does MCP make search obsolete? Not even close]]></title>
    <description><![CDATA[Explore why search engines and indexed search remain the foundation for scalable, accurate, enterprise-grade AI, even in the age of MCP, federated search, and large context windows.]]></description>
    <content:encoded><![CDATA[<p>With the rise of large language models (LLMs), agent frameworks, and new protocols like Model Context Protocol (MCP), a provocative question is starting to surface:</p><strong>Do we still need a search engine at all?</strong><p>If agents can call tools on demand and models can reason over massive context windows, why not just fetch data live from every system and let the LLM figure it out?</p><p>It’s a reasonable question. It’s also the wrong conclusion.</p><p>The reality is that MCP and agent tooling don’t eliminate the need for search. They make the quality of search <strong>more critical than ever</strong>. In this blog, we’ll explore why MCP, federated search, and large context windows don’t replace search engines and why indexes remain the foundational layer for scalable, accurate, enterprise-grade AI.</p><h2><strong>What MCP actually is (and what it is not)</strong></h2><p>MCP is a <strong>coordination protocol</strong>. It standardizes how an agent requests information or actions from external systems.</p><p>What MCP <em>doesn’t</em> do:</p><ul><li><p>Rank results across systems.</p></li><li><p>Understand relevance across heterogeneous data.</p></li><li><p>Normalize schemas or metadata.</p></li><li><p>Data transformations or enrichments at scale.</p></li><li><p>Apply consistent security and permissions.</p></li><li><p>Optimize for latency, cost, or scale.</p></li></ul><p>In other words, <strong>MCP tells agents </strong><em><strong>how</strong></em><strong> to ask for data, not </strong><em><strong>which</strong></em><strong> data matters most</strong>.</p><h2><strong>Modern retrieval requires query intelligence, not just data access</strong></h2><p>In modern enterprise search architectures, retrieval quality is determined long before a query reaches an index. Raw queries — especially those generated by agents — may be incomplete, overly literal, schema-driven rather than intent-driven, and at times syntactically invalid.</p><p>This is why mature search platforms introduce a query intelligence layer that performs query rewriting, entity normalization, synonym expansion, and intent disambiguation before retrieval even begins.</p><p>For example, an agent-generated request such as: “Show severity 2 authentication failures from last sprint” may be rewritten to include authentication synonyms (login, SSO, OAuth), normalized severity mappings, and sprint-to-date-range translation. The result is not just more matches — it is more <em>relevant</em> matches.</p><p>In enterprise AI, retrieval is not a single step. It is a controlled pipeline.</p><p>This distinction is crucial because once MCP-based agents start pulling information live from multiple tools, they recreate a familiar pattern under a new name: <strong>federated search</strong>.</p><h2><strong>MCP-based retrieval is federated search in disguise</strong></h2><p>Federated search isn’t new. Enterprises have tried it for decades.</p><p>The model is simple:</p><ul><li><p>Send the user’s query to multiple systems in parallel (SharePoint, GitHub, Jira, customer relationship management [CRM]).</p></li><li><p>Collect the responses.</p></li><li><p>Merge and present the results.</p></li></ul><p>MCP-driven tool calls follow the same pattern, except that the caller is now an agent instead of a user interface.</p><p>And the same problems resurface.</p><h2><strong>Why federated search breaks down at enterprise scale</strong></h2><ul><li><p><strong>Latency becomes unpredictable:</strong> A federated query is only as fast as its slowest system. Enterprise systems can have wildly different response times and rate limits, so federated queries tend to be <strong>slow and jittery</strong>. Agents must wait for multiple round trips before reasoning can even begin. The result is a laggy experience and unpredictable wait times.</p></li><li><p><strong>Relevance is fragmented:</strong> Because each system ranks results on its own, there’s no unified relevance model. Federated search <strong>cannot apply a single ranking or semantic understanding across all content</strong>, so results often seem disjointed or incomplete. Agents may retrieve <em>correct</em> information but not the <em>most useful</em> information.</p></li><li><p><strong>Context is shallow and incomplete: </strong>Federated systems typically expose only what’s directly accessible through an API call.They rarely surface:</p><ul><li><p>Usage signals, like clicks, dwell time, recency of access, popularity, or authority.</p></li><li><p>Relationships between documents across different systems to correlate the insights.</p></li><li><p>Organizational knowledge beyond a single silo.

This strips agents of the broader context required for high-quality reasoning.
</p></li></ul></li><li><p><strong>Limited filtering and features:</strong> In a federated setup, you can only filter on fields that every system supports (the “lowest common denominator”). If one system doesn’t support a particular filter or facet, you lose that functionality entirely. This severely limits rich search features, like date ranges, categories, or tags.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6f82101d3ef2019b/6a170ca96f7f04f0219148ac/25bb778f4da9a3cb4f0d4e10af66221b8af73900-1376x768.jpg" alt="Federated search workflow" /><h2><strong>The power of an indexed search</strong></h2><p>Search engines achieve millisecond-level retrieval at massive scale by using specialized data structures, including inverted indexes for lexical search and k‑dimensional trees (k-d trees) for vector-based retrieval. The approach is to <strong>crawl or ingest every source into search engines</strong>, creating a central place of company knowledge. This brings big advantages:</p><ul><li><p><strong>Speed by design:</strong> Searching an index is lightning fast. Queries hit inverted indexes and specialized data structures, avoiding the need to poll each backend system.</p></li><li><p><strong>Relevance that compounds over time:</strong> Search engines that support <strong>semantic search </strong>are capable of comprehending the intent, and machine learning models can rerank results for enterprise contexts. In one Elastic <a href="https://www.elastic.co/blog/elastic-generative-ai-experiences?">experiment</a>, Elastic users see more accurate results when combining vector search with a question-answering (QA) model to extract answers. It gives better precision than keyword matching.</p></li><li><p><strong>Advanced features:</strong> Elastic’s <a href="https://www.elastic.co/search-labs/blog/rag-graph-traversal#:~:text=Retrieval,for%20deeper%2C%20more%20contextual%20retrieval">Graph retrieval augmented generation (RAG) solution</a> shows how structuring an index as a knowledge graph can power more contextual retrieval. In other words, indexes aren’t just backward-looking dumps of text; they can also encode relationships and ontologies that let AI connect the dots across documents.</p></li><li><p><strong>Permission-aware search:</strong> Enterprise AI cannot compromise on security. Indexed search allows:</p><ul><li><p><a href="https://www.elastic.co/docs/reference/search-connectors/document-level-security">Document-level security.</a></p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/user-roles#roles">Role-based access control.</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/rag-and-rbac-integration">Permission-aware retrieval for RAG and agents.</a></p></li></ul></li></ul><p>Agents see only what users are allowed to see, without leaking data into model prompts or training. Elasticsearch is suitable for the indexed search layer in the diagram below, as it provides the essential components for context engineering.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb88e668cb4814bee/6a170cab8b73cb5d6918a090/8785e7806616273d086a90b3540273fb26d045ae-1392x768.jpg" alt="Essential components of context engineering, highlighting the Elasticsearch role in the indexed search layer." /><h2><strong>Retrieval consistency through search templates and governed execution</strong></h2><p>At scale, retrieval must be predictable, secure, and repeatable. This is where <a href="https://www.elastic.co/docs/solutions/search/search-templates">search templates</a> become critical.</p><p>Search templates act as retrieval contracts between applications, agents, and the search platform. Instead of dynamically constructing queries at runtime, agents invoke pre-defined retrieval patterns that enforce:</p><ul><li><p>Consistent relevance logic</p></li><li><p>Mandatory security filters</p></li><li><p>Cost and latency guardrails</p></li><li><p>Business-specific ranking rules</p></li><li><p>Explicit index and field scope boundaries</p></li></ul><p>In MCP-driven architectures, this becomes even more important. Agents should not dynamically invent retrieval strategies. Instead, MCP tool calls can map directly to approved search templates, ensuring that every retrieval request adheres to enterprise relevance and governance standards.</p><p>This approach shifts retrieval from ad-hoc query execution to controlled retrieval orchestration.</p><h2><strong>Retrieval is now a multi-layer engineering discipline</strong></h2><p>Modern enterprise retrieval is no longer a simple query-to-index operation. It typically includes multiple coordinated layers:</p><ul><li><p>Query understanding — rewriting, expansion, entity resolution</p></li><li><p>Retrieval strategy selection — hybrid search, vector search, graph retrieval, or synthetic query techniques such as Hypothetical Document Embeddings (HyDE), where the system generates a representative answer or expanded context first and retrieves documents using that richer semantic signal.</p></li><li><p>Execution governance — templates, security enforcement, and performance guardrails</p></li><li><p>Ranking and re-ranking — blending lexical precision, semantic similarity, and interaction-derived relevance signals such as click-through patterns, dwell time, and document usage frequency.</p></li></ul><p>When these layers are implemented upstream, agents receive clean, high-confidence context rather than raw, fragmented data.</p><p>This is what makes large-scale agent systems reliable in production environments.</p><h2><strong>Advanced retrieval techniques improve context quality before reasoning begins</strong></h2><p>Modern retrieval systems increasingly use AI-assisted techniques to improve recall and semantic coverage before ranking is applied.</p><p>One example is <a href="https://medium.com/@nirdiamant21/hyde-exploring-hypothetical-document-embeddings-for-ai-retrieval-cc5e5ac085a6">Hypothetical Document Embeddings (HyDE)</a>. Instead of embedding only the original query, the system first generates a hypothetical answer or expanded context, embeds that representation, and retrieves documents based on that richer semantic signal.</p><p>This is particularly useful in enterprise environments where:</p><ul><li><p>Users or agents may not know the exact terminology</p></li><li><p>Knowledge is distributed across silos</p></li><li><p>Important context is implied rather than explicitly stated</p></li></ul><p>Techniques like HyDE improve the probability that relevant documents are retrieved even when the original query is underspecified.</p><p>This reinforces a key principle of enterprise AI: better context retrieval produces better reasoning outcomes.</p><h2><strong>Agents aren’t data engineers; they’re reasoning systems</strong></h2><p>They shouldn’t be responsible for stitching together raw data, reconciling schemas, or compensating for poor retrieval.</p><p>This is where a search platform such as <strong>Elasticsearch</strong> becomes foundational.</p><p>By ingesting data once and normalizing it upstream (through pipelines, mappings, enrichment processors, and prebuilt indexes), Elasticsearch resolves schema mismatches, joins signals across sources, and materializes retrieval-ready views of the data. At query time, the agent receives clean, ranked, semantically enriched results rather than fragmented raw records.</p><p>For example, instead of an agent pulling independently from CRM, ticketing, and documentation systems and attempting to reconcile customer IDs, timestamps, and formats in real time, Elasticsearch can pre-index these sources into a unified customer interaction index with hybrid (keyword + vector) search and relevance ranking. The agent then queries a single, coherent interface and immediately reasons over the most relevant context.</p><p>This separation of concerns, that is, <strong>Elasticsearch handling data integration and retrieval, and agents focusing on reasoning, planning, and decision-making</strong>,is what makes agent systems scalable, reliable, and production ready.</p><h2><strong>Elastic’s role in the AI stack</strong></h2><p>Elastic sits at the intersection of search and AI by design.</p><ul><li><p><strong>Connectors and crawlers</strong> ingest data continuously from enterprise systems.</p></li><li><p><strong>Semantic and vector search</strong> enable intent-based retrieval.</p></li><li><p><strong>Hybrid search</strong> blends lexical precision with semantic understanding.</p></li><li><p><strong>RAG workflows</strong> ground LLMs in authoritative, permission-aware data.</p></li></ul><p>Elastic does not compete with agents or MCP. It <strong>makes them effective</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt950398e2b4ab71f8/6a170cac2b835fb826f4b26b/193da239544ce858416db845f9fc34c7c0e9b6f9-1920x1080.png" alt="AI-native experiences powered by tools and agents, enabled by the platform, and built on enterprise data." /><h2><strong>Bigger models don’t eliminate retrieval</strong></h2><p>Some have wondered whether huge new LLMs can bypass traditional search, perhaps by letting the model read <em>everything</em> in one go. Large context windows feel powerful, but they introduce:</p><ul><li><p>Higher latency.</p></li><li><p>Higher cost.</p></li><li><p>Lower precision due to noise.</p></li><li><p>A higher propensity for confusion, context clash, and context poisoning.</p></li></ul><p>RAG wins because it filters first and then reasons.In another <a href="https://www.elastic.co/search-labs/blog/rag-vs-long-context-model-llm#:~:text=,context%20approach%20led%20to%20inaccuracies">Elastic Search Labs experiment</a>, RAG achieved answers in about <strong>1 second</strong>, versus 45 seconds for the raw-LM approach, at <strong>1/1250th</strong> the cost, and with far higher accuracy. In other words, giving an LLM a million tokens of documents is slower, more expensive, and actually <em>less precise</em> than filtering through an index first.</p><h2><strong>Conclusion: MCP changes the interface, not the fundamentals</strong></h2><p>MCP is a meaningful step forward in how agents interact with tools. But it doesn’t replace the need for fast, relevant, governed retrieval.</p><p>In enterprise AI:</p><ul><li><p>Context quality determines answer quality.</p></li><li><p>Indexes create that context.</p></li><li><p>Search is the foundation, not the legacy.</p></li></ul><p>Indexes aren’t obsolete in the era of MCP. They’re <strong>the reason that MCP-based agents can work at all</strong>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/future-of-search-engines-indexed-search-mcp</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/future-of-search-engines-indexed-search-mcp</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Relevance]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Dayananda Srinivas]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1caa3ee906789415/6a170cae2b835f9a90f4b26f/5b8af1c3ca51f2c038406c714eb9a71b696bbc5a-1999x1091.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 05 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Using subagents and Elastic Agent Builder to bring business context into code planning]]></title>
    <description><![CDATA[Learn about subagents, how to ensure they have the right information, and how to create a specialized subagent that connects Claude Code to your Elasticsearch data.]]></description>
    <content:encoded><![CDATA[<p><a href="https://code.claude.com/docs/en/sub-agents">Subagents in Claude Code</a> let you offload specialized tasks to separate context windows, keeping your main conversation focused. In this article, you'll learn what subagents are, when to use them, and how to build a retrieval subagent using Elastic Agent Builder that connects your development workflow to business data in Elasticsearch.</p><h2>What are subagents?</h2><p><em>Subagents </em>are specialized assistants that can be called to execute a specific task, using their own context window. They complete a task and give the results to the main agent, preventing it from saving information that isn’t relevant for the rest of the conversation in the context window.</p><p>Their four core principles are:</p><ul><li><p><strong>Context preservation:</strong> Each subagent uses its own context window.</p></li><li><p><strong>Specialized expertise:</strong> Each subagent is designed for a specific task.</p></li><li><p><strong>Reusability:</strong> You can reuse a subagent in different sessions and projects.</p></li><li><p><strong>Flexible access:</strong> You can limit the subagent access to specific tools.</p></li></ul><p>Each subagent can have access to Claude Code tools to work with the terminal, such as glob, read, write, grep, or bash, or to access the internet, like search, fetch, or call external tools with Model Context Protocol (MCP) servers.</p><p>A subagent uses the following schema:</p>---
name: your-sub-agent-name
description: Description of when this subagent should be invoked
tools: tool1, tool2, tool3  # Optional - inherits all tools if omitted
model: sonnet  # Optional - specify model alias or 'inherit'
permissionMode: default  # Optional - permission mode for the subagent
skills: skill1, skill2  # Optional - skills to auto-load
---

Your subagent's system prompt goes here. This can be multiple paragraphs
and should clearly define the subagent's role, capabilities, and approach
to solve problems.

Include specific instructions, best practices, and any constraints
the subagent should follow.<p>You can call subagents implicitly by talking about the task they run, and Claude will call them automatically. For example, you can say, "I want to plan my new functionality."</p><p>You can also call them explicitly by directly asking Claude Code to use a subagent and telling it, "Use the planning subagent to plan my new functionality."</p><p>Another important feature is that subagents are stateful, so once you give one a task, it will generate an ID. This way, when you use it again, you can start from scratch or provide the ID to give it context from its previous tasks.</p><p>You can read the <a href="https://code.claude.com/docs/en/sub-agents">full documentation here</a>.</p><h2>When are subagents used?</h2><p>Subagents are useful when you need to delegate tasks that require specialized context but you don't want to clutter the main chat window. Considering our example of coding, the most common subtasks include:</p><p>Subtask type</p><p>Description</p><p>Typical tools</p><p>Exploration / research</p><p>Searching and analyzing code without modifying it.</p><p>Read, grep, glob</p><p>Planning</p><p>Running deep analysis to create implementation plans.</p><p>Read, grep, glob, bash</p><p>Code review</p><p>Reviewing quality, safety, and best practices.</p><p>Read, grep, glob, bash</p><p>Code modification</p><p>Writing and editing code.</p><p>Read, edit, write, grep, glob</p><p>Testing / debugging</p><p>Running tests and analyzing issues.</p><p>Bash, read, grep, edit</p><p>Retrieval</p><p>Getting information from external sources (APIs, databases).</p><p>MCP tools, bash</p><p>Claude Code includes three built-in agents that showcase these use cases:</p><p></p><ul><li><p><strong>Explore:</strong> Quick agents for read-only search in the codebase. It's great for answering questions like, "Where are the client's errors handled?"</p></li><li><p><strong>Plan:</strong> Research agent that activates in plan mode to analyze the codebase before proposing changes.</p></li><li><p><strong>General-purpose:</strong> The most capable agent for complex tasks that require multiple steps and can include modifications.</p></li></ul><h2>Context management: Ensuring subagents have the right information</h2><p>One of the most important decisions when designing subagents is how to handle context. There are three key considerations:</p><h3><strong>1. Which context the subagent should get</strong></h3><p>The prompt you give to the subagent must contain all of the necessary information to complete the task since the subagent doesn’t have access to the main chat. You need to be specific:</p><ul><li><p>Do NOT say, "Review the code."</p></li><li><p>SAY, "Review the changes to src/auth/index.ts, focusing on JWT token validation."</p></li></ul><p>Providing the exact file name makes a difference between using the read tool against the file directly and making a wide search using grep and thus wasting time and tokens.</p><p>Also consider what not to include. Irrelevant context can distract the subagent or bias results. It’s tempting to ask for multiple things in one pass, but focused tasks yield better results:</p><ul><li><p>Do NOT say, “Review src/auth/<a href="http://index.ts">index.ts</a>. Here is also the database schema and our API docs for reference, fix bugs and suggest improvements about the architecture decisions.”</p></li><li><p>SAY, “Fix the token refresh bug in src/auth/index.ts that's throwing AUTH_TOKEN_EXPIRED unexpectedly.”</p></li></ul><h3><strong>2. What tools to provide</strong></h3><p>Limit the tools to what’s strictly necessary. This improves security, keeps the subagent focused, and reduces unnecessary tool calls and execution costs.</p># For just an analysis agent
tools: Read, Grep, Glob

# For an agent that needs to modify the code
tools: Read, Edit, Write, Grep, Glob<p>If you don't specify a tools field, the subagent inherits all tools from the main agent, including MCP tools.</p><p>You can learn all Claude Code tools <a href="https://code.claude.com/docs/en/how-claude-code-works#tools">here</a>.</p><h3><strong>3. How to keep context between calls</strong></h3><p>Subagents can be resumed using their agentId:</p># First call
&gt; Use the code-analyzer agent to review the authentication module
[Agent completes the analysis and returns agentId: "abc123"]

# Continue with previous context
&gt; Resume agent abc123 and now analyze the authorization module
[Agent continues with the context from the previous chat]<p></p><p>You can ask Claude for the agent ID or find it in <code>~/.claude/projects/{project}/{sessionId}/subagents/</code></p><p>This is especially useful for long research tasks or multistep workflows.</p><p>Another way to keep context consistent is to ask the agent to write a Markdown checklist with what it's doing and its current progress. Then you can execute <code>/clear</code> without losing the initial instruction. In that request, you can define the task granularity or details to retain that make sense for your use case.</p># Task: Review authentication module

## Progress
- [x] Analyzed src/auth/index.ts
- [x] Found JWT validation issue
- [ ] Review authorization module
- [ ] Check rate limiting

## Findings
- Token refresh has race condition in line 42<p>After you clear the conversation, the next agent can pick it up from here. This is very useful when you want an agent to run a script over a list and watch the output record by record.</p><h2>Orchestration patterns</h2><p>It’s important to see subagents as a context optimization mechanism. The way in which you coordinate them determines the efficiency of the whole system. There are different orchestration patterns.</p><h3><strong>Sequential (chaining)</strong></h3><p>Here, a subagent completes a task, and its results feed the next one in a sequence of tasks, similar to traditional Linux piping.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt421eabae16c6057f/6a170cb06234e07fcedb1a43/74a3a376600cd1b7cdd2dddddfed2f00ab131eed-896x94.png" alt="Sequential (or chaining) subagents, each feeding the next in a sequence of tasks." /><p>Call example:</p>&gt; First use the planning agent to design the feature,
&gt; then use the coding agent to implement it,
&gt; finally use the reviewer agent to check the code<h3><strong>Parallel</strong></h3><p>In this pattern, multiple subagents run independent tasks simultaneously. The main Claude Code agent invokes them since <strong>subagents cannot spawn other subagents</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt43666c7c597cf666/6a170cb228671458ed93e375/84eca68d29bf79cf978a8089d3c18972738cd2c1-595x272.png" alt="A parallel subagent pattern, with the main Claude Code agent invoking three subagents." /><p>This approach reduces the execution time for tasks like code review since it allows you to work with the same code from different angles without impacting the running time.</p><h3><strong>Hub-and-spoke (delegation)</strong></h3><p>In this approach, the main agent acts as an orchestrator, delegates tasks to specialized agents, and then consolidates the results.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc869c00340b0ac97/6a170cb3e8fbce396239fcb7/93bb2cc55c435f509b426fbcc090a67c53021684-595x272.png" alt="A hub-and-spoke subagent pattern, where the main agent acts as an orchestra" /><p>This is the pattern we’ll implement in our example: The main Claude Code agent will delegate the gathering of business information to a retrieval agent built with Elastic Agent Builder, while the explore agent will look into local files and the planning agent builds a plan.</p><h2>Why use an agent instead of a single query?</h2><p>Before building our retrieval subagent, it's worth understanding when an agent adds value versus when a simple Elasticsearch Query Language (ES|QL) query suffices.</p><p>If you need a single aggregation, like "What's our most visited page?" just run the query directly. The agent adds value when your question requires:</p><ul><li><p><strong>Multiple queries that build on each other:</strong> The answer from query 1 informs query 2.</p></li><li><p><strong>Cross-index reasoning:</strong> Correlating data from different sources.</p></li><li><p><strong>Ambiguity resolution:</strong> The agent interprets and follows leads.</p></li><li><p><strong>Synthesis:</strong> Combining quantitative data with qualitative knowledge.</p></li></ul><p>Our example will demonstrate all of these capabilities.</p><h2>Agent Builder as subagent</h2><p>Generating code using AI is very quick, but the problem is having a good planning phase to set the boundaries for our coding agent. To help with that, Claude created a subagent that <a href="https://code.claude.com/docs/en/common-workflows#use-plan-mode-for-safe-code-analysis">specializes in planning</a> to perform deep analysis and create a to-do list for the main agent to execute.</p><p>With this flow, you can plan based on what Claude Code can see both in local files and on the internet. However, there's still knowledge available in Elasticsearch that you cannot access via standard tools.</p><p>To access our internal knowledge during the planning phase, we'll create a Claude Code subagent by making a retrieval agent using Agent Builder.</p><p>You can configure the agent using the UI or an API. In this example, we'll use the latter.</p><h3><strong>Prerequisites</strong></h3><ul><li><p><a href="https://code.claude.com/docs/en/setup">Claude Code</a> 2.0.76+</p></li><li><p>Elasticsearch 9.2</p></li><li><p>Elasticsearch <a href="https://www.elastic.co/docs/deploy-manage/api-keys/elasticsearch-api-keys">API key</a></p></li></ul><h3><strong>The scenario: Technical debt sprint planning</strong></h3><p>You're a tech lead. You have two weeks and two developers. Your <code>TECH_DEBT.md</code> lists 12 items. You can tackle maybe three or four. Which ones should you prioritize?</p><p>The complexity is that you need to optimize across multiple dimensions simultaneously:</p><ul><li><p><strong>User impact:</strong> How many users hit this issue?</p></li><li><p><strong>Business impact:</strong> Does it affect paying customers? Enterprise tier?</p></li><li><p><strong>Severity:</strong> Errors? Performance? Just ugly code?</p></li><li><p><strong>Effort:</strong> Quick win or rabbit hole?</p></li><li><p><strong>Dependencies:</strong> Does fixing A unlock fixing B?</p></li><li><p><strong>Strategic alignment:</strong> Does it align with Q1 priorities?</p></li></ul><p>A single query like, "What's the most important tech debt item?" fails because this requires:</p><ol><li><p>Reading <code>TECH_DEBT.md</code> to understand what the 12 items even are.</p></li><li><p>For EACH item, querying <code>error_logs</code>to get error frequency.</p></li><li><p>Cross-referencing with <code>customer_data</code> to see tier breakdown.</p></li><li><p>Checking <code>support_tickets</code>to see complaint volume.</p></li><li><p>Reading <code>engineering_standards</code> in the knowledge base to see whether any items violate core principles.</p></li><li><p>Reading <code>Q1_roadmap</code> to check strategic alignment.</p></li><li><p>Synthesizing all of this into a prioritized recommendation.</p></li></ol><p>This is where a retrieval agent can be helpful in orchestrating multiple queries across different indices and synthesizing the results.</p><h2>Steps</h2><h3><strong>Preparing the test dataset</strong></h3><p>We'll create four indices: a knowledge base with internal documentation, error logs, support tickets, and customer data.</p><p>You can create the indices, index the data, and create the agent using one of the following:</p><ul><li><p><strong>Kibana Dev Tools:</strong> Using the Elasticsearch requests provided below.</p></li><li><p><strong>Jupyter Notebook:</strong> Using the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/notebook.ipynb">complete notebook</a> written for this article.</p></li></ul><h2>Create the indices</h2><p>Open Kibana Dev Tools, and run the following requests to create each index with its mapping and bulk data. Here's an example of the knowledge index structure and data to be indexed:</p>PUT customer_data
{
  "mappings": {
    "properties": {
      "user_id": { "type": "keyword" },
      "customer_tier": { "type": "keyword" },
      "company_name": { "type": "text" },
      "mrr": { "type": "float" },
      "joined_at": { "type": "date" }
    }
  }
}

POST customer_data/_bulk
{"index":{}}
{"user_id":"enterprise_user_01","customer_tier":"enterprise","company_name":"Acme Corp","mrr":2500.00,"joined_at":"2023-01-15"}
{"index":{}}
{"user_id":"enterprise_user_02","customer_tier":"enterprise","company_name":"GlobalTech Inc","mrr":4200.00,"joined_at":"2022-08-20"}
{"index":{}}
{"user_id":"enterprise_user_05","customer_tier":"enterprise","company_name":"DataFlow Systems","mrr":3100.00,"joined_at":"2023-06-01"}
{"index":{}}
{"user_id":"user_001","customer_tier":"free","company_name":"","mrr":0,"joined_at":"2024-03-15"}
{"index":{}}
{"user_id":"user_002","customer_tier":"free","company_name":"","mrr":0,"joined_at":"2024-05-20"}
{"index":{}}
{"user_id":"user_045","customer_tier":"pro","company_name":"SmallBiz LLC","mrr":49.00,"joined_at":"2024-01-10"}
{"index":{}}
{"user_id":"user_089","customer_tier":"pro","company_name":"StartupXYZ","mrr":49.00,"joined_at":"2024-02-28"}<p>Full requests for all indices:</p><ul><li><p><strong>Knowledge index:</strong> <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/elasticsearch_requests/knowledge.txt">knowledge.txt</a></p></li><li><p><strong>Error logs index:</strong> <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/elasticsearch_requests/error_logs.txt">error_logs.txt</a></p></li><li><p><strong>Support tickets index:</strong> <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/elasticsearch_requests/support_tickets.txt">support_tickets.txt</a></p></li><li><p><strong>Customer data index:</strong> <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/elasticsearch_requests/customer_data.txt">customer_data.txt</a></p></li></ul><p>The raw JSON files with the dataset are also available:</p><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/dataset/knowledge.json">knowledge.json</a></p></li><li><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/dataset/error_logs.json">error_logs.json</a></p></li><li><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/dataset/support_tickets.json">support_tickets.json</a></p></li><li><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/dataset/customer_data.json">customer_data.json</a></p></li></ul><h2>Local project files</h2><p>Create the following Markdown (MD) files in your project. These files look like this:</p># Tech Debt Items

## AUTH-001: Token refresh race condition
- **Module**: src/auth/refresh.ts
- **Symptom**: Users randomly logged out
- **Estimate**: 3 days

## EXPORT-002: CSV export timeout on large datasets
- **Module**: src/export/csv.ts
- **Symptom**: Timeout after 30s for &gt;10k rows
- **Estimate**: 2 days

...<p>Full files:</p><p></p><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/TECH_DEBT.md">TECH_DEBT.md</a>: Tech debt items list.</p></li><li><p><a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/REQUIREMENTS.md">REQUIREMENTS.md</a>: FlowDesk Q1 2025 requirements.</p></li></ul><p>This ties directly to the tech debt items and gives the agent clear priorities to work with when cross-referencing with the Elasticsearch data.</p><h2>Create an agent with Agent Builder</h2><p>We'll now create an agent capable of running analytics queries with ES|QL to provide us with app usage information while also capable of searching to provide us info from Knowledge Base (KB) in unstructured text format.</p><p>We're using the <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools#built-in-tools">built-in tools</a> since they cover search and analytics on any index. Agent Builder also supports custom tools for more specialized operations, like scoping an index or adding ES|QL dynamic parameters, but that's beyond our scope here.</p><p>You can create the agent using the curl request in <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/subagents-with-elastic-agent-builder/elasticsearch_requests/create_agent.txt">create_agent.txt</a>.</p>curl -X POST "https://${KIBANA_URL}/api/agent_builder/agents" \
  -H "Authorization: ApiKey ${API_KEY}" \
  -H "kbn-xsrf: true" \
  -H "Content-Type: application/json" \
  -d '{
    "id": "tech-debt-advisor",
    "name": "Tech Debt Prioritization Agent",
    "description": "I help prioritize technical debt by analyzing error logs, support tickets, customer impact, and aligning with engineering standards and roadmap priorities.",
    "avatar_color": "#BFDBFF",
    "avatar_symbol": "TD",
    "configuration": {
      "instructions": "This agent helps prioritize technical debt items. Use the following indices:\n\n- knowledge: Engineering standards, policies, and roadmap priorities\n- error_logs: Production error frequency by module\n- support_tickets: Customer complaints and their urgency\n- customer_data: Customer tier information (enterprise, pro, free)\n\nWhen analyzing tech debt:\n1. Check error frequency in error_logs\n2. Cross-reference affected users with customer_data to understand tier impact\n3. Count support tickets and note urgency markers\n4. Check knowledge base for relevant policies and Q1 priorities\n5. Synthesize findings into prioritized recommendations",
      "tools": [
        {
          "tool_ids": [
            "platform.core.search",
            "platform.core.list_indices",
            "platform.core.get_index_mapping",
            "platform.core.get_document_by_id",
            "platform.core.execute_esql",
            "platform.core.generate_esql"
          ]
        }
      ]
    }
  }'<p>You’ll get this response if everything went OK:</p>{
  "id": "tech-debt-advisor",
  "type": "chat",
  "name": "Tech Debt Prioritization Agent",
  "description": "I help prioritize technical debt by analyzing error logs, support tickets, customer impact, and aligning with engineering standards and roadmap priorities.",
  ...
}<p>The agent will be available in Kibana, so you can now chat with it if you want:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb05460b667943ed6/6a170cb52867142a3193e379/c655ec6b9b1cc2fa1ab3cc13d289e7b96a543284-815x784.png" alt="Chat with a new agent in Kibana, creating a chart with clients sorted by monthly recurring revenue." /><h3><strong>Configure the agent as Claude Code tool</strong></h3><p>The agent we just created will expose an <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/mcp-server">MCP server.</a> Let's add the MCP server to Claude Code using the already-generated API key:</p>claude mcp add --transport http agentbuilder https://${KIBANA_URL}/api/agent_builder/mcp --header "Authorization: ApiKey ${API_KEY}"<p>We can check the connection status using <code>claude mcp get agentbuilder</code>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d9034e6312b1aa3/6a170cb6c1e8a59e95f88318/ba5fbc144f9e29151b8628dffd33dc74b12deece-499x177.png" alt="Code for &quot;Claude MCP get agentbuilder." /><h3><strong>Create a subagent that uses the tool</strong></h3><p></p><p>Now that we have the Agent Builder available as a set of MCP tools, we can create a subagent in Claude Code that will use all or some of those tools, in combination with Claude Code ones.</p><p></p><p>Claude Code recommends using its agent creator tool for this step:</p><p></p><p>1. Type <code>/agents</code> in Claude Code.</p><p>2. Choose <strong>Create new agent</strong>.</p><p>3. Select <strong>Project scope</strong> so that it's only available for this project. (This is the recommended setting to avoid agent overflow.)</p><p>4. Select <strong>Generate with Claude (recommended)</strong>.</p><p>5. Type in the description: "Agent that analyzes technical debt by querying Elasticsearch for error logs, support tickets, customer data, and engineering knowledge base. Use this agent when you need to prioritize tech debt items based on business impact."</p><p>6. In “Select tools,” choose <strong>Advanced options</strong> and select the tools we defined on the agent creation.</p>Individual Tools:
☒ platform.core.search (agentbuilder)
☒ platform.core.list_indices (agentbuilder)
☒ platform.core.get_index_mapping (agentbuilder)
☒ platform.core.get_document_by_id (agentbuilder)
☒ platform.core.execute_esql (agentbuilder)
☒ platform.core.generate_esq (agentbuilder)<p>7. Select <strong>[ Continue ]</strong>.</p><p>Now choose the model. For planning tasks, the recommendation is to use Opus due to its significant reasoning capacity. So let's select that and continue.</p><p>Finally, choose the background color for our subagent text and confirm.</p><p>Claude automatically names our subagent based on the description (for example, <code>tech-debt-analyzer</code>).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc9ab15478a2f3a77/6a170cb8b0367d3ed472bd75/f01ac4c9f30fbcbed7fc69881aae9ff72c4616a0-869x521.png" alt="Code for creating a new subagent" /><h2>Testing the agent</h2><p>Once the agent has been created, we can test it with a complex prioritization question that requires multistep reasoning:</p>&gt; Based on TECH_DEBT.md, which items should we prioritize for our 2-week sprint?
&gt; Use the tech-debt-analyzer agent to check error frequency, customer impact,
&gt; support ticket volume, and alignment with engineering standards.<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt81a604f43091d27a/6a170cb90e2e49471441a165/d76d972ab5b07e6d35bdf3036cb5ee3c080c7156-749x239.png" alt="Code to demonstrate testing the agent with a complex prioritization question." /><p>Watch how the agent orchestrates multiple queries:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt093a3a5425cccb68/6a170cbb7d8d671c6570e775/c49b56c366576406586ba03f694d2bfb09d30895-875x96.png" alt="Code to demonstrate how the agent orchestrates multiple queries." /><p>And will give you a comprehensive analysis of the local files combined with Elasticsearch data:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8d9193f977e5fbbc/6a170cbdc1e8a5779ff8831c/084c532b4c9e993e53810738ae1da1fd4af1f025-1228x693.png" alt="A comprehensive analysis of the local files combined with Elasticsearch data." /><p>This demonstrates why a single query fails and an agent succeeds: It orchestrates five or more queries across different indices, correlates the data, and synthesizes a recommendation that contradicts the naive "fix highest error count" approach.</p><p>By typing <code>/context</code>, we can see how much context each of the MCP tool's definitions uses and our subagent's prompt. Keep an eye on this overhead when creating subagents.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8c6475d26fa91b2d/6a170cbfd7c02246d5de64e5/3a6f528a9cae7b7fbf17f1b97e13c51c78c1b8b4-666x391.png" alt="Code that shows how much context each of the MCP tool's definitions uses and our subagent's prompt." /><h2>Start planning</h2><p>We can now start planning using local files, the internet, and our Elasticsearch knowledge as information sources.</p><p>Ask something like:</p>"Based on our requirements defined in REQUIREMENTS.md, use the planning agent
to create a detailed implementation plan, prioritizing tasks according to
business impact. Use the tech-debt-analyzer agent to query about internal
company knowledge and make analytical queries about error patterns and
customer impact."<p>Note that Claude decides to run the Elasticsearch data analysis and the local documentation reading in parallel, following the hub-and-spoke orchestration pattern.</p><p>After the analysis, you should get a plan that prioritizes based on actual business data rather than on assumptions. This context will make your AI coding experience much more reliable, as you can feed this plan directly to the agent and execute step by step:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt76067b358ead61bf/6a170cc04a531bb59136a9ba/cfa5c6c44425d6e73355116e08082a33699915a3-961x873.png" alt="Data analysis results that provide an implementation plan prioritized based on actual business data rather than on assumptions." /><p>The more details you provide and the more focused the instructions are, the better the quality of the plan will be. If you have an existing codebase, it will suggest the code changes.</p><h2>Conclusion</h2><p>Subagents are a great tool to offload specific tasks where we only need the final result for the main chat (without going through how we got there), keeping the chat flow focused.</p><p>By choosing the right orchestration pattern (sequential, parallel, or hub-and-spoke) and handling the context properly, we can build efficient and maintainable agent systems.</p><p>Elastic Agent Builder and its MCP feature allow us to access our data using a retrieval subagent to facilitate planning and coding by combining local (files, source code), external (internet), and internal (Elasticsearch) sources. The key insight is that agents add value not for simple queries but when you need multistep reasoning that builds on previous results and synthesizes information from multiple sources.</p><h2>Resources</h2><ul><li><p><a href="https://code.claude.com/docs/en/sub-agents">Claude Code Subagents</a></p></li><li><p><a href="https://www.elastic.co/elasticsearch/agent-builder">Elastic Agent Builder</a></p></li><li><p><a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/mcp-server">Agent Builder MCP</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/subagents-with-elastic-agent-builder</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/subagents-with-elastic-agent-builder</guid>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Gustavo Llermaly]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt75bf7ebc5c2c72a8/6a170cc26f7f04f6ba9148b4/bfeb78b687bd930371364ee7dd0341ae90004349-1280x720.png" length="0" type="image/png"/>
    <pubDate>Tue, 03 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[AI agents that perform actions: Automating IT requests with Agent Builder and Workflows]]></title>
    <description><![CDATA[Using  Elastic Agent Builder and Workflows to create an AI agent that automatically performs IT actions, such as laptop refreshes.]]></description>
    <content:encoded><![CDATA[<p>In the world of IT operations, context switching is the enemy of productivity. For internal teams, simple requests, like a laptop refresh or employee onboarding, often require navigating multiple portals, filling out rigid forms, and manually updating information technology service management (ITSM) tools like ServiceNow.</p><p>At a recent <strong>DevFest</strong>, we demonstrated how to bridge the gap between natural language requests and structured IT workflows. By combining <a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder"><strong>Elastic Agent Builder</strong></a> with <a href="https://www.elastic.co/docs/explore-analyze/workflows"><strong>Elastic Workflows</strong></a>, we can create AI assistants that not only answer questions but also perform complex actions.</p><p>In this post, we’ll dive into the architecture from that talk, specifically looking at how we built an automated "Laptop Refresh" workflow. We’ll demonstrate how to configure an agent that collects user requirements and triggers a server-side automation to interact directly with ServiceNow APIs.</p><p><strong>Watch the full breakdown:</strong> This post is based on our presentation at Google DevFest. You can <a href="https://www.youtube.com/watch?v=OzStbTUZqyw">watch the full session here</a> to see the demo in action.</p><h2><strong>The architecture: From chat to fulfillment</strong></h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc15ebf6d2dbf924a/6a170e28a6c2b91dcbe79788/eb42459bfae9c2ac95f2012882ce826db5526705-1600x1000.png" alt="Agent Builder &amp; Workflows architecture: Laptop Refresh automation" /><p><strong>Note:</strong> The technical implementation described in this document is a streamlined version of the full production environment. While the <strong>architecture diagram</strong> provided serves as an accurate structural reference for the actual deployment, the accompanying text and code snippets have been simplified for illustrative purposes and may differ from the final, complex configurations used in the live implementation.</p><p>The goal is to move from a manual, form-heavy process to a conversational interface. Instead of a user navigating a catalog, they simply tell the AI assistant that they’re due for a laptop upgrade.</p><p>As illustrated above, the flow consists of three distinct layers:</p><p><strong>1. Interaction layer (ElasticGPT/Agent Builder):</strong> The user interacts naturally with an interface powered by ElasticGPT. Behind the scenes, Agent Builder processes this conversation, handling intent detection and slot filling, to structure the data and orchestrate interactions with other internal systems.</p><ul><li><p><strong>Intent detection</strong></p><ul><li><p><strong>Mechanism:</strong> System prompt instruction.</p></li><li><p><strong>Implementation:</strong> The agent is explicitly told its single purpose in the <code>MISSION</code> statement. It doesn’t need to "detect" other intents because it’s scoped strictly to IT provisioning.</p><ul><li><p><em><strong>Code reference</strong></em><em>:</em> <code>MISSION: You are a specialized agent designed to collect complete employee onboarding information...</code></p></li></ul></li><li><p><strong>Constraint:</strong> If a user asks about non-IT topics (for example, "What is the weather?"), the <code>MISSION</code> implies that the agent should pivot back to data collection or decline, depending on the large language model’s (LLM's) default safety alignment.</p></li></ul></li><li><p><strong>Slot filling (data collection)</strong></p><ul><li><p><strong>Mechanism:</strong> Phased conversation flow.</p></li><li><p><strong>Implementation:</strong> Instead of asking for all slots at once, the DATA <code>COLLECTION STRATEGY</code> breaks the slots into five logical phases. This prevents the context switching fatigue mentioned above.</p><ul><li><p><em><strong>Code reference:</strong></em><code>PHASE 1: Personal information, PHASE 2: Employment Details, and so on.</code></p></li></ul></li><li><p><strong>Validation:</strong> The prompt enforces immediate validation (for example, <code>Validate inputs immediately</code>), acting as a gatekeeper before moving to the next slot.</p></li></ul></li></ul><p><strong>2. Automation layer ( Workflows):</strong> Once the agent has the data, it triggers a workflow. This workflow handles the logic: checking device eligibility, enforcing policy (for example, "Is the laptop &gt; 3 years old?"), and making API calls.</p><p><strong>3. System of record (ServiceNow):</strong> The workflow reads and writes directly to the ITSM tool to maintain audit trails and initiate fulfillment.</p><h2><strong>Step 1: Configuring the agent</strong></h2><p>The first step is defining the "brain" of the operation using <strong>Agent Builder</strong>. We need an agent that acts strictly within the bounds of IT provisioning. We don't want a general chatbot; we want a data collection machine that feels like a helpful colleague.</p><p>We achieve this via a robust <strong>system prompt</strong>. The prompt dictates the agent's operating protocol, enforcing a step-by-step data collection strategy.</p><p>Here’s the refined structure of the prompt we used. Notice how it enforces validation and logically groups questions to avoid overwhelming the user:</p>MISSION: You are a specialized agent designed to collect complete employee onboarding information for IT equipment provisioning.

OPERATING PROTOCOL:
0. On every new chat, send a welcome message, and directly jump to data collection.

1. DATA COLLECTION STRATEGY:
   - Use a step-by-step approach across 5 clear phases
   - Validate inputs immediately

2. CONVERSATION FLOW:
   PHASE 1: Personal Information (Name, Email, Phone)
   PHASE 2: Employment Details (Job Title, Department, Manager)
   PHASE 3: Location &amp; Shipping (Address, Country)
   PHASE 4: Technical Setup (Laptop Type, Accessories)
   PHASE 5: Confirmation

...

6. SUCCESS COMPLETION:
   After all data is collected and validated, invoke the tool "laptoprefreshworkflow" with the JSON payload.<p>For a sample system prompt or instructions, please refer <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/building-actionable-ai-automating-it-requests-with-agent-builder-and-one-workflow/Try%20it%20yourself%20Agents/service_now_utility_agent.ipynb">here</a>.</p><p>By explicitly instructing the agent to send the data in a specific JSON format at the end of the conversation, we ensure that the input matches exactly what our automation layer expects.</p><h2><strong>Step 2: The automation layer (Workflows)</strong></h2><p>The agent provides the <em>intent</em> and the <em>data</em>, but <strong>Workflows</strong> provides the <em>muscle</em>.</p><p>We define a workflow using a YAML configuration. This workflow acts as the bridge between the AI agent and the ServiceNow REST APIs. It handles authentication, data retrieval, and the ordering process.</p><p>Below is the workflow definition. We’ve refined the code to use secure variable handling for credentials rather than hardcoding them.</p><h3><strong>Workflow inputs</strong></h3><p>First, we define the inputs the workflow expects to receive from the agent:</p>YAML
version: "1"
name: Submit Laptop Refresh Request
enabled: true
triggers:
  - type: manual
inputs:
  - name: userid
    type: string
  - name: preferred-address
    type: string
  - name: laptop-choice
    default: Macbook latest
    type: string
  - name: laptop-keep-or-return
    default: return
    type: string<h3><strong>Interacting with ServiceNow</strong></h3><p>The workflow executes a series of HTTP steps. Crucially, we first need to identify the user's <em>current</em> asset to link the refresh request correctly.</p><p>1. Fetching computer data</p><p>We query the cmdb_ci_computer table in ServiceNow to find the asset currently assigned to the user.</p>YAML
steps:
  - name: snow_get_computer_data
    type: http
    with:
      url: https://elasticdev.service-now.com/api/now/table/ci_computer?assigned_to={{ inputs.userid }}
      method: GET
      headers:
        Accept: application/json
        Content-Type: application/json
        # Best Practice: Use secrets for authorization headers
        Authorization: Basic {{ secrets.servicenow_creds }}
      timeout: 30s<p>2. Adding to cart</p><p>Once we have the asset details and the user's preferences, we don't just create a generic ticket. We use the ServiceNow Service Catalog API to programmatically add the specific item to a cart.</p>YAML
  - name: snow_post_add_item_to_cart
    type: http
    with:
      url: https://elasticdev.service-now.com/example
      method: POST
      headers:
        Accept: application/json
        Content-Type: application/json
        Authorization: Basic {{ secrets.servicenow_creds }}
      body: |
        {
            "sysparm_quantity": 1,
            "variables": {
              "caller_id_common": "{{ inputs.userid }}",
              "current_device": "{{ steps.snow_get_asset.output.data.result.sys_id }}",
              "laptop_keep_or_return": "{{ inputs.laptop-keep-or-return }}",
              "choose_your_laptop": "{{ inputs.laptop-choice }}",
              "shipping_address": "{{ inputs.preferred-address }}"
            }
        }<p>3. Indexing the transaction</p><p>Finally, we want to keep a record of this transaction within Elasticsearch for analytics and future reference. We use the elasticsearch.index step to store the request details immediately after submission.</p>YAML

  - name: index-submission-record
    type: elasticsearch.index
    with:
      index: laptop-refresh-submission-data
      id: "{{ steps.snow_post_submit_order.output.data.result.request_id }}"
      document:
        request-id: "{{ steps.snow_post_submit_order.output.data.result.request_id }}"
        user-id: "{{ inputs.userid }}"
        configuration-item: "{{ steps.snow_get_computer_data.output.data.result[0].sys_id }}"
        laptop-choice: "{{ inputs.laptop-choice }}"
        timestamp: "{{ steps.snow_post_submit_order.output.data.result.sys_created_on }}"<p>For detailed workflow yaml, please refer <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/building-actionable-ai-automating-it-requests-with-agent-builder-and-one-workflow">here</a>.</p><h2><strong>The result</strong></h2><p>By stitching these components together, we create a seamless experience:</p><ol><li><p><strong>The user</strong> chats naturally with the agent to provide details.</p></li><li><p><strong>The agent</strong> structures this unstructured conversation into a JSON object.</p></li><li><p><strong>Workflow</strong> receives the JSON, validates the user's current hardware via ServiceNow, creates the order, and indexes the result.</p></li></ol><p>This approach reduces a process that traditionally took users 5–10 minutes of form navigation into a quick conversation, while ensuring that IT operations retains full visibility and control.</p><p>Video demo: </p><h2><strong>Ready to build?</strong></h2><p>This pattern, using an agent for the interface and using Workflows for the execution, can be applied to almost any ITSM task, from password resets to software provisioning.</p><p>If you’re interested in trying this out, be sure to watch the <a href="https://www.youtube.com/watch?v=OzStbTUZqyw">DevFest talk</a> for the full context, and check out the <a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Elastic AI Agent Builder documentation</a> to get started building your own agents today.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-builder-one-workflow</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-builder-one-workflow</guid>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Sri Kolagani,Ziyad Akmal]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9cd2f96f1ccdd81f/6a170e2a961e69b254c4cfa0/80e98ed860633a0a20abcc55ad10b2854a4e8df0-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 13 Feb 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Managing agentic memory with Elasticsearch]]></title>
    <description><![CDATA[Creating more context-aware and efficient agents by managing memories using Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p>In the emerging discipline of <strong>context engineering</strong>, giving AI agents the right information at the right time is crucial. One of the most important aspects of context engineering is managing an AI’s <strong>memory</strong>. Much like humans, AI systems rely on both a short-term memory and a long-term memory to recall information. If we want large language model (LLM) agents to carry on logical conversations, remember user preferences, or build on previous results or responses, we need to equip them with effective memory mechanisms.</p><p>After all, everything in the context influences the AI’s responses. G<em>arbage in, garbage out</em> holds true.</p><p>In this article, we’ll introduce what short-term and long-term memory mean for AI agents, specifically:</p><ul><li><p>The difference between short- and long-term memory.</p></li><li><p>How they relate to retrieval-augmented generation (RAG) techniques with vector databases, like Elasticsearch, and why careful memory management is necessary.</p></li><li><p>The risks of neglecting memory, including context overflow and context poisoning.</p></li><li><p>Best practices, like context pruning, summarizing, and retrieving only what’s relevant, to keep an agent’s memory both useful and safe.</p></li><li><p>Finally, we’ll touch on how memory can be shared and propagated in multi-agent systems to enable agents to collaborate without confusion using Elasticsearch.</p></li></ul><h2>Short-term versus long-term memory in AI agents</h2><p><em><strong>Short-term memory</strong></em> in an AI agent typically refers to the immediate conversational context or state—essentially, the current chat history or recent messages in the active session. This includes the user’s latest query and recent back-and-forth exchanges. It’s very similar to the information a person holds in mind during an ongoing conversation.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blteb714ce810d1c472/6a170f321949f782cbe7aaf6/4fbcc6f68055b2bccefc4176297a4ca50056dc0d-764x498.png" alt="Short-term &amp; long-term agentic memory" /><p>AI frameworks often maintain this transient memory as part of the agent’s state (for example, using a checkpointer to store the conversation state as covered by <a href="https://docs.langchain.com/oss/python/langgraph/persistence#checkpoints">this example from LangGraph</a>). Short-term memory is <em><strong>session-scoped</strong></em>; that is, it exists within a single conversation or task and is reset or cleared when that session ends, unless explicitly saved elsewhere. An example of session-bound short-term memory would be the <a href="https://help.openai.com/en/articles/8914046-temporary-chat-faq"><strong>temporary chat</strong></a>available in ChatGPT.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4b8680e22d4e1185/6a170f341949f78bbae7aafa/150bdf209cda5ed20b59cddf34e624ad1a8016aa-1100x577.png" alt="AI frameworks memory" /><p><em><strong>Long-term memory</strong></em>, on the other hand, refers to information that persists <strong>across conversations or sessions</strong>. This is the knowledge an agent retains over time, facts it learned earlier, user preferences, or any data we’ve told it to remember permanently.</p><p>Long-term memory is usually implemented by storing and fetching it from an external source, such as a file or vector database that’s outside the immediate context window. Unlike short-term chat history, long-term memory isn’t automatically included in every prompt. Instead, based on a given scenario, the agent must <strong>recall</strong> or retrieve it when relevant tools are invoked. In practice, long-term memory might include a user’s profile info, prior answers or analyses the agent produced, or a knowledge base the agent can query.</p><p>For instance, if you have a travel-planner agent, the <em>short-term memory</em> would contain details of the current trip inquiry (dates, destination, budget) and any follow-up questions in that chat; whereas the <em>long-term memory</em> could store the user’s general travel preferences, past itineraries, and other facts shared in previous sessions. When the user returns later, the agent can pull from this long-term store (for example, the user loves beaches and mountains, has an average budget of INR 100,000, has a bucket list to visit, and prefers to experience history and culture rather than kid-friendly attractions) so that it doesn’t treat the user as a blank slate each time.</p><p>The short-term memory (chat history) provides immediate context and continuity, while long-term memory provides a broader context that the agent can draw upon when needed. Most advanced AI agent frameworks enable both: They keep track of recent dialogue to maintain context <em>and</em> offer mechanisms to look up or store information in a longer-term repository. Managing short-term memory ensures it stays within the context window, while managing long-term memory helps the agent to ground the answers based on prior interactions and personas.</p><h2>Memory and RAG in context engineering</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt98c1514741bea460/6a170f36509168083ce1bbae/46635aa11ceff89b8d6a26ac3e22da52407d82f3-1600x900.png" alt="Memory and RAG in context engineering" /><p><em><strong>How do we give an AI agent a useful long-term memory in practice?</strong></em></p><p>One prominent approach for long-term memory is <em><strong>semantic memory</strong></em>, often implemented via <strong>retrieval-augmented generation (RAG)</strong>. This involves coupling the LLM with an external knowledge store or vector-enabled datastore, like Elasticsearch. When the LLM needs information beyond what’s in the prompt or its built-in training, it performs semantic retrieval against Elasticsearch and injects the most relevant results into the prompt as context. This way, the model’s effective context includes not only the recent conversation (short-term memory) but also pertinent long-term facts fetched on the fly. The LLM then grounds its answer on both its own reasoning and the retrieved information, effectively combining short-term memory and long-term memory to produce a more accurate, context-aware response.</p><p><strong>Elasticsearch </strong>can be used to implement long-term memory for AI agents. Here’s a high-level example of how context can be retrieved from Elasticsearch for long-term memory.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt44f5a6887b0bca32/6a170f37a6c2b9c735e797be/41ccbc7b5171e8170ac300139a963c0708816ba6-1600x900.png" alt="RAG in action" /><p>This way, the agent “remembers” by searching for relevant data rather than by storing everything in its limited prompt, <strong>where it leads to different risks.</strong></p><p><strong>Using RAG with Elasticsearch or any vector stores offers multiple benefits:</strong></p><p>First, it <strong>extends the knowledge</strong> of the model beyond its training cutoff. The agent can retrieve up-to-date information or domain-specific data that the LLM might not know. This is crucial for questions about recent events or specialized topics.</p><p>Second, retrieving context on demand helps reduce hallucinations, especially since LLMs aren’t trained on the proprietary or highly specialized data relative to your niche use case, which is highly likely to expose it to hallucinations. Instead of the LLM guessing or inventing new information as it has been incentivised through evaluation, as highlighted in a recent OpenAI paper (<a href="https://arxiv.org/pdf/2509.04664">Why Language Models Hallucinate</a>), the model can be grounded by factual references from Elasticsearch. Naturally, the LLM depends on the reliability of the data in the vector store to truly prevent misinformation and the relevant data is retrieved as per the core relevance measures.</p><p>Third, RAG allows an agent to work with knowledge bases far larger than anything you could ever fit into a prompt. Instead of pushing entire documents, like long research papers or policy documents, into the context window and risking overload or irrelevant information <a href="https://www.elastic.co/search-labs/blog/agentic-memory-management-elasticsearch#context-poisoning">context poisoning</a> the model’s reasoning, RAG relies on <a href="https://www.elastic.co/search-labs/blog/chunking-strategies-elasticsearch">chunking</a>. Large documents are broken into smaller, semantically meaningful pieces, and the system retrieves only the few chunks most relevant to the query. This way, the model doesn’t need a million-token context to appear knowledgeable; it just needs access to the right chunks of a much larger corpus.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4c90f81a56db0a33/6a170f3960084be7ba3c462e/e6897356c9f0940e35a63d005e9cd20bc33e5dd7-1600x931.png" alt="Evolution of LLM context engineering" /><p>It’s worth noting that as LLM context windows have grown (<a href="https://www.anthropic.com/news/1m-context">some models now support hundreds of thousands or even millions of tokens</a><em>)</em>, a debate arose about whether RAG is “dead.” Why not push all the data into the prompt? If you feel likewise, refer to this wonderful article by my colleagues, Jeffrey Rengifo and Eduard Martin, <a href="https://www.elastic.co/search-labs/blog/rag-vs-long-context-model-llm">Longer context ≠ better: Why RAG still matters</a>. This avoids the “garbage in, garbage out” problem: The LLM stays focused on the few chunks that matter, rather than running through noise.</p><p>That said, integrating Elasticsearch or any vector store into an AI agent architecture provides <strong>long-term memory</strong>. The agent stores knowledge externally and pulls it in as memory context when needed. This could be implemented as an <em>architecture</em>, where after each user query, the agent performs a search on Elasticsearch for relevant info and then appends the top results to the prompt before calling the LLM. The response might also be saved back into the long-term store if it contains useful new information (creating a feedback loop of learning). By using such retrieval-based memory, the agent remains informed and up to date, without having to cram everything it knows into every prompt, even though the context window supports <em>one million tokens</em>. This technique is a cornerstone of context engineering, combining the strengths of information retrieval and generative AI. </p><p>Here’s an example of a managed in-memory conversation state using LangGraph's checkpoint system for short-term memory during the session. (Refer to our <a href="https://github.com/someshwaranM/elastic-context-engineering-short-term-long-term-memory">supporting context engineering app</a>.)</p># Initialize chat memory (Note: This is in-memory only, not persistent)
memory = MemorySaver()

# Create a LangGraph agent
langgraph_agent = create_react_agent(model=llm, tools=tools, checkpointer=memory)

...
...
# Only process and display checkpoints if verbose mode is enabled
if args.verbose:
    # List all checkpoints that match a given configuration
    checkpoints = memory.list({"configurable": {"thread_id": "1"}})
    # Process the checkpoints
    process_checkpoints(checkpoints)<p>Here’s how it stores <strong>checkpoints</strong>:</p>Checkpoint:
Timestamp: 2025-12-30T09:19:41.691087+00:00
Checkpoint ID: 1f0e560a-c2fa-69ec-8001-14ee5373f9cf
User: Hi I'm Som, how are you? (Message ID: ad0a8415-5392-4a58-85ad-84154875bbf2)
Agent: Hi Som! I'm doing well, thank you! How about you? (Message ID: 
56d31efb-14e3-4148-806e-24a839799ece)
Agent:  (Message ID: lc_run--019b6e8e-553f-7b52-8796-a8b1fbb206a4-0)

Checkpoint:
Timestamp: 2025-12-30T09:19:40.350507+00:00
Checkpoint ID: 1f0e560a-b631-6a08-8000-7796d108109a
User: Hi I'm Som, how are you? (Message ID: ad0a8415-5392-4a58-85ad-84154875bbf2)
Agent: Hi Som! I'm doing well, thank you! How about you? (Message ID: 
56d31efb-14e3-4148-806e-24a839799ece)

Checkpoint:
Timestamp: 2025-12-30T09:19:40.349027+00:00
Checkpoint ID: 1f0e560a-b62e-6010-bfff-cbebe1d865f6<p>For long-term memory, here's how we perform semantic search on Elasticsearch to retrieve relevant previous conversations using vector embeddings after summarizing and indexing the checkpoints to Elasticsearch.</p>Functions: 
retrieve_from_elasticsearch() 

# Enhanced Elasticsearch retrieval with rank_window and verbose display
def retrieve_from_elasticsearch(query: str, k: int = 5, rank_window: int = None) -&gt; tuple[List[Dict[str, Any]], str]:
    """
    Retrieve context from Elasticsearch with score-based ranking
    
    Args:
        query: Search query
        k: Number of results to return
        rank_window: Number of candidates to retrieve before ranking (default: args.rank_window)
        
    Returns:
        Tuple of (retrieved_documents, formatted_context_string)
    """
    if not es_client or not es_index_name:
        return [], "Elasticsearch is not available. Cannot search long-term memory."
    
    if rank_window is None:
        rank_window = args.rank_window
    
    try:
        # Check if index exists and has documents
        if not es_client.indices.exists(index=es_index_name):
            return [], "No previous conversations stored in long-term memory yet."
        
        # Get document count
        try:
            doc_count = es_client.count(index=es_index_name)["count"]
            if doc_count == 0:
                return [], "Long-term memory is empty. No previous conversations to search."
        except Exception as e:
            return [], f"Error checking memory: {str(e)}"
        
        # Generate embedding for the query
        try:
            query_embedding = embeddings.embed_query(query)
        except Exception as e:
            return [], f"Error generating embedding: {str(e)}"
        
        # Perform semantic search using kNN with rank_window
        try:
            search_body = {
                "knn": {
                    "field": "vector",
                    "query_vector": query_embedding,
                    "k": k,
                    "num_candidates": rank_window  # Retrieve more candidates, then rank top k
                },
                "_source": ["text", "content", "message_type", "timestamp", "thread_id"],
                "size": k
            }
            
            response = es_client.search(index=es_index_name, body=search_body)
            
            if not response.get("hits") or len(response["hits"]["hits"]) == 0:
                return [], "No relevant previous conversations found in long-term memory."
            
            # Extract documents with scores
            retrieved_docs = []
            for hit in response["hits"]["hits"]:
                source = hit["_source"]
                score = hit["_score"]
                retrieved_docs.append({
                    "content": source.get("content", source.get("text", "")),
                    "message_type": source.get("message_type", "unknown"),
                    "timestamp": source.get("timestamp", "unknown"),
                    "thread_id": source.get("thread_id", "unknown"),
                    "score": score
                })
            
            # Format context string
            context_parts = []
            for i, doc in enumerate(retrieved_docs, 1):
                context_parts.append(doc["content"])
            
            context_string = "\n\n".join(context_parts)
            
            # Verbose display
            if args.verbose:
                rich.print(f"\n[bold yellow]🔍 RETRIEVAL ANALYSIS[/bold yellow]")
                rich.print("="*80)
                rich.print(f"[blue]Query:[/blue] {query}")
                rich.print(f"[blue]Retrieved:[/blue] {len(retrieved_docs)} documents (from {rank_window} candidates)")
                rich.print(f"[blue]Total context length:[/blue] {len(context_string)} characters\n")
                
                for i, doc in enumerate(retrieved_docs, 1):
                    rich.print(f"[cyan]📄 Document {i} | Score: {doc['score']:.4f} | Type: {doc['message_type']}[/cyan]")
                    rich.print(f"[cyan]   Timestamp: {doc['timestamp']} | Thread: {doc['thread_id']}[/cyan]")
                    content_preview = doc['content'][:200] + "..." if len(doc['content']) &gt; 200 else doc['content']
                    rich.print(f"[cyan]   Content: {content_preview}[/cyan]")
                    rich.print("-" * 80)
            
            return retrieved_docs, context_string
            
        except Exception as e:
            return [], f"Error searching memory: {str(e)}"
            
    except Exception as e:
        return [], f"Error accessing long-term memory: {str(e)}"<p>Now that we’ve explored how short-term memory and long-term memory are indexed and fetched using LangGraph’s checkpoints in Elasticsearch, let’s take some time to understand why indexing and dumping the complete conversations can be risky.</p><h2>Risks of not managing context memory</h2><p>As we’re talking much about context engineering, along with short-term and long-term memory, let’s understand what happens if we don’t manage an agent’s memory and context well.</p><p>Unfortunately, many things can go wrong when an AI’s context grows extremely long or contains bad information. As context windows get larger, <strong>new failure modes</strong> emerge, like:</p><ul><li><p><strong>Context poisoning</strong></p></li><li><p><strong>Context distraction</strong></p></li><li><p><strong>Context confusion</strong></p></li><li><p><strong>Context clash</strong></p></li><li><p><strong>Context leakage and knowledge conflicts</strong></p></li><li><p><strong>Hallucinations and misinformation</strong></p></li></ul><p>Let’s break down these issues and other risks that arise from poor context management:</p><h3>Context poisoning</h3><p><em>Context poisoning</em> refers to when incorrect or harmful information ends up in the context and “poisons” the model’s subsequent outputs. A common example is a hallucination by the model that gets treated as fact and inserted into the conversation history. The model might then build on that error in later responses, compounding the mistake. In iterative agent loops, once a false information makes it into the shared context (for example, in a summary of the agent’s working notes), it can be reinforced over and over. </p><p><a href="https://storage.googleapis.com/deepmind-media/gemini/gemini_v2_5_report.pdf">Researchers at DeepMind, in the release of the Gemini 2.5 report</a> (TL;DR, check <a href="https://www.dbreunig.com/2025/06/17/an-agentic-case-study-playing-pok%C3%A9mon-with-gemini.html">here</a>), observed this in a long-running <em>Pokémon</em>-playing agent: If the agent hallucinated a wrong game state and that got recorded into its <em>context </em>(its memory of goals), the agent would form <strong>nonsensical strategies</strong> around an impossible goal and get stuck. In other words, a poisoned memory can send the agent down the wrong path indefinitely.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd56e9e0681f32239/6a170f3b4a531bd79536aa21/3f2facf5aad67613ad557422e09ec23a66adc0ed-1600x1388.png" alt="Context poisoning" /><p>Context poisoning can happen innocently (by mistake) or even maliciously, for instance, via prompt injection attacks where a user or third-party sneaks in a hidden instruction or false fact that the agent then remembers and follows.</p><p><strong>Recommended countermeasures:</strong></p><p>Based on insights from <a href="https://www.wiz.io/academy/data-poisoning">Wiz</a>, <a href="https://zerlo.net/en/blog/what-is-llm-data-poisoning">Zerlo</a>, and <a href="https://www.anthropic.com/research/small-samples-poison">Anthropic</a>, countermeasures for context poisoning focus on preventing bad or misleading information from entering an LLM’s prompt, context window, or retrieval pipeline. Key steps include:</p><ul><li><p>Check the context constantly: Monitor the conversation or retrieved text for anything suspicious or harmful, not just the starting prompt.</p></li><li><p>Use trusted sources: Score or label documents based on credibility so the system prefers reliable information and ignores low scored data.</p></li><li><p>Spot unusual data: Use tools that detect odd, out-of-place, or manipulated content, and remove it before the model uses it.</p></li><li><p>Filter inputs and outputs: Add guardrails so harmful or misleading text can’t easily enter the system or be repeated by the model.</p></li><li><p>Keep the model updated with clean data: Regularly refresh the system with verified information to counter any bad data that slipped through.</p></li><li><p>Human-in-the-loop: Have people review important outputs or compare them against known, trustworthy sources.</p></li></ul><p>Simple user habits also help, resetting long chats, sharing only relevant information, breaking complex tasks into smaller steps, and maintaining clean notes outside the model.</p><p>Together, these measures create a layered defense that protects LLMs from context poisoning and keeps outputs accurate and trustworthy.</p><p>Without countermeasures as mentioned here, an agent might remember instructions, like ignore previous guidelinesor trivial facts that an attacker inserted, leading to harmful outputs.</p><h3>Context distraction</h3><p><em>Context distraction</em> is when a context grows so long that the model overfocuses on the context, neglecting what it learned during training. In extreme cases, this resembles <a href="https://en.wikipedia.org/wiki/Catastrophic_interference"><em>catastrophic forgetting</em></a>; that is, the model effectively “forgets” its underlying knowledge and becomes overly attached to the information placed in front of it. Previous studies have shown that LLMs often lose focus when the prompt is extremely long.</p><p>The Gemini 2.5 agent, for example, supported a million-token window, but once its context grew beyond a certain point (on the order of 100,000 tokens in an experiment), it began to <strong>fixate on repeating its past actions</strong> instead of coming up with new solutions. In a sense, the agent became a prisoner of its extensive history. It kept looking at its long log of previous moves (the context) and mimicking them, rather than using its underlying training knowledge to devise fresh and novel strategies.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt91ea0056bbda6e2d/6a170f3d2b835fdd2bf4b2db/e08e5b6d2e8ec7e3511d455985eed3d7fa6241e0-1352x636.png" alt="Context distraction " /><p>This is counterproductive. We want the model to use relevant context to help reasoning, not override its ability to think. Notably, even models with huge windows exhibit this <a href="https://research.trychroma.com/context-rot"><em>context rot</em></a>: Their performance degrades nonuniformly as more tokens are added. There appears to be an <em>attention budget</em>., Like humans with limited working memory, an LLM has a finite capacity to attend to tokens, and as that budget is stretched, its precision and focus drop.</p><p>As a mitigation, you can prevent context distraction using chunking, engineering the right information, regular context summarization, and evaluation and monitoring techniques to measure the accuracy of the response using scoring.</p><p>These methods keep the model grounded in both relevant context and its underlying training, reducing the risk of distraction and improving overall reasoning quality.</p><h3>Context confusion</h3><p><em>Context confusion</em> is when superfluous content in the context is used by the model to generate a low-quality response.A prime example is giving an agent a large set of tools or API definitions that it might use. If many of those tools are unrelated to the current task, the model may still try to use them inappropriately, simply because they’re present in context. Experiments have found that providing <em>more</em> tools or documents can <em>hurt</em> performance if they’re not all needed. The agent starts making mistakes, like calling the wrong function or referencing irrelevant text. </p><p>In one case, a small <strong>Llama 3.1 8B</strong> model failed a task when given 46 tools to consider but succeeded when given only 19 tools. The extra tools created confusion, even though the context was within length limits. The underlying issue is that any information in the prompt will be <em>attended to</em> by the model. If it doesn’t know to ignore something, that something could influence its output in undesired ways. Irrelevant bits can “steal” some of the model’s attention and lead it astray (for instance, an irrelevant document might cause the agent to answer a different question than asked). Context confusion often manifests as the model producing a low-quality response that integrates unrelated context. Refer to the research paper: <a href="https://arxiv.org/pdf/2411.15399">Less is More: Optimizing Function Calling for LLM Execution on Edge Devices.</a></p><p>It reminds us that more context isn’t always better, especially if it’s not <strong>curated</strong> for relevance.</p><h3>Context clash</h3><p><em>Context clash</em> occurs when <strong>parts of the context contradict each other</strong>, causing internal inconsistencies that derail the model’s reasoning. A clash can happen if the agent accumulates multiple pieces of information that are in conflict. </p><p>For example, imagine an agent that fetched data from two sources: One says <em>Flight A departs at 5 PM</em>, and the other says <em>Flight A departs at 6 PM</em>. If both facts end up in the context, the poor model has no way to know which is correct; it may get confused or produce an incorrect or non-similar answer.</p><p>Context clash also frequently occurs in multiturn conversations where the model’s <strong>earlier attempts</strong> at answering are still lingering in the context along with later refined information.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd86976266867c0ed/6a170f3e66c4f9c785f8c105/500d7a80dc8db1923f9b5ca84728eed64fa296f7-1316x580.png" alt="Context clash" /><p>A <a href="https://arxiv.org/pdf/2505.06120">research study</a> by Microsoft and Salesforce shows that if you break a complex query into multiple chatbot turns (adding details gradually), the final accuracy drops significantly, compared to giving all details in a single prompt. Why? Because the early turns contain partial or incorrect intermediate answers from the model, and those remain in the context. When the model later tries to answer with all info, its <em>memory</em> still includes those wrong attempts, which conflict with the corrected info and lead it off track. Essentially, the conversation’s context clashes with itself. The model may inadvertently use an outdated piece of context (from an earlier turn) that doesn’t apply after new info is added.</p><p>In agent systems, context clash is especially dangerous because an agent might combine outputs from different tools or subagents. If those outputs disagree, the aggregated context is inconsistent. The agent could then get stuck or produce nonsensical results trying to reconcile the contradictions. Preventing context clash involves ensuring the context is <strong>fresh and consistent</strong>,for instance, clearing or updating any outdated info and not mixing sources that haven’t been vetted for consistency.</p><h3>Context leakage and knowledge conflicts</h3><p>In systems where multiple agents or users share a memory store, there’s a risk of information bleeding over between contexts.</p><p>For example, if two separate users’ data embeddings reside in the same vector database without proper access control, an agent answering User A’s query might accidentally retrieve some of User B’s memory. This <em><strong>cross-context leak</strong></em> can expose private information or just create confusion in responses.</p><p>According to the <a href="https://wtit.com/blog/2025/04/17/owasp-top-10-for-llm-applications-2025/">OWASP Top 10 for LLM Applications</a>, multitenant vector databases must guard against such leakage:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte433216805a66d29/6a170f404a531b2c4e36aa25/8f0ccf0b2f7bd6715c14aceee2deffb213d50bd9-1600x936.png" alt="Context leakage" /><p>According to <a href="https://wtit.com/blog/2025/04/17/owasp-top-10-for-llm-applications-2025/">LLM08:2025 Vector and Embedding Weaknesses</a><em>,</em> one of the common risks is context leakage:</p><em>In multi-tenant environments where multiple classes of users or applications share the same vector database, there's a risk of context leakage between users or queries. Data federation knowledge conflict errors can occur when data from multiple sources contradict each other. This can also happen when an LLM can’t supersede old knowledge that it has learned while training, with the new data from Retrieval Augmentation.</em><p>Another aspect is that an LLM might have trouble overriding its <strong>built-in knowledge</strong> with new info from memory. If the model was trained on some fact and the retrieved context says the opposite, the model can get confused about which to trust. Without proper design, the agent could mix up contexts or fail to update old knowledge with new evidence, leading to stale or incorrect answers.</p><h3><strong>Hallucinations and misinformation</strong></h3><p>While <em>hallucination </em>(the LLM making up plausible-sounding but false information) is a known problem even without long contexts, poor memory management can amplify it. </p><p>If the agent’s memory is lacking a crucial fact, the model may just <strong>fill in the gap with a guess</strong>, and if that guess then enters the context (poisoning it), the error persists. </p><p>The OWASP LLM security report <a href="https://wtit.com/blog/2025/04/17/owasp-top-10-for-llm-applications-2025/"><strong>(LLM09:2025 Misinformation)</strong></a> highlights misinformation as a core vulnerability: LLMs can produce confident but fabricated answers, and users may overtrust them. An agent with a bad or outdated long-term memory might confidently cite something that was true last year but is false now, unless its memory is kept up to date. </p><p>Overreliance on the AI’s output (by either the user or the agent itself in a loop) can make this worse. If no one ever checks the info in memory, the agent can accumulate falsehoods. This is why RAG is often used to reduce hallucinations: By retrieving an authoritative source, the model doesn’t have to invent facts. But if your retrieval pulls in the wrong document (say, one that contains misinformation) or if an early hallucination isn’t pruned, the system may propagate that misinformation throughout its actions. </p><p>The bottom line: Failing to manage memory can lead to <strong>incorrect and misleading outputs</strong>, which can be damaging, especially if the stakes are high (for example, bad advice in a finance or medical domain). An agent needs mechanisms to verify or correct its memory content, not just unconditionally trust whatever is in the context.</p><p>In summary, giving an AI agent an infinitely long memory or dumping every possible thing into its context is <em>not</em> a recipe for success.</p><h2>Best practices for memory management in LLM applications</h2><p>To avoid the pitfalls above, developers and researchers devised a number of <strong>best practices for managing context and memory</strong> in AI systems. These practices aim to keep the AI’s working context lean, relevant, and up to date.Here are some of the key strategies, along with examples of how they help.</p><h3>RAG: Use targeted context</h3><p>Much of RAG has already been covered in the earlier section, so this serves as a concise set of practical reminders:</p><ul><li><p>Use targeted retrieval, not bulk loading: Retrieve only the most relevant chunks instead of pushing entire documents or full conversation histories into the prompt.</p></li><li><p>Treat RAG as just-in-time memory recall: Fetch context only when it’s needed, rather than carrying everything forward across turns.</p></li><li><p>Prefer relevance-aware retrieval strategies: Approaches like top-k semantic search, Reciprocal Rank Fusion, or tool loadout filtering help reduce noise and improve grounding.</p></li><li><p>Larger context windows don’t remove the need for RAG: Two highly relevant paragraphs are almost always more effective than 20 loosely related pages.</p></li></ul><p>That said, RAG isn’t about adding more context; it’s about adding the right context.</p><h3>Tool loadout</h3><p><em>Tool loadout</em> is about giving a model only the tools it actually needs for a task. The term comes from gaming: You pick a loadout that fits the situation. Too many tools slow you down; the wrong ones cause failure. LLMs behave the same way, according to the research paper <a href="https://arxiv.org/abs/2411.15399">Less is more</a>. Once you pass ~30 tools, descriptions start overlapping and the model gets confused. Past ~100 tools, failure is almost guaranteed. This isn’t a context window problem, it’s context confusion.</p><p>A simple and effective fix is <a href="https://arxiv.org/abs/2505.03275"><strong>RAG-MCP</strong></a>. Instead of dumping every tool into the prompt, tool descriptions are stored in a vector database and only the most relevant ones are retrieved per request. In practice, this keeps the loadout small and focused, dramatically shortens prompts, and can improve tool selection accuracy by up to 3x.</p><p>Smaller models hit this wall even sooner. The research shows an 8B model failing with dozens of tools but succeeding once the loadout is trimmed. Dynamically selecting tools, sometimes with an LLM first, reasoning about what it thinks it needs, can boost performance by 44%, while also reducing power usage and latency. The takeaway is that most agents only need a few tools, but as your system grows, tool loadout and RAG-MCP become first-order design decisions.</p><h3>Context pruning: Limit the chat history length</h3><p>If a conversation goes on for many turns, the accumulated chat history can become too large to fit, leading to context overflow or becoming too distracting to the model. </p><p><em>Trimming</em> means programmatically removing or shortening less important parts of the dialogue as it grows. One simple form is to drop the oldest turns of the conversation when you hit a certain limit, keeping only the latest <em>N</em> messages. More sophisticated pruning might remove irrelevant digressions or previous instructions that are no longer needed. The goal is to <strong>keep the context window uncluttered</strong> by old news. </p><p>For example, if the agent solved a subproblem 10 turns ago and we have since moved on, we might delete that portion of the history from the context (assuming it won’t be needed further). Many chat-based implementations do this: They maintain a rolling window of recent messages. </p><p>Trimming can be as simple as “forgetting” the earliest parts of a conversation once they’ve been summarized or are deemed irrelevant. By doing so, we reduce the risk of context overflow errors and also reduce <a href="https://www.elastic.co/search-labs/blog/agentic-memory-management-elasticsearch#context-distraction"><strong>context distraction</strong></a>, so the model won’t see and get sidetracked by old or off-topic content. This approach is very similar to how humans might not remember every word from an hour-long talk but will retain the highlights. </p><p>If you’re confused about context pruning, as highlighted by the author Drew Breunig <a href="https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.html#tool-loadout:~:text=Provence%20is%20fast%2C%20accurate%2C%20simple%20to%20use%2C%20and%20relatively%20small%20%E2%80%93%20only%201.75%20GB.%20You%20can%20call%20it%20in%20a%20few%20lines%2C%20like%20so%3A">here</a>, usage of the Provence (`<a href="https://huggingface.co/naver/provence-reranker-debertav3-v1">naver/provence-reranker-debertav3-v1</a>`) model, a lightweight (1.75 GB), efficient, and accurate context pruner for question answering, can make a difference. It can trim large documents down to only the most relevant text for a given query. You can call it in specific intervals.</p><p>Here’s how we invoke the `provence-reranker` model in our code to prune the context:</p># Context pruning with Provence
def prune_with_provence(query: str, context: str, threshold: Optional[float] = None) -&gt; str:
    """
    Prune context using Provence reranker model
    
    Args:
        query: User's query/question
        context: Original context to prune
        threshold: Relevance threshold (0-1) for Provence reranker.
                   If None, uses args.pruning_threshold.
                   0.1 = conservative (recommended, no performance drop)
                   0.3-0.5 = moderate to aggressive pruning
    
    Returns:
        Pruned context with only relevant sentences
    """
    if provence_model is None:
        return context
    
    if threshold is None:
        threshold = args.pruning_threshold
    
    try:
        # Use Provence's process method
        provence_output = provence_model.process(
            question=query,
            context=context,
            threshold=threshold,
            always_select_title=False,
            enable_warnings=False
        )
        
        # Extract pruned context from output
        pruned_context = provence_output.get('pruned_context', context)
        reranking_score = provence_output.get('reranking_score', 0.0)
        
        # Log statistics
        original_length = len(context)
        pruned_length = len(pruned_context)
        reduction_pct = ((original_length - pruned_length) / original_length * 100) if original_length &gt; 0 else 0
        
        if args.verbose:
            rich.print(f"[cyan]📊 Pruning stats: {pruned_length}/{original_length} chars ({reduction_pct:.1f}% reduction, threshold={threshold:.2f}, rerank_score={reranking_score:.3f})[/cyan]")
        
        return pruned_context if pruned_context else context
        
    except Exception as e:
        rich.print(f"[yellow]⚠️ Error in Provence pruning: {str(e)}[/yellow]")
        rich.print(f"[yellow]⚠️ Falling back to original context[/yellow]")
        return context<p>We use the Provence reranker model (`naver/provence-reranker-debertav3-v1`) to score sentence relevance. Threshold-based filtering keeps sentences above the relevance threshold. Also, we introduce a fallback mechanism, where we return to the original context if pruning fails. Finally, statistics logging tracks reduction percentage in verbose mode.</p><h3>Context summarization: Condense older information instead of dropping it entirely</h3><p><em>Summarization</em> is a companion to trimming. When the history or knowledge base becomes too large, you can employ the LLM to generate a brief summary of the important points and use that summary in place of the full content going forward, as we performed in our code above.</p><p>For example, if an AI assistant has had a 50-turn conversation, instead of sending all 50 turns to the model on turn 51 (which likely won’t fit), the system might take turns 1–40, have the model summarize them in a paragraph, and then only supply that summary plus the last 10 turns in the next prompt. This way, the model still knows what was discussed without needing every detail. Early chatbot users did this manually by asking, “Can you summarize what we’ve talked about so far?” and then continuing in a new session with the summary. Now it can be automated. Summarization not only saves context window space but can also reduce <strong>context confusion/distraction</strong> by stripping away extra detail and retaining just the salient facts.</p><p>Here’s how we use OpenAI models (you can use any LLMs) to condense context while preserving all relevant information, eliminating redundancy and duplication.
</p># Context summarization
def summarize_context(query: str, context: str) -&gt; str:
    """
    Summarize context using LLM to reduce duplication and focus on relevant information
    
    Args:
        query: User's query/question
        context: Context to summarize
        
    Returns:
        Summarized context
    """
    try:
        summary_prompt = f"""You are an expert at summarizing conversation context.

Your task: Analyze the provided conversation context and produce a condensed summary that fully answers or supports the user's specific question.

The summary must:
1. Preserve every fact, detail, and information that directly relates to the question
2. Eliminate redundancy and duplicate information
3. Maintain chronological flow when relevant
4. Focus on information that helps answer: "{query}"

Context to summarize:
{context}

Provide a concise summary that preserves all relevant information:"""

        summary = llm.invoke(summary_prompt).content
        
        if args.verbose:
            original_length = len(context)
            summary_length = len(summary)
            reduction_pct = ((original_length - summary_length) / original_length * 100) if original_length &gt; 0 else 0
            rich.print(f"[cyan]📝 Summarization stats: {summary_length}/{original_length} chars ({reduction_pct:.1f}% reduction)[/cyan]")
        
        return summary
        
    except Exception as e:
        rich.print(f"[yellow]⚠️ Error in context summarization: {str(e)}[/yellow]")
        rich.print(f"[yellow]⚠️ Falling back to original context[/yellow]")
        return context<p>Importantly, when the context is summarized, the model is less likely to get overwhelmed by trivial details or past errors (assuming the summary is accurate). </p><p>However, summarization has to be done carefully. A bad summary might omit a crucial detail or even introduce an error. It’s essentially another prompt to the model (“summarize this”), so it can hallucinate or lose nuance. Best practice is to summarize incrementally and perhaps keep some canonical facts unsummarized.</p><p>Nonetheless, it has proven very useful. <a href="https://storage.googleapis.com/deepmind-media/gemini/gemini_v2_5_report.pdf">In the Gemini agent scenario, </a>summarizing the context every ~100k tokens was a way to counteract the model’s tendency to repeat itself. The summary acts like a compressed memory of the conversation or data. As developers, we can implement this by having an agent periodically call a summarization function (maybe a smaller LLM or a dedicated routine) on the conversation history or a long document. The resulting summary replaces the original content in the prompt. This tactic is widely used to keep contexts within limits and distill the information.</p><h3>Context quarantine: Isolate contexts when possible</h3><p>This is more relevant in complex agent systems or multistep workflows. The idea of context segmentation is to split a big task into smaller, isolated tasks, each with its own context, so that you never accumulate one enormous context that contains everything. Each subagent or subtask works on a piece of the problem with a focused context, and then a higher-level agent, or supervisor or coordinator integrates the results.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt09d1eac7442aea2b/6a170f42dc55deb10de00ea7/f2de68c3339883d7658e633af3948f29f427e6cf-1600x900.png" alt="Context quarantine" /><p><a href="https://www.anthropic.com/engineering/multi-agent-research-system">Anthropic’s research strategy uses multiple subagents</a>, each investigating a different aspect of a question, with their own context windows, and a lead agent that reads the distilled results from those subagents. This parallel, modular approach means that no single context window gets too bloated. It also reduces the chance of irrelevant information mixing, each thread stays on topic (no context confusion), and it doesn’t carry unnecessary baggage when answering its specific subquestion. In a sense, it’s like running separate threads of thought that only share their outcomes, not their entire thought process.</p><p>In multi-agent systems, this approach is essential. If Agent A is handling task A and Agent B is handling task B, there’s no reason for either agent to consume the other’s full context unless it’s truly required. Instead, agents can exchange only the necessary information. For example, Agent A can pass a consolidated summary of its findings to Agent B via a supervisor agent, while each subagent maintains its own dedicated context thread. This setup doesn’t require human-in-the-loop intervention; it relies on a supervisory agent with enabled tools with minimal and controlled context sharing.</p><p>Nonetheless, designing your system so that agents or tools operate with minimal necessary context overlap can greatly enhance clarity and performance. Think of it as <strong>microservices for AI</strong>, each component deals with its context, and you pass messages between them in a controlled way, instead of one monolithic context.These best practices are often used in combination. Also, this gives you the flexibility to trim trivial history, summarize important older messages or conversations, offload the detailed logs to Elasticsearch for long-term context, and use retrieval to bring back anything relevant when needed.</p><p>As mentioned <a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents#:~:text=While%20some%20models,to%20the%20LLM">here</a>, the guiding principle is that context is a limited and precious resource. You want every token in the prompt to earn its keep, meaning it should contribute to the quality of the output. If something in memory is not pulling its weight (or worse, actively causing confusion), then it should be pruned, summarized, or kept out.</p><p>As developers, we can now program the context just like we program code, deciding what information to include, how to format it, and when to omit or update it. By following these practices, we can give LLM agents the much-needed context to perform tasks without falling victim to the failure modes described earlier. The result is agents that remember what they should, forget what they don’t need, and retrieve what they require just in time.</p><h2>Conclusion</h2><p>Memory isn’t something you add to an agent; it’s something you engineer. Short-term memory is the agent’s working scratch pad, and long-term memory is its durable knowledge store. RAG is the bridge between the two, turning a passive datastore, like Elasticsearch, into an active recall mechanism that can ground outputs and keep the agent current.</p><p>But memory is a double-edged sword. The moment you let context grow unchecked, you invite poisoning, distraction, confusion, and clashes, and in shared systems, even data leakage. That’s why the most important memory work isn’t “store more,” it’s “curate better”: Retrieve selectively, prune aggressively, summarize carefully, and avoid mixing unrelated contexts unless the task truly demands it.</p><p>In practice, good context engineering looks like good systems design: smaller, sufficient contexts, controlled interfaces between components, and a clear separation between raw and the distilled state you actually want the model to see. Done right, you don’t end up with an agent that remembers everything - you end up with an agent that remembers the right things, at the right time, for the right reason.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agentic-memory-management-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agentic-memory-management-elasticsearch</guid>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Someshwaran Mohankumar]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3bad6b045392e641/6a170f43a29299c189d010cc/80907fd072e72d6ec902470b449c9f337957a0d7-1280x720.png" length="0" type="image/png"/>
    <pubDate>Fri, 16 Jan 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Implementing an agentic reference architecture with Elastic Agent Builder and MCP]]></title>
    <description><![CDATA[Explore an agentic reference architecture with Elastic Agent Builder, MCP, and semantic search to build a security agent for automated threat analysis.]]></description>
    <content:encoded><![CDATA[<p>In this article, we will present a reference architecture for using Elasticsearch with AI capabilities through the <a href="https://www.elastic.co/docs/solutions/search/elastic-agent-builder">Elastic Agent Builder</a>, exposing an <a href="https://modelcontextprotocol.io/docs/getting-started/intro">MCP server</a> to access Agent Builder tools and Elasticsearch data.</p><p>Model Context Protocol (<a href="https://modelcontextprotocol.io/docs/getting-started/intro">MCP</a>) is an open-source standard that enables applications and LLMs to communicate with external systems via <a href="https://modelcontextprotocol.io/specification/2025-06-18/server/tools">MCP tools</a> (programmatic capabilities), and <a href="https://docs.langchain.com/oss/python/langgraph/overview">LangGraph</a> (an extension of <a href="https://docs.langchain.com/oss/javascript/langchain/overview">LangChain</a>) provides the orchestration framework for these agentic workflows.</p><p>We’ll implement an application that can search both internal knowledge (Elasticsearch stored data) and external sources (on the internet) to identify potential and known vulnerabilities related to a specific tool. The application will gather the information and generate a detailed summary of the findings.</p><h2>Requirements</h2><ul><li><p>Elasticsearch 9.2</p></li><li><p>Python 3.1x</p></li><li><p><a href="https://platform.openai.com/api-keys">OpenAI API Key</a></p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/api-keys/elasticsearch-api-keys">Elasticsearch API Key</a></p></li><li><p><a href="https://serpapi.com/users/sign_up?plan=free">Serper API Key</a></p></li></ul><h2>Elastic Agent Builder</h2><p><a href="https://www.elastic.co/docs/solutions/search/elastic-agent-builder">Elastic Agent Builder</a> is a set of AI-powered capabilities for developing and integrating agents that can interact with your Elasticsearch data. It provides a built-in agent that can be used for natural language conversations with your data or instance, and it also supports tool creation, Elastic APIs, A2A, and MCP. In this article, we will focus on using the <a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">MCP server</a> for external access to the Elastic Agent Builder tools.</p><p>To know more about Agent Builder features, you can read <a href="https://www.elastic.co/search-labs/blog/elastic-ai-agent-builder-context-engineering-introduction">this article</a>.</p><h3>Agent Builder MCP feature</h3><p>The <a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server">MCP server</a> is available in the Agent Builder and can be accessed at:</p>{KIBANA_URL}/api/agent_builder/mcp
# Or if you are using a custom Kibana space:
{KIBANA_URL}/s/{SPACE_NAME}/api/agent_builder/mcp<p>The Agent Builder offers <a href="https://www.elastic.co/docs/solutions/search/agent-builder/tools#built-in-tools">Built-in tools</a>, and you can also create your <a href="https://www.elastic.co/docs/solutions/search/agent-builder/tools#custom-tools">custom tools</a>.</p><h2>Reference architecture</h2><p>To get a complete overview of the elements used by an agentic application in an end-to-end workflow, let’s look at the following diagram:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7a1c664318e2c848/6a170bde964cea3c6908bbe8/c5bbba345340bfe5571b17d53b5896d4a3235eac-4720x2560.png" alt="Agent Builder MCP feature reference architecture." /><p>Elasticsearch is at the center of this architecture, functioning as a vector store, providing the embeddings generation model, and also serving the MCP server to access the data via tools. To better explain the workflow, let’s look at the ingestion and the Agent Builder layer separately.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb616cae7400ea5f0/6a170be0dc55debba1e00e27/97a0075ae637d64140ec7ff0d167297723675632-3000x1176.png" alt="Elasticsearch at the center of the architecture, functioning as a vector store, providing the embeddings generation model, and also serving the MCP server to access the data via tools." /><p>Here, the first element is the data that will be stored in Elasticsearch. The data passes through an ingest pipeline, where it is processed by the Elasticsearch ELSER model to generate embeddings and then stored in Elasticsearch.</p><h3>Elastic Agent Builder layer</h3><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt33175b496b636661/6a170be2dc55de7daae00e2b/9bb396bbd4c3baa3be26f9d9e386f4d5405132ab-2180x2560.png" alt="The agent builder layer where the Agent Builder plays a central role by exposing the tools needed to interact with the Elasticsearch data." /><p>On this layer, the Agent Builder plays a central role by exposing the tools needed to interact with the Elasticsearch data. It manages the tools that operate over Elasticsearch indices and makes them available for consumption. Then <a href="https://docs.langchain.com/oss/python/langchain/overview">LangChain</a> handles the orchestration via the MCP client.</p><p>This architecture allows Agent Builder to work as one of many MCP servers available to the client so that the Elasticsearch agent builder can combine with other MCPs. This way, the MCP client can ask cross-source questions and then combine the answers.</p><h2>Use case: Security vulnerability agent</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d7291c0c39e3dbe/6a170be4ab7f08b2bedb9ec4/1b46b29a8cde4645ebaec1f747be4f6888dd8d39-1600x906.png" alt="Agent builder and MCP use case. Building a security vulnerability agent." /><p>The security vulnerability agent identifies potential risks based on a user’s question by combining three complementary layers:</p><p><strong>First</strong>, it performs a <a href="https://www.elastic.co/docs/solutions/search/semantic-search">semantic search</a> with embeddings over an internal knowledge base of past incidents, configurations, and known vulnerabilities to retrieve relevant historical evidence.</p><p><strong>Second</strong>, it searches the internet for newly published recommendations or threat intelligence that may not yet exist internally.</p><p><strong>Finally</strong>, an LLM correlates and prioritizes both internal and external findings, evaluates their relevance to the user’s specific environment, and produces a clear explanation along with potential mitigation steps.</p><h2>Developing the application</h2><p>The application’s code can be found in the attached <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-reference-architecture-for-agentic-applications/notebook.ipynb">notebook</a>.</p><p>You can see the setup for the Python application below:</p># load environment variables
load_dotenv()

ELASTICSEARCH_ENDPOINT = os.getenv("ELASTICSEARCH_ENDPOINT")
ELASTICSEARCH_API_KEY = os.getenv("ELASTICSEARCH_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
SERPER_API_KEY = os.getenv("SERPER_API_KEY")
KIBANA_URL = os.getenv("KIBANA_URL")

INDEX_NAME = "security-vulnerabilities"
KIBANA_HEADERS = {
    "kbn-xsrf": "true",
    "Content-Type": "application/json",
    "Authorization": f"ApiKey {ELASTICSEARCH_API_KEY}",
} # Useful for Agent Builder API calls


es_client = Elasticsearch(ELASTICSEARCH_ENDPOINT, api_key=ELASTICSEARCH_API_KEY) # Elasticsearch client<p>We need to access Agent Builder and create one agent specialized in security queries and one tool to perform semantic search. You need to have the<a href="https://www.elastic.co/docs/solutions/search/agent-builder/get-started"> Agent Builder </a><a href="https://www.elastic.co/docs/solutions/search/agent-builder/get-started"><strong>enabled</strong></a> for the next step. Once it’s on, we’ll use the <a href="https://www.elastic.co/docs/solutions/search/agent-builder/kibana-api#tools">tools API</a> to create a tool that will perform a semantic search.</p>security_search_tool = {
    "id": "security-semantic-search",
    "type": "index_search",
    "description": "Search internal security documents including incident reports, pentests, internal CVEs, security guidelines, and architecture decisions. Uses semantic search powered by ELSER to find relevant security information even without exact keyword matches. Returns documents with severity assessment and affected systems.",
    "tags": ["security", "semantic", "vulnerabilities"],
    "configuration": {
        "pattern": INDEX_NAME,
    },
}

try:
    response = requests.post(
        f"{KIBANA_URL}/api/agent_builder/tools",
        headers=KIBANA_HEADERS,
        json=security_search_tool,
    )

    if response.status_code == 200:
        print("✅ Security semantic search tool created successfully")    
    else:
        print(f"Response: {response.text}")
except Exception as e:
    print(f"❌ Error creating tool: {e}")<p>Configure your tools following the <a href="https://www.elastic.co/docs/solutions/search/agent-builder/tools#best-practices">best practices</a> defined by Elastic for developing Tools. Once created, this tool will be ready to use in the Kibana UI.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt57d9fb62f55979e7/6a170be6509168a2a9e1bb0d/5e5b3282dea07987613d8e8d35c372ca68820e44-1600x381.png" alt="Configuring tools following the best practices defined by Elastic for developing Tools." /><p>With the tool created, we can start writing the code for the ingestion workflow:</p><h3>Ingest pipeline</h3><p>To define the data structure, we need to have a <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-reference-architecture-for-agentic-applications/dataset.json">dataset</a> prepared for ingestion. Below is a sample document for this example:</p>{
    "title": "Incident Report: Node.js Express 4.17 Prototype Pollution RCE",
    "content": "In March 2024, our production Node.js Express 4.17 API gateway experienced a critical prototype pollution vulnerability leading to remote code execution. The attack vector involved manipulating object prototypes through JSON payloads in POST requests. This affected all Express middleware processing user input. Immediate mitigation: upgrade to Express 4.18.2+, implement input validation, use Object.freeze() for critical objects. Related to CVE-2022-24999.",
    "doc_type": "incident_report",
    "severity": "critical",
    "affected_systems": [
      "api-gateway-prod",
      "api-gateway-staging"
    ],
    "date": "2024-03-15"
}<p>For this type of document, we will use the following index mappings:</p>index_mapping = {
    "mappings": {
        "properties": {
            "title": {"type": "text", "copy_to": "semantic_field"},
            "content": {"type": "text", "copy_to": "semantic_field"},
            "doc_type": {"type": "keyword", "copy_to": "semantic_field"},
            "severity": {"type": "keyword", "copy_to": "semantic_field"},
            "affected_systems": {"type": "keyword", "copy_to": "semantic_field"},
            "date": {"type": "date"},
            "semantic_field": {"type": "semantic_text"},
        }
    }
}

if es_client.indices.exists(index=INDEX_NAME) is False:
    es_client.indices.create(index=INDEX_NAME, body=index_mapping)
    print(f"✅ Index '{INDEX_NAME}' created with semantic_text field for ELSER")
else:
    print(f"ℹ️  Index '{INDEX_NAME}' already exists, skipping creation")<p>We are creating a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">semantic_text</a> field to perform semantic search using the information from the fields marked with the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/copy-to">copy_to</a> property.</p><p>With that mapping definition, we can ingest the data using the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk">bulk API</a>.</p>def build_bulk_actions(documents, index_name):
    for doc in documents:
        yield {"_index": index_name, "_source": doc}


try:
    with open("dataset.json", "r") as f:
        security_documents = json.load(f)

    success, failed = helpers.bulk(
        es_client,
        build_bulk_actions(security_documents, INDEX_NAME),
        refresh=True,
    )
    print(f"📥 {success} documents indexed successfully")

except Exception as e:
    print(f"❌ Error during bulk indexing: {str(e)}")<h3>LangChain MCP client</h3><p>Here we’re going to create an MCP client using LangChain to consume the Agent Builder tools and build a workflow with LangGraph to orchestrate the client execution. The first step is to <a href="https://www.elastic.co/docs/solutions/search/agent-builder/mcp-server#configuring-mcp-clients">connect to the MCP server</a>:</p>client = MultiServerMCPClient(
    {
        "agent-builder": {
            "transport": "streamable_http",
            "url": MCP_ENDPOINT,
            "headers": {"Authorization": f"ApiKey {ELASTICSEARCH_API_KEY}"},
        }
    }
)

tools = await client.get_tools()

print(f"📋 MCP Tools available: {[t.name for t in tools]}") # ['platform_core_search',  ... 'security-semantic-search']<p>Next, we create an agent that selects the appropriate tool based on the user input:</p>reasoning = {"effort": "low"}

llm = ChatOpenAI(
    model="gpt-5.2-2025-12-11", reasoning=reasoning, openai_api_key=OPENAI_API_KEY
) # LLM client 

agent = create_agent(
    llm,
    tools=tools,
    system_prompt="""You are a cybersecurity expert specializing in infrastructure security.

        Your role is to:
        1. Analyze security queries from users
        2. Search internal security documents (incidents, pentests, CVEs, guidelines)
        3. Provide actionable security recommendations
        4. Assess vulnerability severity and impact

        When responding:
        - Always search internal documents first using the agent builder tools
        - Provide specific, technical, and actionable advice
        - Cite relevant internal incidents and documentation
        - Assess severity (critical, high, medium, low)
        - Recommend immediate mitigation steps

        Be concise but comprehensive. Focus on practical security guidance.""",
)<p>We’ll use the GPT-5.2 model, which represents OpenAI’s state-of-the-art for agent management tasks. We configure it with low reasoning effort to achieve faster responses compared to the medium or high settings, while still delivering high-quality results by leveraging the full capabilities of the GPT-5 family. You can read more about the GPT 5.2 <a href="https://openai.com/index/introducing-gpt-5-2/">here</a>.</p><p>Now that the initial setup is done, the next step is to define a workflow capable of making decisions, running tool calls, and summarizing results.</p><p>For this, we use LangGraph. We won’t cover LangGraph in depth here; <a href="https://www.elastic.co/search-labs/blog/ai-agent-workflow-finance-langgraph-elasticsearch">this article</a> provides a detailed overview of its functionality.</p><p>The following image shows a high-level view of the LangGraph application.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte293b7cf62f54f8e/6a170be7964cea816908bbec/729295115427ec981a594e873245fa541dd977aa-332x531.png" alt="High-level view of the LangGraph application." /><p>We need to define the application state:</p>class AgentState(TypedDict):
    query: str
    agent_builder_response: dict
    internet_results: list
    final_response: str
    needs_internet_search: bool<p>To better understand how the workflow operates, here is a brief description of each function. For full implementation details, refer to the accompanying <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-reference-architecture-for-agentic-applications/notebook.ipynb">notebook</a>.</p><ul><li><p><strong>call_agent_builder_semantic_search:</strong> Queries internal documentation using the Agent Builder MCP server and also stores the retrieved messages in the state.</p></li><li><p><strong>decide_internet_search:</strong> Analyzes the internal results and determines whether an external search is required.</p></li><li><p><strong>perform_internet_search: </strong>Runs an external search using the <a href="https://serper.dev/">Serper</a> API when needed.</p></li><li><p><strong>generate_response:</strong> Correlates internal and external findings and produces a final, actionable cybersecurity analysis for the user.</p></li></ul><p>With the workflow defined, we can now send a query:</p>query = "We are using Node.js with Express 4.17 for our API gateway. Are there known prototype pollution or remote code execution vulnerabilities?"<p>In this example, we want to evaluate whether this specific version of Express is affected by known vulnerabilities.</p><h4>Research results</h4><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltac164b2086d23589/6a170be9a29299162cd01057/b18a31e42bcd8f4d86bb605f85d4ff77135b0855-1084x517.png" alt="Elastic agent builder and MCP security agent research results." /><p>See the complete response in <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-reference-architecture-for-agentic-applications/notebook.ipynb">this file</a>.</p><p>This response clearly correlates internal and internet findings and provides actionable mitigation steps. It successfully highlights the severity of the vulnerability and offers a structured, security-oriented summary.</p><h3>Extensions and future enhancements</h3><p>This architecture is modular and allows us to extend its capabilities by replacing, improving, or adding components to the existing list. We could add another agent, consumed by the same MCP client. We can also use an automated ingestion workflow with tools such as Logstash, Kafka, or <a href="https://www.elastic.co/docs/reference/search-connectors/self-managed-connectors">Elastic self-managed connectors.</a> Feel free to change the LLM, the MCP client framework, or the embeddings model or add more tools depending on your needs.</p><h2>Conclusion</h2><p>This reference architecture shows a practical way to combine Elasticsearch, the Agent Builder, and MCP to build an AI-driven application. Its structure keeps each part independent, which makes the system easy to implement, maintain, and extend.</p><p>You can start with a simple setup (like the security use case in this article) and scale it by adding new tools, data sources, or agents as your needs grow. Overall, it provides a straightforward path for building flexible and reliable agentic workflows on top of Elasticsearch.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-builder-mcp-reference-architecture-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-builder-mcp-reference-architecture-elasticsearch</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[AI Tools ]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt22bfbe4b04ea2e92/6a170beb60084b717d3c4597/33a57e3f61f9095c99b6d1499175a6edb0d5dfc5-4720x2560.png" length="0" type="image/png"/>
    <pubDate>Wed, 07 Jan 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Building an AI agent for HR with Elastic Agent Builder and GPT-OSS]]></title>
    <description><![CDATA[Discover how to build an AI agent that can answer natural language queries about your employee HR data using Elastic Agent Builder and GPT-OSS.]]></description>
    <content:encoded><![CDATA[<h2>Introduction</h2><p>This article will show you how to build an AI agent for HR using <a href="https://openai.com/index/introducing-gpt-oss/">GPT-OSS</a> and Elastic Agent Builder. The agent can answer your questions without sending data to OpenAI, Anthropic, or any external service.</p><p>We’ll use LM Studio to serve GPT-OSS locally and connect it to Elastic Agent Builder.</p><p>By the end of this article, you’ll have a custom AI agent that can answer natural language questions about your employee data while maintaining full control over your information and model.</p><h2>Prerequisites</h2><p>For this article, you need:</p><ul><li><p><a href="https://www.elastic.co/cloud">Elastic Cloud</a> hosted 9.2, serverless or <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">local</a> deployment</p></li><li><p>Machine with 32GB RAM recommended (minimum 16GB for GPT-OSS 20B)</p></li><li><p><a href="https://lmstudio.ai/">LM Studio</a> installed</p></li><li><p><a href="https://www.docker.com/products/docker-desktop/">Docker Desktop</a> Installed</p></li></ul><h2>Why use GPT-OSS?</h2><p>With a local LLM you have the control to deploy it in your own infrastructure and fine-tune it to fit your own needs. All this while maintaining control over the data that you share with the model, and of course, you don’t have to pay a license fee to an external provider.</p><p>OpenAI <a href="https://openai.com/index/introducing-gpt-oss/">released GPT-OSS</a> on August 5, 2025, as part of their commitment to the open model ecosystem.</p><p>The 20B parameter model offers:</p><ul><li><p><strong>Tool use capabilities</strong></p></li><li><p><strong>Efficient inference</strong></p></li><li><p><strong>OpenAI SDK compatible</strong></p></li><li><p><strong>Compatible with agentic workflows</strong></p></li></ul><p>Benchmark comparison:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt58fab956edb40412/6a170cfcb0367da43a72bd80/29160e3345352088e8213297630882f252b00c47-1600x680.png" alt="" /><h2>Solution architecture</h2><p>The architecture runs entirely on your local machine. Elastic (running in Docker) communicates directly with your local LLM through LM Studio, and the Elastic Agent Builder uses this connection to create custom AI agents that can query your employee data.</p><p>For more details, refer to this <a href="https://www.elastic.co/docs/solutions/observability/connect-to-own-local-llm">documentation</a>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt80db5bb0a797f51b/6a170cfd0e2e492f2c41a16f/a4a886750ff25fa8bb7aefc7448161e52cf73ed3-1600x896.png" alt="" /><h2>Building an AI agent for HR: Steps</h2><p>We’ll divide the implementation into 5 steps:</p><ol><li><p>Configure LM studio with a local model</p></li><li><p>Deploy Local Elastic with Docker</p></li><li><p>Create the OpenAI connector in Elastic</p></li><li><p>Upload employee data to Elasticsearch</p></li><li><p>Build and test your AI Agent</p></li></ol><h2>Step 1: Configure LM Studio with GPT-OSS 20B</h2><p>LM Studio is a user-friendly application that allows you to run large language models locally on your computer. It provides an OpenAI-compatible API server, making it easy to integrate with tools like Elastic without a complex setup process. For more details, refer to the <a href="https://lmstudio.ai/docs/app">LM Studio Docs</a>.</p><p>First, download and install <a href="https://lmstudio.ai/">LM Studio</a> from the official website. Once installed, open the application.</p><h3>In the LM Studio interface:</h3><ol><li><p>Go to the search tab and search for “GPT-OSS”</p></li><li><p>Select the <code>openai/gpt-oss-20b</code> from OpenAI</p></li><li><p>Click download</p></li></ol><p>The size of this model should be approximately <strong>12.10GB</strong>. The download may take a few minutes, depending on your internet connection.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2dc341a6625e34b7/6a170cff839dfa2eb4dcff44/5d01bc4dcb377b5259fc6b521fe2425a31b90ca4-1312x872.png" alt="" /><h4>Once the model is downloaded:</h4><ol><li><p>Go to the local server tab</p></li><li><p>Select the openai/gpt-oss-20b</p></li><li><p>Use the default port 1234</p></li><li><p>On the right panel, go to <strong>Load </strong>and set the Context Length to <strong>40K</strong> or higher</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3704ca1b28465cc4/6a170d00d7c022ed8fde64ef/e546033f916381647b876815b2c1f1ae2a08365f-326x337.png" alt="" /><p>5. Click start server</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7b9170a4945ff857/6a170d0266c4f9ffadf8c0a6/28ee78a3caa84d14e04db3d42f30acbe4d4d005a-1312x872.png" alt="" /><p>You should see this if the server is running.</p>[LM STUDIO SERVER] Success! HTTP server listening on port 1234
[LM STUDIO SERVER] Supported endpoints:
[LM STUDIO SERVER] -&gt;	GET  http://localhost:1234/v1/models
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/responses
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/chat/completions
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/completions
[LM STUDIO SERVER] -&gt;	POST http://localhost:1234/v1/embeddings
Server started.<h2>Step 2: Deploy Local Elastic with Docker</h2><p>Now we’ll set up Elasticsearch and Kibana locally using Docker. Elastic provides a convenient script that handles the entire setup process. For more details refer to the <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">official documentation</a>.</p><h3>Run the start-local script</h3><p>Execute the following command in your terminal:</p>curl -fsSL https://elastic.co/start-local | sh<p>This script will:</p><ul><li><p>Download and configure Elasticsearch and Kibana</p></li><li><p>Start both services using Docker Compose</p></li><li><p>Automatically activate a 30-day Platinum trial license</p></li></ul><h3>Expected output</h3><p>Just wait for the following message and save the password and API key shown; you’ll need them to access Kibana:</p>🎉 Congrats, Elasticsearch and Kibana are installed and running in Docker!
🌐 Open your browser at http://localhost:5601
   Username: elastic
   Password: KSUlOMNr
🔌 Elasticsearch API endpoint: http://localhost:9200
🔑 API key: cnJGX0pwb0JhOG00cmNJVklUNXg6cnNJdXZWMnM4bncwMllpQlFlUTlWdw==
Learn more at https://github.com/elastic/start-local<h3>Access Kibana</h3><p>Open your browser and navigate to:</p>http://localhost:5601<p>Log in using the credentials obtained in the terminal output.</p><h3>Enable Agent Builder</h3><p>Once logged in to Kibana, navigate to <strong>Management </strong>&gt;<strong> AI </strong>&gt;<strong> Agent Builder </strong>and activate the Agent Builder.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0a934bd99fa6a0ce/6a170d046234e019c3db1a5a/92e104cb846c20d875865ded8a3d37f5c7daae9b-1491x1528.png" alt="" /><h2>Step 3: Create the OpenAI connector in Elastic</h2><p>Now we’ll configure Elastic to use your local LLM.</p><h3>Access Connectors</h3><ol><li><p>In Kibana</p></li><li><p>Go to <strong>Project Settings</strong> &gt; <strong>Management</strong></p></li><li><p>Under <strong>Alerts and Insights</strong>, select <strong>Connectors</strong></p></li><li><p>Click Create Connector</p></li></ol><h3>Configure the connector</h3><p>Select <strong>OpenAI</strong> from the list of connectors. LM Studio uses the OpenAI SDK, making it compatible.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt762023c39781eb78/6a170d06a29299a59ed01087/5ac87042e086c7a2bd47a8039e646ec831f0dcc6-923x974.png" alt="" /><p>Fill in the fields with these values:</p><ul><li><p><strong>Connector name: </strong>LM Studio - GPT-OSS 20B</p></li><li><p><strong>Select an OpenAI provider: </strong>Other (OpenAI Compatible Service)</p></li><li><p><strong>URL: </strong><code>http://host.docker.internal:1234/v1/chat/completions</code></p></li><li><p><strong>Default model: </strong>openai/gpt-oss-20b</p></li><li><p><strong>API Key:</strong> testkey-123 (any text works, because LM Studio Server doesn't require authentication.)</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt980e595f80e2be2e/6a170d086f7f0468a19148cc/2084ac32fcf1fb810c8b54ecab1c85a1e3e8905b-672x1302.png" alt="" /><p>To finish the configuration, click <strong>Save &amp; test</strong>.</p><p><strong>Important:</strong> Toggle ON the “<strong>Enable native function calling</strong>”; this is required for the Agent Builder to work properly. If you don’t enable this, you’ll get a <strong><code>No tool calls found in the response</code></strong> error.</p><h3>Test the connection</h3><p>Elastic should automatically test the connection. If everything is configured correctly, you’ll see a success message like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4d2e815dd558f881/6a170d090e2e49076541a177/f567d767f1969c4730c1daa92f651789dc3742ac-1042x812.png" alt="" /><p>Response:</p>{
  "status": "ok",
  "data": {
    "id": "chatcmpl-flj9h0hy4wcx4bfson00an",
    "object": "chat.completion",
    "created": 1761189456,
    "model": "openai/gpt-oss-20b",
    "choices": [
      {
        "index": 0,
        "message": {
          "role": "assistant",
          "content": "Hello! 👋 How can I assist you today?",
          "reasoning": "Just greet.",
          "tool_calls": []
        },
        "logprobs": null,
        "finish_reason": "stop"
      }
    ],
    "usage": {
      "prompt_tokens": 69,
      "completion_tokens": 23,
      "total_tokens": 92
    },
    "stats": {},
    "system_fingerprint": "openai/gpt-oss-20b"
  },
  "actionId": "ee1c3aaf-bad0-4ada-8149-118f52dad757"
}<h2>Step 4: Upload employee data to Elasticsearch</h2><p>Now we’ll upload the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/gpt-oss-with-elasticsearch/hr-employees-bulk.json">HR employee dataset</a> to demonstrate how the agent works with sensitive data. I generated a fictional dataset with this structure.</p><h3>Dataset structure</h3>{
  "employee_id": "0f4dce68-2a09-4cb1-b2af-6bcb4821539b",
  "full_name": "Daffi Stiebler",
  "email": "lscutchings0@huffingtonpost.com",
  "date_of_birth": "1975-06-20T15:39:36Z",
  "hire_date": "2025-07-28T00:10:45Z",
  "job_title": "Physical Therapy Assistant",
  "department": "HR",
  "salary": "108455",
  "performance_rating": "Needs Improvement",
  "years_of_experience": 2,
  "skills": "Java",
  "education_level": "Master's Degree",
  "manager": "Carl MacGibbon",
  "emergency_contact": "Leigha Scutchings",
  "home_address": "5571 6th Park"
}<h3>Create the index with mappings</h3><p>First, create the index with proper mappings. Note that we’re using <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">semantic_text</a> fields for some key fields; this enables semantic search capabilities for our index.</p>​​PUT hr-employees
{
  "mappings": {
    "properties": {
      "@timestamp": {
        "type": "date"
      },
      "employee_id": {
        "type": "keyword"
      },
      "full_name": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "email": {
        "type": "keyword"
      },
      "date_of_birth": {
        "type": "date",
        "format": "iso8601"
      },
      "hire_date": {
        "type": "date",
        "format": "iso8601"
      },
      "job_title": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "department": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "salary": {
        "type": "double"
      },
      "performance_rating": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "years_of_experience": {
        "type": "long"
      },
      "skills": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "education_level": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "manager": {
        "type": "text",
        "copy_to": "employee_semantic"
      },
      "emergency_contact": {
        "type": "keyword"
      },
      "home_address": {
        "type": "keyword"
      },
      "employee_semantic": {
        "type": "semantic_text"
      }
    }
  }
}<h3>Index with Bulk API</h3><p>Copy and paste the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/gpt-oss-with-elasticsearch/hr-employees-bulk.json">dataset</a> into your Dev Tools in Kibana and execute it:</p>POST hr-employees/_bulk
{"index": {}}
{"employee_id": "57728b91-e5d7-4fa8-954a-2384040d3886", "full_name": "Filide Gane", "email": "vhallahan1@booking.com", "job_title": "Business Systems Development Analyst", "department": "Marketing", "salary": "$52330.27", "performance_rating": "Meets Expectations", "years_of_experience": 12, "skills": "Java", "education_level": "Bachelor's Degree", "date_of_birth": "2000-02-07T16:49:32Z", "hire_date": "2023-11-07T13:03:16Z", "manager": "Freedman Kings", "emergency_contact": "Vilhelmina Hallahan", "home_address": "75 Dennis Junction"}
{"index": {}}
{"employee_id": "...", ...}<h3>Verify the data</h3><p>Run a query to verify:</p>GET hr-employees/_search<h2>Step 5: Build and test your AI agent</h2><p>With everything configured, it’s time to build a custom AI agent using Elastic Agent Builder. For more details refer to the <a href="https://www.elastic.co/docs/solutions/search/agent-builder/get-started">Elastic documentation</a>.</p><h3>Add the connector</h3><p>Before we can create our new agent, we have to set our Agent builder to use our custom connector called <code>LM Studio - GPT-OSS 20B</code> because the default one is the <a href="https://www.elastic.co/docs/reference/kibana/connectors-kibana/elastic-managed-llm">Elastic Managed LLM</a>. For that, we need to go to <strong>Project Setting</strong> &gt; <strong>Management</strong> &gt; <strong>GenAI Settings</strong>; now we select the one we created and click <strong>Save</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc42f079c5e756057/6a170d0acf4f2501d9b2d1c7/11e830c3e2fb4c298b020c928fa5422f3397ba08-1600x1152.png" alt="" /><h3>Access Agent Builder</h3><ol><li><p>Go to <strong>Agents</strong></p></li><li><p>Click on <strong>Create a new agent</strong></p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb8e734817c5a7c6a/6a170d0ca929cf867cae0a34/c1e60541563650163f972ac9088dc1ed1de759a7-1600x1054.png" alt="" /><h3>Configure the agent</h3><p>To create a new agent, the required fields are the <strong>Agent ID</strong>, <strong>Display Name</strong>, and <strong>Display Instructions</strong>.</p><p>But there are more customization options, like the Custom Instructions that guide how your agent is going to behave and interact with your tools, similar to a system prompt, but for our custom agent. Labels help organize your agents, avatar color, and avatar symbol.</p><p>The ones that I chose for our agent based on the dataset are:

<strong>Agent ID:</strong> <code>hr_assistant</code></p><p><strong>Custom instructions:</strong></p>You are an HR Analytics Assistant that helps answer questions about employee data.
When responding to queries:
- Provide clear, concise answers
- Include relevant employee details (name, department, salary, skills)
- Format monetary values with currency symbols
- Be professional and maintain data confidentiality<p>
Labels: <code>Human Resources</code> and <code>GPT-OSS</code></p><p>Display name: <code>HR Analytics Assistant</code></p><p>Display description:</p>A specialized AI assistant for Human Resources that helps analyze employee data, compensation, performance metrics, and talent management. Ask questions about employees, departments, salaries, or performance analytics.<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt23fb011e5b4f4d49/6a170d0e7d8d67f47a70e77f/f94bb2bf08497e5e756ca76b30a3a51f42927756-1424x1217.png" alt="" /><p>With all the data in there, we can click on <strong>Save</strong> our new agent.</p><h3>Test the agent</h3><p>Now you can ask natural language questions about your employee data, and GPT-OSS 20B will understand the intent and generate an appropriate response.</p><h4>Prompt:</h4>Which employee is the one with the highest salary in the hr-employees index?<h4>Answer:</h4><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc0c52faacf63b583/6a170d0f0e2e497bfd41a17b/94ad19f80b96304028a59f60beca51dfc9aecc8a-899x631.png" alt="" /><p>The Agent process was:</p><p>1. Understand your question using the GPT-OSS connector</p><p>2. Generate the appropriate Elasticsearch query (using the built-in tools or custom <a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a>)</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte32a8a7e6363c7f2/6a170d115091680077e1bb44/6f2961d0d1b97475f6dda300acee84da540938e6-844x466.png" alt="" /><p>3. Retrieve matching employee records</p><p>4. Present results in natural language with proper formatting</p><p>Unlike traditional lexical search, the agent powered by GPT-OSS understands intent and context, making it easier to find information without knowing exact field names or query syntax. For more details on the agent's thinking process, refer to this <a href="https://www.elastic.co/search-labs/blog/ai-agent-builder-experiments-performance">article</a>.</p><h2>Conclusion</h2><p>In this article, we built a custom AI agent using Elastic’s Agent Builder to connect to the OpenAI GPT-OSS model running locally. By deploying both Elastic and the LLM on your local machine, this architecture allows you to leverage generative AI capabilities while maintaining full control over your data, all without sending information to external services.</p><p>We used GPT-OSS 20B as an experiment, but the officially recommended models for Elastic Agent Builder are referenced <a href="https://www.elastic.co/docs/solutions/search/agent-builder/models#recommended-models">here</a>. If you need more advanced reasoning capabilities, there's also the <a href="https://huggingface.co/openai/gpt-oss-120b">120B parameter variant</a> that performs better for complex scenarios, though it requires a higher-spec machine to run locally. For more details, refer to the <a href="https://openai.com/open-models/">official OpenAI documentation</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/build-an-ai-agent-hr-elastic-agent-builder-gpt-oss</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/build-an-ai-agent-hr-elastic-agent-builder-gpt-oss</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Tomás Murúa]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt664f490053e46e6b/6a170d13b0367d2d7e72bd84/05d2d0513fff67d975f9223d75108aa9f50646bc-1600x914.png" length="0" type="image/png"/>
    <pubDate>Wed, 26 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Getting started with Elastic Agent Builder and Microsoft Agent Framework]]></title>
    <description><![CDATA[Walk through the complete process of creating an agent with Elastic Agent Builder and then explore how to use the agent via the A2A protocol orchestrated with the Microsoft Agent Framework.]]></description>
    <content:encoded><![CDATA[<p>Elastic <a href="https://www.elastic.co/blog/whats-new-elastic-9-2-0">9.2</a> was recently released and includes a new feature called <a href="https://www.elastic.co/elasticsearch/agent-builder">Agent Builder</a>. It enables developers to quickly create AI agents and tools powered by data stored in Elasticsearch. Any tools or agents you create in Agent Builder can be utilized immediately within your own custom AI apps.</p><p>In this blog post we’ll walk through all the steps to use Elastic Agent Builder to create an agent. Then we’ll walk through the process of running an example Python app that uses the Microsoft Agent Framework to orchestrate your Elastic agent.</p><h2>Create an Elastic Serverless project</h2><p>To use Agent Builder you need an Elastic deployment or an Elastic serverless project, so let’s begin by creating an Elastic serverless project. Go to <a href="https://cloud.elastic.co/registration">Elastic Cloud</a> and create a new Elasticsearch Serverless project.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt26ec5e33540a6a05/6a170c6c67045b1ffb45c23a/05da6b45ca88b70181028f394bdcc2c289ca68da-1677x952.gif" alt="elastic-agent-builder-gif" /><h2>Create an index and add data</h2><p>Now that we’ve got an Elastic project, let’s create an index, which is what Elasticsearch uses to store data. Open Developer Tools in Elastic Cloud where we can run a command to create an index.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc53e2236abb894aa/6a170c6d66c4f9dbaaf8c082/ac31098e0d557c7758f180d497b86c90ff50cf66-1976x1099.png" alt="elastic-agent-builder-add-data" /><p>Copy the following PUT command which creates an index named <em>my-docs </em>containing a mixture of fields, and our content leveraging <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">semantic search</a>.</p><p></p>PUT /my-docs
{
  "mappings": {
    "properties": {
      "title": { "type": "text" },
      "content": { 
        "type": "semantic_text"
      },
      "filename": { "type": "keyword" },
      "last_modified": { "type": "date" }
    }
  }
}<p>Paste the PUT command into the input area of the Developer Tools console. Hover your mouse over the command in the console and then click the <strong>Run</strong> button to execute the command.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8393cfc3184ed7f9/6a170c6f4a531b8e2436a9a5/e7f426fd9a5ad6909f81af1246fe84726b7d6596-1980x1103.png" alt="elastic-agent-builder-send-request" /><p>The next step is to add some data to the <em>my-docs</em> index that you just created. Copy and paste the following command into the Develop Tools console.</p>PUT /my-docs/_doc/greetings-md
{
  "title": "Greetings",
  "content": "
# Greetings

## Basic Greeting
Hello!

## Helloworld Greeting
Hello World! 🌎

## Not Greeting
I'm only a greeting agent. 🤷

",
  "filename": "greetings.md",
  "last_modified": "2025-11-04T12:00:00Z"
}<p>Click the command’s <strong>Run </strong>button to execute the command which will add a document to the <em><code>my-docs</code></em> index.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfca5c697878a7669/6a170c71c1e8a5fa58f88308/a8224c4379c88cfb720cb110d13b1c3c27291fc3-1999x1247.png" alt="elastic-agent-builder-run" /><p>As you can see, the command above adds a document named <em>greetings.md</em> that includes the contents of different potential types of greeting responses.</p><p>Now that we’ve got some data in an Elastic index, let’s get a confirmation of what data we have to work with. Using the power of the built-in Elastic AI Agent that is enabled by default in Agent Builder, you can now have a chat about your data. Select <strong>Agents</strong> in the navigation menu.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt12c35b2ef94e2d50/6a170c734a531b34fc36a9a9/5f2ab858f9cb73c40b6ca70c8da60f6d7417db74-1970x1266.png" alt="elastic-agent-builder-agents" /><p>Then simply ask, “What data do I have?”</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4469224308944d2c/6a170c74a929cf8e3cae0a0e/4311fb41a114e932ceacb4d6bb535263cad57479-1708x938.gif" alt="elastic-agent-builder-gif-data" /><p>The default Elastic AI Agent provides a nice summary of the data currently stored in Elastic.</p><h2>Create a tool</h2><p>The next step on this walkthrough journey is to create an agent that can utilize the data stored in Elastic.</p><p>As you’ve seen the default agent in Elastic Agent builder is already useful for chatting with your data but to really give agents custom powers, they need access to tools via the <a href="https://modelcontextprotocol.io/docs/getting-started/intro">Model Context Protocol</a> (MCP). Agent Builder has fully featured tool creation and management functionality that you can use to quickly create custom MCP tools that are hosted right in the same scalable Elastic project as your data.</p><p>Let’s create a tool in Elastic Agent Builder that can access the data now stored in Elastic. Click <strong>+ New</strong> to start a new chat.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltea95e23f23f46050/6a170c76e8fbceb11f39fc9d/f5af8fdb8738b07eceba130e49fdf97478d65646-1636x414.png" alt="elastic-agent-builder" /><p>Then click on <strong>Manage tools</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt09343b252bf64c54/6a170c787d8d675a9e70e766/b8a29be0d6c8fa07deb2c523585c3a6bc67153b0-1999x992.png" alt="elastic-agent-builder-manage-tools" /><p>Click the <strong>+ New tool</strong> button.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd076c54cbb46af32/6a170c7a2b835f80caf4b254/92a91e94cfa071b66761aaa81e48f7b3962cdaea-1970x1128.png" alt="elastic-agent-builder-new-tool" /><p>In the <strong>Create Tool</strong> form, select the <strong>ES|QL </strong>as the tool <strong>Type</strong> and enter the following values.</p><p>For <strong>Tool ID</strong>:</p>example.get_greetings<p>For <strong>Description</strong>:</p>Get greetings doc from Elasticsearch my_docs index.<p>For <strong>Configuration </strong>enter the following query into the <strong>ES|QL Query </strong>text area:</p><p>Your completed <strong>Create a new tool</strong> form should look like the following completed form. Click <strong>Save</strong> to create the tool.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb9361be43f00ef0d/6a170c7c1949f7784ee7aa7f/51698c74174fd6963101eebb7ebe720902209eb5-1406x1271.png" alt="elastic-agent-builder-create-tool" /><h2>Create an Agent and assign it a tool</h2><p>Ah! There’s that feeling of having a new tool and being ready to use it. Agents need tools to give them special abilities beyond what general LLMs can provide and we’ve now got a brand new tool. Let’s create an agent that can put our tool to good use. Select <strong>Agents</strong> in the navigation menu.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1516be4466ab7b94/6a170c7dc1e8a514e5f8830c/f6770cbf2047fed5a5827bfa9f24a3489a1f7deb-1400x500.png" alt="elastic-agent-builder-tools" /><p>Click <strong>Create a new agent</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfdc817e7fe832623/6a170c7f286714389293e359/8fe9bbb5296118c3c044aa943a08f9ec78c82173-1400x763.png" alt="elastic-agent-builder-create-agent" /><p>Based on the name of the tool and the data it’s accessing, you’ve probably already guessed that we’re going to be creating a greeting agent and you’re right! Let’s create a Hello World agent right now.</p><p>In the <strong>New Agent</strong> form, enter the following values.</p><p>For <strong>Agent ID </strong>enter the text:</p>helloworld_agent<p>In the <strong>Custom Instructions </strong>text area enter the following instructions:</p>If the prompt contains greeting text like "Hi" or "Hello" then respond with only the Basic Hello text from your documents.

If the prompt contains the text “Hello World” then respond with only the Hello World text from your documents.

In all other cases where the prompt does not contain greeting words, then respond with only the Not Greeting text from your documents.<p>For <strong>Display name </strong>enter the text:</p>HelloWorld Agent<p>For the <strong>Display description </strong>enter the text:</p>An agent that responds to greetings.<p>Your completed <strong>New Agent</strong> form should look like the following completed form. The next step is to assign the agent the tool we created in the previous step. Click the <strong>Tools </strong>tab.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt116ae96fcd2183a1/6a170c81a929cf7c3cae0a12/e189613957fa86016710764e665a2cd11e98d401-1400x1303.png" alt="elastic-agent-builder-agent-tools" /><p>Select only the <em><code>example.get_greetings</code></em> tool that we created previously. Unselect all the other available tools. This will configure the agent being created to only have access to the tool we’ve created.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltef927cf26297de7c/6a170c83d7c022ee16de64d1/5e1f95fa27afe30c402da5ffde385774e1bc8b5a-1999x1550.png" alt="elastic-agent-builder-example-tool" /><p>Click <strong>Save</strong> to create the agent.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltefbde78961805855/6a170c848b73cbcb2118a080/20e7e03a8e749597f612e0f0087ff2541b3cb930-1999x545.png" alt="elastic-agent-builder-save-new-agent" /><p>You’ll be taken to the Agents list where you can see that the new HelloWorld Agent has been created.We can quickly test out our new agent right inside Agent Builder. Select <strong>Agents</strong> in the navigation menu.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc210a95ca0a13a0e/6a170c8650916808b3e1bb26/e16a5ab26003c2e5b5fc8516427bd2f9b9737a12-1999x704.png" alt="elastic-agent-builder-agent-list" /><p>Select the <strong>HelloWorld Agent</strong> from the Agent Chat agent selector. Enter the prompt “hello world” and you should get back the Hello World text from the <em><code>greetings.md</code></em> document stored in the <em><code>my-docs</code></em> Elastic index.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt69eda206b8d87e06/6a170c87dc55de1414e00e38/2cc98878450a7d3b26ebbfbcc5f9b841b8971c41-1191x566.gif" alt="elastic-agent-builder-gif-hello-world" /><p>Well done. Now that we know our agent is working as expected, let’s explore the immediate development benefit that you get with tools and agents created in Agent Builder. Any tools you create in Agent Builder are usable via MCP by any agent-building platform that supports MCP. Also, any agents you create in Agent Builder are available for use in any agent-building platform that supports the <a href="https://a2a-protocol.org">AgentToAgent</a> (A2A) protocol.</p><h2>Microsoft Agent Framework</h2><p>If you’re interested in trying out new Agent development tools, then there’s a recently announced open-source development kit called the <a href="https://learn.microsoft.com/en-us/agent-framework/overview/agent-framework-overview">Microsoft Agent Framework</a> that you should definitely try out for yourself. The Agent Framework allows you to use the A2A protocol to orchestrate agentic apps that can combine multiple agents running on different hosts to enable solutions that aren’t possible with only a generic GenAI Large Language Model. The Agent Framework is available in Python and C#. Let’s see how we can use the Python-based Agent Framework to call the custom Elastic Agent we just created.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4eeacb3bf7e3256d/6a170c89286714037193e35f/6428e470f3323c2a88c20e126969939a7b616a83-1844x414.png" alt="microsoft-agent-framework" /><h2>Getting started with the Agent Framework in Python</h2><p>Let’s run some code! On your local computer open <a href="https://code.visualstudio.com/download">Visual Studio Code</a> and open a new terminal.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9455a01413b8c277/6a170c8b0c4857291201aad5/11f2ea916bf277e39c98701c8d31e251fcdf6a8b-956x571.png" alt="new-terminal" /><p>In the open terminal, clone the Elastic Search Labs source code repository which contains the <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/agent-builder-a2a-agent-framework">Elastic Agent Builder A2A example app</a>.</p>git clone https://github.com/elastic/elasticsearch-labs<p>In the terminal, cd to change directory to elasticsearch-labs.</p>cd elasticsearch-labs<p>In the terminal, enter the following command to open the current folder in the Visual Studio Code editor.</p>code .<p>In the Visual Studio File Explorer, expand the <code>supporting-blog-content</code> and <code>agent-builder-a2a-agent-framework</code> folders and then open the file named <em>elastic_agent_builder_a2a.py</em> in the text editor.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt38aafb5bbf2b4274/6a170c8dd7c0222ecede64d5/ae1f173f953e6d805cdc2b5ab756c229b3b31793-1428x1044.png" alt="agent-builder-code" /><p>Here’s the contents of <em>elastic_agent_builder_a2a.py </em>that you should see in your text editor:</p>import asyncio
from dotenv import load_dotenv
import httpx
import os
from a2a.client import A2ACardResolver
from agent_framework.a2a import A2AAgent


async def main():
    load_dotenv()
    a2a_agent_host = os.getenv("ES_AGENT_URL")
    a2a_agent_key = os.getenv("ES_API_KEY")

    print(f"Connection to Elastic A2A agent at: {a2a_agent_host}")

    custom_headers = {"Authorization": f"ApiKey {a2a_agent_key}"}

    async with httpx.AsyncClient(timeout=60.0, headers=custom_headers) as http_client:
        # Resolve the A2A Agent Card
        resolver = A2ACardResolver(httpx_client=http_client, base_url=a2a_agent_host)
        agent_card = await resolver.get_agent_card(
            relative_card_path="/helloworld_agent.json"
        )
        print(f"Found Agent: {agent_card.name} - {agent_card.description}")

        # Use the Agent
        agent = A2AAgent(
            name=agent_card.name,
            description=agent_card.description,
            agent_card=agent_card,
            url=a2a_agent_host,
            http_client=http_client,
        )
        prompt = input("Enter Greeting &gt;&gt;&gt; ")
        print("\nSending message to Elastic A2A agent...")
        response = await agent.run(prompt)
        print("\nAgent Response:")
        for message in response.messages:
            print(message.text)


if __name__ == "__main__":
    asyncio.run(main())<p>The code within the main() method demonstrates how to control your Elastic Agent Builder agent using the Agent Framework. It creates an <code>http_client</code> using a URL and API key for the agent which you’ll provide from your Elastic project. Then the Agent Framework’s A2ACardResolver is called with that <code>http_client</code> to get your agent’s A2A agent card based on the <code>relative_card_path</code> of “<code>/helloworld_agent.json</code>” to reference your agent’s <strong>Agent ID </strong>which is “helloworld_agent”. The code then uses the Agent Framework to invoke your agent with the A2A agent card. The final part of the main() method prompts the user of the app for input of a “greeting” and then sends the user input as a prompt to your agent. Based on the instructions and tools specified when you created your agent, the agent’s response is displayed to the app user.</p><h2>Setting your agent URL and API Key as environment variables</h2><p>Make a copy of the file <em>env.example</em> and name the new file <em>.env</em> Edit the newly created <em>.env</em> file to set the values of the environment variables to use specific values copied from your Elastic project.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc8d702407ed2eaec/6a170c8fa6c2b9798be79751/6420978cf4e3edd5ff148dca47554995da3e3f22-1428x603.png" alt="" /><p>First we’ll replace <strong>&lt;YOUR-ELASTIC-AGENT-BUILDER-URL&gt;</strong> with the Agent URL path that you can copy from your Elastic project’s Agent Builder - Tools page. Back in Elastic Agent Builder click <strong>Agents </strong>in the navigation menu.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt09a8b629754c9cae/6a170c90e8fbce3f4739fca6/530ebafbc6327f24cb1a94b7c94d205339db6d28-1191x321.png" alt="" /><p>Select <strong>Manage tools</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt09343b252bf64c54/6a170c787d8d675a9e70e766/b8a29be0d6c8fa07deb2c523585c3a6bc67153b0-1999x992.png" alt="" /><p>Click the <strong>MCP Server</strong> dropdown at the top of the Tools page. Select <strong>Copy MCP Server URL.</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt69a60e413e418922/6a170c9260084b20cf3c45ae/3a9d5bf730b9541013db2b72601202d6a76e15f1-1977x1002.png" alt="" /><p>Back in Visual Studio Code, within the <em>.env file</em>, find where the placeholder text “<strong>&lt;YOUR-ELASTIC-AGENT-BUILDER-URL&gt;</strong>” appears and paste in the copied <strong>MCP Server URL </strong>to replace the placeholder text. Now edit the pasted <strong>MCP Server URL</strong>. Delete the text “mcp” at the end of the URL and replace it with the text “a2a”. The edited URL should look something like this:</p>https://example-project-a123.kb.westus2.azure.elastic.cloud/api/agent_builder/a2a<p>The next placeholder text to replace in the <em>.env</em> file is <strong>&lt;YOUR-ELASTIC-API-KEY&gt;.</strong> We’ll replace it with an actual API Key from your Elastic project. Back in your Elastic project, click <strong>Elasticsearch</strong> in the navigation menu.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta81babf678fe00c8/6a170c946234e0dd7edb1a32/11696609b354e75aa8110987b0d476634ac6b322-1965x663.png" alt="" /><p>Click <strong>Create API key</strong> to create a new API key.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6b6476fdcb6d6715/6a170c966f7f04b6479148a6/67b15b2db4e43d6abe3ec42f3e8692f953aa4731-1995x1038.png" alt="" /><p>Enter a <strong>Name</strong> for the API key and click <strong>Create API key</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta9f36eefdc509494/6a170c98a929cf56c8ae0a16/5e49445a77dc9730dd967f2aa8f8f11f4911eb41-1999x1076.png" alt="" /><p>Click the <strong>copy</strong> button to copy the API key.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5b99693a8a71414d/6a170c9aab7f082955db9ee3/f5b6993bea36493cf5d80867df607a71c2c24bcf-1971x1038.png" alt="" /><p>Back in Visual Studio Code, within the <em>.env</em> file , find where the placeholder text “<strong>&lt;YOUR-ELASTIC-API-KEY&gt;</strong>” appears and paste in the copied API Keyvalueto replace the placeholder text.</p><p>Now we can save the changes we’ve made to the <em>.env</em> file. The edited file should look something like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1fa4f6a82f1d6656/6a170c9ba2929993a6d0107a/91636964285003a3dda01ce43214b74c33492393-1428x601.png" alt="" /><h2>Run the example app</h2><p>It’s time to run the code. To do so, open a new terminal in Visual Studio Code. Click the <strong>Terminal</strong> top level menu and select <strong>New Terminal</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9455a01413b8c277/6a170c8b0c4857291201aad5/11f2ea916bf277e39c98701c8d31e251fcdf6a8b-956x571.png" alt="" /><p>In the new terminal, <code>cd</code> to change directory to the directory containing the agent-<code>builder-a2a-agent-framework</code> example app.</p>cd elasticsearch-labs/supporting-blog-content/agent-builder-a2a-agent-framework<p>In the terminal, create a Python virtual environment by running the following code.</p>python -m venv .venv<p>Activate the virtual environment by running the following command (based on your operating system) in the terminal window:</p><ul><li><p>If you’re running MacOS or Linux, the command to activate the virtual environment is:</p></li></ul>source .venv/bin/activate<ul><li><p>If you’re on Windows, the command to activate the virtual environment is:</p></li></ul>.venv\Scripts\activate<p>The code in the <em>elastic_agent_builder_a2a.py</em> file is powered by the Microsoft Agent Framework and we still need to install it, so let's do that now. Run the following <em>pip</em> command to install the Python based Agent Framework along with its necessary Python packages:</p>pip install -r requirements.txt<p>Hurray! Everything is now in its right place. It’s time for the good feeling fireworks…let’s run it. Run the example code by entering the following command into the terminal:</p>python elastic_agent_builder_a2a.py<p>You should see the agent framework connect to the Elastic Agent. When prompted for a greeting, enter “hello world”. You should see the HelloWorld Agent’s response → Hello World! 🌎</p><p>Top-notch work!</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt725fda509b411bf7/6a170c9dab7f0862ffdb9ee7/f77903e5bcaa0f52bed80d5c8ea23e7c538561d6-1703x1027.gif" alt="" /><p>Building agents and connecting them to tools in Agent Builder gets you immediate operability with the latest agent development platforms like the Microsoft Agent Framework. You now know how to create an Elastic agent and put it to use as a scalable relevant data source, ready to provide custom context to all the AI apps you’ll be building next.</p><p>Try <a href="https://cloud.elastic.co/registration?utm_source=agentic-ai-category&amp;utm_medium=search-labs&amp;utm_campaign=agent-builder">Elastic</a> for free and build some agents today!</p><p>
</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-builder-a2a-with-agent-framework</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-builder-a2a-with-agent-framework</guid>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Jonathan Simon]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt725fda509b411bf7/6a170c9dab7f0862ffdb9ee7/f77903e5bcaa0f52bed80d5c8ea23e7c538561d6-1703x1027.gif" length="0" type="image/gif"/>
    <pubDate>Fri, 21 Nov 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[CI/CD pipelines with agentic AI: How to create self-correcting monorepos]]></title>
    <description><![CDATA[How our team introduced GenAI into CI pipelines to create self-correcting pull requests, automizing the update of hundreds of dependencies in large monorepos]]></description>
    <content:encoded><![CDATA[<p>At Elastic Control Plane, the team behind <a href="http://cloud.elastic.co/">cloud.elastic.co</a> (Elastic Cloud Hosted) and Elastic Cloud Enterprise, we have introduced agentic AI technology into our build pipelines, giving our codebases the self-healing capabilities: Just like axolotls can grow limbs, our Pull Request builds fix themselves. This article shows why we needed to take this step, how we designed and executed it, what we learnt, and the impact of this change in our daily work.</p><p>Large codebases and their maintainers are like organisms, and any change can make them sick: Breaking builds, unit tests, etc. Maintainers, like antibodies, quickly jump in to restore health by removing the resulting bugs, a process that takes time and energy.</p><p>These codebases are built on the shoulders of hundreds of dependencies. We keep them all up to date on the products we host or distribute. This is aligned with our security standards, but comes with significant work volume generated from high rates of update-fix cycles. Each update is a potential germ that can break the build, getting the organism “sick”.</p><h2>Traditional automation to deal with a big problem: Keeping dependencies up-to-date</h2><p>This post is about automation. In this day and age, there is an important distinction to make between:</p><ul><li><p>Traditional automation: Steps performed by software encoding algorithms, driving the process being automated. It is deterministic, and its functionality is limited to what the developer intended the software to do.</p></li><li><p>Generative AI (gen AI) automation: Steps performed by Large Language Models technology from inputs in natural language and driven by prompts. Its results are usually non-deterministic and require human supervision.</p></li></ul><p>Back to the codebase-maintainers as organisms analogy, let’s talk about choosing our experiment subject. And yes, this post is about an experiment. An experiment that was so successful that it started helping our teams before becoming a fully polished internal feature.</p><p>We maintain a considerably large monorepo. We make sure dependencies are up-to-date and free of known vulnerabilities.</p><p>With about 500 actively updated dependencies for our core services, checking and bumping dependency versions is a non-trivial challenge. If performed manually, capable of taking away hundreds of engineering hours a week.</p><p>We adopted the dependency updates management system provided by our internal Engineering Productivity team as traditional automation. <a href="https://www.elastic.co/blog/reducing-cves-in-elastic-container-images">It is based on Renovate</a>.</p><p>Renovate is a simple but effective bot. In a nutshell, it does two things:</p><ol><li><p>Compares our repository dependency catalogs with the upstream packages repositories.</p></li><li><p>Opens Pull Requests when newer versions are found.</p></li></ol><p>In about six months of operation, Renovate authored pull -requests have already bumped 41% of our dependencies, becoming one of our most prolific contributors:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt393c19b2bc601973/6a170ce31949f70216e7aa8b/6b2834659619615131bcc22b3dc530f0349bef30-769x262.png" alt="" /><h2>Self-healing Pull Requests: Fix-Approve-Merge</h2><p>Our integration with Renovate was successful. So much so that it raised the bar enough to swamp the team with PR reviews. Of course, it can be set up to throttle down, but we really want to keep <em>everything</em> up-to-date<em>. </em>It made us aware that such a standard came with an effort price that could have a detrimental impact on project progression time.</p><p>Renovate PRs, as any other in our repository, go through a build pipeline that compiles code, runs unit tests (UTs) and integration tests (ITs), builds Docker images, and publishes them. All this is done using the Buildkite platform with Gradle build steps as demonstrated in the following screenshot:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc827642220d2c761/6a170ce51949f73b3ce7aa8f/b3ed789481006f9b5280a33790c8d3c59a7b519f-1372x579.png" alt="" /><p>Renovate is set up to automatically merge Pull Requests when ITs pass successfully.</p><p>One could expect that most dependency bumps would just go through, especially for patch versions. So, where is the toil coming from? Well, we live in a world where dependency maintainers introduce breaking changes in patch or minor releases, or even worse, where they introduce subtle runtime changes that are not reflected in the interfaces at all (e.g, ZooKeeper’s 3.8.3-&gt;3.8.4 new ACL constraints in existence checks operations - <a href="https://issues.apache.org/jira/browse/ZOOKEEPER-2590">ZOOKEEPER-2590</a> and downstream consequences in Apache Curator clients <a href="https://issues.apache.org/jira/browse/CURATOR-715">CURATOR-715</a>). Here is where the build breaks and humans <strong>are</strong> required.</p><h3>Really, humans?</h3><p>We noticed that these broken PRs were the real bottleneck for automatic updates. Broken as in not compiling at all, failing UTs or ITs. Engineers on our team needed to chime in on these cases invariably. This is an unplanned interruption of work, bringing frequent and production-killing context switches.</p><p>These fixes:</p><ul><li><p>Are usually self-contained: They don’t require software rearchitecture, just adjustments to the changes in the updated libraries. Nor do they require deep creative work.</p></li><li><p>They provide fast feedback loops: Edit-compile-test.</p></li></ul><p>Can you think of an emergent technology targeting repetitive self-contained coding tasks?</p><p>What if each broken PR to review came with proposed code changes fixing it?</p><p>This is exactly the approach we decided to try in an experimentation week. And it worked!</p><h3>But, how?</h3><p>The idea is simple: follow the natural way of working with your workmates. Let AI chime in and contribute to PR branches. That is, integrate AI agents as part of PRs’ workflow.</p><p>This is also an approach pursued in the coding agents industry, with examples like <a href="https://github.blog/news-insights/product-news/github-copilot-meet-the-new-coding-agent/">GitHub Copilot Coding Agent</a> or <a href="https://docs.anthropic.com/en/docs/claude-code/github-actions">Claude Code GitHub actions</a>.</p><p>Both off-the-shelf solutions make it easy to get the agent to interact with the code, but are less flexible.</p><p>We wanted the agent to act on the build errors in a targeted way: This unit test failed with this error for this module, fix that concrete problem, and add a commit to the working branch when you have fixed it.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte01dbaa9ffa77055/6a170ce71949f706c4e7aa93/e52476e705353a99e3446e90855c8b8def5c31ba-898x428.png" alt="" /><p>For that, we needed to:</p><ul><li><p>Be able to feed the AI agent with the concrete error messages and failed build tasks. Including those depending on internal Elastic Cloud services.</p></li><li><p>Give it the agency to run and iterate over failed Gradle steps until a solution is found or it desists.</p></li></ul><p>As in most systems, the key to human Software Engineers’ productivity is fast iteration cycles: Edit-Compile-Test, and that’s what we decided with code editing AI agents.</p><p>For that, we just expanded the set of build steps we described at the beginning of this post with another. A replica of the Gradle build steps with the twist of being controlled by a coding agent with a quite specific prompt:</p> # This step is used to fix compilation and UT failures when builds fail. It uses Claude Code to analyze the build logs and suggest fixes that
  # are then applied to the original branch in the form of commits once verified to have fixed the build. Only branches in elastic/cloud repository
  # can get these commits. If the source branch lives in a different repository, a new branch with the suggested branches will be posted in elastic/cloud.
  #
  # The effect is that automatic updates PR issued by Renovate can self heal. The moment Claude commits to elastic/cloud's PR branch, the build pipeline will be restarted
  # with the fixes.
  # Take into account that these commits are only added if Claude can verify that the fix the build by running the initially broken Gradle tasks.
  # To save time, computational resources and genAI tokens, this step is first run for AMD64 architecture. Chances are that the pushed fixes will also work for ARM64 architecture.
  #
  # NOTE: This step is currently using Claude Code agent but this latter could be replaced by any other genAI agent that can be invoked in headless mode from the command line.
  #
  - label: ":github: :terminal: :gradle: Claude Fix Build (AMD64)"
    key: "claude-fix-build-amd64"
    depends_on: "publishPlatformIndependent-amd64"
    allow_dependency_failure: true
    command: |
      if [ $$(buildkite-agent step get "outcome" --step "publishPlatformIndependent-amd64") != "passed" ]; then
        echo "--- Trying to fix build with Claude Code"
        .buildkite/scripts/claude-fix-build.sh
      else
        echo "--- There is nothing to fix"
      fi
    timeout_in_minutes: 240
    ...
    ...
    ...<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb25d0410d012aea3/6a170ce8c1e8a5c961f88322/3519248466ca50df4939dffe98ab7e1de0b3076b-1448x750.png" alt="" /><p><code>claude-fix-build.sh</code> is the script that invokes the agent through several steps:</p><p>1. Run a pre-hook command, grabbing the credentials we need to interact with Claude and GitHub.</p><p>2. Clone the GitHub target repository.</p><p>3. Get ready to run Gradle build steps. That is preparing the environment for Gradle to be able to build the working code. To this point, this script is just a copy of the scripts we were already using for the regular build steps in the Buildkite workers: <code>publishPlatformIndependent-amd64</code> and <code>publishPlatformIndependent-arm64</code>.</p><p>4. Use Buildkite CLI to try to fetch the previous step, the building step that failed, thus triggering the fix log. As we’ll see soon, this is a performance booster: The AI agent will analyze the file, deduce which build steps failed, and iterate over fixing loops for them.</p># Obtain the result of the build step that failed thus triggering the Claude fix
echo "--- Obtaining Gradle log from the failed build step"
mkdir /tmp/previous_step_artifacts
buildkite-agent artifact download "tmp/gradle_*.log" /tmp/previous_step_artifacts/

if [ -f /tmp/gradle.log ]; then
  echo "Found Gradle logs from previous job, Claude will use them to analyze the build failure:"
  ls -l /tmp/previous_step_artifacts
else
  echo "Couldn't find previous stage gradle log, Claude will run the build commands to evaluate what is failing"
fi<p>5. Install and configure the Claude Code agent CLI tool:</p>echo "--- Installing Claude"
sudo apt update -y
sudo apt install -y curl

curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.3/install.sh | bash
# shellcheck disable=SC1091
\. "$HOME/.nvm/nvm.sh"
nvm install 22
npm install -g @anthropic-ai/claude-code

echo "Configuring Claude Code and running environment"

# This value needs to be passed with the --allowedTools parameter to claude run
export CLAUDE_ALLOWED_TOOLS='Bash,Bash(chmod:*),Bash(git:*),Bash(./gradlew:*),Edit,NotebookEdit,MultiEdit,View,GlobTool,GrepTool,BatchTool,Write,WebFetch,WebSearch'<p>6. Prepare the agent actions log file and replicate its content on stdout (more about this below):</p># Prepare Claude actions log file and replicate its contents in stdout
touch /tmp/claude-actions.log
tail -f /tmp/claude-actions.log &amp;<p>7. Set the prompt, the soul of this integration (we are skipping the details here because they deserve their own section in the post: <a href="https://docs.google.com/document/d/17UFfcn3RKRdzOjcT_iANoBxXuc1TzoGBM_ZMk9oEVBQ/edit?tab=t.0#heading=h.ljo06oa5znpl">The prompt</a>), along with the repository CLAUDE.md file.</p># Claude fix prompt string
CLAUDE_FIX_PROMPT=$(cat &lt;&lt; 'EOF'
The build is failing, you might find a Gradle log ...
...
...
EOF
)<p>8. The original motivation for this feature is to help us fix Pull Requests bumping dependency versions. Pull Requests created by the Renovate version management bot. Our Renovate set-up rebases its PRs branches as it detects changes in upstream main branches. We observed that this behaviour interrupted the agent’s work: More often than not, it was verifying that it fixes with a long-running Gradle task just to have the build canceled by the latest change in the main branch. To avoid this situation, we leveraged the <a href="https://docs.renovatebot.com/configuration-options/#stopupdatinglabel">Renovate “stop updating labels” feature</a> so it would stop pushing changes to branches of PRs with a pre-configured label. In our case, <code>stop-updating</code> .
This is the magical spot where you can see traditional automation shaking hands with generative AI automation to reach a common goal. Our integration is telling Renovate: “Hey, I am in charge now”.</p>function add-pr-label() {
  local label="$1"
  echo "Adding label '$label' to the PR"
  if ! curl -f -X POST \
    -H "Authorization: token $GITHUB_TOKEN" \
    -H "Accept: application/vnd.github.v3+json" \
    -H "Content-Type: application/json" \
    -d "{\"labels\":[\"$label\"]}" \
    https://api.github.com/repos/elastic/cloud/issues/"$BUILDKITE_PULL_REQUEST"/labels; then
    echo "Failed to add label '$label' to the PR"
    exit 1
  fi
}

...
...
...

echo "--- Block further updates from Renovate until the PR is fixed"

# This is done adding a 'stop-updating' label to the PR (https://docs.renovatebot.com/configuration-options/#stopupdatinglabel)
add-pr-label "stop-updating"<p>9. With the next step, this integration is going to start pushing AI-generated commits to a branch candidate, merging a branch that is likely set to be auto-merged upon successful builds. Our team has a core principle: Never commit AI work without human supervision. Therefore, the script makes sure that GitHub PR auto-merge is disabled:</p>echo "--- Making sure auto-merge is disabled for the PR when there are AI contributions"

# Get GraphQL PR node id
PR_NODE_ID=$(curl -s -X POST \
    -H "Authorization: bearer $GITHUB_TOKEN" \
    -H "Content-Type: application/json" \
    -d "{\"query\":\"query { repository(owner: \\\"elastic\\\", name: \\\"cloud\\\") { pullRequest(number: $BUILDKITE_PULL_REQUEST) { id } } }\"}" \
    https://api.github.com/graphql | jq '.data.repository.pullRequest.id' -r
)

# Disable automerge
AUTOMERGE_FAILURE_MSG="Failed to disable auto-merge for the PR. Aborting: It is dangerous to allow auto-merge when there are AI contributions."

if ! AUTOMERGE_NO_ERRORS=$(curl -s -f -X POST \
  -H "Authorization: bearer $GITHUB_TOKEN" \
  -H "Content-Type: application/json" \
  -d "{\"query\":\"mutation { disablePullRequestAutoMerge(input: {pullRequestId: \\\"$PR_NODE_ID\\\"}) { clientMutationId } }\"}" \
  https://api.github.com/graphql | jq -r '.errors | length'); then
  echo "$AUTOMERGE_FAILURE_MSG"
  exit 1
fi

if [ "$AUTOMERGE_NO_ERRORS" -ne 0 ]; then
  echo "$AUTOMERGE_FAILURE_MSG"
  exit 1
fi<p>10. Finally! The agent is invoked with the configurations needed, and the prompt is prepared from the previous steps. The <code>--allowedTools</code> parameter deterministically constrains which commands and actions Claude can use and take, respectively. This is key for safety and was set as part of the configuration generated in step (5). As we’ll see in the prompt description and analysis, we ask Claude to explain its actions as they happen, appending them to the file created in step (6), claude-actions.log . This is crucial for real-time monitoring and post-finalization reporting.</p>echo "--- Claude actions"
claude --allowedTools="$CLAUDE_ALLOWED_TOOLS" -p "$CLAUDE_FIX_PROMPT"<p>11. When the agent finishes its work, some post-processing steps are taken:</p><p>a. Upload the actions log.</p><p>b. Determine the script return code depending on whether it has found a solution to the broken build.</p><p>c. Add success report labels to the pull request. </p>echo "~~~ Post processing"

# Upload claude output log 
buildkite-agent artifact upload /tmp/claude-actions.log

# Evaluate the success of the step in function of the success of the Claude fix
if grep -qF "SUCCESSFUL FIX" ; then
  EXIT_CODE=0
else
  EXIT_CODE=1
fi

if [ $EXIT_CODE != "0" ] ; then
    RED='\033[0;31m'
    NC='\033[0m' # No Color
    echo -e "${RED}Claude was unable to fix the build, check artifacts for details.${NC}"
    add-pr-label "claude-fix-failed"
else
    echo "Claude was able to fix the build"
    add-pr-label "claude-fix-success"
fi

exit $EXIT_CODE<p>One important point about this flow is that the build pipeline is set up to be restarted when new commits are pushed. So, after the agent pushes its changes, the pipeline starts over, and the agent kicks in again only and only if the new iteration fails.Commits control the pipeline flow, pipeline build steps control the invocation of the agent, and this latter commits fixes when it can find them. As a result, closing the circle of a user experience that can be described as human-supervised AI autonomous contribution.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8b120ebb01b836c1/6a170cea6234e000dedb1a51/bfec7ca96e602a651a429e45e4634cb3e5cee199-1030x784.png" alt="" /><p>This approach to AI agent-based coding is to editors with AI batteries (VSCode+Copilot, Cursor, Windsurf…) what a real autonomous car is to a car with cruise control and lane-keep assist.</p><p><em>“This isn't about replacing our work with generated code, nor is it about having an AI buddy making suggestions beside us. Instead, it's about an AI buddy contributing to our codebase through GitHub, offering code change suggestions in an experience akin to open-source collaboration online.”</em></p><h4>The prompt</h4><p>Though it comes with the disadvantages of undetermined behaviour and ambiguity, the big win behind generative AI is that the instructions are self-explanatory. This is the prompt we are currently using (italic black font) with annotations adding context for this post (regular, magenta):</p><p><em>The build is failing, you might find a Gradle log to analyze under /tmp/previous_step_artifacts, but if not, you will have to run the build commands to evaluate what is failing. </em><strong>The reader might recall that on step (4) of the integration script we fetched the Gradle logs from Buildkite so Claude could analyze them. This is where we tell it to look at them and take action.</strong></p><p><em>These commands are './gradlew "--max-workers=$MAX_WORKERS" --console=plain publishForPlatform' and</em></p><p><em>'./gradlew "--max-workers=$MAX_WORKERS" --console=plain publishPlatformIndependent' but you don't need to run them as a first step if the log files under /tmp/previous_step_artifacts exist and contain the necessary information to analyze the failure. </em><strong>…to fall back to building everything from scratch in those cases where those logs are not available.</strong></p><p><em>In any case, you must find which Gradle subtasks are failing and fix them so that the build succeeds. </em>This gives Claude its goal: It must make sure that the build succeeds, <strong>applying whatever changes are necessary to the source (always obeying the constraints set below) code and iterating on the execution of global or local subtasks.</strong></p><p><em>Please:</em></p><p><em>- Analyze their output and apply the necessary fixes to make them succeed. </em></p><p><strong>This integration is our AI contributor buddy, our team treats it as a new hire. A capable one which still needs to learn “our ways”. The code style, preferred techniques, pitfalls to avoid, etc… In a nutshell, what we learn as we contribute is encoded in an ever evolving file of recommendations in the Cloud repository: CLAUDE.MD.Claude code looks for this file by default but we wanted to make it explicit that it must follow its recommendations:</strong></p><p><em>- Follow the recommendations under the CLAUDE.md file within the working directory.</em> </p><p><em>- If concrete Gradle subtasks are failing, iterate over them before attempting execution of the global ones. </em><strong>Boom! Just like this, you can follow the agent's “thinking” and actions in real-time from Buildkite. This is how the contents appended to claude-actions.log are generated.</strong></p><p><em>- Log each action you take in the file /tmp/claude-actions.log as they happen in realtime. Each entry should have a prefix with the current timestamp in the format "Claude Action [YYYY-MM-DD HH:MM:SS]: " and a description of the action taken. Run commands and their outputs should be logged too.</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f8ebc994afbe602/6a170ceccf4f255bb6b2d1c1/489a1092e976e74f457d14cec2d545edbbb923b4-1495x947.png" alt="" /><p><em>- Log in the same file, using the same format, the plans you are going to follow and their outcomes as you finish them.</em></p><p><em>- Always include "--max-workers=$MAX_WORKERS" option in invocations to ./gradlew command.</em><strong>Used by the post-processing step in the integration script, step (11):</strong></p><p><em>- If you get the build fixed, please add a last line to /tmp/claude-actions.log with the string "SUCCESSFUL FIX". Otherwise, add "FAILED FIX" at the end of the file.</em></p><p><strong>Next, the instructions telling Claude to commit its fix changes, only they are successful:</strong> <em>- If you succeed and added the "SUCCESSFUL FIX" line, please commit the changes that fixed the problems and push to the source branch in the same Git repository. The branch name is given by the BUILDKITE_BRANCH environment variable. Your commit messages should be prefixed with the "Claude fix: " string.</em></p><p><em>- To push the changes, you must use Github token authentication like in the following example git push https://token:$GITHUB_TOKEN@github.com/elastic/cloud.git HEAD:&lt;BRANCH_NAME&gt; taking into account that the token is stored in the GITHUB_TOKEN environment variable.</em></p><p><em>- Separate each fix into a different commit, so that the history is clear and understandable.</em></p><p><em>- If Git pushes fail due authentication issues, retry again after 1 minute. If still failing, then after 5 minutes and a last attempt after 10 minutes.</em></p><p><em>- Before pushing to Git: If and only if there are changes to push because the fix was successful, add the label "claude-fix-success" to the PR.</em></p><p><em>- Do not change what is not strictly necessary to fix the build. </em><strong>Humans tend to be indolent, AI even more. We had to introduce this last point as we found that Claude tended to fix the problems of a version bump by, well, … removing the version bump:</strong></p><p><em>- NEVER downgrade versions of dependencies as declared in the version catalogs of elastic/cloud master branches such as gradle/libs.versions.toml, prefer failure to version downgrades relative to the master branch.</em></p><h2>Results and lessons learnt</h2><p>During its first month of operation, and limited to only 45% of dependencies, Self Healing PRs became one of our Cloud GitHub repository’s top contributors. The plugin successfully fixed a total of 24 initially broken PRs, with the Claude author making 22 commits between July 22nd and August 22nd.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd4d32417b2c914c8/6a170ced47d49c57932d8a67/c74c19b53c16fb8685dc71b598b96988d83b1f21-1128x583.png" alt="" /><p>Our estimates indicate that, during this period, its contributions<strong> saved 20 days of active development work for our team</strong>. That’s a considerable amount of reduced toil, repetitive low-value work for engineering that is key for the codebase organism but invisible for its performance metrics.</p><p>Even when it fails to completely fix problems, it nudges things forward, hinting at the way to the solution or trimming the search space.</p><p>We have also learnt that tuning AI agents' behaviour marks the difference between failure and success. Teaching it the team’s skills through contributions to CLAUDE.md repository files made it stop following undesired coding practices, but, more importantly, made it diligent. We have learnt that there is nothing as lazy as an uneducated coding agent. This teaching process never ends, but each added hint and rule translates into a saved day of work.</p><p><em>“Trust but verify” </em></p><p>I started this post highlighting the differences between traditional and generative AI automation. Underterminism in the latter group means that you can never blindly trust the changes proposed by this automation. <strong>Reviews are imperative, and a critical attitude towards the proposed changes is extremely important</strong> for the health of this axolotl-like code organism. The alternative is exposing ourselves to devilish incidents and bugs behind seemingly perfect code with alien ways of being wrong.</p><h2>Conclusion</h2><p>We have seen how CI and AI can work together to satisfy the ever-expanding demands of infosec quality standards.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltac9faff69f27f2a3/6a170cef8b73cb60f818a097/3eecb15806bba67dc983c82522cbe889db37627a-542x440.png" alt="" /><p>This area of intersection between tools and needs made us design an application that behaves pretty much like a human, adding contributions in a collaborative code repository, and that can easily control traditional automation, avoiding bot tug-of-wars.</p><p>This is just the beginning. With this tool at our hand, we are starting to explore alternative direct applications.</p><p>For example, by enabling it on all pull requests, we can just open Pull Requests with incomplete changes, leaving behind less creative tasks such as API specs regeneration or linting; expecting the integration to add the necessary commits to finish these steps that are as important as boring for software engineers.</p><p>We have observed with awe how it has worked around transient problems in ancillary build services, filling the gaps when necessary with proactive execution of approved tools.</p><p>We expect this to go well beyond helping with automatic updates.</p><p>Who knows? Even to the extreme of reverse axolotl PRs: Humans writing interfaces and their unit tests, and letting Self Healing PRs come up with the rest.</p><p>The success of this pilot has traced a plan where we activate the integration for all Renovate PRs in the Cloud repository and possibly expand to all pull requests regardless of their origin.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ci-pipelines-claude-ai-agent</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ci-pipelines-claude-ai-agent</guid>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Pablo Pérez Hidalgo]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd4f368c579d510a6/6a170cf0961e69ef90c4cf6d/7259232be4b223710e21e6cd0082e2270ed07ad3-1600x873.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 30 Sep 2025 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>