<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[ES|QL - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[ES|QL - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/blog/category/esql</link>
    </image>
    <link>https://www.elastic.co/search-labs/blog/category/esql</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/category/esql.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Sat, 10 Oct 2026 23:00:27 GMT</lastBuildDate>
  <item>
    <title><![CDATA[The fastest work is the work you never do: 100x faster sorted queries in Elasticsearch]]></title>
    <description><![CDATA[How Elasticsearch 9.6 lets ES|QL push its current TopN threshold into Lucene, and how that 100x shows up on some sorted queries and barely registers on others.]]></description>
    <content:encoded><![CDATA[<p>In <a href="https://www.elastic.co/elasticsearch">Elasticsearch</a> 9.6, a query filtering 346 million Nginx log documents dropped from around 500 seconds to 5 on a single node. That’s 100x faster, and the same query gains 37x on <a href="https://www.elastic.co/docs/deploy-manage/deploy/elastic-cloud/serverless">Elastic Cloud Serverless</a>. The optimization is called <em>min-competitive TopN</em>.As <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> accumulates result candidates, it raises the minimum value that Lucene has to return, so document ranges that can no longer improve the result are never loaded or decoded. You get the same 500 rows back, with hundreds of millions of documents left unread.</p><h2>A sorted ES|QL query over 346 million log documents</h2>FROM logs_*
| WHERE message LIKE "*request body too large*" OR message LIKE "*queue is full*"
| SORT @timestamp DESC
| LIMIT 500<p>Using real-world Nginx logs, the benchmark searched 346 million docs across 432 GB of total index data. It returned 500 results and took ~500 seconds. That's more than eight minutes.</p><p></p><h3>What’s happening under the hood?</h3><p></p><p>Before this optimization, candidate rows were loaded, decoded, and evaluated by the user filter before TopN could determine which of them could no longer contribute to the final result.
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf091d4d4f50632fe/6abfaca2e7a20edb3234a965/esql-query-pipeline.png" alt="ES|QL query pipeline: Lucene scan, load and decode, filter, local TopN, final TopN, with the TopN decision last" /><p></p><h2>How min-competitive TopN prunes work in ES|QL</h2><p>TopN learns something useful while the query is still running. Once its heap is full, the worst value currently in the heap becomes a competitive threshold: that is, any row that cannot beat that value cannot make the final result.</p><p></p><p>ES|QL can feed that threshold back into earlier stages of execution. Two changes make this work:</p><p></p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/esql-elasticsearch-8-19-9-1?#significant-performance-and-scalability-improvements"><strong>Dynamic Lucene pushdown</strong></a><strong>.</strong> As TopN finds better candidates, its competitive threshold improves. ES|QL pushes that threshold down to Lucene, allowing Lucene to skip candidates that cannot beat the current worst TopN value. This avoids loading, decoding, and evaluating rows that have no chance of making the result.</p></li><li><p><a href="https://github.com/elastic/elasticsearch/pull/142406"><strong>Shared threshold across drivers</strong></a><strong>.</strong> ES|QL processes data in parallel across multiple drivers. If each driver only uses its own local threshold, one driver cannot benefit from better candidates discovered by another. Sharing the best-known competitive threshold means a strong result found by one driver immediately helps all the others prune their remaining work.</p></li></ol><p></p><p>Together, these changes move TopN from a late-stage filter to something that shapes query execution much earlier in the pipeline.</p><h3>How does the min-competitive TopN optimization work?</h3><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7807646d99ccf5ff/6abfad1b0491ebddfc047acb/baseline-topn-query.png" alt="Baseline TopN query scanning all 30 rows across three Lucene segments to return the top 3 sorted results" /><p>The query asks for the top three rows matching a wildcard filter, sorted in descending order, across three segments of 10 rows each. Without the optimization, all 30 rows are loaded, decoded, and evaluated against the filter before TopN sees any of them, at a cost of 30 row loads and 15 TopN operations.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltca82b2e1bee9dfb7/6abfad6b23d193d2985b6ccb/dynamic-lucene-pushdown.png" alt="Dynamic Lucene pushdown cutting rows loaded from 30 to 16 as the TopN min-heap threshold tightens per segment" /><p>With dynamic pushdown, ES|QL maintains the TopN heap as it scans. The heap fills at the end of the first segment with [98, 51, 47], and its worst entry is pushed down to Lucene as the competitive threshold. That cuts the second segment to four rows out of 10, which improves the heap to [98, 83, 77] and tightens the threshold again, leaving just two rows to load in the third segment. This produces the same three results, with 16 row loads and 10 TopN operations.</p><h3>Sharing the TopN threshold across drivers</h3><p>One driver may find excellent recent candidates immediately. Without coordination, another driver doesn't know that and continues working with a weak threshold. Sharing the global worst competitive value lets that discovery benefit all drivers.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt71b29eda542cb155/6abfadc44e42ee2b277fce13/parallel-esql-drivers.png" alt="Parallel ES|QL drivers sharing a global TopN min-heap threshold to prune work across segments and shards" /><p></p><h2>Benchmark results on a local node and Elastic Cloud Serverless</h2><p>The min-competitive TopN optimization ships as generally available (GA) in Elasticsearch 9.6. For real-world Nginx data benchmarks (same query, same dataset, p50 of 51 measurements, with 20 warmup runs discarded):</p><p></p><p>
</p><p><strong>Local node</strong></p><p><strong>Serverless (GCP)</strong></p><p>Before optimization</p><p>~500s</p><p>~520s</p><p>Production implementation</p><p>~5s</p><p>~14s</p><p>Runtime improvement</p><p>~99%</p><p>~97%</p><p>Speedup</p><p>100x</p><p>37x</p><p></p><h3>Benchmark hardware</h3><p></p><p>
</p><p><strong>Local node</strong></p><p><strong>Serverless (GCP)</strong></p><p>CPU</p><p>Apple M4</p><p>Small GCP host</p><p>RAM</p><p>48 GB</p><p>8 GB</p><p>Java Virtual Machine (JVM) heap</p><p>24 GB</p><p>4 GB</p><p>Storage</p><p>Local external TB5 NVMe SSD</p><p>Cloud storage</p><p>Shards</p><p>100</p><p>100</p><h2>When does this query optimization technique help most?</h2><p>The biggest gains happen when TopN becomes competitive early. If good candidates appear at the start of execution, the threshold rises quickly and subsequent scans get pruned aggressively. Three conditions amplify this:</p><p></p><ol><li><p>N is small relative to the candidate population. Finding 500 rows out of 346 million documents gives the optimizer far more room to eliminate work than returning a large fraction of the matches does.</p></li><li><p>The post-Lucene work is expensive enough that skipping a row saves something measurable, like avoiding loading large strings, decompression, or wildcard evaluation.</p></li><li><p>Many candidates pass the coarse Lucene scan in the first place. If almost nothing matches the initial filter, there’s not much unnecessary work left to eliminate.</p></li></ol><p></p><p>If matching rows are rare or if good candidates are discovered late, the threshold takes longer to become useful and the gain can be much smaller.</p><p></p><p>The current implementation applies to <code>@timestamp</code> sorted queries. Queries sorted by other fields don’t yet benefit from this optimization.</p><h2>What's next for ES|QL query performance</h2><p>The same idea has room to run in several directions, including:</p><p></p><ul><li><p><strong>Other numeric sort fields. </strong>The underlying approach isn’t inherently limited to timestamps, so extending it to other numeric sort fields could benefit a much broader range of sorted TopN queries.</p></li></ul><p></p><ul><li><p><strong>Processing promising data first.</strong> How quickly the threshold becomes useful depends on how early TopN sees good candidates. Processing more promising data slices first could establish a strong threshold sooner, allowing more aggressive pruning of subsequent work.</p></li></ul><p></p><ul><li><p><strong>Early termination.</strong> As the threshold strengthens during execution, it becomes possible to recognize when remaining ranges cannot produce a better result and to stop processing them entirely, rather than running through to completion.</p></li></ul><p></p><ul><li><p><strong>Coordination across nodes.</strong> The current optimization shares competitive information across drivers within a node. Broader coordination across nodes or clusters could let strong candidates discovered in one part of a distributed query reduce work elsewhere.</p></li></ul><p></p><ul><li><p><strong>Late materialization and result streaming.</strong> Competitive filtering and late materialization attack the same problem from different angles. Competitive filtering eliminates rows that cannot reach the final TopN as early as possible, while a fetch phase can defer loading expensive field values until the set of likely winners is much smaller. Result streaming could extend this further. If the engine can determine that some leading results can no longer be displaced, those rows could be returned before the full query completes. For interactive investigations, getting the first useful result fast matters as much as total query runtime.</p></li></ul><p></p><h2>What this means for Elasticsearch performance tuning</h2><p>
The min-competitive TopN optimization in Elasticsearch 9.6 shows how much work can be avoided when information discovered during query execution is fed back into earlier stages. In our Nginx log benchmark, that turned a roughly 500-second query into a roughly 5-second query on a single node.</p><p></p><p>This is also a starting point. Extending competitive filtering to more sort fields, finding strong candidates earlier, terminating work sooner, and eventually coordinating these decisions across nodes offer opportunities to push the same idea further.</p><p></p><p>The earlier we know what can still win, the less work we need to do for everything that cannot.</p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-performance-tuning-esql-topn</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-performance-tuning-esql-topn</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Lucene]]></category>
    <dc:creator><![CDATA[Attila Sahi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf092de2db9684039/6abfa9e5701a1a7aa18aa403/faster-query-runtime.png" length="0" type="image/png"/>
    <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[The best LLM writes correct Elasticsearch ES|QL 59% of the time. Here's what breaks the other 41%.]]></title>
    <description><![CDATA[We scored 6,000 ES|QL queries from four models against BIRD's answer key. Most misses come from mismatched join keys, SQL syntax the parser rejects, counting after a one-to-many join, or a value the model guessed.]]></description>
    <content:encoded><![CDATA[<p>We gave four models 500 natural-language questions from BIRD's Mini-Dev set, asked each for one<a href="https://www.elastic.co/docs/reference/query-languages/esql"> Elasticsearch Query Language (ES|QL)</a> query, and graded all 6,000 answers on whether the right rows came back. The best model got 59% correct on a single attempt with nothing but the index mappings, and wrote a query that failed to parse only 7 times out of 500. Grammar is no longer the ceiling. A compact ES|QL reference in the prompt is worth around 10 points to a smaller model and takes 10 points off the strongest one. Of the queries that still miss, roughly half throw a precise Elasticsearch error that a single retry can fix. The rest run cleanly and come back empty, because the model guessed a value that isn't in the data.</p><p>We started from<a href="https://aclanthology.org/2025.acl-long.971/"> Text-to-ES Bench</a> (ACL 2025), which measures how well large language models query Elasticsearch using<a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/querydsl"> Query DSL</a>. Its models pair a DSL query with Python post-processing, and pandas assembles the multi-index answers. We wanted to see what changes when the model can join inside the query.<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join"> <code>LOOKUP JOIN</code></a> joins indices natively, so a multi-index question becomes one statement that Elasticsearch executes end to end.</p><h2>The dataset: BIRD benchmark, Mini-Dev set</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt39ef07d655230eff/6ab9ea2ad4fedf3a7d8c1e3f/image2.png" alt="Cartoon bird holding notes of SQL keywords, illustrating LLM text-to-ES|QL query generation" /><p>Text-to-ES Bench sourced its questions from <a href="https://bird-bench.github.io/"><strong>BI</strong>g Bench for La<strong>R</strong>ge-scale <strong>D</strong>atabase Grounded Text-to-SQL Evaluation (BIRD)</a>, a widely used text-to-SQL benchmark. We do the same, using BIRD's Mini-Dev set: 500 natural-language questions over real relational databases. Every task gives you three things:</p><ol><li><p><strong>A question in English:</strong> The prompt that the model translates.</p></li><li><p><strong>A gold SQL query:</strong> The reference query that produces the answer.</p></li><li><p><strong>The returned rows:</strong> The answer key that you grade against.</p></li></ol><p>The part we rely on is that the ground truth is the rows, not the SQL. We never compare the generated query text against the gold SQL; a prediction is graded only on the rows that it returns. A query that doesn’t parse returns no rows, and no rows is a zero. "The names of the three drivers with the shortest average pit stop" should produce the same answer, whether you computed it in SQL, ES|QL, or by hand. So we keep BIRD's questions and BIRD's answer key and swap only the language that the model writes in.</p><p>Here’s one task, from the <code>student_club</code> database:</p>Question:  List out the full name and total cost that member id "rec4BLdZHS2Blfp4v" incurred?
Evidence:  full name refers to first_name, last_name

Gold SQL:  SELECT T1.first_name, T1.last_name, SUM(T2.cost)
           FROM member AS T1
           INNER JOIN expense AS T2 ON T1.member_id = T2.link_to_member
           WHERE T1.member_id = 'rec4BLdZHS2Blfp4v'

Gold rows: [["Sacha", "Harrison", 866.25]]<p>That <code>[["Sacha", "Harrison", 866.25]]</code> is the answer key. We also pass through BIRD's "evidence" hints (the <code>full name refers to...</code> line above) exactly as the dataset provides them; the questions are written assuming that you have them.</p><h2>Getting BIRD into Elasticsearch</h2><p>Each table becomes its own index, deliberately not denormalized. Flattening the schema in advance would quietly answer the hard part of the question for the model, and we would end up measuring our data modeling rather than its querying.</p><p>Indices on the right side of a <code>LOOKUP JOIN</code> must use <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-lookup-join"><code>lookup index</code> mode</a>:</p>es.indices.create(index=idx, settings={"index.mode": "lookup"}, mappings=mapping)<p>And anything you join or group on needs to be a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword"><code>keyword</code></a> rather than a <code>text</code> field, so text columns become <code>keyword</code> with an <code>ignore_above</code> guard. That keeps them usable in <code>WHERE</code> equality, <code>STATS ... BY</code>, <code>LOOKUP JOIN</code>, and <code>LIKE</code>, while staying under Lucene's term limit:</p>KEYWORD_IGNORE_ABOVE = 8000  # keeps UTF-8 byte length under Lucene's 32766 term limit

props[field] = {"type": "keyword", "ignore_above": KEYWORD_IGNORE_ABOVE}<p>All in, that’s roughly 3.9 million documents across 75 indices, running on Elasticsearch 9.5. If you haven’t built a join like this before, <a href="https://www.elastic.co/search-labs/blog/elasticsearch-esql-lookup-join">this walkthrough of native joins in Elasticsearch</a> covers the index-mode requirements in more depth.</p><h2>How we prompted the models and scored ES|QL accuracy</h2><p>The prompt mirrors BIRD's zero-shot protocol: a schema block, the evidence hint, the question, and an instruction to return only the query. The one deviation is that the schema is presented as Elasticsearch index mappings instead of <code>CREATE TABLE</code> DDL, because the target language is ES|QL.</p>SYSTEM_PROMPT = (
    "You are an expert Elasticsearch ES|QL query writer. You translate a natural-language "
    "question into ONE valid ES|QL query that runs against the provided indices.\n\n"
    "ES|QL is a piped query language: FROM &lt;index&gt; | WHERE ... | STATS ... BY ... | SORT ... | LIMIT ...\n"
    "It is NOT SQL and NOT Elasticsearch Query DSL. To join indices, use LOOKUP JOIN.\n\n"
    "Think step by step, then return ONLY the final ES|QL query."
)<p>Every query gets one attempt. That’s deliberate, because one call with one prompt isolates what the model knows from what a scaffold could recover.</p><p>We ran three prompt variants against four models:</p><ol><li><p><strong>base:</strong> Schema, evidence, question; whatever ES|QL the model already knows.</p></li><li><p><strong>focused skill:</strong> The above, preceded by a compact subset of the <a href="https://github.com/elastic/agent-skills/tree/main/skills/elasticsearch/elasticsearch-esql">ES|QL skill</a>: its <code>SKILL.md</code> overview plus the language reference, the generation tips, and the query patterns. The parts that are irrelevant to relational queries, time series, PromQL, and full-text search are left out.</p></li><li><p><strong>full skill:</strong> The same, but with the complete skill attached and every reference file included.</p></li></ol><p>In both skill variants, the files are pasted into the prompt as a static block. A skill normally reaches the model through a trigger, an agent deciding it needs the reference and loading it. We cut that step out so the variable under test is the content of the reference.</p><p>The models tested were <a href="https://developers.openai.com/api/docs/models/gpt-5.5"><code>gpt-5.5</code></a>, <a href="https://www.anthropic.com/claude/opus"><code>claude-opus-4-8</code></a>, <a href="https://www.anthropic.com/news/claude-sonnet-4-6"><code>claude-sonnet-4-6</code></a>, and <a href="https://developers.openai.com/api/docs/models/gpt-5.4-mini"><code>gpt-5.4-mini</code></a>. </p><p>Scoring runs the generated ES|QL through the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-esql-query"><code>_query</code> API</a> and compares the result to BIRD's gold rows as a set. Row order, duplicate rows, column order, and extra returned columns are ignored. One standard: the query returns the right data.</p><p>We decided to ignore the presentation details because they say nothing about whether the model understood the question. Take the <code>student_club</code> task from above and a model answer that a strict tuple comparison rejects:</p>gold: [["Sacha", "Harrison", 866.25]]
pred: [["Sacha Harrison", 866.25]]<p>The model built a <code>full_name</code> where the gold SQL kept <code>first_name</code> and <code>last_name</code> apart. They produced the same data, same rows. Extra columns are the same class of problem and are more common; a model that answers <code>KEEP atom_id, element</code> when the question only asked for the element has still found the element.</p><h2>How accurate is LLM-written ES|QL?</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3b1ea8a53752be5d/6ab9eb08907ac85efe7fbcbf/image3.png" alt="Bar chart of ES|QL execution accuracy by model across base, focused skill and full skill prompt variants" /><p>Execution accuracy, 500 questions per cell:</p><p><strong>Model</strong></p><p><strong>base</strong></p><p><strong>focused skill</strong></p><p><strong>full skill</strong></p><p>gpt-5.5</p><p>59.0%</p><p>48.8%</p><p>50.4%</p><p>claude-opus-4-8</p><p>40.4%</p><p>50.0%</p><p>49.0%</p><p>claude-sonnet-4-6</p><p>30.2%</p><p>39.2%</p><p>39.6%</p><p>gpt-5.4-mini</p><p>19.8%</p><p>30.4%</p><p>31.6%</p><p>Two things stand out:</p><ol><li><p><strong>The reference is worth about 10 points to every model except the strongest:</strong> Opus, Sonnet, and gpt-5.4-mini gain between 9.0 and 10.6 points from the focused skill; gpt-5.5 loses 10.2. It was already at 59.0% with nothing but the schema, and it misformed only 7 of 500 queries, so it had no grammar problem for a reference to fix, and handing it one cost it more than it could possibly win.</p></li><li><p><strong>The gains are smaller than the error counts suggest:</strong> The skill nearly eliminated the biggest failure in the run: across all 6,000 queries, join errors fell from 939 to 354. Aggregate accuracy didn’t move. The reason shows up one column over, as a class of error that barely existed before: <code>Found ambiguous reference</code> went from 120 to 706. The errors weren’t fixed so much as renamed, and we take that apart in the next section. Documentation teaches a model the grammar of ES|QL; it doesn’t teach it the shape of your data.</p></li></ol><h2>Four ways that ES|QL generation breaks</h2><p>The percentages below are out of the misses. A <em>miss</em> is any query that didn’t come back with the right data, whether it failed to run or it ran and returned the wrong rows. Getting the data right and the shape wrong doesn’t count.</p><p><strong>Failure class</strong></p><p><strong>What to do about it</strong></p><p>Join key names differ on each side</p><p><code>RENAME</code> before the join, match foreign key names to primary key names, or use the 9.2 join predicate</p><p>SQL syntax the ES|QL parser rejects</p><p>Put a one-page syntax card in the prompt</p><p>Counting after a one-to-many LOOKUP JOIN</p><p><code>COUNT_DISTINCT</code> on the entity, or aggregate before the join</p><p>The model guessed a value that isn't in your data</p><p>Sample rows and distinct values in the prompt</p><h3>1. ES|QL LOOKUP JOIN needs the same field name on both sides</h3><p>Join key naming is the largest failure class: 134 of Opus's 298 base misses, roughly 45% of them. SQL joins columns with different names, and BIRD's schemas rely on it (<code>superhero.skin_colour_id</code> joins to <code>colour.id</code>). ES|QL's bare join form does not. <code>LOOKUP JOIN &lt;index&gt; ON &lt;field&gt;</code> takes a single field name that must exist on both sides, closer to SQL's <code>JOIN ... USING</code> than to <code>JOIN ... ON</code>. Elasticsearch 9.2 added a second form that lifts this restriction, but only when the two keys have different names.  That condition is where our run went sideways, and we come back to it at the end of this section.</p><p>A model reaching for SQL join syntax writes this:</p>line 2:51: mismatched input '=' expecting {&lt;EOF&gt;, '|', 'and', ...}<p>Give it the skill, and it learns the <code>ON &lt;field&gt;</code> form and then trips one step later, still reaching for the left-hand name (<code>expense.link_to_budget</code> points at <code>budget.budget_id</code>):</p>line 3:39: Unknown column [link_to_budget] in right side of join<p>Across the run, 212 join queries failed to parse at all, the SQL-shaped <code>ON a = b</code> among them, and another 226 parsed but were rejected for the right-hand name. There are three ways out, depending on what you control.</p><p><strong>1. Rename at query time.</strong> One line before the join:</p><p>One wrinkle: <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/rename"><code>RENAME</code></a> replaces the target column if the name you rename to already exists on the left. Joining <code>superhero</code> to <code>colour</code> is exactly that case, since both carry an <code>id</code>, so move the collision out of the way in the same command (renames apply left to right):</p><p><strong>2. Fix it in the data model.</strong> Give foreign keys the same name as the primary key that they point at, and the problem disappears permanently. It’s cheap at design time, which is why we call it a modeling habit.</p><p><strong>3. Use a join predicate (Elasticsearch 9.2 and later).</strong> Elasticsearch 9.2 added complex join predicates that compare differently named fields directly, the way that SQL taught you to:</p>| LOOKUP JOIN student_club__budget ON link_to_budget == budget_id<p>Every name in a predicate must be unambiguous, so this form doesn’t rescue the case where the key already exists on both sides. That case is most of BIRD, and it’s where our edited skill backfired. Telling the models to prefer the predicate over <code>RENAME</code> dropped gpt-5.5's use of <code>RENAME</code> from 47% of its joins to 11%, and the error that <code>RENAME</code> was preventing showed up in its place:</p>Found ambiguous reference to [id]; matches any of [line 1:1 [id], line 3:15 [id]]<p>Across the run, that error went from 120 to 706, while the join bucket fell from 939 to 354; almost exactly a wash. <a href="https://www.elastic.co/search-labs/blog/esql-elasticsearch-9-2-multi-field-joins-ts-command">The 9.2 release write-up</a> covers the new join forms.</p><h3>2. SQL syntax the ES|QL parser rejects</h3><p>The models that struggle most are the ones reaching for SQL habits. This is where gpt-5.4-mini lost 236 of its 500 base queries and Sonnet lost 181.</p><p><strong>The model writes</strong></p><p><strong>ES|QL wants</strong></p><p><code>WHERE x = 5</code></p><p><code>WHERE x == 5</code></p><p><code>WHERE name = 'Bob'</code></p><p><code>WHERE name == "Bob"</code></p><p><code>CASE WHEN x &gt; 1 THEN 'a' ELSE 'b' END</code></p><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/conditional-functions-and-expressions"><code>CASE(x &gt; 1, "a", "b")</code></a></p><p><code>COUNT(DISTINCT id)</code></p><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/aggregation-functions"><code>COUNT_DISTINCT(id)</code></a></p><p><code>DIVIDE(a, b)</code>, <code>YEAR(d)</code></p><p>These do not exist</p><p>A correlated subquery</p><p>Restructure with <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/stats-by"><code>STATS</code></a> and a join</p><p>Single quotes are the sneakiest of these, because in ES|QL, double quotes delimit strings and single quotes do not. Here Sonnet reached for both habits at once, a bare <code>=</code> and a single-quoted string, and the parser stopped at the quote:</p>line 2:18: token recognition error at: '''<p>This is the one category that in-context documentation helps with. Sonnet's 181 syntax errors became 86, and gpt-5.4-mini's 236 became 150. If you’re pointing a smaller model at ES|QL, a one-page syntax card is the single highest-return thing that you can put in the prompt. A full language reference isn’t better.</p><h3>3. Counting across a one-to-many LOOKUP JOIN</h3><p><code>LOOKUP JOIN</code> fans out the way that a SQL join does; when a row on the left matches several rows in the lookup index, you get one output row per match. A model that forgets this counts the joined rows instead of the entity that the question asked about. This is around 32% of the runs-but-wrong bucket.</p><p>Join a member to their expenses, and you get one row per expense, so a plain count answers <em>How many expenses</em>, when the question asked <em>How many members</em>:</p><p><code>COUNT(*)</code> here counts expense rows. The fix is to count the entity that you actually mean or to aggregate before the join rather than after:</p>| STATS members = COUNT_DISTINCT(member_id)<h3>4. The model guessed a value that isn't in your data</h3><p>The fourth failure class is the model inventing a value that looks plausible and isn’t what’s stored.</p><p>BIRD's databases are full of these. A transaction type is stored as <code>'VYBER'</code>, not <code>'withdrawal'</code>. Ask for withdrawals, and gpt-5.4-mini writes the English word, which matches nothing:</p><p>The query is valid ES|QL and runs cleanly. It just comes back empty, because nothing in that column says <code>withdrawal</code>. The same pattern shows up across the dataset:</p><ul><li><p>A lab result is <code>'negative'</code>, not <code>'-'</code> or <code>false</code></p></li><li><p>Dates frequently live as keyword strings rather than date fields, which breaks any date function that the model reaches for and accounts for roughly 17% to 24% of wrong results on its own</p></li><li><p>Ranking questions confuse position with a stored rank column</p></li></ul><p>A schema block cannot fix this, because a schema tells you that a column is a <code>keyword</code> and never tells you which keywords are in it. The fix is sample values in the prompt: a handful of rows or the distinct values of low-cardinality columns. This is the failure class where retrieval helps and documentation does not.</p><h2>Conclusion: What this means for text-to-ES|QL in production</h2><p>We draw three conclusions from the run:</p><ol><li><p><strong>The newest models already write good ES|QL:</strong> Single call, zero shot, nothing but a schema, and the strongest model still gets the data right on most questions, while almost never writing one that fails to run. ES|QL isn’t exotic to frontier models anymore.</p></li><li><p><strong>A new language feature doesn’t automatically become a win:</strong> Elasticsearch 9.2's join predicate removes the restriction behind our largest failure bucket, and pointing the models at it did cut that bucket from 939 errors to 354. Accuracy didn’t move, because the models spent the savings on a new error. A feature only pays once the guidance says when not to use it.</p></li><li><p><strong>Prompt additions pay off for every model except the strongest:</strong> A compact syntax reference is worth around 10 points to a smaller model, a fuller one doesn’t help more, and the frontier model needs neither. Give it advice it didn’t need, and you can take 10 points off it.</p></li></ol><p>Roughly half of the failures throw a precise, actionable Elasticsearch error (<code>Unknown column [gender_id] in right side of join</code> tells you exactly what to fix), so a single retry with the error text appended should recover a large share of them when using an agent like <a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Elastic Agent Builder</a>. </p><p>The other half fail silently on guessed values, and no retry helps; those need sample rows and value lookups in the prompt. Grammar is no longer the ceiling. Knowing your schema and your values is, and that’s a retrieval problem, a much more tractable one.</p><h2>Resources</h2><ul><li><p>Browse <a href="https://github.com/Delacrobix/Text-to-ES-QL-Results-and-analysis">the benchmark harness and all 6,000 scored queries</a>, or filter them in the <a href="https://delacrobix.github.io/Text-to-ES-QL-Results-and-analysis/">live interactive report</a></p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join"><code>LOOKUP JOIN</code> command reference</a>, including multi-field joins and the predicate form</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/esql-elasticsearch-9-2-multi-field-joins-ts-command">The Elasticsearch 9.2 ES|QL release write-up</a> for the new join forms</p></li><li><p><a href="https://aclanthology.org/2025.acl-long.971/">Text-to-ES Bench (ACL 2025)</a>, the Query DSL benchmark that this work builds on</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/text-to-esql-benchmark</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/text-to-esql-benchmark</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Query Languages]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Jeffrey Rengifo,Gustavo Llermaly]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb263614f6585fc21/6ab9e9d326c31b3764359b35/image1.png" length="0" type="image/png"/>
    <pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Ask Elastic Agent Builder why it's slow: Natural-language trace analysis]]></title>
    <description><![CDATA[Four agent performance questions your Agent Builder traces can answer, covering token spend by model, tool error rates, slow conversation turns, and recent prompts. The ES|QL for each is here, including the type cast SUM() needs.]]></description>
    <content:encoded><![CDATA[<p>Ask Elastic Agent Builder how many tokens your agents burned today, and it writes the Elasticsearch Query Language (ES|QL) and runs it against your OpenTelemetry (OTel) trace data. Then it answers in the chat UI. The same holds for your other agent performance questions: Which tool fails most often? Which conversation turns are slowest? What have users actually been asking? Elastic Agent Builder is designed for this purpose. It allows users to ship agents grounded in your data in minutes. </p><p>If tracing is on, the <code>agent-builder-traces</code> <a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview">skill</a> is already loaded on every agent in your space and there’s nothing to install. The <a href="https://www.elastic.co/search-labs/blog/opentelemetry-tracing-agent-builder">first post</a> in our series covered enabling OTel tracing and the out-of-the-box (OOTB) dashboards, along with threshold alerts. This post is about asking.</p><h2>What is the agent-builder-traces skill?</h2><p><code>agent-builder-traces</code> is a built-in Agent Builder skill. It takes a natural-language question, generates an ES|QL query from it, executes that query against your trace index, and returns a plain-language summary. The index it targets is <code>traces-agent_builder.otel-&lt;space-id&gt;</code>, where <code>&lt;space-id&gt;</code> is the Kibana space that you’re working in. Using the exact space-scoped index pattern, rather than a wildcard, keeps data from other spaces out of your results.</p><p>Under the hood, the skill uses a single inline tool: <code>agent-builder-traces.generate_esql</code>. Instead of calling this tool directly, you ask a question and the agent calls the tool on your behalf, passing your question as the prompt for ES|QL generation. The tool resolves the current space's trace index, builds a query using the default model, executes it against Elasticsearch, and returns the result.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt556675b25cdd19b9/6ab4ad2b69958c862adc8c79/unnamed.png" alt="Elastic Agent Builder chat calling the agent-builder-traces skill to answer a token usage question in plain English" /><h2>Which privacy settings control what your traces capture?</h2><p>Several fields containing sensitive information are off by default. You can enable them in <strong>Gen AI Settings</strong> (<strong>Stack Management</strong> &gt; <strong>AI Assistants</strong>), under the "Agent Builder Traces" section. Expand the <strong>Advanced privacy</strong> settings to find them.</p><p><strong>Setting key</strong></p><p><strong>What it captures</strong></p><p><code>agentBuilder:tracing:includeUserPrompts</code></p><p>User messages</p><p><code>agentBuilder:tracing:includeLlmResponses</code></p><p>Large language model (LLM) response text</p><p><code>agentBuilder:tracing:includeToolDetails</code></p><p>Tool call arguments and results</p><p><code>agentBuilder:tracing:includeSystemPrompt</code></p><p>The agent's system prompt</p><p><code>agentBuilder:tracing:includeRealNames</code></p><p>Real tool, agent, and conversation names (hashed when off)</p><p><code>agentBuilder:tracing:includeRealIds</code></p><p>Real conversation and workflow IDs (hashed when off)</p><p><code>agentBuilder:tracing:includeUserData</code></p><p>Real user IDs and usernames (hashed when off)</p><h3>Why is my trace analysis query returning empty rows?</h3><p>If the skill returns empty rows or tells you that a field is unavailable, check two things: whether the relevant setting is enabled in your configuration, and whether your query time window overlaps with any recorded spans. The skill reports exactly what the query returned; it doesn’t fabricate content when fields are empty..</p><h2>What agent performance questions can you ask?</h2><p>Examples of questions that you can ask include:</p><h3>How many tokens have my agents used, by model?</h3><p><strong>Ask:</strong> <em>How many input and output tokens have my agents used in the last 24 hours, broken down by model?</em></p><p><strong>ES|QL:</strong> </p><p>The <code>TO_LONG()</code> cast around each token field is required. These fields can surface as mixed integer and long types across index generations, and ES|QL’s <code>SUM()</code> needs an explicit numeric conversion before it can aggregate them. If you write your own queries against this index and hit unexpected type errors, this is usually why.</p><h3>Which tools are failing most often?</h3><p><strong>Ask:</strong> <em>What is the error rate for each tool over the last 7 days?</em></p><p><strong>ES|QL: </strong></p><p>The tool error rate query aggregates across spans named <code>‘execute_tool’</code> and compares <code>status.code == "Error"</code> counts to total counts per tool name. The result shows which tools are failing most often. </p><h3>Which conversation turns are slowest?</h3><p><strong>Ask:</strong> <em>Show me the 10 slowest conversation turns in the last hour.</em></p><p><strong>ES|QL: </strong></p><p>The skill queries spans where <code>span.name LIKE "invoke_agent *"</code> and <code>attributes.elastic.inference.span.kind == "CHAIN"</code>. These correspond to individual conversation turns, from the moment a user sends a message to when the agent returns a response. Durations are stored in nanoseconds, so the skill converts to seconds before sorting.</p><h3>What have users been asking my agents?</h3><p><strong>Ask:</strong> <em>What questions have users been asking in the last 30 minutes?</em></p><p><strong>ES|QL:</strong> </p><p>The recent-prompts query returns data only when <code>agentBuilder:tracing:includeUserPrompts</code> is set to <code>true</code>. When the setting is off, the <code>attributes.gen_ai.input.messages</code> field is empty and the skill will suggest that you check your<strong> Include User Prompts</strong> privacy setting.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt207b044228f7d38b/6ab4adecf2c7a075d08b4ba0/unnamed.png" alt="Agent Builder trace analysis showing empty input messages because the Include User Prompts privacy setting is off" /><h2>When should you use Discover or Kibana Lens instead?</h2><p>The skill targets one index pattern: <code>traces-agent_builder.otel-&lt;space-id&gt;</code>. Two scenarios fall outside that boundary.</p><h3>Can the skill query my own application indices?</h3><p>If you want to run ad hoc questions against data that you've indexed yourself, use a general data exploration skill or write ES|QL directly in Discover. The <code>agent-builder-traces</code> skill won’t query outside the Agent Builder traces index.</p><h3>Can the skill build or edit dashboards?</h3><p>The skill focuses on ad hoc queries and generating summary text. If your goal is to build a permanent visualization from your trace data, use Lens or the <code>dashboard-management</code> skill instead. We covered that specific process in depth in the first post in our series.</p><p>The skill sits alongside two other ways of watching your agents.</p><p><strong>Tool</strong></p><p><strong>Best for</strong></p><p><strong>You use it when</strong></p><p><code>agent-builder-traces</code> skill</p><p>Ad hoc agent performance questions</p><p>You notice something and want an answer without leaving the chat</p><p>OOTB dashboards</p><p>Ongoing visibility</p><p>You want trends over time without asking anything</p><p>Threshold alerts</p><p>Automated monitoring</p><p>You define a condition once and get notified when it's breached</p><h2>How to build evaluation pipelines from agent trace data</h2><p>Spans in <code>traces-agent_builder.otel-&lt;space-id&gt;</code> capture real execution data from every conversation turn and LLM call running in your space. That record is the raw material for evaluation pipelines: automated checks on response quality and latency budgets per agent type, along with regression detection when you update a system prompt. The next post in our series covers how to build those eval loops from trace data that Agent Builder is already collecting.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agent-performance-trace-analysis-agent-builder</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agent-performance-trace-analysis-agent-builder</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[Operations]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Meghan Murphy]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc61c395d7859ef8/6ab4acdee2d12a17d37d48c9/unnamed.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Columnar storage isn't a columnar database. What Columnar mode brings to Elasticsearch]]></title>
    <description><![CDATA[Elasticsearch has stored data in columns since 2013, but adding full columnar database capabilities required a new mode.]]></description>
    <content:encoded><![CDATA[<p>When we wrote that Elasticsearch is becoming a columnar database, the sharpest reply we got was that it already is one. That reply is correct on the facts. Doc values, the per-field column store that Elasticsearch inherited from Lucene, arrived in 2013, and nearly every aggregation, sort, and query in Elasticsearch Query Language (ES|QL) has read them since Elasticsearch 2.0 made them the default. Each field's values sit together in their own file on disk. So the interesting question is what else a columnar database needs (rather than whether we store columns), and the answer turns out to be five things.</p><p>Doc values were built to make aggregations, sorting, and grouping possible on a document engine, and they do that job well. <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage">Columnar Mode</a> changes what the columns are for, and this post walks through the five properties that separate storing columns from being a columnar database.</p><h2>Five properties separate a column store from a columnar database</h2><h3>Doc values were an optimization on top of <code>_source</code></h3><p>For most of the last decade, the original JSON document was the source of truth and the columns were a derived convenience. That ordering has consequences throughout the engine.</p><p>Because the engine could always fall back to <code>_source</code>, per-field storage was allowed to be lossy. Text fields had no doc values at all, since text could be reread from the stored document when needed. Even synthetic source, which reconstructs a document from its fields rather than storing a copy, sometimes reads from row-shaped structures to stay faithful to the JSON that arrived, with values that exceed <code>ignore_above</code> and fields that arrived unmapped going into stored fields.</p><p>The result is a clear contract; whatever JSON you send, you get back, and the columns accelerate everything else. For an engine whose job is to return your documents, that’s the right way round. Columnar Mode inverts it. Every field stores itself exactly once as doc values, doc values cannot be turned off, text fields get doc values, too, and the document is reconstructed from the columns when something asks for it.</p><h3>How dictionary encoding handles high-cardinality data</h3><p>Sorted doc values, the default for keyword fields, store a dictionary of distinct values plus one ordinal per document pointing into it. This is an excellent trade when values repeat. A <code>host.name</code> field drawn from a hundred machines, or a status code that’s almost always 200, compresses beautifully and groups quickly.</p><p>It works like the index cards in a warehouse. When 50 crates hold the same product, one card and 50 pointers beats writing the product name 50 times. When 50 crates hold the same product, you can store the product name in the index with the list of 50 crate IDs. When every crate holds something unique, you might as well just put the product name on the crates; the index will help you find what crate you want, but it won't save ink.</p><p>High-cardinality fields describe a lot of real data, including URLs and trace identifiers, along with message bodies. Columnar Mode, which skips the dictionary and compresses the values in blocks instead, uses binary doc values for high-cardinality strings. Which of the two a field gets isn’t something you configure. The engine decides per field, based on the values it sees, so each column is encoded for the data it actually holds instead of one default applied to every field. Pure columnar systems have long carried cardinality in their type system, but usually as something you declare, and you own the consequences when the data shifts underneath it. Here, it’s the engine's job.</p><h3>Why every field builds an inverted index by default</h3><p>By default, a keyword field also builds an inverted index and a numeric field also builds a BKD tree. That happens on every field because at write time the engine doesn’t know which capability you’ll want at read time, significantly increasing the footprint of each field. Those structures also have to be rebuilt during segment merges, which costs CPU exactly when ingest is heaviest.</p><p>Our time series engine (TSDB) is proof of what happens when you stop paying for capability that the workload doesn’t use. Replacing the indices on <code>@timestamp</code> and dimension fields with <em>doc value skippers</em>, which are sparse structures holding the minimum and maximum value for each block of documents, removed 10 bytes of the original 25 bytes per OpenTelemetry (OTel) data point. There was no measurable query regression on time range and dimension filters, and indexing CPU dropped by about 10%, as a bonus.</p><p>Columnar Mode generalizes that default. Fields aren’t indexed unless something needs them to be, with only text-mapped fields keeping their inverted index for fast free-text search.</p><h3>Metadata fields like _id and _routing were row-shaped</h3><p>The fields you never think about followed the same document-first design. The <code>_id</code> field was a stored field plus an inverted index. Custom <code>_routing</code> was a stored field. Sequence numbers were kept for optimistic concurrency control, regardless of whether a workload ever updated a document.</p><p>TSDB deals with all three. It synthesizes <code>_id</code> from the <code>_tsid</code> and <code>@timestamp</code> values that already identify a data point, using a segment-level bloom filter to catch duplicates, which removes 5 bytes per data point with no loss of functionality. It trims sequence numbers once replication no longer needs them, which removes 4 bytes. Add a codec block size increase from 128 to 512 elements for another 2 bytes, and those four changes contribute across versions 9.1 through 9.4 to the 21 bytes that took OTel metrics <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">from 25 bytes per data point down to 3.75</a>.</p><p>Columnar Mode makes those ideas general rather than metrics-specific. By general availability (GA), all metadata fields will store themselves as doc values, while we plan to follow up and add a sort id mode synthesizing the identifier from the index sort fields, in addition to derived fields that will generalize what <code>_tsid</code> does for time series to any set of fields.</p><h3>Columnar query execution in the ES|QL compute engine</h3><p>A column store only pays off if the engine reads it as columns. Aggregations inherited the document-at-a-time shape from search, which is the natural fit for an engine built around documents. Reading columns instead lets the engine hand a whole block of values to a single instruction, and that’s where the numbers below come from.</p><p>The ES|QL compute engine changed that shape, and TSDB again shows the size of the effect:</p><ul><li><p><strong>Vectorized execution</strong> of time series aggregations was worth up to 8x on its own.</p></li><li><p><strong>Decoding on-disk data</strong> straight into the primitive arrays the engine aggregates over, with no intermediate copies, was worth roughly another 10x.</p></li><li><p><strong>Constant blocks</strong> turned repeated values into a form of in-memory run-length encoding.</p></li><li><p><strong>Filter pushdown</strong> moved filters down to Lucene, where skippers can discard whole blocks unopened.</p></li></ul><p>Together with the rest of the block-level query work, query latency improved by up to 160x compared to earlier versions.</p><p>That work continues. Skipper-aware operators, aggregations that group on ordinals and convert to real values as late as possible, and richer per-block summaries are all in progress, and they benefit every index mode because every mode reads doc values underneath.</p><h2>What Columnar Mode changes for logs and analytical data</h2><p>Storing values in columns is a storage detail. A columnar database needs five things:</p><ol><li><p>The columns are the only copy of the data.</p></li><li><p>Each column is encoded for the data it actually holds.</p></li><li><p>Metadata is columnar, too.</p></li><li><p>Fields add indices only when something needs them.</p></li><li><p>The query engine processes blocks of values rather than records or documents.</p></li></ol><p>TSDB reached all five for metrics in Elasticsearch 9.4, which is why the numbers in this post come from metrics rather than from a slide. Columnar Mode applies the same treatment to logs and security telemetry, along with analytical data. It’s in technical preview in Elasticsearch 9.5, with GA targeted for 9.7.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt723063cac8566f76/6aa8fb5a5ceda93f62dae858/unnamed.png" alt="Elasticsearch doc values, inverted index and _source across standard index mode, LogsDB and Columnar Mode" /><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-doc-values-columnar-database</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-doc-values-columnar-database</guid>
    <category><![CDATA[Lucene]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Yannis Roussos]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt25aa087300badaae/6ab3e7938ab0c04cc32254df/diagram-one-field-three-structures.webp" length="0" type="image/webp"/>
    <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How we built PromQL into Elasticsearch]]></title>
    <description><![CDATA[PromQL runs on the same Elasticsearch compute engine as ES|QL, with no plugin and no separate process to operate. Getting there meant changing how the engine evaluates time windows and builds grouping keys.]]></description>
    <content:encoded><![CDATA[<p>More than 80% of the Prometheus Query Language (PromQL) queries in our real-world corpus run on Elasticsearch without modification. Elasticsearch 9.5 makes the PromQL and the Prometheus-compatible API generally available (GA), so you can ingest Prometheus metrics with<a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write"> remote write</a> and query them through the<a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api"> Prometheus HTTP APIs</a> or the<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql"> PROMQL</a> command in Elasticsearch Query Language (ES|QL).</p><p>PromQL compiles to the same compute engine that runs ES|QL and inherits its planner and distributed execution, along with its release process. We didn’t build a second engine for this, and there’s no plugin to install.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9e310070ec244c5d/6aa3928e224d356c5e0dd18e/1.png" alt="PromQL compatibility in Elasticsearch rising from zero to 80% between 9.4 Tech Preview and 9.5 GA" /><p>This post is about how we built it.</p><p>Key takeaways:</p><ul><li><p><strong>One engine:</strong> The implementation combines Elasticsearch’s mature distributed planning, storage, and testing infrastructure with its newer compute engine, which provides a columnar execution runtime. This lets PromQL reuse proven Elasticsearch capabilities while executing through a modern, native vectorized pipeline rather than introducing a separate runtime.</p></li><li><p><strong>One server:</strong> Elasticsearch implements the Prometheus remote write and query APIs directly, so Prometheus-compatible ingest and queries run without any additional plugins.</p></li><li><p><strong>Engineered for efficiency:</strong> Supporting PromQL required new engine primitives for range-aligned evaluation grids, backward-looking windows, dynamic label grouping, pipeline result reshaping, and compact wide aggregation keys. These primitives allow PromQL queries to execute efficiently end to end, with the relevant semantics implemented directly in the compute engine rather than through external post-processing.</p></li><li><p><strong>Compatibility measured in real use:</strong> In addition to Prometheus compliance tests, we built a differential-testing and quality-control pipeline over 2,000 PromQL queries collected from public repositories. </p></li></ul><p>Read more:</p><ul><li><p><a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">Query Prometheus Metrics in Elasticsearch with PromQL</a></p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch">Ship Prometheus Metrics to Elasticsearch with Remote Write</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api">Bringing Fire to Elasticsearch: Adding Native Prometheus APIs</a></p></li></ul><h2><strong>Why run PromQL on Elasticsearch</strong></h2><p>Many teams already store logs and traces in Elasticsearch while running metrics in Prometheus or another dedicated metrics back end.</p><p>Prometheus and its ecosystem are strong and widely adopted, but large deployments can also bring operational sprawl and scaling challenges, along with limited retention. </p><p>So we set out to combine Elastic’s highly optimized time-series database (TSDB) with a best-in-class metrics ecosystem. The result is a smaller observability stack, with fewer systems to operate and metrics storage that scales horizontally and supports long-term retention.</p><h2><strong>One engine: PromQL and ES|QL share the same compute engine</strong></h2><p>We made an early architectural decision not to run a separate PromQL engine next to Elasticsearch.</p><p>PromQL is instead another front end to the Elasticsearch compute engine.</p><p>This puts PromQL in the normal Elasticsearch development lifecycle. It uses the same planner, distributed execution engine, testing infrastructure, and release process as Elasticsearch itself.</p><p>To learn more about Elasticsearch’s query engine, check <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">our blog</a>.</p><p>Like ES|QL’s time series queries that use the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> source command, PromQL is translated into a highly optimized query plan and executed across the cluster. Nodes process columnar batches through vectorized operators, while partial results move through exchanges until the final result is assembled. </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcf3a413e89209805/6aa392b38406d9813bcaa1d8/2.png" alt="PromQL and ES|QL frontends feed one shared Elasticsearch planner and columnar execution DAG across shards" /><p>This also means that PromQL and ES|QL operate over the same execution engine and time-series data. ES|QL can additionally extend a PromQL computation with post-processing that PromQL doesn’t support, such as lookup joins and inline aggregations.</p><p>For example, assume Prometheus request counters are stored in <code>metrics-*</code> and keyed by the <code>service</code> label. A lookup index, <code>service_registry</code>, maps each instance to its owning team and environment and to its service tier:</p><p></p><p></p><p>This architecture requires the execution engine to support PromQL semantics natively and efficiently rather than ES|QL syntax sugar. The following sections describe the changes and new execution primitives we introduced to achieve that.</p><h2><strong>One server: Prometheus remote write and HTTP API built into Elasticsearch</strong></h2><p>Query execution is only half of the story. The Prometheus ecosystem also expects familiar ingest and query APIs.</p><p>Prometheus protocols are the de facto standard for everything metrics in almost every team’s observability stack. So we built the HTTP API directly in the Elasticsearch server, which eliminated the need for a third component and tightened integration stability and performance.</p><p>On the ingest side, we <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">added</a> an endpoint for the <a href="https://prometheus.io/docs/specs/prw/remote_write_spec/">Prometheus remote write</a> protocol. It accepts Snappy-compressed Protocol Buffer messages, maps labels to TSDS dimensions, maps the metric name/value into metric fields, infers counter versus gauge mappings, and writes directly into TSDS. The built-in template is dynamic, so users don’t have to predeclare every Prometheus label or metric.</p><p>On the query side, Elasticsearch <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">exposes</a> Prometheus query APIs. A request enters through the Prometheus endpoint and is executed in the compute engine.</p><h2><strong>Running PromQL efficiently in a columnar engine</strong></h2><p>Sharing an execution engine doesn’t mean treating PromQL as syntax sugar over ES|QL. PromQL has different time, grouping, and response semantics, as well as workload characteristics that matter at scale. Supporting it efficiently required extending the compute engine rather than compensating in the API layer.</p><h3><strong>PromQL time grids: Aligning evaluation steps with TSTEP</strong></h3><p>Time-series query engines are optimized for grouping over time.</p><p>Elasticsearch normally groups timestamps with <code>TBUCKET(...)</code>, which truncates each timestamp to a fixed interval boundary. Truncation is cheap and produces deterministic bucket boundaries. It also makes intermediate results easier to reuse.</p><p>Prometheus defines evaluation points differently. For a range query, timestamps are laid out as fixed steps anchored to the query range, rather than derived by truncating each sample timestamp. Two queries with the same step but different range boundaries can therefore produce different evaluation grids.</p><p>To preserve these semantics, we introduced <code>TSTEP(...)</code>, which derives its grouping grid from the query range and step,  rather than truncating timestamps to globally aligned boundaries.</p><p>PromQL uses <code>TSTEP(...)</code> internally, preserving Prometheus timestamp semantics while still lowering the operation to a native Elasticsearch execution primitive.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt437155bca2029671/6aa3937927a5312436dcbf38/3.png" alt="TSTEP vs TBUCKET in Elasticsearch: PromQL step grid anchored to query start, TBUCKET to fixed boundaries" /><p>Some might see this as a simple problem, but the nuances matter. </p><p>Take, for example, a query that finds a 5m rolling average of a metric: </p><p></p><p></p><p>At evaluation time <code>T</code>, the result represents the average over the preceding five-minute range: </p><p><code>(T - 5m, T]</code></p><p>When the query is executed with a five-minute step, each output value is labeled with the upper end of its corresponding five-minute window.</p><p>Elasticsearch previously lacked this semantic and supported only forward-looking window aggregation functions:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt97aac2fa95cf657d/6aa393b71ade64390142ee2e/4.png" alt="Forward-looking window aggregation where each bucket covers the interval from timestamp T to T plus W" /><p>We rewrote the window-evaluation path so that both ES|QL and PromQL use a common backward-looking windowing implementation:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1a12ceb27d9ab95b/6aa393dcf08ee141fe856719/5.png" alt="Backward-looking PromQL window covering T minus W to T, the range used by rate and avg_over_time" /><p>Together, <code>TSTEP(...)</code> and backward-looking windows preserve the two time semantics that matter for PromQL range evaluation.</p><h3><strong>Dynamic label grouping: How PromQL </strong><strong><code>without()</code></strong><strong> resolves at runtime</strong></h3><p>Time grids determine when a PromQL expression is evaluated. Aggregation determines which input series are combined and which labels identify each output series.</p><p>For most analytical query engines, that identity is known when the query is planned. The planner can allocate grouping columns and choose an aggregation strategy. It also carries a fixed key through the execution pipeline.</p><p>That’s how ES|QL works:</p><p></p><p></p><p>The output series are grouped by an explicit key <code>(cluster, namespace)</code>.</p><p>PromQL can express the same operation in the opposite direction:</p><p></p><p></p><p>Now we know which dimensions <em>not</em> to use. We don’t necessarily know the full grouping key until the query is executed. This is a small language difference with significant execution consequences. </p><p>One possible implementation is to discover every label used by the metric, subtract <code>instance</code> and <code>pod</code>, and rewrite the expression into an ordinary <code>by(...)</code> aggregation. That adds a <a href="https://www.elastic.co/search-labs/blog/esql-metrics-info-ts-info-time-series-catalog">discovery phase</a> before planning. It also becomes inefficient for high-dimensional metrics where only a subset of all possible dimensions may have useful value in a particular series. Most queries need only a small subset of the available dimensions, so carrying the entire dimension universe as an aggregation key wastes memory and adds bookkeeping overhead.</p><p>We instead extended the time-series execution path with dynamic grouping columns. The engine loads dimensions as the series are read and applies the exclusions per time series. This avoids making the grouping schema a prerequisite for planning and avoids carrying a large sparse set of grouping columns through aggregation.</p><h3><strong>Dimension packing: Keeping wide PromQL grouping keys cheap</strong></h3><p>The <a href="https://prometheus.io/docs/practices/rules/#aggregation">idiomatic</a> way of writing PromQL aggregations involves heavy use of <code>without(...)</code> over <code>by(...)</code>:</p><p></p><p></p><p>Excluding labels rather than explicitly listing them makes dashboards and alerts resilient to schema evolution. If a new label is added to the metric, the query continues to preserve it unless it’s explicitly excluded.</p><p>The consequence for the engine is that effective grouping keys can be wide. Many of those labels often have low cardinality, yet each still participates in every aggregation stage.</p><p>In a columnar engine, each grouping column is normally represented as a separate vector. Ten grouping labels therefore mean 10 vectors flowing through every aggregation operator:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2f491d0d9fdb82f2/6aa3940027a53105ccdcbf3c/6.png" alt="Elasticsearch columnar page: rows split into typed blocks with delta and ordinal dictionary compression" /><p>Dimension fields are declared in the index mapping; the planner knows the schema up front, and grouping keys stay narrow and predictable.</p><p>In the columnar engine, each grouping label is carried as a separate vector or block. A key with 10 labels therefore requires 10 vectors to be read, hashed, compared, and retained by aggregation operators. As key width grows, so does the amount of data and bookkeeping that must move through the pipeline:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt13b25c686bbb46e0/6aa39429de2395f2e2e86d1d/7.png" alt="PromQL aggregation without packing: each grouping label hashed as a separate block into 64-byte keys" /><p>To avoid paying that per-column cost for every additional label, we introduced dimension packing. Before aggregation begins, the engine encodes the full grouping key into a single compact representation. Hash and comparison operations run on the packed key rather than on each block independently:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5a75855ef51c017b/6aa394468406d95667caa1de/8.png" alt="Dimension packing in Elasticsearch encodes PromQL grouping labels into one 16-byte key before hashing" /><p>Packing lets hashing and comparison operate on a single compact key rather than an increasing number of grouping blocks, making aggregation overhead less sensitive to key width. Because the engine is shared, ES|QL time-series queries will benefit from this optimization as well.</p><h3><strong>Building the Prometheus HTTP API response inside the pipeline</strong></h3><p>Unlike Elasticsearch's ES|QL column-oriented response format, the Prometheus response is row-oriented.  The Prometheus API returns one result row per time series, with its samples represented as timestamp-value pairs. </p><p>To support a compatible API layer, we had to regroup in the HTTP layer converting the columnar results into boxed row objects and accumulate them in map- and list-based structures until the complete Prometheus response could be produced. </p><p>We replaced this with the <code>TimeSeriesCollapse</code> compute operator. It groups rows by series and aligns samples to the query’s fixed step grid. It emits the reshaped result as ordinary columnar pages containing one row per series, with aligned multi-valued timestamp and value blocks. And it preserves compact, vectorized block representation throughout the pipeline: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt99e2ee3137bc2fa6/6aa39496d29b4e02e81da7b6/9.png" alt="TimeSeriesCollapse operator reshapes five columnar rows into two Prometheus time series per output page" /><p>The HTTP layer can now serialize those blocks directly, avoiding the maps, lists, boxed objects, and associated allocations required by the earlier implementation.</p><h2><strong>Testing PromQL compatibility against 2,000 real queries</strong></h2><p>Prometheus <a href="https://github.com/prometheus/compliance">compliance tests</a> were our starting point.</p><p>Even though they gave us a strong baseline, they didn’t tell us how frequently individual PromQL features appear in real workloads. To complement that baseline, we built a second test corpus from over 2,000 PromQL queries collected from public repositories. </p><p>We then classified those queries by the language features and expression patterns they exercise:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf008821c47e46040/6aa394b68406d90c5bcaa1e2/10.png" alt="PromQL feature use across 2,000 real queries: aggregations 59.67%, selectors 57.55%, rate functions 45.21%" /><p>For each compatible query shape, we run the same query against Elasticsearch and Prometheus and compare the results. </p><p>In addition to that, we actively rely on <a href="https://en.wikipedia.org/wiki/Fuzzing">fuzz testing</a>, which catches issues that unit tests alone are unlikely to expose, including differences in timestamp alignment, label retention, aggregation behavior, range-vector evaluation, and response encoding.</p><h2><strong>Which PromQL functions and APIs are supported in 9.5</strong></h2><p>Since 9.4 (technical preview), PromQL support in Elasticsearch has expanded substantially. In Elasticsearch 9.5, both the<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql"> <code>PROMQL</code></a> command in ES|QL and the<a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api"> Prometheus HTTP APIs</a> are generally available (GA), with more than 80% of the PromQL workflows in our real-world corpus now running without modification:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6706c16409dccfd7/6aa394d327a5312acfdcbf41/1.png" alt="PromQL compatibility in Elasticsearch rising from zero to 80% between 9.4 Tech Preview and 9.5 GA" /><p>The main additions since technical preview are:</p><p><strong>Feature</strong></p><p><strong>Example</strong></p><p><strong>Status</strong></p><p>Prometheus remote write ingest</p><p><code>POST /_prometheus/api/v1/write</code></p><p>GA in 9.5</p><p>Range queries</p><p><code>/api/v1/query_range</code></p><p>GA in 9.5</p><p>Instant queries</p><p><code>/api/v1/query</code></p><p>GA in 9.5</p><p>Metric metadata and build info</p><p><code>/api/v1/metadata</code>, <code>/api/v1/status/buildinfo</code></p><p>GA in 9.5</p><p>Native histogram functions</p><p><code>histogram_quantile</code>, <code>histogram_count</code>, <code>histogram_sum</code></p><p>GA in 9.5</p><p>Per-selector offset modifiers</p><p><code>[5m] offset 1h</code></p><p>GA in 9.5</p><p>Top-level <code>or</code> operator</p><p><code>rate(a[5m])</code> or <code>rate(b[5m])</code></p><p>GA in 9.5, up to eight operands</p><h3><strong>Prometheus remote write ingest</strong></h3><p>Elasticsearch <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write">accepts</a> Prometheus remote write (v1) messages directly:</p><p></p><p></p><p>Snappy-compressed Protocol Buffer messages are decoded, and labels are mapped to TSDS dimensions. Metric names and values are written into the time-series index. The built-in template is dynamic, so users don’t have to predeclare every Prometheus label or metric.</p><h3><strong>Range and instant queries through the Prometheus HTTP API</strong></h3><p>Both range and instant query endpoints are <a href="https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api">available</a>:</p><p></p><p></p><p></p><p></p><p>Range queries return matrices evaluated over a time window, and instant queries return vectors evaluated at a single timestamp. These endpoints can be used by Kibana, Grafana, or Prometheus-compatible alerting tools, and custom dashboards.</p><h3><strong>Metric metadata and build info endpoints</strong></h3><p>Elasticsearch <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api#promql-http-api-metadata">exposes</a> metadata about available metrics and a build-info endpoint:</p><p></p><p></p><p></p><p></p><p>The metadata endpoint returns metric types and help text, and the build-info endpoint returns the Prometheus-compatible server version. Grafana and other tools use these endpoints for feature detection and UI behavior.</p><h3><strong>Native histogram functions: histogram_quantile, count, and sum</strong></h3><p>Elasticsearch supports the <a href="https://www.elastic.co/docs/reference/query-languages/promql/functions/histogram">main PromQL operations</a> over native histograms:</p><p></p><p></p><p></p><p></p><p>Native histograms adapt their bucket layout to the data, providing useful precision across a wide value range without requiring users to configure every bucket boundary in advance. Classic histograms continue to work alongside native histograms.</p><h3><strong>Per-selector offset modifiers in PromQL</strong></h3><p>Offset modifiers shift a selector’s time window backward:</p><p></p><p></p><p>This returns the request rate from one hour earlier. Per-selector offsets are commonly used to compare current traffic, latency, or resource usage with an earlier baseline, such as the same period one week ago.</p><h3><strong>Top-level </strong><strong><code>or</code></strong><strong> operator in PromQL</strong></h3><p>Elasticsearch supports the top-level PromQL <code>or</code> operator:</p><p></p><p></p><p>In PromQL, <code>or</code> isn’t a Boolean operation. It performs a union between two sets of time series. Results from the left side are retained; a series from the right side is added only when its label set doesn’t match a series already returned by the left side. This is useful during migrations where the same logical metric may exist under an old and a new name.</p><p>The implementation follows Prometheus’s left-side precedence rules and preserves the <code>__name__</code> label. Top-level chains of up to eight operands are supported.</p><h2><strong>PromQL features not yet supported in Elasticsearch</strong></h2><p>GA doesn’t mean complete PromQL compatibility. Some less common and more complex parts of PromQL remain unsupported. These gaps now define the next phase of the work: </p><p><strong>Feature</strong></p><p><strong>Example</strong></p><p><strong>Status</strong></p><p>Advanced vector matching</p><p><code>on(instance) group_left</code></p><p>Planned</p><p>Sorting and ranking</p><p><code>topk</code>, <code>bottomk</code>, <code>limitk</code>, <code>sort</code>, <code>sort_desc</code></p><p>Planned</p><p>Label manipulation</p><p><code>label_replace</code>, <code>label_join</code></p><p>Planned</p><p>Absolute time modifier</p><p><code>@ 1710000000</code></p><p>Planned</p><p>Mixed-offset compound expressions</p><p><code>rate(...) - rate(... offset 1h)</code></p><p>Planned</p><p>Alerting and target endpoints</p><p><code>/api/v1/alerts</code>, <code>/api/v1/targets</code></p><p>Out of scope</p><h3><strong>Advanced </strong><a href="https://prometheus.io/docs/prometheus/latest/querying/operators/#group-modifiers"><strong>vector matching</strong></a><strong> with </strong><strong><code>on()</code></strong><strong> and </strong><strong><code>group_left</code></strong></h3><p>Some binary operations that require Prometheus vector matching aren’t yet part of GA.</p><p>For example, this query divides per-instance request rates by a per-instance capacity metric:</p><p></p><p></p><p>The <code>on(instance)</code> clause specifies which labels identify matching series. <code>group_left</code> permits many request-rate series to match a single per-instance capacity series, while retaining the labels from the higher-cardinality left-hand side.</p><p>These expressions are common when joining a detailed metric with metadata or a lower-cardinality capacity metric. Basic binary expressions are supported where applicable, while the remaining vector-matching forms are planned work.</p><h3><strong>Sorting and ranking: </strong><strong><code>topk</code></strong><strong>, </strong><strong><code>bottomk</code></strong><strong>, and </strong><strong><code>sort</code></strong></h3><p>Prometheus sorting and ranking functions are also not yet part of GA:</p><p></p><p></p><p>This returns the 10 services with the highest request rate. Similar queries are widely used in “top offenders” dashboards for traffic, latency, errors, and resource consumption.</p><p>The remaining functions include:</p><p></p><p></p><p></p><p></p><p></p><p></p><h3><strong>Label manipulation with label_replace and label_join</strong></h3><p>PromQL can construct or rewrite labels during query evaluation. These functions are particularly useful when dashboard variables, naming conventions, or label schemas don’t match exactly:</p><p></p><p></p><p>This creates an <code>environment</code> label from the <code>cluster</code> label.</p><p>Another common example combines existing labels into a display-oriented label:</p><p></p><p></p><p>This produces a <code>target</code> label, such as <code>payments/api-7f6d9</code>. <code>label_replace(...)</code> and <code>label_join(...)</code> aren’t yet included in GA.</p><h3><strong>Advanced time modifiers: The </strong><strong><code>@</code></strong><strong> modifier and mixed offsets</strong></h3><p>Several advanced time modifiers and expression forms remain outside the GA scope.</p><p>For example, an absolute <code>@</code> modifier evaluates a selector at a fixed Unix timestamp rather than at the query’s normal evaluation time:</p><p></p><p></p><p>This is useful for comparisons against a fixed historical point.</p><p>PromQL also permits expressions in which the two sides use different offsets:</p><p></p><p></p><p>This compares current traffic with traffic one hour earlier. Per-selector <code>offset</code> is available in GA, but not every combination of offsets and compound expressions is part of GA yet.</p><h3><strong>Prometheus API endpoints not yet implemented</strong></h3><p>In addition, the Prometheus HTTP API surface isn’t yet fully complete. Notably:</p><p>Alerting metadata through:</p><p></p><p></p><p>used by tools that inspect active alert state.</p><p>Target discovery through:</p><p></p><p></p><p>used to inspect scrape targets, health, and labels.</p><p>These endpoints concern Prometheus server and scrape-target state rather than querying metrics stored in Elasticsearch.</p><p>For the full list of limitations, see the <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-limitations#promql-limitations-form-post">PromQL limitations</a> page. </p><h2><strong>Try PromQL in Elasticsearch 9.5</strong></h2><p>To query Prometheus metrics in Elasticsearch 9.5 or Serverless, see the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql">PromQL documentation</a> and the <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api">Prometheus HTTP API reference</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/promql-elasticsearch-compute-engine</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/promql-elasticsearch-compute-engine</guid>
    <category><![CDATA[Query Languages]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Sergey Sidorov,Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf690e82ea51bbeec/6aa37da41ade6445cc42eded/unnamed.png" length="0" type="image/png"/>
    <pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Query rewrite rules in Elasticsearch: 2.3x faster wildcard scans]]></title>
    <description><![CDATA[A second rule makes empty-string filters 1.6x faster. It reads string lengths straight from the offset array and never touches the compressed bytes. Both rules came from the same habit of running real queries and hunting for the special case.]]></description>
    <content:encoded><![CDATA[<p>Lucene query rewrite rules make two string scan queries in Elasticsearch's <a href="https://www.elastic.co/docs/reference/elasticsearch/columnar">columnar mode</a> 2.3x and 1.6x faster. Both rules spot a query shape at runtime and swap in a cheaper implementation. For a wildcard query like <code>*google*</code>, that's a substring search in place of the automaton. A filter like <code>SearchPhrase != ''</code> can skip Zstd decompression, because it only needs string lengths that are sitting in an offset array.</p><p>Columnar mode is Elasticsearch's analytics-optimized <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage">columnar storage</a> mode, built for scan-heavy workloads, like log analytics. In this mode, keyword fields don't get an inverted index by default, so term and wildcard queries scan doc values. <a href="https://www.elastic.co/search-labs/blog/docvaluesskippers-lucene-range-queries">DocValuesSkippers</a> (zone maps) already trim how much data a scan touches, but these rewrites cut the cost of what's left. </p><h2>How Lucene's query rewrite mechanism works</h2><p>In Lucene, every query has the option to implement a <code>rewrite</code> method that returns another query. This method returns a query with the same semantics but a different implementation. The query engine repeatedly calls the <code>rewrite</code> method until the returned query doesn’t change. This final query is the one that’s actually evaluated. Importantly, the <code>rewrite</code> can see the actual query arguments and specialize the implementation based on these.</p><p>For example, in a query looking for documents where a string field contains the value "foo", the <code>rewrite</code> method knows that the term we’re searching for is "foo". In theory, <code>rewrite</code> could replace the general query class with something specific to "foo". For example, the original query class <code>ScanningBinaryDocValuesTermQuery</code> could be replaced with <code>FooQuery</code>. Now this rule probably wouldn't be helpful, but it gives a sense for the level of specialization that’s achievable with rewrite rules.</p><h3>Rewrite rules and query optimization in database systems</h3><p>It's worth placing rewrite rules in the larger context of database systems. Lucene and Elasticsearch aren’t the first systems to use transformation rules to optimize queries. Most (or maybe all) database systems use some kind of rule system during query optimization. The most influential rewrite rule system was in IBM's <a href="https://dl.acm.org/doi/10.1145/141484.130294">Starburst</a> database. This system's core contribution was extensibility; for example, it was possible to add new data types and storage methods, along with (most importantly to us) optimizer rewrite rules.</p><p>Each rule consisted of two parts:</p><ol><li><p><strong>A condition function:</strong> A predicate determining whether the rule applies to the current query graph.</p></li><li><p><strong>An action function:</strong> The transformation that rewrites the query plan into a more optimal form.</p></li></ol><p>A rule engine applied matching rules until a stopping condition was met.</p><p>Though Lucene's <code>rewrite</code> method is superficially different from these condition and action functions, it achieves the same goal. It checks whether certain conditions match, and if they do, it applies the rewrite by returning a new query. If conditions don’t match, the <code>rewrite</code> returns <code>this</code>, replacing the query with itself; that is, choosing not to apply the rule.</p><h3>Why these rules live in Lucene, not the ES|QL query optimizer</h3><p>Elasticsearch actually contains a separate rewrite rule system within the <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> optimizer. This operates on the high-level structure of a query; for example, doing predicate pushdown to avoid unnecessary computation on documents that will be filtered out. But it’s still useful to have the rule system within Lucene. Since Lucene acts as the storage layer for ES|QL (and classic <code>_search</code>) queries, it’s easier to express rewrites that take advantage of the physical data format in Lucene rather than in a higher-level optimizer.</p><h2>A query rewrite rule for wildcard queries: Simpler code, no automaton</h2><p><a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-wildcard-query">Wildcard queries</a> support the <code>?</code> and <code>*</code> operators to match any character once or any character multiple times. These operators can appear any number of times in a wildcard query. As with regexes, to evaluate whether a string matches a wildcard query, we build an automaton from the query string and then use the string bytes to do state transitions through the automaton. This is relatively fast, but if you have to evaluate it for every document, the latency really adds up.</p><p>But maybe we don't always have to run an automaton. Consider a query like <code>*foo*</code>. How would you implement this if you were writing a simple query engine to find matching strings in a list of strings? Pretty much every programming language has the tool you want built in: a method that finds a substring within a given string. This function doesn't need a complicated automaton; it probably just consists of a couple of <code>for</code> loops.</p><p>Now of course we couldn't use this function to implement an arbitrary wildcard query, but we don't have to. The rule rewrite system isn't for the general form. It's for implementing special cases, and it can see the specific query. It knows that we’re looking for <code>*foo*</code> and realizes that this specific case doesn't require the heavyweight automaton machinery. And it can do the same for any query that starts and ends with a <code>*</code>, with some term in the middle.</p><p>The following pseudo-code shows the pattern. At the top, we have the generic <code>WildcardQuery</code>. It has two notable fields: the query string (for example, <code>*foo*</code>) and the automaton built for that query. The <code>matches</code> method checks whether the field value for a given <code>docId</code> is a match by using it to evaluate the state transitions of the automaton. More interestingly, its rewrite method checks whether the query matches our special case. We show this with a regex that checks whether the query string starts with a <code>*</code>, has any non-<code>*</code>characters at least once, and then ends in a <code>*</code>. If so, we return the special case as a <code>ContainsQuery</code> and pass in the inner query string (since it doesn't care about the <code>*</code>s). The <code>ContainsQuery</code> then just does a simple <code>contains</code> check to see whether the term bytes are somewhere within the value bytes.</p>class WildcardQuery(query, automaton, docValues):

    boolean matches(docId):
        value = docValues.loadValue(docId)
        return automaton.matches(value)

    Query rewrite():
        if query matches r"^\*[^*]+\*$":
            return ContainsQuery(query[1:-1], docValues)
        return self


class ContainsQuery(term, docValues):

    boolean matches(docId):
        value = docValues.loadValue(docId)
        return value.contains(term)<h3>Benchmarking the wildcard rewrite on ClickBench Q20</h3><p>The wildcard rewrite is straightforward, but does it actually work? Yes, we can use the <a href="https://benchmark.clickhouse.com/">ClickBench</a> benchmark, which has several queries of this form. Query 20 (Q20) is <code>FROM hits | WHERE URL LIKE "*google*" | STATS count = COUNT(*)</code>. It's exactly the query shape that this rule matches: a string match against the wildcard query <code>*google*</code>. And since the query is just counting, we can see exactly how well this technique works. It turns out to be quite effective. Q20 saw a 1.75x improvement on median latency of hot query times, with no filter cache. All benchmarks in this post were run on an Intel Core i9-13900H.</p><h3>Adding SIMD to the substring search: 1.75x to 2.3x</h3><p>But can we do better? Yes, switching to a simple contains check opens up a new possibility. Instead of using the two for loops, we can swap scalar logic for single instruction, multiple data (SIMD) logic. Elasticsearch uses the <a href="https://openjdk.org/jeps/438">Panama vector API</a> (see our <a href="https://www.elastic.co/blog/accelerating-vector-search-simd-instructions">post on SIMD in Elasticsearch</a>), which lets us implement the contains check in SIMD. This works particularly well for longer strings that can take advantage of the wide SIMD registers; for strings under 24 characters, we still use the scalar approach. With this change, we saw another 1.32x improvement, resulting in a total speedup of 2.3x over the automaton-based approach.</p><h2>A query rewrite rule for empty strings: Less data, no decompression</h2><p>One benefit of Lucene-based rules is that they’re low level and can fit to the data format. That’s the case for this rule, which applies to string data. </p><h3>How columnar storage encodes string data</h3><p>In Elasticsearch's standard mode, string values are stored by document; this is a row-major format. But in columnar mode, unsurprisingly, the data is stored in columnar-major format. A column of string data is stored in chunks. Each chunk contains many string values and consists of an array of integer offsets and a (Zstd-compressed) blob of the strings' bytes. For a string at index <code>i</code>, <code>offsets[i]</code> points to the offset in the decompressed byte blob where the string starts. So the length of string <code>i</code> can be computed from <code>offset[i+1]-offsets[i]</code>. (There's a dummy extra offset at the end, so we can easily compute the length of the last string). The following diagram shows how a chunk with the strings ‘Feta’, ‘Asiago’, ‘’, ‘Stilton’, and ‘Brie’ is encoded.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt046cea94361c6305/6a9ee0c3508cca28e2a15ff9/image1.png" alt="Columnar storage chunk with an offsets array and byte blob, showing an empty string as two equal offsets" /><h3>Why a term query has to decompress the chunk</h3><p>Now that we understand the columnar format, let's get back to query optimization. First, consider a term query for the query <code>foo</code>. We’re looking for documents where a given string field exactly matches the string <code>foo</code>. So how do we implement this on a string column in the above format? The algorithm is straightforward:</p>docId = 0
for chunk in chunks:
    bytes = zstd_decompress(chunk.bytes)
    for i in range(len(chunk.offsets) - 1):
        value = bytes[chunk.offsets[i] : chunk.offsets[i+1]]
        if value == term:
            yield docId
        docId++<p>The bottleneck is the Zstd decompression step. But there's not much we can do about that; if we want to check the bytes, we have to decompress the chunks. But remember, we aren't trying to optimize the general case, we’re looking for special cases. (In reality, you don't just try to think up special cases. These optimizations came about by first running a useful query, realizing that it could be faster, and then looking for ways to improve it.)</p><h3>Rewriting the empty string query as a length check</h3><p>One special case we found that’s worth improving is a query for the term <code>""</code>. Admittedly, it's a silly term, but empty strings are all over the place. Since they're rarely useful, we usually filter them out with a query like <code>term != ""</code>. Thankfully, this is a query we can optimize.</p><p>Consider the above algorithm for the empty string term. The line <code>if value == term</code> is a bit weird; we’re asking <em>Does this value equal the empty string?</em> We can do that, but there are no bytes to compare, so the check unwinds:</p><ol><li><p>We only need to know whether the value has length 0.</p></li><li><p>If we only need the length, we don't need to look up the value in the decompressed chunk.</p></li><li><p>If we never look up a value, we don't need any bytes from the chunk at all.</p></li><li><p>If we need no bytes from the chunk, we don't need to decompress it.</p></li></ol><p>All we need are the lengths, and those live in the offsets array. It's compressed, too, but with cheap integer compression rather than Zstd, which is much faster.</p><p>With this realization, we can rewrite empty string term queries. The one new operation we need is <code>docValues.loadLength(docId)</code>, which reads directly from the offset array without touching the compressed bytes. After the previous example, this should look familiar. The most interesting part is <code>TermEqualsQuery.rewrite</code>; it finds the empty string special case and replaces the query with the simpler version that only checks the length.</p>class TermEqualsQuery(term, docValues):

    boolean matches(docId):
        value = docValues.loadValue(docId)  # requires Zstd decompression
        return value == term

    Query rewrite():
        if term == "":
            return LengthEqualsQuery(0, docValues)
        return self


class LengthEqualsQuery(queryLen, docValues):

    boolean matches(docId):
        length = docValues.loadLength(docId)  # reads only from offset array
        return length == queryLen<h3>Benchmarking the empty string rewrite: 1.6x faster</h3><p>Now let's see how this stacks up. There aren't any pure-scan ClickBench queries that use this rule as directly as Q20 does for the previous rule, so we'll make our own. Consider the query: <code>FROM hits | WHERE SearchPhrase != '' | STATS count(*)</code>. On this query, we see a 1.6x speedup, which is a great improvement for a fairly uncomplicated change. Better yet, ES|QL can take advantage of <code>loadLength</code> directly. Any time that ES|QL accesses a string's <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/string-functions/byte_length"><code>BYTE_LENGTH</code></a>, without needing the string itself, the request uses this same specialized length loading to avoid unnecessary decompression.</p><h2>What makes a good query rewrite rule</h2><p>The two rules covered here follow the same shape: identify that a query is a special case, and then swap it for a cheaper implementation. But they reduce cost in different ways. </p><p></p><p>
</p><p><strong>Wildcard rule</strong></p><p><strong>Empty string rule</strong></p><p>Query shape detected</p><p><code>*term*</code></p><p><code>field == ""</code></p><p>Replaced with</p><p>SIMD substring search</p><p>Length check on the offsets array</p><p>Cost reduced</p><p>Algorithmic work</p><p>Data access</p><p>Speedup</p><p>2.3x</p><p>1.6x</p><p>The underlying pattern is worth noting: finding a query that leaves performance on the table, finding a special case that can be optimized, and swapping in a cheaper implementation. The hard parts are finding queries that uncover these opportunities for optimization and then identifying the special cases. The actual fix is often relatively straightforward, as both rules here show. Our work on columnar mode has provided many opportunities to run interesting queries and hunt down exactly these kinds of wins.</p><p>That's also why extensibility in a rule system is so important. These rules can't be built into a database from the start; they're found through an incremental discovery process. Lucene's rewrite system makes that practical. As columnar mode grows to handle new workloads, rules like these will keep emerging.</p><p>To try columnar mode and the optimizations described in this article, use Elastic Cloud Serverless or Elasticsearch 9.5 or later, where columnar mode is available as a technical preview.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/query-rewrite-columnar-storage-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/query-rewrite-columnar-storage-elasticsearch</guid>
    <category><![CDATA[Lucene]]></category>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Parker Timmins,Martijn Van Groningen]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt838d10c743cf1f3e/6a9ee03c8936813a883df5d5/image2.png" length="0" type="image/png"/>
    <pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Introducing SPARKLINE in ES|QL: Spot trends at a glance]]></title>
    <description><![CDATA[Spot trends across thousands of groups at a glance without leaving your workflow. ES|QL's new SPARKLINE function turns aggregations into trend lines. One array per row, zero effort.]]></description>
    <content:encoded><![CDATA[<p>When you run a <code>STATS ... BY</code> query and get back dozens or hundreds of results (log patterns, hosts, services, status codes), the counts alone don't tell you what's happening <em>over time</em>. Is the error count or log pattern climbing or settling down? Is it within the usual range? To answer those questions today, you either build a separate time-series visualization or eyeball the numbers and hope for the best.</p><p>This can be a lot of effort, time, and context switching for what should be just a glance. In this blog, we explain how Elasticsearch Query Language’s (ES|QL’s) SPARKLINE works, what it does, and how to get started.</p><h2>How it works</h2><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/aggregation-functions/sparkline">SPARKLINE</a> is an ES|QL aggregate function with a straightforward signature:</p><p><code>SPARKLINE(aggregation, key, buckets, from, to)</code></p><ul><li><p><strong><code>aggregation</code></strong><strong>:</strong> Expression that calculates the y-axis value, including any supported aggregation: <code>COUNT(*)</code>, <code>SUM(bytes)</code>, <code>AVG(latency)</code>, or others.</p></li><li><p><strong><code>key</code></strong><strong>:</strong> Date expression from which to derive buckets.</p></li><li><p><strong><code>buckets</code></strong><strong>:</strong> Target number of buckets.</p></li><li><p><strong><code>from</code></strong><strong> / </strong><strong><code>to</code></strong><strong>:</strong> The time range boundaries. (In Kibana, they bind to the time picker via query parameters.)</p></li></ul><p>Under the hood, SPARKLINE buckets the time range, computes the aggregation per bucket, and packs the results into a single ordered array. Empty buckets are zero-filled so every group shares the same x-axis grid for fast, easy visual comparison.</p><p>The function composes naturally with <code>STATS ... BY</code>, so you can combine it with any grouping.</p><h2>Where sparklines shine</h2><p>The first place you'll see SPARKLINE in action is Discover's log pattern analysis, starting 9.5. When you run a <code>CATEGORIZE</code> query, Discover constructs the SPARKLINE query under the hood and renders trend lines next to each pattern. You don't write <code>SPARKLINE</code> yourself here; Discover handles it when you use the “identify patterns” option in the ES|QL editor. The result is immediate: You scan dozens of log patterns and instantly see which ones are spiking <em>right now</em> versus which have been steady all day.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte16b62f513b5915d/6a9904382707c505d6329e29/unnamed.png" alt="A dashboard shows a bar chart of hourly event counts above a table of results. The chart displays activity across several days, and the table lists each row with a count value, a sparkline, and pattern tags." /><p>Here’s the query under the hood:</p><p>Consider a platform team investigating how they can cut logging costs. They point Discover at tens of millions of documents and let <code>CATEGORIZE</code> cluster them into patterns. Two patterns rise to the top: verbose lifecycle messages like "fetching resource..." and "completed resource...", each with several millions of hits. The sparklines next to those rows tell the rest of the story: flat, steady streams running around the clock. Not incident-driven. Not bursty. Just constant noise, silently consuming hundreds of terabytes of storage.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt366155f45da3169b/6a9904777f53142f28c63114/unnamed.png" alt="A dashboard shows a bar chart of hourly result counts above a table of rows. The chart displays consistent activity over several days, and the table lists each row with a count value, a sparkline, and pattern tags." /><p>The fix in cases like this is usually straightforward: Adjust log levels, drop the pattern at ingest, or route it to a cheaper tier. The hard part was always <em>finding</em> it. With log pattern analysis and inline sparklines in ES|QL, that discovery takes seconds instead of hours. </p><p>Elastic’s internal site reliability engineering (SRE) teams routinely use log pattern analysis successfully, using every ES|QL enhancement. </p><p>These are just some concrete examples. Break down request counts by region, and see which regions are trending up. Compare latency across container IDs. Monitor queue depths by consumer group. The uses are endless.</p><h2>Simple by design</h2><p>We deliberately kept SPARKLINE focused:</p><ul><li><p><strong>It's an aggregate function</strong>, not a new command. It composes with existing <code>STATS ... BY</code> syntax, so there's nothing new to learn structurally.</p></li><li><p><strong>It returns data that the consumer can render in context.</strong> The function produces an array of values. How those values are rendered, as a mini-chart in Discover, a line in a notebook, or a JSON array in an API response, is up to the consumer.</p></li><li><p><strong>It fills empty buckets.</strong> Every group gets the same number of values aligned to the same time grid. This is a deliberate choice: Sparklines are most useful when you can compare shapes across rows at a glance, and that requires consistent alignment.</p></li></ul><p>The pattern is always the same: one query, many trend lines, instant visual triage.</p><h2>What's next</h2><p>SPARKLINE ships as a <strong>technical preview</strong> in Elasticsearch 9.5. Future work includes <strong>rendering sparklines in more ES|QL surfaces</strong>, beyond the initial integration, with the <code>CATEGORIZE</code> context in Kibana. This includes dashboards but also Elastic Observability use cases like the following:</p><p>In application performance monitoring (APM) workflows, engineers routinely analyze rate, errors, and duration (RED) metrics to understand service health. The challenge is that aggregate numbers hide dimensional outliers. A service might look healthy overall, but one region, one container, or one newly deployed version could be quietly degrading.</p><p>Today, the Elastic APM UI lets you break down metrics by transaction name, but root-cause analysis requires slicing by arbitrary dimensions: availability zone, service version, container ID, cloud region. SPARKLINE can make this practical. Break down error rate by <code>service.version</code>, and instantly see which version's trend line diverges from the pack.</p><h2>Get started</h2><p>SPARKLINE is available in Elasticsearch 9.5 as a technical preview. Try it with the <a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql-rest">ES|QL _query API</a> or in Kibana's Discover. Check the <a href="https://elastic.co/docs/reference/query-languages/esql/functions-operators/aggregation-functions/sparkline">SPARKLINE function reference</a> for the full syntax and supported types.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-sparkline-function</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-sparkline-function</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Daniel Rubinstein]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb8ad1babd1641235/6a9901a43481c241218af226/unnamed.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Know your facts: How Elasticsearch AI Indices let agents skip the reading and keep the answer]]></title>
    <description><![CDATA[A technical walkthrough of precomputing facts into an Elasticsearch AI Index, so agents answer from a single ES|QL query instead of reading whole documents, with fewer tokens and lower latency.]]></description>
    <content:encoded><![CDATA[<p>Pulling whole documents into an agent's context to answer one question is expensive, and the cost compounds with every miss. In this walkthrough, we precompute the facts instead. A Kibana workflow distills each document into a fact-level Knowledge Indicator (KI), stored in an Elasticsearch AI Index and retrieved with a single Elasticsearch Query Language (ES|QL) query. On the same question, an agent answering from KIs reached the same grounded answer using fewer tokens and lower latency than reading raw documents, without loading a single full document into context. These facts are precomputed once and then stored for use by future agents when they encounter similar queries. This is Part 2 of our series on building context with AI indices; <a href="https://www.elastic.co/search-labs/blog/ai-index-building-context-agents">Part 1</a> covered routing agents to the right index.</p><p>Managing context depends on good retrieval. Rather than have agents rediscover the same content for every question, burning tokens by retracing similar steps over and over again, Elastic’s agentic AI capabilities enable us to precompute these details and store them in a structured, searchable form, and they let agents load that context directly. We call this precomputed unit of context a Knowledge Indicator.</p><p>The default agentic retrieval augmented generation (RAG) pattern does the opposite. It retrieves whole documents and dumps them into the model's context at query time, paying for that retrieval in tokens and latency on every single question. Precomputing the answer as a KI moves that cost out of the hot path and does it once.</p><h2>How it works: AI Index, Kibana Workflows, and the query-ki skill</h2><p>Building context through AI indices has three main parts: the AI Index (a special Elasticsearch index where KIs live), Kibana Workflows to create your KIs, and a <code>query-ki</code> skill to help agents directly query KIs using ES|QL: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt201b0bf84c5f5002/6a8fef16ecdaa77015050aa9/unnamed.png" alt="AI Index architecture: Kibana Workflows write Knowledge Indicators, agents read them via the query-ki ES|QL skill" /><p>This blog post is similar to Part 1 in that we’re using the same core building blocks. But in this post, we’re demonstrating a very different use case. Instead of precomputing index metadata, we’re distilling specific <em>facts</em> from our indexed documents that may be used to directly answer agents’ questions without subsequent searches. We've also provided a <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/precomputed-context-technical-walkthrough-part-2/index-facts-kis.ipynb">notebook</a>, if you'd like to create the same KIs yourself, end to end, as you go through these examples. </p><h3>Prerequisites: Elasticsearch Serverless and an LLM API key</h3><p>This tutorial assumes you have:</p><ol><li><p>An Elasticsearch Serverless project. You can <a href="https://cloud.elastic.co/registration?onboarding_token=search&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">sign up for a trial</a> if you don't have one.</p></li><li><p>An API key to access your Elasticsearch project.</p></li><li><p>An OpenAI-compatible large language model (LLM) API key, to access AI indices via Deep Agents scripts.</p></li></ol><h2>Load the BrowseComp-Plus sample corpus into Elasticsearch</h2><p>First, we’ll need some sources. Sources can be data that already exists in your Elasticsearch indices or external data accessed via connectors or ES|QL data sources. 
For this blog, we’ll create an index, <code>browsecomp-plus</code>, to hold our example data, with the following mappings:</p>{
  "browsecomp-plus": {
    "mappings": {
      "_meta": {
        "description": "BrowseComp-Plus corpus: ~100k human-verified web documents (news articles, Wikipedia entries, institutional pages) used as a reasoning-intensive browsing/QA retrieval benchmark. BM25-only index."
      },
      "properties": {
        "docid": {
          "type": "keyword",
          "meta": {
            "description": "Stable corpus document id."
          }
        },
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document text: title, date, and body content."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document title (from the document's front matter)."
          }
        },
        "url": {
          "type": "keyword",
          "meta": {
            "description": "Source URL the document was crawled from."
          }
        }
      }
    }
  }
}<p>and populate it with a small sample of BrowseComp-Plus data via the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk"><code>_bulk</code> API</a>. You can use the supporting <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/precomputed-context-technical-walkthrough-part-2/index-facts-kis.ipynb">notebook</a> to load a sample of this data in your project. </p><h2>Create the AI Index that stores your KIs</h2><p>Just like in Part 1, the first step is to create an AI Index:</p>PUT ai-index-idx-my-corpus<p>This is preconfigured with the same required mappings as we listed out in Part 1. We perform hybrid search here using <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_text</code></a> out of the box.</p><h2>How agents retrieve KIs using ES|QL</h2><p>A KI is a document in the AI Index. What makes KIs useful is <em>retrieval</em>, or querying the AI Index to find the right content. This query is packaged within a small, portable skill that’s harness-agnostic and can be run in any agent harness. </p><p>Here’s a sample <code>query-ki</code> skill:</p> ---
name: query-ki
description: &gt;-
  Retrieve Knowledge Indicators (precomputed context) from the Elasticsearch AI
  Index before answering. Use it to find which index to search (routing profiles)
  or to look up precomputed facts without reading source documents. Trigger on any question that depends on specific facts, names, dates, or on choosing a data source.
allowed-tools: esql_query
---

# Retrieving Knowledge Indicators

Knowledge Indicators (KIs) live in Elasticsearch indices named <code>ai-index-*</code>.
Retrieve them by calling the <code>esql_query</code> tool with the query below. Substitute
the user's question for <code>&lt;query&gt;</code>, and <code>corpus_entry</code> as the <code>&lt;ki_type&gt;</code> for facts.

```esql
FROM ai-index-idx-* METADATA _id, _index, _score
| WHERE type == "&lt;ki_type&gt;"
| FORK
    (WHERE MATCH(content, "&lt;query&gt;") OR MATCH(description, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
    (WHERE MATCH(content.semantic, "&lt;query&gt;") OR MATCH(description.semantic, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
| FUSE
| SORT _score DESC
| KEEP title, content, description, tags
| LIMIT 5
```

Ground your answer in what the query returns, and cite the KI titles you used. If
nothing relevant comes back, say so rather than guessing.<p>Save this as<code>skills/query-ki/SKILL.md</code>.</p><p>Here’s what this skill is doing: </p><ul><li><p>We’re defining <code>corpus_entry</code> as our KI use case.</p></li><li><p>We’re performing a hybrid ES|QL search on our AI indices, filtering by the appropriate <code>type</code>, using reciprocal rank fusion (RRF) as the default method to fuse results.</p></li><li><p>The KI results will directly ground the agent’s answer when determining what facts are relevant to the users’ query.</p></li></ul><p>When we say that AI indices and KIs are <em>harness-agnostic</em>, it’s because the skill is just instructions plus a query. It will work in Elastic Agent Builder, a Kibana workflow agent, Claude Code, or any other harness. We’ll be using Deep Agents for examples of how to query it outside the Kibana ecosystem. Since an AI Index is, at its core, an Elasticsearch index, you can also explore your data directly. </p><h2>Precompute facts as KIs for agentic RAG</h2><p>In this example, we extract actual facts so agents can retrieve an answer without consuming a full document. We generate one fact-based KI per selected document, though the actual number and structure of KIs you generate are completely customizable.</p><p>We'll use a sample of the <a href="https://github.com/texttron/BrowseComp-Plus">BrowseComp-Plus</a> corpus, indexed into a <code>browsecomp-plus</code> index, with <code>docid</code>, <code>url</code>, <code>title</code>, and <code>text</code> fields.</p><h3>Baseline: Retrieving whole documents with RRF</h3><p>As a baseline, here's a simple <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion">RRF</a> query:</p>POST /_query?format=txt
{
  "query": """
    FROM browsecomp-plus METADATA _score, _id, _index
    | FORK
        (WHERE match(title, "What was the actress who played Torvi from Vikings also known for?") | SORT _score DESC | LIMIT 100)
        (WHERE match(text,  "What was the actress who played Torvi from Vikings also known for?") | SORT _score DESC | LIMIT 100)
    | FUSE // uses RRF by default
    | SORT _score DESC
    | KEEP _id, title, text
    | LIMIT 10
  """
}<p>This drops several hundred words of raw body text into the model's context. It may work, but it's expensive, and the cost compounds with every miss.</p><h3>Build the Kibana workflow</h3><p>The workflow below reads a batch of documents with a single ES|QL query and writes one fact-level KI per document into the AI Index. Each iteration runs two steps: <code>generate_ki</code> distills a raw document into a structured KI, and <code>sink_ki</code> writes it to the AI Index keyed on <code>docid</code> so reruns are idempotent.</p><p>Copy and paste the following YAML into the <a href="https://www.elastic.co/docs/explore-analyze/workflows">Elastic Workflows</a> editor:</p>version: '1'
name: browsecomp-plus-doc-ki
description: Query the BrowseComp-Plus corpus with ES|QL, generate a KI per doc with an AI agent, and bulk-write each into the AI Index as a corpus_entry.
enabled: true
tags:
  - precomputed-context
  - browsecomp-plus
triggers:
  - type: manual
steps:
  - name: query_corpus
    type: elasticsearch.esql.query
    with:
      # WHERE drops empty bodies and restricts to the curated KI_DOCIDS -- the
      # specific documents this example's question depends on -- so the workflow
      # generates only a handful of KIs instead of one per corpus document.
      # SUBSTRING keeps the prompt bounded (a full body would blow the context window).
      # Column order drives the foreach.item[N] indices:
      #   item[0]=docid  item[1]=title  item[2]=url  item[3]=text
      query: &gt;
        FROM browsecomp-plus
        | WHERE text IS NOT NULL AND docid IN ("11589", "50639", "64501", "41758", "57766", "84983", "82008")
        | KEEP docid, title, url, text
        | EVAL text = SUBSTRING(text, 1, 12000)

  - name: loop_corpus_docs
    type: foreach
    foreach: '{{ steps.query_corpus.output.values }}'
    steps:
      # Turn the raw doc into a retrieval-optimized Knowledge Indicator.
      - name: generate_ki
        type: ai.agent
        timeout: 300s
        with:
          message: &gt;
            You are a knowledge engineer building a Knowledge Indicator (KI)
            for an enterprise document-retrieval corpus. A KI is a compact,
            high-signal record that a hybrid (BM25 + semantic) search engine
            and an AI agent use to FIND and JUDGE the source document without
            reading it in full.

            Read the document below and extract a faithful, richly structured KI.
            Follow these rules strictly:
            - Be 100% grounded: never state anything not supported by the text.
            - Prefer concrete, named specifics (people, organizations, products,
              dates, places, figures) over vague phrasing.
            - Write for retrieval, not prose flourish. No marketing language.
            - If a field cannot be determined from the text, return an empty
              string or empty array rather than guessing.

            Document ID: {{ foreach.item[0] }}
            Original Title: {{ foreach.item[1] }}
            Source URL: {{ foreach.item[2] }}
            Document Body:
            {{ foreach.item[3] }}
          schema:
            type: object
            properties:
              title:
                type: string
                description: A concise, specific, human-readable title (&lt;= 12 words).
              summary:
                type: string
                description: A dense 3-5 sentence factual summary capturing the document's main claims, named entities, and conclusions. PRIMARY semantic search surface.
              answers_questions:
                type: array
                items:
                  type: string
                description: 2-5 natural-language questions this document can authoritatively answer.
              key_entities:
                type: array
                items:
                  type: string
                description: 3-10 salient named entities (people, organizations, products, places, dates) explicitly mentioned in the text.
              topics:
                type: array
                items:
                  type: string
                description: 3-8 short topic/category labels.
              tagline:
                type: string
                description: A single ultra-short phrase (&lt;= 6 words) as a quick-reference label.
            required:
              - title
              - summary
              - answers_questions
              - key_entities
              - topics

      # Direct bulk write to the AI Index. The explicit <code>index</code> action row sets
      # _id = docid so re-runs upsert in place (idempotent). <code>index:</code> in <code>with</code>
      # supplies the default target index for the bulk request.
      - name: sink_ki
        type: elasticsearch.bulk
        with:
          index: ai-index-idx-my-corpus
          operations:
            - index:
                _id: '{{ foreach.item[0] }}'
            - '@timestamp': '{{ execution.startedAt | date: "%Y-%m-%dT%H:%M:%S.%LZ" }}'
              type: corpus_entry
              title: '{{ foreach.item[1] | default: steps.generate_ki.output.structured_output.title }}'
              tags:
                - browsecomp-plus
              references:
                uri: '{{ foreach.item[2] }}'
              attributes:
                docid: '{{ foreach.item[0] }}'
                url: '{{ foreach.item[2] }}'
                source_index: browsecomp-plus
                tagline: '{{ steps.generate_ki.output.structured_output.tagline }}'
                topics: '{{ steps.generate_ki.output.structured_output.topics | json }}'
                answers_questions: '{{ steps.generate_ki.output.structured_output.answers_questions | json }}'
                key_entities: '{{ steps.generate_ki.output.structured_output.key_entities | json }}'
              content: &gt;
                === SOURCE / PROVENANCE ===
                Backing Elasticsearch index: browsecomp-plus
                Document ID (docid): {{ foreach.item[0] }}
                Source URL: {{ foreach.item[2] }}
                Retrieve the full original document with ES|QL:
                FROM browsecomp-plus | WHERE docid == "{{ foreach.item[0] }}"
                === KNOWLEDGE INDICATOR ===
                {{ steps.generate_ki.output.structured_output.summary }}
                Questions this document answers: {{ steps.generate_ki.output.structured_output.answers_questions | join: " | " }}
                Key entities: {{ steps.generate_ki.output.structured_output.key_entities | join: ", " }}
              description: &gt;
                {{ steps.generate_ki.output.structured_output.tagline }}.
                Topics: {{ steps.generate_ki.output.structured_output.topics | join: ", " }}.
                Entities: {{ steps.generate_ki.output.structured_output.key_entities | join: ", " }}.<p>Here’s what this workflow is doing: </p><ul><li><p><code>query_corpus</code> runs an ES|QL query against the <code>browsecomp-plus</code> index, applying some rules, like dropping documents with empty bodies and trimming each body to 12,000 chars so the agent prompt stays inside the context window.</p></li><ul><li><p>Note: In this example, we’re cherry-picking some concrete KI IDs, because generating KIs for every document in the index would take a long time, and we want this exercise to be short for those following along.</p></li></ul><li><p><code>loop_corpus_docs</code> iterates over every returned document, running the following two steps per document: </p></li><ul><li><p><code>generate_ki</code> reads the document and calls an LLM to emit a strictly grounded, structured KI.</p></li><li><p><code>sink_ki</code> bulk-writes each KI into the AI Index (<code>ai-index-idx-my-corpus</code>) as a KI of type <code>corpus_entry</code>. It forces <code>_id</code> to be the same as the document’s <code>docid</code> so rerunning the workflow is idempotent.</p></li></ul></ul><p>To summarize, this workflow turns each raw corpus document into a compact, searchable metadata record that agents can find and judge without reading the full source into the context window.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5a37c99585270959/6a8ff23da1b20b401c8728c7/unnamed.png" alt="Kibana Workflow browsecomp-plus-doc-ki: query_corpus, generate_ki and sink_ki write a corpus_entry KI to the AI Index" /><p>This workflow is used for example purposes, and the same <code>foreach</code> caveat as in Part 1 applies. For scale, use <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/composition"><code>workflow.executeAsync</code></a> or native parallel support. The <a href="https://www.elastic.co/docs/explore-analyze/workflows/reference/cheat-sheet">cheat sheet</a> is useful for optimizing Workflows. There could also be cost and efficiency gains in production by using <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/ai-steps#ai-prompt"><code>ai.prompt</code></a> or by choosing different models with which to create KIs. </p><h3>Inspect the KIs in your AI Index</h3><p>Once the workflow runs, you can query the AI Index to browse what was written:</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt626b3ba4dc3a52c0/6a8ff268971ef9107f537cb5/unnamed.png" alt="ES|QL query in Kibana Discover returning five corpus_entry Knowledge Indicators from an Elasticsearch AI Index" /><p>Here’s an example of what one of the KI documents looks like: </p>{
  "_index": "ai-index-idx-my-corpus",
  "_id": "57766",
  "_version": 1,
  "_seq_no": 0,
  "_primary_term": 1,
  "found": true,
  "_source": {
    "@timestamp": "2026-08-05T20:35:39.034Z",
    "type": "corpus_entry",
    "title": "Vikings (TV series) - Wikipedia",
    "tags": [
      "browsecomp-plus"
    ],
    "references": {
      "uri": "https://en.wikipedia.org/wiki/Vikings_%28TV_series%29"
    },
    "attributes": {
      "docid": "57766",
      "url": "https://en.wikipedia.org/wiki/Vikings_%28TV_series%29",
      "source_index": "browsecomp-plus",
      "tagline": "Ragnar Lothbrok's rise and legacy",
      "topics": """["Historical drama television","Viking Age","Norse mythology and sagas","Canadian-Irish co-production","Television cast and production","Medieval Scandinavia"]""",
      "answers_questions": """["When did the Vikings TV series premiere and on which network?","Who created and wrote the Vikings TV series?","Where was the Vikings TV series filmed?","Who are the main cast members of Vikings?","What historical and literary sources inspired the Vikings TV series?"]""",
      "key_entities": """["Michael Hirst","Travis Fimmel","Katheryn Winnick","History Channel","Amazon Prime Video","Ashford Studios","County Wicklow, Ireland","Vikings: Valhalla","Ragnar Lodbrok","Wardruna"]"""
    },
    "content": """=== SOURCE / PROVENANCE === Backing Elasticsearch index: browsecomp-plus Document ID (docid): 57766 Source URL: https://en.wikipedia.org/wiki/Vikings_%28TV_series%29 Retrieve the full original document with ES|QL: FROM browsecomp-plus | WHERE docid == "57766" === KNOWLEDGE INDICATOR === Vikings is a historical drama television series created and written by Michael Hirst, co-produced between Canada and Ireland, that premiered on the History Channel on March 3, 2013, and concluded on March 3, 2021, after 6 seasons and 89 episodes. The series is inspired by the sagas of legendary Norse hero Ragnar Lodbrok — drawing on 13th-century texts Ragnars saga Loðbrókar and Ragnarssona þáttr, as well as Saxo Grammaticus' Gesta Danorum — and follows Ragnar's rise from farmer to Scandinavian king, then the exploits of his sons across England, Scandinavia, Kievan Rus', the Mediterranean, and North America. Principal cast includes Travis Fimmel as Ragnar Lothbrok, Katheryn Winnick as Lagertha, Gustaf Skarsgård as Floki, and Alexander Ludwig as Bjorn Ironside, among many others. The series was filmed entirely in Ireland at Ashford Studios and County Wicklow, with additional location shoots in Iceland, Morocco, Norway, and Canada; the first season budget was US$40 million. A sequel series, Vikings: Valhalla, premiered on Netflix on February 25, 2022. Questions this document answers: When did the Vikings TV series premiere and on which network? | Who created and wrote the Vikings TV series? | Where was the Vikings TV series filmed? | Who are the main cast members of Vikings? | What historical and literary sources inspired the Vikings TV series? Key entities: Michael Hirst, Travis Fimmel, Katheryn Winnick, History Channel, Amazon Prime Video, Ashford Studios, County Wicklow, Ireland, Vikings: Valhalla, Ragnar Lodbrok, Wardruna
""",
    "description": """Ragnar Lothbrok's rise and legacy. Topics: Historical drama television, Viking Age, Norse mythology and sagas, Canadian-Irish co-production, Television cast and production, Medieval Scandinavia. Entities: Michael Hirst, Travis Fimmel, Katheryn Winnick, History Channel, Amazon Prime Video, Ashford Studios, County Wicklow, Ireland, Vikings: Valhalla, Ragnar Lodbrok, Wardruna.
"""
  }
}<h3>Query KIs from LangChain Deep Agents</h3><p>We’ll use <a href="https://docs.langchain.com/oss/python/deepagents/overview">LangChain Deep Agents</a> with an OpenAI-compatible key to show that AI indices and KIs will work with any agent harness, inside and outside of Kibana’s Agent Builder ecosystem. </p><p>First, let’s create <code>facts_baseline_agent.py</code> to measure our baseline before applying KIs: </p># Example question: What was the actress who played Torvi from Vikings also known for?
import os
import sys
import time
from elasticsearch import Elasticsearch

from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM browsecomp-plus | LIMIT 5'.
               Full-text search syntax: WHERE MATCH(field, "value") — not field MATCH "value".
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


@tool
def get_mapping(index: str) -&gt; dict:
    """Return the field mapping for an Elasticsearch index or pattern."""
    return es.indices.get_mapping(index=index).body


baseline_agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query, get_mapping],  # no query-ki skill
    system_prompt=(
        "You are a research assistant answering questions about a document corpus "
        "stored in the Elasticsearch index <code>browsecomp-plus</code> (fields: docid, url, "
        "title, text). You have NOT memorized the corpus. Answer by querying the raw "
        "index directly with ES|QL via the esql_query tool. "
        "Full-text search syntax: WHERE MATCH(field, \"value\") — never use field MATCH \"value\". "
        "Use get_mapping if you are unsure of field names. Ground your answer strictly "
        "in the rows returned, and cite the docid or url you used."
    ),
)

start = time.perf_counter()
result = baseline_agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>When I ran this baseline query, <em>What was the actress who played Torvi from Vikings also known for?</em>, it output the following: </p>% python facts_baseline_agent.py "What was the actress who played Torvi from Vikings also known for?"

--- Tool calls ---
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Torvi Vikings actress") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Georgia Hirst") | LIMIT 10'}
  [read_file] {'file_path': '/large_tool_results/toolu_bdrk_01SVEwHdtc9uE1zdfgCwchb5', 'offset': 0, 'limit': 50}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Torvi Vikings TV") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Vikings series cast actress") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Michael Hirst Vikings") | LIMIT 10'}
  [esql_query] {'query': 'FROM browsecomp-plus | WHERE MATCH(text, "Torvi") | LIMIT 20'}
  [read_file] {'file_path': '/large_tool_results/toolu_bdrk_01PS3ryzML5BstqMXnV4c65N', 'offset': 0, 'limit': 100}
Total: 8

--- Usage ---
Tokens: 386187 (input 384940, output 1247)
Latency: 44.86s

--- Answer ---
Perfect! I found the information. Based on the document from docid 11589 (Georgia Hirst's Wikipedia page), I now have the answer to the question.

**Georgia Hirst**, the actress who played Torvi in the History Channel series Vikings (2014-2020), is also known for:

1. **Her role as Becky in the zombie horror movie "Ravers" (2018)** - This film had its world premiere at the 2018 London FrightFest Film Festival.

2. **Being the daughter of Michael Hirst**, the creator and writer of Vikings. This is a notable connection that's frequently mentioned about her.

3. **Her role in the short film "Dungeness" (2018)**, directed by Stephen Gates, where she was the sole actor.

4. **Her role as Grace in the interactive rom-com "Five Dates" (2020)**.

5. **Her advocacy work** - In 2017, she made an appearance on Sky News urging young women to get smear tests after being diagnosed with precancerous cells at age 22.

6. **Her philanthropy** - She joined The Celebs in 2021 to record a cover of The Beatles' "Let It Be" in support of the Mind charity.

So while she's primarily known for her role as Torvi in Vikings (where she appeared in 63 episodes across seasons 2-6), she's also recognized for her work in horror films like "Ravers," her health advocacy, and being part of the Hirst family that created the show.<p>(Note: Deep Agents automatically adds the <code>read_file</code> tool to handle paginated tool results, which is why it shows up in the output.) </p><p>Next, let’s create an agent that knows how to use our <code>query-ki</code> skill, <code>facts_ki_agent.py</code>: </p># Example question: What was the actress who played Torvi from Vikings also known for?
import os
import sys
import time
from elasticsearch import Elasticsearch
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent
from deepagents.backends.filesystem import FilesystemBackend

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM ai-index-idx-* | LIMIT 5'.
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


# FilesystemBackend loads skills from disk, relative to root_dir.
backend = FilesystemBackend(root_dir=".", virtual_mode=False)

agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query],
    skills=["skills"],
    backend=backend,
    system_prompt=(
        "You are a research assistant answering questions about a document corpus. "
        "You have NOT memorized the corpus. When a question depends on specific facts, "
        "names, dates, or events, use the query-ki skill to retrieve Knowledge "
        "Indicators before answering. Ground your answer strictly in what it returns, "
        "and cite the KI titles you used."
    ),
)

start = time.perf_counter()
result = agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>This agent will query the KI indices to generate the answer, rather than pulling every document into context. </p><p>When I ran these results against the same query, here was the output: </p>% python facts_ki_agent.py "What was the actress who played Torvi from Vikings also known for?"

--- Tool calls ---
  [read_file] {'file_path': './query-ki/SKILL.md', 'limit': 1000}
  [esql_query] {'query': 'FROM ai-index-idx-* METADATA _id, _index, _score\n| WHERE type == "corpus_entry"\n| FORK\n    (WHERE MATCH(co
Total: 2

--- Usage ---
Tokens: 27625 (input 27037, output 588)
Latency: 15.22s

--- Answer ---
Based on the retrieved information, **Georgia Hirst** played Torvi in Vikings. She is also known for:

1. **Ravers (2018)** - A zombie horror film where she played the character Becky. This was her debut film and premiered at the London FrightFest Film Festival.

2. **Being the daughter of Michael Hirst** - The creator and writer of Vikings itself, making her connection to the show a notable family affair. Her older half-sister Maude Hirst also appeared in Vikings as Helga.

3. **Cervical cancer awareness advocacy** - She has publicly advocated for cervical cancer screening after being diagnosed with precancerous cells at age 22 and successfully recovering through treatment.

4. **Charity work** - In 2021, she participated in a celebrity cover of The Beatles' "Let It Be" in support of the Mind charity (a mental health organization), alongside Anne Hegerty, Ivan Kaye, Eunice Olumide, and Shona McGarty.

**Sources cited:** "Georgia Hirst" and "Georgia Hirst - Wikipedia" Knowledge Indicators from the AI Index.<h2>How much can precomputing facts reduce agent token usage?</h2><p>Both agents had similar conclusions, but they took far different paths to get there: </p><p>The same question and the same grounded answer result in 93% fewer tokens and two tool calls instead of eight, when answering from KIs.</p><p>
</p><p>Baseline (No AI Index)</p><p>With AI Index</p><p>Total tool calls</p><p>8</p><p>2</p><p><code>read_file</code> calls</p><p>2</p><p>1</p><p><code>esql_query</code> calls</p><p>6, all against the <code>browsecomp-plus</code> index</p><p>1, from <code>ai-index-idx-*</code></p><p>Tokens consumed</p><p>386,187</p><p>27,625</p><p>Latency</p><p>44.86s</p><p>15.22s</p><p>Answer</p><p>Grounded, correct</p><p>Grounded, correct</p><p>Exact tool call counts, latency, and answers will vary between runs and using different agents. </p><p>Both agents produced solid, grounded answers. The difference is cost. Querying KIs from the AI Index cut token use by 93% and cut latency by roughly two thirds. Here’s how both paths went, side by side:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt52fac74907997882/6a8ff2a049d4293b02a4fd64/unnamed.png" alt="Agentic RAG tool calls: 8 calls and 386,187 tokens without Knowledge Indicators, 2 calls and 27,625 tokens with them" /><p>That was in-depth, but it shows what AI indices and Workflows do together: the same answer, at a fraction of the tokens.</p><h2>Build precomputed context in Elasticsearch Serverless</h2><p>This walkthrough shows how to generate more sophisticated KIs based on documented facts and query them for knowledge retrieval use cases using Elasticsearch primitives. </p><p>Managing context is critical in agentic search systems. And at its core, context is a retrieval problem. AI indices help you manage context within the Elastic Stack. Try it out in Serverless, and let us know what you think in our <a href="https://discuss.elastic.co/top?period=monthly">Discuss forums</a> or the <code>#stack-kibana</code> channel in our <a href="https://elasticstack.slack.com/signup#/domain-signup">Community Slack</a>.</p><p>We’d also love to hear from you about what use cases you’d like to solve using AI indices.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/agentic-rag-precomputed-facts-ai-index</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/agentic-rag-precomputed-facts-ai-index</guid>
    <category><![CDATA[AI Tools ]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Kathleen DeRusso,Matt Nowzari ,Apostolos Matsagkas,Peter Pišljar]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3e1939e01169bb08/6a8fedbec8ced9f736055f59/1.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Taming PUNKs: How ES|QL queries Elasticsearch fields it was never told about]]></title>
    <description><![CDATA[In Elasticsearch 9.5, ES|QL can query unmapped fields. It reads them from _source or returns nulls, so a query keeps working when a field drops out of the mapping and you avoid a reindex that takes hours.]]></description>
    <content:encoded><![CDATA[<p>How do you make an analytical query engine use data that it cannot know exists? You “just” read the query, since everything that the user asks for is right there. Right?</p><p>In Elasticsearch 9.5, <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> queries no longer fail when a field isn't in the mapping. The new <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-unmapped-fields"><code>unmapped_fields</code> setting</a> lets queries load values from <code>_source</code> or fill with <code>nulls</code>, so queries keep working even when a backing index changes and a field goes missing, and can use unmapped data without reindexing. Here’s how we built that: the design choices and the edge cases (including a class of fields we nicknamed PUNKs), along with the testing strategies that gave us the confidence to ship it in general availability (GA).</p><h2>Why ES|QL queries fail when a field is unmapped</h2><p>You built a visualization using an ES|QL query. You refined it, and the query grew. You’re at 15 chained commands and counting, but it does <em>just</em> the right thing. It works, and your dashboard is <em>useful</em>.</p><p>Your query uses an index from a remote cluster, say <code>my-remote:logs-foo</code>. But actually, <code>logs-foo</code> is an alias, and at some point, the remote cluster makes it point to a different backing index. The new index is missing a field that’s used in your query, and your query and visualization break.</p><p>Or maybe you have an already fairly large index, and while building ES|QL queries on top of it, you realize that you’d like to use a field in the indexed documents that unfortunately never made it into the index mapping. You could reindex the data, but that would take hours.</p><p>ES|QL’s <code>unmapped_fields</code> setting is meant to deal with these types of situations.</p><p>If your query looks like this:</p><p>and <code>some_field</code> is unmapped, ES|QL’s default behavior is to fail with a verification exception.</p><p>You can use the <code>unmapped_fields</code> setting to instead either fill <code>some_field</code> with <code>null</code>s or read it from the document’s <code>_source</code>, like so:</p><h2>How ES|QL resolves queries with field caps</h2><p>Before we jump into the inner workings of <code>unmapped_fields</code>, we have to look into how ES|QL resolves queries regularly. Let’s consider the above query:</p><p>We said that if <code>some_field</code> isn’t in the mapping for <code>index</code>, ES|QL will reject the query. How does it make that decision?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7ba9af4b4b550018/6a8ee83fe41d7fea88654d42/image4.png" alt="ES|QL query resolution flow: analyzer checks index mappings, unresolved fields fail with Unknown column error" /><h3>How field caps tells ES|QL which fields exist</h3><p>In a typical schema-on-write fashion, Elasticsearch clusters maintain mappings with their respective indices. As a first step, ES|QL makes an internal request to the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-field-caps">field caps endpoint</a> to determine which fields the <code>index</code> has. It then passes the query, together with the field caps response, to the query planner, which consists essentially of the query analyzer (unrelated to analyzers of text fields) and query optimizer. The analyzer makes sense of raw names, like <code>some_field</code>, and notices that they correspond to index fields (or not). If all went well, the query is then passed on to the optimizer, which rewrites the query for efficiency, before it’s handed to the compute engine for execution.</p><h3>How the analyzer resolves field names in the query plan</h3><p>Let’s zoom in to the analyzer. The parsed query is represented in a tree structure, and the analyzer partially rewrites it, one command at a time, until it either has resolved all references or not.</p><p>For illustration, let’s use a somewhat more complex query and see how the analyzer would resolve it:</p><p>The parsed tree is actually a chain here, and it looks something like this:</p><p></p><p>The analyzer then moves up through the query tree to try and resolve the field names used in every command.</p><p>This is a simplified version of how we represent parse trees in tests and when debugging. The bottom of the chain corresponds to the <code>FROM</code> command and contains a list of all mapped fields that we know about, obtained from the field caps endpoint. (The <code>{f}</code> suffix marks an actually mapped field for better distinction later.)</p><p>The two <code>EVAL</code> nodes on top of it correspond to the remaining commands, and their fields are still unresolved, expressed by the question mark <code>?</code> in front of the name. At this point, the analyzer still has to check whether they correspond to existing index fields.</p><p>For the <code>EVAL</code> that defines <code>uppercased_mapped</code>, it can see that the previous command outputs <code>mapped_field</code>, so the unresolved <code>?mapped_field</code> marker can be replaced by a real field reference:</p><p>Next, it encounters the topmost <code>EVAL</code>, which defines <code>uppercased_unmapped</code>. The previous tree nodes produce only two fields: <code>[mapped_field, uppercased_mapped]</code>. The reference <code>?unmapped_field</code> thus has to remain unresolved. We bail here and emit the verification exception to the user.</p><h2>How unmapped_fields LOAD and NULLIFY work</h2><h3>Adding unmapped fields to the query plan</h3><p>When using <code>unmapped_fields=”NULLIFY”</code> or <code>”LOAD”</code>, we do something else; we act as if the field was actually in the index. The analyzer adds <code>unmapped_field</code> to the <code>From</code> node and marks it as unmapped to signal to the compute engine that this has to be read from <code>_source</code> or filled with <code>null</code>s. Let’s express this with a <code>{u}</code> (for <strong>u</strong>nmapped):</p><p>After amending the <code>From</code>, the analyzer can continue trying to resolve the topmost <code>Eval</code> node. It sees that the upstream nodes produce the fields <code>[mapped_field, unmapped_field, uppercased_mapped]</code> and thus <code>unmapped_field</code> can be correctly resolved:</p><p></p><p>The query plan is now fully resolved and can be passed down the regular optimization-execution pipeline. Other than the actual value extraction mechanism, everything stays the same. Schematically, the workflow looks like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt59bf2d2da201d00a/6a8ee8df8658b77c28469356/image2.png" alt="ES|QL unmapped fields flow: analyzer retries with NULLIFY or LOAD instead of failing on an unresolved field" /><h3>Example: enabling unmapped fields with the SET directive</h3><p>To give an example, let’s fire up a cluster and create an index with non-dynamic mappings.</p>PUT /index
{                                 
  "mappings": {
    "dynamic": false,
    "properties": {
      "mapped_field": {"type": "keyword"}
    }
  }
}

POST /index/_doc?refresh
{
  "mapped_field":"foo"
  "unmapped_field": "bar"
}<p>We can run the example query, above:</p>POST /_query
{
  "query": """
           FROM index
           | EVAL uppercased_mapped = TO_UPPER(mapped_field)
           | EVAL uppercased_unmapped = TO_UPPER(unmapped_field)
           """
}<p>This should result in the error message:</p><p><code>Unknown column [unmapped_field], did you mean [mapped_field]?</code></p><p>To make things work, we can prepend <code>SET unmapped_fields=”...”;</code> with <code>LOAD</code> or <code>NULLIFY</code>:</p>POST /_query
{
  "query": """
           SET unmapped_fields="LOAD";
           FROM index
           | EVAL uppercased_mapped = TO_UPPER(mapped_field)
           | EVAL uppercased_unmapped = TO_UPPER(unmapped_field)
           """
}

 mapped_field  |unmapped_field |uppercased_mapped|uppercased_unmapped
---------------+---------------+-----------------+-------------------
foo            |bar            |FOO              |BAR<h3>Inspecting the analyzer's rewrite steps</h3><p>If you want to see what the query analyzer is doing to the parse tree, you can log the query rewrite steps, like so:</p>PUT /_cluster/settings"
{
  "transient" : {
    "logger.org.elasticsearch.xpack.esql.analysis.Analyzer.changes": "TRACE"
  }
}<p>This will log a line containing <code>Rule rules.ResolveUnmapped applied with change…</code> You’ll see that <code>unmapped_field</code> is added to the bottom of the parse tree as described above.</p><h2>Why we have to infer the schema</h2><p>Of course, this isn’t the only possible method to deal with unmapped fields. Here are some alternatives:</p><ol><li><p>We could also scan or probe the documents in <code>index</code> to determine that their <code>_source</code> actually has the <code>unmapped_field</code>.</p></li><li><p>We could disable verifications in the analyzer and make the compute engine blindly pass unmapped fields through individual computation steps.</p></li></ol><p>The first alternative front-loads more work to understand the <em>actual</em> schema of an index and thus generally increases latency. It doesn’t scale to large, highly distributed datasets. The second alternative isn’t viable since it means a large-scale change to how ES|QL’s compute engine is built, because it passes around streams of data with fixed columns from one operator to another.</p><p>In contrast, the approach we chose is neatly compatible with ES|QL’s existing optimization pipeline.</p><p>The trade-off is that the analyzer has to correctly <em>infer</em> a schema based on the actual index mappings (obtained from the field caps endpoint) and additional fields used inside the query. </p><p>This isn’t always straightforward. There were two main challenges:</p><ol><li><p>There are many different query shapes and commands that can be used. The mechanism needs to detect unmapped fields, update the proper <code>FROM</code> command, and pass the new field through the halfway resolved plan correctly in all cases.</p></li><li><p>There are many different mappings we have to deal with, and we specifically need to make our feature work correctly when mappings <em>change over time</em> on top of that.</p></li></ol><p>In the following, we’ll focus on <code>LOAD</code>, although some problems (generally many fewer) also apply to <code>NULLIFY</code>.</p><h3>Which index to load unmapped fields from for LOOKUP JOIN and FORK</h3><p>To briefly illustrate the first problem, here are some choices we needed to make:</p><ul><li><p>Which index do we load from when using <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join">lookup joins</a>? This one?</p><p>The <code>unmapped_field</code> cannot be attributed to both indices. We chose <code>index</code> since this is where we expect mappings to change more often than in lookup indices.</p></li><li><p>Similarly, how do we deal with subqueries and views or the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fork"><code>FORK</code> command</a>? In the following:</p><p>one fork branch triggers loading of an unmapped field. Is it also present in the other fork branch? (Yes, it should be, but it’s not obvious and is specifically not true if the two <code>FORK</code>s are replaced by independent subqueries.)</p></li></ul><h3>Two principles to keep queries working</h3><p>The second problem, diversity of mappings and their evolution over time, is a far bigger driver for complexity. We strove for two basic usability principles:</p><ul><li><p>Queries that work in the default mode should generally still work when using <code>unmapped_fields=”NULLIFY”</code> and <code>”LOAD”</code>.</p></li></ul><ul><li><p>Queries that work when all fields are mapped should generally still work with <code>NULLIFY</code> and <code>LOAD</code> when a field becomes unmapped and vice versa.</p></li></ul><h3>The type of unmapped fields and inadvertent type conflicts</h3><p>Let’s talk about data types to see where this leads to complexity. First, when using <code>unmapped_fields=”LOAD”</code>, we need to assume a data type for unmapped fields. We chose <code>KEYWORD</code>, which allows us to avoid type conflicts when reading from <code>_source</code>. One document can contain <code>”unmapped_field”: “foo”</code>, and another can contain <code>”unmapped_field”: 123.4</code>. It’s fine because we treat both as strings.</p><p>However, this is a violation of the second principle when a non-<code>KEYWORD</code> field happens to go unmapped. Consider this query:</p><p>If <code>some_field</code> becomes unmapped, we’ll have to assume that the <code>KEYWORD</code> type and the query will fail with a type conflict.</p><p>Type conflicts <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-multi-index#esql-multi-index-invalid-mapping">aren’t new</a> and can be dealt with by using explicit casts in the query, like so:</p><p>It would be great if ES|QL just inferred a useful type to cast to, but this is something for the future.</p><h3>Type conflicts with partially unmapped fields, or: making PUNKs well behaved</h3><p>In addition to fully unmapped fields, <em>partially unmapped</em> fields are everywhere and should also work with <code>LOAD</code>. Let’s look at a query that uses multiple indices.</p><p>Let’s say that there are indices <code>index</code> and <code>index_without_some_field</code>, containing just the following documents.</p>// index1
{
  "some_field": "foo"
}

// index2
{
  "some_field": "bar"
}<p>Now let’s consider the query:</p>FROM index, index_without_some_field<p>and assume that <code>some_field</code> is unmapped in <code>index_without_some_field</code>. This will return:</p>some_field
-------------
 foo
 null<p>because ES|QL doesn’t load unmapped fields per default.</p><p>Of course, when setting <code>unmapped_fields=”LOAD”</code>, we want to load from <code>_source</code> for <code>index_without_some_field</code>:</p>SET unmapped_fields="LOAD";
FROM index, index_without_some_field

 some_field
-------------
 foo
 bar           // loaded from _source<p>As with fully unmapped fields, the case is simple when <code>some_field</code> is mapped as <code>KEYWORD</code> in <code>index</code>. When loading from <code>_source</code> for <code>index_without_some_field</code>,  we treat the field as <code>KEYWORD</code> as well, so there’s no conflict.</p><h3>What makes a field a PUNK</h3><p>The case is less clear when <code>some_field</code>is partially unmapped and the mapped leg is of a type other than <code>KEYWORD</code>. Such fields caused a lot of trouble until we found the best solution, which makes their acronym quite fitting: <strong>p</strong>artially <strong>u</strong>nmapped <strong>n</strong>on-<strong>k</strong>eyword fields, or PUNKs.</p><p>Unfortunately, PUNKs are far from being esoteric. For instance, it’s very natural to filter on a PUNK:</p><p>If <code>some_field</code> is mapped as <code>INTEGER</code> in <code>index</code>, the type conflict looks like this:</p><ul><li><p>Mapped as an <code>INTEGER</code> in <code>index</code>.</p></li><li><p>Unmapped in <code>index_without_some_field</code> and thus treated as <code>KEYWORD</code>.</p></li></ul><p>This can again be resolved manually by providing an explicit cast:</p><p>But this is far from acceptable. Even queries that work fine without <code>NULLIFY</code> and <code>LOAD</code> typically have <em>some</em> PUNKs; the unmapped leg is simply treated as <code>null</code> then. Both guiding principles are violated if <code>LOAD</code> requires an explicit cast here.</p><h3>Casting implicitly to the mapped type</h3><p>The solution is to introduce an implicit cast to the mapped type. In this case, we know that <code>some_field</code> is an <code>INTEGER</code> in <code>index</code>, and thus we treat it essentially as if the user wrote:</p><p>This means that queries that work without <code>LOAD</code> keep working. (ES|QL may even give you more data because we load the unmapped leg of PUNKs from <code>_source</code>.) Queries that used to work when a field is fully mapped also keep working when it goes unmapped in some (but not all) of its indices without having to alter the query in any way.</p><p><strong>Behavior</strong></p><p><strong>Default</strong></p><p><strong><code>NULLIFY</code></strong></p><p><strong><code>LOAD</code></strong></p><p>Unmapped field in query</p><p>Query fails</p><p>Query runs</p><p>Query runs</p><p>Values returned</p><p>None</p><p><code>null</code></p><p>Read from <code>_source</code> </p><p>Assumed type</p><p>n/a</p><p><code>NULL</code></p><p><code>KEYWORD</code></p><p>Partially unmapped field (PUNK)</p><p>Unmapped leg is <code>null</code></p><p>Unmapped leg is <code>null</code></p><p>Cast to the mapped type</p><p>Pushdown optimization</p><p>Full</p><p>Full</p><p>Per-node where fully mapped</p><h2>Don't throw it all away: Keeping ES|QL query optimization with unmapped fields</h2><p>There's one more thing to get right; that is, to make sure that optimizations still work correctly with <code>LOAD</code>. Consider the previous query:</p><p>ES|QL’s optimizer aggressively pushes down such <code>WHERE</code> filters and turns them into Lucene queries, so the compute engine doesn’t perform unnecessary work.</p><p>For this query, evaluating the filter in the compute engine would require fetching each and every document from the index; meaning, a full scan, very slow. If <code>some_field</code> was mapped as an <code>INTEGER</code> in both indices, we would instead perform a Lucene query, which looks like this:</p>{
  "range": {
    "some_field": {
      "gt" : 10,
      "boost" : 0.0
    }
  }
}<p>The compute engine then doesn’t have to load each document separately and check whether it matches the filter. Documents with <code>some_field &lt;= 10</code> are never fetched from the Lucene index, which is very efficient at this kind of filtering. Nice.</p><h3>Why filter pushdown is unsafe for unmapped fields</h3><p>If <code>some_field</code> is unmapped in <code>index_without_some_field</code>, however, it’s wrong to narrow the documents down using the same Lucene query, as Lucene interprets an unmapped <code>some_field</code> as <code>null</code> and thus no documents from <code>index_without_some_field</code> will ever match. This edge case is easy to miss, and it doesn’t help that there are several flavors of similar pushdowns. For instance, in the query:</p><p>the compute engine pushes even the counting to Lucene. Again, this is only correct if <code>some_field</code> is fully mapped.</p><p>This means that such optimizations can't apply to unmapped fields. It would be disappointing if a query used hundreds of indices and only one of them happened to not map <code>some_field</code>, causing the whole query to run unoptimized.</p><h3>How the local optimizer recovers the fast path</h3><p>Luckily, this problem has a solution, too. ES|QL actually has multiple optimizer runs:</p><ol><li><p>First, a preliminary optimizer run on the node handling the <code>_query</code> request.</p></li><li><p>Then, a second, local optimizer run on every node we fan out to because we need to fetch documents from its shards.</p></li></ol><p>The workflow after the initial optimization looks more like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb2c907ea31960f96/6a8eeabe36492afa7685fa36/image1.png" alt="" /><p>If the current node happens to map <code>some_field</code> in all shards, the local optimizer detects this situation and treats <code>some_field</code> like any other fully mapped field, including performing Lucene queries to greatly narrow down the dataset to be processed. In fact, data nodes process <code>LIMIT</code> queries like:</p><p>in batches of shards (to avoid loading too much data too eagerly), which includes a full local optimizer run per batch. This makes it even more likely to encounter batches where <code>some_field</code> is fully mapped, allowing ES|QL to run a fast Lucene query.</p><h2>Is it working now? Testing unmapped_fields across every ES|QL query shape</h2><p>As we have seen from the optimizer issues above, problems can hide in plain sight, even for very simple queries. Because <code>unmapped_fields=”LOAD”</code> can affect each and every kind of query, the surface area for bugs is essentially all of ES|QL.</p><p>Accordingly, getting good test coverage was tricky and challenged us to refine our testing strategies.</p><h3>Reusing spec tests with unmapped_fields</h3><p>Conveniently, ES|QL has an extensive corpus of test queries, together with expected result sets; we call them <em>spec tests</em> because they’re written using a simple text specification language, which looks roughly like this:</p>simpleEval
row a = 1 | eval b = 2
;

a:integer | b:integer
1         | 2
;<p>This lets us create new tests out of the existing ones by introducing slight variations. For instance, any existing test that runs without <code>SET unmapped_fields=”...”</code> should produce the exact same results when run with <code>SET unmapped_fields=”NULLIFY”</code>.</p><p>It also helped find major issues early in the development process, especially for <code>NULLIFY</code>. The <code>LOAD</code> setting changes the meaning of queries much more dramatically, limiting the usefulness of this approach. However, ES|QL also uses what we call <em>generative testing</em>; that is, we string together random commands, run the query, and then check whether the server reports a bug. This approach cannot confirm the correctness of results, but it still helped greatly with finding query types that didn’t work properly and resulted in some kind of error. (Property-based tests would be a refinement in the future by running the queries against a reference implementation. This way, correctness of results can also be checked.)</p><h3>Testing type conflicts across different mappings</h3><p>In the end, one of the most important testing dimensions was using different indices with various mappings in the same query. (Recall how, above, we had to deal with type conflicts to come up with a solid approach for PUNKs? It doesn’t end there; all kinds of type conflicts are more complex with <code>LOAD</code>.) Since we couldn’t automatically generate correct expected results, ES|QL’s test suite had to grow by adding more than 10,000 lines of CSV spec tests. Fortunately, adding such tests is a well-suited task for an AI agent, which has cut down the effort dramatically. (Of course, the test results were still reviewed by humans.)</p><p>All testing strategies together provided us with good confidence for the GA release of <code>unmapped_fields</code> with Elasticsearch 9.5.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-unmapped-fields-deep-dive</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-unmapped-fields-deep-dive</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Alexander Spies]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f8079c053533ad8/6a8ee79f8658b748a0469342/image4.png" length="0" type="image/png"/>
    <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Three SLOs every search team needs: monitoring search latency, availability and quality with OpenTelemetry]]></title>
    <description><![CDATA[Your OpenTelemetry search spans already carry the signals for SLOs, burn rate alerts, anomaly detection and incident response, and this post shows how to build all four in Elastic Observability.]]></description>
    <content:encoded><![CDATA[<p>Every search request your API handles already emits an OpenTelemetry span with latency, error, and result count data. You built that instrumentation for product analytics. Turns out it also gives you search monitoring for free. This post takes those spans and turns them into three SLOs (99% of queries under 250ms, 99.9% availability, zero-results rate below 15%), then layers on alerting, anomaly detection and an incident response workflow, all with Elastic Observability's built-in tooling. If you instrumented your search API following Blogs 2-4, you can set this up in an afternoon.</p><h2>What you'll discover</h2><p>In this post, you'll learn how to:</p><ul><li><p>Use Elastic APM's built-in views to explore search latency and throughput and to explore errors.</p></li><li><p>Define Service Level Objective (SLOs) for search, including latency targets and availability, along with search quality.</p></li><li><p>Create SLOs in Kibana that track your search health over time, with burn rate alerting.</p></li><li><p>Build operational dashboards with Elasticsearch Query Language (ES|QL) that show latency percentiles and time breakdowns and that include trends.</p></li><li><p>Set up alerts for latency regressions and error spikes, along with zero-results rate increases.</p></li><li><p>Establish an incident response pattern for search degradation.</p></li></ul><h2>What you'll need</h2><ul><li><p>Search instrumentation from <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blogs 2–4</a> (search spans with <code>search.*</code>  attributes in Elastic).</p></li><li><p>Kibana access with permissions to create SLOs and alert rules.</p></li><li><p>Basic understanding of SLOs. (We'll explain the search-specific parts.)</p></li><li><p>An Elastic cluster with an Enterprise subscription, an <a href="https://www.elastic.co/cloud?utm_campaign=G-TXT-EMEA-UK+CA-Core-EN-Lead_Gen-CloudTrials-BR&amp;utm_content=Brand-Cloud&amp;utm_source=google&amp;utm_medium=cpc&amp;device=c&amp;utm_term=elastic%20cloud%20trial&amp;utm_id=701610000005lJVAAY&amp;gad_source=1&amp;gad_campaignid=22979576770&amp;gbraid=0AAAAADrDgoKn2OUpnHv5-QMO2ZrcBzj4K&amp;gclid=Cj0KCQjwjb3SBhDgARIsAMKiWziycspjFFEKHsgcIEsVdAYu6qNwIrTL27hQoMuX4eZbcm9oFox0tBYaAmH1EALw_wcB">Elastic Cloud trial</a>, or <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">a local deployment</a> with the trial activated.</p></li></ul><h2>Why search monitoring matters beyond cluster health</h2><p>Search is the primary navigation path for a significant share of visitors, and search-initiated sessions tend to show stronger purchase intent than browse sessions. A search outage is a revenue event, rather than a minor feature degradation. A latency regression from 100ms to 500ms changes user behavior before anyone files a ticket.</p><p>Most platform teams monitor search at the infrastructure level, checking whether the Elasticsearch cluster is healthy and whether nodes are responding. They also determine whether the disk is full. These are all necessary but not sufficient. A cluster can be green while search quality silently degrades; for example, queries returning stale data after a bad index deployment or latency creeping up as the index grows. This could also include zero-results rates climbing because a synonym list wasn't updated.</p><p>The gap is between "search is up" and "search is working well."</p><p><strong>A note on examples:</strong> As we have throughout this series, we use ecommerce search for concrete examples, but these reliability patterns apply equally to any search application, including content platforms, internal knowledge bases, job boards, and support portals.</p><h3>OpenTelemetry search spans as monitoring signals</h3><p>If your team followed <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blogs 2–4</a> in this series, every search request already emits an OTel span with <code>search.*</code> attributes. These include the query text, result count, Elasticsearch execution time, and error status. Those spans land in <code>traces-generic.otel-default</code>  in Elastic.</p><p><strong>Following along with code?</strong> The <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">reference project</a> has all the instrumentation from <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blogs 2–4</a>. Generate traffic, and then follow along with the SLO and dashboard setup below. See <code>queries/blog6_reliability.esql </code> for ready-to-run queries.</p><p>The search team built that instrumentation for <em>product analytics</em>; that is, understanding what users search for and measuring click-through rates (CTRs) and conversion rates. They’ve prioritized relevance work, but the same spans contain everything you need for operational monitoring. Span duration gives you a latency signal, and <code>search.result_count == 0</code> value reflects quality. Span errors point to availability signals.</p><p>This post shows how to put this operational value to work, beginning with what you can see right now in Kibana and then building SLOs and alerting on top of it, along with incident response.</p><h2>Search monitoring out of the box with Elastic APM</h2><p>Before building anything new, let's look at what Elastic APM already gives you out of the box.</p><p>If your search API is instrumented with Elastic Distribution of OpenTelemetry (EDOT) (as in <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a>), it appears automatically as a service in the Elastic APM UI. Open <strong>Observability </strong>&gt;<strong> APM </strong>&gt;<strong> Services</strong> in Kibana, and select your search service (named <code>search-analytics-demo</code> if you're using the reference project). You'll immediately see the following:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt901b364e81dae2dc/6a8d6f3e0897900dc2ef9425/image6.gif" alt=" Elastic APM service overview page for an OpenTelemetry-instrumented search API showing Overview, Transactions and Errors tabs" /><h3>Elastic APM service overview for search</h3><p>The service overview page shows latency distribution and throughput over time, along with error rate, and doesn’t require configuration. You can see at a glance whether search is healthy, and the time-series charts make regressions obvious. If latency crept up after yesterday's deployment, you'll see it here.</p><h3>Trace waterfall: breaking down search request latency</h3><p>Click into any transaction, and you'll see the <em>trace waterfall</em>, which is a visual breakdown of every span in the request. For a search API call, this typically shows:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb005066de9a98d55/6a8d6f82b74e9d482044a3b3/blog6-trace-waterfall-span-details.gif" alt="Elastic APM transaction view for POST /api/search showing search latency distribution, throughput, and failed request rate" /><p>The waterfall makes the invisible visible. You can see that the 503ms API response breaks down into HTTP handling, a 241ms query rules lookup, and a 260ms Elasticsearch query,  plus the custom <code>search</code> span (36ms) carrying all of our <code>search.*</code> attributes. Click any span, and the metadata flyout shows exactly what was captured: <code>search.query: "usb hub"</code><code>,</code> <code>search.result_count: 33</code>, <code>search.took_ms: 43</code>, the index name, hit IDs, and more.</p><p><a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a> discussed the gap between <code>search.took_ms </code> and span duration. The waterfall shows you exactly where that gap lives, without writing any queries.</p><h3>Automatic search error capture with OpenTelemetry</h3><p>One of the most valuable things OTel auto-instrumentation gives you is <em>automatic error capture</em>. When an Elasticsearch query fails because of issues like a tripped circuit breaker or a timeout, or if an index isn’t found,  the span records the exception type and message, along with the stack trace. <a href="https://www.elastic.co/search-labs/blog/search-conversion-tracking-opentelemetry">Blog 4</a> mentioned this as a side benefit of span-based conversion tracking; here it becomes an operational lifeline.</p><p>The errors tab on your service page automatically aggregates these, grouped by error type and frequency. The instrumentation captures the details for you, so you don't need custom error handling or logging. During an incident, this is often the fastest way to understand what's actually failing.</p><h3>Service map: search API and Elasticsearch dependencies</h3><p>The <a href="https://www.elastic.co/docs/solutions/observability/apm/service-map">service map</a> shows dependencies between your search API and Elasticsearch, making it easy to see whether a latency problem is in your service or in the cluster it depends on.</p><p>All of this is available the moment you deploy the instrumentation from <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a>, without building any dashboards or writing any queries. This is the foundation everything else in this post builds on.</p><h2>Defining service level objectives for search</h2><p>An SLO defines <em>good enough</em> in measurable terms. You define what <em>working</em> means, measure it continuously, and alert when you're burning through your error budget too fast, instead of reacting when something breaks.</p><p>Elastic Observability has a <a href="https://www.elastic.co/guide/en/observability/current/slo.html">built-in SLO framework</a> that handles <a href="https://www.elastic.co/guide/en/observability/current/slo.html#slo-important-concepts">Service Level Indicator</a> (SLI) calculation and budget tracking. It also takes care of burn rate alerting. You create SLOs directly in Kibana. No ES|QL or custom pipelines are required for the core indicators.</p><h3>Three SLOs every search service needs</h3><p>Navigate to <strong>Observability </strong>&gt; <strong>SLOs</strong> in Kibana, and click <strong>Create SLO</strong>. The <a href="https://www.elastic.co/guide/en/observability/current/slo-create.html">SLO creation workflow</a> walks you through three steps: Define the SLI (what to measure), set the objective (the target), and describe the SLO.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2ea1fc523a27cb57/6a8d7037bc1d3a69d7701bba/blog6-slo-creation-form.gif" alt="Creating a search latency SLO in Kibana showing SLI preview chart, rolling time window, and 99% target objective setting" /><h4>1. Search latency SLO: 99% of queries under 250ms</h4><p>Indicator type: Elastic APM latency target: 99% of searches complete in under 250ms.</p><p>The Elastic APM latency indicator is purpose-built for this. Select your search service (<code>search-analytics-demo</code>), and set the threshold to 250ms.  Elastic handles the rest, including calculating the percentage of transactions below the threshold and tracking your error budget over time.</p><p>Note: The Elastic APM latency indicator measures all HTTP transactions for the <code>search-analytics-demo</code> service, including health checks and click/cart/checkout endpoints, along with static asset requests, not just the <code>POST /api/search</code> endpoint. For a search-only latency SLO, use a custom Kibana Query Language (KQL) indicator with <code>name: "search" AND attributes.search.query: *</code> on the <code>traces-generic.otel-default</code><code> index</code>. The Elastic APM indicator is still valuable for whole-service health. For full coverage, combine both.</p><p>This measures end-to-end span duration; that is, what the user actually experiences. If your Elasticsearch query takes 50ms but the user waits 300ms because of network overhead or slow application logic, this SLO catches it. Use <code>search.took_ms</code> in the Elastic APM waterfall to diagnose <em>where</em> the latency lives when the SLO starts burning.</p><h4>2. Search availability SLO: 99.9% success rate</h4><p>Indicator type: Elastic APM availability <strong>t</strong>arget: 99.9% of searches succeed.</p><p>The Elastic APM availability indicator calculates the percentage of successful transactions for your service. When the Elasticsearch client throws an exception or the search endpoint returns a 5xx, the span's status records an error and this SLO counts it.</p><p>Note: Like the latency SLO, the Elastic APM availability indicator covers all HTTP transactions on <code>search-analytics-demo</code>, not just <code>POST /api/search</code>. Click/cart/checkout errors will consume this budget. For a search-only availability SLO, use a custom KQL indicator with <code>name: "search" AND attributes.search.query: *</code> for good events and <code>name: "search"</code> as the total query.</p><p>An 0.1% error budget on 100,000 daily searches means that you can tolerate 100 errors per day. That's tight, but search errors are hard failures and the user gets nothing. Availability SLOs should be stricter than latency SLOs.</p><h4>3. Search quality SLO: tracking zero-results rate</h4><p>Indicator type: Custom KQL target: 85% of searches return at least one result (zero-results rate &lt; 15%); index: <code>traces-generic.otel-default</code>; <strong>g</strong>ood query: <code>name: "search" AND attributes.search.result_count &gt; 0</code>; total query:<code>name: "search" AND attributes.search.query: *</code>.</p><p></p><p>Note on KQL versus ES|QL: The SLO framework uses KQL for its indicator filters rather than ES|QL. KQL uses <code>field: value</code>syntax and is the same language you see in the Kibana search bar. The ES|QL queries throughout this series are for ad hoc analysis and dashboards; KQL here is the SLO indicator's document filter. Both query the same <code>traces-generic.otel-default</code> index.</p><p>This is the SLO that surprises most teams. A search that returns an empty result set isn't an error; HTTP status is 200 and the span status is OK. Plus, no exception was thrown. But from the user's perspective, it failed. They asked for something and got nothing.</p><p>The quality SLO uses the custom KQL indicator type because it relies on our custom <code>search.result_count</code> attribute, which the built-in Elastic APM indicators don't know about. But the SLO framework handles everything else, including budget tracking and burn rate calculation, along with alerting.</p><p>A sudden spike in zero-results rate, such as from 12% to 40% over an hour, is almost always an infrastructure event, like a failed index deployment or a mapping change that broke queries. It could also be a synonym list misconfiguration. That's an operational problem, not a relevance problem.</p><h3>Reading your search health in the SLO overview</h3><p>Once the latency, availability and quality SLOs are created, the SLO overview page shows your search health at a glance:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9e9ed72bf5840749/6a8d706fd05d4c5a8957f946/blog6-slo-overview-drill-in.gif" alt="Search Latency SLO detail showing 50% observed against 99% objective, burn rate windows, historical SLI, and error budget" /><p></p><p>Each SLO shows the current value, the target, the remaining error budget, and the burn rate. Green means <em>healthy</em>:  Search availability is at 100%, and Search quality is just above its 85% target. Red means <em>violated</em>: Search latency is at 50% against a 99% objective, with the burn rate breached at 200x the sustainable rate. When a budget bar starts shrinking faster than expected, you know something changed, even before users complain.</p><p>Clicking into the SLO detail shows burn rate across multiple time windows (1h, 6h, 24h, 72h) and the historical SLI trend. It also shows remaining error budget. For the latency SLO, the Elastic APM latency indicator tracks the percentage of transactions below your 250ms threshold. For the quality SLO, the custom KQL indicator uses <code>traces-generic.otel-default</code> with the good query filtering for <code>attributes.search.result_count &gt; 0</code>. This is where the custom <code>search.*</code> attributes from <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a> pay off, since they're the foundation of meaningful SLOs.</p><h2>Burn rate alerting for search SLOs</h2><p>When you create an SLO through the Kibana UI, a default burn rate alert rule is automatically created. This is where the real operational value lives.</p><p><a href="https://www.elastic.co/guide/en/observability/current/slo-burn-rate-alert.html">Burn rate alerts</a> improve on threshold alerts ("error rate &gt; 1%"), which are noisy and miss slow degradation. : Burn rate alerts measure how fast you're consuming your error budget relative to the SLO window.</p><p>A burn rate of 1.0 means that you're spending budget at exactly the sustainable rate, but a burn rate of 10.0 means that you're burning 10x too fast and you'll exhaust the budget in 1/10th of the window.</p><p>The default burn rate rule uses a multi-window approach, with four severity levels:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5794fbcc015d4076/6a8d709811115542160cfc88/blog6-burn-rate-alert-config.gif" alt="Selecting alert rule types in Elastic Observability including anomaly detection, APM anomaly, custom threshold and SLOs" /><p></p><p><strong>Severity</strong></p><p><strong>Burn rate</strong></p><p><strong>Long window</strong></p><p><strong>Short window</strong></p><p><strong>What it means</strong></p><p>Critical (page)</p><p>&gt; 14.4x</p><p>1 hour</p><p>5 minutes</p><p>Exhausts budget in ~50 hours</p><p>High (ticket)</p><p>&gt; 6.0x</p><p>6 hours</p><p>30 minutes</p><p>Exhausts budget in ~5 days</p><p>Medium (review)</p><p>&gt; 3.0x</p><p>24 hours</p><p>120 minutes</p><p>Exhausts budget in ~10 days</p><p>Low (awareness)</p><p>&gt; 1.0x</p><p>72 hours</p><p>360 minutes</p><p>Trending toward exhaustion</p><p>The short window prevents alerting on brief spikes that self-resolve, and the long window catches sustained degradation. Together, they balance responsiveness with alert fatigue.</p><h3>Routing search alerts to PagerDuty, Slack and Jira</h3><p>Alerts are only useful if they reach the right people in the right tools. Elastic's <a href="https://www.elastic.co/guide/en/kibana/current/alerting-getting-started.html">alerting framework</a> supports a wide range of <a href="https://www.elastic.co/guide/en/kibana/current/action-types.html">connectors</a> out of the box, including:</p><ul><li><p><strong>Incident management:</strong> PagerDuty, Opsgenie, xMatters for on-call routing.</p></li><li><p><strong>Chat:</strong> Slack, Microsoft Teams for team notifications.</p></li><li><p><strong>Case management:</strong> <a href="https://www.elastic.co/guide/en/kibana/current/jira-action-type.html">Jira</a>, <a href="https://www.elastic.co/guide/en/kibana/current/servicenow-action-type.html">ServiceNow</a> for automatic ticket creation when SLOs breach.</p></li><li><p><strong>Custom:</strong> Webhooks for integrating with any system via HTTP.</p></li></ul><p>You can also use Elastic's built-in <a href="https://www.elastic.co/guide/en/kibana/current/cases.html">cases</a> to track incidents directly within Kibana, linking alerts and traces in one place, along with investigation notes, with push to Jira or ServiceNow when escalation is needed.</p><p>A typical routing setup:</p><p></p><p><strong>Alert</strong></p><p><strong>Severity</strong></p><p><strong>Channel</strong></p><p>Latency SLO burn rate &gt; 14.4</p><p>Page</p><p>PagerDuty</p><p>Availability SLO burn rate &gt; 14.4</p><p>Page</p><p>PagerDuty + Slack</p><p>Quality SLO burn rate &gt; 6</p><p>Ticket</p><p>Jira (auto-create) + Slack</p><p>CTR anomaly (machine learning [ML] job)</p><p>Notification</p><p>Slack (search team)</p><h2>Anomaly detection for search quality</h2><p>Some search degradations are gradual shifts that slip past threshold-based alerts, rather than sudden spikes. A relevance regression after a model update might reduce CTR by 15% over a week, and latency might creep up by 5ms per day as the index grows. These are real problems, but they don't trigger burn rate alerts until it's too late.</p><p>Elastic's <a href="https://www.elastic.co/guide/en/machine-learning/current/ml-ad-overview.html">anomaly detection</a> is built for exactly this. It learns normal patterns in your search metrics and flags deviations automatically, and you don’t have to configure any thresholds. </p><h3>Detecting search quality degradation with ML anomaly detection</h3><ul><li><strong>Latency anomalies:</strong></li></ul><p> <a href="https://www.elastic.co/docs/reference/machine-learning/ootb-ml-jobs-apm">Elastic APM anomaly detection</a> can be enabled directly from the Elastic APM UI for your search service. It learns the typical latency distribution, including daily and weekly patterns, and alerts when behavior deviates. A gradual 5ms/day creep will eventually register as anomalous before it hits your SLO threshold.</p><ul><li><strong>CTR drops:</strong></li></ul><p>A relevance regression is invisible to traditional monitoring; latency is fine and errors are zero, plus the result counts are normal, but the ranking changed and users aren't clicking. Anomaly detection on click volume per query is a practical proxy: When a query that normally receives 20 first-click events per hour drops to 5, something likely changed.</p><p>To set this up: In <strong>Kibana</strong> → <strong>Machine Learning</strong> → <strong>Anomaly Detection</strong>, create a new job. Use the <strong>Multi-metric</strong> wizard, and select <code>traces-generic.otel-default</code> as the index. Configure a <code>count</code> detector on <code>attributes.search.first_click</code> split by <code>attributes.search.query</code>. This creates a per-query click-volume baseline and alerts when individual query engagement drops outside the expected range.</p><p>Note: This job detects click-volume anomalies per query, not CTR (which requires dividing clicks by searches). Click volume is a useful proxy (a CTR regression usually manifests as a drop in absolute click count), but be aware that a traffic surge with flat click volume would show as a CTR drop without triggering this alert. For true CTR anomaly detection, use a scheduled ES|QL transform to materialize hourly CTR values and run anomaly detection on the computed ratio.</p><p>Route the resulting ML alert rule to your Slack search channel.</p><ul><li><strong>Throughput shifts:</strong></li></ul><p>A sudden drop or unexpected surge in search volume can indicate upstream problems (like load balancer changes or traffic shifts) or downstream issues (such as search becoming unresponsive or users retrying).</p><p>Configure <a href="https://www.elastic.co/guide/en/machine-learning/current/ml-configuring-alerts.html">ML anomaly alert rules</a> to route these to your notification channels. These complement your SLO burn rate alerts; burn rates catch budget consumption, and anomaly detection catches pattern changes.</p><h2>Building a search monitoring dashboard with ES|QL</h2><p>SLOs tell you <em>whether</em> search is healthy. When they indicate a problem, you need a dashboard that tells you <em>why</em>.</p><p>The search team and the on-call team need different views of the same data. A search engineer wants query-level detail, such as which queries have low CTR and which ones return nothing. They’re also interested in where to invest in relevance. But an on-call SRE wants the operational picture, including whether search is fast and whether it’s up. It also wants to know whether search is degrading, and if so, since when.</p><h3>Search monitoring panels for the on-call dashboard</h3><p>Build it in <a href="https://www.elastic.co/guide/en/kibana/current/dashboard.html">Kibana dashboards</a> using <a href="https://www.elastic.co/guide/en/kibana/current/lens.html">Kibana Lens</a> panels. Lens supports <a href="https://www.elastic.co/docs/explore-analyze/visualize/esorql">ES|QL as a data source</a>, so the queries from <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blogs 2–4</a> can power dashboard panels directly. The key panels include:</p><ul><li><p><strong>Search throughput over time:</strong>  A sudden drop is often the first sign of a problem.</p></li><li><p><strong>Latency percentiles (p50, p95, p99) over time:</strong> When they diverge (p50 flat, p99 spikes), you have a subset of slow queries.</p></li><li><p><strong>Error rate over time:</strong> Spikes here mean hard failures.</p></li><li><p><strong>Zero-results rate over time:</strong> A step change upward, especially correlated with a deployment, means something changed in the index or query pipeline.</p></li></ul><p>The ES|QL for each panel follows the patterns from earlier blogs. For example, a latency percentile panel:</p>FROM traces-generic.otel-default
| WHERE name == "search"
AND attributes.search.query IS NOT NULL
| EVAL bucket = DATE_TRUNC(5 minutes, @timestamp)
| STATS
    p50 = PERCENTILE(attributes.search.took_ms, 50),
    p95 = PERCENTILE(attributes.search.took_ms, 95),
    p99 = PERCENTILE(attributes.search.took_ms, 99)
BY bucket
| SORT bucket<p>This includes three lines on one chart. When they diverge, such as when p50 stays flat but p99 spikes, you likely have a subset of queries that are slow while the majority are fine. That's a different diagnosis than all queries slowing down (cluster-level pressure).</p><h3>Drill-down panels: slowest queries and top zero-result queries</h3><p>For investigation, add a few detail panels, such as:</p><p><strong>Slowest queries:</strong> A table showing the queries with the highest p95 latency and their search volume. During an incident, this narrows the problem from "search is slow" to "these specific queries are slow."</p><ul><li><p><strong>Top zero-result queries:</strong> A table showing which queries most frequently return nothing. When zero-results rate spikes, this panel immediately shows which queries are responsible.</p></li></ul><p>These drill-down panels use the same ES|QL patterns as Blogs <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">2</a> and <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">3</a>, just surfaced on a persistent dashboard instead of run ad hoc.</p><h2>Search incident response using OpenTelemetry traces</h2><p>As an example, an alert fires, noting that search latency has spiked. What happens now?</p><p>The trace data from Blog 2's instrumentation gives you a structured path from symptom to root cause.</p><h3>Step 1: Assess scope</h3><p>Start at the on-call dashboard, and get answers to the basics:</p><ul><li><p><em>When did it start?</em> Narrow the time range to the degradation window.</p></li><li><p><em>How bad is it?</em> Is p50 affected (all queries slow) or just p99 (a subset)?</p></li><li><p><em>Is it just search?</em> Check the Elastic APM service map to determine whether the Elasticsearch dependency is also degraded.</p></li></ul><h3>Step 2: Find the problem queries</h3><p>If the problem is a subset of queries (p99 spike but p50 is fine), use the slowest queries panel or run:</p>FROM traces-generic.otel-default
| WHERE name == "search"
AND attributes.search.query IS NOT NULL
  AND attributes.search.took_ms &gt; 100
| STATS
    count = COUNT(*),
    avg_ms = ROUND(AVG(attributes.search.took_ms), 0),
    max_ms = MAX(attributes.search.took_ms)
BY attributes.search.query
| SORT count DESC
| LIMIT 10<p>Adjust the <code>100</code> ms threshold to match your environment's normal range. It can be lower for a fast cluster or higher if your data volume makes 100ms typical. This narrows the problem from "search is slow" to "these specific queries are slow." That's the difference between restarting the cluster and investigating a specific query pattern.</p><h3>Step 3: Drill into the trace waterfall</h3><p>Pick a slow query, and open it in the Elastic APM trace view. The waterfall shows exactly where time was spent. (Refer back to the trace waterfall GIF above to see a real example of a <code>POST /api/search</code>  trace broken down into its component spans.)</p><p>The overhead gap between <code>search.took_ms</code> (Elasticsearch time) and span duration (end-to-end time) is your diagnostic tool:</p><p></p><p><strong>Scenario</strong></p><p><strong>search.took_ms</strong></p><p><strong>Span duration</strong></p><p><strong>Diagnosis</strong></p><p>Elasticsearch slow</p><p>400ms</p><p>430ms</p><p>Elasticsearch problem: Check slow log, cluster metrics.</p><p>App slow</p><p>50ms</p><p>350ms</p><p>Application / network overhead: Check serialization, network.</p><p>Both slow</p><p>400ms</p><p>700ms</p><p>Multiple issues: Investigate both.</p><p></p><p>If the problem is in Elasticsearch, drill into the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-profile.html">Search Profile API</a> or <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/monitor-elasticsearch-cluster.html">cluster monitoring</a>. If it's application overhead, look at the spans around the search span in the waterfall.</p><h3>Step 4: Correlate with events</h3><p>Check whether the degradation correlates with:</p><ul><li><p><strong>Deployments:</strong> Did someone deploy a new version of the search service or push a new index?</p></li><li><p><strong>Cluster events:</strong> Is Elasticsearch under memory pressure, or are there long garbage collection pauses? Or maybe the disk is I/O saturated?</p></li><li><p><strong>Network:</strong> Is latency between the search service and Elasticsearch elevated?</p></li></ul><p>Elastic Observability's unified platform makes this correlation straightforward because traces, logs, metrics, and infrastructure data all live in the same Kibana instance. You're adding filters in the same interface, rather than switching between tools.</p><h2>Going further: Infrastructure metrics and cost attribution</h2><p>This post focuses on what you get from trace data; that is, the spans your search API already emits. But OTel and Elastic Observability support a wider instrumentation picture that becomes valuable as your search infrastructure matures.</p><ul><li><p><strong>Infrastructure metrics.</strong> Adding host and container metrics (like CPU, memory, disk I/O, and network) alongside your traces lets you correlate search performance with infrastructure use. When p99 latency spikes, you can immediately see whether the Elasticsearch nodes are under memory pressure and whether garbage collection  pauses are increasing. You can also check whether disk I/O is saturated, and you can do all this in the same Kibana interface. The <a href="https://www.elastic.co/guide/en/fleet/current/elastic-agent-installation.html">Elastic Agent</a> collects these automatically for your infrastructure, and the <a href="https://www.elastic.co/guide/en/observability/current/analyze-hosts.html">infrastructure monitoring UI</a> surfaces them alongside your Elastic APM data.</p></li></ul><ul><li><p><strong>Total cost attribution (TCA).</strong> With infrastructure metrics flowing alongside traces, you can start attributing infrastructure costs to specific services and operations. How much compute does your search service consume? How does that correlate with query volume? If a new ranking model doubles CPU usage per query, you can see the cost impact directly. This is particularly valuable for teams running search on cloud infrastructure where costs scale with resource consumption; understanding the cost per search helps justify infrastructure investment and identify optimization opportunities.</p></li></ul><ul><li><p><strong>Logs correlation.</strong> OTel auto-instrumentation injects trace context (such as trace ID and span ID) into your application logs. This means that when you're investigating a slow search in the trace waterfall, you can click through to the exact log lines from that request, including Elasticsearch slow log entries and application debug output. It also includes error details that don't fit in span attributes. The <a href="https://www.elastic.co/guide/en/observability/current/application-logs.html">logs correlation</a> feature automatically ties them together.</p></li></ul><p>These are natural next steps once you have traces working. Each one extends the same unified platform, without new tools or separate pipelines.</p><h2>How search analytics and search monitoring share one data pipeline</h2><p>Here's how it all fits together:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt540076694fa1cbcc/6a8d717236492a486785f109/image3.png" alt="Search monitoring workflow: out-of-the-box APM features, SLO and alert definition, dashboard building, incident response" /><p></p><p>The data flows from the instrumentation you built in Blog 2. This one investment supports two audiences: The search team gets product analytics (<a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blogs 2–5</a>), and the SRE team gets operational monitoring (this post). Neither team needs separate data pipelines.</p><h2>Getting started with search monitoring in Elastic</h2><p>This is the last post in the series, and it brings us full circle. <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry">Blog 1</a> describes the vision: Instrument search once with OTel, and send spans to Elastic. Then use ES|QL to answer any question about search behavior. <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blogs 2–4</a> build the instrumentation and analytics, and <a href="https://www.elastic.co/search-labs/blog/search-analytics-relevance-click-streams">Blog 5</a> shows how to feed that data back into relevance improvements. This post shows how the same data powers operational monitoring, including SLOs, alerting, anomaly detection, and incident response.</p><p>The key takeaway for search engineers is that the instrumentation you built for analytics already generates the signals. The SLOs and alerts are built-in capabilities of Elastic Observability, as are the dashboards. You're closer to production-grade search monitoring than you might think. Observability isn't a separate discipline you need to learn from scratch. </p><p>If you've been following along and built the instrumentation from Blogs 2–4, start here:</p><ol><li><p><strong>Open Elastic APM:</strong> Look at your search service, and explore a trace waterfall. You can also check the errors tab.</p></li><li><p><strong>Create three SLOs:</strong> Latency (Elastic APM latency), availability (Elastic APM availability), and quality (custom KQL for zero-results).</p></li><li><p><strong>Enable anomaly detection:</strong> One click in the Elastic APM UI for latency anomalies.</p></li><li><p><strong>Build the on-call dashboard:</strong> Four Lens panels with the queries from this post.</p></li></ol><p>By the end of an afternoon of work, your search service can have the same observability coverage as any other critical production system.</p><h2>Get started</h2><h3>Working code</h3><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">Reference project:</a> Working code for the entire blog series; clone, configure, and run.</p></li></ul><h3>Elastic APM and traces</h3><ul><li><p><a href="https://www.elastic.co/guide/en/apm/guide/current/apm-overview.html">Elastic APM Overview:</a> Elastic APM concepts and trace analysis.</p></li><li><p><a href="https://www.elastic.co/guide/en/observability/current/apm-ui.html">Elastic APM UI:</a> Service overview, transactions, dependencies, errors.</p></li><li><p><a href="https://www.elastic.co/docs/solutions/observability/apm/service-map">Service Maps:</a> Dependency visualization and health.</p></li><li><p><a href="https://www.elastic.co/guide/en/observability/current/open-telemetry.html">OpenTelemetry Integration:</a> OTel ingestion in Elastic.</p></li></ul><h3>SLOs and alerting</h3><ul><li><p><a href="https://www.elastic.co/guide/en/observability/current/slo.html">SLOs in Elastic Observability:</a> Creating and managing SLOs.</p></li><li><p><a href="https://www.elastic.co/guide/en/observability/current/slo-create.html">Create an SLO:</a> SLI types (like Elastic APM latency, Elastic APM availability, or custom KQL), time windows, budgeting.</p></li><li><p><a href="https://www.elastic.co/guide/en/observability/current/slo-burn-rate-alert.html">SLO Burn Rate Alerts:</a> Multi-window burn rate alerting.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/action-types.html">Alert Connectors:</a> Slack, PagerDuty, webhook, email integrations.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/alerting-getting-started.html">Alerting Framework:</a> Rule types and configuration.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/jira-action-type.html">Jira Connector:</a> Automatic ticket creation from alerts.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/servicenow-action-type.html">ServiceNow Connector:</a> ITSM integration.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/cases.html">Elastic Cases:</a> Built-in incident tracking with external push.</p></li></ul><h3>Anomaly detection</h3><ul><li><p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-ad-overview.html">Anomaly Detection Overview:</a> Unsupervised time series anomaly detection.</p></li><li><p><a href="https://www.elastic.co/docs/reference/machine-learning/ootb-ml-jobs-apm">Elastic APM Anomaly Detection:</a> Enable ML for latency, throughput, error rate.</p></li><li><p><a href="https://www.elastic.co/guide/en/machine-learning/current/ml-configuring-alerts.html">ML Anomaly Alert Rules:</a> Alerting on detected anomalies.</p></li></ul><h3>Dashboards and ES|QL</h3><ul><li><p><a href="https://www.elastic.co/guide/en/kibana/current/dashboard.html">Kibana dashboards:</a> Building operational dashboards.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/lens.html">Lens:</a> Visualization editor.</p></li><li><p><a href="https://www.elastic.co/docs/explore-analyze/visualize/esorql">ES|QL in Lens:</a> ES|QL-powered dashboard panels.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL Overview:</a> Language reference.</p></li></ul><h3>Elasticsearch operations</h3><ul><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/index-modules-slowlog.html">Slow Log Configuration:</a> Threshold-based query slow logging.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-profile.html">Search Profiling:</a> Profile API for query execution analysis.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/monitor-elasticsearch-cluster.html">Monitoring Elasticsearch:</a> Cluster stats, search rate, latency.</p></li></ul><h3>From this series</h3><ul><li><p><a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry">Modern search analytics with OpenTelemetry:</a> The vision.</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Instrument your search API </a>: Search spans and <code>search.*</code> attributes.</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Measuring search quality </a>: CTR, MRR, click distribution.</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/search-conversion-tracking-opentelemetry">From clicks to conversions</a>: Conversion tracking and revenue attribution.</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/search-analytics-relevance-click-streams">Personalizing search from behavior</a>: Judgment lists, rank features, Learning To Rank (LTR).</p></li></ul><p><em>This is the final post in a six-part series on search analytics with OpenTelemetry and Elastic. Start from the beginning: </em><a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry"><em>Modern search analytics with OpenTelemetry,</em></a><em> or to start building, jump to </em><a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql"><em>Instrument your search API.</em></a></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/opentelemetry-search-monitoring-slos</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/opentelemetry-search-monitoring-slos</guid>
    <category><![CDATA[Operations]]></category>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Matthew Adams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb4ac8bfdb49b5f32/6a8d5eb7da6aeaa5fa37aae2/image1.png" length="0" type="image/png"/>
    <pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Skip the mapping explosion: ES|QL queries schemaless JSON keys without dynamic mapping]]></title>
    <description><![CDATA[Flattened fields turn Elasticsearch into a schema-on-read store where you index schemaless data under one mapping, then use ES|QL's FIELD_EXTRACT to pull out any JSON key you need to filter, group or join on, with predicates pushed into the columnar store.]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a> now reads<a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/flattened"> <code>flattened</code> fields</a>. FIELD_EXTRACT pulls any key out of a schemaless JSON object so you can filter, group, sort and join on keys you never mapped. The planner pushes those predicates into the columnar store rather than parsing the whole blob per row, which means dynamic JSON keys from OTel attributes, log labels, user metadata or whatever else you didn't want to map individually are queryable without causing a<a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/mapping-limit"> mapping explosion</a>.</p><h2>Why dynamic mapping breaks down with schemaless data</h2><p>Elasticsearch wants a schema. Every field you index has a mapping that defines its type, its analyzer (for text), and how it is stored. That schema makes storage compression efficient, with fast search and cheap aggregation. It becomes a liability the moment your data stops looking like a database table.</p><p>Consider the shapes that show up in real systems:</p><ul><li><p>Log events where every service adds its own attributes. One emits <code>labels.region</code>, another <code>labels.k8s.pod</code>, and another <code>labels.tenant_id</code>.</p></li><li><p>OpenTelemetry resource attributes, where the set of keys is defined by whatever agent happened to send the span.</p></li><li><p>User-supplied metadata bags, feature flags, or tagging systems, where the keys are open-ended by design.</p></li></ul><p>If you map each key as its own field, the mapping grows unbounded. This is the classic "mapping explosion." Thousands of dynamically created fields inflate the cluster state, slowing down every mapping update and eventually hitting the field limit. Each field also carries overhead in the index. You pay a structural cost for keys you didn’t plan for and may query only once.</p><p>The naive escape hatch is to store the whole object as a string and give up on querying its contents. That trades one problem for another: you keep the data but lose the ability to filter or group by anything inside it.</p><p>The <code>flattened</code> field type is the better path. You can index an entire JSON object under a single mapped field, keeping the keys queryable with almost none of the mapping-explosion cost. </p><p>This post covers how a flattened field stores the data on a disk and how ES|QL reads it back.</p><h2>How flattened fields index dynamic JSON keys under one mapping</h2><p>Map one field as flattened:</p>PUT logs
{
 "mappings": {
"properties": {
"labels": { "type": "flattened" }
   }
 }
}<p>Then write arbitrary nested JSON into it:</p>POST logs/_doc
{
 "labels": {
   "region": "us-east-1",
   "k8s": { "pod": "web-7f9", "node": "ip-10-0-0-3" },
   "retries": 4
 }
}<p>There is exactly one field in the mapping, <code>labels</code>, no matter how many keys appear across your documents. The cluster state doesn’t grow when a new key shows up. The subkeys remain individually searchable. You can reference <code>labels.region</code> or <code>labels.k8s.pod</code> in queries, even though neither was ever declared.</p><p>The catch and central tradeoff is that every leaf value is a keyword. The number 4 above is indexed as the string <code>"4"</code>. There is no numeric typing, no date parsing, and no range math on dynamic keys. Flattened fields exchange per-field richness for schema flexibility. That fact explains almost every design decision that follows.</p><h2>How Elasticsearch stores flattened field JSON keys on disk</h2><p>There are two types of queries on flattened fields: an unkeyed query on the root flattened field, and a keyed query on a specific subfield. Following the previous example, a query of the form <code>labels: "us-east-1"</code> matches a value under <em>any</em> key, while the query<code>labels.region: "us-east-1"</code> matches a value under the <em>specific</em> <code>region</code>key.</p><p>To support these two distinct query formats, the flattened mapper writes each leaf value into two distinct Lucene fields.</p><p>Take this document:</p>{ "labels": { "region": "us-east-1", "k8s": { "pod": "web-7f9" } } }<p>The mapper produces:</p><ul><li><p>A root field under labels, holding the bare values:</p></li></ul>us-east-1
web-7f9<ul><li><p>A keyed field under labels._keyed, holding the flattened key concatenated with its value:</p></li></ul>region\0us-east-1
k8s.pod\0web-7f9<p>In the keyed field, nested objects are dot-flattened into a single key (k8s.pod), and the key is joined to its value with a reserved NULL byte (\0) as the separator. Keys that contain a NULL byte are rejected at parse time, so the separator is always unambiguous. To find the value, you split on the first NULL.</p><p>These two fields make both query shapes work:</p><ul><li><p><code>labels: "us-east-1"</code> matches a value under <em>any</em> key, so it searches the root field.</p></li><li><p><code>labels.region: "us-east-1"</code> matches a value under a <em>specific</em> key. It rewrites the query to the term <code>region\0us-east-1</code> and searches the keyed field.</p></li></ul><h2>Query restrictions on flattened field subkeys</h2><p>Because every key's terms live in one sorted list, the keyed field cannot answer every query shape a plain keyword field can. Three restrictions follow:</p><ol><li><p>No fuzzy, regexp, or wildcard on a specific subkey. Nothing about the layout makes them impossible,  but the pattern would have to be combined with the key prefix so it can’t walk past the key boundary. The mapper doesn’t do that at this writing.</p></li><li><p>Any range query on a subkey needs at least one bound. Elasticsearch already has a query for "this field has some value here, whatever it is": the <code>exists</code> query. On a flattened subkey it runs as a prefix query on key\0, which sweeps every term belonging to that key. A range with neither bound would sweep exactly the same terms. Rather than support two spellings of one scan, the mapper rejects the boundless range and asks for the <code>exists</code> query.</p></li><li><p>A range query on a subkey needs the field to be indexed. A flattened field can be mapped with <code>index: false</code>, which skips the inverted index and keeps only doc values, the columnar per-document storage covered in the next section. Exact-match queries survive that. With no terms to look up, Elasticsearch scans the doc values column instead, which is slower but gives the same answer. Range queries have no equivalent fallback, so a range on a subkey of an unindexed flattened field throws an Exception.</p></li></ol><p>All three are limits on the Lucene query the mapper is willing to build, and where you notice them depends on how you query.</p><p>On the search API, where you name <code>labels.region</code> directly, they come back as errors.</p><p>In ES|QL you will not see them as errors at all. There, the same limits decide only whether a predicate is pushed into Lucene or runs in the compute engine on the extracted column. </p><p>This is pushed to a term query on the keyed field:</p><p>This is not, so the filter runs per row on the extracted keyword:</p><p>Same answer, more work. That distinction is the subject of the second half of this post.</p><h2>Under the hood: how range queries stay inside key boundaries</h2><p>This part is internal. You don’t need it to use the field, but it explains where the bounds rule comes from.</p><p>For a handful of documents in one segment, the shared term list looks like this:</p>k8s.pod\0web-7f9
region\0us-east-1
region\0us-west-2
tenant_id\0acme<p>Each key owns a contiguous slice of that list. For example, all values for key "region" are clustered together in an ordered sublist. A range with both bounds set encodes each bound the same way a term is encoded, so a lower bound of "us-east" on the region key becomes region\0us-east and an upper bound of "us-west" becomes region\0us-west. Both endpoints already carry the key prefix, so the scan can’t leave the region slice. Nothing special is needed.</p><p>The half-open case is problematic. Handing Lucene a lower bound of <code>region\0us-east</code> with no upper bound would scan to the end of the term list, straight through <code>tenant_id\0acme</code> and every other key that sorts after region. So the mapper substitutes a sentinel for the missing side:</p><ul><li><p>A missing lower bound becomes key\0, inclusive. That’s the encoding of the empty value, and it’s the first term in the key's slice.</p></li><li><p>A missing upper bound becomes key\1, exclusive. Byte 0x01 is the next byte after the 0x00 separator, so it sorts after every key\0value term and before the first term of any other key.</p></li></ul><p>A one-sided range is therefore boxed into [key\0, key\1), which makes it exactly as safe as a closed one.</p><p>The upper sentinel also covers the case where one key is a prefix of another. If an index holds both region and regionx, then region\1 still sorts below regionx\0eu-west-1, because 0x01 is smaller than the x that follows the shared region prefix. A range on region cannot leak into regionx.</p><h3>How flattened fields use the inverted index and doc values</h3><p>Each leaf value can be written into two Lucene structures. Both are enabled by default, but can be disabled by the mapping configuration.</p><ul><li><p>The inverted index (when the field is indexed). 
Two untokenized keyword terms are indexed per value, one on the root path and one on the keyed path. This powers term, prefix, and range searches.</p></li><li><p>Doc values (when <code>doc_values</code> is enabled). 
A columnar, document-ordered structure. This powers sorting, aggregations, and ES|QL reads.</p></li></ul><p>The inverted index answers the question: "Which documents contain this term?" </p><p>Doc values answer another question: "For this document, what are the values?" Doc values are laid out column by column so a scan touches only the bytes it needs. A flattened field uses both inverted indexes and doc values, so it can serve search and analytics from the same field.</p><p>The nature of the index means that the root field is only present in some cases. When the inverted index is disabled by the mapping, any value search requires a linear scan of the doc values for the searched value. In this case, when performing a search on the root field, there isn’t much additional overhead compared to just scanning the keyed field and ignoring the key markers. So the flattened mapper skips writing the root field, relying on the keyed field for both root and keyed queries.</p><h3>Why flattened fields switched from dictionary to binary doc values</h3><p>Historically, flattened field doc values used Lucene's <code>SortedSetDocValues</code>, which is a dictionary-compressed format. This means that every <code>key\0value</code> value indexed across all documents per segment is stored in one big sorted, deduplicated set of values. Each document tracks a list of ordinals into that value set.</p><p>This dictionary approach provides great compression for low-cardinality fields that tend to have repeated values. It is byte-efficient to store a value only once, and then refer to it by a single integer value. However, there is overhead associated with building and maintaining that dictionary of values, and that overhead is wasted effort when operating on high-cardinality fields that don’t repeat values.</p><p>Because flattened fields are a catch-all type, their cardinality tends to be very high. So while the dictionary approach works, it’s not the most compact or scan-friendly layout for this data, especially in time-series indices where flattened bags are common and storage pressure is real.</p><p>Recent versions switched the storage to use Lucene’s <code>BinaryDocValues</code>. This format just stores a literal binary blob for each document, which is compressed using Zstandard by our doc values codec when written to disk.</p><p>This new binary format provides an additional benefit: it allows us to maintain original array ordering without any overhead. Flattened fields support the mapping parameter <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/flattened#flattened-params">preserve_leaf_arrays</a>, which affects how multivalued fields are returned when using <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-source-field#synthetic-source">synthetic source</a>. When configured to <code>preserve_leaf_arrays: exact</code>, returned values preserve the order, duplicates, and nulls from the original source value.</p><p>The dictionary encoding inherent to sorted-set doc values means the returned values are sorted, deduplicated, and de-nulled. To implement <code>preserve_leaf_arrays</code>, flattened fields have traditionally used an additional sidecar field, tracking the required metadata to reconstruct the original source value. However, the nature of binary doc values means that this sidecar field is no longer needed. The values are just stored and returned as indexed.</p><h3>Limitations of schema on read with flattened fields</h3><p>The limitations below follow from the keyword-only rule and the shared keyed field:</p><ul><li><p>No numeric, date, or Boolean typing on dynamic keys. <code>100</code> and <code>"100"</code> are the same term.</p></li><li><p>No fuzzy, regexp, or wildcard on a specific subkey.</p></li><li><p>No multi-fields (<code>fields</code>) or <code>copy_to</code> on the flattened field.</p></li><li><p>A <code>depth_limit</code> (default 20) on how deeply nested the object can be.</p></li></ul><p>If you need real typing for a <em>known</em> key, flattened now supports explicitly <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/flattened#flattened-properties">mapped subfields</a>: you declare individual keys with real types under a <code>properties</code> block, and those keys are indexed by their own typed mapper instead of the keyed field. You get numeric ranges on <code>labels.status_code</code>, while everything else in <code>labels</code> stays dynamic and keyword-only. This is the escape hatch for the handful of keys you actually know about.</p><h2>Schema on read: querying flattened field JSON keys in ES|QL</h2><p>Search has supported flattened fields for years. ES|QL, the newer, piped query language built on a columnar compute engine, now reads flattened fields, too. Support began in the Technical Preview of Elasticsearch 9.5.0. It has two distinct pieces.</p><h3>What ES|QL returns when you select a flattened field</h3><p>ES|QL uses a dedicated data type for flattened fields, rather than folding it into <code>keyword</code>. When you select the root, you get the whole object back as a JSON string:</p>labels:flattened
{"k8s.pod":"web-7f9","region":"us-east-1"}<p>Keys come back sorted. You can carry this value through a query, count it, group by it, and run the multi-value and comparison functions on it. But an opaque JSON blob is not usually what you want to filter or aggregate on. For that, you need to reach inside it and process its contents.</p><h3>How FIELD_EXTRACT reads JSON keys from flattened fields</h3><p>There is no dotted-path syntax for dynamic keys in ES|QL. You cannot write <code>labels.region</code> for an unmapped key, because to the engine the flattened root is a single leaf value, not a set of columns. Instead you use a function:</p><p>The absence of a dotted-path syntax is a tentative limitation. The keyed field already addresses individual subkeys, so a more natural syntax for reaching into a flattened root is something we are planning to support.</p><p>FIELD_EXTRACT(field, path) takes a flattened field and a key, and returns a keyword. The rules are easier to see against a document. Take this one:</p>POST logs/_doc
{
 "labels": {
   "region": "us-east-1",
   "k8s": { "pod": "web-7f9", "node": "ip-10-0-0-3" },
   "tags": ["prod", "canary"],
   "retries": 4
 }
}<p>And this query:</p><p>The result is:</p>region     | pod     | k8s  | tags            | retries | namespace
us-east-1  | web-7f9 | null | [prod, canary]  | 4       | null<p>Four things to take from this:</p><ul><li><p>The dot is part of the key, not a navigation operator. The mapper already collapsed the nested object into the flat key k8s.pod, so "k8s.pod is a direct lookup, not a walk from k8s to pod.</p></li><li><p>Matching is exact. "k8s" returns null because there is no leaf stored at k8s, only at k8s.pod and k8s.node. For the same reason, "host" will not find "host.name", and matching is case-sensitive, so "Region" will not find "region".</p></li><li><p>Arrays come back multi-valued. "tags" yields a multi-valued keyword you can <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/mv_expand">MV_EXPAND</a>, count, or filter on. A missing key yields null.</p></li><li><p>Everything is a keyword. "retries" comes back as the string "4", and a Boolean leaf comes back as "true" or "false".</p></li></ul><p>JSONPath syntax is rejected outright, at parse time rather than per row. Both FIELD_EXTRACT(labels, "['k8s.pod']") and FIELD_EXTRACT(labels, "tags[0]") fail with <em>field_extract path must be a literal flattened sub-field name</em>.</p><p>Once extracted, the value is an ordinary <code>keyword</code> column. You can filter on it, group by it, sort by it, or use it as the join key in a <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join"><code>LOOKUP JOIN</code></a>:</p><h3>How ES|QL pushes flattened field predicates into the columnar store</h3><p>The obvious way to implement FIELD_EXTRACT would be to read the whole flattened root out of storage, render it as a JSON string, hand it to the compute engine, and parse it once per row to pull out a single key. But that means reading every key to use once and paying for a JSON parse on every document. So this obvious implementation is not performant.</p><p>ES|QL avoids this whenever possible. Between storage and the compute engine sits the <em>block loader</em>, the step that turns stored data into the columnar blocks the engine operates on. FIELD_EXTRACT hooks into that step instead of running after it.</p><p>When ES|QL loads a column, the flattened field type inspects the request. If the request is an extraction of a single constant key and the field has doc values, it routes straight to the keyed doc-values loader, which reads the key\0value entries for only that key out of the columnar structure. The column that arrives at the compute engine already contains only that key's values. The JSON string is never built and never parsed, and the other keys in the object are never read.</p><p>When extraction can’t be fused into the block loader, for example, because the key is computed per row or the root is the output of another function such as CASE, ES|QL falls back to the parse-per-row path. The results are identical either way. Only the cost changes.</p><p>The comparison can push down further. A predicate like <code>FIELD_EXTRACT(labels, "region") == "us-east-1"</code> can be pushed to Lucene as a term query against the synthetic keyed field, the same <code>region\0us-east-1</code> term the search path uses. So a filter on an extracted subkey can be answered by the inverted index (if available), and the projection can be answered by doc values, exactly like a first-class field, even though the key was never in the mapping.</p><p>Ordering comparisons push down, too. The four, single-sided comparators (&gt;, &gt;=, &lt;, &lt;=) and closed BETWEEN-style ranges all become a range query on that same synthetic keyed field. This is where the key\0 / key\1 sentinels earn their keep: the single-sided forms are only pushable because the mapper can box an open bound inside the key. The pushed range is treated as a candidate, and the predicate is re-evaluated on the extracted keyword column afterwards, so multi-valued keys don’t slip through.</p><p>The values are keywords, so the ordering is lexicographic, not numeric. FIELD_EXTRACT(labels, "retries") &gt; "10" compares strings, which means "9" is greater than "10". If you need numeric ranges on a key, map it explicitly under properties, or cast the value in ESQL.</p><p>Explicitly mapped subfields behave differently on purpose. Because they carry real types, comparison semantics diverge from the keyword path, so they are loaded and compared through their own typed mapper rather than fused into the keyed loader. And when you select a flattened root that has mapped subfields, ES|QL loads it from <code>_source</code>, so every leaf renders as a string and no keys are dropped silently.</p><h2>When to use flattened fields vs. dynamic mapping in Elasticsearch</h2><p>Use <code>flattened</code> fields when:</p><ul><li><p>The set of keys is open-ended or unknown ahead of time.</p></li><li><p>You would otherwise cause a mapping explosion.</p></li><li><p>Keyword-level filtering and grouping on the values is enough, and you don’t need numeric or date semantics on the dynamic keys.</p></li><li><p>You have a few keys that <em>do</em> need real types. Map those explicitly under <code>properties</code>, and let the rest stay dynamic.</p></li></ul><p>Avoid it, or map fields normally, when the schema is stable and you need full-text analysis, numeric aggregation, or date math across the board.</p><p>Remember that flattened is not a dumping ground for JSON you have given up on. It’s a real columnar-and-inverted store for schemaless data, and with ES|QL support, it’s now a first-class analytical citizen. You can keep the messy, unmapped parts of your data messy, and still filter, group, join, and aggregate across them as needed.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/schema-on-read-esql-json-keys</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/schema-on-read-esql-json-keys</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <category><![CDATA[Lucene]]></category>
    <dc:creator><![CDATA[Jordan Powers,Dima Leontyev]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb3fca0aedcfba9ab/6a8419fae41d7f68b46522b9/unnamed.png" length="0" type="image/png"/>
    <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Ask the source: Scaling code search to a billion lines with Elasticsearch and Elastic Agent Builder]]></title>
    <description><![CDATA[Sourcerer matches Claude Code and Codex on code retrieval quality and searches up to thousands of times faster than grep. Every answer links back to the exact files and lines across repos and versions.]]></description>
    <content:encoded><![CDATA[<p>Code is the source of truth for its own behavior; it’s always authoritative and never outdated. Definitive answers live at a specific commit in a particular repository, but enterprise deployments depend on many versioned projects working together. Understanding how it all works is a major code search effort.</p><p><em>Does App A v1.2.3 support Feature X? Is it compatible with App B v9.8.7 when running on Kubernetes? Will I need more JVM heap space?</em></p><p>These are the kinds of questions that our field teams handle constantly. Answering them is harder than it looks. Documentation offers context, but it's an abstraction that can't anticipate every possible question. When we hit one it doesn't cover, our options are to interrupt an engineer who should be developing code or to hunt through that code ourselves. Often we don't have the time or expertise to navigate that much of it.</p><p>Coding agents do this well for a single repository on your laptop. Wouldn't it be great if we could scale that to our entire code estate? As a field engineer, I wanted that capability to serve my customers: agentic code intelligence across every project, dependency, platform, and version that we support. And I wanted it grounded in linked citations and always available to everyone as a service.</p><p>So I built it with <a href="https://www.elastic.co/elasticsearch">Elasticsearch</a> and <a href="https://www.elastic.co/elasticsearch/agent-builder">Elastic Agent Builder</a> and packaged it into a command line interface (CLI). I released it under an Apache 2.0 license and called it <a href="https://github.com/elastic/sourcerer">Sourcerer</a>. This blog post reports multiple performance benchmarks of Sourcerer as a code research agent and walks through the design and rationale of its implementation.</p><h2>Sourcerer</h2><p><a href="https://github.com/elastic/sourcerer">Sourcerer</a> explores code like a frontier coding agent, searching across many versioned repositories as fast as it would in a single repository, and it generates answers with linked citations that establish trust.</p><p>Sourcerer consists of:</p><ol><li><p>A set of configuration files for tools, skills, and agents in Agent Builder.</p></li><li><p>A set of index templates to store and search code from Git commit snapshots.</p></li><li><p>A CLI to install those assets and index and prune commit snapshots from remote Git repositories.</p></li></ol><p>At Elastic, we're using Sourcerer to support our customers with verifiable information about our software directly from the source. Our internal deployment has indexed over a billion lines of code from our own public and private repositories. It also includes our core dependencies, such as Apache Lucene and OpenJDK, along with our common integrations, like Kubernetes and OpenTelemetry. Our solution architects, customer architects, consulting architects, and support engineers no longer have to hunt for answers in documentation or reach out to our engineers who should be building software rather than supporting it.</p><h2>Code search benchmarks</h2><h3>Agentic code retrieval</h3><p><a href="https://arxiv.org/abs/2606.07297">SWE-Explore</a> is a new benchmark, published on June 5, 2026, by Zhang et al., that evaluates "how well coding agents explore, localize, and rank repository context." It appears to be the only benchmark that specifically tests agentic code retrieval quality. I ran the benchmark with Sourcerer to see how it performs and compares to the other coding agents from the original paper, and again with Claude Code to measure and compare its token usage and task durations with Sourcerer's.</p><h4>Retrieval scores</h4><p>Sourcerer performs as well as frontier coding agents on relevance metrics for code retrieval (see Figure 1). The composite retrieval score is the arithmetic mean of all retrieval metrics weighed by their Pearson correlations (<em>r</em>) as reported in the paper; I did this to rank the agents by a measurement of overall retrieval quality. Sourcerer trailed Claude Code by 0.002 and surpassed Codex by 0.022 on a 0.0–1.0 scale, which should be interpreted as a statistical tie, given that the results vary slightly on each run due to the indeterminism of large language models (LLMs).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt12808f4ed4f16659/6a7deb00cbb9ed3b2b53f944/image2.png" alt="" /><p></p><p>Generally, all the coding agents, including Sourcerer, performed well on precision metrics and suboptimally on recall metrics, although recall metrics had lower Pearson correlations and thus less importance. Table 1 shows Sourcerer's retrieval scores alongside the scores of the other agents tested in the original paper (page 8, table 6). "SignalReg" is the inverse of what the authors called "NoiseReg" (that is, 1 – NoiseReg); I did this to keep that metric consistent with the other metrics whose ranges imply that higher is better.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd9283fe7e6209465/6a7deb0ec2602368e0b38642/image4.png" alt="" /><h4>Token usage and task duration</h4><p>Zhang et al. didn’t publish metrics for token usage or task durations, so I ran the benchmark again to capture those metrics for Claude Code. Given the high token cost of running the benchmark with any agent that uses an LLM, I opted to test only one coding agent, and Claude Code was the one that I expected most people would find useful as a comparison.</p><p>Compared to Claude Code, Sourcerer used ~13.9% more tokens to complete all 848 benchmark tasks. Sourcerer used 156,379,977 input tokens and 1,612,175 output tokens, while Claude Code used 137,629,435 input tokens and 1,137,176 output tokens. Sourcerer took ~8.9% longer to complete all 848 benchmark tasks. Sourcerer took 40,717 seconds, and Claude Code took 37,460 seconds.</p><p>I view these results on token usage and task duration as an acceptable modest tax in exchange for efficiently searching across multiple repositories and versions. That said, there’s room to explore optimizations to Sourcerer's tools, skills, and system prompt, the harness of Agent Builder, or the search engine of Elasticsearch and Lucene.</p><h4>Single-repo vs. multi-repo scope</h4><p>Critically, the SWE-Explore benchmarks only measure the retrieval scores, token usage, and task duration of agents searching within the boundaries of a single commit snapshot of a repository for any given task. This is the typical search space of a development coding agent. Sourcerer's intended scope is much broader, covering many commit snapshots of many repositories. The benchmark on "search speed and scalability," covered next in this report, shows Sourcerer's unique advantage when searching across many repositories.</p><h4>Retrieval benchmark methodology</h4><p><a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-a.-swe-explore-benchmark-configuration">Appendix A</a> explains the configuration of these benchmarks in detail.</p><p>Sourcerer searched all indexed representations of the <a href="https://huggingface.co/datasets/SWE-Explore-Bench/SWE-Explore-Bench">SWE-Explore-Bench dataset</a> (see <a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-a.-swe-explore-benchmark-configuration">Appendix A</a>). Both benchmark runs used the same GPT-5.4 model that was used in the original paper. Sourcerer communicated with GPT-5.4 through <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service (EIS)</a>, while Claude Code communicated with GPT-5.4 through a shim proxy to be compatible with the OpenAI API.</p><p>I made a best effort to ensure that Sourcerer's benchmark task prompt was like-for-like with Claude Code's (see <a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-a.-swe-explore-benchmark-configuration">Appendix A</a>). Both agents received identical instructions for their roles and tasks, along with output formats, and differed only in their brief harness-specific instructions. I instructed Sourcerer not to use its repo discovery skill and instead gave it explicit repo filtering instructions, ensuring that it was on a level playing field with Claude Code, which already receives the resolved directories. An alternative could have been to leave Sourcerer's repo discovery skill active while instructing Claude Code to find the repository in a filesystem that has all the repositories for the benchmark. I left Sourcerer's system prompts and skills, in addition to its tools, unmodified from their defaults, given that we're comparing the two harnesses overall, and much of which in Claude Code is closed source and not visible or controllable anyway.</p><h3>Code search speed and scalability</h3><p>Coding agents tend to use external tools to match substrings or regular expressions as a first line of retrieval. LLMs are trained to use shell commands, like <code>ls</code> and <code>grep</code>, when exploring code on a filesystem. Claude Code's own built-in <a href="https://code.claude.com/docs/en/tools-reference#grep-tool-behavior">Grep</a> tool invokes <a href="https://github.com/BurntSushi/ripgrep"><code>ripgrep</code></a>. This is the behavior I wanted to reproduce in Elasticsearch.</p><p>Elasticsearch has a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field type that can scale regular expression matching to billions of documents. Sourcerer mimics the inputs and outputs of <code>grep</code> in Elasticsearch using <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.grep.yml"><code>sourcerer.code.grep</code></a>, an Elasticsearch Query Language (ES|QL) tool that performs an<a href="https://www.elastic.co/docs/reference/query-languages/sql/sql-like-rlike-operators"><code>RLIKE</code></a> query on a<a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field of an index where each document has the contents of a single line of code. Sourcerer also provides<a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.search.yml"><code>sourcerer.code.search</code></a>, which performs a BM25-ranked <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/search-functions/match"><code>MATCH</code></a> query against an analyzed text field, for relevance-ranked discovery rather than exact substring retrieval.</p><p>I benchmarked the speed of both approaches against the speed of <code>ripgrep</code> and <code>grep</code> on a filesystem using two different corpus sizes and pattern rarities. Corpus sizes included one commit (7,215,509 lines of code) and 52 commits (200,325,684 lines of code) from the<a href="https://github.com/elastic/elasticsearch"> elastic/elasticsearch</a> repository. The single commit covers the release tag for v9.4.3. The 52 commits cover the latest patch release tag for every major and minor release from v6.0.1 to v9.4.3. Sourcerer searched the corpus as indexed in Elasticsearch, while <code>ripgrep</code>  and <code>grep</code> searched the corpus as stored on a filesystem, reflecting their respective use cases. The regular expression patterns included one that appears rarely among the commits (DiskBBQ) and one that appears commonly among the commits (XContentType).</p><h4>sourcerer.code.grep</h4><p>The first search I benchmarked was a rare pattern for DiskBBQ that appears only in some commits:</p><p><code>.*[dD][iI][sS][kK][-_]?[bB][bB][qQ].*</code></p><p>Search latency (in seconds) spanning a single commit (605 matches found from 7,215,509 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>0.069s</p><p>0.124s</p><p>0.167s</p><p>0.020s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>0.027s</p><p>0.029s</p><p>0.046s</p><p>0.005s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>0.788s</p><p>0.800s</p><p>0.809s</p><p>0.005s</p><p>~6.5x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>0.081s</p><p>0.088s</p><p>0.109s</p><p>0.009s</p><p>~3.0x slower</p><p><code>grep</code></p><p>Cold</p><p>3.565s</p><p>3.580s</p><p>3.821s</p><p>0.055s</p><p>~28.9x slower</p><p><code>grep</code></p><p>Warm</p><p>0.825s</p><p>0.827s</p><p>0.833s</p><p>0.002s</p><p>~28.5x slower</p><p>Search latency (in seconds) spanning 52 commits (1,041 matches found from 200,325,684 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>0.159s</p><p>0.164s</p><p>0.270s</p><p>0.028s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>0.027s</p><p>0.031s</p><p>0.053s</p><p>0.006s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>22.356s</p><p>22.364s</p><p>22.459s</p><p>0.026s</p><p>~136.4x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>16.017s</p><p>16.123s</p><p>16.297s</p><p>0.058s</p><p>~520.1x slower</p><p><code>grep</code></p><p>Cold</p><p>101.962s</p><p>102.386s</p><p>104.281s</p><p>0.752s</p><p>~624.3x slower</p><p><code>grep</code></p><p>Warm</p><p>73.082s</p><p>73.507s</p><p>74.915s</p><p>0.536s</p><p>~2,371.2x slower</p><p>Table 2. p0/p50/p100/stdev retrieval speeds of <code>sourcerer.code.grep</code>, <code>ripgrep</code> , and <code>grep</code>, under cold and warm caches, at two corpus scopes (20 runs per method per cache state; three warmup runs discarded before each warm-cache measurement). All percentiles computed via linear interpolation. Ratios are computed against <code>sourcerer.code.grep</code>'s p50 at the matching cache state. See the “Methodology” section for cache definitions and query/command syntax.</p><p>The <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field type indexes trigrams of each value and uses them as a filter to narrow the candidate set for a regular expression before verifying full matches. For a selective pattern like this one (605 matches out of 7.2 million lines, 1041 out of 200 million) this approaches sublinear time complexity relative to corpus size. <code>ripgrep</code>  and <code>grep</code> both perform scans with linear time complexity, with <code>ripgrep</code>  using multithreading and single instruction, multiple data–accelerated (SIMD-accelerated) literal prefiltering to speed up searches, but neither has a mechanism to skip the vast majority of a corpus the way that an indexed trigram search can.</p><p>This shows up starkly in how each approach scales. Going from the single-commit corpus to the all-commits corpus is a 27.8x increase in line count. <code>sourcerer.code.grep</code> warm-cache time barely moves from 0.029 seconds to 0.031 seconds. ripgrep's warm-cache time goes from 0.089 seconds to 16.1 seconds, a 183x change; and grep's goes from 0.83 seconds to 1.2 minutes, an 89x change. Both filesystem tools scale worse than linearly with corpus size on this hardware, while the indexed approach is nearly flat.</p><p>The second search I benchmarked was a common pattern for XContentType that appears in all commits:</p><p><code>.*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*</code></p><p>Search latency (in seconds) spanning a single commit (7,999 matches found from 7,215,509 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>0.158s</p><p>0.172s</p><p>0.224s</p><p>0.015s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>0.055s</p><p>0.057s</p><p>0.140s</p><p>0.023s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>0.794s</p><p>0.804s</p><p>0.824s</p><p>0.007s</p><p>~4.7x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>0.093s</p><p>0.095s</p><p>0.096s</p><p>0.001s</p><p>~1.7x slower</p><p><code>grep</code></p><p>Cold</p><p>3.385s</p><p>3.399s</p><p>3.421s</p><p>0.010s</p><p>~19.8x slower</p><p><code>grep</code></p><p>Warm</p><p>0.681s</p><p>0.683s</p><p>0.686s</p><p>0.001s</p><p>~12.1x slower</p><p>Search latency (in seconds) spanning 52 commits (290,662 matches found from 200,325,684 lines of code):</p><p><strong>Retrieval method</strong></p><p><strong>Cache</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>vs. sourcerer.code.grep</strong></p><p><code>sourcerer.code.grep</code></p><p>Cold</p><p>1.587s</p><p>1.644s</p><p>1.756s</p><p>0.047s</p><p>-</p><p><code>sourcerer.code.grep</code></p><p>Warm</p><p>1.480s</p><p>1.541s</p><p>1.627s</p><p>0.040s</p><p>-</p><p><code>ripgrep</code> </p><p>Cold</p><p>22.426s</p><p>22.440s</p><p>22.548s</p><p>0.028s</p><p>~13.6x slower</p><p><code>ripgrep</code> </p><p>Warm</p><p>16.021s</p><p>16.276s</p><p>16.498s</p><p>0.119s</p><p>~10.6x slower</p><p><code>grep</code></p><p>Cold</p><p>97.699s</p><p>97.922s</p><p>99.150s</p><p>0.404s</p><p>~59.6x slower</p><p><code>grep</code></p><p>Warm</p><p>68.883s</p><p>69.184s</p><p>70.374s</p><p>0.360s</p><p>~44.9x slower</p><p>Table 3. p0/p50/p100/stdev retrieval speeds for the pattern <code>.*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*</code> under the same conditions as Table 2.</p><p>The relative search latencies of <code>sourcerer.code.grep</code> compared to <code>grep</code> shows why the pattern's match count matters as much as the corpus size:</p><p><strong>Corpus</strong></p><p><strong>Corpus size</strong></p><p><strong>Cache</strong></p><p><strong>Speed of sourcerer.code.grep with a rare pattern (DiskBBQ)</strong></p><p><strong>Speed of sourcerer.code.grep with a common pattern (XContentType)</strong></p><p>1 commit</p><p>7,215,509 lines</p><p>Cold</p><p>~28.9x faster</p><p>~19.8x faster</p><p>1 commit</p><p>7,215,509 lines</p><p>Warm</p><p>~28.5x faster</p><p>~12.1x faster</p><p>52 commits</p><p>200,325,684 lines</p><p>Cold</p><p>~624.3x faster</p><p>~59.6x faster</p><p>52 commits</p><p>200,325,684 lines</p><p>Warm</p><p>~2,371.2x faster</p><p>~44.9x faster</p><p>Table 4. This table shows how much faster <code>sourcerer.code.grep</code> was compared to <code>grep</code> when searching across two different corpus sizes and two different pattern rarities.</p><p>The pattern with far more matches shows a dramatically smaller Elasticsearch advantage, most strikingly at all-commits scope, where the advantage drops from 2,371x to 45x. The reason is visible in the absolute numbers: <code>sourcerer.code.grep</code>'s warm-cache time at all-commits scope jumps from 0.031 seconds (DiskBBQ) to 1.541 seconds (XContentType), a 49.7x increase for a 279x increase in match count, while <code>grep</code>'s warm-cache time barely changes (73.5 seconds to 69.2 seconds, effectively flat, since it scans the same number of bytes regardless of how many of them match). The <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field's trigram index has sublinear-in-corpus-size behavior that comes specifically from narrowing the candidate set before verification. Once a pattern matches hundreds of thousands of lines, the bottleneck shifts from narrowing candidates to collecting and serializing all of them, a cost that scales with match count rather than corpus size. <code>sourcerer.code.grep</code> still wins by a wide margin even in this less favorable case, but the margin depends heavily on how selective the search is, not just how large the corpus is.</p><h4>sourcerer.code.search</h4><p>Sourcerer's other retrieval tool, <code>sourcerer.code.search</code>, performs a BM25-ranked <code>MATCH</code> query rather than an exact-substring regex match. Its speed on both patterns is included below for reference.</p><p><strong>Pattern</strong></p><p><strong>Corpus scope</strong></p><p><strong>p0</strong></p><p><strong>p50</strong></p><p><strong>p100</strong></p><p><strong>stdev</strong></p><p><strong>Matches found</strong></p><p>DiskBBQ</p><p>One commit</p><p>0.013s</p><p>0.017s</p><p>0.018s</p><p>0.002s</p><p>288</p><p>DiskBBQ</p><p>52 commits</p><p>0.014s</p><p>0.020s</p><p>0.046s</p><p>0.007s</p><p>493</p><p>XContentType</p><p>One commit</p><p>0.062s</p><p>0.063s</p><p>0.154s</p><p>0.020s</p><p>7,177</p><p>XContentType</p><p>52 commits</p><p>2.203s</p><p>2.359s</p><p>2.611s</p><p>0.125s</p><p>249,737</p><p>Table 5. p0/p50/p100/stdev retrieval speeds of <code>sourcerer.code.search</code>, warm cache only, for both patterns at both corpus scopes (20 runs per row; cold-cache figures omitted; see Methodology). Match counts are <code>sourcerer.code.search</code>'s own, not the regex-based methods' BM25 matches on tokens rather than substrings, so these aren’t directly comparable to Tables 2–4.</p><h4>What drives the speed advantage</h4><p><code>sourcerer.code.grep</code> outperformed <code>ripgrep</code> and <code>grep</code> at every corpus scale and pattern rarity tested, along with every cache state tested. But it wasn’t by a fixed margin. The advantage ranged from ~3–30x on a common pattern to over 2,300x on a rare pattern, because indexed and brute-force search respond to different things. The trigram narrowing of the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field does less work as a pattern gets more selective, while <code>grep</code> and <code>ripgrep</code> do the same amount of work regardless of how much of the corpus happens to match. <code>sourcerer.code.search</code> is a third option for the cases where the exact string isn't known at all.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcc34bd5e704c0112/6a7deb29b8c2e6ffc9be51a1/image1.png" alt="" /><p></p><p>Figure 2: Sourcerer's <code>sourcerer.code.grep</code> tool outperformed <code>ripgrep</code> and <code>grep</code> in every combination of corpus size and pattern rarity benchmarked, sometimes by multiple orders of magnitude. Its outperformance was strongest when searching rare patterns in large corpora and weakest when searching common patterns in small corpora.</p><p>These search speed benchmarks show what <em>scalable</em> means in practice for an enterprise code search agent. It's being able to search the histories of any number of repositories at interactive speeds, just like how a developer coding agent searches the working state of a single repository on a filesystem.</p><h4>Speed benchmark methodology</h4><p><a href="https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch#appendix-b.-code-search-speed-and-scalability-benchmark-configuration">Appendix B</a> explains the configuration of this benchmark. Note that while the Elasticsearch deployment had two data nodes each with the same specs as the virtual machine used for <code>ripgrep</code> and <code>grep</code>, the benchmark consisted of one index with one primary shard and one replica shard. Each search ran on a single shard, which means that <code>sourcerer.code.grep</code> had the same amount of vCPUs and memory available per search as <code>ripgrep</code> and <code>grep</code>, despite having twice as much capacity across the overall deployment. Having two data nodes actually incurred a slight latency <em>penalty</em> compared to just one data node. For brevity, I've omitted results from the benchmark with one data node. We don't recommend single-node deployments in production, so it's worth including the realistic latency overhead that comes with a multi-node deployment in this benchmark.</p><p>I compared all retrieval methods under both cold and warm caches, 20 cold runs and 20 warm runs per method. The definitions of cold and warm caches weren’t like-for-like between the Elasticsearch and filesystem benchmarks. For <code>ripgrep</code> and <code>grep</code>, a <em>cold cache</em> meant performing a full-page cache drop (<code>sync; echo 3 &gt; /proc/sys/vm/drop_caches</code>) immediately before each cold run; and a <em>warm cache</em> meant three discarded warmup executions immediately followed by the 20 measured runs, with no cache drops in between. For ES|QL, a <em>cold cache</em> meant calling <code>POST /_cache/clear</code> before each run. This only clears Elasticsearch's internal caches, not the OS page caches of the data nodes, which can't be cleared by hand on Elastic Cloud Hosted (ECH). A <em>warm cache</em> for ES|QL meant three discarded warmup queries immediately before the 20 measured runs, to help ensure that the caches on both data nodes would be warm. <code>sourcerer.code.search</code>'s cold-cache figures are omitted from Table 5 for the same reason discussed elsewhere in this post: Its cold-cache measurements came back statistically indistinguishable from its own warm-cache measurements, evidence that the OS-level cache-clearing limitation affects it more than it affects <code>sourcerer.code.grep</code>, which showed a consistent, physically sensible cold/warm gap throughout.</p><h3>Code search indexing throughput</h3><p>I didn't conduct a formal benchmark of indexing throughout. I'll share my general observations instead.</p><p>Typically, I see a sustained indexing throughput of 20K–25K lines per second on data nodes that each have ~16GiB RAM and ~8 vCPU on c4a-highcpu instances on Google Cloud Platform (GCP). That includes writing to a primary shard and its replica shard. I've seen throughput as high as ~60K lines per second on <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a> with <a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-tier-autoscaling">Search Power</a> set to "Performant."</p><p>An engineering team at Elastic compared Sourcerer's indexing throughput to semantic code search implementations that used either sparse vector generation (<a href="https://www.elastic.co/docs/explore-analyze/machine-learning/nlp/ml-nlp-elser">Elastic Learned Sparse EncodeR [ELSER])</a> or dense vector embedding generation (<a href="https://jina.ai/models/jina-embeddings-v5-text-small/">Jina</a>). Sourcerer indexed ~25x faster than <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/nlp/ml-nlp-elser">.elser-2-elastic</a> and ~15x faster than <a href="https://jina.ai/models/jina-embeddings-v5-text-small/">.jina-embeddings-v5-text-small</a>, while retrieval quality was similar among all of them. More concretely, what took ~6 hours to index with ELSER took ~14 minutes to index with Sourcerer.</p><h2>Code search solution design and rationale</h2><p>The remainder of this blog post explains the rationale for my design decisions of Sourcerer, giving expert insights for practitioners of Elasticsearch and generative AI (GenAI).</p><h3>Goals</h3><p>Ultimately, we want an agent that answers questions about deployed software and its supporting infrastructure by searching the primary sources of truth (the code itself) and generating verifiable responses that cite those sources so they can be trusted. Inspired by <a href="https://arxiv.org/abs/2605.15184">the success of coding agents with grep</a>, my main functional goal for Sourcerer was to reproduce the search behavior of a coding agent and generate responses with citations, all using Agent Builder. My nonfunctional goals were to keep it fast and scalable, as well as accurate, when searching across many versioned repositories, while maintaining acceptable costs and ease of use. Of these goals, reproducing the search behaviors of coding agents would be the most consequential, as it would dictate the access pattern, index design, and query design, plus their effects on nonfunctional goals.</p><h3>Access pattern</h3><p>From a human perspective, the intended access pattern is simple: We expect to ask natural language questions about software and receive plain language answers grounded in the source of truth. From the perspective of the agent handling those questions, the intended access pattern is to reproduce the search behavior of coding agents to find what it needs. The LLMs used by coding agents are heavily trained to explore code with shell commands, like<code>ls</code>or <code>find</code>, <code>grep</code> or <code>ripgrep</code>, <code>cat</code>, <code>head</code>, <code>tail</code>, and so on. <a href="https://code.claude.com/docs/en/tools-reference">Claude Code's built-in tools</a>, such as <a href="https://code.claude.com/docs/en/tools-reference#glob-tool-behavior">Glob</a> and <a href="https://code.claude.com/docs/en/tools-reference#grep-tool-behavior">Grep</a>, provide similar functions.</p><p>I chose to go with the grain of how models are trained. So my intended access pattern for Sourcerer was to reproduce the names, inputs, and outputs of shell commands as <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools/esql-tools">ES|QL tools in Agent Builder</a>. That way, the Agent Builder harness would allow an LLM to use its trained intuition to achieve similar results as a frontier harness, like Claude Code or Codex, without having to fill the LLM's limited context window with instructions for using a different search interface. This was the main context engineering problem to solve with Sourcerer. Solving it would enable faster searches that scale across many repositories at once, allowing an agent to answer questions about deployments in which many different versioned software projects work together.</p><h3>Index design</h3><p>I decided on three index templates to fulfill this access pattern:</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/index_templates/sourcerer-v2-refs.json"><code>sourcerer-refs</code></a>: Each document indexes the high-level metadata for a single Git reference or <em>ref</em> identified by its unique commit hash, which can have a tag name or branch name associated with it. A ref represents an entire snapshot of a repository at a point in time. This is a small index. The agent mainly uses this to discover the repositories and snapshots that are available to search.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/index_templates/sourcerer-v2-files.json"><code>sourcerer-files</code></a>: Each document indexes the metadata for a single file of a given ref. This is a larger index. The agent mainly uses this to navigate files and directories using <code>ls</code> semantics.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/index_templates/sourcerer-v2-lines.json"><code>sourcerer-lines</code></a>: Each document indexes the contents of a single, numbered line of code for a given file of a given ref. Yes, every line of code becomes a document. This is the largest index (but perhaps not as large as you might expect). The agent mainly uses this to search and view code using <code>grep</code> and <code>cat</code> semantics.</p></li></ul><p>The design is almost entirely denormalized. Each index has the same namespacing fields for fast, joinless filtering. Each indexed ref stores all files and lines from its commit snapshot, rather than storing diffs and reconstructing them at search time or storing unique files and lines with a mutable array of ref names associated with each. These choices trade duplicative storage (the cheapest compute resource) for faster searches and less segment merging pressure.</p><h4>Namespacing</h4><p>All three indices use four fields to namespace the ref, file, or line of code:</p><ul><li><p><code>git.host</code>: A Git hosting provider (for example, github, gitlab).</p></li><li><p><code>git.org</code>: An account name (for example, elastic).</p></li><li><p><code>git.repo</code>: A Git repository (for example, elasticsearch, kibana).</p></li><li><p><code>git.commit</code>: A commit hash, stored as the full 40-character SHA-1 digest for integrity.</p></li></ul><p>The document <code>_id</code> hashes for files and lines are also namespaced by <code>{git.host}</code>, <code>{git.org}</code>, <code>{git.repo}</code>, and <code>{git.commit}</code> to allow for idempotent indexing. That means you can safely rerun an indexing job without duplicating any documents.</p><p>Likewise, the index names are namespaced with the same semantics, using tildes (<code>~</code>) as a reliable separator since it's a disallowed character in Git repository names and organization names:</p><ul><li><p><code>sourcerer-v*-files~{git.host}~{git.org}~{git.repo}</code></p></li><li><p><code>sourcerer-v*-lines~{git.host}~{git.org}~{git.repo}</code></p></li></ul><p>This namespace convention has many benefits:</p><ul><li><p>Agents can quickly narrow the search space for refs and files, along with lines of code, by these common scoping fields, keeping searches fast and focused.</p></li><li><p>The semantics reflect common permission boundaries. You can reproduce the access policies of your Git hosting provider by implementing your choice of index-level security and/or document-level security based on host or organization or based on repository.</p></li><li><p>You can instantly delete indices for a whole repository or organization, or for a host, without an expensive <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-delete-by-query"><code>_delete_by_query</code></a>.</p></li><li><p>The index names are future proofed for different levels of granularity. Sourcerer might eventually allow indexing code by <code>{git.host}</code>, <code>{git.host}~{git.org}</code>, or <code>{git.host}~{git.org}~{git.repo}~{git.commit}</code> for selective shard sizing optimizations. The query syntax would be unaffected because they target index aliases (<code>sourcerer-files</code> and <code>sourcerer-lines</code>), not individual indices.</p></li></ul><h4>Settings</h4><p>Three index settings help to optimize storage costs and search speed:</p><ul><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/sorting">Index sorting</a> gives faster searches and better compression at the cost of reduced indexing throughput. Each index sorts documents on disk by <code>git.host</code>, <code>git.org</code>, <code>git.repo</code>, <code>git.commit</code>. File and line documents are further sorted by <code>file.path</code>, and line documents are further sorted by <code>line.number</code>.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-source-field#synthetic-source">Synthetic <code>_source</code></a> discards <code>_source</code> and instead reconstructs it as needed when reindexing. None of the queries access <code>_source</code>, which makes it dead weight. Enabling this setting reclaims ~50% storage space in the files and lines indices. While it requires an Enterprise license, enabling it on a non-licensed deployment won’t prevent the index from being created; instead the setting will be ignored.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/index-modules#index-codec"><code>best_compression</code></a> is a fallback for Elastic deployments that lack an Enterprise license to use synthetic <code>_source</code>. It provides decent compression for <code>_source</code> (~11% storage savings by my observations) in exchange for a modest tax on indexing throughput (~15% slower), while search speeds are essentially unaffected because the queries don't fetch <code>_source</code>.</p></li></ul><p>I use <a href="https://www.elastic.co/docs/manage-data/data-store/aliases">index aliases</a> to support zero-downtime upgrades when reindexing to a new schema.</p><h4>Mappings</h4><p>The indices mainly use <code>keyword</code> fields. They facilitate efficient filtering with basic wildcard support and aggregations, along with optimal storage usage and indexing throughput.</p><p>The <code>line.content</code> field is indexed both as a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field and as a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/text"><code>text</code></a> field with <a href="https://www.elastic.co/docs/manage-data/data-store/text-analysis">tokenization</a> and <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/similarity">similarity</a> settings tuned for code search. This gives agents the option to search code using familiar and effective <code>grep</code>-like regular expressions on the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field or using the inverted index of the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/text"><code>text</code></a> field to return only the highest ranking matched lines to reduce token usage and increase search speed. Both options are remarkably fast and scalable, taking milliseconds to finish in most cases.</p><h4>Shards</h4><p>While shards are becoming less relevant with the rise of stateless platforms like <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a>, I want the solution to accommodate all deployment modes of Elasticsearch. So I've given attention to the effect of the index design on the number and size of primary shards. Based on my observations, I expect that most users won’t need to give much attention to shards.</p><p>By default, Elasticsearch enforces a <a href="https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/size-shards#shard-count-per-node-recommendation">soft limit of 1,000 shards per data node</a> (including replicas). That means this solution will hit a soft limit of just under 250 repositories indexed per data node, because each repository is written to a files index and a lines index, each with one primary shard and one replica. Additionally, there’s a conventional best practice of limiting shard sizes to ~50GB, which affects how many refs you can index per repository. Both of these limits can be pushed a bit. But they reveal that this solution design really is optimized for its intended use case of searching the commit snapshots of supported, deployed software. You wouldn't use Sourcerer to index every repository on the Internet, and you shouldn't use it to index every ephemeral development branch. Plus, you should decide how many refs are worth retaining for each repository.</p><p>Repository-level granularity of indices appears to strike the right balance of shard counts and shard sizes. For reference, Kibana is one of the largest repositories on GitHub (<a href="https://stacey-gammon.github.io/repo-stats/">source</a>). I observed its shard size to be a manageable 60GB–75GB when retaining only the latest patch release for every major and minor version release from v6.0.0 to v9.5.0. That's great coverage for the Elastic deployments we see in the wild. If that's one of the largest repositories out there, you can expect just about any other repository to fit in a single shard as long as you have a reasonable retention policy, which I discuss in the next section (“Pruning”).</p><h4>Pruning</h4><p>By default, Sourcerer retains everything you index. Pruning lets you delete old refs to prevent unbounded growth. You can define ref retention policies based on the age of refs and the number of refs indexed in the repo. You can also define these policies based on the number of semantic versions indexed in the repo at any level of granularity for major, minor, patch, build, and prerelease versions. Some common configurations are to retain only the latest commit of the default branch or the most recent patch release tag for every major and minor version release tag.</p><h3>Elastic Agent Builder tools</h3><p>With the index design in place, we can review the tools that query those indices.</p><p><a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/tools/esql-tools">ES|QL tools in Agent Builder</a> are parameterized ES|QL queries with descriptions to guide the agent's use of them. Sourcerer has tools for several purposes: repo discovery, file discovery, code search, and code display. These tools reproduce the names, inputs, and outputs of shell commands that coding agents prefer to use when exploring code, making them intuitive enough for the LLM to use with minimal instructions passed into its context window.</p><h4>Repo discovery</h4><p>These are typically the first tools that the agent calls. Unlike most coding agents, which search within a single repository on a filesystem, Sourcerer is aware that its search space likely has multiple repositories and versions, and so its first step is to decide which repos and refs to scope its searches to.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.repos.list.yml"><code>sourcerer.repos.list</code></a>: Lists the repos that are available to search.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.repos.search.yml"><code>sourcerer.repos.search</code></a>: Lists the repos whose file contents best match a given query.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.refs.list.yml"><code>sourcerer.refs.list</code></a>: Lists the repos and refs that are available to search.</p></li></ul><h4>File discovery</h4><p>All file discovery tools support glob matching (<code>*</code> and <code>**</code>) on file paths.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.ls.yml"><code>sourcerer.files.ls</code></a>: Lists files and directories that match a given pattern.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.tree.yml"><code>sourcerer.files.tree</code></a>: Lists files and directories that match a given pattern in a tree-like format.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.wc.yml"><code>sourcerer.files.wc</code></a>: Counts lines, words, characters, bytes, and longest lines for each matching file.</p></li></ul><h4>Code search</h4><p>All file code search tools support glob matching (<code>*</code> and <code>**</code>) on file paths.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.grep.yml"><code>sourcerer.code.grep</code></a>: Searches lines of code using <a href="https://www.elastic.co/docs/reference/query-languages/sql/sql-like-rlike-operators"><code>RLIKE</code></a> on a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code></a> field for rapid execution of regular expressions.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.code.search.yml"><code>sourcerer.code.search</code></a>: Searches lines of code using <a href="https://www.elastic.co/docs/reference/query-languages/sql/sql-functions-search#sql-functions-search-match"><code>MATCH</code></a> on a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/text"><code>text</code></a> field that has been tuned for code search.</p></li></ul><h4>Code retrieval</h4><p>All file code retrieval tools concatenate the desired lines of any matching file and return them as a single, contiguous block of code in <code>grep -n</code> format, which is a format preferred by coding agents, including Claude Code's built-in <a href="https://code.claude.com/docs/en/tools-reference#grep-tool-behavior">Grep</a> tool. This lets the agent see a faithful representation of file contents with line-level attribution for precise citations, without requiring the agent to reconstruct the contents or infer line numbers through reasoning.</p><p>To illustrate, here's how the <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.head.yml"><code>sourcerer.files.head</code></a> tool formats the first five lines of <a href="https://raw.githubusercontent.com/elastic/kibana/refs/tags/v9.5.0/README.md">Kibana's <code>README.md</code></a> file, reconstructed from five documents from the lines index:</p>1:# Kibana
2:
3:Kibana is the open source interface to query, analyze, visualize, and manage your data stored in Elasticsearch.
4:
5:- [Getting Started](#getting-started)<p>All file code retrieval tools support glob matching (<code>*</code> and <code>**</code>) on file paths. Agents typically use these tools to display the contents of a single file.</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.cat.yml"><code>sourcerer.files.cat</code></a>: Concatenates and displays all lines for each matching file.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.head.yml"><code>sourcerer.files.head</code></a>: Concatenates and displays the first <code>n</code> lines for each matching file.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.tail.yml"><code>sourcerer.files.tail</code></a>: Concatenates and displays the last <code>n</code> lines for each matching file.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_tools/sourcerer.files.read_lines.yml"><code>sourcerer.files.read_lines</code></a>: Concatenates and displays the range of lines between two given line numbers for each matching file.</p></li></ul><h3>Agent Builder skills</h3><p>With the tools implemented, we can review the skills that guide the agent's proper use of them. Agents typically invoke the follow skills in this order:</p><ul><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/repo-discovery/SKILL.md"><code>sourcerer-repo-discovery</code></a>: Guides the agent in discovering and selecting the repositories that are available to search for a given prompt. While this is typically the first skill an agent invokes, the agent might return to it when tracing dependencies from other repositories or when answering questions that span multiple repositories or versions.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/ref-resolution/SKILL.md"><code>sourcerer-ref-resolution</code></a>: Guides the agent in resolving the names of tags or branches to their unique, immutable commit hashes. This lets the agent reliably filter its searches to a single commit snapshot.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/code-search/SKILL.md"><code>sourcerer-code-search</code></a>: Guides the agent in exploring code, with basic best practices on when and how to use the available tools for maximum efficiency.</p></li><li><p><a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/skills/code-citations/SKILL.md"><code>sourcerer-code-citations</code></a>: Guides the agent in citing files, directories, lines of code, and ranges of lines of code. Sourcerer auto-generates more specific citation skills for each major Git hosting provider, so that its citation links conform to the URL formats of each respective host.</p></li></ul><h3>Agent system prompt</h3><p>The final packaging of the agent comes with a <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/elastic/agent_builder_agents/sourcerer.yml">system prompt</a> that succinctly describes the agent's role and its high-level instructions. The configure file that defines the system prompt also defines the tools and skills that are made available to the agent, so that it can only execute what we permit it to execute.</p><h3>Sourcerer CLI</h3><p>The <a href="https://github.com/elastic/sourcerer">Sourcerer CLI</a> assists with setup and indexing, along with pruning to keep operations simple. Configuration is managed through a <a href="https://github.com/elastic/sourcerer/blob/main/specs/sourcerer-yml.md"><code>sourcerer.yml</code></a> configuration file.</p><ul><li><p><code>sourcerer setup</code>: Idempotently loads the index templates and Agent Builder configurations, in addition to Kibana dashboards. This is typically a one-time operation and takes a few seconds.</p></li><li><p><code>sourcerer index</code>: Checks for new refs that match patterns defined in <a href="https://github.com/elastic/sourcerer/blob/main/specs/sourcerer-yml.md"><code>sourcerer.yml</code></a> and then idempotently indexes them, skipping any refs that have already been indexed or that qualify for pruning based on retention policies. It calls <code>git</code> to clone repos and to list remote refs, as well as to  check out refs.</p></li><li><p><code>sourcerer prune</code>: Checks the retention policies for any refs that qualify for pruning and then deletes them from all three indices using <code>_delete_by_query</code>.</p></li></ul><p>You can easily schedule indexing and pruning using external schedulers, like cron. For our internal use at Elastic, we maintain <a href="https://github.com/elastic/sourcerer/blob/main/specs/sourcerer-yml.md"><code>sourcerer.yml</code></a> files in a private Git repository and schedule indexing and pruning with <a href="https://github.com/features/actions">GitHub Actions</a>. I prefer to index code frequently, while pruning outside of normal working hours to prevent agents from suddenly losing context in the middle of a conversation.</p><h3>Feature license summary</h3><p>For transparency, here’s a summary of the license levels for all non–open source software (non-OSS) features referenced in this solution design.</p><p>Subscription features:</p><ul><li><p><a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Agent Builder</a> (optional; you can query the indices from a different harness using the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/">Elasticsearch API</a>).</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-source-field#synthetic-source">Synthetic <code>_source</code></a> (optional).</p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/controlling-access-at-document-field-level#document-level-security">Document-level security</a> (optional).</p></li></ul><p>Free features, proprietary to Elastic (not OSS as defined by the Open Source Initiative [OSI]):</p><ul><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a>.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/keyword#wildcard-field-type"><code>wildcard</code> field</a>.</p></li></ul><p>A subscription gives you the magic of "everything just works" with Agent Builder, along with resource optimizations, finer security permissions, and platform support. Without a subscription, you can still index and prune code with the Sourcerer CLI and search the code using the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/">Elasticsearch API</a>.</p><h2>Conclusion</h2><p>Sourcerer demonstrates that Elasticsearch + Agent Builder is an exceptional solution for agentic enterprise code intelligence, as evidenced in many ways:</p><ul><li><p><strong>Accurate:</strong> Sourcerer's agentic code retrieval quality is at parity with frontier coding agents as demonstrated by its performance on the academic benchmark SWE-Explore.</p></li><li><p><strong>Fast:</strong> Sourcerer's query speed ranges from milliseconds in a single repository to multiple seconds across a billion lines of code. Indexing is also much faster than could be achieved with vector embedding generations.</p></li><li><p><strong>Scalable:</strong> Elasticsearch sharding enables horizontal scaling, making it possible to search the current and historical states of an entire enterprise software estate. Alternatively, the stateless architecture of <a href="https://www.elastic.co/cloud/serverless">Serverless</a> naturally scales without having to plan shards.</p></li><li><p><strong>Resilient:</strong> Elasticsearch replication enables high availability to keep the agent operational 24/7. Likewise, the stateless architecture of <a href="https://www.elastic.co/cloud/serverless">Serverless</a> naturally provides high availability.</p></li><li><p><strong>Polyglot:</strong> Sourcerer's indexing and retrieval methods are completely language-agnostic and tolerant of malformed code.</p></li><li><p><strong>Secure:</strong> Elasticsearch <a href="https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/controlling-access-at-document-field-level">document-level security</a> enforces access policies for humans and agents at the organization, repository, and commit levels.</p></li><li><p><strong>Efficient:</strong> Sourcerer demonstrates an efficient use of storage, memory, compute, and token consumption for its intended use case.</p></li><li><p><strong>Manageable:</strong> The Sourcerer CLI, and its use of native <code>git</code> commands and Elastic REST APIs, makes it easy to get started with and operate, as well as schedule.</p></li><li><p><strong>Universal:</strong> Organizations from all industries build software, and Git is the source of truth for ~85% of them (<a href="https://fosspost.org/git-market-share-statistics/">source</a>). Sourcerer has value to all of these organizations. </p></li></ul><p>Sourcerer is currently less than two months old, and the benchmark results suggest that there’s room to improve recall and token efficiency, along with task duration. I'll continue to work on this project in the near future.</p><h2>Try it yourself</h2><p>Sourcerer depends on Elasticsearch and Kibana. You can get started with those in a couple ways:</p><ul><li><p><a href="https://www.elastic.co/cloud/cloud-trial-overview">Elastic Cloud</a> (includes an Enterprise trial).</p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">Local setup with Docker</a> (includes an Enterprise trial).</p></li></ul><p>Then you can get started with <a href="https://github.com/elastic/sourcerer">Sourcerer</a>.</p><p>I built this with <a href="https://www.elastic.co/elasticsearch/agent-builder">Agent Builder</a>. What will you build?</p><h2>References</h2><p>Sen, S. (2026). <em>Is Grep All You Need? How Agent Harnesses Reshape Agentic Search</em>. arXiv.<a href="https://arxiv.org/abs/2605.15184"> https://arxiv.org/abs/2605.15184</a></p><p>Zhang, S. (2026). <em>SWE-Explore: Benchmarking How Coding Agents Explore Repositories</em>. arXiv.<a href="https://arxiv.org/abs/2606.07297"> https://arxiv.org/abs/2606.07297</a></p><h2>Appendices</h2><h3>Appendix A. SWE-Explore benchmark configuration</h3><p>Benchmark environment:</p><ul><li><p>Sourcerer version: v1.0.0 (commit hash: <a href="https://github.com/elastic/sourcerer/tree/26d2e84e3f9e4c1e1598a48532eba9a739465c3e">26d2e84e3f9e4c1e1598a48532eba9a739465c3e</a>)</p></li><li><p>Elastic Cloud Hosted (ECH):</p></li><ul><li><p>Region: GCP - Los Angeles (us-west2)</p></li><li><p>CPU Optimized hardware (c4a-highcpu)</p></li><li><p>3x Elasticsearch data nodes (each with 16GiB RAM, 8 vCPUs)</p></li><li><p>2x Kibana instances (each with 2GiB RAM)</p></li><li><p>Elastic stack version: v9.5.0 (commit hash: <a href="https://github.com/elastic/elasticsearch/tree/dbedef007f580447413782705acb9afec41f945b">dbedef007f580447413782705acb9afec41f945b</a>)</p></li></ul><li><p>Claude Code version: 2.1.202</p></li><li><p>LLM: GPT-5.4</p></li></ul><p>Indices as reported by <code>GET /_cat/indices</code>:</p>index                                          pri rep docs.count docs.deleted store.size
sourcerer-v1-files~ansible~ansible               1   1     234423            0     18.5mb
sourcerer-v1-files~apache~druid                  1   1      46720            0      4.1mb
sourcerer-v1-files~apache~lucene                 1   1      47932            0      4.2mb
sourcerer-v1-files~astral-sh~ruff                1   1      49021            0      9.1mb
sourcerer-v1-files~astropy~astropy               1   1      38994            0      6.1mb
sourcerer-v1-files~axios~axios                   1   1        339            0       89kb
sourcerer-v1-files~babel~babel                   1   1      24798            0      4.9mb
sourcerer-v1-files~briannesbitt~carbon           1   1      14054            0      1.2mb
sourcerer-v1-files~burntsushi~ripgrep            1   1        408            0    232.4kb
sourcerer-v1-files~caddyserver~caddy             1   1       2162            0    533.9kb
sourcerer-v1-files~django~django                 1   1    1329708            0     91.7mb
sourcerer-v1-files~element-hq~element-web        1   1      29426            0      6.1mb
sourcerer-v1-files~facebook~docusaurus           1   1       9288            0        2mb
sourcerer-v1-files~faker-ruby~faker              1   1       1130            0    207.2kb
sourcerer-v1-files~fastlane~fastlane             1   1      13663            0      1.7mb
sourcerer-v1-files~flipt-io~flipt                1   1      10518            0    873.8kb
sourcerer-v1-files~fluent~fluentd                1   1       3340            0    744.4kb
sourcerer-v1-files~fmtlib~fmt                    1   1       1243            0    133.7kb
sourcerer-v1-files~future-architect~vuls         1   1       3952            0      828kb
sourcerer-v1-files~gin-gonic~gin                 1   1        217            0     76.6kb
sourcerer-v1-files~gohugoio~hugo                 1   1      11545            0      1.2mb
sourcerer-v1-files~google~gson                   1   1       1432            0    349.4kb
sourcerer-v1-files~gravitational~teleport        1   1     100317            0     16.3mb
sourcerer-v1-files~hashicorp~terraform           1   1      13511            0      1.5mb
sourcerer-v1-files~immutable-js~immutable-js     1   1        458            0      120kb
sourcerer-v1-files~internetarchive~openlibrary   1   1      35448            0      5.9mb
sourcerer-v1-files~javaparser~javaparser         1   1       5154            0      1.3mb
sourcerer-v1-files~jekyll~jekyll                 1   1        720            0    272.9kb
sourcerer-v1-files~jordansissel~fpm              1   1        167            0     98.3kb
sourcerer-v1-files~jqlang~jq                     1   1       1172            0    291.4kb
sourcerer-v1-files~laravel~framework             1   1      29021            0      4.9mb
sourcerer-v1-files~matplotlib~matplotlib         1   1     133013            0     19.9mb
sourcerer-v1-files~micropython~micropython       1   1      15631            0        3mb
sourcerer-v1-files~mrdoob~three.js               1   1       9858            0      1.1mb
sourcerer-v1-files~mwaskom~seaborn               1   1        633            0      181kb
sourcerer-v1-files~navidrome~navidrome           1   1      10301            0      2.1mb
sourcerer-v1-files~nlohmann~json                 1   1       1090            0    281.9kb
sourcerer-v1-files~nodebb~nodebb                 1   1     104805            0     14.5mb
sourcerer-v1-files~nushell~nushell               1   1       8990            0      1.7mb
sourcerer-v1-files~pallets~flask                 1   1        251            0    120.1kb
sourcerer-v1-files~php-cs-fixer~php-cs-fixer     1   1       8300            0        1mb
sourcerer-v1-files~phpoffice~phpspreadsheet      1   1      17140            0      3.6mb
sourcerer-v1-files~preactjs~preact               1   1       3446            0    772.5kb
sourcerer-v1-files~projectlombok~lombok          1   1      24153            0        4mb
sourcerer-v1-files~prometheus~prometheus         1   1       3416            0    644.9kb
sourcerer-v1-files~protonmail~webclients         1   1      84927            0     14.9mb
sourcerer-v1-files~psf~requests                  1   1        928            0    329.4kb
sourcerer-v1-files~pydata~xarray                 1   1       5114            0    910.6kb
sourcerer-v1-files~pylint-dev~pylint             1   1      24945            0      3.8mb
sourcerer-v1-files~pytest-dev~pytest             1   1       8880            0      1.3mb
sourcerer-v1-files~qutebrowser~qutebrowser       1   1      29453            0      4.6mb
sourcerer-v1-files~reactivex~rxjava              1   1       1959            0    523.7kb
sourcerer-v1-files~redis~redis                   1   1      12646            0        2mb
sourcerer-v1-files~rubocop~rubocop               1   1      15349            0      3.3mb
sourcerer-v1-files~scikit-learn~scikit-learn     1   1      39030            0      6.2mb
sourcerer-v1-files~sharkdp~bat                   1   1       2526            0    789.3kb
sourcerer-v1-files~sphinx-doc~sphinx             1   1      55638            0      8.8mb
sourcerer-v1-files~sympy~sympy                   1   1     112572            0     16.7mb
sourcerer-v1-files~tokio-rs~axum                 1   1       1474            0    365.9kb
sourcerer-v1-files~tokio-rs~tokio                1   1       5175            0        1mb
sourcerer-v1-files~tutao~tutanota                1   1       9267            0      1.6mb
sourcerer-v1-files~uutils~coreutils              1   1       3226            0    806.6kb
sourcerer-v1-files~valkey-io~valkey              1   1       6681            0      1.4mb
sourcerer-v1-files~vuejs~core                    1   1       2785            0      750kb
sourcerer-v1-lines~ansible~ansible               1   1   23067027            0      6.9gb
sourcerer-v1-lines~apache~druid                  1   1   10230691            0      1.4gb
sourcerer-v1-lines~apache~lucene                 1   1    9857695            0      1.5gb
sourcerer-v1-lines~astral-sh~ruff                1   1    5895776            0    797.9mb
sourcerer-v1-lines~astropy~astropy               1   1   16855521            0      5.1gb
sourcerer-v1-lines~axios~axios                   1   1     104915            0     17.3mb
sourcerer-v1-lines~babel~babel                   1   1     661054            0     96.4mb
sourcerer-v1-lines~briannesbitt~carbon           1   1    2256429            0      339mb
sourcerer-v1-lines~burntsushi~ripgrep            1   1     121406            0     20.5mb
sourcerer-v1-lines~caddyserver~caddy             1   1     423391            0     60.4mb
sourcerer-v1-lines~django~django                 1   1  187595831            0     26.4gb
sourcerer-v1-lines~element-hq~element-web        1   1    5435563            0        1gb
sourcerer-v1-lines~facebook~docusaurus           1   1    1098502            0      173mb
sourcerer-v1-lines~faker-ruby~faker              1   1     281240            0     72.8mb
sourcerer-v1-lines~fastlane~fastlane             1   1    3774359            0    515.3mb
sourcerer-v1-lines~flipt-io~flipt                1   1    2318144            0    339.1mb
sourcerer-v1-lines~fluent~fluentd                1   1     616236            0     87.9mb
sourcerer-v1-lines~fmtlib~fmt                    1   1     525144            0     78.7mb
sourcerer-v1-lines~future-architect~vuls         1   1    1340951            0    410.3mb
sourcerer-v1-lines~gin-gonic~gin                 1   1      42186            0      6.2mb
sourcerer-v1-lines~gohugoio~hugo                 1   1    1323287            0    415.2mb
sourcerer-v1-lines~google~gson                   1   1     269864            0     82.1mb
sourcerer-v1-lines~gravitational~teleport        1   1   35961497            0      5.3gb
sourcerer-v1-lines~hashicorp~terraform           1   1    1913651            0    284.3mb
sourcerer-v1-lines~immutable-js~immutable-js     1   1     132082            0     41.1mb
sourcerer-v1-lines~internetarchive~openlibrary   1   1    6505234            0    903.1mb
sourcerer-v1-lines~javaparser~javaparser         1   1     756387            0    233.7mb
sourcerer-v1-lines~jekyll~jekyll                 1   1      57072            0      9.2mb
sourcerer-v1-lines~jordansissel~fpm              1   1      32285            0      8.8mb
sourcerer-v1-lines~jqlang~jq                     1   1     373367            0     56.1mb
sourcerer-v1-lines~laravel~framework             1   1    4680025            0      1.2gb
sourcerer-v1-lines~matplotlib~matplotlib         1   1   24426468            0      3.6gb
sourcerer-v1-lines~micropython~micropython       1   1    2104199            0      320mb
sourcerer-v1-lines~mrdoob~three.js               1   1    5519778            0   1014.6mb
sourcerer-v1-lines~mwaskom~seaborn               1   1     219103            0       35mb
sourcerer-v1-lines~navidrome~navidrome           1   1    1293741            0    405.1mb
sourcerer-v1-lines~nlohmann~json                 1   1     174859            0     53.5mb
sourcerer-v1-lines~nodebb~nodebb                 1   1    6621106            0      2.2gb
sourcerer-v1-lines~nushell~nushell               1   1    1424029            0    405.3mb
sourcerer-v1-lines~pallets~flask                 1   1      34538            0     11.1mb
sourcerer-v1-lines~php-cs-fixer~php-cs-fixer     1   1    1272280            0    357.6mb
sourcerer-v1-lines~phpoffice~phpspreadsheet      1   1    1999435            0      291mb
sourcerer-v1-lines~preactjs~preact               1   1     975966            0    274.3mb
sourcerer-v1-lines~projectlombok~lombok          1   1    1480367            0    470.7mb
sourcerer-v1-lines~prometheus~prometheus         1   1    1132908            0    484.6mb
sourcerer-v1-lines~protonmail~webclients         1   1   22247802            0      3.2gb
sourcerer-v1-lines~psf~requests                  1   1     337782            0    144.1mb
sourcerer-v1-lines~pydata~xarray                 1   1    2413091            0    714.7mb
sourcerer-v1-lines~pylint-dev~pylint             1   1    1188057            0    351.7mb
sourcerer-v1-lines~pytest-dev~pytest             1   1    1882734            0    532.2mb
sourcerer-v1-lines~qutebrowser~qutebrowser       1   1    6734755            0        1gb
sourcerer-v1-lines~reactivex~rxjava              1   1     486774            0    135.3mb
sourcerer-v1-lines~redis~redis                   1   1    3551629            0        1gb
sourcerer-v1-lines~rubocop~rubocop               1   1    2899319            0    763.3mb
sourcerer-v1-lines~scikit-learn~scikit-learn     1   1   10940081            0      3.2gb
sourcerer-v1-lines~sharkdp~bat                   1   1     317845            0      118mb
sourcerer-v1-lines~sphinx-doc~sphinx             1   1   15102452            0      3.9gb
sourcerer-v1-lines~sympy~sympy                   1   1   45538891            0     13.8gb
sourcerer-v1-lines~tokio-rs~axum                 1   1     146486            0     40.6mb
sourcerer-v1-lines~tokio-rs~tokio                1   1    1068294            0    293.9mb
sourcerer-v1-lines~tutao~tutanota                1   1    2932455            0    985.3mb
sourcerer-v1-lines~uutils~coreutils              1   1     478330            0    127.7mb
sourcerer-v1-lines~valkey-io~valkey              1   1    1876977            0    575.5mb
sourcerer-v1-lines~vuejs~core                    1   1     684268            0    190.9mb
sourcerer-v1-refs                                1   1        847            0    457.5kb<p>Benchmark task prompt:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3b4aec652d09215e/6a816866f6ab8740b5d16c31/Screenshot_2026-08-16_at_10.35.51_a.m..png" alt="" /><p>Table 6. This table compares <a href="http://google.com/url?q=https://github.com/Qiushao-E/SWE-Explore-Bench/blob/main/explorers/claude_code.py%23L30-L47&amp;sa=D&amp;source=docs&amp;ust=1786869254292102&amp;usg=AOvVaw3XAbMzUTtNbrrEmcvVpLY9">Claude Code's task prompt </a>used in the original paper and <a href="https://github.com/elastic/sourcerer/blob/main/src/sourcerer/commands/benchmark/swe_explore_bench/explorer.py#L145-L177">Sourcerer's task prompt</a> used in this benchmark run. The highlights indicate where Sourcerer's task prompt differed from Claude Code's: red indicates an instruction from Claude Code's task prompt that wasn't used by Sourcerer's, and green indicates an instruction that was unique to Sourcerer's task prompt.</p><h3>Appendix B. Code search speed and scalability benchmark configuration</h3><p>Benchmark environment:</p><ul><li><p>Sourcerer version: v2.0.0 (commit hash: <a href="https://github.com/elastic/sourcerer/tree/dcc373b3bf7b24168e00ed7b172aea7227f705f9">dcc373b3bf7b24168e00ed7b172aea7227f705f9</a>)</p></li><li><p>Elastic Cloud Hosted (ECH) for ES|QL:</p></li></ul><ul><li><p>Region: GCP - Los Angeles (us-west2)</p></li><li><p>CPU Optimized hardware (c4a-highcpu)</p></li><li><p>2x Elasticsearch data nodes (each with 16GiB RAM, 8 vCPUs)</p></li><li><p>1x Kibana instances (2GiB RAM)</p></li><li><p>Elastic stack version: v9.5.0 (commit hash: <a href="https://github.com/elastic/elasticsearch/tree/8d4246a64bc255212407b1b313fe402391299c88">8d4246a64bc255212407b1b313fe402391299c88</a>)</p></li></ul><ul><li><p>Virtual machine for <code>ripgrep</code> and <code>grep</code>:</p></li><ul><li><p>Region: GCP - Los Angeles (us-west2-a)</p></li><li><p>CPU Optimized hardware (c4a-highcpu-8-lssd)</p></li><li><p>16GiB RAM</p></li><li><p>Image: projects/ubuntu-os-cloud/global/images/ubuntu-2404-noble-arm64-v20260717</p></li><li><p>Provisioned IOPS: 3300</p></li><li><p>Provisioned throughput: 215</p></li></ul></ul><p>Indices as reported by <code>GET /_cat/indices</code>:</p>index                                           pri rep docs.count docs.deleted store.size
sourcerer-v2-lines~github~elastic~elasticsearch   1   0  200387235            0     44.8gb<p>Cluster settings adjusted to avoid truncating the returned match count on frequent patterns:</p>PUT /_cluster/settings
{
  "persistent": {
    "esql.query.result_truncation_max_size": 1000000
  }
}<p>Regular expression used by <code>ripgrep</code> and <code>grep</code>:</p><ul><li><p>DiskBBQ: <code>.*[dD][iI][sS][kK][-_]?[bB][bB][qQ].*</code></p></li><li><p>XContentType: <code>.*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*</code></p></li></ul><p>ES|QL query syntax used for <code>sourcerer.code.grep</code>:</p><p>ES|QL query syntax used for <code>sourcerer.code.search</code>:</p><p>ES|QL query parameters used in all tests:</p><ul><li><p><code>git_host="github"</code></p></li><li><p><code>git_org="elastic"</code></p></li><li><p><code>git_repo="elasticsearch"</code></p></li><li><p><code>n=1000000</code></p></li></ul><p>ES|QL query parameters used for specific tests:</p><ul><li><p>Corpus with single commit: <code>git_commit="45f6a06b1b441b41fe711059b8720013173e7c89"</code></p></li><li><p>Corpus with all commits: <code>git_commit="*"</code></p></li><li><p><code>sourcerer.code.grep</code> pattern for rare pattern (DiskBBQ): <code>regex=".*[dD][iI][sS][kK][-_]?[bB][bB][qQ].*"</code></p></li><li><p><code>sourcerer.code.grep</code> pattern for common pattern (XContentType): <code>regex=".*[xX][cC][oO][nN][tT][eE][nN][tT][tT][yY][pP][eE].*"</code></p></li><li><p><code>sourcerer.code.search</code> pattern for common pattern (DiskBBQ): <code>q="DiskBBQ"</code></p></li><li><p><code>sourcerer.code.search</code> pattern for common pattern (XContentType): <code>q="XContentType"</code></p></li></ul><p><code>n</code>  was set high enough to recover the true, uncapped match count for each pattern (605 and 1,041 matches for DiskBBQ;  7,999 and 290,662 matches for XContentType), confirmed in each case by comparing <code>hits_returned</code> against <code>ripgrep</code>'s and <code>grep</code>'s own match counts on the same underlying data.</p><p>Command syntax used for <code>ripgrep</code> :</p><p><code>rg -n --no-ignore -j 6</code></p><p>Command syntax used for <code>grep</code>:</p><p><code>grep -E -r -n --binary-files=without-match --exclude-dir=.git</code> </p><p>Directories of each cloned repository snapshot and the approximate total sizes of their nonbinary files tracked by Git, which is the search space of <code>ripgrep</code> and <code>grep</code>:</p>54M     elastic-elasticsearch-v6.0.1
55M     elastic-elasticsearch-v6.1.4
56M     elastic-elasticsearch-v6.2.4
79M     elastic-elasticsearch-v6.3.2
83M     elastic-elasticsearch-v6.4.3
89M     elastic-elasticsearch-v6.5.4
95M     elastic-elasticsearch-v6.6.2
99M     elastic-elasticsearch-v6.7.2
100M    elastic-elasticsearch-v6.8.23
99M     elastic-elasticsearch-v7.0.1
99M     elastic-elasticsearch-v7.1.1
135M    elastic-elasticsearch-v7.10.2
137M    elastic-elasticsearch-v7.11.2
140M    elastic-elasticsearch-v7.12.1
144M    elastic-elasticsearch-v7.13.4
147M    elastic-elasticsearch-v7.14.2
149M    elastic-elasticsearch-v7.15.2
154M    elastic-elasticsearch-v7.16.3
157M    elastic-elasticsearch-v7.17.29
103M    elastic-elasticsearch-v7.2.1
105M    elastic-elasticsearch-v7.3.2
109M    elastic-elasticsearch-v7.4.2
113M    elastic-elasticsearch-v7.5.2
118M    elastic-elasticsearch-v7.6.2
124M    elastic-elasticsearch-v7.7.1
127M    elastic-elasticsearch-v7.8.1
132M    elastic-elasticsearch-v7.9.3
149M    elastic-elasticsearch-v8.0.1
152M    elastic-elasticsearch-v8.1.3
173M    elastic-elasticsearch-v8.10.4
182M    elastic-elasticsearch-v8.11.4
187M    elastic-elasticsearch-v8.12.2
191M    elastic-elasticsearch-v8.13.4
196M    elastic-elasticsearch-v8.14.3
205M    elastic-elasticsearch-v8.15.5
230M    elastic-elasticsearch-v8.16.6
233M    elastic-elasticsearch-v8.17.10
241M    elastic-elasticsearch-v8.18.8
250M    elastic-elasticsearch-v8.19.18
149M    elastic-elasticsearch-v8.2.3
153M    elastic-elasticsearch-v8.3.3
155M    elastic-elasticsearch-v8.4.3
159M    elastic-elasticsearch-v8.5.3
161M    elastic-elasticsearch-v8.6.2
164M    elastic-elasticsearch-v8.7.1
168M    elastic-elasticsearch-v8.8.2
170M    elastic-elasticsearch-v8.9.2
231M    elastic-elasticsearch-v9.0.8
244M    elastic-elasticsearch-v9.1.10
254M    elastic-elasticsearch-v9.2.8
267M    elastic-elasticsearch-v9.3.7
305M    elastic-elasticsearch-v9.4.3]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/code-search-sourcerer-elasticsearch</guid>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Dave Moore]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf9ea6549b4ea8708/6a7deaae0da673fe0c57c6ea/image3.png" length="0" type="image/png"/>
    <pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Search relevance from click streams: Using Learn To Rank and behavioral signals with OpenTelemetry]]></title>
    <description><![CDATA[Learn how to turn click streams and behavioral signals from OpenTelemetry search analytics into judgment lists, rank features and Learn To Rank models that make search relevance improve over time.]]></description>
    <content:encoded><![CDATA[<p>Every click on a search result is an implicit relevance judgment, and conversions are a stronger signal. The search analytics you've been capturing through OpenTelemetry contain the behavioral data to improve search relevance. Techniques start simple, with fixes you can ship this week and build toward Learn To Rank (LTR) models trained on real click data. The instrumentation that surfaces problems also generates the training data to fix them.</p><h2>What you'll discover</h2><p>In this post, you'll learn how to:</p><ul><li><p>Build judgment lists from click data to evaluate and improve relevance.</p></li><li><p>Apply basic search tuning (field weights, boosts, query rules) informed by analytics.</p></li><li><p>Create rank features from behavioral signals, like popularity and conversion rate.</p></li><li><p>Understand LTR and how click data becomes training data.</p></li><li><p>Close the feedback loop between analytics and relevance improvement.</p></li></ul><h2>What you'll need</h2><ul><li><p>Search analytics data from the previous blogs in this <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">series</a> (<a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">search</a>, <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">click</a>, and <a href="https://www.elastic.co/search-labs/blog/search-conversion-tracking-opentelemetry">conversion</a> spans in Elastic).</p></li><li><p>An Elasticsearch index with product data (the reference project includes sample data with rank features).</p></li><li><p>Familiarity with Elasticsearch queries (BM25 [Elasticsearch's default text scoring algorithm], <code>rank_feature</code>, function scores).</p></li></ul><h2>Search relevance techniques at a glance</h2><p><strong>Technique</strong></p><p><strong>Effort</strong></p><p><strong>What it improves</strong></p><p><strong>When to use</strong></p><p><strong>Field weight tuning</strong></p><p>Low</p><p>Relevance for attribute-rich queries</p><p>First step; quick wins</p><p><strong>Query rules</strong></p><p>Low–medium</p><p>Specific high-value queries</p><p>Known bad results for specific terms</p><p><strong>Rank features</strong></p><p>Medium</p><p>Blending behavioral signals with text score</p><p>Popularity, conversion, freshness boosts</p><p><strong>LTR</strong></p><p>High</p><p>Systematic ranking from click data</p><p>When you have ≥5k labeled query-doc pairs</p><h2>From search analytics to search relevance improvements</h2><p>Over the past three posts, you built a full instrumentation pipeline, including search spans with <code>search.*</code> attributes and click tracking with position data and click-through rate (CTR) / Mean Reciprocal Rank (MRR) metrics. This pipeline also includes conversion spans tying searches to revenue. And it all sits in <code>traces-generic.otel-default</code>, queryable with Elasticsearch Query Language (ES|QL).</p><p></p><p>Now you have dashboards, and you know that your CTR is 28%. You also know which queries generate revenue and which ones users abandon. Plus, you can tell your product manager exactly where the funnel leaks.</p><p></p><p>Now what?</p><p></p><p>The real value of search analytics is using behavioral data to improve relevance and close feedback loops. It’s also important to make search learn from its users. Your experience tells you that measurement is necessary but not sufficient, and a dashboard that shows poor ranking doesn't improve that ranking. </p><p></p><p></p><p></p><p>This post covers four practical areas:</p><p></p><ol><li><p><strong>Judgment lists:</strong> The foundation for evaluating and improving relevance.</p></li><li><p><strong>Basic search tuning:</strong> Field weights, boosts, and query rules informed by analytics.</p></li><li><p><strong>Rank features for personalization:</strong> Feeding behavioral signals back into ranking.</p></li><li><p><strong>LTR:</strong> Training machine learning (ML) models on click data to optimize ranking automatically.</p></li></ol><p></p><p>Each one draws directly from the <code>search.*</code> attributes you're already collecting, and they don’t require any new instrumentation.</p><h2>What are judgment lists?</h2><p>Before diving into specific techniques, it's worth understanding the concept that ties them all together: <em>judgment lists</em>.</p><p></p><p>A judgment list is a set of query-document pairs with relevance grades: For a given query, how relevant is each document? They look like this:</p><p></p><p><strong>Query</strong></p><p><strong>Document</strong></p><p><strong>Grade</strong></p><p><strong>Label</strong></p><p>"wireless headphones"</p><p>SKU-001 (Sony WH-1000XM5)</p><p>3</p><p>Highly relevant</p><p>"wireless headphones"</p><p>SKU-042 (AirPods Max)</p><p>2</p><p>Relevant</p><p>"wireless headphones"</p><p>SKU-099 (Wired earbuds)</p><p>0</p><p>Not relevant</p><p></p><p>Judgment lists serve three purposes:</p><p></p><ul><li><p><strong>Evaluating current relevance.</strong> Given a set of queries and known-good results, how well does your search rank them? Metrics like <a href="https://en.wikipedia.org/wiki/Discounted_cumulative_gain">Normalized Discounted Cumulative Gain</a> (NDCG) and the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-rank-eval.html">Rank Eval API</a> use judgment lists to score your ranking quality. This gives you a baseline before making changes.</p></li></ul><p></p><ul><li><p><strong>Measuring the impact of changes.</strong> When you adjust field weights or add synonyms, judgment lists tell you whether the change helped or hurt. The same is true when you change boosting rules. Run the same evaluation before and after: If NDCG went up, the change improved relevance for those queries.</p></li></ul><p></p><ul><li><p><strong>Training LTR models.</strong> LTR algorithms need labeled training data, such as, <em>For this query, these documents are relevant and those aren't</em>. Judgment lists are that training data.</p></li></ul><h3>Manual vs. automated judgment lists</h3><p>Traditionally, judgment lists are created by human assessors who manually rate documents for a set of test queries. This works well for a small number of high-value queries (your top 50 searches, for example), but it doesn't scale. A large catalog with thousands of distinct queries and frequent inventory changes makes manual assessment impractical.</p><p></p><p>This is where your click data becomes valuable. Every click is an implicit relevance judgment; that is, a signal that for a given query, a given document was relevant enough to engage with. Conversions are even stronger signals. By aggregating this data, you can build judgment lists automatically from real user behavior, at a scale that manual assessment can't match.</p><p></p><p>The trade-off is noise. Clicks are influenced by position bias (users click higher-ranked results more often regardless of relevance) and presentation effects. They’re also influenced by accidental clicks. We'll cover techniques to handle this noise later in this post. But the key insight is that approaches like personalization and LTR lean toward automated judgment lists because it's not scalable to manually create lists for every query or user segment, nor for every inventory change.</p><p></p><p>The <a href="https://www.elastic.co/search-labs/blog/judgment-lists-search-query-relevance-elasticsearch">judgment lists guide on Search Labs</a> covers the concept in depth, including how to structure lists for evaluation with the Rank Eval API.</p><h2>How search analytics identify search relevance problems</h2><p>The most immediate use of your analytics data is to identify and fix specific relevance problems. This only requires using data to direct manual improvements, no personalization or ML.</p><h3>Finding problem queries with search analytics</h3><p>The ES|QL queries from the click in your application give you per-query CTR and MRR. Sort by search volume descending and CTR ascending to find your highest-impact relevance failures:</p>FROM traces-generic.otel-default
| WHERE ((name == "search" AND attributes.search.query IS NOT NULL)
    OR attributes.search.first_click == true)
  AND attributes.search.query IS NOT NULL
| STATS
    searches = COUNT(CASE(name == "search" AND attributes.search.query IS NOT NULL, 1)),
    clicked = COUNT(CASE(attributes.search.first_click == true, 1))
  BY attributes.search.query
| EVAL ctr_pct = ROUND(100.0 * clicked / searches, 1)
| WHERE searches &gt; 5
| SORT ctr_pct ASC, searches DESC
| LIMIT 20<p>This is the CTR-by-query query re-sorted to surface problems first. The <code>WHERE searches &gt; 5</code> filter removes one-off queries that would dominate the low-CTR list with small sample noise.</p><ul><li><p><strong>Zero-CTR queries with results. </strong> Your ranking is returning content, but none of it’s compelling. These queries often benefit from synonym expansion. You can also improve them with boosting rules or pinned results.</p></li><li><p><strong>Low-CTR, high-volume queries.</strong> These are the biggest relevance investment opportunities, and they affect the most users.</p></li><li><p><strong>Low-MRR, high-CTR queries.</strong> Users find what they need but have to scroll for it. In these instances, the relevant documents exist, but they're ranked wrong.</p></li></ul><h3>Try it: Close the loop in the reference project</h3><p>If you have click data in the reference project(<code>python generate_traffic.py --blog 3 --sessions 100</code>), you can run the problem-query query above against <code>traces-generic.otel-default</code> and observe your lowest-CTR queries.</p><p></p><p>The reference project's <code>app.py</code> already uses <code>rank_feature</code> boosting: <code>rank_features.popularity</code>, <code>rank_features.conversion_rate</code>, <code>rank_features.margin_score</code>, and <code>rank_features.freshness</code> are indexed on every product and blended into the BM25 score at query time. To see the effect of adjusting a boost:</p><p></p><ol><li><p>Open <code>reference/app.py</code>, and find the <code>"should"</code> clause in <code>_build_search_query()</code>.</p></li><li><p>Change the <code>"boost"</code> value on <code>rank_features.popularity</code> from <code>2</code> to <code>5</code>.</p></li><li><p>Restart the server (<code>python app.py</code>) and rerun <code>generate_traffic.py --blog 3 --sessions 50</code>.</p></li><li><p>Compare the top-queries and click position distribution in ES|QL before and after.</p></li></ol><p></p><p>This is informed manual tuning using the behavioral data you collected, not ML. That's the pattern for the rest of this post: First, ES|QL tells you what's wrong. Then, configuration changes and, eventually, LTR models fix it.</p><h3>Field weights, boosts, and scoring for search relevance</h3><p>The simplest tuning lever in Elasticsearch is adjusting how different fields contribute to the relevance score. A <code>multi_match</code> query across <code>title</code>, <code>description</code>, and <code>brand</code> fields can weight the title higher because a match there is usually more relevant. Your analytics data tells you where these weights are wrong: If users consistently click products with titles that don't match the query but with descriptions that do, your title boost may be too aggressive.</p><p></p><p>Beyond field weights, you can incorporate business metrics directly into scoring. A common example is <em>boosting by profit margin</em>; that is, products with higher margins rank slightly higher when relevance scores are similar. The <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/rank-feature.html"><code>rank_feature</code> field type</a> is designed for exactly this: Index a numeric signal (margin, popularity, recency) alongside each document, and the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-rank-feature-query.html"><code>rank_feature</code> query</a> blends it with text relevance at query time. For more complex scoring combinations, like decay functions, weighted field values, and scripts, the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-function-score-query.html"><code>function_score</code> query</a> gives you full control. The <a href="https://www.elastic.co/search-labs/blog/function-score-query-boosting-profit-popularity-elasticsearch">boosting by profit and popularity guide on Search Labs</a> walks through this in detail, including the trade-offs between additive and <a href="https://www.elastic.co/search-labs/blog/bm25-ranking-multiplicative-boosting-elasticsearch">multiplicative boosting</a> approaches.</p><h3>Query rules for targeted search relevance fixes</h3><p>For targeted interventions, <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/searching-with-query-rules">query rules</a> let you pin, boost, or exclude specific results for specific queries, and no model training is required. They're the search equivalent of a manual override.</p><p></p><p>Your analytics data tells you exactly where to apply them. When you see a zero-CTR query, like "returns policy", that consistently returns product results instead of the returns page, you can pin the returns page at position 1 for that query. If you notice a high-revenue query where the best-selling product appears at position 4, you can boost it.</p><p></p><p>Query rules are valuable precisely because they're simple. They solve known problems immediately while you build toward more sophisticated approaches. The <a href="https://www.elastic.co/search-labs/blog/elasticsearch-query-rules-ui-introduction">query rules tutorial</a> on Search Labs walks through the setup.</p><h2>Building judgment lists from click data</h2><p>The basic tuning above handles individual problems. To improve relevance systematically, you need judgment lists, and your click data can build them automatically.</p><h3>Converting click data into relevance grades</h3><p>The simplest approach is to count clicks per query-document pair and assign graded relevance:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
  AND attributes.search.query IS NOT NULL
| STATS
    click_count = COUNT(*),
    avg_position = AVG(attributes.search.result_click_position)
  BY attributes.search.query, attributes.search.result_click_id
| SORT attributes.search.query, click_count DESC<p>This gives you a table of (<code>query</code>, <code>document</code>, <code>click_count</code>, <code>avg_position</code>) tuples. The grading step maps click counts to relevance levels:</p><p><strong>Click count</strong></p><p><strong>Suggested grade</strong></p><p><strong>Label</strong></p><p>0</p><p>0</p><p>Not relevant (never clicked for this query)</p><p>1</p><p>1</p><p>Marginally relevant</p><p>2–3</p><p>2</p><p>Relevant</p><p>4+</p><p>3</p><p>Highly relevant</p><p></p><p>The thresholds depend on your traffic volume. For a high-traffic site, you might need 10+ clicks before calling something "highly relevant." For lower traffic, even two or three clicks is a meaningful signal. The point is that click frequency across users is a stronger signal than any single click.</p><p></p><p>You can strengthen judgments further by incorporating conversion data, a document that gets clicked <em>and</em> added to cart is a stronger relevance signal than one that gets clicked and abandoned:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action IN ("click", "add_to_cart")
  AND attributes.search.query IS NOT NULL
| STATS
    clicks = COUNT(CASE(attributes.search.action == "click", 1)),
    carts = COUNT(CASE(attributes.search.action == "add_to_cart", 1))
  BY attributes.search.query, attributes.search.result_click_id
| SORT attributes.search.query, clicks DESC<p>A document with five clicks and three cart additions is a stronger candidate for grade 3 than one with five clicks and zero cart additions.</p><h3>Handling position bias in click data</h3><p>There's a problem with raw click counts: <em>position bias</em>. Users click position 1 more often because they <em>see</em> it first, not necessarily because it's the most relevant result. Blog 3 introduced this concept when we discussed click position distribution and referenced the foundational work by <a href="https://www.cs.cornell.edu/people/tj/publications/joachims_etal_05a.pdf">Joachims et al. (2005)</a>.</p><p></p><p>Position bias matters for judgment lists because it means raw click data over-weights whatever the current ranking happens to surface first. If you train an LTR model on biased judgments, it learns to replicate the existing ranking, which defeats the purpose.</p><p></p><p>Two practical approaches to handle this:</p><p></p><ul><li><p><strong>Position-normalized click rates.</strong> Instead of raw click counts, calculate click-through rate <em>per position</em>. A document clicked 3 out of 10 times when shown at position 5 is arguably more relevant than one clicked 5 out of 10 times at position 1. Position 5 gets less visibility, so a higher click rate there is a stronger relevance signal.</p></li></ul><ul><li><p><strong>Skip-above heuristics.</strong> When a user clicks position 3 but skips positions 1 and 2, those skipped documents are implicitly judged "less relevant" for that query. This is the core insight from Joachims et al.: Skipped-above results provide negative training signals. You can extract these pairs from your click data:</p></li></ul>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
  AND attributes.search.query IS NOT NULL
| STATS
    min_click_position = MIN(attributes.search.result_click_position),
    max_click_position = MAX(attributes.search.result_click_position),
    click_count = COUNT(*)
  BY attributes.search.query_id, attributes.search.query
| WHERE min_click_position &gt; 1
| SORT click_count DESC<p>Queries where the minimum click position is greater than 1 are sessions where the user skipped the top result(s). The documents at those skipped positions, for those queries, are candidates for grade 0 in your judgment list, because the user saw them and chose something lower.</p><p></p><p>For production judgment list generation, the <a href="https://www.elastic.co/search-labs/blog/elasticsearch-learning-to-rank-introduction">LTR tutorial on Search Labs</a> walks through the complete pipeline. The <a href="https://www.elastic.co/search-labs/blog/training-learning-to-rank-models-elasticsearch-ubi-data">training LTR models with user behavior data</a> guide covers the specific workflow of deriving training data from click-through behavior, including the Clicks Over Expected Clicks (COEC) algorithm for debiasing. And the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/notebooks/search/08-learning-to-rank.ipynb">LTR Jupyter notebooks</a> provide working Python code for feature extraction and model training.</p><h2>Rank features for search relevance personalization</h2><p>Basic tuning applies the same ranking to every user. Personalization means adapting results based on who’s searching (their history, preferences, or segment). Rank features are the most accessible way to do this in Elasticsearch.</p><h3>Building behavioral signals from click streams</h3><p>You can aggregate your click and conversion data at the document level to produce signals that reflect overall or segment-specific popularity:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action IN ("click", "add_to_cart")
| STATS
    clicks = COUNT(CASE(attributes.search.action == "click", 1)),
    carts = COUNT(CASE(attributes.search.action == "add_to_cart", 1))
  BY attributes.search.result_click_id
| EVAL cart_rate_pct = ROUND(100.0 * carts / clicks, 1)
| SORT clicks DESC
| LIMIT 20<p>These document-level signals (click popularity, cart rate, and conversion rate) are indexed as <code>rank_feature</code> fields on each product document. The workflow:</p><p></p><ol><li><p>Run the ES|QL query above periodically (daily or weekly).</p></li><li><p>Write the results back to a <code>click_popularity</code> or <code>conversion_rate</code> <code>rank_feature</code> field on each product document.</p></li><li><p>Use a <code>rank_feature</code> query to blend text relevance with behavioral signals.</p></li></ol><p></p><p>The <code>rank_feature</code> query applies a saturation function by default: Initial popularity gains matter most, diminishing as values get large. This prevents a single viral product from dominating all queries. You can tune the function's pivot point to control how much influence the feature has relative to text relevance.</p><h3>Personalizing rank features by user segment</h3><p>The query above produces <em>global</em> popularity; that is, it’s the same for every user. Personalization comes from segmenting these signals. If you're tracking <code>enduser.pseudo.id</code> or <code>user.id</code>, you can compute features per user cohort:</p><p></p><ul><li><p><strong>Category affinity:</strong> How often does this user click products in "electronics" versus "clothing"?</p></li><li><p><strong>Price sensitivity:</strong> Does this user tend to click and convert on higher- or lower-priced items?</p></li><li><p><strong>Brand preference:</strong> Which brands does this user engage with most?</p></li></ul><p></p><p>These become additional rank features, applied at query time based on who's searching. The <a href="https://www.elastic.co/search-labs/blog/personalized-search-elasticsearch-ltr">personalized search with LTR</a> guide on Search Labs walks through training per-user ranking models. For a lighter approach, the <a href="https://www.elastic.co/search-labs/blog/ecommerce-search-relevance-cohort-aware-ranking-elasticsearch">cohort-aware ranking guide</a> shows how to use multiplicative boosting to personalize at the segment level without ML, just analytics-derived weights applied to rank features at query time.</p><h3>Combining multiple signals with the linear retriever</h3><p>When you're blending text relevance with semantic search and behavioral rank features, you need a way to combine them. Elasticsearch's <a href="https://www.elastic.co/search-labs/blog/linear-retriever-hybrid-search">linear retriever</a> gives you precise control over how different query types contribute to the final ranking. It computes a weighted sum of normalized scores, so you can say <em>text relevance matters 60%, semantic similarity 30%, popularity 10%</em> and adjust those weights based on your analytics. This is unlike Reciprocal Rank Fusion (RRF), which only considers relative rank positions.</p><p></p><p>This is particularly useful for personalization because you can vary the weights per user segment. A returning customer might get more weight on purchase history, while a first-time visitor gets more weight on global popularity.</p><p></p><p>If you're also using semantic search with embedding models, the <a href="https://www.elastic.co/search-labs/blog/jina-embeddings-v3-elastic-inference-service">Elastic Inference Service</a> (EIS) now offers GPU-accelerated embedding generation, including multilingual models, directly within Elastic Cloud, making it straightforward to add a semantic retriever alongside your text and behavioral signals.</p><h2>How Learn To Rank uses click streams to optimize ranking</h2><p>Rank features handle individual signals, but LTR handles all of them at once. It trains an ML model to combine features (like text relevance, popularity, CTR, recency, margin, or user affinity) into a single ranking function optimized for your users.</p><h3>What you need for Learn To Rank in Elasticsearch</h3><p>Elasticsearch has supported <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/learning-to-rank.html">native Learning to Rank since version 8.12</a> as an Enterprise subscription feature. The pipeline:</p><p></p><ol><li><p><strong>Judgment list.</strong> Query-document pairs with relevance grades (built from click data, as above).</p></li><li><p><strong>Feature extraction.</strong> Numeric signals for each query-document pair (BM25 score, popularity, CTR, recency, price, margin, user segment features).</p></li><li><p><strong>Model training.</strong> Typically XGBoost or LambdaMART, trained on your judgment list with features.</p></li><li><p><strong>Deployment.</strong> Upload the trained model to Elasticsearch via Eland, and use it as a rescorer.</p></li></ol><p></p><p>The click and conversion data from this series feeds steps 1 and 2, and the judgment list is your training labels. The document-level features from the rank features section above (click popularity, conversion rate, margin) are additional features alongside text relevance scores.</p><h3>Why automated judgment lists from click data scale better</h3><p>This is where the scalability argument for automated judgment lists becomes concrete. A manually curated judgment list might cover your top 100 queries well, but personalized ranking needs judgment data across thousands of queries and multiple user segments. You can't hire assessors to rate results for "wireless headphones" separately for electronics enthusiasts, budget shoppers, and professional audio engineers.</p><p></p><p>Automated judgment lists from click data scale to every query your users actually run and update as inventory and user behavior change. They can also be segmented by cohort. The <a href="https://www.elastic.co/search-labs/blog/training-learning-to-rank-models-elasticsearch-ubi-data">training LTR models with user behavior data</a> guide demonstrates this end-to-end workflow, showing how to go from raw click events to trained ranking models.</p><p></p><p>For teams looking to close this loop even further, the <a href="https://www.elastic.co/search-labs/blog/agentic-search-relevance-autotuning-elasticsearch">agentic autotuning approach</a> demonstrates using an AI agent to continuously monitor search quality and generate judgment lists from user interactions. It automatically retrains LTR models, turning the feedback loop into an autonomous system.</p><h3>What Learn To Rank requires: Traffic, pipeline, and evaluation</h3><p>LTR requires enough traffic to generate meaningful judgment lists and engineering time to build and maintain the pipeline. It also requires ongoing evaluation to ensure that the model improves over time. But for search applications with sufficient volume, it's the most effective way to make ranking learn from user behavior. The <a href="https://www.elastic.co/search-labs/blog/elasticsearch-learning-to-rank-introduction">LTR introduction on Search Labs</a> covers the full scope of what's involved.</p><h2>Elasticsearch Relevance Studio: Visual search relevance tuning</h2><p>Between query rules (manual, targeted) and LTR (ML, systemic), there's a middle ground: visual relevance tuning. <a href="https://elastic.github.io/relevance-studio/#/">Elasticsearch Relevance Studio</a> is a tool for comparing and tuning search configurations side by side. It lets you adjust boost values, field weights, and query structures while seeing the results update in real time.</p><p></p><p>Analytics data, especially CTR and MRR, tells you which queries to focus on. Start from a ranked list of problem queries (the low-CTR, high-volume queries from the ES|QL analysis above), and work through them systematically in Relevance Studio, instead of guessing which searches need tuning.</p><p></p><p>The workflow:</p><p></p><ol><li><p><strong>Identify problem queries.</strong> Run the CTR-by-query and MRR-by-query analyses.</p></li><li><p><strong>Open those queries in Relevance Studio.</strong> See current results alongside the tuned version.</p></li><li><p><strong>Adjust field weights and boosting.</strong> Experiment with configuration changes.</p></li><li><p><strong>Evaluate with judgment lists.</strong> Use the Rank Eval API to confirm that the change improves NDCG for your test queries.</p></li><li><p><strong>Measure the impact in production.</strong> Rerun the analytics after deploying changes, and compare CTR / MRR.</p></li></ol><p></p><p>This before-and-after loop is where analytics and tuning connect. With the data analytics you can prioritize the queries that affect the most users and verify that changes actually helped. Without this data, tuning is guesswork, and you're adjusting weights without knowing which queries matter or how to measure improvement. </p><p></p><p>If you are looking at tuning and evaluating relevance you should also investigate the <a href="https://elastic.github.io/relevance-studio/#/">Elasticsearch Relevance Studio</a> project which lets you directly compare search strategies. </p><h2>Closing the loop between search analytics and search relevance</h2><p>Search relevance improvement follows a consistent loop: measure quality with ES|QL analytics, apply changes (query rules, field weights, rank features or LTR models), then measure again.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbc895531c8a5ff31/6a7ef484af7924b0b425af35/image1.png" alt="Diagram showing the search analytics feedback loop from click streams through judgment lists to search relevance improvement" /><p>The key insight is that the same instrumentation that measures search quality also generates the data to improve it. Your search spans produce the queries to analyze, and your click spans produce the judgment data for evaluation and LTR training. Plus, your conversion spans tell you which improvements matter most to the business.</p><h3>Measuring search relevance improvements after changes</h3><p>After deploying a change, like a new query rule or updated field weights, or deploying an LTR model, you need to know whether it helped:</p><p></p><p><strong>Metric</strong></p><p><strong>What it measures</strong></p><p><strong>How to compare</strong></p><p>CTR trend</p><p>Whether more users are clicking results after the change</p><p>Run the CTR-by-query ES|QL query for the week before and after; compare per-query percentages</p><p>MRR trend</p><p>Whether clicks are shifting toward higher-ranked positions</p><p>Compare mean reciprocal rank per query before and after using the MRR-by-query analysis</p><p>Conversion rate trend</p><p>Whether more clicks are turning into purchases or cart additions</p><p>Run the click-to-conversion ES|QL query for both periods; compare cart rate percentages</p><p>Revenue per query</p><p>Whether the business impact of the change is positive</p><p>Compare revenue attributed to affected queries via conversion spans before and after deployment</p><p></p><p>Run the same ES|QL queries before and after. If your change was a query rule for "laptop bag", compare that query's CTR and MRR from the week before to the week after.</p><h3>A/B testing with experiment attributes</h3><p>The <code>feature_flag.key</code> attribute  is designed for exactly this. Route a percentage of traffic to a new ranking configuration, and set <code>feature_flag.key</code> to the experiment name and optionally <code>feature_flag.result.variant</code> to the variant (for example, <code>"control"</code> or <code>"variant-boost-v2"</code>). Then propagate it to click spans the same way you propagate `query_id`, and compare metrics between groups:</p>FROM traces-generic.otel-default
| WHERE ((name == "search" AND attributes.search.query IS NOT NULL)
    OR attributes.search.first_click == true)
  AND attributes.feature_flag.key IS NOT NULL
| STATS
    searches = COUNT(CASE(name == "search" AND attributes.search.query IS NOT NULL, 1)),
    clicked = COUNT(CASE(attributes.search.first_click == true, 1))
  BY attributes.feature_flag.key, attributes.feature_flag.result.variant
| EVAL ctr_pct = ROUND(100.0 * clicked / searches, 1)<p>This gives you an A/B comparison of CTR by experiment variant. You don’t need a separate experimentation platform for basic comparisons, although you’ll want a proper framework for statistical rigor on sample sizes and significance.</p><h3>Monitoring for search relevance regressions</h3><p>Once you've identified your high-value queries, that is, the ones that drive revenue and have been tuned for good engagement, you need to protect them. Relevance regressions on your top 20 revenue-generating queries are business-critical incidents.</p><p></p><p>This is where monitoring comes in. In the next blog in the series, we'll cover setting up alerts on these metrics: CTR drops on high-revenue queries and MRR regressions after deployments. We’ll also cover conversion rate anomalies. The same ES|QL queries that power your analytics dashboards can drive alerting rules; the feedback loop covers measurement and improvement. It also provides operational protection.</p><h2>Search relevance improvement roadmap: Week 1 to Quarter 1</h2><p>Here's a practical starting point. You don't need to implement LTR on day one. The approaches build on each other:</p><p></p><ul><li><p><strong>Week 1: Evaluate and fix known problems.</strong> Run the CTR-by-query analysis. Find your zero-CTR queries with high search volume, and fix the worst ones with synonyms or pinned results via query rules. You could also use adjusted field weights. Use judgment lists (even a small manual one for your top queries) to confirm that the changes improve NDCG before deploying. You should get immediate, targeted impact.</p></li></ul><p></p><ul><li><p><strong>Month 1: Add behavioral rank features.</strong> Extract document-level click popularity and conversion rates from your analytics. Index them as <code>rank_feature</code> fields, and blend behavioral signals with text relevance using the linear retriever or `rank_feature` queries. Then consider simple business boosts like margin. This lifts baseline quality across all queries without manual intervention per query.</p></li></ul><p></p><ul><li><p><strong>Quarter 1: Automate with LTR.</strong> Once you have enough click data (typically several weeks of production traffic), build judgment lists automatically from click and conversion data. Train an LTR model that combines text features, behavioral features, and business features, and then deploy it as a rescorer. The ranking now learns from user behavior and improves as you collect more data.</p></li></ul><p></p><p>At each stage, measure the impact with the same metrics. CTR and MRR are your scorecards, as is conversion rate. If a change doesn't move them, it didn't help, regardless of how sophisticated the approach.</p><h2>What's next: Search reliability engineering and SLOs</h2><p>We've covered the full sequence: instrument search, measure quality, track conversions, and improve relevance. The missing piece is making sure it all keeps working.</p><p></p><p>Next, we close the series with search reliability engineering, including Service Level Objectives (SLOs) for search quality and alerting on metric regressions, along with operational dashboards that catch problems before users notice them. The same metrics you've been building become the basis for search health monitoring, turning your analytics pipeline into a reliability system.</p><p></p><h2>Get started</h2><h3>Working code</h3><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">Reference project:</a> Working code for the entire blog series; clone, configure, and run.</p></li></ul><p></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/search-analytics-relevance-click-streams</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/search-analytics-relevance-click-streams</guid>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Relevance]]></category>
    <dc:creator><![CDATA[Matthew Adams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt259007788af667fe/6a7ee686ef5bef8db54fa878/image2.png" length="0" type="image/png"/>
    <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Skip the stateful OTel Collector: Elasticsearch 9.5 natively stores both metric temporalities]]></title>
    <description><![CDATA[Ingest cumulative and delta OpenTelemetry metrics under the same metric name while ES|QL and PromQL queries auto-detect temporality per series, with no new syntax or conversion pipelines required.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch 9.5 natively stores both cumulative and delta OpenTelemetry (OTel) counters and histograms, even when mixed for the same metric name. You ingest via OpenTelemetry Protocol (OTLP) and Elasticsearch preserves the temporality metadata automatically.<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"> ES|QL TS</a> and<a href="https://www.elastic.co/docs/reference/query-languages/promql"> PromQL</a> queries detect the temporality per series and interpret the data correctly, without new syntax, configuration changes to your OTel SDKs or stateful OTel Collector conversion. Existing queries and downsampled data continue to work as expected.</p><h2>What is metric temporality in OpenTelemetry?</h2><p>Metrics stores usually receive client-side, pre-aggregated metrics. For example, if an application records request response times, it won’t send each individual response time as a single data point to your metrics back end. Instead, the application (or rather the OTel SDK) pre-aggregates those raw response times into counters or histograms. These pre-aggregated values are then exported at a periodic interval, dramatically reducing the number of data points. <em>Temporality</em> is about how this pre-aggregation works. There are two temporality models: <em>cumulative</em> and <em>delta</em>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1829137fe767c17d/6a7edfc40da673bd3357cae8/image6.png" alt="Diagram showing how delta and cumulative temporality represent the same OTel counter metric data points differently" /><h3>Cumulative temporality in OTel metrics</h3><p>With <em>cumulative temporality</em>, each data point represents the totalamount of change in the metric value since the process started. Values monotonically increase, with occasional reset to 0 (for example, when the process restarts).</p><p>Take a counter tracking the total CPU time consumed by a Java Virtual Machine (JVM):</p><p><strong>Timestamp</strong></p><p><strong>Value</strong></p><p><strong>Meaning</strong></p><p>10:01</p><p>12.4s</p><p>12.4s total CPU time since start</p><p>10:02</p><p>13.1s</p><p>13.1s total CPU time since start</p><p>10:03</p><p>13.9s</p><p>13.9s total CPU time since start</p><p>To compute the rate of change between 10:01 and 10:02, we subtract: <code>13.1 - 12.4 = 0.7s</code> of CPU time was consumed in that interval. Dividing by the time range of the interval gives us the <code>rate</code>. This is the default temporality for counters in both Prometheus and OTel.</p><h3>Delta temporality in OTel metrics</h3><p>With <em>delta temporality</em>, each data point represents the change since the last measurement. Values are independent of each other. In other words, after each export, the OTel SDK resets all values for all series.</p><p>The same raw observations from the cumulative example above would look as follows with delta temporality.</p><p><strong>Timestamp</strong></p><p><strong>Value</strong></p><p><strong>Meaning</strong></p><p>10:01</p><p>0.5s</p><p>0.5s of CPU time in this interval</p><p>10:02</p><p>0.7s</p><p>0.7s of CPU time in this interval</p><p>10:03</p><p>0.8s</p><p>0.8s of CPU time in this interval</p><p>To compute the rate or increase, we can use the value directly, without any subtraction.</p><h3>Trade-offs between cumulative and delta OpenTelemetry metrics</h3><p>Both temporalities have practical trade-offs:</p><ul><li><p><strong>Resilience to data loss: </strong>Cumulative counters are self-describing: If you miss an export, the next data point still gives you the correct total. Delta values are incremental, so a lost data point means that the corresponding increase is lost.</p></li><li><p><strong>Metric producer memory footprint: </strong>For cumulative temporality, the OTel SDKs need to keep a state for every series in memory. For delta temporality, the footprint is much lower. There, the SDKs only need to keep track of counters or histograms which changed since the last export. If there are a lot of counters or histograms and many of them don’t increase each period, this difference can be quite substantial.</p></li><li><p><strong>Aggregation across restarts: </strong>Cumulative counters require reset detection logic, which in edge cases can fail: If the metric value decreases, it’s detected as a reset. We assume that the application was restarted and the counter started from 0 again. This can be missed if the first reported counter value after the restart is higher than before the restart. A concrete example:</p></li><ul><li><p>The service consumes 1 second CPU time and restarts.</p></li><li><p>After the restart, the service performs a CPU-intensive task and consumes 2 seconds of CPU time before the metric is exported again.</p></li><li><p>The metric back end just sees 1 followed by 2 as the metric value. It never observes a decrease and therefore misses the reset.</p></li></ul></ul><p>Delta values don't have this problem since each value is independent.</p><p>If you’re using histograms, the trade-offs have an even bigger effect:</p><p><strong>Trade-off</strong></p><p><strong>Cumulative</strong></p><p><strong>Delta</strong></p><p>Histogram size</p><p>Buckets accumulate across exports, consuming more storage</p><p>Buckets reset each export, producing smaller histograms</p><p>Min/max accuracy</p><p>Approximated from buckets for custom time ranges (tracked values represent extremes since process start)</p><p>Exact per-export minimum and maximum values</p><p>Query performance</p><p>Faster: only the first and last value in a time range plus resets are needed</p><p>Slower: all histograms in the queried range must be combined</p><p>OpenTelemetry supports both models and lets you choose per SDK via the <code>OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE</code> environment variable.</p><h2>Why native temporality support eliminates OTel Collector workarounds</h2><p><a href="https://prometheus.io/docs/concepts/metric_types/#counter">Prometheus</a> and most other metrics back ends pick a side: All metrics have to be either cumulative or delta. Elasticsearch previously followed that pattern, too, with native storage of cumulative counters and delta histograms, and workarounds for everything else. Delta counters were stored as gauges, functional but without native counter semantics for rate queries. And cumulative histograms were unsupported.</p><p>One workaround for unsupported temporalities is to configure your metric producers (for example, <a href="https://opentelemetry.io/docs/languages/">OTel SDKs</a>) to produce data with the temporality that your back end supports. In large-scale deployments, this can be a very challenging task. And sometimes this isn’t even possible (for example, if you consume OTLP metrics from third-party services).</p><p>Another workaround is to convert the temporality prior to ingestion. In the OTel Collector, you would typically use the <a href="https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/processor/cumulativetodeltaprocessor/README.md">cumulative-to-delta processor</a>, which comes with a big warning sign about <em>statefulness</em>. The conversion is inherently stateful, requiring ordered delivery of metric series to the same collector and persisted state across restarts. In practice, it works, but at scale, it comes with a lot of deployment headaches.</p><p>With Elasticsearch 9.5, you can skip the conversion pipeline entirely. Elasticsearch natively stores and queries metric data with both temporalities. It doesn’t require any stateful conversion required or explicit configuration of your OTel SDKs.</p><h2>Demo: ingesting cumulative and delta OTel metrics side by side</h2><p>To demonstrate the temporality support, we'll reuse a demo setup from our <a href="https://www.elastic.co/search-labs/blog/otel-histogram-metrics-esql">OTel histogram metrics ES|QL blog post</a>: a Java <a href="https://github.com/renaissance-benchmarks/renaissance">Renaissance</a> benchmark instrumented with the <a href="https://opentelemetry.io/docs/zero-code/java/agent/">OTel Java agent</a>. The twist this time: We run two instances of the benchmark, each configured with a different temporality:</p><ul><li><p><strong><code>renaissance-delta</code></strong><strong>: </strong>Exports metrics with delta temporality.</p></li><li><p><strong><code>renaissance-cumulative</code></strong><strong>: </strong>Exports metrics with cumulative temporality.</p></li></ul><p>Both instances report the same metrics under the same service name <code>renaissance</code>, but with different <code>service.instance.id</code> values. Here’s the relevant section of the <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/elasticsearch-temporality-demo/docker-compose.yml">docker-compose.yml</a> that can be found in <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/elasticsearch-temporality-demo">the companion code</a>:</p>renaissance-delta:
  environment:
    OTEL_SERVICE_NAME: renaissance
    OTEL_RESOURCE_ATTRIBUTES: "service.instance.id=delta-instance"
    OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE: delta
    OTEL_EXPORTER_OTLP_METRICS_DEFAULT_HISTOGRAM_AGGREGATION: BASE2_EXPONENTIAL_BUCKET_HISTOGRAM

renaissance-cumulative:
  environment:
    OTEL_SERVICE_NAME: renaissance
    OTEL_RESOURCE_ATTRIBUTES: "service.instance.id=cumulative-instance"
    OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE: cumulative
    OTEL_EXPORTER_OTLP_METRICS_DEFAULT_HISTOGRAM_AGGREGATION: BASE2_EXPONENTIAL_BUCKET_HISTOGRAM<p>To run the demo yourself, you'll also have to fill out the <a href="https://www.elastic.co/docs/reference/opentelemetry/managed-inputs/managed-otlp-endpoint">managed OTLP endpoint URL</a> and the corresponding API key:</p>OTEL_EXPORTER_OTLP_ENDPOINT: https://&lt;cluster-endpoint&gt;
OTEL_EXPORTER_OTLP_HEADERS: "Authorization=ApiKey &lt;base64 api key&gt;"<p>After starting the demo with <code>docker compose up --build</code>, both instances will start reporting metrics to Elasticsearch.</p><h3>Querying OTel counter metrics with ES|QL and PromQL</h3><p>Let's query the first few raw data points of <code>jvm.cpu.time</code> for both instances to see the different temporalities in action:</p><p>This gives us the first five data points for each service instance:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte236c69ee14c534f/6a7ee18f2888399d7307f09f/image2.png" alt="ES|QL query results showing raw cumulative and delta OTel metrics for jvm.cpu.time from two service instances" /><p>The benchmark consumes CPU at a nearly constant rate. This is directly visible based on the delta temporality data: The values are nearly constant between exports. In contrast, the cumulative temporality values grow over time, as they represent the total CPU usage of the benchmark instance.</p><p>Now let's have a look at how to properly query this metric using PromQL:</p>PROMQL sum by (service.instance.id) (rate(jvm.cpu.time))<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf7ed76cee486794e/6a7ee0250e035cf4d8b865a2/image4.png" alt="PromQL rate query showing CPU time per service instance with cumulative and delta OTel metrics overlaid" /><p>The screenshot shows that both benchmark instances consume a nearly constant of 1 to 1.2 number of CPU cores with some variance. This query works because we made our <code>rate</code> implementation respect the temporality: Every time series (so every service instance in our case) stores the temporality as a metric dimension. The <code>rate</code> implementation looks at this dimension and interprets the data accordingly: For delta temporality, values are summed up; for cumulative temporality, a difference computation is done. This all happens automatically in the background, without requiring any changes to your queries.</p><p>We’ve adapted <code>rate</code>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#increase"><code>increase</code></a>, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#irate"><code>irate</code></a> to work this way. The same applies when using those functions in ES|QL TS queries:</p><p>Because Elasticsearch tracks the temporality as a dimension, you can have multiple series with different temporalities for the same metric, just like in the demo use case. Aggregating across series also works as expected, because at that point <code>rate</code>, <code>increase</code>, or <code>irate</code> already took care of normalizing the data:</p>PROMQL sum(rate(jvm.cpu.time))<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6b8eb2e5ee6199bc/6a7ee06a28883935c907f097/image3.png" alt="PromQL chart showing total CPU time aggregated across both cumulative and delta OTel metrics instances" /><h3>Querying OTel histogram metrics across temporalities</h3><p>Metric temporality applies to histograms in the same way it applies to counters: histogram buckets are effectively a set of counters, each tracking values in a specific range.As in our histogram demo, we use exponential histograms, where bucket boundaries adapt automatically to minimize relative error.</p><p>Due to this similarity, histograms can also be cumulative or delta. Either the counter per bucket is reset after each metric export or the cumulative count carries over between exports.</p><p>Let's query the median major garbage collection (GC) duration for our benchmark instances, which is a histogram metric:</p>PROMQL histogram_quantile(0.5,  sum by (service.instance.id) (increase(jvm.gc.duration{jvm.gc.action=~".*major.*"})))<p>Or the equivalent ES|QL query:</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb02046819944c5ff/6a7ee0f6dcb4372a2b2d1232/image5.png" alt="Median major GC duration queried across cumulative and delta OpenTelemetry histogram metrics per instance" /><p>Again, both queries will automatically load the temporality per series and interpret the histograms accordingly. In PromQL, this is handled by the <code>increase</code> function. Note that in ES|QL, you don't explicitly call <code>increase</code> on histograms. The <code>TS</code> command automatically handles the temporality-aware merging of histograms when you use aggregation functions, like <code>PERCENTILE</code>, <code>MEDIAN</code>, or <code>AVG</code>.</p><h2>How Elasticsearch stores metric temporality in TSDB</h2><p>Elasticsearch's time series database (TSDB) stores metric temporality in a dedicated dimension field on each document. The <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/time-series#index-time-series-temporality-field"><code>index.time_series.temporality_field</code></a> index setting lets you specify which field carries the temporality information. The field must be a <code>keyword</code> field with <code>time_series_dimension: true</code> and the permissible values <code>"delta"</code> or <code>"cumulative"</code>.</p><p>As soon as this setting is present on a time series index, ES|QL and PromQL will load the corresponding field when performing temporality-dependent aggregations. If the field isn’t present or has no value on a document, we fall back to defaults based on the type of the corresponding metric: counters default to cumulative, and histograms default to delta. This matches the historical behavior and ensures existing queries and existing data continue to work without changes.</p><p>When you ingest metrics via the OTLP endpoint, Elasticsearch automatically adds a <code>temporality</code> dimension field to each document, populated from the <a href="https://opentelemetry.io/docs/specs/otel/metrics/data-model/#temporality">OTLP AggregationTemporality</a> metadata. For custom (neither OTLP nor <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write">Prometheus remote write</a>) ingestion, you’ll have to manually set up the <code>index.time_series.temporality_field</code> setting and populate your temporality dimension.</p><p>The temporality is also respected during downsampling: As it’s a dimension, it’s preserved automatically and used to compute the aggregate values.</p><h2>Getting started with mixed-temporality OTel metrics in Elasticsearch</h2><p>With Elasticsearch 9.5, cumulative versus delta is no longer a decision you have to get correct at the start. Ingest both temporalities side by side, even for the same metric name, and let ES|QL and PromQL handle the rest. You can switch between both without having to touch your queries. For more details, see the <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/metric-temporality">metric temporality documentation</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/otel-metrics-cumulative-delta-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/otel-metrics-cumulative-delta-elasticsearch</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Integrations]]></category>
    <dc:creator><![CDATA[Jonas Kunz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfebe6e53bb1ad4e4/6a7eded6b591027803eeca82/image1.png" length="0" type="image/png"/>
    <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Building context in Elasticsearch: how AI Indices power smarter agents using fewer tokens]]></title>
    <description><![CDATA[Store AI agent context in an AI Index and power smarter agents using fewer tokens. Step-by-step walkthrough with ES|QL and Kibana Workflows included.]]></description>
    <content:encoded><![CDATA[<p>Agents burn tokens exploring your data before they answer anything, inspecting mappings, sampling documents, probing which index to use. Elasticsearch AI Indices let you precompute that work once and store it as a Knowledge Indicator (KI): a structured, searchable record agents retrieve directly instead of rediscovering from scratch. This walkthrough shows you how to build the full pipeline: create an AI Index, generate routing KIs with a Kibana Workflow, and wire them to any agent harness via a portable ES|QL skill. We've also provided a <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/building-context-technical-walkthrough-part-1">notebook</a> if you'd like to run it yourself end to end as you go through the examples in this blog. This is Part 1 in a blog series providing a technical walkthrough to managing your context through KIs and AI indices. </p><p>While AI indices will be included in future Stack releases, today we recommend using Serverless.</p><h2>How it works: AI Index, Kibana Workflows, and the query-ki skill</h2><p>Building context in this walkthrough has three moving parts:</p><ol><li><p>An <strong>AI Index</strong>, where KIs live. It's a regular Elasticsearch index or data stream with a specific naming convention triggering component templates to configure the right mappings automatically.</p></li><li><p><strong>Kibana Workflows</strong>, which read from your data sources, run an LLM to structure content into KIs, and write those KIs into the AI Index.</p></li><li><p>A <strong><code>query-ki</code></strong><strong> skill</strong>, a skill that queries KIs directly from the AI Index using ES|QL, and that a chat agent can call as a tool.</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfb84f57dc630f26b/6a7b37268bcd80ea30262453/image4.png" alt="Architecture: Kibana Workflows write KIs to an AI Index, agents read AI agent context via a query-ki skill" /><p></p><h3>Prerequisites</h3><p>This tutorial assumes you have:</p><ol><li><p>An Elasticsearch Serverless project. You can <a href="https://cloud.elastic.co/registration?onboarding_token=search&amp;cta=cloud-registration&amp;tech=trial&amp;plcmt=article%20content&amp;pg=search-labs">sign up for a trial</a> if you don't have one.</p></li><li><p>An API key to access your Elasticsearch project. </p></li></ol><h2>Create sample indices for agent routing</h2><p>First, we’ll need some sources. Sources can be data that already exists in your Elasticsearch indices, or external data accessed via connectors or ES|QL data sources. </p><p>For this blog, we’ll create some indices with example data. We’ll start with an example using three datasets: <a href="https://huggingface.co/datasets/BeIR/fiqa">BEIR/fiqa</a> (financial), <a href="https://huggingface.co/datasets/BeIR/nfcorpus">beir-nfcorpus</a> (biomedical/nutrition), and <a href="https://huggingface.co/datasets/BeIR/scifact">beir-scifact</a> (scientific fact-checking). Each index is populated with its own <code>_meta.description</code>. </p><p>Here are the mappings we define for these indices: </p>{
  "beir-fiqa": {
    "mappings": {
      "_meta": {
        "description": "FiQA: financial question answering corpus from StackExchange Finance community posts and web crawls. Covers investments, banking, taxes, and market analysis. BM25-only index."
      },
      "properties": {
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document body text."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document or article title."
          }
        }
      }
    }
  }
}


{
  "beir-nfcorpus": {
    "mappings": {
      "_meta": {
        "description": "NFCorpus: biomedical information retrieval corpus from NutritionFacts.org. Contains nutrition science and medical research documents on diet, disease, and health interventions. BM25-only index."
      },
      "properties": {
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document body text."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document or article title."
          }
        }
      }
    }
  }
}


{
  "beir-scifact": {
    "mappings": {
      "_meta": {
        "description": "SciFact: scientific fact-checking corpus of biomedical research abstracts used to verify factual claims in peer-reviewed literature. BM25-only index."
      },
      "properties": {
        "text": {
          "type": "text",
          "meta": {
            "description": "Full document body text."
          }
        },
        "title": {
          "type": "text",
          "meta": {
            "description": "Document or article title."
          }
        }
      }
    }
  }
}<p>Then, using the above convenience scripts, load a handful of documents into each index with the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk"><code>_bulk</code></a> API.</p><p>Now imagine an agent with a question and the indices we’ve just created. The agent has no idea which one is relevant at the start. Without pre-computed context, it either performs exploratory lookups (mappings, test searches) to figure out which source to use, or searches all three and hopes the merged results contain something useful. Either approach costs tokens, and if you multiply that inefficiency across every query an agent makes, it adds up.</p><h2>Create your AI Index</h2><p>Before generating any KIs, you need an index to store them. We call this an <strong>AI Index</strong>.</p><p>The naming convention is what triggers automatic configuration. Any index whose name starts with <code>ai-index-idx-</code> is a regular index; <code>ai-index-ds-</code> is a data stream. You’ll want to choose data streams for observability use cases, time series data, and when recency is important. Conversely, standard indices are a good choice for static data that will exist for a long while, where recency is not as much of a concern, and may need to occasionally be updated on demand. This naming convention is required for AI indices. </p><p>When Elasticsearch sees the <code>ai-index-</code> prefixes, it automatically applies component templates that configure the right mappings and settings.</p><p>Creating an AI Index is a single call:</p>PUT ai-index-idx-my-corpus<p>To see exactly what the component templates applied, inspect the mappings:</p>GET ai-index-idx-my-corpus/_mapping<p>The response shows the fields every AI Index gets out of the box:</p>{
  "ai-index-idx-my-corpus": {
    "mappings": {
      "properties": {
        "@timestamp": {
          "type": "date"
        },
        "attributes": {
          "type": "flattened"
        },
        "content": {
          "type": "text",
          "fields": {
            "semantic": {
              "type": "semantic_text",
              "inference_id": ".jina-embeddings-v5-text-small"
            }
          }
        },
        "description": {
          "type": "text",
          "fields": {
            "semantic": {
              "type": "semantic_text",
              "inference_id": ".jina-embeddings-v5-text-small"
            }
          }
        },
        "references": {
          "properties": {
            "uri": {
              "type": "keyword"
            }
          }
        },
        "tags": {
          "type": "keyword"
        },
        "title": {
          "type": "text",
          "fields": {
            "semantic": {
              "type": "semantic_text",
              "inference_id": ".jina-embeddings-v5-text-small"
            }
          }
        },
        "type": {
          "type": "keyword"
        }
      }
    }
  }
}<p><code>title</code>, <code>description</code>, and <code>content</code> are each a <code>text</code> field with a <code>.semantic</code> sub-field of type <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">semantic_text</a>, supporting hybrid retrieval.</p><p>Data stream indices (<code>ai-index-ds-*</code>) additionally carry a default 90-day data retention policy. This blog uses a standard index (<code>ai-index-idx-*</code>).</p><h2>Index Metadata as a Knowledge Indicator</h2><p>The target use case for this example is how the <code>query-index-metadata-ki</code> skill can route an agent to the correct Elasticsearch index, even when index or field names are vague. This reduces mistakes from choosing the wrong index or formulating queries based on incomplete schema exploration.</p><p>Since we're creating KIs for our own indices, we can give the LLM a head start: annotate index mappings with human-written <code>_meta.description</code> content. The workflow generates better KIs with more context to work from.</p><p>To address this, we'll manually create a <a href="https://www.elastic.co/docs/reference/kibana">Kibana Workflow</a> that profiles each index and writes routing KIs into the AI Index. The workflow chains four steps:</p><p></p><p><strong>Step</strong></p><p><strong>Type</strong></p><p><strong>What it does</strong></p><p><code>get_mapping</code></p><p><code>elasticsearch.request</code></p><p>Read the mapping, including <code>_meta.description</code> and per-field descriptions.</p><p><code>sample_docs</code></p><p><code>elasticsearch.search</code></p><p>Pull a few real documents so the profile reflects actual value shapes.</p><p><code>profile_index</code></p><p><code>ai.agent</code></p><p>Generate a structured index profile as structured output.</p><p><code>sink_index_ki</code></p><p><code>elasticsearch.bulk</code></p><p>Write the profile into the AI Index as a KI.</p><p>Paste the following YAML into the <a href="https://www.elastic.co/docs/explore-analyze/workflows">Workflows</a> editor:</p>version: '1'
name: beir-index-profile-ki
description: Profile an index into an index-selection Knowledge Indicator.
enabled: true
tags:
  - context-management
  - index-selection

triggers:
  - type: manual

consts:
  indices:
    - beir-fiqa
    - beir-nfcorpus
    - beir-scifact

steps:
  - name: loop_indices
    type: foreach
    foreach: '{{ consts.indices | json }}'
    iteration-on-failure:
      continue: true
    steps:
      - name: get_mapping
        type: elasticsearch.request
        with:
          method: GET
          path: '/{{ foreach.item }}/_mapping'

      - name: sample_docs
        type: elasticsearch.search
        with:
          index: '{{ foreach.item }}'
          size: 3
          query:
            match_all: {}

      - name: profile_index
        type: ai.agent
        timeout: 120s
        with:
          message: &gt;
            You are a data steward building an INDEX PROFILE for an enterprise
            data catalog. Downstream, an AI agent uses these profiles to decide
            WHICH Elasticsearch index to query for a given user question -- this
            is an index-SELECTION aid, not a place to answer the question itself.

            You are given (a) the index name, (b) its Elasticsearch mapping
            including human-written descriptions in `_meta.description` and each
            field's `meta.description`, and (c) a few sample documents. Produce a
            faithful, decision-useful profile. Rules:
            - Ground everything in the provided mapping + samples. Never invent
              fields, values, or purpose. If unknown, use an empty string/array.
            - Optimize for routing: make it obvious what kinds of questions this
              index can authoritatively answer, and what it canNOT.
            - Prefer concrete field names and real example values from the
              samples over vague phrasing.
            - For joins, surface shared keys (e.g. *_id fields) that link this
              index to sibling indices, since cross-index questions hinge on them.

            Index name: {{ foreach.item }}

            Elasticsearch mapping (JSON):
            {{ steps.get_mapping.output | json }}

            Sample documents (JSON):
            {{ steps.sample_docs.output.hits.hits | map: '_source' | json }}
          schema:
            type: object
            properties:
              display_name:
                type: string
                description: A concise human-readable name for what this index represents (&lt;= 8 words).
              purpose:
                type: string
                description: 2-4 sentences describing what this index stores and its role. PRIMARY semantic surface for matching a question to this index.
              answers_questions:
                type: array
                items:
                  type: string
                description: 3-7 representative natural-language questions this index can authoritatively answer.
              does_not_contain:
                type: array
                items:
                  type: string
                description: 1-4 things a searcher might wrongly expect here but that live elsewhere, to prevent mis-routing.
              key_fields:
                type: array
                items:
                  type: string
                description: 3-10 of the most query-relevant fields as "field_name - what it is".
              when_to_use:
                type: string
                description: A single crisp routing heuristic - when should an agent pick THIS index? (&lt;= 30 words).
              example_esql:
                type: string
                description: One realistic, runnable ES|QL query against this index answering one of answers_questions.
            required:
              - display_name
              - purpose
              - answers_questions
              - key_fields
              - when_to_use

      - name: sink_index_ki
        type: elasticsearch.request
        with:
          method: PUT
          path: '/ai-index-idx-my-corpus/_doc/{{ foreach.item | url_encode }}'
          body:
            '@timestamp': '{{ "now" | date: "%Y-%m-%dT%H:%M:%S.%LZ" }}'
            type: index_metadata_entry
            title: '{{ steps.profile_index.output.structured_output.display_name | default: foreach.item }}'
            tags:
              - index-profile
              - '{{ foreach.item }}'
            attributes:
              display_name: '{{ steps.profile_index.output.structured_output.display_name }}'
              purpose: '{{ steps.profile_index.output.structured_output.purpose }}'
              when_to_use: '{{ steps.profile_index.output.structured_output.when_to_use }}'
              answers_questions: '{{ steps.profile_index.output.structured_output.answers_questions | json }}'
              does_not_contain: '{{ steps.profile_index.output.structured_output.does_not_contain | json }}'
              key_fields: '{{ steps.profile_index.output.structured_output.key_fields | json }}'
              example_esql: '{{ steps.profile_index.output.structured_output.example_esql }}'
              source_index: '{{ foreach.item }}'
            content: &gt;
              === SOURCE / PROVENANCE ===
              This is an INDEX PROFILE for routing/index-selection.
              Backing Elasticsearch index: {{ foreach.item }}
              Inspect it directly with ES|QL:
              FROM {{ foreach.item }} | LIMIT 10
              === WHAT THIS INDEX IS ===
              {{ steps.profile_index.output.structured_output.purpose }}
              Questions this index can answer: {{ steps.profile_index.output.structured_output.answers_questions | join: " | " }}
              When to use this index: {{ steps.profile_index.output.structured_output.when_to_use }}
              Example query:
              {{ steps.profile_index.output.structured_output.example_esql }}
            description: &gt;
              Index profile: {{ steps.profile_index.output.structured_output.display_name }}.
              Does NOT contain: {{ steps.profile_index.output.structured_output.does_not_contain | join: "; " }}.
              Key fields: {{ steps.profile_index.output.structured_output.key_fields | join: "; " }}.<p>Let's walk through what this workflow does. We loop over three specified indices with a <code>foreach</code> loop. For each:</p><ol><li><p><code>get_mapping</code> fetches the Elasticsearch index mappings, including any <code>_meta.description</code> annotations we added earlier.</p></li><li><p><code>sample_docs</code> pulls 3 real documents. Concrete examples give the LLM much better signal than schema alone.</p></li><li><p><code>profile_index</code> calls <code>ai.agent</code> with the index name, mappings, and sample documents. The LLM returns structured output describing the index's purpose, key fields, and an example ES|QL query showing how to use it.</p></li><li><p><code>sink_index_ki</code> writes the result into the AI Index as a KI of type <code>index_metadata_entry</code>, keyed on the index name so re-runs are idempotent.</p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfe75c3b235a08332/6a7b387e28883919bf07de96/image1.png" alt="Kibana Workflow generating Knowledge Indicators: get_mapping, sample_docs, profile_index, and sink to AI Index" /><p>A few things to point out: </p><ul><li><p>This workflow hard-codes a specific set of indices. In practice, you could derive the list from an index pattern or a dynamic source. </p></li><li><p>The <code>foreach</code> loop also runs iterations sequentially, which is fine for this guide but slow in production because each iteration involves an LLM call. For scale, use <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/composition">workflow.executeAsync</a> or native parallel support. The <a href="https://www.elastic.co/docs/explore-analyze/workflows/reference/cheat-sheet">cheat sheet</a> has tips on both.</p></li><li><p>In the <code>profile_index</code> step, the agent prompt is the special sauce. This is what shapes the accuracy and usefulness of the KIs. </p></li><li><p>Using <a href="https://www.elastic.co/docs/explore-analyze/workflows/steps/ai-steps#ai-prompt"><code>ai.prompt</code></a> can improve workflow efficiency (and cost) if you don’t need to load other tools. </p></li><li><p>Cost can be controlled in multiple ways. Richer prompts and structured output often result in higher token utilization, and of course the model you choose significantly impacts total costs. The <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service</a> (EIS) can be a great playground to test different models against the <code>profile_index</code>’s <code>ai.agent</code> step to compare how different models stack up against each other when generating KIs. </p></li></ul><h3>Query your AI Index to verify Knowledge Indicators</h3><p>Once the beir-index-profile-ki workflow runs, query the AI Index directly in the Discover tab using the following ES|QL query to confirm what got written:</p><p></p><p>This will result in the following output: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7a497f27ccbb343c/6a7b38b81f7b5adf6678fdf1/image2.png" alt="ES|QL query results from an AI Index showing three index metadata Knowledge Indicators in Kibana Discover" /><h3>Build a portable skill to retrieve AI agent context</h3><p>Retrieval is a critical component in an AI index. A KI is a document in the AI Index, and finding one is a single ES|QL query. We package that query as a small, portable skill so any agent can call it, regardless of the harness it runs in.</p><p>We write the skill as a SKILL.md: a YAML header with a name and description, followed by markdown instructions. This is the same Agent Skills format that many harnesses, including Claude Code, LangChain's Deep Agents, and others, load directly. </p><p>The harness reads the header content up front, and only pulls in the full instructions when a question matches the description. The one thing the skill asks of the harness is a way to run ES|QL against Elasticsearch.</p><p>Here is a sample <code>query-index-metadata-ki</code> skill: </p>---
name: query-index-metadata-ki
description: &gt;-
  Retrieve Knowledge Indicators (pre-computed context) from the Elasticsearch AI
  Index before answering. Use it to find which index to search (routing profiles).
  Trigger on any question that depends on choosing a data source.
allowed-tools: esql_query
---

# Retrieving Knowledge Indicators

Knowledge Indicators (KIs) live in Elasticsearch indices named `ai-index-*`.
Retrieve them by calling the `esql_query` tool with the query below. Substitute
the user's question for `&lt;query&gt;`, and `index_metadata_entry` as the `&lt;ki_type&gt;` for routing profiles.

```esql
FROM ai-index-idx-* METADATA _id, _index, _score
| WHERE type == "&lt;ki_type&gt;"
| FORK
    (WHERE MATCH(content, "&lt;query&gt;") OR MATCH(description, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
    (WHERE MATCH(content.semantic, "&lt;query&gt;") OR MATCH(description.semantic, "&lt;query&gt;")
     | SORT _score DESC | LIMIT 20)
| FUSE
| SORT _score DESC
| KEEP title, content, description, tags
| LIMIT 5
```

Ground your answer in what the query returns, and cite the KI titles you used. If
nothing relevant comes back, say so rather than guessing.<p>Let’s break down what the skill is doing: </p><ul><li><p>We’re defining <code>index-metadata-entry</code> as a KI type/use case.</p></li><li><p>We’re performing a hybrid ES|QL search on our AI indices, filtering by the appropriate <code>type</code> using RRF as the default method to fuse results.</p></li><li><p>The KI results will directly ground the agent’s answer when determining what indices are relevant to the query.</p></li></ul><p>Because the skill is just instructions plus a query, it travels wherever your agent does. You can point the same file at a Kibana Workflow agent, Claude Code, LangChain Deep Agents, or any other harness without changing a line of it.</p><h3>Connect your AI Index to an agent harness </h3><p>We want to demonstrate how you can use AI indices to query your data with any harness. For these examples, we’ll use LangChain Deep Agents and an OpenAI-compatible key, but any other agent harness can be easily substituted in, including Elastic Agent Builder.</p><p>First, let’s create a baseline to see how an agent will perform without using KIs:</p># Example question: Is there scientific evidence that vitamin D supplementation prevents cancer?
import os
import sys
import time
from elasticsearch import Elasticsearch
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM beir-fiqa | LIMIT 5'.
               Full-text search syntax: WHERE MATCH(field, "value") — not field MATCH "value".
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


@tool
def get_mapping(index: str) -&gt; dict:
    """Return the field mapping for an Elasticsearch index or pattern."""
    return es.indices.get_mapping(index=index).body


baseline_agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query, get_mapping],
    system_prompt=(
        "You are a research assistant with access to three Elasticsearch indices: "
        "beir-fiqa, beir-nfcorpus, and beir-scifact. "
        "You do NOT know which index is relevant for a given question. "
        "Use get_mapping to inspect an index's description and fields, "
        "then query the most relevant one with esql_query. "
        "Ground your answer strictly in what the queries return."
    ),
)

start = time.perf_counter()
result = baseline_agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>Here’s a modified example that could run the same agent, but now with the ability to search AI indices to return KIs: </p># Example question: Is there scientific evidence that vitamin D supplementation prevents cancer?
import os
import sys
import time
from elasticsearch import Elasticsearch
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from deepagents import create_deep_agent
from deepagents.backends.filesystem import FilesystemBackend

if len(sys.argv) &lt; 2:
    sys.exit(f'Usage: python {sys.argv[0]} "your question"')

es = Elasticsearch(os.environ["ES_URL"], api_key=os.environ["ES_API_KEY"])


@tool
def esql_query(query: str) -&gt; list[dict] | str:
    """Execute an ES|QL query against Elasticsearch and return the matching rows.

    Args:
        query: A complete ES|QL query string, e.g. 'FROM beir-fiqa | LIMIT 5'.
               Full-text search syntax: WHERE MATCH(field, "value") — not field MATCH "value".
    """
    try:
        resp = es.esql.query(query=query, format="json")
        cols = [c["name"] for c in resp["columns"]]
        return [dict(zip(cols, row)) for row in resp["values"]]
    except Exception as e:
        return f"ES|QL error: {e}"


backend = FilesystemBackend(root_dir=".", virtual_mode=False)

agent = create_deep_agent(
    model=ChatOpenAI(  # any OpenAI-compatible endpoint; configure via LLM_* env vars
        base_url=os.environ.get("LLM_BASE_URL", "https://openrouter.ai/api/v1"),
        model=os.environ.get("LLM_MODEL", "anthropic/claude-sonnet-4.5"),
        api_key=os.environ["LLM_API_KEY"],
    ),
    tools=[esql_query],
    skills=["skills"],
    backend=backend,
    system_prompt=(
        "You are a research assistant with access to several Elasticsearch indices. "
        "You do NOT know which index is relevant for a given question. "
        "Before searching, always use the query-ki skill with type 'index_metadata_entry' "
        "to retrieve the routing profile for the right index, then query that index directly. "
        "Ground your answer strictly in what the queries return and cite the KI you used for routing."
    ),
)

start = time.perf_counter()
result = agent.invoke(
    {
        "messages": [
            {
                "role": "user",
                "content": sys.argv[1],
            }
        ]
    }
)
latency = time.perf_counter() - start

print("\n--- Tool calls ---")
for m in result["messages"]:
    if isinstance(m, AIMessage) and m.tool_calls:
        for tc in m.tool_calls:
            print(f"  [{tc['name']}] {str(tc['args'])[:120]}")
total = sum(
    len(m.tool_calls)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.tool_calls
)
print(f"Total: {total}\n")

print("--- Usage ---")
input_tokens = sum(
    (m.usage_metadata or {}).get("input_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
output_tokens = sum(
    (m.usage_metadata or {}).get("output_tokens", 0)
    for m in result["messages"]
    if isinstance(m, AIMessage) and m.usage_metadata
)
print(f"Tokens: {input_tokens + output_tokens} (input {input_tokens}, output {output_tokens})")
print(f"Latency: {latency:.2f}s\n")

print("--- Answer ---")
print(result["messages"][-1].content)<p>This agent will always query the KI indices to get the answer.</p><h2>How much do Knowledge Indicators reduce agent token usage?</h2><p>Since we’re using agents, the results of these scripts are non-deterministic. However, when I ran these results against the query <code>Is there scientific evidence that vitamin D supplementation prevents cancer?</code>, both agents led to the same conclusion, but they took different paths to get there: </p><p>
</p><p>Baseline (No AI Index)</p><p>With AI Index</p><p>Total tool calls</p><p>12</p><p>8</p><p><code>read_file</code> calls</p><p>0</p><p>2</p><p><code>get_mapping</code> calls</p><p>3</p><p>0</p><p><code>esql_query</code> calls</p><p>9</p><p>6 </p><p>Total indices queried</p><p>2 (bounced between <code>beir-scifact</code> and <code>beir-nfcorpus</code>)</p><p>1 (<code>beir-nfcorpus</code>)</p><p>Tokens consumed</p><p>167,763</p><p>92,711</p><p>Latency</p><p>39.58s</p><p>36.15s</p><p>Answer</p><p>Grounded, correct</p><p>Grounded, correct</p><p>The KI answers were both grounded and correct, but an interesting datapoint is the fact that the overall tool usage and token utilization was smaller when using KIs (latency was roughly equivalent). Here’s how both paths went, side by side: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcce26759ded49327/6a7b3983288839683507dea0/image5.png" alt="Agent comparison: 12 tool calls without KIs vs 8 with KIs, 45% fewer tokens, same grounded answer" /><h2>Run the full AI Index pipeline in Serverless</h2><p>This walkthrough offered a deep dive into the build-it-yourself version of AI indices and KIs. In production, you wouldn't hand-write these workflows; a setup agent would generate them, and a feedback loop would refine KIs from the agent's own traces. But the primitives are exactly what you just used: extract KIs with a workflow, store them in an AI Index, and retrieve them with a skill.</p><p>Managing context is key to a relevant and efficient agentic search system, and AI indices are a way to manage this context with the full power of the Elastic stack. Try it out in Serverless and let us know what you think in our <a href="https://discuss.elastic.co/top?period=monthly">Discuss forums</a> or the <code>#stack-kibana</code> channel in our <a href="https://elasticstack.slack.com/signup#/domain-signup">Community Slack</a>! </p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-index-building-context-agents</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-index-building-context-agents</guid>
    <category><![CDATA[AI]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Kathleen DeRusso,Matt Nowzari ,Apostolos Matsagkas,Peter Pisljar]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt82d8e9495cde2e30/6a7b37020c5aa95ac5f88c47/image3.png" length="0" type="image/png"/>
    <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch as one platform: What a second data system really costs]]></title>
    <description><![CDATA[Running search, analytics, metrics, logs, and vector retrieval in five systems costs more than five licenses. Here's what one platform looks like in practice.]]></description>
    <content:encoded><![CDATA[<p>If you count the data engines in your stack, you’ll find that full-text search runs in one system, while analytics runs in a warehouse. You’ll also see that metrics live in a time series database and logs are in an aggregator. And vector retrieval sits in its own dedicated vector store. That’s five engines for five shapes of the same operational data, and the licenses are the cheapest part.</p><p>We wrote about the architecture behind consolidating those shapes in <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage">Why Elasticsearch is becoming a columnar database</a>. This post is about the other side of the ledger. What does the split actually cost, and what does "search and analytics on one platform" mean in terms clear enough to hold up in a proof of concept?</p><h2>What running five data engines actually costs</h2><p>Anyone who has kept two sets of books for the same business knows where the hours go. Writing the second ledger is quick, but making the two agree is what takes the week.</p><p>Each engine needs its own ingest path, so the same events get parsed and shipped twice. That gives you two sets of failure modes and two backlogs to drain when a broker slows down. It also gives you drift: A field rename lands in one copy before the other, and for a while, the two systems disagree about the same hour of data. Reconciling that disagreement is real engineering work that rarely appears in the business case.</p><p>Each engine also brings a query language, and the syntax is the small part. The cost is everything written in that language, including dashboards, alert rules, saved queries, runbooks, and the operational knowledge of the person on call this week. Two languages means two of all of it, maintained in parallel.</p><p>Then there’s correlation. You find the failing request in the log aggregator. You move to the warehouse to chart how often it happened this week and then to the metrics store to check whether the host was saturated at the time. Each of those moves is a join performed by hand, by a person under time pressure, and every one of them adds minutes to the incident.</p><p>Retention compounds all of this. Each system gets its own lifecycle policy, so the cheap system ends up keeping data that the expensive one dropped weeks ago. When you finally need the history, it lives in the engine that cannot answer your question quickly.</p><h2>Why search and analytics were split across two systems</h2><p>The split was a reasonable response to a real constraint. Document engines and columnar engines were built to answer different questions. A <em>document engine</em> is good at finding the records that match a query and ranking them by relevance, and a <em>columnar engine</em> is good at reading three columns out of 50 and aggregating them across billions of rows. For years, running both was the reasonable answer, because no single system was credible at both jobs. The introduction of a full columnar engine in Elasticsearch 9.5 makes the split optional rather than necessary. For the full history and what changed, check out the <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage">columnar post</a>.</p><h2>What one platform means in practice</h2><p>Three things have to be true before "one platform" means anything.</p><ol><li><p>The first is one <em>query language</em>. Elasticsearch Query Language (ES|QL) runs across logs, metrics, traces, security events, and documents, which means a single skill set and one set of dashboards. Because it’s one language rather than a federation of several, queries compose: A time series aggregation can sit in the same query as <code>LOOKUP JOIN</code> or <code>INLINE STATS</code>, which systems built around PromQL alone cannot do.</p></li><li><p>The second is one <em>storage substrate</em>. Doc values, the column store that has been inside Elasticsearch since 2013, is what every index mode reads and writes underneath. Various modes tune that substrate for different shapes of data. Elasticsearch's time series engine (TSDB) went fully columnar in 9.4, which is the clearest evidence so far that the approach works. Columnar Mode and Columnar Logs, both in technical preview in Elasticsearch 9.5, extend the same treatment to analytical and log-shaped data.</p></li><li><p>The third is one <em>operational story</em>. It’s the same cluster, with the same APIs, integrations, agents, access control, backup, and upgrade path that you already run.</p></li></ol><p>Several index modes share one column store, and one language queries all of them inside a single cluster. That’s a narrower claim than "one engine for everything," and it’s the one that holds up when someone tests it.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc2d95f952dcd1adb/6a79ad4acbb9ed60cd53e4ce/image1.png" alt="" /><h2>When is a specialized database still the right choice?</h2><p>A specialized system will always win a specialized benchmark. If your workload is one shape with a single query pattern, there’s a purpose-built engine that beats us on it, and we would rather say so than pretend otherwise. The argument for consolidation applies to data that spans more than one shape, which describes most production estates.</p><p>That said, being the general-purpose platform doesn’t mean settling for second place on every shape. We’ve climbed this curve once already. TSDB has stored metrics since Elasticsearch 8.7, and the early work concentrated on storage efficiency rather than on competing with dedicated metrics stores.</p><p>A run of releases from Elasticsearch 9.1 through 9.4 then turned it into a columnar metrics engine: OpenTelemetry (OTel) metrics now land at 3.75 bytes per data point, down from 25 a year earlier, which is 2.5x less storage than Prometheus and 2x less than ClickHouse. Gauge average and counter rate queries run up to 30x faster than Prometheus and Mimir, and on the high cardinality benchmark, Elasticsearch scans four hours of data across half a million time series in under two seconds, where the other systems needed more than 30 seconds. The full methodology and per-query results are in <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">how we rebuilt Elasticsearch as a leading columnar metrics datastore</a>.</p><p>Columnar Mode is at the start of that same curve. Technical preview lands in Elasticsearch 9.5 and general availability (GA) in 9.6, with phased improvements to storage, ingest, and query performance in the releases that follow. Metrics took a year of that work to get where they are, and we expect logs and analytical data to follow the same path rather than a shorter one.</p><h2>Which workloads should stay on the document modes</h2><p></p><p>Some workloads should stay exactly where they are. </p><ol><li><p>Search-first applications, like product catalogs and knowledge bases, where the answer is the 10 most relevant documents, are what the existing document modes do well, and none of those modes are deprecated. </p></li><li><p>Frequent individual document updates - these workloads suit the document modes, too. </p></li><li><p>Nested data models - The same is true of data models that genuinely depend on nested structure, because Columnar Mode flattens fields into key/value pairs and doesn’t support the nested field type. In Elasticsearch 9.5, <code>semantic_text</code> and <code>dense_vector</code> fields aren’t available in Columnar Mode; a columnar profile for vector retrieval comes later.</p></li></ol><p>Adoption is per index and opt-in. Existing indices continue to behave as they do today, and your APIs don’t change. Plus, your dashboards don’t break.</p><h2>Where Columnar Mode is today: Elasticsearch 9.5 preview, 9.6 GA</h2><p>Storage and performance numbers are coming in a separate technical deep dive, and the public roadmap issues for <a href="https://github.com/elastic/roadmap/issues/290">Columnar Mode</a> and <a href="https://github.com/elastic/roadmap/issues/291">Columnar Logs</a> are the right place to tell us what your workload needs.</p><p>If you want a real number for your own five-tool tax, start by counting two things: ingest pipelines carrying the same events to more than one destination, and dashboards that answer the same question in two different query languages. In most estates, the second number is the one that surprises people.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/data-platform-consolidation-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/data-platform-consolidation-elasticsearch</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Columnar]]></category>
    <dc:creator><![CDATA[Yannis Roussos,Bharath Aleti]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4696c73e3407b15b/6a79ae477724f2098d0460a0/good_oone.png" length="0" type="image/png"/>
    <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[One ES|QL query instead of two: WHERE IN subquery replaces the copy-paste loop in Elasticsearch]]></title>
    <description><![CDATA[ES|QL's WHERE clause can filter by another Elasticsearch subquery's results instead of a static ID list you copied by hand, with nested subqueries, NOT IN and compound conditions built in.]]></description>
    <content:encoded><![CDATA[<p>The <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/where"><code>WHERE</code></a> clause can now filter by the results of another query. If you've been running one query to find suspicious users or failing services, copying the IDs, then pasting them into a second query, you can stop. One ES|QL statement does the whole job: the subquery builds the filter list from live data, and it stays current every time you run it. The feature ships as a technical preview in Elasticsearch 9.5 and supports nesting, <code>NOT IN</code> and compound <code>AND</code><code>/</code><code>OR</code> conditions.</p>  The <code>WHERE IN</code> subquery may change or be removed in a future release. Elastic will work to fix any issues, but technical preview features aren’t subject to the support Service Level Agreement (SLA) of official general availability (GA) features.<h2>Static ID lists vs. dynamic filtering with ES|QL's WHERE clause</h2><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt112e1ace022eb1a1/6a75c6d0b966e18d6563ef05/image1.png" alt="Infographic comparing manual ID copying across queries vs. ES|QL WHERE IN subquery dynamic filtering in a single pipeline" /><p></p><p><strong>Static ID list filtering</strong></p><p><strong>Dynamic filtering with WHERE IN subquery</strong></p><p>Run query A to find the IDs you care about</p><p>Outer query asks the main question</p><p>Copy the values by hand</p><p>Subquery builds the filter list from live data</p><p>Paste them into a static <code>WHERE IN</code> list</p><p><code>WHERE</code> field <code>IN</code> (subquery) applies the filter live</p><p>Run query B with the hardcoded list</p><p>Results stay current with the data</p><p>Repeat the whole process when the data changes</p><p>No copied list to maintain</p><p><em>Figure 1. The old copy-paste loop collapses into a single dynamic filter.</em></p><h2>How the WHERE IN subquery replaces static filter lists</h2><p>Traditional <code>IN</code> filtering is still useful when the list is small and static:</p>FROM logs-*
| WHERE status_code IN (401, 403, 429)<p>But many real investigations don’t start with a tidy list. They start with a question: <em>Which users are suspicious</em>, <em>which hosts are noisy</em>, <em>which services are failing</em>, or <em>which accounts crossed a threshold?</em></p><p>That’s where the <code>WHERE IN</code> subquery becomes useful. The list is produced by ES|QL rather than typed by hand.</p><h2>WHERE IN subquery example: filtering logs by suspicious users</h2>FROM logs-*
| WHERE user.name IN (
    FROM auth-logs-*
    | WHERE event.action == "login_failed"
    | STATS failed_attempts = COUNT(*) BY user.name
    | WHERE failed_attempts &gt;= 10
    | KEEP user.name
  )<p>Read it like this: <em>Show me log events for users who appear in the list of users with at least 10 failed login attempts.</em></p><p>The outer query asks the main question, and the subquery builds the dynamic filter list, eliminating the need to copy and paste.</p><p>The subquery can target a different index or index pattern from the outer query.</p><h2>Filtering without subqueries: the manual ID copy workflow</h2><p>Imagine that you want to inspect traffic for the top failing services. First you run:</p>FROM service-logs-*
| WHERE status_code &gt;= 500
| STATS failures = COUNT(*) BY service.name
| SORT failures DESC
| LIMIT 5<p>Then you copy the five service names and paste them into another query:</p>FROM service-logs-*
| WHERE service.name IN ("checkout", "payments", "search", "profile", "orders")<p>That’s fine once, but less fine when the top five change every hour.</p><h2>Dynamic filtering with a WHERE IN subquery</h2>FROM service-logs-*
| WHERE service.name IN (
    FROM service-logs-*
    | WHERE status_code &gt;= 500
      AND @timestamp &gt;= now() - 2 days
    | STATS failures = COUNT(*) BY service.name
    | SORT failures DESC
    | LIMIT 5
    | KEEP service.name
  )
  AND status_code &gt;= 500
  AND @timestamp &gt;= now() - 2 days
| KEEP @timestamp, service.name, status_code, message<p>The subquery finds the top failing services from the last two days, and the outer query returns the log events for those services. One query builds the full picture.</p><h2>Excluding values with NOT IN subqueries in ES|QL</h2><p>Sometimes the interesting question is about what doesn’t belong:</p>FROM access-logs-*
| WHERE user.name NOT IN (
    FROM known-users
    | WHERE user.name IS NOT NULL
    | KEEP user.name
  )<p>That pattern is useful for exclusion checks, gap analysis, and workflows that ask for the things outside an approved or expected set.</p><h2>Nested subquery chains in ES|QL's WHERE clause</h2><p>An <code>IN</code> subquery replaces the literal value list with a query in parentheses. The inner query runs first and returns a single column, and the outer <code>WHERE</code> filters against it. Because that inner query is a full pipeline, it can contain its own <code>IN</code> subquery, which lets you express a chain of lookups that would otherwise require three separate queries and two rounds of copy-paste.</p>FROM orders
| WHERE customer_id IN (
    FROM customers
    | WHERE region_id IN (
        FROM regions
        | WHERE tier == "priority"
        | KEEP region_id
      )
    | KEEP customer_id
  )
| STATS revenue = SUM(amount) BY customer_id<p>Read it inside out. The innermost query finds priority regions, and the middle query finds customers in those regions, while the outer query sums revenue for those customers. Each layer is a normal ES|QL pipeline, so each one can filter, aggregate, or sort on its own before handing a clean column up to the layer above.</p><h2>Combining WHERE IN subqueries with AND and OR conditions</h2><p>Because an <code>IN</code> subquery is a Boolean condition, it composes with <code>AND</code> and <code>OR</code> like any other predicate. You can require membership in two independent sets or accept membership in either:</p>FROM orders
| WHERE customer_id IN (FROM vip_customers | KEEP customer_id)
  AND product_id IN (FROM discontinued_products | KEEP product_id)
| KEEP order_id, customer_id, product_id, amount<p>The <code>AND</code> combination finds orders placed by VIP customers for products that are being discontinued. Swap <code>AND</code> for <code>OR</code>, and you get orders that match either condition. Each subquery runs its own pipeline, so the two sets are computed independently and then combined by the Boolean operator.</p><h2>Merging multiple indices into one WHERE IN subquery</h2><p>The query inside an <code>IN</code> subquery is a full pipeline, so its <code>FROM</code> command can reference more than one subquery. Each branch runs its own pipeline, and the <code>FROM</code> command merges the rows from all branches into one result set. <code>KEEP host_id</code> selects the single column that the outer filter needs. This is useful when the values you want to filter against live in several indices with different schemas. For more details on how subqueries in the <code>FROM</code> command handle indices with different schemas, see <a href="https://www.elastic.co/search-labs/blog/esql-subquery-from">Three indices walk into a FROM clause: ES|QL subqueries in Elasticsearch</a>.</p>FROM alerts
| WHERE host_id IN (
    FROM
      (FROM prod_hosts    | WHERE region == "us-east"),
      (FROM staging_hosts | WHERE region == "us-east"),
      (FROM edge_hosts    | WHERE region == "us-east")
    | KEEP host_id
  )
| STATS alert_count = COUNT(*) BY host_id<p>The <code>IN</code> subquery combines matching host IDs from three indices, prod, staging, and edge, into one value list. The outer query then counts alerts for any host in that combined set. Adding a fourth source means adding one more branch, with no change to the outer query.</p><h2>Using an Elasticsearch subquery inside each FROM branch</h2><p>A <code>FROM</code> subquery gives each index its own branch with its own <code>WHERE</code>, and that <code>WHERE</code> can use an <code>IN</code> subquery. This is how you apply the same dynamic filter across several indices that each have their own schema.</p>FROM
  (FROM orders  | WHERE customer_id IN (FROM vip_customers | KEEP customer_id)),
  (FROM refunds | WHERE customer_id IN (FROM vip_customers | KEEP customer_id))
| STATS total_events = COUNT(*) BY customer_id<p>Each branch filters its index down to VIP customers before the two branches combine, so the final aggregation runs over a single normalized set of rows.</p><h2>When to use ES|QL WHERE IN subqueries</h2><ul><li><p>Investigations that start by finding risky users, hosts, accounts, or services.</p></li><li><p>Operational dashboards where the interesting entities change over time.</p></li><li><p>Top-N follow-up queries, such as events for the five noisiest services.</p></li><li><p>Set comparison workflows, especially with <code>NOT IN</code>.</p></li><li><p>Queries that would otherwise need glue code just to pass values from one step to the next.</p></li></ul><h2>Requirements and constraints for WHERE IN subqueries</h2><ul><li><p>Return exactly one column from the <code>IN</code> subquery.</p></li><li><p>Use <code>KEEP</code> at the end of the subquery so the comparison field is obvious.</p></li><li><p>Make sure the outer field and the subquery field have compatible types.</p></li><li><p>If the subquery uses <code>SORT</code>, add an explicit <code>LIMIT</code>, as unbounded <code>SORT</code> isn’t supported in ES|QL yet.</p></li><li><p>Use this for membership filtering. If you need columns from both sides, a <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-lookup-join">JOIN`</a> may be the better tool.</p></li></ul><h2>Why ES|QL dynamic filtering replaces manual query workflows</h2><p>The <code>WHERE IN</code> subquery turns a manual workflow into a declarative one. You can let one query build the filter for another query directly inside the <code>WHERE</code> command, instead of asking ES|QL for a list, copying it somewhere else, and hoping it stays fresh. </p><p>Your <code>WHERE</code> clause now has a better way to handle <em>Filter this by whatever that query finds.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/dynamic-filtering-esql-where-in-subquery</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/dynamic-filtering-esql-where-in-subquery</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Query Languages]]></category>
    <dc:creator><![CDATA[Fang Xing]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8588803197e341c5/6a75c709f124644e306fd73e/good_oone.png" length="0" type="image/png"/>
    <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[One query, three data sources: ES|QL subqueries get FROM, TS and ROW]]></title>
    <description><![CDATA[Filter application logs by live metric behavior and combine indexed data with inline test values. Your filter lists pull from time-series data on the fly, so nothing is hard-coded.]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> subqueries now support three source commands in Elasticsearch 9.5: <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/from"><code>FROM</code></a> for indexed data, <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> for time-series metrics, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/row"><code>ROW</code></a> for inline literal values. You can use them individually or combine all three in a single query, filtering log data by live metric behavior or mixing real and synthetic rows without any index setup.</p><p>If you've been exploring ES|QL, you might have noticed that a <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-subquery">subquery</a> used to feel like it had exactly one front door: <code>FROM</code>. And it made sense. Most of the time, queries start with a simple directive to go fetch documents from an index. In an earlier <a href="https://www.elastic.co/search-labs/blog/dynamic-filtering-esql-where-in-subquery">post</a>, we taught the <code>WHERE</code> command a new trick: <code>IN</code> subqueries. And before that, <a href="https://www.elastic.co/search-labs/blog/esql-subquery-from">subqueries showed up in the <code>FROM</code> command</a> to combine data sources. In both instances, every subquery started the same way, with <code>FROM</code>.</p><p>But not every useful query starts with a bulk document fetch. Sometimes you need to evaluate time-series semantics. Other times, you just need to whip up a tiny inline row for testing. Sometimes, the absolute best input to a filter is a dynamic query that builds the list for you on the fly, rather than a static, hard-coded list.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4453fe9ed8dad10b/6a75c5db7724f2073a045965/image3.png" alt="Diagram of three ES|QL subquery source commands with log filtering flow for FROM union and WHERE IN placements" /><p></p><p><strong>Source command</strong></p><p><strong>Reads from</strong></p><p><strong>Best for</strong></p><p><strong>Requirements</strong></p><p>FROM</p><p>Indexes, data streams, aliases, views</p><p>Stored document lookups, live filter lists</p><p>None (works with any index)</p><p>TS</p><p>Time-series data streams</p><p>Metric aggregations with counter-reset handling</p><p>TSDS with <code>index.mode: time_series</code></p><p>ROW</p><p>Inline literal values</p><p>Test cases, seed values, synthetic placeholders</p><p>None (no index needed)</p><p>The <code>FROM</code>, <code>TS</code> and <code>ROW</code> source commands are generally available (GA) in Elasticsearch 9.5, while the <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-in-subquery"><code>WHERE IN</code> subquery</a> remains in technical preview in 9.5.</p><h2>How ES|QL subquery source commands work</h2><p>Think of subqueries as having two specific placements and three different engines.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5542afed68fff9e7/6a75c5f6cd87c322296bcb39/image1.png" alt="ES|QL subquery source commands grid showing FROM, ROW and TS examples in both FROM union and WHERE IN filter placements" /><p></p><h3>Where subqueries go: FROM and WHERE placements</h3><ul><li><p><strong>Inside the </strong><strong><code>FROM</code></strong><strong> command:</strong> The subquery is an independent result source, contributing its rows to the outer query. Fields that exist in one source but not the other are gracefully filled with null values.</p></li><li><p><strong>Inside the </strong><strong><code>WHERE</code></strong><strong> command:</strong> The <code>IN</code> subquery contributes dynamic values to be used as a predicate or filter.</p></li></ul><h3>Three source commands for starting a subquery</h3><h4>Door #1: FROM (the classic door)</h4><p><code>FROM</code> is the familiar workhorse. You use it when the subquery should read stored data from indices, data streams, aliases, or views. One of its best use cases is the "stop-copy-pasting-IDs" pattern. Instead of running one query, manually copying the output values, and pasting them into the filter of another query, the subquery becomes your live filter list.</p>FROM employees
| WHERE emp_no IN (FROM high_value_accounts
                   | KEEP emp_no
                  )
| KEEP emp_no, first_name, last_name<h4>Door #2: ROW (the tiny door)</h4><p><code>ROW</code> is the lightweight option that requires absolutely no index setup. It allows you to build rows completely out of literal, inline values. This makes <code>ROW</code> useful for small seed values, test cases, allow/deny lists, or one-off "what if?" scenarios. In the query below, <code>ROW</code> is the perfect way to staple a synthetic sentinel row or placeholder directly onto real data.</p>FROM
(FROM access_logs
   | WHERE status == 500
| KEEP cluster, status),
  (ROW cluster = "synthetic", status = 0)
| SORT status
| KEEP cluster, status<h4>Door #3: TS (the time-series door)</h4><p>The <code>TS</code> command targets time-series data streams and enables time-series aggregation functions. Why not just use <code>FROM</code> for metrics? <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> is uniquely optimized for time-series data, and it natively handles tricky scenarios, like counter resets on process restarts and uneven metric publish intervals. Using TS as a subquery lets you filter your application logs by what your metrics are saying. For example, imagine asking ES|QL: <em>Show me log events only from clusters whose peak metric throughput crossed 800</em>:</p>FROM access_logs
| WHERE cluster IN (TS k8s_metrics
                    | STATS peak = MAX(bytes_in) BY cluster
                    | WHERE peak &gt; 800
| KEEP cluster
                   )
| SORT cluster, path
| KEEP cluster, status, path<p>In this scenario, the filter is reacting to live metric behavior rather than being hard-coded.</p><h2>Combining FROM, TS and ROW in one query</h2><p>Because subquery placements and source commands are independent, you can freely mix and match them. You can throw all three doors into a single <code>FROM</code> union to generate a cohesive table containing real logs, a live metrics summary, and a synthetic row.</p>FROM
(FROM access_logs
   | KEEP cluster, status),
  (TS k8s_metrics
   | STATS peak = MAX(bytes_in) BY cluster),
  (ROW cluster = "synthetic")
| STATS log_events = COUNT(status), peak = MAX(peak) BY cluster
| SORT cluster
| KEEP cluster, log_events, peak<h2>ES|QL subquery constraints</h2><p>A few constraints to keep in mind before using subqueries:</p><ul><li><p><strong><code>IN</code></strong><strong> subqueries demand one column:</strong> If a subquery feeds an <code>IN</code> operator, it must project exactly one column. Use the <code>KEEP</code> command to make that explicitly clear.</p></li><li><p><strong><code>TS</code></strong><strong> requires a time series data stream (TSDS):</strong> The <code>TS</code> command only works on data stored in a <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">TSDS</a>, which uses <code>index.mode: time_series</code>.</p></li></ul><h2>The takeaway: when to use each source command</h2><p>Subqueries in ES|QL provide a structured way to compose queries by using the results of one query as the input to another. Choosing the appropriate source command (<code>FROM</code>, <code>ROW</code>, or <code>TS</code>) lets you combine data and generate inline values. It also lets you filter dynamically without duplicating query logic. For more details and additional examples, see the ES|QL subqueries documentation.</p><ul><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-subquery">Subquery in FROM command</a></p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-in-subquery">Subquery in WHERE command</a></p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-subquery-source-commands</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-subquery-source-commands</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Query Languages]]></category>
    <dc:creator><![CDATA[Fang Xing]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt71cfd2ab01e3262e/6a75c5beaec90746295fd3b7/image2.png" length="0" type="image/png"/>
    <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[From search to checkout in 20 lines of code: building a 4-stage conversion funnel with OpenTelemetry]]></title>
    <description><![CDATA[Add cart and purchase tracking to your search analytics pipeline and use ES|QL to answer the question every product manager asks: which search queries drive the most revenue?]]></description>
    <content:encoded><![CDATA[<p>Your product manager wants to know which searches drive the most revenue. You can tell them what users search for (<a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">check out our second blog</a>) and what they click (<a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">in our third blog</a>), but not what they buy. Two new span types, add-to-cart and purchase, complete a four-stage funnel from search query to checkout, built on the same search.* attributes and ES|QL queries you've been using since Blog 2. About 20 lines of code, and every relevance decision you make gets a revenue number attached to it.</p><h2>What you'll discover</h2><p>In this post, you'll learn how to:</p><ul><li><p>Add conversion tracking (add-to-cart and purchase spans) with <code>search.*</code> attributes that tie back to the originating search.</p></li><li><p>Build a full search-to-revenue funnel: search → click → add-to-cart → purchase.</p></li><li><p>Write Elasticsearch Query Language (ES|QL) queries to calculate conversion rates, revenue per query, and average order value.</p></li><li><p>Identify where users drop off and which team should own each drop-off point.</p></li><li><p>Attribute revenue to specific search queries for prioritizing relevance work.</p></li></ul><h2>What you'll need</h2><ul><li><p>Click tracking from <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a> (search + click spans flowing to Elastic via OTel-native ingestion).</p></li><li><p>Backend endpoints for add-to-cart and checkout events (example code provided).</p></li><li><p>Basic understanding of ecommerce conversion funnels.</p></li></ul><h2>Why search revenue attribution matters</h2><p>Your product manager walks into a meeting and asks, "Which searches are driving the most revenue?"</p><p>You can tell them what users search for (<a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blog 2</a>) and what they click (<a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>). But you can't tell them what they <em>buy</em>. The gap between "clicked a result" and "purchased a product" is where the business case for search investment lives, and right now it's invisible.</p><p>This post closes the loop. By adding two more span types (add-to-cart and purchase), you get a full funnel from search to revenue, built on the same <code>search.*</code> attributes and ES|QL queries you've been using since <a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry">Blog 2</a>.</p><h3>What search conversion tracking answers for your team</h3><p>Here's what conversion tracking lets you answer and why each question matters to different people on your team.</p><ul><li><p><strong>Which searches drive revenue?</strong> This is the product manager's question. When you can attribute dollars to specific queries, you can prioritize relevance work by business impact. A query with mediocre click-through rate (CTR) but high conversion value is more important than a high-CTR query that never leads to a purchase.</p></li></ul><ul><li><p><strong>Where do users drop off?</strong> The funnel from search to purchase has four stages: search, click, add-to-cart, and purchase. Each drop-off points to a different problem. High click-to-cart drop-off suggests that product pages aren't convincing. High cart-to-purchase drop-off is checkout friction rather than a search problem. Knowing <em>where</em> users abandon tells you <em>which team</em> should fix it.</p></li></ul><ul><li><p><strong>Which queries to protect?</strong> Once you know that "laptop bag" generates $12,000/month in attributed revenue, you treat it differently. Any relevance change that touches high-revenue queries gets extra scrutiny. You can set up monitoring (Blog 6, coming soon) to alert when conversion rates drop for your top-earning searches.</p></li></ul><p>Two new instrumentation points total about 20 lines of code, and they follow the same pattern as you’ve used before. You add attributes to spans, and query them with ES|QL.</p><p><strong>Following along with code?</strong> The <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">reference project</a> has conversion tracking ready to enable. Uncomment the Blog 4 sections in <code>app.py</code> and <code>app.js</code>, restart, and then generate traffic with <code>python generate_traffic.py --blog 4</code>.</p><h2>The four-stage search conversion funnel</h2><p>Before we write any code, here's the shape of what we're building. Each stage is an instrumentation point, and each creates spans in <code>traces-generic.otel-default</code>:</p>search          →  click              →  cart.add         →  checkout.complete
(Blog 2)           (Blog 3)              (this post)         (this post)
query_id=abc       query_id=abc          query_id=abc        query_id=abc
user_query=...     click_position=1      product_id=...      order_total=$149
result_count=15    product_id=...        quantity=1           item_count=2<p>The thread running through the entire chain is <code>search.query_id</code><code>.</code> The same identifier you derived from the trace ID in <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a> and used to link clicks to searches in <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a> now carries through to add-to-cart and purchase events. This is what makes revenue attribution possible: You can trace a purchase back to the search that started the journey.</p><h2>Add conversion tracking spans with OpenTelemetry</h2><h3>Add-to-cart span instrumentation</h3><p>When a user adds a product to their cart from a search results page (or from a product detail page they reached via search), you create a <code>cart.add</code> span. This captures the moment that intent turns into action.</p>@app.post("/api/cart/add")
async def add_to_cart(event: AddToCartRequest):  # reference project uses CartEvent
    with tracer.start_as_current_span("cart.add") as span:
        span.set_attribute("search.action", "add_to_cart")
        span.set_attribute("search.result_click_id", event.object_id)
        span.set_attribute("search.result_click_position", event.position)
        span.set_attribute("search.query_id", event.query_id)
        span.set_attribute("enduser.pseudo.id", event.client_id)
        span.set_attribute("cart.quantity", event.quantity)
        if event.price is not None:
            span.set_attribute("cart.price", event.price)
        if event.user_query:
            span.set_attribute("search.query", event.user_query)<p>This follows the same pattern as click tracking in <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>; that is, an independent span linked to the originating search via <code>query_id</code>. The new attributes are <code>cart.quantity</code> and <code>cart.price</code> which let you aggregate revenue at query time.</p><h3>Purchase span instrumentation</h3><p>When the user completes checkout, you create a <code>checkout.complete</code> span. This is the revenue event, the one that answers the product manager's question.</p>@app.post("/api/checkout")
async def checkout(event: CheckoutRequest):  # reference project uses CheckoutEvent
    with tracer.start_as_current_span("checkout.complete") as span:
        span.set_attribute("search.action", "purchase")
        span.set_attribute("checkout.order_id", event.order_id)
        span.set_attribute("checkout.total_amount", event.total_amount)
        span.set_attribute("checkout.item_count", len(event.items))
        span.set_attribute("enduser.pseudo.id", event.client_id)
        if event.query_id:  # last search in journey
            span.set_attribute("search.query_id", event.query_id)
            span.set_attribute("search.query", event.user_query)<p></p><ul><li><p><strong>Spans capture errors automatically.</strong> This is a side benefit of using OTel spans for conversion events. If an add-to-cart or checkout call throws an unhandled exception, the span's status is automatically set to <code>ERROR</code> and the exception details are recorded. These errors are business-critical (a broken checkout flow means lost revenue), and they show up immediately in Elastic APM's error tracking, service maps, and alerting. You get conversion analytics <em>and</em> operational monitoring from the same instrumentation, with no extra code.</p></li></ul><ul><li><p><strong><code>checkout.total_amount</code></strong><strong> is the revenue number.</strong> This is what you'll aggregate in ES|QL to get revenue-by-query. It represents the order total, not the price of a single item.</p></li></ul><ul><li><p><strong><code>search.query_id</code></strong><strong> is conditional.</strong> Not every purchase originates from search. Users browse categories and follow promotional links. Then they return to their cart days later. The <code>if event.query_id:</code> guard ensures that you only attribute purchases to search when there's a genuine connection. Purchases without a <code>query_id</code> still get recorded; they just don't appear in search attribution queries.</p></li></ul><ul><li><p><strong><code>search.query</code></strong><strong> is set on both cart and purchase spans.</strong> This is a deliberate denormalization. You could join back to searches via <code>query_id</code> to get the query text, but while ES|QL does support <code>LOOKUP JOIN</code>, it’s likely not the right choice here due to the need to optimize the lookup index. By putting the query text directly on conversion spans, your revenue-by-query and cart-by-query queries are single, straightforward aggregations.</p></li></ul><h3>Span attributes for cart and purchase events</h3><p><strong>Add-to-cart attributes:</strong></p><p><strong>Attribute</strong></p><p><strong>Type</strong></p><p><strong>Required</strong></p><p><strong>Purpose</strong></p><p><code>search.action</code></p><p>string</p><p>yes</p><p><code>"add_to_cart"</code></p><p><code>search.result_click_id</code></p><p>string</p><p>yes</p><p>Product document ID</p><p><code>search.result_click_position</code></p><p>int</p><p>yes</p><p>Position in results when added</p><p><code>search.query_id</code></p><p>string</p><p>yes</p><p>Links to originating search</p><p><code>enduser.pseudo.id</code></p><p>string</p><p>yes</p><p>Client/device identifier (OTel Semantic Conventions  [SemConv])</p><p><code>cart.quantity</code></p><p>int</p><p>yes</p><p>Quantity added</p><p><code>cart.price</code></p><p>float</p><p>optional</p><p>Product price</p><p><code>search.query</code></p><p>string</p><p>recommended</p><p>The search query text (for cart-by-query analysis)</p><p><strong>Purchase attributes:</strong></p><p><strong>Attribute</strong></p><p><strong>Type</strong></p><p><strong>Required</strong></p><p><strong>Purpose</strong></p><p><code>search.action</code></p><p>string</p><p>yes</p><p><code>"purchase"</code></p><p><code>checkout.order_id</code></p><p>string</p><p>yes</p><p>Unique order identifier</p><p><code>checkout.total_amount</code></p><p>float</p><p>yes</p><p>Order total</p><p><code>checkout.item_count</code></p><p>int</p><p>yes</p><p>Number of items purchased</p><p><code>enduser.pseudo.id</code></p><p>string</p><p>yes</p><p>Client/device identifier (OTel SemConv)</p><p><code>search.query_id</code></p><p>string</p><p>recommended</p><p>Links to originating search (last search in journey)</p><p><code>search.query</code></p><p>string</p><p>recommended</p><p>The search query text (for revenue-by-query analysis)</p><p>These follow the same <code>search.*</code> namespace from Blogs 2 and 3, with new <code>cart.*</code> and <code>checkout.*</code> prefixes for conversion-specific data. With OTel-native ingestion, all attributes are stored under <code>attributes.*</code> with dot notation preserved. That means no more mapping strings to <code>labels.*</code> and numbers to <code>numeric_labels.*</code>. A <code>search.query</code> attribute is queryable as <code>attributes.search.query</code>. Likewise, <code>checkout.total_amount</code> is queryable as <code>attributes.checkout.total_amount</code>.</p><h3>User identity: connecting search to purchase across sessions</h3><p>You'll notice that <code>enduser.pseudo.id</code> appears on every span type in this series. It's the <em>minimum identity level</em>; that is, a persistent identifier stored in the browser's localStorage that ties events to a device across sessions.</p><p>For conversion tracking, identity becomes more important. You need to connect a search on Monday to a purchase on Tuesday, or correlate cart additions across tabs. Our schema supports three identity levels, aligned with OTel semantic conventions:</p><p><strong>Attribute</strong></p><p><strong>Persistence</strong></p><p><strong>Purpose</strong></p><p><code>enduser.pseudo.id</code></p><p>Permanent (localStorage)</p><p>Device/browser identifier (OTel SemConv)</p><p><code>session.id</code></p><p>Per-visit (sessionStorage)</p><p>Groups events within a single visit (OTel SemConv)</p><p><code>user.id</code></p><p>Account (auth system)</p><p>Authenticated user (OTel SemConv)</p><p>For the funnel queries in this post, <code>enduser.pseudo.id</code> is sufficient. It links the journey from search to purchase within a browser. If your users authenticate, adding <code>user.id</code> enables cross-device attribution (searched on mobile, purchased on desktop) and richer personalization. <code>session.id</code> helps disambiguate when the same client has multiple active sessions.</p><p>All three are optional on interaction spans. Start with <code>enduser.pseudo.id</code>, and add the others when your use case requires them. The important thing is consistency: Use the same identifiers across search, click, and conversion spans so the joins work.</p><h3>Frontend integration</h3><p>The front end needs to propagate <code>query_id</code> through the user journey. When the user clicks a search result, you already have <code>query_id</code> from the search response (<a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>). The key is carrying it forward.</p>// CLIENT_ID: persistent browser identifier from localStorage (set up in Blog 3)
// const CLIENT_ID = localStorage.getItem("search_client_id") || ...

// On add-to-cart from a search result page
fetch('/api/cart/add', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    object_id: product.id,
    position: product.resultPosition,  // from search results
    query_id: product.queryId,         // from search response
    client_id: CLIENT_ID,              // persistent browser identifier → enduser.pseudo.id
    quantity: 1,
    price: product.price,
  })
});

// On checkout completion
fetch('/api/checkout', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    order_id: order.id,
    total_amount: order.total,
    items: order.items,
    client_id: CLIENT_ID,              // same identifier as search and click spans
    query_id: lastSearchQueryId,       // last search in session
    user_query: lastSearchQuery,
  })
});<p>The front end sends <code>client_id</code> as an HTTP field name; and the back end maps it to the OTel semantic convention <code>enduser.pseudo.id</code> when setting span attributes. The <code>query_id</code> propagation is the critical piece. Store it alongside the product in the cart data structure so it survives navigation between pages. We'll discuss the design challenges of this in the attribution section below.</p><p></p><p><strong>Using the reference project?</strong> The reference app's <code>frontend/app.js</code> wires up add-to-cart buttons, but the checkout flow isn’t implemented in the browser UI. It goes through the traffic generator. To simulate conversion events, run: <code>python generate_traffic.py --blog 4 --sessions 100</code>. This sends a realistic mix of searches, clicks, cart additions, and purchases to all three backend endpoints.</p><h3>Verify that conversion events are arriving</h3><p>Before building funnel queries, confirm that both span types are flowing to Elastic:</p><p>Add-to-cart events:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "add_to_cart"
| KEEP attributes.search.result_click_id, attributes.search.query_id,
       attributes.cart.quantity
| LIMIT 5<p>Purchase events:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "purchase"
| KEEP attributes.checkout.order_id, attributes.checkout.total_amount,
       attributes.checkout.item_count, attributes.search.query_id,
       attributes.search.query
| LIMIT 5<p>If these return rows, you have the full funnel instrumented. If not, check the same things as always: OpenTelemetry Protocol (OTLP) endpoint, auth token, and span export.</p><h2>Funnel analysis with ES|QL</h2><h3>Count search, click, cart and purchase events in a single query</h3><p>You can count all four funnel stages in a single ES|QL query using the same <code>COUNT(CASE(...))</code>pattern from the CTR query in <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>:</p>FROM traces-generic.otel-default
| WHERE (name == "search" AND attributes.search.query IS NOT NULL)
    OR attributes.search.first_click == true
    OR attributes.search.action IN ("add_to_cart", "purchase")
| STATS
    searches  = COUNT(CASE(name == "search" AND attributes.search.query IS NOT NULL, 1)),
    clicked   = COUNT(CASE(attributes.search.first_click == true, 1)),
    carts     = COUNT(CASE(attributes.search.action == "add_to_cart", 1)),
    purchases = COUNT(CASE(attributes.search.action == "purchase", 1))
| EVAL
    click_rate    = ROUND(100.0 * clicked   / searches, 1),
    cart_rate     = ROUND(100.0 * carts     / searches, 1),
    purchase_rate = ROUND(100.0 * purchases / searches, 1)<p><strong>Example output:</strong></p><p><strong>searches</strong></p><p><strong>clicked</strong></p><p><strong>carts</strong></p><p><strong>purchases</strong></p><p><strong>click_rate</strong></p><p><strong>cart_rate</strong></p><p><strong>purchase_rate</strong></p><p>146</p><p>41</p><p>28</p><p>12</p><p>28.1%</p><p>19.2%</p><p>8.2%</p><p>You get all four counts and three conversion rates from a single query. The <code>WHERE</code> clause pulls all four span types into one result set, and <code>COUNT(CASE(...))</code> counts each type separately. This is the same technique that made the CTR query in <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a> so clean.</p><p>Notice that we use <code>search.first_click</code>for the click stage rather than counting all click events. This gives you the number of searches that received at least one click (the same definition used for CTR in <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>). Without this, a search with three clicks would inflate the click count to three while only counting as one search, making the funnel numbers misleading. Each stage now represents a unique progression: how many searches happened, how many of those got clicked, how many led to a cart addition, and how many resulted in a purchase.</p><p>You can turn this into a funnel visualization using Kibana Lens. Run each stage count as a separate ES|QL query, save them as dashboard panels, and arrange them as a horizontal bar chart with the four stages on the y-axis and counts on the x-axis. The drop-off at each step becomes immediately visible.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcc9d695520cd0aff/6a7458d85751aa8c3e7e0a4f/image2.png" alt="Kibana dashboard showing search analytics: conversion funnel, revenue by query, click-through rate and latency SLOs" /><p>The dashboard above (from <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry">the first blog in the series</a>) shows this in practice: The <strong>Conversion Funnel</strong> panel in the lower left uses a horizontal bar chart to visualize the four stages, and the <strong>Top Queries by Revenue</strong> table alongside it shows revenue attribution. You can build these directly from the ES|QL queries in this post.</p><h3>What drop-off rates tell you and who owns each bottleneck</h3><p>Each transition in the funnel tells you something specific:</p><ul><li><p><strong>Search to click (CTR):</strong> You measured this in <a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>. Low CTR means that results aren't compelling. This is a relevance problem.</p></li></ul><ul><li><p><strong>Click to cart:</strong> The user engaged with a result but didn't add it to their cart. This could mean that the product page isn't persuasive or the price isn't competitive. It could also mean that the item was out of stock. It's often <em>not</em> a search problem. It’s possible that the search worked (the user clicked), but something downstream lost them.</p></li></ul><ul><li><p><strong>Cart to purchase:</strong> The user committed to buying but didn't complete checkout. Complicated forms, unexpected shipping costs, and payment issues cause this <em>checkout friction</em>. This is almost never a search problem, but it's useful to know where the funnel leaks so you don't waste time optimizing relevance when checkout is the bottleneck.</p></li></ul><p>The diagnostic pattern is straightforward:</p><p><strong>Drop-off point</strong></p><p><strong>Likely cause</strong></p><p><strong>Who owns it</strong></p><p>Search to click</p><p>Relevance / ranking</p><p>Search team</p><p>Click to cart</p><p>Product page / pricing / availability</p><p>Product / merchandising</p><p>Cart to purchase</p><p>Checkout UX / payment / shipping</p><p>Checkout / growth</p><p>This is one of the most valuable things that full-funnel data gives you: the ability to point at the right problem. When the VP asks, "Why aren't searches converting?", you can show whether the bottleneck is relevance, product pages, or checkout.</p><h2>Revenue attribution by search query</h2><p>Here's the query your product manager actually wants:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "purchase"
  AND attributes.search.query IS NOT NULL
| STATS
    purchase_count = COUNT(*),
    total_revenue = SUM(attributes.checkout.total_amount)
  BY attributes.search.query
| SORT total_revenue DESC<p>This gives you a ranked list of queries by the revenue they generated. The <code>search.query</code> attribute on purchase spans (the denormalization we discussed earlier) makes this a single aggregation query, without joins or subqueries.</p><h3>How to act on search revenue  data</h3><ul><li><p><strong>Protect high-revenue queries.</strong> If "laptop bag" generates the most revenue, any relevance change that affects that query gets extra scrutiny. You might add it to a regression test suite or pin specific results with <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/searching-with-query-rules">query rules</a>. You might even set up an alert when its conversion rate drops (Blog 6).</p></li></ul><ul><li><p><strong>Prioritize relevance investment.</strong> The queries at the top of this list are where relevance improvements have the most business impact. A 10% CTR improvement on a query that generates $500/month in revenue is worth more than a 50% improvement on one that generates $20.</p></li></ul><ul><li><p><strong>Identify missed opportunities.</strong> Cross-reference with the top-queries data from <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a>. A query with high search volume but no purchase attribution is either a browsing query (informational intent) or a conversion gap worth investigating.</p></li></ul><h3>Top revenue queries: where to focus relevance investment</h3><p>To focus your relevance team's efforts, pull the top revenue-generating queries:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "purchase"
  AND attributes.search.query IS NOT NULL
| STATS
    purchases = COUNT(*),
    revenue = SUM(attributes.checkout.total_amount)
  BY attributes.search.query
| SORT revenue DESC
| LIMIT 10<p>Cross-reference this with the per-query CTR from Blog 3. A query with high revenue but low CTR is underperforming; even small relevance improvements have outsized business impact. A query with high CTR but no purchase attribution might be informational (such as users browsing but not buying). The queries at the top of <em>both</em> lists deserve the most attention from your relevance team.</p><h3>Average order value by search query</h3><p>Average order value (AOV) tells you how much purchases are worth, in addition to which queries convert to those purchases. Queries with high AOV are your premium-intent searches; ranking improvements there have the biggest per-purchase impact:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "purchase"
  AND attributes.search.query IS NOT NULL
| STATS
    purchase_count = COUNT(*),
    total_revenue = SUM(attributes.checkout.total_amount),
    avg_order_value = ROUND(AVG(attributes.checkout.total_amount), 2)
  BY attributes.search.query
| SORT avg_order_value DESC
| LIMIT 10<p>A query with high AOV but low volume is a different opportunity than high volume + low AOV. The first means premium intent from a small audience (consider featured results or dedicated landing pages); the second means broad reach with budget buyers (price sensitivity may be limiting conversion more than relevance).</p><h2>Search revenue attribution: limitations and workarounds</h2><p>Revenue attribution from search is valuable, but it's imperfect. Understanding the limitations helps you set appropriate expectations and design around them.</p><h3>Last-touch attribution in multi-search journeys</h3><p>Users rarely search once and buy. A typical journey might look like:</p><ol><li><p>Search "laptop bag": Browse results and click a few.</p></li><li><p>Search "laptop bag leather": Refine the search.</p></li><li><p>Search "laptop sleeve 15 inch": Try a different angle.</p></li><li><p>Add to cart from the third search's results.</p></li><li><p>Purchase.</p></li></ol><p>With the instrumentation above, this purchase attributes to the third search, the one whose <code>query_id</code> was on the cart item. The first two searches contributed to the journey but get no credit.</p><p>This is <em>last-touch attribution</em>, and it's the simplest model that works within a single <code>query_id</code> linkage. It's not perfect, but it's concrete and unambiguous. The alternative, that is,tracking every <code>query_id</code> in a user's session and distributing credit, adds significant complexity to both instrumentation and analysis.</p><p>For most teams, last-touch is a good starting point. If you need multi-touch attribution later, the raw data is there. You can query all searches and clicks for a given <code>client_id</code> within a time window and reconstruct the full journey:</p>FROM traces-generic.otel-default
| WHERE attributes.enduser.pseudo.id == "client-abc-123"
  AND (name == "search"
    OR attributes.search.action == "click"
    OR attributes.search.action == "add_to_cart"
    OR attributes.search.action == "purchase")
| KEEP @timestamp, name, attributes.search.action,
       attributes.search.query, attributes.search.query_id,
       attributes.search.result_click_id
| SORT @timestamp ASC<p>This reconstructs a user's full search journey in chronological order. It's useful for debugging individual sessions, even if you don't build automated multi-touch attribution.</p><h3>Cross-session attribution limits with query_id</h3><p>A user searches for "wireless headphones" on Monday, clicks a few results, leaves, and comes back on Wednesday to buy. The <code>query_id</code> from Monday's search is long gone, since it was a property of that specific search request.</p><p>This is a fundamental limitation of <code>query_id</code>-based attribution. It works within a session (or more precisely, within the scope where the front end retains the <code>query_id</code>). It doesn't work across sessions.</p><p>For cross-session attribution, you'd need a different approach, typically a user-level event store where you associate product views, cart additions, and purchases with a persistent user ID and then look back in time to find the originating search. That's a more complex analytics pipeline and is outside the scope of what we're building here.</p><p>The practical impact is that your search-attributed revenue will be an <em>undercount</em>. Some purchases that were genuinely influenced by search won't carry a <code>query_id</code>. This is fine for relative comparisons (such as, <em>Which queries generate </em>more<em> revenue than others?</em>), even if the absolute numbers are conservative.</p><h3>How to persist query_id from search to checkout</h3><p>A few practical decisions affect how far your <code>query_id</code> propagation reaches:</p><ul><li><p><strong>Store </strong><strong><code>query_id</code></strong><strong> in the cart.</strong> When a user adds a product to their cart, persist the <code>query_id</code> alongside the item. This way, even if the user navigates away and comes back to checkout later (within the same session), the attribution survives.</p></li></ul><ul><li><p><strong>Don't overwrite </strong><strong><code>query_id</code></strong><strong> on re-search.</strong> If a user adds a product from search A and then searches again and adds another product from search B, each cart item should keep its own <code>query_id</code>. The purchase event carries the last search's <code>query_id</code> as a summary, but per-item attribution gives you richer data.</p></li></ul><ul><li><p><strong>Accept the limitations.</strong> Not every purchase will have search attribution. Direct navigation, category browsing, promotional links, and returning customers who go straight to their cart will all produce purchases without a <code>query_id</code>. That's correct behavior, not missing data.</p></li></ul><h3>Spans vs. log events for conversion tracking</h3><p>Conversion events can emit both an OTel span and a UBI-compatible log event, the same dual-signal pattern used for click tracking in Blog 3. The log event is actually richer for conversions. It can contain the full items list with per-item <code>query_id</code> attribution, which doesn't map cleanly to flat span attributes.</p><p>The span gives you the simple, aggregatable view (total revenue by query), and the log gives you the detailed, per-item view (which specific products from which specific searches). For the funnel queries in this post, spans are sufficient. If you need item-level attribution analysis, the log events in <code>logs-generic.otel-default</code> have the detail you need, and ES|QL queries them the same way, just against a different index pattern.</p><h2>What's next: turning conversion data into relevance improvements</h2><p>We now have the complete instrumentation picture, with four span types: <code>search</code> (<a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">Blog 2</a>), <code>search.result.click</code> (<a href="https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql">Blog 3</a>), <code>cart.add</code>, and <code>checkout.complete</code> (this post). These capture the full user journey from query to purchase. Every span lives in <code>traces-generic.otel-default</code>, and every metric is queryable with ES|QL. The <code>search.query_id</code> thread ties the entire funnel together.</p><p>But measuring the funnel is only half the story. The real payoff is using this data to make search better.</p><p>In a later blog in this series, we take everything we've built and turn it into relevance improvements. Click positions and conversion data become judgment lists for <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/learning-to-rank.html">Learning To Rank</a>, and per-query CTR and revenue become rank features for boosting. Plus, high-revenue queries get protective monitoring. And tools like <a href="https://elastic.github.io/relevance-studio/#/">Elasticsearch Relevance Studio</a> give you a visual interface for tuning the searches that matter most, using exactly the data you're now collecting.</p><p>The instrumentation you've built in Blogs 2–4 is a feedback loop, beyond analytics: Measure, improve, and measure again.</p><h2>Get started with search conversion tracking</h2><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">Reference project:</a> Working code for the entire blog series; clone, configure, and run.</p></li><li><p><a href="https://github.com/elastic/elastic-otel-python">Elastic Distribution of OpenTelemetry for Python:</a> EDOT for Python.</p></li><li><p><a href="https://www.elastic.co/docs/solutions/observability/apm/opentelemetry">OpenTelemetry with Elastic:</a> How to send OTel data to Elastic APM.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL documentation:</a> Query language reference.</p></li><li><p><a href="https://www.ubisearch.dev/">UBI Standard:</a> Reference schema for search event structure.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/query-rules.html">Query rules:</a> Pin, boost, or exclude results for specific queries.</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/search-conversion-tracking-opentelemetry</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/search-conversion-tracking-opentelemetry</guid>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Relevance]]></category>
    <dc:creator><![CDATA[Matthew Adams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9492cc64bcc1d0dc/6ab3ead098a4e3c62b53f8a0/diagram-building-a-4-stage-funnel.webp" length="0" type="image/webp"/>
    <pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch ES|QL brings full-text search to data you never indexed]]></title>
    <description><![CDATA[MATCH and TO_TEXT bring full-text search to data you never indexed. Search computed columns, unmapped fields and federated sources in ES|QL.]]></description>
    <content:encoded><![CDATA[<p>ES|QL <code>MATCH</code> now runs full-text search on data you never indexed. Computed columns, unmapped fields, strings assembled on the fly, even federated data sitting in S3. The new <code>TO_TEXT</code> function tells ES|QL to treat any string as analyzable text, so <code>MATCH</code> can tokenize, case-fold and term-match values that exist only for the lifetime of a query. This goes beyond the <code>LIKE</code> and <code>RLIKE</code> pattern matching that most query engines offer for unindexed strings: it's real analysis. Available now in Elastic Cloud Serverless and as a technical preview in Elasticsearch 9.5.</p><h2>How MATCH and TO_TEXT enable full-text search on any ES|QL expression</h2><p>Let's start with a query that was impossible in Elasticsearch 9.4, which uses <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/eval">the <code>EVAL</code> command</a>:</p><p>In this example, <code>summary</code> has no mapping or analyzer configuration. It’s also not associated with any inverted index. It exists only for the lifetime of this query, but now you can search it anyway. Two additions make this work.</p><p>First, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/search-functions/match"><code>MATCH</code></a> now accepts any expression as its first argument, not just a mapped field. That includes columns produced by <code>EVAL</code> and function results used inline. It also includes unmapped fields loaded directly from the original document. Furthermore, all data types normally accepted by <code>MATCH</code> are supported in this new use case.</p><p>The second part of this is the new <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/type-conversion-functions/to_text"><code>TO_TEXT</code></a> function, which is the first ES|QL conversion function that produces output of type <code>text</code>. Until now, <code>text</code> columns could only come from indexed mapped fields, and all strings produced by ES|QL expressions were <code>keyword</code> values rather than <code>text</code>. The distinction matters because <code>MATCH</code> treats the two differently: <code>text</code> values are analyzed, while <code>keyword</code> values are compared exactly, mirroring how a <code>MATCH</code> query on an indexed keyword field rewrites to a term query. <code>TO_TEXT(x)</code> is how you tell ES|QL: <em>treat this string as full text</em>.</p><p>This ships as a technical preview in Elasticsearch 9.5, and as such, it has some limitations:</p><ul><li><p>It’s currently filtering only. A <code>MATCH</code> on an expression doesn't contribute to the relevance score yet; only matches on indexed fields affect the score.</p></li><li><p>Querying options like <code>fuzziness</code> and others aren't yet supported when matching an expression.</p></li><li><p>Runtime text is analyzed with the standard analyzer. This isn’t configurable yet.</p></li></ul><p>Work is underway to address these limitations.</p><h2>Why use full-text search instead of LIKE or RLIKE in ES|QL?</h2><p>ES|QL already had two ways to search strings without an index: <code>LIKE</code> (wildcard patterns) and <code>RLIKE</code> (regular expressions). Both work on any string expression, so it's fair to ask what <code>MATCH</code> adds. The answer is <a href="https://www.elastic.co/docs/manage-data/data-store/text-analysis">analysis</a>, a more advanced form of search which uses techniques such as stemming and synonyms. It also uses stopword handling.</p><p><code>LIKE</code> is simple substring matching, without any understanding of the words that comprise a string. Say, for example, you're looking for log messages about a fox:</p><p>This misses <code>"Fox spotted near the henhouse"</code> due to the capitalization, while matching <code>"Outfoxed by the competition"</code>, which isn't about a fox at all. It fails in both directions, with false negatives on capitalization and false positives on substrings buried inside other words.</p><p>Regular expressions can patch the case problem, but the word-boundary problem gets ugly fast. Something like:</p><p>And even that's not right yet. It misses a fox at the end of a sentence followed by <code>!</code> or <code>?</code>, and it says nothing about tabs, quotes, or parentheses. Each fix makes the pattern longer, and the next person to read the query has to reverse-engineer what it’s actually doing.</p><p><code>MATCH</code> makes the problem go away, because it runs both the query and the value through an <a href="https://www.elastic.co/docs/reference/text-analysis/analyzer-reference">analyzer</a>, which tokenizes text into lowercase terms and then matches term against term:</p><p>This query will match values like <code>"The quick brown fox"</code> and <code>"FOX spotted near the henhouse"</code> but not <code>"Outfoxed by the competition"</code> or <code>“FOXTROT protocol enabled"</code>, regardless of any punctuation surrounding the words. Of course, this all works for multi-term queries, like <code>MATCH(TO_TEXT(message), "brown fox")</code>, too, just the way you’d expect it to.</p><p>Work is underway to enable the use of the <a href="https://www.elastic.co/docs/reference/text-analysis/analysis-lang-analyzer">36 dedicated language analyzers</a>, with support for natural languages on data that was never indexed or mapped.</p><h2>Full-text search use cases for unindexed and unmapped data</h2><p>The examples above searched values computed from <a href="https://www.elastic.co/docs/manage-data/data-store/mapping">mapped fields</a>. . The more interesting use cases for ES|QL <code>MATCH</code> on expressions involve data that was never searchable at all. Let's walk through a few.</p><h3>How to search unmapped fields in ES|QL without adding a mapping</h3><p>Sometimes you deliberately leave a field out of your mappings, such as a verbose stack trace or a raw request payload. You might even leave out a debug blob. Indexing one of these would cost disk and heap space on every document and wouldn’t be worth it for a field you might query once a quarter.</p><p>That decision has always been final, because <a href="https://www.elastic.co/search-labs/blog/esql-unmapped-fields">unmapped fields</a> were invisible to queries entirely. In Elasticsearch 9.5, you can use <code>SET unmapped_fields="load"</code> to make ES|QL load unmapped fields directly from the source document as keywords. Follow that up by wrapping it in <code>TO_TEXT</code>, and now you can run full-text search on it:</p><p>Here, <code>stack_trace</code> was never mapped. Every value is fetched from the original documents and analyzed on the fly. They’re matched row by row. That’s real work, and it will never be as fast as an <a href="https://www.elastic.co/docs/manage-data/data-store/index-basics">inverted index</a> lookup. But now, that field you didn't index is no longer unsearchable. You get to keep the mapping small for the everyday case and still answer the once-a-quarter question when it matters.</p><h3>Full-text search on a keyword field without reindexing</h3><p>Keyword fields can do a lot. They give you exact matching, fast aggregations, and sorting, which is why so many fields end up mapped that way. But mappings are decided when data arrives, and it’s easy to end up in a situation where you want to do something different with your data than you had originally intended. Maybe <code>product_name</code> was mapped as a <code>keyword</code> because the dashboards aggregate on it, and then, after receiving a year’s worth of product data, someone wants to be able to search within <code>product_name</code> values.</p><p>The old answer was to change the mapping to <code>text</code> (or add a multi-field) and reindex everything. This can be both time-consuming and costly, and in many cases, users simply won’t want to bother with it. The new answer is one function call:</p><p><code>TO_TEXT</code> converts the <code>keyword</code> values to <code>text</code> on the fly, so <code>MATCH</code> analyzes them instead of comparing them exactly. This allows you to query a <code>keyword</code> field without creating a mapping or reindexing the source document. If the search becomes an everyday query, indexing the field as <code>text</code> is still the right long-term move, but <code>TO_TEXT</code> gets you an answer today, without any extra work.</p><h3>Searching the same field across indices with different mappings</h3><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-multi-index">ES|QL can span many indices</a>, and the same field doesn't need to always look the same in all of them. When the same field has different types in different indices, ES|QL treats it as a union type, and a conversion function resolves the conflict. Let’s consider an example in which the <code>message</code> field has type <code>text</code> in this year's index template but was a <code>keyword</code> in last year's:</p><p>Every value is analyzed at query time, whether it came from the <code>text</code> index or the <code>keyword</code> one. The keyword values from the older indices are tokenized and lowercased like everything else, so "connection reset" finds “Connection RESET by peer”, no matter which index it lives in.</p><p>Another interesting case is when a field is mapped in only one index but also present (and unmapped) in the other:</p><p>There's a nuance worth calling out here. If <code>error_details</code> is mapped in <code>logs-2026</code> but not in <code>logs-2025</code>, Elasticsearch cannot push this query down to <a href="https://lucene.apache.org/">Lucene</a>, because the indices where the field is unmapped would silently return no matches. Instead, the planner notices that the field is potentially unmapped and evaluates the whole <code>MATCH</code> row by row, wherever the rows came from. You don't have to know which of your indices have the field mapped; the query just answers the question.</p><h2>How ES|QL analyzes text at query time without an inverted index</h2><p>When ES|QL plans a <code>MATCH</code> against an expression, it analyzes the query string once, up front, into a set of terms. How each row is then evaluated depends on the expression's type:</p><p><strong>Expression type</strong></p><p><strong>Processing</strong></p><p><strong>Matching behavior</strong></p><p><code>text</code> (via <code>TO_TEXT</code>)</p><p>Analyzer tokenizes value into lowercase terms</p><p>Token-against-token comparison; a row matches if any token equals any query term (<code>OR</code> semantics)</p><p><code>keyword</code>,<code>ip</code>, <code>date</code>, numeric</p><p>No analysis; query constant converted once to the native type</p><p>Exact comparison per row</p><p>Both paths bypass Lucene entirely and evaluate values row by row. The non-text path mirrors exactly what a match query does when pushed down to Lucene against those field types, so the semantics stay consistent regardless of whether your query hits an index.</p><p>An inverted-index lookup does its work at ingest time and never touches non-matching documents at query time. A runtime <code>MATCH</code> does that analysis at query time, for every row that reaches it. One is fast because the work already happened; the other is flexible because the data doesn't need to have been indexed at all.</p><h2>What's next for ES|QL full-text search</h2><p>Everything in this post is the first installment of a larger effort to make search in ES|QL work on anything, not just on what you indexed ahead of time. The limitations called out earlier are actively being worked on, and the roadmap goes further:</p><ul><li><p><strong>Scoring.</strong> Runtime matches will contribute to <code>_score</code>, so you can sort by relevance even when the data was never indexed.</p></li><li><p><strong><code>MATCH_PHRASE</code></strong><strong> on expressions.</strong> Already available in Elastic Cloud Serverless, and coming to the Elastic Stack in 9.6.</p></li><li><p><strong>Configurable analyzers.</strong> Analyzer support for <code>MATCH</code> and <code>MATCH_PHRASE</code> on expressions, enabling language analyzers, stemming, and synonyms at query time.</p></li><li><p><strong>Match options.</strong> Options like <code>fuzziness</code> and <code>operator</code> for runtime matches.</p></li><li><p><strong>Vector search.</strong> Generating embeddings per row and running k-nearest neighbors (kNN) on runtime <code>dense_vector</code> expressions, bringing semantic search to unindexed data, too.</p></li></ul><h2>Try ES|QL full-text search on expressions today</h2><p>You can try runtime search today. It's available now in Elastic Cloud Serverless, where new ES|QL capabilities land first, and it ships as a technical preview in Elasticsearch 9.5. Start with the <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/search-functions">search functions</a> reference, and check the <a href="https://www.elastic.co/docs/reference/query-languages/esql/limitations">ES|QL limitations</a> page for the current boundaries. It's a technical preview because we want your feedback: If you search something that was never indexed and it surprises you, either positively or negatively, <a href="https://www.elastic.co/community">we'd love to hear about it</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/full-text-search-unindexed-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/full-text-search-unindexed-data</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <dc:creator><![CDATA[Kevin Corcoran,Ioana Tagirta]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a74e633d05585a8/6a730774b8c2e64c3ebe0fd4/image1.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Close enough is fast enough: How ES|QL Fast mode makes Kibana dashboards up to 100x faster]]></title>
    <description><![CDATA[Fast mode samples a fraction of the data instead of scanning all of it. This release also brings click-to-filter for ES|QL charts, query-powered controls, and cleaner metric and bar chart layouts.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch Query Language (ES|QL) STATS queries on Kibana dashboards now run up to 100x faster. ES|QL Fast mode, in general availability (GA) in Kibana 9.5, samples a fraction of the data rather than scanning every row, and results stay within a 90% confidence interval. With Fast mode, ES|QL charts pick up click-to-filter and Discover drilldowns. Plus, controls can pull their values from an ES|QL query, and metric and bar chart defaults are cleaner. This builds on the dashboard improvements<a href="https://www.elastic.co/search-labs/blog/kibana-dashboards-improvements"> shipped in 9.4</a>. The <a href="https://www.elastic.co/search-labs/blog/dashboards-as-code-kibana-api">Dashboards API</a> and <a href="https://www.elastic.co/search-labs/blog/ai-dashboards-kibana-vega-lite">AI dashboards and Vega-Lite charts</a> also go GA in this release.</p><h2>ES|QL charts performance and interactivity in Kibana dashboards</h2><h3>How ES|QL Fast mode runs dashboard queries up to 100x faster</h3><p>For common analytical tasks, like trend tracking, top-host identification, and capacity overviews, trading a small margin of accuracy for dramatically faster results is the right call, especially since not every question needs an exact answer.</p><p><a href="https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-1">Elastic Search 9.4 introduced approximate ES|QL queries</a> as a syntax-only command in technical preview. Now, 9.5 makes approximation GA and adds<a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql-kibana#esql-kibana-fast-mode-toggle"> <strong>Fast mode</strong></a>, a UI toggle in Dashboards and Discover that enables<a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-query-approximation"> approximate ES|QL STATS queries</a> without writing any query syntax. This makes Kibana one of the first tools to offer smart sampling with automatic extrapolation as a simple switch.</p><p>Fast mode is an Enterprise-only feature and is off by default. Dashboard authors can save their preferred state with the dashboard, and individual queries can override the toggle with <code>SET approximation=true</code> or <code>false</code> inline.</p><p>When switched on, ES|QL STATS queries target a fixed sample size (defaulting to 1,000,000 rows for grouped aggregations and 100,000 rows otherwise) rather than scanning the full dataset.<a href="https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-1"> Benchmarks show heavy aggregations running up to 100x faster</a> on large datasets, with results that are typically highlyaccurate, defaulting to a 90% confidence interval. Approximation only applies to STATS commands where results can remain accurate. When accuracy cannot be ensured (such as with small datasets or aggregations like MAX, MIN, or COUNT_DISTINCT), Kibana automatically falls back to exact execution, even with Fast mode enabled.</p><p>Further improvements to how charts communicate that results are approximate are coming in future releases.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt473f75b25868afa1/6a719ccf75ed4699484a85dc/image2.png" alt="Kibana Fast mode toggle set to ON showing the approximation tooltip on a dashboard with metric panels" /><h3>Click-to-filter and Discover drilldowns for ES|QL charts</h3><p>Two of the most popular interactions for data view charts are also landing now for ES|QL-based visualizations.</p><ul><li><p><strong>Discover drilldowns</strong> now work on ES|QL panels. When a user clicks a data point or uses Explore in Discover, filters are translated to ES|QL <code>WHERE</code> clauses and Kibana Query Language (KQL) queries are carried over automatically. </p></li><li><p><strong>Click-to-filter also works for renamed fields:</strong> if your query renames a column (<code>STATS BY node = k8s.node.name</code>), Kibana now resolves the alias back to the indexed field, so the filter applies correctly.</p></li><li><p><strong>Tooltips:</strong> When filtering genuinely can't work (for example, because the field was computed entirely within the query and doesn't exist in the index), Kibana now shows a tooltip explaining why, so users know that it's a query limitation.Beyond interactivity, ES|QL layers now have the same <strong>Use global filters</strong> toggle (gear icon on the layer header) as data-view-backed visualizations. When you turn it off, the layer's query runs independently of dashboard-level filters, just like form-based layers already do. This is useful for reference lines, thresholds, or baselines that shouldn't change when you filter the dashboard. And ES|QL metric charts now support a background chart, matching the styling option already available for data view metrics.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt960ef7848c551ff0/6a719cfc3931bc448f28de69/image9.gif" alt="Kibana dashboard in edit mode with Host, OS, Cloud Provider and Region controls populated by ES|QL queries" /><p>Beyond interactivity, ES|QL layers now have the same <strong>Use global filters</strong> toggle (gear icon on the layer header) as data-view-backed visualizations. When you turn it off, the layer's query runs independently of dashboard-level filters, just like form-based layers already do. This is useful for reference lines, thresholds, or baselines that shouldn't change when you filter the dashboard. And ES|QL metric charts now support a background chart, matching the styling option already available for data view metrics.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf557d2a4578741bc/6a719e42e35d0253ce0312ec/image8.gif" alt="Kibana dashboard showing metric charts with Default density layout and preset style templates applied" /><p></p><p>Upcoming releases aim to keep adding the remaining functionality to ES|QL visualizations, such as multilayer support and saving visualizations to the library.</p><h3>Identify which Kibana panels use an ES|QL variable</h3><p><a href="https://www.elastic.co/search-labs/blog/kibana-dashboard-interactivity-variable-controls-overview">Variable controls</a> are among the most popular ES|QL-only features, since they let you parameterize chart queries to switch between fields, time intervals, or groupings without duplicating panels. On a dashboard with many panels and controls, though, it can be hard to tell which visualizations a variable actually affects. In edit mode, you can now click an ES|QL variable control's label to identify all related panels that consume the variable. Variables with no related panels display a warning to make it easier to audit wiring before saving.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5e565e80cda5cfce/6a71a13eb966e1768c63d81b/image5.gif" alt="Kibana dashboard with Host, OS and Cloud Provider controls filtering CPU and memory charts in view mode" /><h2>Kibana metric and bar chart layout defaults</h2><h3>Metric chart preset layouts and density options</h3><p>The metric chart appearance panel now offers preset layouts: <strong>Top</strong>, <strong>Middle</strong>, <strong>Bottom</strong>, and <strong>Custom</strong>. When you pick a template, the layout snaps into place. If you need fine-grained control, switch to <strong>Custom</strong>.</p><p>Metrics used to pack values tightly, which is great for data-dense dashboards but hard to scan when a metric stands alone. Elastic Cloud 9.5 adds a <strong>Density</strong> style option under <strong>Style &gt; Details &gt; Other</strong>, with two presets: <strong>Compact</strong> (the previous layout) and <strong>Default</strong> (more padding, larger typography). Newly created metrics use <strong>Default</strong>, and existing saved charts keep <strong>Compact</strong> until you change them.</p><p><strong>Attribute</strong></p><p><strong>Compact</strong></p><p><strong>Default</strong></p><p>Padding</p><p>Tight, minimal spacing</p><p>More generous whitespace</p><p>Typography</p><p>Smaller text</p><p>Larger text</p><p>Best for</p><p>Data-dense dashboards with many metrics side by side</p><p>Standalone metrics or dashboards with fewer panels</p><p>New charts</p><p>Must be selected manually</p><p>Applied automatically</p><p>Existing charts</p><p>Preserved until changed</p><p>Must be selected manually</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7434a5894cd1a048/6a719e8478b5febf378f12bc/image6.gif" alt=" Kibana dashboard edit mode showing the Settings gear icon on an ES|QL metric panel with global filter controls" /><h3>Responsive bar chart labels in Kibana</h3><p>Labels in horizontal bar charts used to grow unchecked, so on smaller screens, a chart with long category names could become unreadable. Bar labels now get a max width and middle-truncate automatically, so the beginning and end of a label stay visible even when the full text doesn't fit. This works by default, with no configuration needed. In 9.6, we’re planning many more improvements to bar charts and labels.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt85aa55cba5eb3a78/6a719ea65f2918842f13c2cb/image7.gif" alt="Before and after comparison of Kibana horizontal bar chart labels truncating responsively on smaller screens" /><h2>Kibana dashboard controls populated by ES|QL queries</h2><p><a href="https://www.elastic.co/docs/explore-analyze/visualize/add-controls#create-and-add-options-list-and-range-slider-controls">Controls</a> are the most user-friendly way to filter a dashboard, and most dashboards use them. One of the longest-standing requests from users has been the ability to prefilter the values that a control shows. ES|QL queries make that possible and open a much wider set of possibilities, like chaining controls in new ways using variables. Regardless of how the values are populated, controls filter every panel on the dashboard, including ES|QL and data view visualizations.</p><p>Controls can now be populated from an ES|QL query instead of selecting a data view field directly. The <strong>Create control</strong> flyout adds a <strong>Select a field / Write a query</strong> toggle. You can write an ES|QL query that returns a single column and run it, and then the control derives its options from the result. Queries can reference dashboard <a href="https://www.elastic.co/docs/explore-analyze/visualize/add-variable-controls">variables</a> through the <code>?variable</code> syntax, enabling flexible chaining between controls.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltacb5b26c2b31f3e6/6a719ebdded0cff1f5f49290/image1.png" alt="Kibana Edit control flyout showing an ES|QL query populating a Host options list on a dashboard" /><h2>Coming soon: Progress bar visualization for Kibana tables</h2><p>A new progress bar visualization type is available in Elastic Cloud Serverless and is planned for general availability in 9.6. Progress bars show a value relative to a goal or maximum, which is useful for many O11y metrics, like CPUs, memory, or Service Level Agreement (SLA) tracking.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt991d8a3c7d03068d/6a719ed7c2c8edb0c708b82a/image4.png" alt="Kibana table visualization with progress bar cell decoration showing Average Bytes per request path" /><h2>What's next for Kibana dashboards and ES|QL visualizations</h2><p>Upcoming releases will keep pushing on better defaults, improving the ES|QL visualization experience, and adding new chart types. If you have a pain point or a feature request, select <strong>Submit feedback</strong> in the top menu; we're listening.</p><h2>How to try ES|QL Fast mode and the new Kibana dashboard features</h2><p>If you use <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a>, you may already be using these changes. Otherwise, upgrade to 9.5, and then create a dashboard or open an existing one. Many updates apply automatically to new visualizations, while layout and style options appear in edit mode. If you aren't on Elastic Cloud yet, <a href="https://cloud.elastic.co/registration">start a trial</a> and explore the latest Kibana dashboards there.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/kibana-dashboards-esql-fast-mode</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/kibana-dashboards-esql-fast-mode</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Analytics]]></category>
    <dc:creator><![CDATA[Teresa Alvarez Soler]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8b0336f51b4694f0/6a719cab5e874b5b0e1ab976/image3.png" length="0" type="image/png"/>
    <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Prompt to dashboard in under a minute, 5x cheaper: AI dashboards and custom Vega-Lite charts in Kibana]]></title>
    <description><![CDATA[Describe your metrics in natural language and Kibana's AI chat generates ES|QL-backed dashboards and Vega-Lite charts, from scatter plots to conditional formatting and custom tooltips.]]></description>
    <content:encoded><![CDATA[<p>Kibana's<a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/chat"> AI chat</a> now builds full<a href="https://www.elastic.co/docs/explore-analyze/visualize/esorql"> Elasticsearch Query Language–backed (ES|QL-backed)</a> dashboards from a natural-language prompt in under a minute. In Elastic 9.5, this moves to general availability (GA) (<a href="https://www.elastic.co/search-labs/blog/ai-dashboard-generation-elastic-agent-kibana">technical preview in 9.4</a>) with error recovery that retries failed queries, 5x lower ES|QL generation costs through tiered model routing, and interactive filter controls. This release also adds<a href="https://www.elastic.co/docs/explore-analyze/visualize/custom-visualizations-with-vega"> Vega-Lite</a> chart creation through natural language, including scatter plots, box plots, conditional formatting, and custom tooltips that you'd normally have to hand-code in JSON.</p><h2>What's new in Kibana's AI dashboard creation</h2><p><strong>Capability</strong></p><p><strong>Technical preview (9.4)</strong></p><p><strong>GA (9.5)</strong></p><p>Error handling</p><p>No retry on failed ES|QL queries</p><p>Automatic retry up to three times with query inspection and adjustment</p><p>ES|QL generation cost</p><p>All queries routed through the primary model</p><p>Tiered model routing, up to 5x cheaper</p><p>Time range</p><p>Fixed default window</p><p>Automatic selection based on data time distribution</p><p>Filter controls</p><p>Not supported</p><p>Automatically added for the most relevant fields</p><p>Vega-Lite charts</p><p>Not supported</p><p>Natural-language creation, including scatter plots, box plots, conditional formatting, custom tooltips</p><p>Chart editing</p><p>Not supported</p><p>Edit existing Vega-Lite panels through natural language</p><h3>Automatic error recovery for AI dashboard generation</h3><p>In the technical preview, the agent didn’t retry failed ES|QL queries. In 9.5, it detects query errors and retries up to three times, inspecting each error and adjusting the query before giving up. In practice, this eliminates the majority of empty-panel issues and produces dashboards that render correctly on the first try.</p><h3>Why is AI dashboard creation cheaper in Elastic 9.5?</h3><p>Not every step in dashboard generation needs the same level of reasoning. In 9.5, ES|QL generation routes through a lighter model by default and falls back to the primary model only when needed. If your <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/connectors">connector</a> uses Anthropic's Claude Opus 4.8, that means 5x cheaper for ES|QL generation across all panels.</p><h3>Automatic time range selection based on your data</h3><p>Dashboards are only useful when they show the right window of data. The agent now applies improved logic to pick a time range that makes sense for the data it's querying, unless the user asks for a specific time range. It considers the data's time distribution and adjusts accordingly, whether that means the last hour for a live incident or the last 90 days for a trend analysis, rather than defaulting to a fixed window.</p><h3>Automatic filter controls on AI-generated dashboards</h3><p>Dashboard creation now supports controls; that is, interactive filters that let viewers narrow a dashboard by field values without editing the underlying queries. When generating a dashboard, the agent automatically adds controls at the top for the fields most relevant to filter by.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4a5520e66ae76001/6a719a218a155220ed6498e7/image3.png" alt="Kibana AI chat generating an ES|QL-backed host metrics dashboard with automatic filter controls in 71 seconds" /><h2>Vega-Lite charts from natural language: Chart types and formatting beyond the defaults</h2><p><a href="https://vega.github.io/vega/">Vega</a> and <a href="https://vega.github.io/vega-lite/examples/">Vega-Lite</a> support a wide range of chart types and customizations in Kibana. With 9.5, you can build them from plain language instead of writing the code yourself. </p><h3>Scatter plots, box plots, and more Vega-Lite chart types</h3><p>Scatter plots, box plot charts, faceted small multiples, bubble charts, and composition charts (like combining histograms with heatmaps), among many others, are supported by <a href="https://vega.github.io/vega-lite/examples/">Vega-Lite</a>. A prompt like <em>Show me a scatter plot of response time vs. request size, colored by service name</em> produces a Vega-Lite panel with the right data mappings. They use Kibana's default color palettes to blend with the rest of the dashboard.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf1651fe7ac903d47/6a719a49f124649f746fc1b4/image5.png" alt="Kibana dashboard with four Vega-Lite charts: box plot, bubble chart, faceted small multiples, and heatmap." /><h3>Conditional formatting, custom tooltips, and labels on standard charts</h3><p>Even for chart types that are already native to dashboards, like bar, line, or area, sometimes you need more control than the default capabilities offer. Vega-Lite through the chat fills that gap. Some examples include:</p><ul><li><p><strong>Conditional color formatting:</strong> Color data points above a threshold differently; for example, turning data points in a line or bars red when some metric spikes beyond your Service Level Objective (SLO). Ask the agent something like <em>Turn any points above 500ms red for my line chart.</em> </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1211784e9033557/6a719a6ab966e1736163d7d1/image1.png" alt="Vega-Lite line and bar charts in Kibana with conditional colour formatting showing data points above a threshold in red" /><p></p></li><li><p><strong>Custom marks and labels:</strong> Add emojis, symbols, or inline text labels to data points for at-a-glance status indicators.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte3c59c3748484bc1/6a719aa02888394fdc07bac1/image2.png" alt="Lite horizontal bar chart in Kibana with emoji flag labels and custom tooltip showing requests by country" /><p></p></li><li><p><strong>Custom tooltips:</strong> Enrich hover states with additional metrics, context, or computed values that aren't part of the chart's axes. Ask something like <em>Add a tooltip that shows total record counts and the % per bar.</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt43f241f4382b5b90/6a719ab7ded0cf3367f49275/image4.png" alt="Vega-Lite stacked bar chart in Kibana with custom tooltip showing total records and percentage of total by extension" /><p></p></li></ul><p>This also works for editing existing Vega charts. If you have a Vega-Lite panel that needs a tweak like changing a color scale, adjusting an axis, or switching the mark type, describe the change in chat instead of digging into the JSON code.</p><h2>How we built natural-language Vega-Lite generation in Kibana</h2><p>Generating a Vega-Lite chart from a sentence is not a one-shot <em>ask the model for JSON</em> prompt. We built a small agentic pipeline that turns natural-language intent into a validated, data-backed chart.</p><p>When a request comes in, the agent first determines whether Vega-Lite is the right fit. For Vega-Lite requests, it grounds the visualization in a real ES|QL query against Elasticsearch and then uses a model to generate the Vega-Lite code. Before rendering, the result goes through a normalization layer that corrects the schema and binds the canonical query. It also applies render-safety transformations. </p><p>A few design choices make this workflow reliable:</p><ul><li><p><strong>Typed tool calling</strong>: Chart creation is a structured tool invocation rather than free-form Vega-Lite pasted into the conversation.</p></li><li><p><strong>Constrained generation</strong>: The model generates Vega-Lite code within a defined schema, making the output more predictable and easier to validate.</p></li><li><p><strong>Curated examples:</strong> Structural patterns, such as faceting, layered marks, and heatmaps, provide guidance without copying the underlying data.</p></li><li><p><strong>Execute-and-verify loops</strong>: Queries are executed before chart authoring, and validation failures trigger corrective retries for ES|QL generation.</p></li></ul><h2>Try AI dashboard creation and Vega-Lite charts in Kibana</h2><p>To try natural-language dashboard creation and Vega-Lite charts, upgrade to <strong>Elastic 9.5</strong> (or <a href="https://cloud.elastic.co/registration">start a free trial</a>), and open the <strong>chat</strong> in Kibana. Then ask it to build a dashboard from your data. For Vega-Lite, try asking for a chart type you've wanted but never built, like a scatter plot or a bubble chart. If the result isn't quite right, tell the agent what to change. It iterates with you.</p><p>This requires an Enterprise license. <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/chat#get-started">Get started</a>.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-dashboards-kibana-vega-lite</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-dashboards-kibana-vega-lite</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Marta Bondyra,Teresa Alvarez Soler]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt54406ad0378bc5fc/6a7199ffed03ccee0dac9d7c/image6.png" length="0" type="image/png"/>
    <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[15 lines of click tracking code that tell you what search logs can't ]]></title>
    <description><![CDATA[Three ES|QL queries calculate click-through rate, mean reciprocal rank and click position distribution from your search click data, so you can pinpoint which queries need relevance tuning and where ranking improvements will have the most impact.]]></description>
    <content:encoded><![CDATA[<p>Search volume and latency tell you that search is working, not that it's useful. About 15 lines of OpenTelemetry (OTel) instrumentation lets you track clicks on search results and then query click-through rate (CTR), Mean Reciprocal Rank (MRR), and click position distribution with Elasticsearch Query Language (ES|QL) against the same traces index that your search spans already live in. You'll wire click tracking to your existing <code>search.query_id</code> and write the queries that show which searches need relevance tuning.</p><h2>What you'll discover</h2><p>In this post, you'll learn how to:</p><ul><li><p>Add client-side click tracking that links clicks back to their originating search via <code>search.query_id</code>.</p></li><li><p>Calculate CTR; that is, the percentage of searches that produce at least one click.</p></li><li><p>Calculate MRR; that is, how far down the results users click on average.</p></li><li><p>Analyze click position distribution to see the full shape of user engagement.</p></li><li><p>Write ES|QL queries for all three metrics against  <code>traces-generic.otel-default</code>.</p></li><li><p>Identify which specific queries need relevance tuning.</p></li></ul><h3>What you'll need</h3><ul><li><p>A working OTel instrumentation setup from Blog 2 (search spans with <code>search.*</code> attributes flowing to Elastic via OTel-native ingestion).</p></li><li><p>A front end that can send click events (JavaScript example provided).</p></li><li><p>Familiarity with the <code>attributes.*</code> field mapping from Blog 2.</p></li></ul><h2>Why search logs alone can't measure search quality</h2><p>In the <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql">second blog</a> of the series, we instrumented search requests and ran six ES|QL queries against the data. We can see what users search for, which queries return nothing, and how fast search is.</p><p>But there's a blind spot. A search that returns 15 results looks healthy from the server side. Every metric we have says it worked. But if nobody clicks any of those results, your ranking has a problem, and none of the queries from Blog 2 will tell you.</p><p>This is the gap between <em>results returned</em> and <em>results that are useful</em>. Search volume, zero-results rate, and latency measure the mechanics of search, but they don't measure whether search is actually helping users find what they need.</p><p>To answer that question, you need a second instrumentation point: <em>click tracking</em>.</p><p>If you’re following along with code, the <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">reference project</a> has click tracking ready to enable. Uncomment the Blog 3 sections in <code>app.py</code> and <code>frontend/app.js</code>, restart, and then generate traffic with <code>python generate_traffic.py --blog 3</code>.</p><h2>How click data measures search relevance</h2><p>Before we write any code, here's what click data lets you measure and why each metric matters:</p><ul><li><p><strong>CTR</strong> answers the most basic engagement question: <em>What percentage of searches result in at least one click?</em> If your CTR is low, users are seeing results but not finding them compelling enough to engage. Establish your own baseline once you have data; CTR varies considerably across product category, query type, and industry vertical.</p></li></ul><ul><li><p><strong>MRR</strong> goes deeper: <em>When users do click, where in the results are they clicking?</em> An MRR of 1.0 means every user clicks the top result (a perfect ranking). An MRR of 0.5 means the average click is at position 2. Low MRR with high CTR is particularly telling. It means that users are finding what they need, but your ranking is making them work for it.</p></li></ul><ul><li><p><strong>Click position distribution</strong> shows the full shape of where users click. A healthy search engine shows most clicks at position 1 with a sharp drop-off. A flat distribution across positions 1–5 means that your ranking isn't differentiating well. Per-query distributions reveal exactly which searches need relevance tuning.</p></li></ul><p>Together, these metrics move you from <em>Did search work?</em> (Blog 2) to <em>Did search work well?</em>, and they pinpoint exactly where to invest in relevance improvements. Later in the series, we'll show how to turn these metrics into concrete actions: building judgment lists for Learning To Rank (LTR), tuning relevance with tools like <a href="https://elastic.github.io/relevance-studio/#/">Elasticsearch Relevance Studio</a>, and evaluating changes with the Rank Eval API. But first, you need the data.</p><p>All three metrics require just one new instrumentation point: about 15 lines of code.</p><h2>Add click tracking</h2><p>Click tracking captures what happens after the results appear. When a user clicks a search result, we create a new span with attributes describing the interaction, including which document they clicked, where it appeared in the results, and which search produced it.</p><p>Here's the code:</p># Track which query_ids have already received a click
_clicked_queries: set[str] = set()

@app.post("/api/events")
async def track_event(event: EventRequest):  # reference project uses ClickEvent
with tracer.start_as_current_span("search.result.click") as span:
        span.set_attribute("search.action", "click")
        span.set_attribute("search.result_click_id", event.object_id)
        span.set_attribute("search.result_click_position", event.position)
        span.set_attribute("search.result_click_type", event.object_id_type)
        span.set_attribute("search.query_id", event.query_id)
        span.set_attribute("enduser.pseudo.id", event.client_id)

# First click per search — enables single-query CTR
if event.query_id not in _clicked_queries:
            span.set_attribute("search.first_click", True)
            _clicked_queries.add(event.query_id)

if event.user_query:
            span.set_attribute("search.query", event.user_query)<p>Let's unpack what matters.</p><h3>Click spans are separate traces</h3><p>This is the key architectural difference from Blog 2. Search spans are created synchronously during the API request: The user searches, the span opens, Elasticsearch responds, and the span closes. Click spans are <em>asynchronous</em>. The user searches, gets results, browses the page, and might click 30 seconds later (or they might never click).</p><p>That means click spans aren't children of the search span's trace. They're independent traces, linked to the originating search through <code>search.query_id</code>. This is the same <code>query_id</code> we derived from the trace ID in Blog 2, and it now serves as the join key between searches and clicks across <code>traces-generic.otel-default</code>.</p><h3>Choosing the right OTel signal for clicks</h3><p>OTel gives you three signal types, and clicks could be modeled as any of them. Each has strengths:</p><p><strong>Signal</strong></p><p><strong>Index</strong></p><p><strong>Weight</strong></p><p><strong>Best for…</strong></p><p>Spans</p><p><code>traces-generic.otel-default</code></p><p>Full trace context</p><p>Same-index queries with search spans</p><p>Logs</p><p><code>logs-*</code></p><p>Lighter weight</p><p>Log-centric pipelines, high volume</p><p>Span events</p><p><code>logs-generic.otel-default</code></p><p>Lightest instrumentation</p><p>Attaching to existing spans</p><p></p><ul><li><p><strong>Spans</strong> land in <code>traces-generic.otel-default</code> alongside search spans, are fully queryable in ES|QL, appear in Kibana APM views, and carry timing information. Since our search spans are already in <code>traces-generic.otel-default</code>, using spans for clicks means you can query searches and clicks together in a single ES|QL statement, and no cross-index joins are needed.</p></li><li><p><strong>Log records</strong> are also independently queryable in ES|QL, living in <code>logs-*</code>. If you use the same <code>search.*</code> attribute names, the analytics queries are almost identical; just change the index pattern. Logs are lighter weight (no trace context overhead) and are a natural fit if your team already has a log-centric observability pipeline. One of Elastic's strengths here is that traces, logs, and metrics all land in the same platform and are all queryable with ES|QL, so choosing logs over spans doesn't mean giving up any query capability.</p></li><li><p><strong>Span events </strong>are lightweight at instrumentation time (attached to an existing span in the OTel API). In Elastic's OpenTelemetry Protocol (OTLP) ingestion pipeline, span events are written as separate documents to <code>logs-*</code> data streams (for example <code>logs-generic.otel-default</code>), and they’re independently queryable in the same way as logs. They’re a good, lightweight option but might require more code changes than logs, which can even pull in logs from legacy code.</p></li></ul><p>In this series, we use <em>spans </em>because they keep searches and clicks in the same index with the simplest query path. But if you're at high volume and want to optimize for cost, or if your organization already routes OTel logs to Elasticsearch, the log-based approach works well; the <code>search.*</code> attribute schema is the same either way, and ES|QL queries against <code>logs-*</code> follow the same patterns you'll see below.</p><h3>The <code>search.first_click</code> attribute</h3><p><code>search.first_click</code> is a Boolean set only on the first click for a given <code>query_id</code>. It exists for one reason: accurate CTR calculation without post-processing.</p><p>CTR is defined as the percentage of searches with at least one click. Without <code>search.first_click</code>, you'd need to deduplicate clicks by <code>query_id</code> at query time: grouping, counting distinct values, and subquerying. By marking the first click at instrumentation time, the ES|QL query becomes a simple count.</p><p>The set above is demo-only. It grows unbounded and breaks with multiple API replicas. The <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">reference implementation</a> uses a thread-safe time-to-live (TTL) dict with a 30-minute expiry window (<code>_is_first_click()</code> in <code>app.py</code>). For production with multiple replicas, use a shared external cache (Redis, Memcached) keyed by <code>query_id</code> with a TTL matching your session window.</p><h3>Where to track first click: Front end vs. back end</h3><p>The <code>search.first_click</code> deduplication could live in either the front end or the back end. Both are valid, and here are the trade-offs:</p><ul><li><p><strong>Frontend tracking</strong> is simpler to implement. The browser already knows the current query and whether the user has clicked before. No server-side state is required, you don’t have to worry about multiple replicas, and it works without any backend changes. The downside is that browser state is ephemeral; a page refresh, multiple tabs, or an ad blocker can interfere with accurate tracking.</p></li></ul><ul><li><p><strong>Backend tracking</strong> (our approach) gives you a single source of truth. All click events flow through one place, so the deduplication is consistent regardless of what the client does. It also means that the analytics logic is colocated with the instrumentation code, which simplifies reasoning about data quality. The trade-off is that the back end needs to maintain state: the <code>_clicked_queries</code> set. For a single-instance API, this is trivial; for multiple replicas behind a load balancer, you'd use a shared TTL cache (Redis or similar).</p></li></ul><p>We chose backend tracking here because we want the analytics data to be authoritative. This click data will later feed into relevance tuning and judgment lists, where accuracy matters. But if you're starting simple or running a client-side–only setup, frontend tracking is a perfectly reasonable first step. The <code>search.first_click</code> attribute works the same way regardless of where you set it.</p><h3>Sending click events from the browser</h3><p>The browser sends click events to the back end when a user clicks a result. It needs three things from the search response: the document ID, the position, and the <code>query_id</code>.</p>// Generate a persistent client ID once per browser (stored in localStorage)
const CLIENT_ID = localStorage.getItem("search_client_id")
    || (() =&gt; {
const id = crypto.randomUUID();
        localStorage.setItem("search_client_id", id);
return id;
    })();

// On result click
fetch('/api/events', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    object_id: product.id,
    position: index + 1,       // 1-indexed
query_id: lastQueryId,     // from the most recent search response
client_id: CLIENT_ID,      // persistent browser identifier → enduser.pseudo.id
user_query: currentQuery,
    object_id_type: 'product', // optional; defaults to "product" on the backend
})
});<p><code>CLIENT_ID</code> is generated once and stored in <code>localStorage</code>. It survives page reloads and gives you a stable <code>enduser.pseudo.id</code> without requiring a login. The back end maps <code>client_id</code> → <code>enduser.pseudo.id</code> on the span.</p><p>Positions are 1-indexed; that is, the first result is position 1, not 0.</p><h3>Click tracking OTel attributes and ES|QL field mapping</h3><p></p><p><strong>Attribute</strong></p><p><strong>Type</strong></p><p><strong>Required</strong></p><p><strong>Purpose</strong></p><p><code>search.action</code></p><p>string</p><p>yes</p><p>Event type: <code>"click"</code></p><p><code>search.result_click_id</code></p><p>string</p><p>yes</p><p>Document ID clicked</p><p><code>search.result_click_position</code></p><p>int</p><p>yes</p><p>Position in results (1-indexed)</p><p><code>search.query_id</code></p><p>string</p><p>yes</p><p>Links to originating search</p><p><code>enduser.pseudo.id</code></p><p>string</p><p>yes</p><p>Client/device identifier</p><p><code>search.first_click</code></p><p>boolean</p><p>recommended</p><p><code>true</code> if first click for this <code>query_id</code></p><p><code>search.result_click_type</code></p><p>string</p><p>recommended</p><p>Object type: <code>"product"</code>, <code>"article"</code></p><p><code>search.query</code></p><p>string</p><p>recommended</p><p>The search query text (for queryability)</p><p>These follow the same <code>search.*</code> namespace we established in Blog 2. With OTel-native ingestion, attributes map directly to <code>attributes.*</code> fields in ES|QL:</p><p></p><p><strong>OTel attribute</strong></p><p><strong>ES|QL field</strong></p><p><code>search.action</code></p><p><code>attributes.search.action</code></p><p><code>search.result_click_id</code></p><p><code>attributes.search.result_click_id</code></p><p><code>search.result_click_position</code></p><p><code>attributes.search.result_click_position</code></p><p><code>search.query_id</code></p><p><code>attributes.search.query_id</code></p><p><code>search.first_click</code></p><p><code>attributes.search.first_click</code></p><p>With OTel-native ingestion, <code>search.first_click</code> is stored as a native Boolean; you query it with <code>== true</code>, not <code>== "true"</code>, and no string coercion is needed.</p><h3>Verify that clicks are arriving</h3><p>Before calculating metrics, confirm that click spans are flowing to APM:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
| KEEP attributes.search.result_click_id, attributes.search.result_click_position,
       attributes.search.query_id, attributes.search.query
| LIMIT 5<p>If this returns rows, you're ready for analytics. If not, check the same things as we looked at in Blog 2: OTLP endpoint, auth token, and span export.</p><p><strong>Note:</strong> Results in this post are illustrative, generated by running <code>python generate_traffic.py --blog 3 --sessions 50</code> on the <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">reference project</a>. Running Blog 3 traffic adds click events <em>and</em> additional search sessions on top of the 62 from Blog 2, so cumulative search counts will exceed 62. Your exact numbers will vary based on session count and the random nature of the traffic simulator. The metric calculations and ES|QL patterns are what to focus on.</p><h2>CTR</h2><p><strong>CTR</strong> is the primary signal for search relevance, answering the question that Blog 2 couldn't: <em>Are users engaging with the results?</em></p><p><strong>CTR = searches with at least one click / total searches * 100</strong></p><p>CTR is a binary per-search metric: Either a search got clicked or it didn't. The maximum is 100%.</p><h3>Overall search CTR with ES|QL</h3>FROM traces-generic.otel-default
| WHERE (name == "search" AND attributes.search.query IS NOT NULL)
OR attributes.search.first_click == true
| STATS
    searches = COUNT(CASE(name == "search" AND attributes.search.query IS NOT NULL, 1)),
    clicked = COUNT(CASE(attributes.search.first_click == true, 1))
| EVAL ctr_pct = ROUND(100.0 * clicked / searches, 1)<p><strong>Result:</strong> 41 clicked searches out of 146 total. <strong>CTR: 28.1%</strong></p><p>This is a single query that pulls both search spans and first-click spans from <code>traces-generic.otel-default</code>. The <code>OR</code> in the <code>WHERE</code> clause brings both into one result set. <code>COUNT(CASE(...))</code> counts each type separately, and <code>EVAL</code> does the division.</p><p>This works because of <code>search.first_click</code>. Without it, you'd be counting raw clicks (a user who clicks three results on one search would inflate the count). The deduplication happened at instrumentation time; the query stays simple.</p><h3>CTR by search query: Finding your worst relevance failures</h3><p>The overall number is useful for dashboards. The per-query breakdown is where you find problems.</p>FROM traces-generic.otel-default
| WHERE ((name == "search" AND attributes.search.query IS NOT NULL)
OR attributes.search.first_click == true)
AND attributes.search.query IS NOT NULL
| STATS
    searches = COUNT(CASE(name == "search" AND attributes.search.query IS NOT NULL, 1)),
    clicked = COUNT(CASE(attributes.search.first_click == true, 1))
BY attributes.search.query
| EVAL ctr_pct = ROUND(100.0 * clicked / searches, 1)
| SORT searches DESC
| LIMIT 20<p>This shows CTR broken down by query text. </p><h3>What CTR tells you (and what it doesn't)</h3><ul><li><p><strong>High searches + zero clicks:</strong> These are the worst relevance failures. Fix these first.</p></li><li><p><strong>High searches + low CTR:</strong> Results appear, but they aren't compelling. Check ranking.</p></li><li><p><strong>Low CTR + high zero-results rate:</strong> This is a double problem; either no results or bad results.</p></li><li><p><strong>CTR trend over time:</strong> This measures the impact of relevance changes.</p></li></ul><p>CTR doesn't measure satisfaction. A user who clicks position 1, bounces back, and then clicks position 3 still counts as one clicked search. For a fuller picture, you need to know <em>where </em>they're clicking. That's what MRR measures.</p><h3>CTR vs. clicks per search</h3><p><strong>CTR</strong> (what we just calculated) is capped at 100%. This is the industry-standard definition.</p><p><strong>Clicks per search</strong> is total clicks divided by total searches. It can exceed 1.0. For example, a search where the user clicks three results scores 3.0. It measures engagement depth, which is useful but different. If you need it, count all click spans (not just <code>first_click</code>) divided by search spans.</p><h2>MRR</h2><p><strong>MRR</strong> tells you <em>where users </em>click. It measures how far down the results list users go before finding something worth clicking.</p><p><strong>MRR = average of (1 / click_position) across all clicks</strong></p><p>The reciprocal rank transforms click positions into a 0–to–1 scale, where higher is better:</p><p></p><p><strong>Click position</strong></p><p><strong>Reciprocal rank</strong></p><p>1</p><p>1.000</p><p>2</p><p>0.500</p><p>3</p><p>0.333</p><p>5</p><p>0.200</p><p>10</p><p>0.100</p><p></p><h3>Overall search MRR with ES|QL</h3><p>MRR can be calculated two ways, depending on what you want to measure:</p><ul><li><p><strong>All-click MRR:</strong> Averages the reciprocal rank of every click and reflects overall click quality, including repeated interactions.</p></li><li><p><strong>First-click MRR:</strong> Averages only the first click per search (using <code>search.first_click == true</code>). It’s more comparable to traditional information retrieval (IR) evaluation and aligns with how you computed CTR.</p></li></ul><p>For consistency with CTR and alignment with judgment-list workflows in Blog 5, we prefer first-click MRR:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
AND attributes.search.first_click == true
| EVAL reciprocal_rank = 1.0 / attributes.search.result_click_position
| STATS mrr = ROUND(AVG(reciprocal_rank), 3)<p><strong>Result:</strong> MRR = <strong>0.495</strong></p><p>This is decent but shows room for improvement. An MRR of 0.495 means the average first click lands around position 2. It isn’t a crisis, but there are queries where ranking can be improved.</p><h3>MRR by search query: Finding your worst-ranked results</h3><p>Like CTR, the per-query breakdown is where the actionable data lives. To surface the worst-ranked queries first, sort ascending.</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
AND attributes.search.first_click == true
AND attributes.search.query IS NOT NULL
| EVAL reciprocal_rank = 1.0 / attributes.search.result_click_position
| STATS
    mrr = ROUND(AVG(reciprocal_rank), 3),
    clicks = COUNT(*)
BY attributes.search.query
| SORT mrr ASC
| LIMIT 20<p>This reveals which queries have the worst ranking. A query with multiple clicks and low MRR means that the ranking is consistently poor for that search; that is, users find results, but they have to dig for them.</p><h3>What is a good MRR score for search?</h3><ul><li><p><strong>MRR &gt; 0.8:</strong> This ranking is solid; users usually click position 1–2.</p></li><li><p><strong>MRR 0.5–0.8:</strong>  This is decent, but there’s room for improvement.</p></li><li><p><strong>MRR &lt; 0.5:</strong> This is a ranking problem, and users are scrolling past top results.</p></li><li><p><strong>MRR drop after a change:</strong> This is ranking regression that should be investigated immediately.</p></li><li><p><strong>Low MRR + high CTR:</strong> Users are finding things, but they have to work for it.</p></li></ul><p>That last pattern is particularly interesting. High CTR with low MRR means your results are relevant (users are clicking), but your ranking isn't surfacing the best results first. It's an optimization opportunity, not a crisis.</p><h3>MRR limitations: Position bias and multi-click sessions</h3><p>MRR is heavily influenced by the gap between position 1 and position 2 (1.0 versus 0.5). Positions 5 and beyond barely move the average. This means that MRR is most sensitive to whether your top result is good, which is often exactly what you want to optimize.</p><p>MRR also only measures clicks, not satisfaction. With <em>all-click MRR</em>, a user who clicks position 1, bounces, and then clicks position 3 contributes two data points, but only the second was useful. The <em>first-click MRR</em> queries above avoid this by counting only the first click per search via <code>search.first_click == true</code>.</p><h2>Click position distribution</h2><p>Click position distribution shows you the full picture of where in the results users are engaging.</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
| STATS click_count = COUNT(*) BY attributes.search.result_click_position
| SORT attributes.search.result_click_position ASC<p><strong>Results:</strong></p><p><strong>Position</strong></p><p><strong>Clicks</strong></p><p>1</p><p>21</p><p>2</p><p>10</p><p>3</p><p>7</p><p>4</p><p>4</p><p>5</p><p>3</p><p></p><p>This is a reasonable distribution: 21 of 45 clicks (47%) land on position 1, with a tapering tail. If you paste this query into Discover's ES|QL editor, Kibana auto-generates a bar chart that makes the shape immediately visible.</p><h3>How to read click position distribution for search relevance</h3><ul><li><p><strong>Sharp dropoff after position 1:</strong> The ranking is effective, and the top result is usually right.</p></li><li><p><strong>Flat across positions 1–5:</strong> The ranking isn't differentiating well, and all positions are equally likely to be clicked.</p></li><li><p><strong>Spike at position 3+ for specific queries:</strong> Those queries have ranking problems.</p></li><li><p><strong>No clicks beyond position 5:</strong> Users don't scroll far. Top 5 ranking matters most.</p></li></ul><h3>Click position distribution by search query</h3><p></p><p>To see the shape for specific queries:</p>FROM traces-generic.otel-default
| WHERE attributes.search.action == "click"
AND attributes.search.query IS NOT NULL
| STATS click_count = COUNT(*)
BY attributes.search.query, attributes.search.result_click_position
| SORT attributes.search.query, attributes.search.result_click_position<p>A query where all clicks land on position 1 has perfect ranking. A query where clicks spread across positions 1–5 needs relevance tuning.</p><h3>Position bias and click models</h3><p>One caveat: Click position distribution is influenced by <em>position bias</em>; that is, users see position 1 first, so it gets clicked more regardless of relevance. A click at position 1 isn't necessarily more relevant, just more visible.</p><p>This is a well-studied problem in information retrieval. <em>Click models</em> are statistical models that attempt to separate genuine relevance from position bias in click data. The foundational work by <a href="https://www.cs.cornell.edu/people/tj/publications/joachims_etal_05a.pdf">Joachims et al. (2005)</a> showed that users are significantly biased toward higher-ranked results, and, in proposed methods like skip-above analysis (if a user clicks position 3 but skips positions 1 and 2), those skipped results are likely less relevant for that query.</p><p>For the metrics in this post, you don't need to implement a full click model. The key insight is practical: Compare distributions <em>between queries</em> rather than treating absolute position counts as ground truth. If query A has 80% of clicks at position 1 and query B has clicks spread across positions 1–5, query B's ranking is worse, even accounting for position bias. Later in the series, when we look at building judgment lists for LTR, position bias correction becomes more important, and the click data you're collecting here is exactly what those models need as input.</p><h2>CTR, MRR, and click distribution: Reading search quality metrics together</h2><p>These three metrics offer three different angles on search result quality:</p><p></p><p><strong>Metric</strong></p><p><strong>What it measures</strong></p><p><strong>Our value</strong></p><p><strong>Interpretation</strong></p><p><strong>CTR</strong></p><p>Do users click at all?</p><p>28.1%</p><p>Moderate: Roughly a third of searches get engagement, but there’s room to improve.</p><p><strong>MRR</strong></p><p>Where do they click?</p><p>0.495</p><p>Decent: The average click is around position 2, but ranking can be improved.</p><p><strong>Distribution</strong></p><p>What's the shape?</p><p>47% at position 1</p><p>Reasonable drop-off: The top result wins most but isn’t dominant.</p><p>Together, they tell a coherent story. For our demo data, search is performing adequately: Users are engaging with results and can find what they need, but the ranking has room to improve. The CTR of 28% and MRR of 0.495 are realistic starting points for a new search implementation without tuning.</p><p>Where they're most valuable is in combination at the query level. The queries to fix first are those with <strong>high volume + low CTR + low MRR</strong>; that is, lots of users are searching, few are clicking, and those who do click are scrolling deep. That's where relevance investment has the highest return.</p><h3>Using click data for LTR and relevance tuning</h3><p>These metrics don't just tell you how search is performing; they're the foundation for making it better. The click data you're now collecting feeds directly into relevance improvement workflows:</p><ul><li><p><strong>Judgment lists for LTR:</strong> Click positions and frequencies become graded relevance labels for training machine learning (ML) ranking models. A document clicked at position 1 across many queries is a strong positive signal.</p></li><li><p><strong>Relevance tuning tools:</strong> Per-query CTR and MRR tell you exactly which queries to focus on in tools like <a href="https://elastic.github.io/relevance-studio/#/">Relevance Studio</a> or the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-rank-eval.html">Rank Eval API</a>, which scores your ranking against expected results.</p></li><li><p><strong>Query rules and boosting:</strong> Zero-CTR queries with results are candidates for pinning, boosting, or synonym rules.</p></li></ul><p>We'll cover these applications in detail in Blog 5. For now, the important thing is that the instrumentation you've built here is doing double duty: It measures search quality <em>and</em> provides the training data to improve it.</p><h2>Next in the series: Conversion tracking from search to purchase</h2><p>We can now measure whether users find results (Blog 2) and whether they engage with them (this post). But a click isn't a conversion. A user who clicks a product and then abandons the page didn't get what they needed.</p><p>In the next post, we add <em>conversion tracking</em>, the third instrumentation point that closes the loop from search to purchase. It’s the same pattern: Add <code>search.*</code> attributes to add-to-cart and checkout spans, query with ES|QL, and answer the question your product manager actually cares about: <em>Which searches drive revenue?</em></p><h2>Get started</h2><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">Reference project:</a> Working code for the entire blog series (clone, configure, and run).</p></li><li><p><a href="https://github.com/elastic/elastic-otel-python">Elastic Distribution of OpenTelemetry Python:</a> EDOT Python.</p></li><li><p><a href="https://www.elastic.co/docs/solutions/observability/apm/opentelemetry">OpenTelemetry with Elastic:</a> How to send OTel data to Elastic APM.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL documentation:</a> Query language reference.</p></li><li><p><a href="https://www.ubisearch.dev/">UBI standard:</a> Reference schema for search event structure.</p></li></ul><p><em>This is the third post in a </em><a href="https://www.elastic.co/search-labs/blog/series/search-analytics-opentelemetry"><em>series on search analytics with OpenTelemetry and Elastic</em></a><em>. Next up: From clicks to conversions: Conversion tracking, funnel analysis, and revenue attribution.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/search-click-tracking-opentelemetry-esql</guid>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Relevance]]></category>
    <dc:creator><![CDATA[Matthew Adams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3f04684f0d65c705/6a6efbf02888390e8607b2c2/image1.png" length="0" type="image/png"/>
    <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How Elasticsearch detects multiple change points in time series with 0.99 recall]]></title>
    <description><![CDATA[ES|QL's CHANGE_POINT command finds structural shifts, variance changes and spikes in any metric in ~1ms, without tuning anything per series.]]></description>
    <content:encoded><![CDATA[<p>The current generation of agentic models are remarkably good system troubleshooters. Given a hypothesis and the means to test it, they can reason about a failing service much the way a seasoned SRE does: form a theory, look for corroborating evidence, discard it when the data disagrees, and narrow in on a root cause. Their main limitation is not their ability to reason but their reach: they can only investigate what their tools let them see.</p><p>This is where Elasticsearch earns its place in the stack. It is already where a great deal of operational telemetry lives (logs, metrics, traces, events), and it exposes them through an expressive query and aggregation layer. That makes it a natural tool for an agent debugging a live system: it can slice by attribute, aggregate over time, correlate across signals, and drill from a symptom down to the documents that produced it.</p><p>We've been building an agentic layer on top of Elasticsearch that continuously monitors a system and root-causes issues as they arise. A recurring primitive in that workflow is time series event analysis. For example, given an error rate, a p99 latency, a queue depth and a throughput counter, tell me whether something happened, what it was and when. A transient spike in errors, a regime change in latency, and a step up in CPU usage are the signatures of the underlying fault, and they're typically what an agent examines first as it forms and tests hypotheses.</p><p>Elasticsearch has shipped a single-change-point aggregation for some time. It answers "did this series change?" with one verdict and the most significant change it found. That's a good fit for a dashboard, but less so for an agent, which often wants to interrogate a long window and enumerate everything of interest in it: the error spike at 02:14, the latency regime shift at 02:30, the throughput dip while the pod was being rescheduled. So we upgraded the capability to detect and report multiple events of multiple kinds in a single series. At the same time, we took the opportunity to further harden it to work reliably against whatever telemetry the agent points it at. This post describes how it works.</p><h2>Why single change point detection isn't enough for agents</h2><p>Concretely, we want a single entry point that takes a numeric time series and returns a small list of interesting events, each with a type, a location, a significance, and some key characteristics. We care about three classes of event, because they map onto three different kinds of underlying fault:</p><p>Event type</p><p>What changes</p><p>Detection channel</p><p>Example fault</p><p>Structural change</p><p>Level (step) or slope (trend) shifts to a new sustained regime</p><p>Value channel</p><p>Config push doubles baseline latency; memory leak turns a flat curve into a ramp</p><p>Distribution change</p><p>Noise level (variance) shifts while the mean holds steady</p><p>Dispersion channel</p><p>Service responds erratically at the same average latency</p><p>Point anomaly</p><p>Isolated spike or dip against a stable background</p><p>Value channel (pulse detector)</p><p>Single burst of errors; one-minute throughput drop during GC pause</p><p>The hard part is not detecting any one of these on clean, well-behaved data. The hard part is doing it on arbitrary telemetry without per-series tuning. The agent does not know in advance whether the series it is examining is near-constant, smoothly drifting, <a href="https://en.wikipedia.org/wiki/Homoscedasticity_and_heteroscedasticity">heteroscedastic</a> (quiet in places and noisy in others), sparsely populated, or has a magnitude of . It is not scalable to hand-pick parameters for every series it needs to analyze. Whatever we build has to be robust to all of that while maintaining excellent recall and precision. If it fails to detect important events it runs the risk of missing key corroborating evidence for a working hypothesis. Conversely, an analysis tool that reports an event for every minor fluctuation will pollute the context the agent reasons over.</p><p>The design goals, in priority order, are: correct on diverse data out of the box, parsimonious (report only what matters), and cheap enough to run interactively across many series.</p><h2>How PELT and BIC power change point detection</h2><p>Change-point detection is a well studied field. The classical offline formulation searches for the segmentation of a series that minimizes a penalized cost: a per-segment goodness-of-fit term plus a penalty for each added break to stop the optimizer from putting a boundary between every pair of points. Solved naively, this is combinatorial, but PELT (<a href="https://arxiv.org/pdf/1101.1438">Pruned Exact Linear Time, Killick et al.</a>) finds the optimal partition in roughly linear time by using a dynamic program to prune candidate boundaries that can never be optimal. On the labeling side, comparing nested models by an information criterion such as the Bayesian Information Criterion (BIC) gives a principled, scale-aware way to decide whether a candidate break is really important and what sort of change it constitutes.</p><p>These are good building blocks, and we use them. However, the textbook recipe assumes more than telemetry gives you. It typically assumes a single change type (a mean shift), a known and stationary noise level, and reasonably benign numerics. Real telemetry violates all three: variance changes matter as much as mean changes, the noise level is unknown, often heavy-tailed and changing, and the data spans extreme magnitudes and degenerate cases, such as perfectly constant segments, that wreck an ill-conditioned polynomial fit or a fixed-variance cost. Most of the engineering I describe below is about closing that gap.</p><h2>Splitting one time series into three detection channels</h2><p>Rather than trying to find one detector that does everything, we run three focused detectors and then merge their findings. Two of the three are the same structural detector applied to two different views of the data, or channels.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blted8fd73801bac283/6a6a33ed15fc5c4ade9e493d/02fa7892061a4343a7720c7269c8df82ef67603f-1508x1132.png" alt="Elasticsearch time series split into value and dispersion channels detecting step changes, spikes and variance shifts" /><p>Two detectors run on the value channel: a structural detector that flags step and trend changes, and a pulse detector that identifies point spikes and dips. The dispersion channel, a windowed measure of spread, is fed to a second copy of the structural detector, where a variance change shows up as a level shift and is relabeled a distribution change. The two channels are complementary by construction: a step is a level shift the value channel flags but only has a single large first difference the dispersion channel ignores, while a variance change is invisible to the value channel yet shows up clearly in the dispersion channel. A thin orchestration layer then merges and de-duplicates the streams.</p><p>Keeping the three concerns separate makes each one tractable. A mean-shift detector and a variance-shift detector pull in opposite directions if you try to fuse them; a point-anomaly detector and a regime detector need opposite robustness settings. Separated, each can be tuned to its job.</p><h3>PELT with a scale-free cost</h3><p>For the structural channel, we fit each candidate segment with a low-order polynomial (constant or linear) and score it with the profiled-variance Gaussian cost. If a segment  of length  has residual sum of squares , the cost is</p><p>This is the negative log-likelihood of a Gaussian segment after profiling out the variance, i.e., after substituting the maximum-likelihood estimate  back into the log-likelihood. The total objective PELT minimizes is the sum of segment costs plus a per-break penalty,</p><p>Here,  is the BIC complexity term (number of parameters times  where  is the number of values in the time series) and the scale factor lets us trade sensitivity against parsimony in one place.</p><p>Using the profiled cost rather than a cost against a fixed global variance is a deliberate and important choice. The global noise level of telemetry is unreliable: on a smoothly varying series the natural estimate (the spread of first differences) can collapse toward zero, and a fixed-variance cost then treats every wiggle as enormously significant and over-segments. The profiled cost depends only on the ratio , so it is invariant to the absolute scale and immune to that failure. A small floor on  keeps the logarithm finite, so a zero-residual segment is not rewarded without bound.</p><h3>Keeping the fit stable using robust local weighting</h3><p>Before PELT runs, we robustly down-weight points so that an excursion does not create spurious breaks or drag a segment boundary onto itself. Crucially, these weights enter into the weighted residual moments, so they shape not just the segment fit but its residual variance, and so the segment costs themselves. Each point gets a Cauchy weight, , of its residual  from a rolling-median baseline against a robust scale : points near the local median keep full weight and points far from it are progressively discounted. A nice side effect of measuring excursions from the median on a window centered on each point is that clean structural breaks do not get down-weighted at all, because the majority of values land on the same side of the break as the point whose residual is being computed. So a sustained regime keeps full weight, but a lone spike does not.</p><p>There is a nice justification for scoring a weighted Gaussian cost when what we really want is robustness to a heavy tail. The Cauchy weight is exactly the <a href="https://en.wikipedia.org/wiki/Iteratively_reweighted_least_squares">iteratively reweighted least squares</a> (IRLS) weight of its loss: , so the weighted normal equations  are identical to the Cauchy M-estimator's estimating equations . A weighted-mean (or weighted-line) fit at those weights is therefore a <a href="https://en.wikipedia.org/wiki/M-estimator">Cauchy M-estimate</a>, not a Gaussian one. The cost we actually evaluate inherits the same properties. Because  is concave in , its tangent at the current residual lies above it. This gives a pointwise bound  with  the weight at the tangent point; summing, the weighted residual sum of squares  is a tangent upper bound on the total Cauchy loss, touching it in both value and gradient at the weights' anchor point. Minimizing the weighted RSS is thus one step of a <a href="https://en.wikipedia.org/wiki/MM_algorithm">majorize–minimize scheme</a> that provably decreases the true Cauchy objective, and the profiled-variance cost we feed the BIC is that majorizer standing in for the Cauchy deviance. The only approximation is that we anchor the weights once, at the rolling-median baseline, rather than iterating IRLS to its fixed point; this is exact for the inliers that sit near the baseline, and correct for gross outliers, whose vanishing weight removes them from the cost wherever the bound is loosest.</p><p>Finally, the trick that makes this work on heteroscedastic data is that the residual is judged primarily against a <em>local</em> robust scale, not a <em>global</em> one: the MAD of residuals in the same sliding window. On a series that is quiet in one stretch and noisy in another, this means a spike in the quiet stretch that is multiple local sigmas is correctly suppressed. The local MAD can collapse on quiet stretches, so we use a backstop that is a fraction of a global composite of robust scales and a floor related to the quantization error for discrete series and numerical precision otherwise.</p><h3>From candidates to labeled events using BIC verification</h3><p>PELT gives a globally optimal penalized segmentation, so we take its boundaries as candidates and verify each one. For a candidate at index  we look at the window of length  spanning to its nearest neighboring candidates and compare a no-change null against step and trend alternatives by BIC,</p><p>where  counts the fitted parameters (the same parameter count as in the PELT penalty above). We map the BIC gain of an alternative over the null to a significance via , to turn a threshold into a decision boundary). We treat this  as a significance score for ranking and thresholding, not as a calibrated tail probability.</p><p>At this stage we allow higher-degree models to avoid splitting smoothly varying trends. These are problematic in PELT itself because it considers short segments, which they overfit. We keep the most parsimonious alternative that clears the significance threshold and survives a persistence check. The persistence check re-scores with the weights immediately around the candidate muted, and if the evidence collapses, the "change" was driven by a few extreme points – an excursion, not a regime change – and we reject it. The polynomial order is applied symmetrically to the null and to each side of the split, so the alternative is always the same model class merely split at the candidate, and therefore strictly more flexible. The whole process can be thought of as Bayesian model selection with a preference for the null.</p><p>When no candidate survives, we still say something useful: we report the series as "stationary" (best no-change model is a constant) or "non-stationary" (best model has a slope), with the trend direction. For an agent, "this series is cleanly trending up over the window" is itself a finding.</p><h3>Detecting distribution changes with a dispersion channel</h3><p>A variance change is indirectly visible to the mean channel – worse, the robust weighting there actively mutes the excursions that signal it. So we detect it on a separate dispersion channel and reuse the exact same structural detector, because on this channel a variance change is just an ordinary level change.</p><p>The channel is built from one sample per non-overlapping window. Within a window, we take the <a href="https://en.wikipedia.org/wiki/Interquartile_range">inter-quartile range</a> of the first differences, rescaled to a standard-deviation equivalent (), and pass it through :</p><p>Then . Three choices matter here. First-differencing cancels level and slope, so a mean step contributes a single large difference rather than inflating the whole window, and a steady ramp produces a flat channel. Non-overlapping windows keep the samples independent; overlapping windows <a href="https://en.wikipedia.org/wiki/Autocorrelation">autocorrelates</a> the channel and makes the segmenter over-detect. And the IQR is used rather than the median (which is too robust and will miss a window that is 40% noisy then flatlines) or the raw standard deviation (which is not robust enough since one spike's two large differences inflate the window). Because the dispersion channel is a fraction of the original length, the verifier there is restricted to a lower-order null so a genuine low-high-low variance bump is not absorbed.</p><p>The functional form  is worth dwelling on, because each part earns its place. The log makes the channel respond to ratios of noise level rather than absolute differences. Variance changes in telemetry are typically multiplicative: a regime is "twice as noisy". On a raw-scale channel, a doubling shows up as an enormous absolute jump at a high baseline and a negligible one at a low baseline, so an additive step-cost detector would find variance changes trivially in loud series and miss them in quiet ones. Under a log, a factor- change in scale is the same offset  wherever it occurs, which is exactly the additive-step behavior the structural detector is built for. The " is a soft floor. A bare  diverges to  as the scale goes to zero, which is precisely what happens on a near constant stretch, and would manufacture a huge spurious step at the first noisy window after it. Conversely,  is finite and smooth at zero, behaves linearly () while the noise is small, and recovers the multiplicative  behavior once the noise is appreciable. This gives graceful degradation instead of a singularity, and with no tuned epsilon to pick. Note that squaring the scale (using a variance instead) only doubles the dynamic range; it makes no difference to the detector either way, since  differs only by a constant the threshold absorbs.</p><h3>Detecting point anomalies as excursions from a local baseline</h3><p>Spikes and dips are detected as point excursions from the local rolling-median baseline. Working from the local residual rather than raw values means level structure is removed and smooth curvature is tracked; even for time series that change significantly the detector is sensitive to significant local deviations.</p><p>The pipeline is a generous proposer followed by a strict gate:</p><ol><li><p>Propose every point whose residual exceeds a threshold number of robust sigmas. The scale is the larger of the global first-difference noise (which stays meaningful on smooth data where most residuals are exactly zero) and a composite of robust scales of the residuals (which inflates once a frequent large-residual population appears).</p></li><li><p>Merge adjacent same-sign candidates into excursions, dropping any that span a full minimum segment: that is a regime, and is owned by the structural channels.</p></li><li><p>Rank and cap the excursions by peak <a href="https://en.wikipedia.org/wiki/Standard_score">z-score</a>, keeping the top , so a pathological series cannot drown the output.</p></li><li><p>Gate using one shared null: build a <a href="https://en.wikipedia.org/wiki/Kernel_density_estimation">Gaussian KDE</a> from the series with all of the retained excursions removed, and keep an excursion only if its peak's Bonferroni-corrected upper-/lower-tail probability under that null clears the threshold.</p></li></ol><p>Removing all the tested excursions from the single null at once is a key trick. The leave-one-out alternative – score each excursion against a null containing the others – lets the largest spike and dip mask everything else. Removing them together means several genuinely distinct excursions are each judged against the remainder and all survive, while a recurring population is still rejected.</p><p>The proposer and the gate ask deliberately different questions, and that distinction drives two further choices. The proposer works on residuals from the rolling median since it wants recall, and a residual is what tells you a point stands out from its local neighborhood. The gate is value-based: it asks "is this magnitude one we see at other times in the series?", so a spike to a level that recurs elsewhere — such as periodic batch jobs — is suppressed even though it is a large local residual. Those are the right semantics for an agent, but they expose a heteroscedasticity problem, because telemetry noise is almost always a function of magnitude. Periodic spikes can sit orders of magnitude above the background, and a single KDE bandwidth fitted to the whole value range is then far too narrow up in the high tail. So ordinary large values come back as significant, a steady source of false positives.</p><p>The fix is a <a href="https://en.wikipedia.org/wiki/Variance-stabilizing_transformation">variance-stabilizing transform</a>. We run the value gate in  space, where  is a robust measure of the spread of the background.  is linear for  and logarithmic for , which turns a multiplicative (magnitude-dependent) spread into a roughly constant one, so a single bandwidth is valid across orders of magnitude. It is also odd and finite at zero, so exact zeros and sign changes (dips below a small baseline) need no special handling, unlike a bare log. Crucially, it is monotone and so does not change what is tested (for any monotone function , ) so it only fixes the estimate of that tail probability.</p><p>One subtlety closes the loop. The KDE null and the kernel bandwidth are taken from different scales, on purpose. The null is the stabilized background values: so it models any mode in the data, which is what makes a recurring large magnitude unsurprising. However, the bandwidth is taken from the stabilized residuals, not the stabilized values, because a genuine level change makes the value distribution bimodal, and a bandwidth computed from that bimodal spread would balloon, masking a real spike sitting on top of a shifted regime. The residual removes the step, so the bandwidth always reflects within-regime noise and the gate stays sensitive to a deviation that is extreme relative to its own neighborhood if it is also outside the envelope for the series as a whole.</p><h3>Merging structural, distribution, and point anomaly events</h3><p>Finally, an orchestration layer merges the structural, distribution, and point-anomaly event streams. Structural and distribution events that mark the same regime boundary are de-duplicated to the more significant one (a boundary that shifts both level and spread is one event, not two). Pulses are a separate stream added after de-duplication, because a spike that lands on a structural boundary is a real, separate finding and must not be suppressed. Everything is then mapped back from the internal value-array index space to source-bucket indices.</p><h2>Handling extreme magnitudes and edge cases</h2><p>What makes this usable as an unattended tool is a collection of defensive choices for the cases that break naive implementations:</p><ul><li><p>Variance computed as  loses all precision at large magnitudes: a constant series at  can manufacture phantom change points purely from floating-point error. We center every PELT input by a constant offset first; the polynomial RSS is invariant to that shift in exact arithmetic, but the working magnitudes drop from  to .</p></li><li><p>Using raw indices as the regressor results in poor condition polynomial fits: the largest moment is , which is about  for a cubic over a 2000-point window. This trips the SVD singularity guard and silently degrades the fit. Mapping  affinely onto  leaves the fit identical (RSS is invariant under reparametrization) but every moment becomes .</p></li><li><p>Scale-free cost, as described, means the segmentation cost doesn't depend on tuning a noise estimate.</p></li><li><p>Down-weighting wants a primarily <em>local</em> scale: suppress whatever is anomalous in its own neighbourhood. It uses the maximum of MAD and a small global floor on the differences from the rolling median. Conversely, spike/dip detection wants a <em>global</em> scale: we care about global outliers. It uses a composite of robust scales of all differences. Using the wrong one in either place produces characteristic failures – irrelevant spikes in a quiet segment, or locally large excursions creating spurious breaks – and maintaining separate channels allows us to pick appropriately.</p></li><li><p>Using one p-value threshold, <a href="https://en.wikipedia.org/wiki/Bonferroni_correction">Bonferroni-corrected</a> by the number of candidates, applied consistently across all three detectors, means that "how surprised should I be" is consistent everywhere.</p></li></ul><p>The recurring theme is that the difference between a detector that works in a notebook and one that works on a firehose of production time series is mainly in handling the edge cases gracefully.</p><h2>Performance: ~1ms per series on a single core</h2><p>Detection cost is dominated by PELT. Its segment cost is a profiled-variance linear fit, which we evaluate in constant time from prefix-summed weighted moments rather than maintaining a regression per candidate boundary, so a single segment cost is a handful of array reads and a 2×2 solve. Cost grows a little faster than linearly with series length since PELT's pruned candidate set does not stay constant on noisy data. Therefore, to bound the worst case on very long series, we downsample ahead of detection: above a cap (2000 samples), the series is collapsed into macro-buckets, keeping two samples per bucket: the median and the largest local deviation. This is inspired by the <a href="https://www.vldb.org/pvldb/vol7/p797-jugel.pdf">M4 downsampling scheme</a>, but because we need only the median and the largest excursion for structural-change and outlier detection, respectively, we can then afford to double the bucket resolution. The downsampled series carries its original bucket indices, so every reported event still maps back to a real source bucket; below the cap it is a no-op. The whole analysis is a single pass over the (possibly downsampled) series with no per-series configuration. This lets the agent call it freely across many signals.</p><p>In absolute terms, this means the analysis is comfortably interactive. On a single core, post-warmup, a typical series of 140–350 buckets is analysed in about 1 ms (≈220,000 buckets/s), and a series long enough to hit the downsample cap (say 5,000 buckets, collapsed to 2,000) takes about 40 ms, which is the effective worst case per call. For an agent issuing a handful of these calls per investigation, and parallelising across the many series in a `BY` query, the latency is negligible.</p><h2>Evaluation on synthetic and production telemetry</h2><h3>Synthetic benchmark</h3><p>We evaluate first on a synthetic generator with known ground truth, because it lets us measure the things that matter precisely. The generator produces random time series with diverse behaviors, which we group into three families. The <strong>positive</strong> family injects known events of each type: clean and noisy step changes (up and down, SNR around 10, at several positions), trend onsets and ramps (including a flat–ramp–flat sequence with two boundaries), variance changes with a constant mean (single steps and a low–high–low bump), and isolated spikes and dips. The <strong>null</strong> family ideally produces no event: stationary noise, perfectly constant series, smooth quadratic drift and clean ramps, and periodic signals. (A variance change with a constant mean is not null: it is a distribution change, an abrupt step on the dispersion channel, and so it belongs in the positive family above. The related null requirement, that such a change does not surface on the value channel as a step, is checked separately.) The <strong>adversarial</strong> family stresses the robustness machinery: a perfectly flat series at a magnitude of  (for which a naive variance arithmetic manufactures phantom breaks here through catastrophic cancellation) and, conversely, a genuine spike or step in a noisy baseline as high as  (which must still be found and located, the high baseline notwithstanding); a spike on top of a step; a within-regime spike after a 100 level jump; a recurring train of equal peaks (a population, not individual spikes); wide sustained excursions (a structural change, not a spike or dip); and fuzzed random level-shift series. On these series we track recall per event type, precision and the false-positive rate on the null family, localization error, parsimony (events per series and adherence to the count limit), and invariance under constant offsets, rescaling and extreme magnitudes.</p><p>The table below summarizes what the suite tests.</p><p>Event family</p><p>Representative scenarios</p><p>Required outcome</p><p>Localization tolerance</p><p>Step</p><p>clean / noisy, up / down, single and multiple, several positions</p><p>detected</p><p>≤ 4–8 buckets</p><p>Trend</p><p>slope change and flat–ramp–flat, clean and in noise</p><p>detected</p><p>≤ 8–12 buckets</p><p>Distribution</p><p>variance step (mean constant), low–high–low bump</p><p>detected</p><p>≤ 1 dispersion window</p><p>Spike / dip</p><p>isolated, multiple distinct, on heavy-tailed and high-magnitude series, within-regime after a step</p><p>detected, capped at max(5, 2% of n)</p><p>≤ 2 buckets</p><p>Null series</p><p>stationary noise, constant, smooth drift / ramp, periodic</p><p>no event reported</p><p>n/a</p><p>Invariance / robustness</p><p>constant offset, rescaling, 10^5–10^9 magnitudes, spike-on-step, recurring population, wide excursion</p><p>result unchanged / no spurious event</p><p>n/a</p><p>Running this over 400 series per scenario, just over half a million buckets in total, and scoring with the same two categories we use for the real-data evaluation below (any regime change versus point spikes and dips) gives:</p><p>Event type</p><p>Recall</p><p>Precision</p><p>Median localization error</p><p>Structural change</p><p>0.994</p><p>0.664</p><p>0</p><p>Spike / dip</p><p>0.847</p><p>0.746</p><p>0</p><p>Two things stand out. When an event is detected, it is placed essentially exactly: the median localization error is zero buckets for both categories; and the point-wise accuracy, 0.998, is directly comparable to the 0.995 we report on real telemetry below: the overwhelming majority of buckets are correctly left unmarked.</p><p>The precision figures are lower than on real data, and understandably so. A sixth of this population is a <em>hostile</em> null family (periodic signals, smooth drift, clean ramps) chosen precisely because they tempt a detector into a spurious break, and every false alarm on them counts against precision. The resulting false-positive rate is 0.5% per bucket, with 15% of null series carrying at least one spurious event. We treat this as the conservative end of the range: on the real-telemetry mix below, where the null series are less adversarial, precision rises to 0.85 (structural) and 0.90 (spikes/dips). Finally, offset and scale invariance holds on all 400 series: the same events, to within a few buckets, whether the series is shifted by a constant or rescaled by up to three orders of magnitude.</p><p>That raw precision also misses how the result is consumed. The agent reads events most-significant-first, so a false positive only does harm if it outranks a genuine one, and by and large it does not. The median p-value of a true positive is about , against about  for a false positive: the genuine events are typically overwhelmingly more significant. Concretely, if we keep only the top  events per series by significance (with  the true count) precision rises to 0.94, and a randomly chosen true positive is more significant than a randomly chosen false positive 93% of the time. It is not a perfectly clean separation: a strong periodicity or a sharp curve genuinely can contain a significant-looking break, which is why that figure is 0.94 rather than 1. However, the ranking is reliable enough that an agent reading from the top, or applying a stricter significance cut-off, sees the real events first and the false alarms as a lower-significance tail. This is also why exposing the detector's significance to the agent (see <a href="https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql#whats-next-for-es|ql-time-series-analysis">What's next</a>) matters more than squeezing the raw precision higher.</p><h3>Real cloud telemetry</h3><p>To evaluate on real data, we scraped around 300 metrics from our production cloud environment. These cover HTTP status-code counts, failed memory allocations, memory usage, network usage, page faults, CPU usage, and throttling metrics, measured both per instance and aggregated across the fleet as a whole. Their values range over more than 12 orders of magnitude, and they display a variety of behaviors including ramps, periodicity, step changes, distribution changes, and trend changes. The figure below shows a sample of series together with the detections we make on them.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte0bc6fbc138cca1c/6a6a33ee065b162280702003/2c30fd3e2f97e32b5b47bf66a923f73b08ec015d-2510x1284.png" alt="Grid of 12 production cloud telemetry series with detected structural changes, distribution changes and anomalies marked by Elasticsearch's change point detector" /><p>To get a sense of the accuracy on this set data, we labeled a subset of 150 most interesting time series, marking the visually clearest features in each. This labeling is not necessarily optimized for our target use case, where false negatives are typically more problematic than false positives: an agent consumes these results as part of a broader investigation and can pull additional information to corroborate them. Even so, we find excellent agreement with human judgment on these series. Since each series comprises between 140 and 350 points, the point-wise accuracy, at 0.995, is extremely high: the great majority of points are correctly identified as neither a change, a spike, nor a dip. But the more telling metrics are recall and precision on the human-labeled points, shown in the table below. The human labels did not attempt to categorize each change, so we break the results down only into "structural changes" and "spikes / dips" — the same two categories, and the same point-wise accuracy and recall/precision metrics, as the synthetic benchmark above.</p><p>Event type</p><p>Recall</p><p>Precision</p><p>Structural change</p><p>0.94</p><p>0.89</p><p>Spike / dip</p><p>0.97</p><p>0.92</p><p>It is worth covering exactly why we get disagreements. These largely fall into three categories: spikes and dips in context, isolated breaks, and small-magnitude breaks in stable series. We deliberately do not try to detect spikes and dips that are unusual only in their immediate context; that is, not globally unusual but visually significant relative to an inferred periodicity in the data, for example. Trying to account for these without fully modeling the seasonality in the data hurt precision more than it helped recall, and we have a separate persistent anomaly-detection process that builds more complete models of baseline behavior over time. Isolated breaks are an artifact we decided to live with: PELT's cost function tends to isolate a change point with a few values intermediate between the two regimes, because absorbing it into either neighboring span inflates that span's cost. Humans are good at judging such situations visually and assign a single change point. Finally, small-magnitude changes are simply not visually obvious. We detect them deliberately and regard this as a strength of a quantitative approach, since they are often the early precursors of an incident whose later, larger effects drown them out.</p><h2>How the agent uses change point results in ES|QL</h2><p>To the agent, all of this is one tool call: it points it at a collection of time series and gets back a typed, located, ranked list of events. That list is small by construction, which means it drops cleanly into the model's context without crowding out everything else it is reasoning about. Because the result contains multiple events, a single call over a window can hand the agent the whole local story – "error spike at 02:14, latency regime change at 02:30, throughput dip at 02:31" – and let it correlate across signals to a root cause.</p><p>Operationally, we expose this through both the ES|QL <code>CHANGE_POINT</code> <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/change-point">command</a> and the <code>change_point</code> <a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-change-point-aggregation">aggregation</a>. Currently, we have not extended the <code>change_point</code> aggregation to return multiple change points since it breaks backwards compatibility of the output schema. It just returns the most significant event. We don't have the same restriction for ES|QL since it returns change points annotated onto the table rows to which they apply. We do plan to revisit the output schema for both ES|QL and the aggregation in a later version. We'd like to migrate to optionally returning significance in log-space, which doesn't underflow, and including a short verbal description of each change, which we expect to help agents when seeing just the change points themselves.</p><p>ES|QL is Elasticsearch's piped query language, and <code>CHANGE_POINT</code> runs the detector as one stage in a pipeline. Its <code>BY</code> clause enables it to analyze many series at once (one per group) so the agent can, in a single query, segment every service's latency or every host's error rate side by side rather than issuing a call per series. The actual leverage, compared to the <code>change_point</code> aggregation, is composability: the events come back as ordinary rows in the pipeline, so the agent then has the entire ES|QL language to manipulate them downstream. It can filter to a window, join change points against deploy markers, count events per service, rank by significance, feed the survivors into a further aggregation, and so on.</p><p>For example, suppose the agent wants to find out whether any servers have recently seen a sudden CPU spike or a prolonged step change in CPU usage over the last 12 hours, and whether that might point to a load-balancing issue. It could use the following query:</p><p>Here it's using a <code>STATS ... BY host.pod</code> to see how the detected events cluster across other dimensions of the data, such as the Kubernetes pod, and so judge whether they share a common cause.</p><h2>What's next for ES|QL time series analysis</h2><p>As far as detecting events of interest in time series, the foundation is in place: a single, robust, parsimonious tool that turns a raw telemetry series into the short list of events that actually matter, which is exactly the kind of reach an agentic SRE needs. Going forward, we plan to explore the best mechanism for feeding the detector's uncertainty to the agent, so that a borderline event can be flagged as "worth a second look" rather than silently included or dropped. Also, this is the first of several analytical tools we plan to build into the ES|QL query language to enable agents to triage and RCA issues more effectively; so stay tuned for further updates.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[ML Research]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Thomas Veasey]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt097721c03648a84e/6a6a33ef0a222b4c70877f32/8f6b95800c65fe389d3e8d8281e8e8dc351f734d-992x342.png" length="0" type="image/png"/>
    <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How to instrument your search API with OpenTelemetry and query it with ES|QL]]></title>
    <description><![CDATA[Add custom attributes to OpenTelemetry spans and run six ES|QL queries that reveal your top searches, zero-result rate and slowest queries.]]></description>
    <content:encoded><![CDATA[<p>Instrument your search API with about 20 lines of OpenTelemetry (OTel) code, and Elasticsearch Query Language (ES|QL) can tell you what people are searching for, how often they get nothing back, and how fast search actually runs. This builds directly on <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry">the first post in this series</a>, where we made the case for using OpenTelemetry over a bespoke analytics pipeline. Here, we wire up a FastAPI search endpoint with custom <code>search.*</code> attributes and run six ES|QL queries against the resulting trace data. On our demo cluster, 17.7% of searches came back empty, a gap we found within minutes of turning the instrumentation on. No separate logging pipeline is required. It's the same spans, attributes, and query language you're probably already running somewhere else in Elastic.</p><h3>What you'll discover</h3><p>In this post, you'll learn how to:</p><ul><li><p>Set up OpenTelemetry in an example Python FastAPI back end.</p></li><li><p>Add custom <code>search.*</code> attributes to your search spans in ~20 lines of code.</p></li><li><p>Understand how OTel-native ingestion maps attributes to queryable data.</p></li><li><p>Write six ES|QL queries against real trace data: top queries, zero-results rate, which queries return nothing, average and max latency, slow query investigation, and search volume over time.</p></li><li><p>Turn those queries into saved Kibana visualizations.</p></li></ul><h3>What you'll need</h3><ul><li><p>An Elastic Cloud deployment (or self-managed with OTel-native ingestion enabled). Examples tested on Elastic Stack 9.x.</p></li></ul><h2>From concept to code: Building the search analytics instrumentation</h2><p>In the <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry">first post</a>, we made the case for using OpenTelemetry to capture search analytics. The idea: Add <code>search.*</code> attributes to your existing OTel spans, send them to Elastic APM, and query them with <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL</a>.</p><p>Now let's build it.</p><p><strong>Want working code?</strong> A <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">companion reference project</a> accompanies this series. It's a minimal FastAPI app with the exact instrumentation described below. Clone it, add your Elastic Cloud credentials, and you'll have search analytics data flowing in 10 minutes. Each blog stage maps to a commented-out code block you can enable as you progress.</p><h3>An OpenTelemetry primer for search API developers</h3><p>If you've been building search systems but haven't worked with OpenTelemetry before, here's the minimum you need to know.</p><p>OTel is an open standard for collecting observability data, traces, metrics, and logs from your applications. It's vendor-neutral: You instrument your code once, and send data to any compatible back end.</p><p>The core concept is the <em>span</em>. A span represents a single operation, for example, an API call, a database query, or a search request. Every span has a start time, an end time (the difference is the <em>span duration</em>), and <em>attributes</em>, which are key-value pairs that describe what happened.</p><p>Spans nest inside each other to form <em>traces</em>. A trace is a tree of spans that represents one end-to-end request. When a user searches, the trace might look like: browser request → API handler → search logic → Elasticsearch query. Each step is a span, and the parent-child relationships show you exactly where time was spent. This is <em>distributed tracing</em>, which works across services and network boundaries, so a single trace can follow a request from front end to back end to database and back.</p><p><strong>Why traces instead of logs?</strong> You could log <code>"search query=headphones results=15 took=120ms"</code> and parse it later. But a log line is flat; it can't show you that the 120ms Elasticsearch time sat inside a 250ms API call, revealing 130ms of overhead in your application layer. Traces give you hierarchy, timing, and correlation across services. For search analytics, that means you can see not just <em>what</em> happened but also <em>where</em> time was spent and <em>how</em> operations relate to each other.</p><p>For this post, we don't need to understand the full OTel ecosystem. We just need three things:</p><ol><li><p><strong>Create a span</strong> when a search request happens.</p></li><li><p><strong>Add attributes</strong> to that span, describing the search (such as <code>search.query</code> or <code>result_count</code>).</p></li><li><p><strong>Send the span</strong> to Elastic, where we can query it with ES|QL.</p></li></ol><p>That's it. If you can call <code>span.set_attribute("key", value)</code>, you can build search analytics.</p><h3>What you'll build</h3><p>By the end of this post, every search request in your API will emit an OTel span that looks like this:</p>span.name:                     "search"
search.query:                  "wireless headphones"
search.result_count:           15
search.query_id:               "e2afdb85eb63382e..."
search.took_ms:                165<p>And you'll run six ES|QL queries against real data to answer questions that your team is already asking, using about 20 lines of instrumentation code in total.</p><h2>Install the OTel SDK</h2><p>We're using Python and FastAPI here. The same pattern applies to any language with an OTel SDK; the concepts are identical, only the imports change.</p><p>Elastic provides the <a href="https://github.com/elastic/elastic-otel-python">Elastic Distribution of OpenTelemetry Python (EDOT)</a>, which bundles the standard OTel SDK with sensible defaults, early access to Elastic-contributed improvements, and a single <code>configure_opentelemetry()</code> call that handles all the boilerplate. We recommend it:</p>pip install elastic-opentelemetry \
    opentelemetry-instrumentation-fastapi \
    opentelemetry-instrumentation-elasticsearch<p>Three packages, two roles:</p><ul><li><p><code>elastic-opentelemetry</code> EDOT: The OTel API, SDK, and OpenTelemetry Protocol (OTLP) exporter in one package, preconfigured for Elastic.</p></li><li><p><code>opentelemetry-instrumentation-fastapi</code>: Auto-instruments HTTP endpoints (automatic spans for every request).</p></li><li><p><code>opentelemetry-instrumentation-elasticsearch</code>: Auto-instruments Elasticsearch client calls (automatic spans for every query).</p></li></ul><p>The auto-instrumentation packages are doing real work here. Without writing a single line of tracing code, you already get HTTP request spans and Elasticsearch query spans. What we're adding is the search-specific context that turns generic traces into analytics.</p><p><strong>Using the standard OTel SDK instead? </strong>Replace <code>elastic-opentelemetry</code> with <code>opentelemetry-api</code>, <code>opentelemetry-sdk</code>, and <code>opentelemetry-exporter-otlp-proto-http</code>. You'll need to wire up the <code>TracerProvider</code>, <code>OTLPSpanExporter</code>, and <code>BatchSpanProcessor</code> manually (about 10 extra lines). Everything else in this post works the same either way. See the <a href="https://www.elastic.co/docs/solutions/observability/apm/opentelemetry">Elastic OTel guide</a> for the full setup.</p><h2>Configure the connection</h2><p>OTel uses environment variables for connection configuration. You need four to get started:</p><p>Variable</p><p>Purpose</p><p>Example</p><p>`OTEL_EXPORTER_OTLP_ENDPOINT`</p><p>Managed OTLP (mOTLP) endpoint URL</p><p>`https://my-deployment.ingest.us-central1.gcp.elastic-cloud.com`</p><p>`OTEL_EXPORTER_OTLP_HEADERS`</p><p>Authentication</p><p>`Authorization=ApiKey &lt;your-api-key&gt;`</p><p>`OTEL_SERVICE_NAME`</p><p>Service name (shown in Kibana APM)</p><p>`search-analytics-demo`</p><p>`OTEL_RESOURCE_ATTRIBUTES`</p><p>Other resource attributes</p><p>`service.version=1.0.0`</p><p>Where to find these values: In Elastic Cloud, your mOTLP endpoint follows the pattern <code>https://&lt;deployment&gt;.ingest.&lt;region&gt;.gcp.elastic-cloud.com</code>. You can find it in the Elastic Cloud console under your deployment's details or in Kibana at the APM integration page (<code>/app/home#/tutorial/apm</code>) under the <strong>OpenTelemetry</strong> tab. Your API key can be created from Kibana's Stack Management &gt; API Keys or via the Elasticsearch Create API Key API. For self-managed deployments, you can use the <a href="https://www.elastic.co/docs/reference/edot-collector">EDOT Collector</a> as an intermediary that receives OTLP and forwards to Elasticsearch.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb920fc0478908c55/6a6a3402c40efb7565d30950/86c9ed08f305c168429eff0c0ad059c4d8262197-1440x708.png" alt="Kibana APM integration page showing OpenTelemetry configuration settings for search API instrumentation" /><p>Set them in your environment or <code>.env</code> file:</p>export OTEL_EXPORTER_OTLP_ENDPOINT="https://my-deployment.ingest.us-central1.gcp.elastic-cloud.com"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=ApiKey &lt;your-api-key&gt;"
export OTEL_SERVICE_NAME="search-analytics-demo"
export OTEL_RESOURCE_ATTRIBUTES="service.version=1.0.0"<h2>Initialize the tracer</h2><p>With EDOT and environment variables configured, initialization wires up three things: the tracer provider and two auto-instrumentation packages that automatically create spans for every HTTP request and every Elasticsearch query:</p>from opentelemetry import trace
from elastic_opentelemetry import configure_opentelemetry
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.instrumentation.elasticsearch import ElasticsearchInstrumentor

def init_otel(app):
    configure_opentelemetry()
    FastAPIInstrumentor.instrument_app(app)
    ElasticsearchInstrumentor().instrument()

tracer = trace.get_tracer("search-api")<p><code>configure_opentelemetry()</code> reads the <code>OTEL_*</code> environment variables and sets up the tracer provider, exporter, and batch processor with Elastic-optimized defaults, including using the HTTP exporter automatically, which is what the mOTLP endpoint requires. If you see connection errors with vanilla OTel, the most common cause is accidentally using the gRPC exporter instead of HTTP.</p><p>The two instrumentors patch the FastAPI and <code>elasticsearch-py</code> libraries at startup so every request and every Elasticsearch call automatically generates a span, without any code changes to individual endpoints.</p><h2>Instrument your search API</h2><p>Here's where it gets interesting. This is the code that turns a generic API endpoint into a search analytics source:</p>from opentelemetry import trace

tracer = trace.get_tracer("search-api")

@app.post("/api/search")
def search(request: SearchRequest):
    with tracer.start_as_current_span("search") as span:
        # Set attributes BEFORE the query
        # (available even if the query fails)
        query_id = format(span.get_span_context().trace_id, "032x")
        span.set_attribute("search.query", request.query)
        span.set_attribute("search.query_id", query_id)

        results = es.search(
            index="products",
            body=build_query(request)
        )

        # Set attributes AFTER the query
        total_hits = results["hits"]["total"]["value"]
        span.set_attribute("search.result_count", total_hits)
        span.set_attribute("search.took_ms", results["took"])

        # Include query_id in the response so the frontend can link
        # click and conversion events back to this search
        return {
            **format_response(results),
            "query_id": query_id,
        }<p>A few things worth unpacking.</p><ul><li><p><code>start_as_current_span</code> creates the span and sets it as the active span in the current context. This matters because the Elasticsearch client instrumentation picks up the active span and nests its own spans underneath it. You get a span hierarchy automatically.</p></li><li><p><code>search.query_id</code> is derived from the trace ID. Every trace already has a unique identifier, and we're reusing it as the query identifier. There’s no UUID generation or database sequence. When we add click tracking later, clicks will reference this same <code>query_id</code> to link back to the search that produced the results.</p></li><li><p><code>search.result_count</code> does double duty. It tells you how many results came back, and when it's zero, you know you have a content gap. There’s no need for a separate boolean flag: Just filter on <code>result_count == 0</code> in your queries.</p></li></ul><p>Note that the application name is no longer set as a span attribute; it's the <code>service.name</code> resource attribute, configured once via <code>OTEL_SERVICE_NAME</code>. This is the standard OTel approach: Resource attributes describe the service, and span attributes describe the operation.</p><h3>Normalize search queries before analyzing them</h3><p>Notice that we're storing <code>request.query</code> as is. That means "Laptop Bag", "laptop bag", and " laptop bag " will be counted as three different queries when you aggregate with <code>STATS ... BY attributes.search.query</code>.</p><p>For cleaner analytics, normalize before setting the attribute:</p>span.set_attribute("search.query", request.query.strip().lower())<p>Lowercasing and trimming whitespace is enough for most cases. If you need the original phrasing (for display or debugging), store it in a separate attribute, like <code>search.query.original</code>. But start simple. You can always add the raw version later if you find you need it.</p><h3>What a search API trace looks like</h3><p>Once this is running, a single search request produces this trace:</p>HTTP POST /api/search          (root — auto-instrumented by FastAPI)
└── search                     (our span — search.* attributes live here)
    ├── info                   (ES client — auto-instrumented)
    ├── query_rules.get_ruleset (ES client)
    └── search                 (ES client — the actual Elasticsearch query)<p>The auto-instrumented spans give you HTTP latency and Elasticsearch query detail. Your <code>search</code> span in the middle ties them together with the business context: what the user searched for, how many results came back, how long Elasticsearch took.</p><h3>The search span attributes you need to capture</h3><p>Here's the full set of attributes we're capturing on the search span:</p><p>Attribute</p><p>Type</p><p>When set</p><p>Purpose</p><p>`search.query`</p><p>string</p><p>Before query</p><p>The query as the user entered it</p><p>`search.query_id`</p><p>string</p><p>Before query</p><p>Unique identifier, derived from trace ID</p><p>`search.result_count`</p><p>int</p><p>After query</p><p>Total matching results (0 = zero-result search)</p><p>`search.took_ms`</p><p>int</p><p>After query</p><p>Elasticsearch execution time in milliseconds</p><p>`search.query_response_hit_ids`</p><p>string[]</p><p>After query</p><p>Document IDs returned (optional; enables per-result analytics)</p><p>`feature_flag.key`</p><p>string</p><p>Before query</p><p>A/B test flag name (optional; pair with `feature_flag.result.variant` for the assigned variant; enables per-variant click-through rate (CTR) comparison)</p><p>We use the <code>search.*</code> namespace following OTel's convention of domain-specific prefixes (<code>http.*</code>, <code>db.*</code>, <code>messaging.*</code>). While there aren't standardized search conventions in OTel yet, <code>search.*</code> is self-describing and vendor-neutral. The naming is informed by the <a href="https://www.ubisearch.dev/">User Behavior Insights (UBI)</a> Standard, which defines a schema for search events. We reference it for structure without coupling to it. Where established OTel conventions exist, like <code>feature_flag.key</code> (flag name) and <code>feature_flag.result.variant</code> (assigned variant) for A/B experiments, we reuse them rather than inventing custom attributes.</p><p>You'll notice that <code>enduser.pseudo.id</code> isn't in the search span table above. We don't need it for query analytics, but you'll add it as soon as you introduce click tracking in our third blog; it ties click events back to a specific browser session, enabling per-user CTR and Mean Reciprocal Rank (MRR). Blog 3 adds <code>enduser.pseudo.id</code> (browser-generated, persistent across sessions) to link clicks back to searches. Our fourth blog focussing on revenue attribution documents <code>session.id</code> and <code>user.id</code> as optional extensions for authenticated users who want cross-device attribution.</p><h2>How OpenTelemetry attributes become queryable ES|QL fields</h2><p>Before we start querying, you need to understand how OTel attributes map to Elasticsearch fields. With OTel-native ingestion into Elastic, the mapping is straightforward.</p><p>OTel attribute</p><p>Type</p><p>ES|QL field</p><p>`search.query`</p><p>string</p><p>`attributes.search.query`</p><p>`search.result_count`</p><p>int</p><p>`attributes.search.result_count`</p><p>`search.took_ms`</p><p>int</p><p>`attributes.search.took_ms`</p><p>`search.query_id`</p><p>string</p><p>`attributes.search.query_id`</p><p>`feature_flag.key`</p><p>string</p><p>`attributes.feature_flag.key`</p><p>With OTel-native ingestion, attribute names preserve their dot notation under <code>attributes.*</code>. All types live in the same namespace; there’s no split between string and numeric fields. Booleans are stored as native booleans, not strings. If you've used Elastic APM's classic ingestion before, you'll appreciate the simplicity: What you set in code is what you query.</p><h2>Running ES|QL queries in Kibana Discover</h2><p>With spans flowing to Elastic, open Kibana and go to <strong>Discover</strong> (in the left sidebar under <strong>Analytics</strong>, or use the global search bar and type "Discover"). By default, you'll see the KQL query bar, a filter language familiar from Kibana dashboards. Click <strong>Try ES|QL</strong> in the top right to switch to the ES|QL editor. Unlike KQL (which filters documents) or the JSON query DSL (which requires nested objects), ES|QL is a piped language: Each <code>|</code> step transforms the previous output, making aggregations like <code>STATS count BY field</code> read naturally from left to right.</p><p>The editor gives you a full-width text area where you type piped queries. Results appear as both a table and an auto-generated chart; Kibana picks a sensible visualization based on your query shape. For <code>STATS ... BY</code> queries, you'll get a bar chart automatically.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt839635c0a21e4916/6a6a3403d57c1d4bd1c13ef0/d284fdad7c58576f79209e72b7579b37f4ebbe63-1440x708.png" alt="Kibana Discover ES|QL query results showing search count and average results by search query" /><p>Set the time range wide enough to capture your data (top-right date picker). If you're just getting started, try "Last 30 days".</p><h2>Six ES|QL queries for search analytics</h2><p>Everything below runs against <code>traces-generic.otel-default</code>, the index where Elastic's OTel-native ingestion automatically stores trace data. You don't need to create this index; it's provisioned by Elastic when the first OTLP span arrives.</p><p>These queries ran against our live demo cluster: 62 searches, ~20 distinct queries, latency range 77ms–153ms.</p><p><strong>Note:</strong> Results below are illustrative. When you run the <a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">reference project</a> and generate traffic with <code>python generate_traffic.py --blog 2 --sessions 50</code>, your exact numbers and top queries will vary based on session count and random query selection. The query patterns and ES|QL syntax are what matter here.</p><h3>Query 1: Are spans arriving?</h3><p>Start simple. Count your search spans.</p>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
  AND name == "search"
| STATS total_searches = COUNT(*)<p><strong>Result:</strong> 62.</p><p>If this returns zero, your spans aren't arriving. Check your OTLP endpoint and API key and that <code>init_otel()</code> is being called before any requests. The <code>name == "search"</code> filter ensures that you're counting your custom spans, not the auto-instrumented Elasticsearch client spans (which are also named "search").</p><h3>Query 2: What are users searching for?</h3><p>The first question every search team asks.</p>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
  AND attributes.search.query != ""
  AND name == "search"
| STATS
    search_count = COUNT(*),
    avg_results = ROUND(AVG(attributes.search.result_count), 0)
  BY attributes.search.query
| SORT search_count DESC
| LIMIT 20<p><strong>Results:</strong></p><p>Query</p><p>Searches</p><p>Average results</p><p>laptop</p><p>9</p><p>8</p><p>headphones</p><p>7</p><p>12</p><p>running shoes</p><p>6</p><p>5</p><p>"laptop" was the most popular query, with nine searches. There were around 20 distinct queries total (including zero-result ones).</p><p>The <code>avg_results</code> column tells you whether popular queries are actually returning content. A query with high volume and low results is a relevance problem worth investigating. If your query has high volume and high results, check whether users are actually clicking. We'll get to that in the next blog focussing on measuring search quality with click data.</p><h3>Query 3: What percentage of searches return nothing?</h3>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
  AND name == "search"
| STATS
    total = COUNT(*),
    zero_results = COUNT(CASE(attributes.search.result_count == 0, 1))
| EVAL zero_rate_pct = ROUND(100.0 * zero_results / total, 1)<p><strong>Result:</strong> 17.7% (11 out of 62 searches returned nothing).</p><p>We're using <code>attributes.search.result_count == 0</code>, a straightforward numeric comparison. No separate boolean attribute is needed when you already have the count.</p><p>A zero-results rate above 10% is worth investigating. Every zero-result search is a user who asked for something and got nothing back. Some of those are junk queries, but others reveal real content gaps or query parsing failures.</p><h3>Query 4: Which queries return nothing?</h3><p>The rate tells you there's a problem. This query tells you where.</p>FROM traces-generic.otel-default
| WHERE attributes.search.result_count == 0
  AND name == "search"
| STATS occurrences = COUNT(*) BY attributes.search.query
| SORT occurrences DESC
| LIMIT 20<p><strong>Results:</strong></p><p>Query</p><p>Occurrences</p><p>quantum physics calculator</p><p>4</p><p>unicorn saddle</p><p>3</p><p>holographic projector</p><p>2</p><p>time machine parts</p><p>2</p><p>Three different failure modes: "quantum physics calculator" and "unicorn saddle" are out-of-catalog queries you'll never be able to serve. This is useful to know but nothing to fix. "holographic projector" might be a real emerging category worth considering. "time machine parts" is probably noise. In a production catalog, these would be mixed with legitimate zero-result queries that <em>are</em> fixable, like missing synonyms, phrasing mismatches, or product gaps.</p><p>Repeated zero-result queries are the highest-priority fixes. One-off failures are usually noise.</p><h3>Query 5: How fast is search?</h3>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
  AND name == "search"
| STATS
    avg_ms = ROUND(AVG(attributes.search.took_ms), 0),
    max_ms = MAX(attributes.search.took_ms)<p><strong>Result:</strong> Average 81ms, max 153ms.</p><p><code>search.took_ms</code> captures Elasticsearch's self-reported execution time, the <code>took</code> field from the search response. This is different from <em>span duration</em>, the wall-clock time from when the span started to when it ended (as we covered in the primer above). Span duration measures end-to-end time, including network round trips, serialization, and application logic. You want both: Comparing them tells you where overhead lives. If <code>took_ms</code> is 50ms but the span duration is 200ms, the extra 150ms is network or application overhead, not a query problem.</p><p>This attribute also keeps your analytics portable. If you're using OTel log records instead of spans (a lighter-weight alternative we mention in <a href="https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry">the first blog in the series</a>), there's no span duration. <code>took_ms</code> is the only timing signal you have.</p><p>We'll go deeper on search performance monitoring (Service Level Objectives [SLOs], alerting on latency regressions, and using this data for operational dashboards) in the last blog in the series focussing on Search Reliability Engineering.</p><p>Want to find the slow queries specifically?</p>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
  AND name == "search"
| STATS
    avg_ms = ROUND(AVG(attributes.search.took_ms), 0),
    max_ms = MAX(attributes.search.took_ms),
    search_count = COUNT(*)
  BY attributes.search.query
| SORT avg_ms DESC
| LIMIT 20<p>A query with high average latency and high result count is hitting many documents; consider query optimization. High latency with low results might mean complex filters or slow aggregations. Outlier max values are often cold caches or cluster issues.</p><h3>Query 6: How does search volume change over time?</h3><p>Counts, rates, and latencies tell you the <em>what</em>. Volume over time tells you the <em>when</em>: W<em>hen did traffic spike, when did it drop, and when did that zero-results rate jump?</em></p>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
  AND name == "search"
| EVAL bucket = DATE_TRUNC(5 minutes, @timestamp)
| STATS searches = COUNT(*) BY bucket
| SORT bucket<p><code>DATE_TRUNC(5 minutes, @timestamp)</code> rounds each timestamp down to the nearest 5-minute boundary. The result is a time series that Kibana's Lens can render as a bar chart or line, showing your search traffic pattern for any time window.</p><p>Narrow the bucket for higher granularity (<code>1 minute</code>), widen it for trend analysis (<code>1 hour</code>, <code>1 day</code>). When you add this to a dashboard alongside your zero-results rate, you can answer: <em>Did zero-results spike because traffic changed or because something broke?</em></p><h2>Verify that your search API instrumentation is working</h2><p>If you're using the reference project, the full setup is:</p>git clone https://github.com/elastic/elasticsearch-labs.git
cd elasticsearch-labs/supporting-blog-content/search-analytics-otel
cp .env.example .env           # fill in ELASTICSEARCH_URL, ELASTIC_API_KEY,
                               # OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_HEADERS
python3 -m venv venv &amp;&amp; source venv/bin/activate
pip install -r requirements.txt
python load_data.py             # index products into Elasticsearch
python app.py                   # starts on http://localhost:8000<p>Then trigger a search:</p>curl -X POST http://localhost:8000/api/search \
  -H "Content-Type: application/json" \
  -d '{"query":"laptop"}'<p>Wait 5–10 seconds for the <code>BatchSpanProcessor</code> to flush, and then open Kibana → Discover → switch to ES|QL mode and run:</p>FROM traces-generic.otel-default
| WHERE attributes.search.query IS NOT NULL
| LIMIT 5<p>You should see rows with <code>attributes.search.query</code>, <code>attributes.search.result_count</code>, and <code>attributes.search.took_ms</code>.</p><p>If no rows appear, check in order:</p><ol><li><p><code>OTEL_EXPORTER_OTLP_ENDPOINT</code> points to the mOTLP endpoint, not your Elasticsearch URL.</p></li><li><p><code>OTEL_EXPORTER_OTLP_HEADERS</code> includes <code>Authorization=ApiKey &lt;your-key&gt;</code>.</p></li><li><p><code>OTEL_TRACES_SAMPLER=always_on</code> is set (default sampler may drop spans).</p></li><li><p>Kibana → Observability → APM → Services shows <code>search-analytics-demo</code> (confirms export is working).</p></li></ol><h2>Turn ES|QL results into Kibana visualizations</h2><p>The bar chart that Discover auto-generates from your ES|QL results is a good start, but you can customize it. Click the <strong>pencil icon</strong> in the top-right corner of the chart to open the inline Lens editor.</p><p>From here you can:</p><ul><li><p>Change chart type (bar, line, area, pie, table, metric).</p></li><li><p>Adjust axes and add breakdown dimensions.</p></li><li><p>Save the visualization to a dashboard.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0e8869398e3cb98b/6a6a3404c40efb00b9d30954/e3db0028b0dff2ce5a524b62f600bfcfd8820f22-1440x708.png" alt="Kibana Lens configuration panel for a bar chart visualization of ES|QL search analytics data" /><p>This is the path from ad hoc ES|QL exploration to a persistent dashboard panel. You don't need to build visualizations from scratch; Discover and Lens handle the chart rendering from your query results.</p><p>Lens is Elastic's drag-and-drop visualization editor, and it's more capable than this quick workflow suggests. You can build multilayer charts, combine metrics with breakdowns, add reference lines, and design full dashboards that mix ES|QL panels with traditional aggregation-based visualizations. For search analytics, that means you can put top queries, zero-results trends, and latency percentiles side by side in a single view.</p><p>To go deeper:</p><ul><li><p><a href="https://www.elastic.co/guide/en/kibana/current/lens.html">Lens documentation</a>: A full guide to the visualization editor.</p></li><li><p><a href="https://www.elastic.co/docs/explore-analyze/visualize/esorql">ES|QL in Lens</a>: Using ES|QL queries as data sources for dashboard panels.</p></li><li><p><a href="https://www.elastic.co/guide/en/kibana/current/dashboard.html">Kibana Dashboards</a>: Building and sharing operational dashboards.</p></li><li><p>Use an <a href="https://www.elastic.co/docs/explore-analyze/ai-features/agent-builder/agent-builder-dashboards-and-visualizations">Agent built in Kibana to make visualizations for you</a></p></li></ul><p>We'll build a full search analytics dashboard in a later post.</p><h2>How sampling affects search analytics accuracy</h2><p>Most application performance monitoring (APM) configurations sample traces to control costs, capturing 10% or 25% of requests. For application monitoring, that's fine. For search analytics, it's a problem.</p><p>If you're sampling at 10%, your "total searches" count is 90% lower than reality. Your zero-results rate is still accurate (it's a ratio), but volume counts are off.</p><p>Two approaches:</p><p><strong>Approach 1: Configure 100% sampling for search endpoints.</strong> Your search API probably handles far fewer requests than your main application, so the data volume increase is manageable. The simplest way is through environment variables:</p># 100% sampling (capture every trace)
export OTEL_TRACES_SAMPLER=always_on

# Or sample a percentage (e.g. 50%)
export OTEL_TRACES_SAMPLER=traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.5<p>These are head-based sampling decisions made at the start of each trace. They apply globally to the service, which is fine if your search API is a dedicated service. If search shares a service with other endpoints and you need per-endpoint sampling rules, you can implement a custom <code>Sampler</code> in the OTel SDK that inspects the span name or attributes before deciding.</p><p>More sophisticated routing (sampling differently per endpoint, dropping noisy spans, or making decisions after a trace completes [tail-based sampling]) typically involves deploying an OTel Collector (such as the <a href="https://www.elastic.co/docs/reference/edot-collector">EDOT Collector</a>) as an intermediary between your application and Elastic. That's a valuable architecture pattern, but it's beyond the scope of this post. See the <a href="https://opentelemetry.io/docs/collector/">OTel Collector documentation</a> and <a href="https://www.elastic.co/docs/reference/edot-collector/modes">Elastic's EDOT Deployment</a> for more on collector-based sampling and routing architectures.</p><p><strong>Approach 2: Upscale in your queries.</strong> If you know the sampling rate, multiply: <code>EVAL estimated_total = total_searches * 10</code>. Ratios and averages stay correct; only absolute counts need adjustment.</p><p>For more on sampling strategies generally, see the <a href="https://opentelemetry.io/docs/concepts/sampling/">OTel sampling documentation</a>.</p><p>For the queries in this post, we used 100% sampling.</p><h2>What's next: Adding click tracking to search analytics</h2><p>The six ES|QL queries in this post answer: <em>What do users search for, what returns nothing, how fast is search, and when does traffic spike?</em> They're all derived from a single instrumentation point: the search span.</p><p>But they can't tell you whether users are finding what they need. A search that returns 15 results looks healthy from the server side. But if nobody clicks any of those results, your ranking has a problem.</p><p>In the next post, we add <em>click tracking</em>, a second span that captures which result the user clicked and where it appeared in the list. If you've been running the reference project, you already have 62 search spans; the next post builds directly on that data. With searches and clicks linked together, we'll calculate:</p><ul><li><p><strong>Click-through rate (CTR):</strong> What percentage of searches result in a click.</p></li><li><p><strong>Mean Reciprocal Rank (MRR):</strong> How far down the results users have to scroll.</p></li><li><p><strong>Click position distribution:</strong> The shape of where users click.</p></li></ul><p>The pattern is the same: You add attributes to spans and query them with ES|QL. You get a richer view of the same data without introducing new infrastructure.</p><h2>Resources to get started with search analytics on OpenTelemetry</h2><ul><li><p><a href="https://github.com/elastic/elasticsearch-labs/tree/main/supporting-blog-content/search-analytics-otel">Reference project</a>: Working code for the entire blog series; clone, configure, run.</p></li><li><p><a href="https://github.com/elastic/elastic-otel-python">EDOT Python</a>: Elastic distribution of OpenTelemetry for Python.</p></li><li><p><a href="https://www.elastic.co/docs/solutions/observability/apm/opentelemetry">OpenTelemetry with Elastic</a>: How to send OTel data to Elastic APM.</p></li><li><p><a href="https://opentelemetry.io/docs/languages/python/">OpenTelemetry Python SDK</a>: Upstream SDK documentation and instrumentation guides.</p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL documentation</a>: Query language reference.</p></li><li><p><a href="https://www.ubisearch.dev/">UBI Standard</a>: Reference schema for search event structure.</p></li></ul><p><em>This is the second post in a series on search analytics with OpenTelemetry and Elastic. Next up: Measuring search quality: Click tracking, CTR, MRR, and click position analysis.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/search-analytics-opentelemetry-esql</guid>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Matthew Adams]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf69e1ba73bd5d402/6a6a340599442cd9e0df1d94/1794a179d9536e693a0982634da59aff209c9d68-1280x720.png" length="0" type="image/png"/>
    <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Piping Hot: Bringing ES|QL to Your Grafana Dashboards Using the Elasticsearch Plugin]]></title>
    <description><![CDATA[You can now write ES|QL queries in Grafana with the Elasticsearch plugin. Learn how to enable it and write pipe-based queries directly in the Grafana UI.]]></description>
    <content:encoded><![CDATA[<p>The Elasticsearch data source is one of the most popular plugins in the Grafana ecosystem, and it now ships ES|QL support as an experimental feature, available starting in Grafana 13.0. ES|QL is Elasticsearch's modern pipe-based query language that enables querying logs, metrics, and traces, and using Elasticsearch as a native <a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">Prometheus PromQL</a> data source, all directly from the Grafana query editor. Built by Elastic in collaboration with Grafana Labs and contributed upstream to the Grafana open source project, this integration is enabled through a single feature flag (<code>elasticsearchESQLQuery = true</code>) that unlocks a Monaco-powered editor with syntax highlighting, autocompletion, and inline error messages. We'll walk through how to enable it and write your first queries for log analysis, time series visualization, and metrics aggregation.</p><p></p><h2>Context</h2><p>Elasticsearch is one of the top plugins used with the Grafana UI. Until now, the Elasticsearch plugin only supported Lucene and raw Query DSL for querying. <a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a> is Elastic's modern pipe-based query language, designed for analytics on log, metrics, and trace data. Its intuitive syntax makes it easier to filter, aggregate, and transform data compared to Query DSL or Lucene.</p><p>This has been tracked as a community request since 2024: <a href="https://github.com/grafana/grafana/issues/81765">grafana/grafana#81765</a>.</p><h2>How to enable ES|QL support in Grafana</h2><p>The feature is behind the <code>elasticsearchESQLQuery</code> feature flag. To turn it on, add the following to your <code>grafana.ini</code>:</p>[feature_toggles]
elasticsearchESQLQuery = true<p>Restart Grafana after saving. The feature is available starting with Grafana 13.0.</p><h2>ES|QL in the Grafana query editor</h2><p>Once the flag is enabled, the Elasticsearch query editor gains a <strong>Query language</strong> selector. You can switch between Lucene, Raw DSL, and ES|QL from the same editor panel.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc54a93910c32b6cf/6a4774fa74bff774dfa0697c/c7e83f6b498d8dbd8755e5d3086d0f44ad52e529-799x192.webp" alt="" /><p>When you select ES|QL, the editor switches to a Monaco-powered code editor, the same engine that powers VS Code. You get syntax highlighting, error highlighting, and basic autocompletion out of the box.</p><p><strong>Smart index pre-population</strong> makes getting started quick: if an index pattern is configured in your data source settings, clicking into the ES|QL editor for the first time auto-inserts <code>FROM &lt;index&gt;</code>. You can change or delete it freely. If no index is configured, the <code>FROM</code> clause is left blank.</p><h2>Running your first queries</h2><h3>Count log entries by severity</h3><p>This is a good first query to confirm ES|QL is working and to get a feel for the syntax.</p><p>Run it in the <strong>Raw Data</strong> panel type to see a table with counts per log level:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd6dea511499d72fd/6a4774fd2d406b4255ba39d1/74740e7fd1033c6e2d4c87b0c57f049360b9fcba-734x788.webp" alt="" /><h3>Browse the latest errors</h3><p><code>WHERE</code>, <code>KEEP</code>, <code>SORT</code>, and <code>LIMIT</code> make it easy to filter down to exactly the fields and rows you care about.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta276e207b97fe2d0/6a477500bfe8a289994da165/68b3dca2777f2dc81cb67bc28dc34fa8a961186f-726x909.webp" alt="" /><h3>Log volume over time</h3><p>Use <code>BUCKET</code> to group log counts into hourly intervals. This works well with the <strong>Metrics</strong> panel type, which can render it as a time series graph.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta68256304a4f0fec/6a4775023eae5d83abd13f90/510dee6a05a2de4af72b0b768411f6e54bd25a6f-722x865.webp" alt="" /><h3>Top hosts by log activity</h3><p>Identifying the most active hosts is a common operations task. <code>STATS</code> with <code>BY</code> and <code>LIMIT</code> makes it concise.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc01282b3689f3db9/6a47750531bdbbb2898b436a/e6be9497d82a46fc97d20417f0b6d85b6e418558-762x777.webp" alt="" /><h3>The TS command for time series metrics</h3><p>For metrics data in <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">Time Series Data Streams (TSDS)</a>, the <code>TS</code> source command in ES|QL enables metrics analytics and time series aggregation.</p><p>Examples:</p><ul><li><p><code>RATE()</code>: rate of change over time</p></li><li><p><code>AVG_OVER_TIME()</code>: average value over a sliding window</p></li><li><p><code>INCREASE()</code>: total increase over a period</p></li><li><p><code>DELTA()</code>: difference between first and last value</p></li><li><p><code>LAST_OVER_TIME()</code>: most recent value in a window</p></li></ul><p>The pattern follows a two-level aggregation: an inner function applied per individual time series, then an outer function aggregating across groups (for example, per host or per service).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt747d48504b7d826c/6a4775084887b866f94256f0/56fa77cdc4c1aa061fc27b54a781d4f8f5ad97ca-719x988.webp" alt="" /><p>You can also use <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions/avg_over_time"><code>AVG_OVER_TIME()</code></a> to compute the average value of a metric over a sliding window, then split the results by host and 10-minute buckets:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt411e15b564d3fcb9/6a47750b043a1f264c499021/5737d00f387383ec9a8d03a7d2588b28bd8cad00-708x908.webp" alt="" /><p><code>TS</code> also runs queries through the ES|QL vectorized compute engine. Internal benchmarks show performance improvements of an order of magnitude or more compared to equivalent Query DSL queries for TSDS-backed data.</p><p>Reference:</p><ul><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts">TS command documentation</a></p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions">Time series aggregation functions</a></p></li></ul><h2>Inline error messages</h2><p>ES|QL queries run against the <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-rest"><code>/_query</code></a> <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-rest">HTTP endpoint</a> on Elasticsearch. If your query has a syntax error or references a non-existent field, Elasticsearch returns a structured error response. The plugin surfaces this directly in the query editor as an inline message, so you see exactly what went wrong right in the Grafana UI.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb34a3ab9edbe00a1/6a47750dae1fcb770098685b/8e4b1a1ab78fe95c948f6e1f5ff8fe4a67507766-723x360.webp" alt="" /><p>In the example above, <code>host.nam</code> is missing the final <code>e</code>. Elasticsearch catches this as a verification exception and returns the field name that could not be resolved. That message appears inline, right below the query editor.</p><h2>Technical details</h2><p>Under the hood, the plugin handles ES|QL and other query types on separate code paths. ES|QL queries go to the <code>/_query</code> endpoint with <code>Content-Type: application/json</code>. Lucene and Query DSL queries continue to use <code>/_msearch</code> with <code>Content-Type: application/x-ndjson</code>.</p><p>This separation is intentional: <code>/_query</code> returns a different response shape that the plugin parses independently before passing data to Grafana panels.</p><h2>Try it out!</h2><p>That also means this is a good moment to try it and give feedback.</p><p>The <a href="https://github.com/grafana/grafana/pull/117798">upstream PR</a> and the <a href="https://github.com/grafana/grafana/issues/81765">original tracking issue</a> are public. If you run into problems or have requests, both are open for comments.</p><h2>Next steps</h2><ul><li><p>Enable the feature in your Grafana 13.0 instance with <code>elasticsearchESQLQuery = true</code></p></li><li><p>Try the example queries above against your own indices</p></li><li><p>For metrics data, give the ES|QL <code>TS</code> command a spin, against an Elasticsearch 9.2 or Serverless data source.</p></li><li><p>Read the full <a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL overview</a> to explore what else the language can do</p></li></ul><p>If you are not yet on Elasticsearch, you can start a free trial at <a href="https://cloud.elastic.co/registration">Elastic Cloud</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-grafana-elasticsearch-plugin</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-grafana-elasticsearch-plugin</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Cauê Marcondes]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2ec58a4e301413fb/6a477510bfe8a2121e4da169/17625309931e2620ff7c0584556f8bd62033483d-1376x768.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Don't leave metrics on the table: query them with the ES|QL TS command]]></title>
    <description><![CDATA[Recalibrate your mental model for time series queries: learn why FROM can produce inaccurate results for metrics, how TS fixes that, and when to use each command.]]></description>
    <content:encoded><![CDATA[<p>If you use ES|QL for logs and traces, <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/from"><code>FROM</code></a> is probably second nature, but on metrics it can return numerically wrong answers. A query like <code>FROM metrics-* | STATS SUM(request_count)</code> adds up cumulative counter values across every sample on every host. The result grows without bound and isn't a rate, a count, or anything else useful. <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> fixes that by grouping samples into time series first, then exposing functions like <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-rate"><code>RATE</code></a>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-avg_over_time"><code>AVG_OVER_TIME</code></a>, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-last_over_time"><code>LAST_OVER_TIME</code></a> that operate per series.</p><p>For a high-level tour of metrics analytics across ES|QL and Discover, see <a href="https://www.elastic.co/observability-labs/blog/metrics-explore-analyze-with-esql-discover">Explore and Analyze Metrics with Ease in Elastic Observability</a>. This post zooms in on the mechanics.</p><p>Here is the mental model in five bullets:</p><ul><li><p><code>FROM</code> treats every document as an independent row. That is right for events, but metric aggregations often need the time series that each row belongs to.</p></li><li><p><code>TS</code> adds that time series context: it groups and aggregates data points by time series before any other aggregation runs, and enables functions like <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-rate"><code>RATE</code></a>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-avg_over_time"><code>AVG_OVER_TIME</code></a>, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-last_over_time"><code>LAST_OVER_TIME</code></a>.</p></li><li><p>A <code>TS | STATS</code> query normally has two aggregation phases. The inner phase reduces samples inside each time series; the outer phase groups and combines those per-series results.</p></li><li><p>The default inner aggregation is <code>LAST_OVER_TIME</code>, which is why <code>TS metrics | STATS AVG(cpu_usage)</code> and <code>FROM metrics | STATS AVG(cpu_usage)</code> can return different numbers.</p></li><li><p>Use <code>TS</code> to query a time series data stream (<a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">TSDS</a>). Use <code>FROM</code> for events and raw document inspection.</p></li></ul><h2>What is a time series, really?</h2><p>A time series is a sequence of <code>(timestamp, value)</code> data points identified by the metric name and a unique set of dimension values.</p><p>For example, <code>request_count</code> reported every 30 seconds by host <code>h1</code> in data center <code>dc1</code> is one time series. The same metric on host <code>h2</code> in <code>dc1</code> is a different time series.</p><p>In a time series data stream, every metric document carries an internal <code>_tsid</code> field that uniquely identifies a time series. Samples that share a <code>_tsid</code> belong to the same time series and are stored sequentially, sorted by timestamp.</p><p>That storage layout enables efficient per-series aggregations. It also explains why <code>TS</code> only works on time series data streams. Other index modes have no notion of a time series, so the per-series operations <code>TS</code> relies on have no such identifier to attach to. <code>FROM</code> does not support those operations, which is what the next section is about.</p><h2>Why FROM leaves metrics on the table</h2><p>Consider a counter named <code>request_count</code> collected every 30 seconds from three hosts.</p><p>A counter is a cumulative metric: each sample is the running total since the process started reporting it. For <code>request_count</code>, a value of <code>1,000</code> means "this time series has observed 1,000 requests so far", not "1,000 requests happened since the previous sample". Counters reset to zero on process restart, so a sample of <code>4</code> right after <code>1,004</code> is a fresh count, not negative traffic. The ES|QL <code>RATE</code> function computes the per-second change within a time series and handles resets without glitches.</p><p>You want to calculate the total request rate across all hosts, bucketed by 5 minutes.</p><p>If you are used to writing ES|QL over event data, you might start with this query:</p><p>The chart it produces looks plausible at first: a line that goes up over time. But the number on the y-axis is the sum of every cumulative counter value reported in the bucket. Each host contributes its own running total, repeatedly, once per sample. Because the query uses <code>SUM</code> on those cumulative values, the result is not a rate, it is not the number of requests in the bucket, and it grows without bound even if the application stops receiving requests.</p><p><code>request_count</code> is a monotonically increasing counter, so its raw values represent "how many requests have ever happened on this host", not how many happened in the bucket. The right computation is "how much did this counter increase per second on each host, then sum across hosts." <code>FROM</code> cannot express that operation directly. It can group rows by fields, but it has no built-in notion of "the same time series over time" and no way to ask for the change of a counter within each time series. It also cannot use sliding-window time series functions such as <code>RATE(request_count, 5m)</code>, which we will come back to below.</p><p><code>TS</code> was introduced for this purpose, providing a succinct syntax to express time series aggregations:</p><p><code>RATE(request_count)</code> runs per time series and produces a per-second rate that handles counter resets correctly. <code>SUM</code> then adds those rates across hosts.</p><h2>Two aggregation phases: inner and outer</h2><p>Every <code>TS | STATS</code> query has two distinct aggregation phases.</p><p>Let's make that concrete with a query that calculates the request rate per data center:</p><p>The diagram below shows how <code>TS</code> evaluates this query. It first reduces samples inside each time series, then groups and combines those per-series values into one result per <code>datacenter</code> and time bucket.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte5b1f3d827f95ad1/6a1706f0961e6910afc4ce92/1d3d765ab8e27e539ac6bedf3e7444632a4bbe7e-3050x900.png" alt="Inner and outer aggregation phases of a TS|STATS query" /><p>The phases are:</p><p><strong>Inner (within a time series).</strong> Runs separately for each time series. It collapses many <code>(timestamp, value)</code> data points within a bucket into a single value per time series per bucket by applying the inner aggregation function, such as <code>RATE</code> in the example above. Functions: <code>RATE</code>, <code>AVG_OVER_TIME</code>, <code>MAX_OVER_TIME</code>, <code>LAST_OVER_TIME</code>, <code>STDDEV_OVER_TIME</code>, and so on. The full list is on the <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions">time series aggregation functions</a> page.</p><p><strong>Outer (across time series, the "grouping" phase).</strong> Combines the per-series values into a single value per group per bucket. Functions: <code>SUM</code>, <code>AVG</code>, <code>MAX</code>, <code>MIN</code>, percentiles, and the rest of the <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/aggregation-functions">regular ES|QL aggregates</a>.</p><p>In <code>SUM(RATE(request_count)) BY datacenter, TBUCKET(5m)</code>:</p><ul><li><p><code>RATE(request_count)</code> is the inner aggregation. It runs per time series.</p></li><li><p><code>SUM(...)</code> is the outer aggregation. It combines time series within the same <code>datacenter</code> and bucket.</p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/grouping-functions#esql-tbucket"><code>TBUCKET(5m)</code></a> defines the bucket boundaries (equivalent to <code>BUCKET(@timestamp, 5m)</code>).</p></li></ul><p>The outer aggregation is optional. If you only need the per-time-series result, use the time series aggregation function directly:</p><p>That query keeps the per-series rate for each bucket instead of wrapping it in <code>SUM</code>, <code>AVG</code>, or another aggregate across time series.</p><h2>The default inner aggregation: LAST_OVER_TIME</h2><p><code>TS</code> has to reduce raw samples inside each time series before it can run the outer aggregation. That means every metric field in a <code>TS | STATS</code> aggregation needs an inner aggregation, even when the query does not spell one out.</p><p>Consider a metric named <code>cpu_usage</code>. It is a gauge: a metric that captures a value at a point in time and can move up and down freely. A sample of <code>0.42</code> means "this host is at 42% CPU at this time". For a gauge, the natural "value in this bucket" is the most recent sample.</p><p>That is what ES|QL fills in for you. If you write <code>TS metrics | STATS AVG(cpu_usage) BY host.name, TBUCKET(5m)</code>, the implicit inner aggregation is <code>LAST_OVER_TIME(cpu_usage)</code> and the query is equivalent to:</p><p>For each time series, <code>LAST_OVER_TIME</code> picks the latest sample in the bucket. Then <code>AVG</code> averages across time series.</p><p>It is also why the same-looking query against <code>FROM</code> and <code>TS</code> can return different numbers. <code>FROM</code> averages every individual document. <code>TS</code> averages one value per time series per bucket. If your hosts publish at slightly different rates, those averages diverge. For example, in a five-minute bucket, a host that publishes every second contributes 300 documents while a host that publishes every two minutes contributes only two or three. With <code>FROM | STATS AVG(cpu_usage)</code>, the chatty host dominates the average. With <code>TS</code>, each time series is reduced to one bucket value first, so the outer average gives each host one value to contribute.</p><p>If you want the average value during the bucket instead of the latest value, make the inner aggregation explicit:</p><p><code>AVG_OVER_TIME</code> averages all CPU utilization samples within each time series. The outer <code>AVG</code> then averages those per-series values across matching hosts. That makes the result sample-weighted within each time series, then equally weighted across time series. Use this when you care about how the value behaved during the bucket, not just where it ended up.</p><p>The same rule applies to peaks and troughs. For a peak CPU chart, use <code>MAX(MAX_OVER_TIME(cpu_usage))</code>, not just <code>MAX(cpu_usage)</code>. The inner <code>MAX_OVER_TIME</code> finds the peak within each time series; the outer <code>MAX</code> finds the peak across matching time series.</p><p>Counters work the other way around. Their sample value is a running total, so the latest sample on its own is rarely meaningful. For a counter, the inner aggregation you almost always want is <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-rate"><code>RATE</code></a> for a per-second rate, or <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions"><code>INCREASE</code></a> for the total change in the bucket. Falling back on the default <code>LAST_OVER_TIME</code> gives you the most recent cumulative value, which is the trap the FROM query in the previous section walked into.</p><p>Pick the inner function deliberately. The outer function is the easy part.</p><h2>When to use TS, when to use FROM</h2><p>A practical rule of thumb:</p><ul><li><p>Use <code>TS</code> for metric aggregations against a <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">time series data stream</a>. It is the source command designed for that data, and it applies per-series semantics by default.</p></li><li><p>Use <code>FROM</code> for events: logs, traces, audit records, transactions. Each row is independent. There is no time series context.</p></li></ul><p><code>FROM</code> still works on TSDS indices and is occasionally useful, for example when you want to inspect raw metric documents without per-series grouping. For dashboards, alerts, and any kind of charting, <code>TS</code> is the right default.</p><p>If you first need to discover which metrics or time series exist in the data, use <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/metrics-info"><code>METRICS_INFO</code></a> or <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts-info"><code>TS_INFO</code></a> after <code>TS</code> and before <code>STATS</code>. See <a href="https://www.elastic.co/search-labs/blog/esql-metrics-info-ts-info-time-series-catalog">ES|QL METRICS_INFO and TS_INFO: Catalog your time series data</a> for a deeper walkthrough.</p><h2>Post-process TS results with ES|QL</h2><p>The first <code>STATS</code> command is the boundary between time series processing and regular ES|QL processing. Before that first <code>STATS</code>, <code>TS</code> needs to keep the data grouped by <code>_tsid</code>, so commands that change row order or shape are not allowed. After that first <code>STATS</code>, the output is a regular ES|QL table. You can sort it, limit it, join lookup data, enrich it, or compute derived columns.</p><p>For example, this query calculates average CPU per host and bucket, finds the maximum bucketed average for each host, and returns the ratio:</p><h2>Sliding windows for the inner aggregation</h2><p>Time series aggregation functions accept a second argument: the window size for the inner phase.</p><p>This computes the rate over a 5-minute sliding window, but reports a value every minute. It is useful when you want a smoother chart at fine bucket sizes.</p><p>The window is the ES|QL counterpart to a PromQL <a href="https://prometheus.io/docs/prometheus/latest/querying/basics/#range-vector-selectors">range vector selector</a>: <code>RATE(app.requests, 5m)</code> serves the same purpose as <code>rate(app_requests[5m])</code>.</p><h2>Gotchas worth knowing</h2><p>A few things in <code>TS</code> can seem surprising, especially when coming from the events-based <code>FROM</code> mental model. None of these are bugs; most are direct consequences of the per-series model. Here is what to watch for.</p><p><strong><code>COUNT(*)</code></strong> <strong>is rejected.</strong> Say you want to know how many samples were collected per service in each bucket. The instinct from <code>FROM</code> is <code>COUNT(*)</code>, but <code>TS</code> rejects it: there is no plain "row" once data is grouped by time series, so a row count has no defined meaning. Pick what you actually want to count:</p><ul><li><p>Number of samples per service: <code>STATS samples = SUM(COUNT_OVER_TIME(cpu_usage)) BY service.name, TBUCKET(5m)</code>. The inner <code>COUNT_OVER_TIME</code> counts samples per time series; the outer <code>SUM</code> adds them across the time series in the group.</p></li><li><p>Number of distinct hosts reporting per service: <code>STATS hosts = COUNT_DISTINCT(host.name) BY service.name, TBUCKET(5m)</code>. This counts unique label values across time series.</p></li></ul><p><strong>You cannot sort, limit, lookup join, or enrich before</strong> <strong><code>STATS</code></strong><strong>.</strong> <code>TS metrics | SORT @timestamp | STATS ...</code> will fail. The grouping by <code>_tsid</code> must happen first, before anything else can run. Filter with <code>WHERE</code> if you need to narrow the scope. After the first <code>STATS</code>, the output is regular ES|QL and you can pipe it through any command, as shown in the previous section.</p><p><strong>Gauge vs counter mapping.</strong> Time series functions are sensitive to the metric type set in the field mapping. <code>RATE</code> only works on counters; <code>*_OVER_TIME</code> functions are intended for gauges. If you build TSDS mappings by hand, pay special attention to this part.</p><p>This can be a source of friction for Prometheus users. Prometheus metric type metadata is not always available in the data Elasticsearch receives, so the metric type may have to be inferred from naming conventions (<code>_total</code> for counters, and so on). Those heuristics are imperfect, and a misclassified metric is rejected by the function that should accept it. The deeper mechanics, including how Prometheus Remote Write maps metric types into TSDS, are covered in <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">How Prometheus Remote Write Ingestion Works in Elasticsearch</a>.</p><p>Explicit converter functions (gauge-to-counter and counter-to-gauge) are on the roadmap to make these cases easier to recover from at query time.</p><p><strong>Kibana charts go empty when you zoom in too far.</strong> In Kibana, <code>TBUCKET</code> adapts to the date picker, so zooming in shrinks the bucket size. When the bucket size drops below the data's collection interval, every other bucket has no sample, <code>RATE</code> and the rest return null, and the chart silently goes blank. Elastic is evaluating mitigations such as a runtime warning when the bucket size is too small, a configurable minimum bucket size, or automatic widening of the window or bucket size.</p><h2>Wrap up</h2><p>For metric queries, start with <code>TS</code> unless you specifically need raw documents. Then choose the inner aggregation based on what the value should mean inside each time series: <code>RATE</code> for counters, <code>LAST_OVER_TIME</code> for current gauge values, and explicit <code>*_OVER_TIME</code> functions for peaks, averages, minimum values, or distributions.</p><p>Once the per-series value is right, the outer aggregation is the familiar part: group and reduce those time series into the chart, alert, or table you need.</p><p>For the full reference, see the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts">command docs</a> and the list of <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions">time series aggregation functions</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-ts-command-querying-metrics</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-ts-command-querying-metrics</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt81635cd2cc7703b3/6a1706ef1949f7a977e7a95c/e2eb1ba006612a352f1317c1621e4ebc5b2a12b6-1376x768.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[ Approximate queries in Elasticsearch ES|QL: 100x faster on billions of records, with built-in confidence intervals]]></title>
    <description><![CDATA[ES|QL now supports approximate query execution. Add one line to your queries, and get results orders of magnitude faster, with built-in confidence intervals that tell you exactly how much to trust them.]]></description>
    <content:encoded><![CDATA[<p>Add one line to any Elasticsearch Query Language (ES|QL) query, and get answers 100x+ faster on billions of documents. Your gains grow as your data grows. Built-in confidence signals tell you when results carry formal guarantees and when they’re best estimates.</p><h2>One line: Speed that scales with your data</h2><p>On billions of documents, analytical queries push against a real efficiency-precision trade-off. We’ve been hard at work pushing back. Our <a href="https://www.elastic.co/observability-labs/blog/elasticsearch-logsdb-storage-evolution">native columnar support</a> is one of the best there are. ES|QL itself is a fast, purpose-built analytical engine, <a href="https://www.elastic.co/search-labs/blog/esql-swiss-hash-stats">getting ever smarter at aggregation execution</a>. And Elasticsearch ships a steady stream of efficiency innovations, like <a href="https://www.elastic.co/search-labs/blog/Elasticsearch-sorting-speed-up">Block k-dimensional (BKD) tree pruning</a>, with more landing all the time.</p><p>But even with all of that, native approximate queries really shine through. Starting in Elasticsearch 9.4, ES|QL supports approximate query execution. All you have to do is add one line to your queries: Prepend <code>SET approximation = true</code>. Now Elasticsearch will automatically sample a subset of your data, run the aggregation on that sample, extrapolate the results, and report confidence intervals. All transparently.</p>SET approximation = true;
FROM logs-*
| STATS count = COUNT(*) BY time = BUCKET(@timestamp, 5 MINUTE)
| SORT time<p>Your existing query stays unchanged. The <code>SET</code> directive tells Elasticsearch to handle the sampling, extrapolation, and statistical validation for you. No query rewriting, no manual sampling math, no guessing at sampling probabilities.</p><p><code>SET approximation = true</code> is a forward-compatible directive. Today, it speeds up the most heavily used aggregations. As we expand support to more capabilities, your existing queries benefit automatically. Queries that aren’t yet approximated run exactly without errors; a warning header explains why.</p><h2>How much faster?</h2><p>On the ClickBench benchmark, well-behaved analytical queries ran on average 23x faster with confidence intervals enabled. Individual queries hit <strong>~100x</strong>. Disabling confidence interval computation, the highest-leverage queries land near <strong>~300x</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4ce2803349f14b4c/6a1710bb964cea663b08bcb3/fd32f0fddea4dd09a1ceddb9b9e226ac83794117-1999x1551.png" alt="" /><p>The advantage grows as your datasets grow. Approximate-mode cost is capped by the configured sample size, while exact-execution cost scales with row count. Doubling your index doubles exact-query time but barely changes approximate-query time for the same accuracy! This is a beautiful property of the underlying math, not an engineering trick, and it’s why approximation gets more valuable as you scale.</p><p>Speedup also depends on query shape, grouping cardinality, and sample size. See the FAQ for the full set of factors and tuning tips. Read “Fast approximate ES|QL” in two parts ([Part 1], [Part 2]), straight from the creators of the feature.</p><h2>What you get back</h2><p>The response includes your original aggregated values, automatically scaled to represent the full dataset; a <code>COUNT</code> on a 1% sample comes back as the estimated total, not the sample count. Column names and types are preserved (backwards-compatible). Plus two additional columns per approximated value:</p><ul><li><p><strong>Confidence interval:</strong> A range that bounds the true value at the configured confidence level (default 90%). For example, a count of 268,473 with interval [264,444–273,179] means you can be 90% confident the true count falls in that range.</p></li><li><p><strong>Certified flag:</strong> A Boolean indicating whether the confidence interval for that value meets formal statistical guarantees. When certified is <code>true</code>, the data distribution allows us to rely on the results at face value. When <code>false</code>, the approximation is still often close but we can’t claim the same formal guarantees, typically because the distribution may be highly skewed or involve too few documents in a group. Think of it as the difference between "statistically proven" and "best estimate."</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7f8356b68f5b8b31/6a1710bc0c48574b5501ab72/db6da3424aa6ee8b789d975f591afe7ca3a92501-1999x392.png" alt="" /><p>This is a deliberate design choice: Consumers that don't care about the confidence metadata can choose to not compute them at all (see “Granular control when you need it”) or ignore the extra columns and use the results exactly as before. Consumers that do care (like AI agents that read the results programmatically) get everything they need without a second query.</p><h2>Use cases for approximate queries in ES|QL: Where this matters</h2><h3>AI agents and agentic workflows</h3><p>Approximate queries don’t just speed up agent queries; they enable a <em>scan-then-enhance</em> investigation pattern that wasn’t practical at scale before. An agent can sweep billions of documents in sub-second time, identify candidates, and zoom in for exact answers, all inside a single reasoning loop. The <code>certified</code> flag turns approximation into a decision signal: Proceed at face value when it’s <code>true</code>, escalate to an exact query when it’s <code>false</code> and the step needs a tight guarantee. As ES|QL becomes the foundation for agentic analytics in Elastic, approximation is the speed layer that makes investigation possible at this scale.</p><h3>Dashboards and charts on large datasets</h3><p>Dashboards that aggregate weeks or months of data can become sluggish as data volumes grow. With <code>SET approximation = true</code>, the same dashboard loads faster. In the future, Kibana will inject the setting transparently, so users won't need to know it's happening; they’ll just see faster charts.</p><h3>Log pattern analysis in ES|QL at scale</h3><p><code>CATEGORIZE</code>, <code>GROK</code>, and regex-heavy conditions are among the most compute-intensive parts of ES|QL because they require nontrivial compute per document. With approximate execution enabled, these large-scale pattern and exploration workflows become practical on very large indices.</p><h3>Exploratory analysis and hypothesis testing</h3><p>When you're exploring data to form hypotheses, for example, "Which services have the highest error rates this week?", you rarely need exact counts. You need shapes, relative magnitudes, and outliers. Approximate mode gives you those at interactive speed, and the confidence intervals tell you when to switch back to exact mode for the final answer.</p><h2>How approximate queries work in ES|QL, without the math</h2><p>The speedup is real engineering, not a query-planner trick. Sampling happens at the Lucene layer: Elasticsearch reads only the documents in the sample, so I/O and compute savings are proportional to the sampling rate. The aggregation runs on the sample, and the result is automatically scaled to represent the full dataset.</p><p>Confidence intervals are computed by a bootstrap procedure over multiple sub-partitions of the sample: statistically rigorous, not a heuristic or a guess. This is what backs the <code>certified</code> flag: When the methodology’s assumptions are met, the intervals carry formal guarantees.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc957f508a88f6244/6a1710bed7c022a893de6586/a9bcb929a90d1b3f5777fbe121d01e078cf99b9c-1999x1574.png" alt="" /><h2>Granular control when you need it</h2><p>The defaults are designed to work well out of the box, but you can tune them:</p>SET approximation = {"rows": 500000, "confidence_level": 0.95};
FROM logs-*
| STATS count = COUNT(*), avg_duration = AVG(duration) BY service.name<ul><li><p><strong>rows:</strong> How many documents to sample (default: 100,000 for ungrouped queries, 1,000,000 for grouped). More rows means higher accuracy and longer runtime.</p></li><li><p><strong>confidence_level:</strong> The confidence level for intervals. Defaults to: 0.9. Set it to a higher level for an increased probability that the value is within the confidence interval.</p></li><li><p><strong>Skip confidence intervals for maximum speed:</strong> Set <code>confidence_level</code> to <code>null</code>, and Elasticsearch returns just the point estimates, adding another 2–5x speed on top of approximate execution. This is how the highest-leverage queries land near <strong>300x</strong>.</p></li></ul><h2>What's next</h2><p><code>SET approximation = true</code> is a forward-compatible directive. As we add support for <code>FORK</code>, <code>JOIN</code>, chained <code>STATS</code>, and additional aggregations, your existing queries automatically benefit.</p><p>Future work also includes tighter integration with Kibana so dashboards and Discover can enable approximation automatically and improved handling of highly skewed grouping fields.</p><p>Additionally, we’ll make approximate queries natively accessible to agents, so they can opt into fast execution as part of their analytics tools and reasoning loop.</p><h2>Get started</h2><p>Approximate queries are available in Elasticsearch 9.4 as a technical preview on the Enterprise subscription tier. Add <code>SET approximation = true;</code> to the beginning of your query, and see the difference. Check the <a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL SET command reference</a> for configuration options.</p><p>
<strong>FAQ</strong></p><p><strong>What is approximate query execution in Elasticsearch?</strong></p><p>Approximate query execution is a mode where Elasticsearch samples a subset of your data, runs the aggregation on the sample, and extrapolates the result to represent the full dataset. You get back the estimated value plus a confidence interval showing how much to trust it. It's controlled by a single <code>SET</code> directive prepended to your existing ES|QL query; no query rewriting required.</p><p><strong>How do I speed up ES|QL aggregations without reducing my data retention?</strong></p><p>Just add <code>SET approximation = true</code> to your query. Approximate execution samples at query time, not at index time. Your data stays fully indexed, fully retained, and queryable both exactly and approximately. Elasticsearch handles sampling and extrapolation on the fly. Drop the directive any time you want exact results; nothing about the underlying data changes.</p><p><strong>How much faster are approximate queries?</strong></p><p>On the ClickBench benchmark, aggregation-heavy ES|QL queries that are well-suited to sampling typically run 10–40x faster with confidence intervals enabled, with individual queries reaching 100x or more. Disabling confidence interval computation (<code>SET approximation = {"confidence_level": null}</code>) adds another 2–5x on top, so the highest-leverage queries hit nearly 300x. The advantage grows with dataset size: Sampling cost is capped by the configured sample size, while exact execution cost scales with the row count, so the bigger your index, the bigger the win for the same precision.</p><p><strong>How accurate are approximate queries? Can I trust the results?</strong> </p><p>Each approximated value comes back with two signals: a confidence interval (a range bounding the true value at a configurable confidence level) and a certified Boolean flag. When certified is <code>true</code>, the confidence interval carries formal statistical guarantees. When <code>false</code>, the result is still often close, but the data distribution didn't meet the assumptions required for a formal guarantee. Accuracy depends on data characteristics and query shape, not on document count, so speedup gains increase as your dataset grows.</p><p><strong>What does the speedup depend on?</strong></p><p>Five main factors:</p><ul><li><p>Dataset size. <strong>Larger datasets produce larger speedups</strong>, for the reason described above (exact scans grow with N; sampled scans don’t).</p></li><li><p>Query shape. Queries that scan a lot to compute relatively little (large <code>STATS</code>, especially <code>MEDIAN</code> and <code>PERCENTILE</code>) benefit most. Queries that are already cheap (small <code>WHERE</code> filters matching few rows, or simple counts that hit indexed summary statistics) see little speedup.</p></li><li><p>Grouping cardinality and distribution. Well-distributed <code>BY</code> fields with healthy per-group sample counts benefit cleanly. Very sparse or highly skewed grouping (for example, a near-unique field or a long tail of rare values) can erode the gain because rare groups end up with too few sampled documents.</p></li><li><p>Confidence interval computation. Computing intervals adds overhead. Set <code>confidence_level</code> to <code>null</code>, and you trade interval reporting for an additional 2–5x speedup.</p></li><li><p>Sample size. The defaults (100k for ungrouped <code>STATS</code>, 1M for <code>STATS … BY</code>) work well for most queries. Increasing rows improves accuracy on high-cardinality grouping at the cost of some speedup; decreasing it does the reverse.</p></li></ul><p><strong>Can I use approximate queries for log analysis and pattern detection?</strong> </p><p>Yes. <code>CATEGORIZE</code>, <code>GROK</code>, and regex-heavy conditions are among the most compute-intensive operations in ES|QL because they require per-document processing. With <code>SET approximation = true</code>, these operations run on a sampled subset instead of the full index, making large-scale log pattern analysis and exploration fast on very large datasets.</p><p><strong>Do I have to rewrite my ES|QL queries to use approximate mode?</strong> </p><p>No. Prepend <code>SET approximation = true</code> to your existing query. The aggregation expressions, column names, and output types stay the same. The response adds two columns per approximated value (the confidence interval and the certified flag), but existing consumers that don't use those columns see no breaking change.</p><p><strong>What aggregations does approximate mode support in 9.4?</strong> </p><p><code>COUNT</code>, <code>SUM</code>, <code>AVG</code>, <code>WEIGHTED_AVG</code>, <code>MEDIAN</code>, <code>PERCENTILE</code> (except extremes), <code>MEDIAN_ABSOLUTE_DEVIATION</code>, and <code>STD_DEV</code> (with caveats for highly skewed distributions). More coverage on the way.</p><p><strong>Will I get the same result twice for the same query?</strong></p><p>Not exactly. Approximate execution randomly samples documents at query time, so successive runs of the same query return slightly different point estimates and confidence intervals. The variation between runs is small relative to the confidence interval each run reports. If you need bit-for-bit reproducibility, run the exact query. For dashboards, depending on the use case, the variation can typically be smaller than the visual resolution of the chart.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/approximate-queries-esql-analytics</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/approximate-queries-esql-analytics</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Aris Papadopoulos]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc957f508a88f6244/6a1710bed7c022a893de6586/a9bcb929a90d1b3f5777fbe121d01e078cf99b9c-1999x1574.png" length="0" type="image/png"/>
    <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[From averages to any percentile: Elasticsearch ships native exponential histogram support in ES|QL]]></title>
    <description><![CDATA[Query any percentile at any time. Elasticsearch natively stores OTel exponential histograms and lets you analyze distributions in ES|QL without fixed buckets or lossy conversions.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch adds native support for OpenTelemetry exponential histograms in ES|QL. Unlike fixed-bucket histograms, exponential histograms dynamically adapt to your data — giving you accurate percentile estimates (median, p99, any percentile you want) at query time with guaranteed error bounds. No more pre-defining buckets, no more lossy conversions. </p><p>Just send your OTel metrics to the <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-otlp">Elasticsearch OTLP/HTTP endpoint</a> and they're stored using the new <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/exponential-histogram">exponential_histogram</a> type and queryable immediately. Already have historical data stored in the classic histogram type? A simple ::exponential_histogram cast in your ES|QL queries handles the migration transparently. Already using <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/downsampling-time-series-data-stream">downsampling</a>? Both histogram field types are now fully supported.</p><h2>Histogram metrics</h2><p>When dealing with metrics (in OpenTelemetry or Prometheus, for instance), counters and gauges are the most common metric types. Gauges allow you to monitor values that rise or fall (e.g., CPU utilization). Counters allow you to, well, count things, such as the total number of HTTP requests your service is handling. Counters normally just increase in value, with a few exceptions when they reset, like when a server reboots.</p><p>In the case of counters, you can additionally collect a counter measuring the total sum of your HTTP response times, which allows you to derive the average response time by dividing that sum by the total number of requests. However, average response times provide limited insights into the collected data and the system behavior. The best insights are gained by analyzing the collected metric distribution, e.g., through median and percentile calculations. This is where counters fall short.</p><p>In the past, workarounds have been applied: For example, classic Prometheus-style histograms attempt to capture the distribution using a set of counters. By defining fixed buckets (e.g., one for response times in the range <code>[0s, 1s)</code>, one for <code>[1s, 4s)</code>, and so on) and associating a counter with each, we can at least estimate percentiles broadly. However, the key problem here is that we have to know the distribution of our data up front to properly define these buckets.</p><p>To that end, the OpenTelemetry community has come up with a better solution: exponential histograms. Exponential histograms assign collected values to buckets, just like classic Prometheus-style histograms. The key differentiator is that these buckets vary dynamically based on the collected values. The name "exponential" comes from the fact that the bucket sizes increase exponentially: we use small buckets for small values and wider buckets for larger values. You can find an excellent introduction in the <a href="https://opentelemetry.io/blog/2022/exponential-histograms/">OpenTelemetry exponential histograms introduction</a>.</p><p>Note that in addition to classic histograms, Prometheus also added <a href="https://prometheus.io/docs/specs/native_histograms/">native histograms</a>, which directly map to OTel <a href="https://prometheus.io/docs/specs/native_histograms/#opentelemetry-interoperability">exponential histograms</a>. Native histograms have their own <a href="https://prometheus.io/docs/specs/native_histograms/#promql">PromQL syntax</a>. We are actively working on adding support for that syntax to the <a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">Elasticsearch PromQL implementation</a>, so that you can directly query exponential histograms using PromQL.</p><h2>Demo setup</h2><p>Let's start by collecting some histogram metrics to show how they can be stored and analyzed in Elasticsearch using ES|QL.</p><p>We'll focus on a Java JVM metric: garbage collection durations. OpenTelemetry defines the <a href="https://opentelemetry.io/docs/specs/semconv/runtime/jvm-metrics/#metric-jvmgcduration">jvm.gc.duration</a>, which is a histogram-typed metric. The <a href="https://github.com/open-telemetry/opentelemetry-java-instrumentation">OpenTelemetry Java agent</a> natively supports collecting this metric.</p><p>We'll spin up a JVM running a <a href="https://renaissance.dev/">Renaissance benchmark</a> to put it under stress. We'll start that JVM with the vanilla OpenTelemetry Java agent attached and have it send the metrics directly to Elasticsearch.</p><p>You can find the ready-to-run Docker-compose file <a href="https://github.com/JonasKunz/es-histogram-demo">here</a>. You'll just need to insert your <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-otlp">Elasticsearch OTLP/HTTP endpoint</a> and API key in the <code>docker-compose.yml</code>:</p>OTEL_EXPORTER_OTLP_ENDPOINT: https://&lt;elasticsearch url&gt;/_otlp
OTEL_EXPORTER_OTLP_HEADERS: "Authorization=ApiKey &lt;base64 API key&gt;"<p>Note that you don't have to use this demo setup. We even encourage you to try it with your own application. Here are the other important OpenTelemetry agent settings the demo already includes, which you should include too if you're bringing your own app:</p>OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE: delta
OTEL_EXPORTER_OTLP_METRICS_DEFAULT_HISTOGRAM_AGGREGATION: BASE2_EXPONENTIAL_BUCKET_HISTOGRAM
OTEL_INSTRUMENTATION_RUNTIME_TELEMETRY_ENABLED: "true"<p>Let's step through them:</p><ul><li><p><em>Temporality preference</em>: OpenTelemetry supports both cumulative and delta-based histograms. Cumulative means that the histogram is only cleared after an application restart, while delta clears it after each export. At the time of writing, Elasticsearch only supports delta temporality for histograms. We are actively working on supporting cumulative histograms as well.</p></li><li><p><em>Default Histogram Aggregation</em>: By default, OpenTelemetry exports histograms in the Prometheus-style fixed bucket format. Since we want to reap the benefits of exponential histograms, we tell the agent to use them instead.</p></li><li><p><em>Runtime Telemetry enabled</em>: This tells the agent to actually collect the detailed JVM metrics, which include <code>jvm.gc.duration</code>.</p></li></ul><p>Now we are ready to go! We'll let the application run in the background and switch over to Kibana to analyze the GC metric.</p><h2>Querying with ES|QL</h2><p>Now let's open up Kibana and navigate to "Discover". There we'll switch to <a href="https://www.elastic.co/docs/explore-analyze/discover/try-esql">ES|QL mode</a>, and start querying the collected data:</p><p>As a response, we now see the metric panel shown below. If you don't see any data, make sure to double-check the Kibana <a href="https://www.elastic.co/docs/explore-analyze/query-filter/filtering#set-time-filter">time range filter</a>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt00f1fc071c411284/6a170849acf08862a9be9a98/b863b2e272ac6584ac193661a6c4419abffdd243-729x190.png" alt="ES|QL metric panel showing the total count of jvm.gc.duration samples" /><p>This number represents the total number of garbage collection operations that happened in our test application during the selected time frame.</p><p>Similarly, we can query the total time spent on those garbage collection operations:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc0fc971c91c0f5b9/6a17084b6f7f04323b914792/eda37c5fa244a42258bb452d18f5cbab3ff76eaf-717x190.png" alt="ES|QL metric panel showing the sum of jvm.gc.duration values in the selected time range" /><p>So we have roughly 270k garbage collections, which in total took 713 seconds. Given these two numbers, we can now compute the average if we are still fluent in primary school-level math. Even if not, you can just let ES|QL do that for you:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbfd0b056ac38d0b2/6a17084c60084b3be93c44d5/9fc5d601604b05378beeb4e6e94b613f95fe2fbc-712x188.png" alt="ES|QL metric panel showing the average jvm.gc.duration value" /><p>Now we know that the average garbage collection operation took about 3 milliseconds. However, Java experts might know that there are different kinds of garbage collections happening, which can have significantly different pause times. Fortunately the OpenTelemetry metric comes with attributes, which allow us to slice the data accordingly:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcedc19c4c217bb91/6a17084e14b2708f3ce3c5a4/535d44cdb2ec3ed2ac9afe6b259d7b69c9167bbd-989x476.png" alt="ES|QL bar chart showing the average jvm.gc.duration grouped by jvm.gc.action" /><p>As expected, major garbage collections take a lot more time per collection than minor ones, at least on average. So far, we have done nothing you couldn't also achieve by just using counters. Let's now use histograms to understand the actual distribution of the GC latency. We'll look at the data over time (by grouping using <code>TBUCKET</code>) and focus on the major garbage collections:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1c18e5836bec77a4/6a17084f47d49cc38e2d8937/ce62a4498d5a6e3fcc3bc85ea78f45a18f6e7576-1016x477.png" alt="ES|QL line chart of min, median, p99 and max jvm.gc.duration for major garbage collections" /><p>The graph now shows us the minimum, maximum, median and 99th percentile for major garbage collections. Note that we aren't bound to only querying the median and the 99th percentile. We can query any percentile we'd like to see, as these are estimated at query time from the raw exponential histograms.</p><h2>A note on backwards compatibility</h2><p>So far, we have seen how you can use the new shiny toy in Elasticsearch and ES|QL: exponential histograms. However, since this has just reached general availability (GA) in the 9.4 release, what about your historical data?</p><p>Before exponential histograms were added, Elasticsearch was already capable of storing OpenTelemetry histograms in the <code>histogram</code> field type. To do so, we converted them to a different data structure supported by the <code>histogram</code> field type: <a href="https://github.com/tdunning/t-digest/blob/main/docs/t-digest-paper/histo.pdf">T-Digest</a>. T-Digest provides good accuracy for extreme percentiles (e.g., 99th percentile) at the cost of accuracy for percentiles in the middle of the distribution, such as the median. In contrast, exponential histograms provide a guaranteed upper bound on the relative error for every percentile. As conversions always introduce errors, we are happy to now have native support for exponential histograms, allowing you to collect and analyze your metrics end-to-end without unnecessary conversions.</p><p>But still, what should you do if you have historical data and still want to query it? Thanks to <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-multi-index#esql-multi-index-union-types">ES|QL union types</a>, the answer is actually easy: You just have to add a <code>::exponential_histogram</code> suffix to the histogram metrics in your queries:</p><p>When this query encounters <code>histogram</code> fields, it will attempt to convert them to exponential histograms. When operating on <code>exponential_histogram</code> fields, the <code>::exponential_histogram</code> cast has no effect. Note that this also works with mixed data sets: if your backing indices use both types, the query will just do the right thing.</p><p>So if you are building queries or dashboards that you expect to run on pre-9.4 ingested data, we recommend that you simply add: <code>::exponential_histogram</code> casts.</p><h2>Wrapping up</h2><p>Native support for OpenTelemetry exponential histograms in Elasticsearch gives you better metric fidelity and more flexible analysis in ES|QL. In this blog post, we have shown you how to easily ingest and analyze your histogram metrics with ES|QL using various aggregations and the impact exponential histograms have.</p><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/exponential-histogram">Exponential histograms</a> are <strong>generally available</strong> in Elasticsearch basic starting with the 9.4.0 release. They will be available in Elastic Cloud <a href="https://www.elastic.co/cloud/serverless">Serverless</a> a few weeks after the 9.4.0 release, once <a href="https://www.elastic.co/docs/reference/opentelemetry/motlp">mOTLP</a> (the managed observability OTLP intake) switches to use the Elasticsearch OTLP endpoint. We'll update this blog post and add a note on the Elastic Cloud Serverless release notes when that happens.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/otel-histogram-metrics-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/otel-histogram-metrics-esql</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Jonas Kunz]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blteb9d15d1bec33675/6a1708511949f75484e7a985/f44560ece4dcc46e6a01826b597e094169e99691-848x477.png" length="0" type="image/png"/>
    <pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Three indices walk into a FROM clause: ES|QL subqueries in Elasticsearch]]></title>
    <description><![CDATA[ES|QL subqueries give each data source its own pipeline and filters, eliminating CASE chains, restoring predicate pushdown, and making multi-index queries extensible by design.]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language</a> (ES|QL) now has <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-subquery">subqueries in </a><a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-subquery"><code>FROM</code></a>. Three indices, different schemas, one query; each source gets its own pipeline with its own filters and transforms. No more <code>CASE</code> chains. No more client-side stitching. Add a fourth source? Add a fourth branch; zero changes to the existing three.</p><h2>The problem: Heterogeneous data, one query</h2><p>Consider a production incident investigation. Errors are spread across three microservices: an API gateway, a payments service, and an auth service, each with different field names and different conventions. Before subqueries, combining them in a single ES|QL query meant cramming everything into one <code>FROM</code> with <code>CASE</code> chains:</p><p>This is brittle and slow. The disjunctive <code>OR</code> prevents predicate pushdown; every index scans every condition. Every <code>CASE</code> chain grows with every source. Copy it into five dashboards and three alert rules, and you have eight places to update when anything changes.</p><h2>The fix: Independent pipelines</h2><p>Subqueries replace the monolithic <code>FROM</code> + <code>CASE</code> pattern. Each data source gets its own complete pipeline:</p><p>The gateway branch only scans for HTTP 500s. The payments branch only looks at transaction statuses. The auth branch only checks login failures. Because each branch has its own <code>WHERE</code>, the optimizer pushes filters independently into each index, restoring the predicate pushdown that a single <code>FROM</code> with <code>OR</code> conditions prevents. Fields that exist in one branch but not another are filled with <code>null</code>.</p><p>Adding a fourth service means adding a fourth branch. Existing branches don't change.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltad1d619d46212715/6a170d64e8fbce71fd39fcd8/f210b98824cb34c514b6cdb2d44b2bc81a51bb7b-1999x1084.png" alt="Diagram comparing two approaches. The left side is labeled “Before: EVAL + CASE: Sequential &amp; Brittle” and shows a vertical sequence with components titled “Complex Filtering (OR Conditions),” “EVAL source = CASE (dataset == …),” “Sequential Processing Chains,” “Linear Transformation,” and “Final Aggregation,” with callouts for “Choke Point” and “Latency Bottleneck.” The right side is labeled “After: Subqueries: Parallel &amp; Optimized” and shows three parallel branches for “weblogs-,” “applogs-,” and “securitylogs-*,” each with its own WHERE clause and EVAL source assignment, feeding into a “Unified Output Stream (UNION ALL Semantics).” A code block appears below the branches." /><h2>Save it as a view</h2><p>This is where subqueries and <a href="https://www.elastic.co/search-labs/blog/elasticsearch-esql-logical-views">logical views</a> combine. Wrap the subquery above in a named view, with one API call:</p>PUT _query/view/error_triage
{
  "query": "FROM (FROM svc-gateway-* | WHERE ...) , (FROM svc-payments-* | WHERE ...) , (FROM svc-auth-* | WHERE ...)"
}<p>Now consumers just write <code>FROM error_triage | STATS error_count = COUNT(*) BY service</code>. Three indices, three pipelines, one name. If you have 10 dashboards and five alert rules consuming this pattern, that's 15 copies of the same logic today; with a view, it's one definition and zero consumer-side edits when you add a fourth service. See <a href="https://www.elastic.co/search-labs/blog/elasticsearch-esql-logical-views">Elasticsearch ES|QL Views</a> for the full views deep dive.</p><h2>What you can do inside a branch</h2><p>Each branch supports the full ES|QL pipeline: <code>WHERE</code>, <code>EVAL</code>, <code>STATS</code>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join"><code>LOOKUP JOIN</code></a>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/enrich"><code>ENRICH</code></a>, and more. See the <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-subquery">subquery documentation</a> for the complete list.</p><h2>Aggregate different metrics, and then combine</h2><p>Each branch can compute its own summary before results are merged. This is useful when different indices track the same concept under different field names:</p><p>Both branches produce <code>avg_latency</code> and <code>hour</code>, but each computes it from a different source field. The combined result is a single table you can chart or alert on, without normalizing field names at ingest time. This pattern is impossible with a single <code>FROM</code>; you can't compute different aggregations per index without subqueries.</p><h2>Subqueries vs. FORK</h2><p>ES|QL also has <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fork"><code>FORK</code></a> (now generally available), which creates parallel execution branches from the same input. The distinction:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6dd3fd7c3afbea5a/6a170d66b0367d56f572bd97/b5728fa2ce08df32b67ef9368c4d75d04b74ab95-1999x992.png" alt="Chart titled “Subqueries vs. FORK” with two columns comparing Subqueries and FORK across four rows: Data Sources, Processing, Mental Model, and Use Case. The Subqueries column shows different indices labeled Index A and Index B, independent FROM pipelines labeled pipeline 1 and pipeline 2, the mental model “Many inputs → one combined result,” and a use case combining web, app, and security logs with corresponding icons. The FORK column shows the same input data labeled Index A, transforms applied to the same rows labeled Transform A and Transform B, the mental model “One input → many analyses,” and a use case of full‑text search and KNN on the same index with magnifying glass and graph icons." /><p>Different indices → subqueries. Same data, different analyses → FORK.</p><h2>How this compares</h2><p>If you're coming from other query languages, here's how ES|QL subqueries stack up at the time of writing:</p><p><strong>Splunk SPL/SPL2</strong> has <code>append</code> and <code>multisearch</code> in classic SPL, and SPL2 adds a <a href="https://help.splunk.com/en/splunk-cloud-platform/search/spl2-search-reference/union-command/union-command-examples">union command</a> that merges events from multiple datasets (the closest analogue to ES|QL subqueries). Federated Search extends this across remote Splunk deployments (analogous to CCS). The differences are in how the engine handles each branch: ES|QL subqueries give each branch independent predicate pushdown, meaning filters are pushed into each index's shard-level structures separately. SPL2 <code>union</code> merges datasets but optimization across branches is limited to what the search scheduler can parallelize. Wrapping ES|QL subqueries in a <a href="https://www.elastic.co/search-labs/blog/elasticsearch-esql-logical-views">view</a> gives you engine-level encapsulation with role-based access control (RBAC); Splunk's equivalent is saved searches and macros, which are text substitution expanded at parse time.</p><p><strong>SQL databases</strong> have <code>UNION ALL</code>, which is the closest analog. The difference is that SQL <code>UNION ALL</code> typically requires matching column counts and types at parse time. ES|QL subqueries are more forgiving; columns that exist in one branch but not another get null-padded automatically, which matters when your sources have different schemas (the norm in observability data). SQL views solve the reuse problem similarly, but ES|QL views are cluster-level objects, not database-scoped; they work across <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-cross-clusters.html">cross-cluster search</a> boundaries.</p><p><strong>Grafana / Datadog / other dashboarding tools</strong> handle multisource composition at the visualization layer: Run separate queries, merge in the panel. This works for display but breaks for alerting, downstream queries, and anything that needs a single result set programmatically. ES|QL subqueries push the composition into the engine, so alerts, views, and API consumers all get the same unified result.</p><p>Capability</p><p>Splunk SPL/SPL2</p><p>SQL UNION ALL</p><p>Dashboard-layer merge</p><p>ES|QL subqueries</p><p>Independent filters per source</p><p>SPL2 `union` merges datasets; optimization is scheduler-level</p><p>Yes</p><p>N/A (separate queries)</p><p>Yes; parallel with pushdown</p><p>Schema mismatch handling</p><p>Manual field normalization</p><p>Strict column matching</p><p>Manual in panel config</p><p>Automatic null-padding</p><p>Engine-level reuse</p><p>Text macros (parse-time expansion)</p><p>Database-scoped views</p><p>Dashboard variables</p><p>Cluster-level views with RBAC</p><p>Works for alerts + API</p><p>Limited (summary indexing)</p><p>Yes</p><p>No; display only</p><p>Yes</p><p>Add a source</p><p>Edit every macro/saved search</p><p>Add a UNION branch</p><p>Add a panel query</p><p>Add a branch; existing branches unchanged</p><h2>Current constraints</h2><p>In the Tech Preview release, subqueries are non-correlated; branches run independently and can't reference the outer query. They're supported in <code>FROM</code> only (not <code>TS</code>), and <code>FORK</code> can't be used inside or after subqueries. See the <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-subquery">subquery documentation</a> for details.</p><h2>What's next for subqueries</h2><p><a href="https://github.com/elastic/roadmap/issues/60"><code>WHERE</code></a><a href="https://github.com/elastic/roadmap/issues/60"> subqueries</a> — <code>WHERE field IN (FROM other_index | ...)</code> and other correlated forms — will extend the composition model from <code>FROM</code> into filtering. This brings the familiar SQL pattern of nested filtering to ES|QL.</p><h2>Try it</h2><p>Subqueries in <code>FROM</code> are available as a Tech Preview. Try them in <a href="https://www.elastic.co/kibana">Kibana</a> Dev Tools or Discover. We'd love your feedback; file a <a href="https://github.com/elastic/elasticsearch/issues">GitHub issue</a> with the <code>ES|QL</code> label.</p><p><em>ES|QL subqueries in FROM are a Tech Preview feature. Tech Preview features are subject to change and are not covered by the support SLA of GA features. The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-subquery-from</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-subquery-from</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Tyler Perkins]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltddbbebe635e925fd/6a170d681949f7318ae7aaa5/2eb755dd2b2b69b8e0e8867a0da85940eb744176-1280x720.png" length="0" type="image/png"/>
    <pubDate>Wed, 06 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch ES|QL views: One query to rule twelve dashboards]]></title>
    <description><![CDATA[With ES|QL views, you only need one query for multiple dashboards. Define it once and let Elasticsearch keep everything in sync.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch Query Language (ES|QL) now has <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-views">logical views</a>. Define a query once, and reference it by name in <code>FROM</code>, like an index. Twelve dashboards, one definition, zero copy-paste. Update the view, and every consumer gets the change automatically.</p><p>Views don't store data; they re-execute on every read, so results always reflect the current data and the current definition. If you've used views in SQL databases, this will feel familiar. The difference: ES|QL views are engine-level virtual indices stored at the Elasticsearch cluster level, not saved query text that gets expanded client-side. They appear in <a href="https://www.elastic.co/kibana">Kibana</a> autocomplete, support <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-cross-clusters.html">cross-cluster search</a> (CCS), and are governed by dedicated role-based access control (RBAC) privileges.</p><h2>A simple view</h2><p>A view can wrap any ES|QL query. Start with a straightforward filter — HTTP 500 errors from the API gateway:</p><p>Now anyone can write <code>FROM error_triage</code> without knowing the index pattern or filter condition:</p><p>The query is defined once. Consumers reference a name.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdab7918cab2c71e5/6a1706e8cf4f2555e2b2d0d2/68ff5a52b0f3ed3dfaa07d2af6e7f08a8c9c0f55-1999x702.png" alt="Dashboard interface showing a query editor at the top with the text “FROM error_triage | STATS error_count@error_triage | SORT error_count DESC.” A bar chart displays error counts for three services: payments, gateway, and auth. Below the chart, a results table lists the same services with error counts of 194 for payments, 37 for gateway, and 19 for auth." /><p>Views support full create, read, list, update, and delete (CRUD) via the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-esql-put-view"><code>_query/view REST API</code></a>.</p><h2>Update propagation</h2><p>Say the team decides <code>error_triage</code> should also capture client errors, not just 500s. Update the definition in place:</p><p>Every dashboard panel, alert rule, and ad-hoc query using <code>FROM error_triage</code> immediately reflects the broader filter. No saved objects to hunt down. No stale copies. Change once, update everywhere.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcd475eb5e0a666f8/6a1706ea6234e0e4dedb1973/aad2b22fd4d45f8db1140ae93429c6b9ca345031-1999x440.png" alt="Side‑by‑side comparison with the headings “Without Views” on the left and “With Views” on the right. Under “Without Views,” a central document connects to multiple window icons with tangled red arrows and the caption “Manual find‑and‑replace across saved objects.” Under “With Views,” a central document labeled “critical_errors” connects to similar window icons with green arrows and green check marks, along with the caption “Change once, update everywhere automatically.”" /><h2>Nested views</h2><p>Views can reference other views, enabling layered abstractions. Create views for suspicious IPs and threat intelligence, and then compose them:</p><p>Security teams query <code>FROM security_overview</code> without knowing the underlying data model. They're also shielded from any changes made to <code>suspicious_ips</code> by its owner; the abstraction boundary is real, not syntactic.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt74d422d44d683f0e/6a1706ec67045b3abd45c13e/fd8f17e0a5bd80a73737f0fe08f9e3480d70ad68-1822x642.png" alt="Diagram titled “From Text to Topography,” showing three stacked layers. The bottom layer, labeled “Bedrock – Raw Indices,” contains “svc-auth-*” and “threat-intel.” The middle layer, labeled “Base Views,” contains “suspicious_ips” and “known_threats.” The top layer, labeled “Operational View,” contains “security_overview.” Arrows point upward from each lower layer to the next. A text box on the right states that consumers only query the top layer and changes below cascade upward instantly." /><h2>Multisource views with subqueries</h2><p>A view can wrap any ES|QL query, including multisource compositions, using <a href="https://www.elastic.co/search-labs/blog/esql-subquery-from"><code>subqueries in FROM</code></a>. Each subquery branch queries one service independently (its own filters, its own field normalization), and the results combine automatically:</p><p>Consumers just write:</p><p>Two indices, two independent pipelines, one name. To add a third service later, add a third branch; existing branches don't change, and every downstream dashboard and alert reflects the update automatically. For a deep dive on subquery syntax and what you can do inside each branch, see <a href="https://www.elastic.co/search-labs/blog/esql-subquery-from">Three Indices Walk Into a FROM Clause</a>.</p><h2>How views work under the hood</h2><p>When you write <code>FROM view_name</code>, ES|QL resolves the view's stored query and executes it inline. Views are re-executed on every read, so results always reflect the current data and the current definition.</p><p>Views share a namespace with indices, aliases, and data streams. A view cannot have the same name as any of these (enforced at creation time). This keeps <code>FROM my_name</code> unambiguous regardless of whether the name resolves to a view, an index, or an alias.</p><h2>Security model</h2><p>Views are governed by four dedicated RBAC privileges: <code>create_view</code>, <code>read_view_metadata</code>, <code>delete_view</code>, and <code>manage_view</code>. Elasticsearch checks the privileges of the user running the query (invoker security), not the user who defined the view. The user querying a view needs permissions on both the view and its underlying indices.</p><h2>Kibana integration</h2><p>Views appear in Discover's ES|QL editor autocomplete alongside indices. ES|QL-based dashboard panels work with views transparently. In the initial Tech Preview release, view management is API-only. A Kibana UI for creating and managing views is planned.</p><h2>Cross-cluster search</h2><p>A view's definition can reference remote indices using <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-cross-clusters.html">CCS syntax</a>:</p><p>Consumers query <code>FROM cross_cluster_errors</code> without knowing which clusters are involved.</p><h2>Current constraints</h2><p>In the Tech Preview release, view management is API-only and SET directives can't appear inside view definitions; the caller applies them when querying. Subquery-based views can't be nested inside other multisource <code>FROM</code> expressions. See the <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-views#esql-views-limitations">views documentation</a> for the full list.</p><h2>What's next for views</h2><p>Views today are always fresh; they re-execute on read. <a href="https://github.com/elastic/roadmap/issues/49">Materialized views</a> flip that tradeoff: Pre-compute once, read instantly. Think pre-aggregated rollup views for Service Level Agreement (SLA) dashboards that load in milliseconds instead of scanning raw data on every refresh. A Kibana CRUD UI for views, including a "Save as View" workflow in Discover, is also planned.</p><h2>Try it</h2><p>Logical views are available as a Tech Preview. Try them in <a href="https://www.elastic.co/kibana">Kibana</a> Dev Tools or Discover. We'd love your feedback; file a <a href="https://github.com/elastic/elasticsearch/issues">GitHub issue</a> with the <code>ES|QL</code> label.</p><p><em>ES|QL logical views are a Tech Preview feature. Tech Preview features are subject to change and are not covered by the support SLA of GA features. The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-esql-logical-views</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-esql-logical-views</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Tyler Perkins]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5417d25f64c77933/6a1706ed1949f76551e7a958/852bff427ac62b79974d88e27ce9670dc132bc46-1280x720.png" length="0" type="image/png"/>
    <pubDate>Tue, 05 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch ES|QL query builder for JavaScript and TypeScript: Fluent, type-safe query construction]]></title>
    <description><![CDATA[Exploring the Elasticsearch ES|QL query builder for JavaScript and TypeScript and explaining how to build ES|QL queries with practical examples.]]></description>
    <content:encoded><![CDATA[<p>We're pleased to announce that the Elasticsearch Query Language (ES|QL) query builder is now available for JavaScript and TypeScript. It's a fluent, type-safe library that lets you construct ES|QL queries with method chaining, automatic value escaping, and full integrated development environment (IDE) support; no more raw string concatenation.</p><p>Learn how to get started with practical examples you can use right away.</p><h2>Elasticsearch ES|QL query builder for JavaScript and TypeScript</h2><p>If you've ever built an ES|QL query in JavaScript, you've probably written something like this:</p>const query = `FROM logs-*
| WHERE status_code &gt;= ${minStatus}
  AND host.name == ${hostname}
  AND @timestamp &gt;= "${startDate}"
| STATS error_count = COUNT(*) BY status_code
| SORT error_count DESC
| LIMIT 10`<p>It looks fine until <strong><code>hostname</code></strong> is<strong><code>O'Brien's server</code></strong> and the whole thing blows up with a parse error. Or until a user passes <strong><code>"; DROP INDEX logs</code></strong> into a search field and you realize you've been building queries with raw string concatenation this entire time.</p><p>There's a better way. The ES|QL query builder for JavaScript and TypeScript lets you write queries like this instead:</p>import { ESQL, E, f } from '@elastic/elasticsearch-esql-dsl'

const query = ESQL.from('logs-*')
  .where(E('status_code').gte(minStatus))
  .where(E('host.name').eq(hostname))
  .where(E('@timestamp').gte(startDate))
  .stats({ error_count: f.count() })
  .by('status_code')
  .sort(E('error_count').desc())
  .limit(10)<p>Values are escaped automatically. You get autocomplete in your editor. And you can see exactly what the query does, without mentally parsing a template literal.</p><p>ES|QL query builders are already available across Elastic's language clients, including Python, Ruby, and others. This article focuses on the JavaScript and TypeScript version, walking through practical examples you can start using today.</p><h2>Getting started</h2><p>Install the package:</p>npm install @elastic/elasticsearch-esql-dsl<p>Here’s a minimal query:</p>import { ESQL, E } from '@elastic/elasticsearch-esql-dsl'

const query = ESQL.from('employees')
  .where(E('still_hired').eq(true))
  .sort(E('last_name').asc())
  .limit(10)

console.log(query.render())<p>This renders:</p>FROM employees
| WHERE still_hired == true
| SORT last_name ASC
| LIMIT 10<p>To run it against Elasticsearch:</p>import { Client } from '@elastic/elasticsearch'

const client = new Client({ node: 'http://localhost:9200' })
const response = await client.esql.query({ query: query.render() })<p>That’s it. No string interpolation, no manual escaping.</p><h2><strong>Building a real query, step by step</strong></h2><p>Let's walk through a realistic scenario: You're building a dashboard that analyzes web server error logs. We'll start simple and layer on features.</p><h3><strong>Step 1: Filter error logs</strong></h3>import { ESQL, E } from '@elastic/elasticsearch-esql-dsl'

const errors = ESQL.from('logs-*')
  .where(E('status_code').gte(400))
  .limit(100)FROM logs-*
| WHERE status_code &gt;= 400
| LIMIT 100<h3><strong>Step 2: Add a computed column</strong></h3><p>Your timestamps are in milliseconds, but you want response time in seconds:</p>const errors = ESQL.from('logs-*')
  .where(E('status_code').gte(400))
  .eval({ response_secs: E('response_time_ms').div(1000) })
  .limit(100)FROM logs-*
| WHERE status_code &gt;= 400
| EVAL response_secs = response_time_ms / 1000
| LIMIT 100<h3><strong>Step 3: Aggregate errors by status code</strong></h3>import { f } from '@elastic/elasticsearch-esql-dsl'

const errorBreakdown = ESQL.from('logs-*')
  .where(E('status_code').gte(400))
  .stats({
    error_count: f.count(),
    avg_response: f.avg('response_time_ms'),
  })
  .by('status_code')
  .sort(E('error_count').desc())FROM logs-*
| WHERE status_code &gt;= 400
| STATS error_count = COUNT(*), avg_response = AVG(response_time_ms) BY status_code
| SORT error_count DESC<p>The <strong><code>f</code></strong> namespace gives you access to 150+ ES|QL function wrappers: aggregations, string functions, date functions, math, geo, and more. They all return chainable expressions, so you can use them anywhere you'd use <strong><code>E()</code></strong>.</p><h3><strong>Step 4: Use date functions for time-based analysis</strong></h3>const hourlyErrors = ESQL.from('logs-*')
  .where(E('status_code').gte(400))
  .eval({ hour: f.dateTrunc('@timestamp', '1 hour') })
  .stats({ error_count: f.count() })
  .by('hour')
  .sort(E('hour'))FROM logs-*
| WHERE status_code &gt;= 400
| EVAL hour = DATE_TRUNC(@timestamp, "1 hour")
| STATS error_count = COUNT(*) BY hour
| SORT hour<h3><strong>Step 5: Branch queries safely</strong></h3><p>Every method returns a new query object. The original is never mutated. This means you can build a base query and branch it for different views:</p>const base = ESQL.from('logs-*')
  .where(E('status_code').gte(400))
  .where(E('@timestamp').gte('2026-01-01T00:00:00Z'))

const byStatus = base
  .stats({ count: f.count() })
  .by('status_code')
  .sort(E('count').desc())

const byHost = base
  .stats({ count: f.count() })
  .by('host.name')
  .sort(E('count').desc())
  .limit(20)

const recent = base
  .sort(E('@timestamp').desc())
  .keep('@timestamp', 'status_code', 'url.path', 'message')
  .limit(50)<p>Three different queries, one shared base. Change the filter on <strong><code>base</code></strong><strong>,</strong> and all three update. This is especially useful for dashboards where multiple panels query the same dataset with different aggregations.</p><h2><strong>Three ways to write expressions</strong></h2><p>The domain‑specific language (DSL) gives you flexibility in how you write conditions. Here's the same WHERE clause written three different ways:</p><p><strong>Raw strings:</strong> When you're writing a quick one-off:</p>.where('status_code &gt;= 400 AND host.name == "web-01"')<p><strong>The </strong><strong><code>E()</code></strong><strong> expression builder: </strong>When you want type safety and autocomplete:</p>import { and_ } from '@elastic/elasticsearch-esql-dsl'

.where(and_(
  E('status_code').gte(400),
  E('host.name').eq('web-01')
))<p><strong>The </strong><strong><code>esql</code></strong><strong> template tag: </strong>-When you want safe interpolation of dynamic values:</p>import { esql } from '@elastic/elasticsearch-esql-dsl'

const minStatus = 400
const host = 'web-01'
.where(esql`status_code &gt;= ${minStatus} AND host.name == ${host}`)<p>All three produce the same ES|QL. Pick whichever fits your situation: raw strings for simple cases, <strong><code>E()</code></strong> when building expressions programmatically, and the template tag when mixing literal ES|QL with dynamic values.</p><h2><strong>Keeping queries safe</strong></h2><p>If any part of your query comes from user input, you need to think about injection. ES|QL supports parameter binding, and the DSL makes it straightforward:</p>function searchLogs(userQuery: string) {
  const query = ESQL.from('logs-*')
    .where(E('message').eq(E('?')))
    .limit(100)

  return client.esql.query({
    query: query.render(),
    params: [userQuery],
  })
}<p>The <strong><code>?</code></strong> placeholder is replaced server-side by Elasticsearch, so the user's input never touches the query string. No escaping, no injection risk.</p><h2><strong>Beyond the basics</strong></h2><p>Once you're comfortable with the core commands, the DSL supports every advanced ES|QL feature:</p><p><strong>Hybrid search with FORK and FUSE:</strong></p>const results = ESQL.from('articles')
  .fork(
    ESQL.branch()
      .where(f.match('title', 'elasticsearch'))
      .sort(E('_score').desc())
      .limit(50),
    ESQL.branch()
      .where(f.knn('embedding', 10))
      .sort(E('_score').desc())
      .limit(50),
  )
  .fuse('RRF')
  .limit(10)<p><strong>Data enrichment:</strong></p>const enriched = ESQL.from('logs-*')
  .enrich('ip_lookup')
  .on('client.ip')
  .with('geo.city', 'geo.country')<p><strong>Conditional aggregation:</strong></p>const stats = ESQL.from('employees')
  .stats({
    eng_avg: f.avg('salary').where(E('dept').eq('Engineering')),
    sales_avg: f.avg('salary').where(E('dept').eq('Sales')),
    total: f.count(),
  })<p><strong>AI/machine learning (ML) integration:</strong></p>const summarized = ESQL.from('docs')
  .completion('Summarize this document')
  .with({ inferenceId: 'my-llm' })<p>For the full list of commands and functions, check out the <a href="https://www.elastic.co/docs/reference/elasticsearch/clients/javascript-dsl">ES|QL query builder documentation</a>.</p><h2><strong>What's next</strong></h2><p>This is the initial release of <strong><code>@elastic/elasticsearch-esql-dsl</code></strong>. You can find the package on <a href="https://www.npmjs.com/package/@elastic/elasticsearch-esql-dsl">npm</a>, explore the source on <a href="https://github.com/elastic/elasticsearch-dsl-js">GitHub</a>, and read the full documentation in the repository. If you run into issues or have feature requests, open an issue; we're actively developing this and want to build what JavaScript and TypeScript developers actually need.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-query-builder-javascript-typescript</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-query-builder-javascript-typescript</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Margaret Gu]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta47c29c9d8d76fbf/6a170f30b0367d006972bdd6/d8cc9dc5b2bcae4c589b402d62a5b7c8c6d63fb7-720x420.png" length="0" type="image/png"/>
    <pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Fast approximate Elasticsearch ES|QL - part II]]></title>
    <description><![CDATA[Explaining the approach we use to obtain fast approximate Elasticsearch ES|QL queries and the testing we did of error estimation.]]></description>
    <content:encoded><![CDATA[<p>As we discussed in our <a href="https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-1">previous blog</a>, we’re introducing fast approximate <a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql">ES|QL</a> <code>STATS</code> queries, which will be available in version 9.4 of Elasticsearch and the Elastic Stack. This feature allows users to estimate an expensive analytics query, often orders of magnitude faster than running the full query, by relaxing the constraint that it returns the exact value. We believe this has many uses; for example, we’re planning to integrate it into Kibana to obtain fast chart previews where possible.</p><p>In order for you to be able to trust our estimates, we provide error estimates. Furthermore, since there are edge cases in error estimation, we certify when the estimated value and error are trustworthy. In this blog post, we’ll dive into the theory for approximating and estimating the error in such queries, as well as discuss the testing we’ve done.</p><h3>Background</h3><p>In order to estimate ES|QL <code>STATS</code> queries efficiently, we make use of a property that’s shared by many statistics: Their estimates computed from a large number of independent samples from a dataset approach their true value. In the case of an index with some field  we can think of the true value of a statistic as its value computed for a random variable with uniform discrete distribution on . In the following we denote this quantity ; it can be things like <code>AVG</code>, <code>MEDIAN</code>, and so on. If we make  independent draws from , denoted , such that each value is selected with probability , we have  independent copies of this random variable. The property we rely on means that a sample statistic value  computed from  approaches  as  becomes large. For example, if  is the mean of some metric values then  as  becomes large. Indeed, for many statistics the limiting error distribution is known to be normal. Furthermore, it only depends on the distribution of , the size of the sample  and the type of the aggregation . This means supported <code>STATS</code> queries can be approximated with fixed accuracy independent of the index size .</p><p>It is easy to pick values at random from a Lucene index: create a filter that takes exponentially distributed jumps through the dataset, where the expected jump size is controlled by the desired sample probability. The AND of this filter and any other Lucene query can be performed extremely efficiently, since AND’ing filter queries is one of the things for which it is well optimized. In our other post, we discussed some real-world query examples to give a sense of the speedup we obtain for different levels of accuracy.</p><p>So far, we've only discussed obtaining an estimate of a query. While such a point estimator can be useful, without knowing anything about its error those uses are limited. We found that ES|QL has existing capabilities that make it relatively easy to incorporate cheap, flexible, and accurate error estimation at the same time. We'll discuss this next.</p><h3>Error estimates</h3><p>We view providing an accurate understanding of the uncertainty in our estimates as crucial for users to be able to trust the approximation. While having the option to quickly estimate an ES|QL query alone can be useful in certain situations, we wanted to provide a richer feature that allows clients to make intelligent choices. For example, if an approximate query is being used to preview a chart and the error is only a couple of pixels, there’s little point in running another expensive query to redraw it.</p><p>The way we've chosen to represent error is by a confidence interval: the -central confidence interval, to be precise. This can be expressed in terms of the <a href="https://en.wikipedia.org/wiki/Cumulative_distribution_function">cumulative density</a>, , of the statistic being estimated. Specifically, it's the interval which contains the true value of the statistic with probability  whose endpoints are  and . Confidence interval calculations are surprisingly subtle. There are also important constraints for our use case that make standard approaches undesirable. Next, we’ll take a look in more detail at the motivation and the design for the approach we’ve adopted.</p><p>A key requirement of the whole project is to dramatically accelerate expensive analytics queries. It’s therefore vital that the overhead of estimating uncertainty isn’t too large compared to estimating the query result itself. We also want the feature to be as general as possible, but “isolated” within the language. In other words, ES|QL is a flexible language, and we want estimation to work with as much of it as possible. At the same time, we don’t want to introduce a cross-cutting feature that incurs development costs on every new feature we ship.</p><p>With these considerations in mind, we chose to estimate confidence intervals by partitioning the sample set and computing the query output on each subsample. This is reminiscent of bootstrap; however, since we ensure that each partition receives a disjoint random subset of the sample data, we know that they comprise true estimates of the statistic distribution. To achieve the best possible estimate of the statistic itself, we still compute its value on the full sample. For example, to estimate the mean and its distribution the process can be expressed as follows:</p><p>This introduces a complication to account for the discrepancy between the count of values used to estimate a query statistic and used to sample its distribution. This is a downside; however, there are some significant advantages.</p><p>Most of the work in analytic queries resides in computing the aggregate statistics: post-processing after a <code>STATS</code> reduction acts on a far smaller table, and the cost is often relatively small. In this scheme, every row in the input data to the <code>STATS</code> command is processed exactly twice compared to just estimating the statistic. Therefore, roughly speaking we pay a fixed overhead that's the same order of magnitude as the cost of estimating the query in order to estimate its uncertainty. Since we often achieve multiple orders of magnitude speedup on the exact query, this is acceptable.</p><p>Because this process uses a plain old table, with extra columns for the distribution samples, we can pass the whole table through any ES|QL pipeline and compute confidence intervals on the final results. For example, if we include <code>EVAL square_avg = avg * avg</code> in the pipeline above, we'd have exactly the same <code>square_avg</code>, <code>square_avg_0</code>, …, <code>square_avg_B-1</code> extra values. At the end of the pipeline, we have samples from the distribution of the original statistics and all quantities that are computed using them. Therefore, we can apply our standard confidence interval machinery to reduce the table and convert samples into confidence intervals for derived quantities as well. This whole process is essentially transparent to the rest of the ES|QL language, and as we showed above, can be achieved by query rewriting.</p><h3>The confidence interval calculation</h3><p>We have independent samples of the statistic distribution . However, they're computed with fewer values than our estimate . We also have a relatively small number of distribution samples, to avoid the count discrepancy being too large, and so we don’t inflate the table too much. We therefore prefer a parametric approach for estimating confidence intervals.</p><p>The errors in the statistics for which we support estimation tend to normal distributions in the limit they're computed from many values. So a natural choice, the standard interval, is to estimate the mean and standard deviation from the samples and report the corresponding normal confidence intervals . Here,  denotes the standard normal distribution function. For heavy-tailed data and statistical functions that are sensitive to outliers, such as <code>STD_DEV</code>, convergence to normality can be slow, resulting in poorly calibrated intervals.</p><p>Briefly, in order to assess the quality of the intervals, one can examine their calibration. Specifically, one computes a quantity called the <a href="https://en.wikipedia.org/wiki/Coverage_probability">coverage</a>. For a central confidence interval, it should contain the true statistic value roughly  times for  trials. In fact, since we seek the central confidence interval, we can make the stronger statement that the true value should be above, or below, the confidence interval endpoints in roughly  out  trials. The empirical coverage is this fraction computed for a large number of trials. It allows us to compare alternative approaches by simulation. We return to this when we report our test results.</p><p>In order to obtain better confidence intervals, we tried a couple of different approaches: the <a href="https://en.wikipedia.org/wiki/Cornish%E2%80%93Fisher_expansion">Cornish-Fisher</a> correction of quantiles and an adaptation of <a href="https://en.wikipedia.org/wiki/Bootstrapping_(statistics)#Deriving_confidence_intervals_from_the_bootstrap_distribution">bias-corrected accelerated</a> (BCa) confidence intervals. Simulation showed BCa provided more robust calibration across a range of confidences, so this is the approach we selected. The basic idea, which was introduced by Efron, is to assume that there exists a monotonic transformation of the underlying statistic  which, when applied to a distribution sample normalizes its distribution:</p><p>Here, ,  and  is the standard normal random variable. This is clearly a relaxation of the assumption that the statistic itself is normally distributed, which is used to derive the standard interval. In fact, this family includes many distributions, since  is only constrained to be monotonic. (You can think of  as a first-order Taylor expansion of the case that the variance is an arbitrary function of the true parameter value. This further relaxes the assumption that the normalizing transformation also stabilizes the variance.) The nice thing about this ansatz is that  never needs to be explicitly computed, and there exist standard approaches for estimating the parameters  and  from the distribution samples.</p><p>To handle  one simply arranges for the estimate to land at the median of transformed distribution. If we assume the cumulative distribution function in theta space is  then , where  is the estimated statistic value, and as before  is the standard normal distribution function. Typically,  is approximated by the empirical distribution function, computed indirectly by bootstrap. However, somewhat surprisingly, extensive simulation showed that we obtained better calibrated intervals using a normal approximation to our sample values, i.e.  with  and  their empirical mean and standard deviation, respectively.</p><p>To complete the procedure, one can rearrange (1) to derive  quantiles for  as follows:</p><p>where  is the standard normal z-score for quantile . Typically, one uses the inverse empirical cumulative density estimate of  to convert quantiles back to a confidence interval. However, because we have a mismatch between the count of values used to compute distribution samples and the query estimate, we need to do some sort of scaling. Exploring options by simulation, we again found it best to use a normal approximation, , where  is the number of distribution samples we use. This is just applying the usual scaling of variance by .</p><p>Efron showed that in the case  is distributed as , i.e. that it depends only on the true value , then the acceleration  can be estimated without any knowledge of . In particular, . By assumption, our statistics tend to normal distributions with mean . Since skew is translation and scale invariant, this gives that , i.e. one sixth of the skew of our distribution samples. One thing this glosses over is the dependence of skew, and therefore acceleration, on sample size. We know it tends to zero as the count increases. In fact, skew also asymptotes to zero as  and so we also adjust acceleration to be  to account for the count mismatch between the samples  and estimate .</p><p>Although we significantly improve the calibration of confidence intervals by using a better methodology, we still see issues in the case that the underlying distribution has very heavy tails for some of the supported <code>STATS</code> functions. Therefore, we introduce some additional guard rails we discuss next.</p><h3>Guard rails</h3><p>To avoid the user having to understand too much about edge cases, we provide additional safeguards that surface when we've been unable to confirm  that the distribution samples behave as we expect. This typically happens when the statistic isn’t computed from a sufficient number of values given the metric distribution. It's exacerbated by very skewed metric data and certain aggregation functions, such as the <code>STD_DEV</code>, which are sensitive to outliers.</p><p>We have some global constraints on the minimum count of values used to estimate a statistic for which we'll certify it. For example, if any bucket is empty, then we can’t rely on the distribution samples. This is because ES|QL allows mixing approximate statistics, which treat empty buckets differently. For example, consider the following query:</p><p>There is no self-contained way of correctly assigning a value to <code>mix</code> for empty buckets, since summing requires that we treat them as zero, in which case we bias our estimate of <code>avg</code>. Alternatively, ignoring empty buckets introduces bias in the <code>sum</code>. There is also a global minimum count of values for which we’ve verified our certification method is sufficiently reliable; this is 10.</p><p>We explored a variety of additional tests to certify the results. These were based on both tests of the underlying data distribution, specifically <a href="https://en.wikipedia.org/wiki/Heavy-tailed_distribution#Hill.27s_tail-index_estimator">Hill’s estimator</a>, as well as the statistic’s distribution properties. If the true distribution of the statistic is sufficiently normal, then our estimate and confidence interval calculation behaves as we expect: The interval is well calibrated and the interval width is representative of the actual error. Therefore, in the end, we chose to use a test based on the p-value for distribution samples’ <a href="https://en.wikipedia.org/wiki/Skewness">skewness</a> and <a href="https://en.wikipedia.org/wiki/Kurtosis">kurtosis</a> versus a normal distribution null hypothesis. To certify a result, we require that the two tail p-values are greater than 0.05 for both tests. As we show below, we found this test was well aligned to our actual needs: to distinguish results for which the estimate and its confidence interval are more and less reliable.</p><p>There's a simple trick we can use to boost the accuracy of the accuracy of the test: Create multiple independent distribution samples and use a vote. Given a test to certify results with a failure rate , the distribution of the count of  failures for  tests is  for the case the null hypothesis, that the estimate is trustworthy, is true. For example, for the majority vote assuming  and  then the significance of the test is , i.e. we fail to certify fewer than 1% of trustworthy results. Note that we can compute multiple trials relatively easily using different seeds for the <code>RANDOM</code> bucket identifier.</p><p>This additional check allows us to certify that we trust our estimates and their errors. We surface this information in the approximate query results. When we can’t certify results, they won’t necessarily be inaccurate, but they should be treated with more caution.</p><h3>Testing</h3><p>The two main aims of the testing we discuss here were to understand the calibration of the confidence intervals and to see how well they characterize the statistics' estimation errors. The count function is particularly well behaved, its error distribution is binomial, so the majority of our testing focused on metric aggregations. We study smooth distributions but make sure we cover a range of tail behaviors. The presence of outliers is the key factor that reduces the accuracy of estimated statistics. For example, if an outlier isn’t sampled at all, it can significantly affect the value of some statistics.</p><p>We explored a range of light-tailed distributions, such as uniform and normal, and skewed and heavy-tailed distributions, such as exponential, log-normal, Cauchy, and Pareto. For each family of distribution, we used multiple parameterizations, focusing primarily on varying the scale parameter. In total, we had 24 distinct data distributions. Figure 1 shows some example sample distributions from this set. Note that we’ve truncated the charts to remove extreme outliers, which are present for both the Cauchy and log-normal distributions.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6e7b7df23a7e1faa/6a170ca0c1e8a505d7f88312/fb088c17f9755c0d1b3173fb917f0af2c0f83847-1712x950.png" alt="" /><p>For each data distribution, we evaluated 14 different sample sizes, ranging from 1000 to 500000. Then, for each sample set, we evaluated <code>AVG</code>, <code>COUNT</code>, <code>MEDIAN_ABSOLUTE_DEVIATION</code>, <code>MEDIAN</code>, <code>PERCENTILE([25, 75, 90, 95, 99])</code>, <code>SUM</code> and <code>STD_DEV</code> at two levels of confidence, 50% and 90%. In total, we have around 7500 distinct experiments. For each experiment, we assessed the interval calibration using 100 runs and counting the number of times the true statistic lands in the confidence interval. This gives us a binomially distributed estimate for the true confidence interval coverage. The variation we expect in the estimated coverage changes slightly with the level of confidence; for example, at 50% we expect to see values mainly between 0.44 and 0.56, and for 90% we expect to see values mainly between 0.86 and 0.94 using 100 trials.</p><p>Figure 2 shows <a href="https://en.wikipedia.org/wiki/Box_plot">box plots</a> for the empirical coverage for the two confidence levels computed from all experiments. In all cases, the confidence intervals are reasonably well calibrated. Extreme percentiles are biased for small sample sizes, which leads to increased outlier counts for small sample sizes. As a rule of thumb, you’d want roughly  samples to ensure that you have enough samples in the appropriate tail.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte33fb7bc54844a83/6a170ca2839dfa7a19dcff38/02a5375025e811ba18c4e823e1d984261bbf6f42-631x763.png" alt="" /><p>Next, we examine the degree to which the confidence intervals capture the typical size of the estimate error. To do this, we examine the distribution of the ratio of the estimated statistics' error and half the confidence interval width for all certified results. The higher the confidence, the wider the interval, so different confidence levels shift the mean of this distribution. Figure 3 shows this distribution computed for the 90% confidence interval. As expected, the distribution is roughly normal, albeit with a tail of some larger errors. We see in all cases the confidence interval width gives the order of magnitude of the estimated statistics' actual errors.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3d4bc25885b66945/6a170ca360084b472d3c45b4/2d4ab88a07910edac7e8406ae4942694751f0090-1000x600.png" alt="" /><p>We’ve shown that certified results are nearly always reliable; however, we’d also like some insight into the proportion of results which we fail to certify that are actually reliable, to confirm that the test aligns with our objective. We use <em>reliable</em> here in the fairly strong sense that the confidence interval is well calibrated. Specifically, for the 50% and 90% confidence intervals, we count the proportion of uncertified results for which the confidence interval empirical calibration has an acceptable margin of error, given the number of trials used to estimate it. Using this procedure, the false positive rate across all experiments is around 1%. This agrees well with the failure rate we expect by chance, given our test parameters, and confirms the assumption underlying the test.</p><p>Finally, to better understand the difference between certified and uncertified results, Figure 4 shows the error distribution of the ratio of the estimated statistics' errors and half the 90% confidence interval for the reliable and unreliable results separately. Note that we truncated the range for uncertified intervals.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0c4f729eb4c1127c/6a170ca5961e69c63cc4cf66/c0cd6d42f5d061aac15767539209a7c443ed1acd-1000x600.png" alt="" /><h3>Wrapping up</h3><p>In this post, we present the background behind our approach for quickly estimating ES|QL queries and providing an indication of their errors. To do this, we developed an effective confidence interval mechanism that allows us to provide error estimates. Our approach also allows us to estimate confidence intervals for quantities derived from sampled statistics via other pipeline operations. Quantifying the error comes with a relatively small overhead compared to just estimating the query. Finally, we developed a statistical test to certify results we return. Values that aren’t certified can still be accurate, but we’re less confident in them.</p><p>As well as testing the feature on a range of real-world use cases, which we discuss in <a href="https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-1">our companion post</a>, we tested the error estimation by extensive simulation across a range of data characteristics, sample sizes, aggregation functions, and confidence levels. This showed confidence intervals are well calibrated, and the interval itself provides a good approximation of the actual error we observe in the estimates. Finally, we showed that we were able to certify intervals with a low false negative rate.</p><p>We’re planning to integrate this feature into other stack capabilities in the future, so stay tuned.

</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-2</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-2</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Thomas Veasey,Jan Kuipers]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0c4f729eb4c1127c/6a170ca5961e69c63cc4cf66/c0cd6d42f5d061aac15767539209a7c443ed1acd-1000x600.png" length="0" type="image/png"/>
    <pubDate>Fri, 17 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Fast approximate Elasticsearch ES|QL - part I]]></title>
    <description><![CDATA[Introducing the work we've done on a fast approximate querying mode for Elasticsearch ES|QL. In many cases, it allows us to achieve orders of magnitude latency reductions while providing accurate estimates.]]></description>
    <content:encoded><![CDATA[<p>Analytics workloads typically involve summarizing large volumes of data into a much smaller number of statistics. The Elasticsearch Query Language (ES|QL) implements this capability using the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/stats-by">STATS command</a>. This allows you to select various aggregation functions and apply them to the previous query results, as well as grouping the results by one or more ES|QL expressions. This is a flexible operation that, coupled with ES|QL querying capabilities, allows one to perform <a href="https://en.wikipedia.org/wiki/MapReduce">MapReduce</a> on data stored in collections of Elasticsearch indices.</p><p>One of the key requirements for a pleasant user experience is that these operations are performed quickly. Large language model–based (LLM) agents also introduce new <a href="https://arxiv.org/pdf/2509.00997">higher bandwidth and speculative query patterns</a> that can potentially benefit from different optimization strategies.</p><p>In this two-part blog series, we discuss an optimization approach we’re introducing to ES|QL in version 9.4 of Elasticsearch and the Elastic Stack, which exploits a relaxation of the problem. Rather than trying to get exact values for aggregates, we allow ourselves to return approximate values, together with some characterization of their error. A key benefit of approximation is that it breaks the dependency between performance and dataset size: The accuracy with which one can approximate a query doesn’t depend on the original dataset size but, principally, its data characteristics and the query itself. As we’ll see later, this allows us to achieve some dramatic performance improvements.</p><p>In our<a href="https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-2"> next blog post</a>, we will discuss the theory behind our approach and the validation we’ve done of its statistical properties. Here, we introduce the syntax and give a sense of how it’s achieved using standard ES|QL and query rewriting. You can explore its performance on a subset of the popular <a href="https://github.com/ClickHouse/ClickBench">ClickBench</a> benchmark. Finally, we discuss some limitations and gotchas that are worth understanding when you use query approximation.</p><h3>Syntax and behavior</h3><p>So how do you actually use it?</p><p>That’s it. You simply introduce the new line <code>SET approximation=true;</code> and write your <code>STATS</code> query pipeline as usual. Below, we discuss some advanced configuration options and some limitations around the <code>agg(...)</code> and <code>commands</code>. However, essentially, we choose defaults so that this will typically provide useful approximations while achieving significant speedups.</p><p>With this change, you’ll see some differences in the query results. Let’s look at a concrete example to illustrate this. Suppose the raw query is as follows:</p><p>The results might look something like this:</p>item_category        | count
---------------------+------
Household Essentials | 5165
Kitchen              | 2132
Storage              | 1121
Home Decor           | 877
Furniture            | 357<p>Approximating this query introduces some extra columns for each quantity that’s estimated:</p>item_category | count | _approximation_confidence_interval(count) | _approximation_certified(count)
--------------+-------+-------------------------------------------+--------------------------------
Essentials    | 5150  | [5100, 5250]                              | true
Kitchen       | 2150  | [2100, 2200]                              | true
Storage       | 1120  | [1100, 1150]                              | true
Home Decor    | 880   | [860, 900]                                | true
Furniture     | 330   | [310, 350]                                | true<p>The count column now contains an estimate, and you’ll see it’s somewhat different from the exact values above. The <code>_approximation_confidence_interval(count)</code> column defaults to the central 90% confidence interval for the <code>count</code> estimate and the <code>_approximation_certified(count)</code> column indicates if we’re highly confident that the results and their confidence interval are trustworthy. In outline, the <em>confidence interval</em> is an interval we expect has a high probability (0.9) of containing the true value for the quantity being estimated. The <em>certified column</em> indicates the distribution of the approximation is behaving as we expect. When the result isn’t certified, it’s often still accurate, but our test of the properties of its distribution hasn’t been able to confirm this. These quantities are discussed in more detail in our second post.</p><h3>Implementation</h3><p>An approximate query is rewritten before query execution using random sampling and extrapolation. Let’s take a look at the query of the previous section. The part of the rewritten query responsible for obtaining the best estimate looks like:</p><p>The query samples a fraction of the data, and therefore the final count has to be extrapolated by scaling up with the inverse of the sample probability. Extrapolation clearly depends on the underlying aggregation function, and we handle this appropriately for all functions we support.</p><p>To obtain the sample probability, we're setting a fixed <code>number_of_rows</code> to be processed by the <code>STATS</code> command. In this case, the probability is calculated as follows:</p><p>This query is executed before the final approximate query is executed.</p><p>As well as this best estimate, confidence intervals and a statistical test used to certify that the value distribution is behaving as we expect also need to be computed. The intervals are computed using a variant of the <a href="https://blogs.sas.com/content/iml/2017/07/12/bootstrap-bca-interval.html">bias-corrected and accelerated bootstrap confidence interval</a> (BCa) method. Therefore, the data needs to be partitioned into B buckets, which are used in turn to compute the intervals. Omitting some implementation details, this approximate query looks like:</p><p>To certify the estimate and confidence interval, there should be enough data, and the distribution of the bucket values should tend to normality.</p><p>Some queries can be efficiently computed using only summary statistics maintained in the index. To handle these correctly, where sampling is both slower and inaccurate, we updated the physical query planner, since detecting this case requires information that’s only available where the data resides. When the planner detects this is possible, it simply executes the query as normal. Such queries are typically fast anyway, and there’s no real side effect, so you don’t need to worry about this when using approximation; however, you’ll see that confidence intervals for such queries always have zero length, indicating the results are exact.</p><h3>Results</h3><p>To explore the performance improvements, we use <a href="https://github.com/ClickHouse/ClickBench">ClickBench</a>. This is a benchmark for analytics workloads for database management systems (DBMS). It comprises approximately 100 million rows, with a focus on clickstream and traffic analysis, web analytics, machine-generated data, structured logs, and events data. The benchmark also defines 43 queries that are typical of ad-hoc analytics and real-time dashboards.</p><p>Some of the queries aren’t suitable for approximation. For example, we don’t support approximating the unique count of a categorical value or computing the minimum and maximum of a metric value. We also don’t care about queries targeting search alone, for which Elasticsearch has excellent performance in any case. We therefore exclude these types of query from our evaluation. Finally, we also want to test a few additional aggregation functions, such as percentiles, which are not well represented in the original query set, so add some variants of the original metric queries to this end.</p><p>Queries in the benchmark are written using standard SQL and so need porting to use ES|QL syntax. This translation is fairly straightforward. Here’s an example:</p>SELECT SUM(AdvEngineID), COUNT(*), AVG(ResolutionWidth) FROM hits<p>becomes:</p><p>when rewritten in ES|QL.</p><p>For running all benchmarks, we use an Elastic Cloud Hosted instance with 870GB disk, 29GB Ram, and 4 vCPUs, in effect, an Amazon Elastic Compute Cloud (EC2) i3.xlarge instance. In the following results, we simply compare ES|QL with and without query approximation. Extensive results on a range of different hardware setups and datastores can be found <a href="https://benchmark.clickhouse.com/">here</a>. Even with significantly constrained test hardware (matching the vCPUs of the smallest setup), our approximation approach achieves competitive results against much larger systems.</p><p>We run each query and its approximation five times in a random order, clearing the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-clear-cache">query cache</a> between each run. We report the average run time over all five runs. While clearing the cache should be sufficient to avoid most of the advantage of running second, we wanted to avoid any possible accidental prewarming effects, which is why we alternate.</p><p>The results break down into four categories:</p><ol><li><p>Queries which are rewritten to use index summary statistics (three queries).</p></li><li><p>Queries that perform well (13 queries).</p></li><li><p>Queries with high cardinality partitioning (seven queries).</p></li><li><p>Queries with restrictive filters (12 queries).</p></li></ol><p>Roughly speaking, for these four categories, approximate querying is: equivalent (1); faster and accurate (2); faster but unreliable (3); and slightly slower (4), compared to exact querying, respectively.</p><p>For category 1, the planner automatically detects that we’re able to perform the query using summary statistics, and we end up executing the queries in the same way. To do this, we need information that’s only available on the data nodes, so we perform the rewrite only after we've estimated the sample probability. Because we're able to do this very efficiently, the overhead is small (around 10–15%). In both cases, the results are exact.</p><p>Queries in category 2 run on average 23 faster if estimating the values and computing confidence intervals and 72 faster if just estimating the values, which you can select as follows: <code>SET approximation={"confidence_level":null}</code>. These headline figures hide quite some variation in the impact of approximation on performance. The table below shows some queries sampled from the range of speedups we see:</p><p>Query</p><p>Baseline / ms</p><p>Approximate with CI / ms</p><p>Approximate without CI / ms</p><p>3</p><p>1725</p><p>145</p><p>15</p><p>10</p><p>4340</p><p>1721</p><p>56</p><p>13</p><p>32912</p><p>6106</p><p>3821</p><p>21</p><p>46739</p><p>3284</p><p>2139</p><p>22</p><p>252505</p><p>6478</p><p>5019</p><p>Here are the corresponding queries:</p><p>We'll return to the accuracy of the approximation in the next blog post, but to give a sense of this, we plot below the exact and approximate values for a sample run for query 13:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5f6acbc3a2c9e58f/6a170ee34a531bbd2c36aa17/9ab83c13f42f88253a242d78339356f4a7c48700-2094x1358.png" alt="approximation-of-clickbench-query-13" /><p>For category 3, we get an average speedup of . However, the results of queries in this category can miss some partitions and often have large estimation errors. Approximation can still be valuable for such queries, particularly in the context of agentic workflows, but requires larger sample sizes than out default if accuracy is important. As we discuss in the next section, we provide an API to explicitly control the sample size. If the source dataset is sufficiently large, this can be increased and approximation will still yield significant performance improvements. The table below shows a couple of query examples for this category:</p><p>Query</p><p>Baseline / ms</p><p>Approximate with CI / ms</p><p>Approximate without CI / ms</p><p>15</p><p>8256</p><p>1187</p><p>124</p><p>17</p><p>70641</p><p>2109</p><p>982</p><p>Here are the corresponding queries:</p><p>Finally, category 4 queries use selective filters and end up being executed exactly, but they run slightly slower because of the work done in the query rewrite stage. Typically, all these queries run fast anyway, so the absolute slowdown is small. On average, they run approximately 14% or 370ms slower than the “without” sampling for our test setup.</p><h3>Limitations and best practices</h3><p>It’s worth explicitly mentioning some limitations. In particular, the following queries are not currently supported:</p><ol><li><p>Queries using the <code>TS</code> source command.</p></li><li><p>Queries using the <code>FORK</code> or <code>JOIN</code> processing command.</p></li><li><p>Pipelines which use two or more <code>STATS</code> commands.</p></li><li><p>The <code>ABSENT</code>, <code>PRESENT</code>, <code>DISTINCT_COUNT</code>, <code>MIN</code>, <code>MAX</code>, <code>TOP</code>, <code>ST_CENTROID_AGG</code> and <code>ST_EXTENT_AGG</code> aggregation functions.</p></li></ol><p>We plan to lift some of these restrictions in future releases, such as approximating queries using <code>TS</code>, <code>FORK</code> and <code>JOIN</code>; however, some are intrinsic. For example, while there’s prior art for estimating the <a href="https://en.wikipedia.org/wiki/Generalized_extreme_value_distribution">minimum and maximum</a> of a metric dataset or the count of unique values of a categorical dataset (see, for example, <a href="https://arxiv.org/pdf/2202.02800">this</a> paper), they require making certain distributional assumptions, either explicitly or implicitly. In summary, we view trying to automatically provide estimates of these statistics as being too open to accidental misuse.</p><p>For the expert user, we provide another route: ES|QL supports using the <code>SAMPLE</code> command directly. This allows one to obtain “point estimates” of any query, albeit with no attempt to correct for the impact of sampling or quantify error. For example:</p><p>computes the unique count of the value field on a sample of roughly 1/100th of the dataset. The sample probability can be adjusted to get a sense of how this is asymptoting, or more sophisticated estimation procedures can use <code>STATS COUNT() BY value</code> to estimate the frequency profile of the data.</p><p>There are a couple of cases that are more problematic for sampling. If a very restrictive filter is applied in the query, then sampling is of little value, since few rows match anyway. In such cases, we discover that we’d have to sample too large a proportion of the rows to estimate the query in the rewrite phase. In this case, we revert to running the query without sampling and its result is exact. However, the search procedure to determine the fraction of rows to sample comes with some overhead. One therefore pays a penalty, albeit less than the original query cost, for no benefit. If you know in advance that the query is expected to match relatively few rows, it's best to run it without approximation.</p><p>The second case only applies when computing <code>STATS</code> partitioned by some expression. If the cardinality of this expression is very high, then even if many rows are searched, individual statistics may be computed from a small number of rows. Some cases are more problematic than others. Sorting by ascending count, that is, finding the rarest partitions, can be impossible to estimate in a single query if heavy hitters would require us to sample most of the dataset to find them. For this particular case, heavy hitting partitions can be estimated first and sometimes efficiently excluded by updating the query. In general, infrequent partitions may be lost in the sampling process, and their statistics' estimation errors can be high. It’s worth noting that we won’t attempt to estimate any statistic for which we have fewer than 10 samples, and we simply drop them from the result set. In the case of very high cardinality <code>BY</code> clause, for example, a field whose value is unique for every row, this means the query can return no results. If you find approximate query results are too inaccurate, you have the option to increase the sample size, which by default is 1,000,000 for <code>STATS</code>, which uses grouping and 100,000 otherwise. Currently, this needs to be done manually, and we provide the following API for this:</p><p>Occasionally, functions significantly alter the distribution characteristics of the quantities they act on. A contrived example is the following:</p><p>If the variation in the estimate <code>sl</code> is much larger than  we expect the distribution of <code>csl</code> to be mainly flat in the interval  with peaks near both endpoints. In this particular case, it’s not clear that the central confidence interval is a particularly useful concept, since the modes of the distribution lie outside almost all central confidence intervals. In any case, just observing the samples of <code>csl</code>, our standard confidence interval machinery won’t reliably characterize this distribution and it will underestimate the variability of <code>csl</code>. However, our statistical test should detect this problem, and the result won’t be certified.</p><p>Finally, we note that Elasticsearch implements some query optimization strategies that ideally <a href="https://github.com/elastic/elasticsearch/issues/138151">need to account for the fact that sampling is taking place</a>. These rewrite the query at the Lucene level and the preprocessing involved in this rewrite can be relatively expensive. Accelerating an expensive string matching operation by first building a suitable data structure makes sense if the query needs to process every row, but if it processes only a small fraction of them, the trade-off is different. This is something we plan to enhance in future.</p><h3>Conclusions</h3><p>In this blog post, we introduced a new form of query optimization we’re bringing to ES|QL that enables dramatically faster querying by relaxing the constraint that the results are exact. We found on ClickBench that we were able to accurately estimate query values and their confidence intervals up to 100 times faster and values alone up to 250 times faster than we can compute them exactly. Furthermore, we expect this advantage to grow as the dataset size increases, because the approximation accuracy is independent of the dataset size. This feature works with many features of the ES|QL language and is enabled by simply prepending <code>SET approximation=true;</code> to the query to estimate.</p><p>As well as providing a point estimate, we also estimate confidence intervals and indicate whether we think that the underlying assumptions used to compute these are satisfied. This allows us to certify the results if the results are reliable. We explain the theory behind this feature and discuss the testing of its accuracy in our <a href="https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-2">next post</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-1</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/fast-approximate-esql-part-1</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Jan Kuipers,Thomas Veasey]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltba968f10a7cac60c/6a170ee50c48571b5301ab34/17afc59be8a46957a341faec1f44c9cb0a221894-1918x1176.png" length="0" type="image/png"/>
    <pubDate>Thu, 16 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[New Elasticsearch ES|QL plugin for IntelliJ IDEA]]></title>
    <description><![CDATA[Build and run Elasticsearch ES|QL queries in your IDE with the new plugin for IntelliJ IDEA.]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> is Elasticsearch’s piped query language, designed for intuitive data querying and manipulation. Refer to our <a href="https://www.elastic.co/blog/getting-started-elasticsearch-query-language">getting started guide</a> to learn more.</p><p>The Elasticsearch Java client <a href="https://www.elastic.co/search-labs/blog/esql-queries-to-java-objects">supports ES|QL queries</a> through the DSL, but currently it treats queries as simple strings, with no dedicated helper; and while <a href="https://www.elastic.co/kibana">Kibana</a> offers an excellent <a href="https://www.elastic.co/docs/explore-analyze/query-filter/languages/esql-kibana">UI to build ES|QL queries</a>, we’re aware that sometimes having everything needed to write applications in the integrated development environment (IDE) offers a better experience. So, until the Java client extends its type support to ES|QL, we wrote an Intellij IDEA plugin that autocompletes, syntax checks, shows documentation, and executes ES|QL queries.</p><p>The plugin currently supports Java, Kotlin, and plain text files, in case the Java Virtual Machine (JVM) isn’t your thing.</p><p>Check it out in the <a href="https://plugins.jetbrains.com/plugin/28898-elasticsearch-es-ql">JetBrains Marketplace page</a> and in the <a href="https://github.com/elastic/esql-idea-plugin">GitHub repository</a>, for more information.</p><h2>Prerequisites</h2><ul><li><p>IDE: Intellij IDEA version &gt;= 253 (community or ultimate)</p></li></ul><h2>Usage</h2><p>Install the plugin in Intellij IDEA like you would with every other plugin, so either from the <a href="https://plugins.jetbrains.com/plugin/28898-elasticsearch-es-ql">JetBrains marketplace</a> or by going to Settings -&gt; Plugins -&gt; Marketplace and searching “esql”.</p><p>The following examples are written using Java, but Kotlin is also supported and the usage is pretty much the same.</p><p>Create a text block string, write “ES|QL” in a simple comment above it, and you’re done.</p>// ES|QL
String query = """
""";<p>If you see the Elastic logo appearing on the left:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt987846e694123956/6a170bc8dc55decc5ce00e21/620d47b0c241271ab9bf727c37d3ab5f4137ca44-417x55.png" alt="Code editor showing an ES|QL comment and a string variable being initialized for a query, with the Elastic icon in the gutter." /><p>then everything is working, and you’re ready to write your queries.</p><p>Why text blocks and not simple strings? The ES|QL syntax accepts quotes in various contexts, and escaping them would trigger other errors in the syntax checker, so we decided on text blocks to keep things simple.</p><p>It’s even simpler for txt files, as you can just add the comment and start writing the query right below:</p><h3>Connecting to a server instance</h3><p>The plugin can be connected to an Elasticsearch server instance to fetch indices and field names, which will then be added to the autocompletion options. Look for the Elastic logo on the bottom left of of the screen (or wherever you keep your tools), and configure your connection to any server instance:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5069560e513f3ff7/6a170bcaa29299108cd0104e/9f5109542c921ad523458b9156551bf1fca7d41a-418x269.png" alt="Elasticsearch connection panel showing a “local” dropdown, status marked “Not connected,” and a Connect button in a dark-themed interface." /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf73cededd56ba9cc/6a170bcba6c2b93aabe79739/9fe37a53d97f67cbdd8623793bce768e8e2f9ced-577x298.png" alt="Dialog box for adding an Elasticsearch connection, showing fields for name, URL, API key, refresh rate, and buttons to test or confirm the connection." /><h3>Autocomplete</h3><p>Start typing while in the text block to automatically open the autocompletion popup, which will return a list of acceptable commands/values to continue writing the query correctly. If you want to manually trigger autocompletion, <code>ctlr+space</code> is the IDE’s shortcut to use:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt229d9a3bc33091f4/6a170bcc60084b84183c4590/987a927ab0e682bb1f9d07c934dd4254a769db20-584x252.png" alt="Java editor showing an ES|QL query with an autocomplete menu listing keywords like WHERE, STATS, DISSECT, FORK, and KEEP." /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1359cdac9ed83fa4/6a170bce961e6982e1c4cf4e/5f8e379ea1cca345b38d0b7a0c2a873db6de624f-584x252.png" alt="Java editor showing an ES|QL query with an autocomplete panel listing field suggestions such as field, name, title, vector, and string." /><h3>Syntax check</h3><p>The plugin will highlight errors in queries, explaining what to fix:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc935a693e4fcf2eb/6a170bcf67045b724745c216/15e91eab6c00fb49cfa5c1c6e270beeba534afc3-812x252.png" alt="Java editor showing an ES|QL query with an invalid keyword after a pipe operator and a tooltip explaining the syntax error." /><h3>Documentation</h3><p>Hovering with the cursor over commands will display documentation describing what the command can be used for and its correct syntax:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4aca4151e89f1d1b/6a170bd167045b597945c21a/c78295103c2188a3615edb2e001acbaf17523656-1072x627.png" alt="ava editor showing an ES|QL query alongside a documentation panel explaining how the WHERE clause works, including syntax, parameters, and examples." /><h3>Running the query</h3><p>Once connected to a server instance, you can run queries by clicking on the green button beside the Elastic icon: The results will be displayed in the tool window:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt96849d8f88f75832/6a170bd38b73cb7e2118a066/bff17bee7269dfb6f30140c9cd78dc203990f21c-1130x441.png" alt="IDE window showing an ES|QL query in a text file and a results panel below it listing returned book records with columns like author, title, year, and ID." /><p>Or if you’re writing an application, you can use the Java client like so:</p>// ES|QL
String query = """
	FROM my-index
| SORT year DESC
| LIMIT 10
""";

try (ElasticsearchClient client = ElasticsearchClient.of(e -&gt; e
                .host(serverUrl)
                .apiKey(apiKey))) {

client.esql().query(QueryRequest.of(qr -&gt; qr.query(query)));

}<p>Check our previous <a href="https://www.elastic.co/search-labs/blog/esql-queries-to-java-objects">ES|QL Java Client article</a> for a complete example of mapping ES|QL results to Java objects.</p><h2>How does it work?</h2><p>There’s no AI involved; the plugin is based on the ES|QL <a href="https://www.antlr.org/">ANTLR</a> grammar for autocompletion and syntax check, and it uses the <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch docs</a> to show documentation.</p><h2>Conclusion</h2><p>The plugin is still experimental, so feel free to report any bug or feature request on the <a href="https://github.com/elastic/esql-idea-plugin">Github repository</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-plugin-intellij-idea</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-plugin-intellij-idea</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Java]]></category>
    <dc:creator><![CDATA[Laura Trotta]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt050f40a347c2ebb8/6a170bd414b2706a45e3c644/91366de35a1b66860ce0d126c8a83e5b25b678f0-1280x720.png" length="0" type="image/png"/>
    <pubDate>Mon, 13 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[LINQ to Elasticsearch ES|QL: Write C#, query Elasticsearch]]></title>
    <description><![CDATA[Exploring the new LINQ to Elasticsearch ES|QL provider in the Elasticsearch .NET client, which allows you to write C# code that’s automatically translated to ES|QL queries.]]></description>
    <content:encoded><![CDATA[<p>Starting with <strong>v9.3.4</strong> and <strong>v8.19.18</strong>, the Elasticsearch .NET client includes a <a href="https://learn.microsoft.com/en-us/dotnet/csharp/linq/">Language Integrated Query (LINQ) </a>provider that translates C# LINQ expressions into <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">Elasticsearch Query Language (ES|QL)</a> queries at runtime. Instead of writing ES|QL strings by hand, you compose queries using <code>Where</code>, <code>Select</code>, <code>OrderBy</code>, <code>GroupBy</code>, and other standard operators. The provider takes care of translation, parameterization, and result deserialization, including per-row streaming that keeps memory usage constant, regardless of result set size.</p><h2>Your first query</h2><p>Start by defining a plain old CLR object (POCO) that maps to your Elasticsearch index. Property names are resolved to ES|QL column names through standard <code>System.Text.Json</code> attributes, like <code>[JsonPropertyName]</code>, or through a configured <code>JsonNamingPolicy</code>. The same <a href="https://www.elastic.co/docs/reference/elasticsearch/clients/dotnet/source-serialization">source serialization</a> rules that apply across the rest of the client apply here as well.</p>using System.Text.Json.Serialization;

public class Product
{
    [JsonPropertyName("product_id")]
    public string Id { get; set; }

    public string Name { get; set; }

    public string Brand { get; set; }

    [JsonPropertyName("price_usd")]
    public double Price { get; set; }

    [JsonPropertyName("in_stock")]
    public bool InStock { get; set; }
}<p>With the type in place, a query looks like this:</p>var minPrice = 100.0;
var brand = "TechCorp";

await foreach (var product in client.Esql.QueryAsync&lt;Product&gt;(q =&gt; q
    .From("products")
    .Where(p =&gt; p.InStock &amp;&amp; p.Price &gt;= minPrice &amp;&amp; p.Brand == brand)
    .OrderByDescending(p =&gt; p.Price)
    .Take(10)))
{
    Console.WriteLine($"{product.Name}: ${product.Price}");
}<p>The provider translates this into the following ES|QL:</p><p>A few details to note:</p><ul><li><p><strong>Property name resolution:</strong> <code>p.Price</code> becomes <code>price_usd</code> because of the <code>[JsonPropertyName]</code> attribute, and <code>p.Brand</code> becomes <code>brand</code> following the default camelCase naming policy.</p></li><li><p><strong>Parameter capturing:</strong> The C# variables <code>minPrice</code> and <code>brand</code> are captured as named parameters (<code>?minPrice</code>, <code>?brand</code>). They’re sent separately from the query string in the JSON payload, which prevents injection and enables server-side query plan caching.</p></li><li><p><strong>Streaming:</strong> <code>QueryAsync&lt;T&gt;</code> returns <code>IAsyncEnumerable&lt;T&gt;</code>. Rows are materialized one at a time as they arrive from Elasticsearch.</p></li></ul><p>You can also inspect the generated query and its parameters without executing it:</p>var query = client.Esql.CreateQuery&lt;Product&gt;()
    .Where(p =&gt; p.InStock &amp;&amp; p.Price &gt;= minPrice &amp;&amp; p.Brand == brand)
    .OrderByDescending(p =&gt; p.Price)
    .Take(10);

Console.WriteLine(query.ToEsqlString());
// FROM products | WHERE (in_stock == true AND price_usd &gt;= 100) | SORT price_usd DESC | LIMIT 10

Console.WriteLine(query.ToEsqlString(inlineParameters: false));
// FROM products | WHERE (in_stock == true AND price_usd &gt;= ?minPrice AND brand == ?brand) | SORT price_usd DESC | LIMIT 10

var parameters = query.GetParameters();
// { "minPrice": 100.0, "brand": "TechCorp" }<h2>How does this work? A quick LINQ refresher</h2><p>The mechanism that makes LINQ providers possible is the distinction between <code>IEnumerable&lt;T&gt;</code> and <code>IQueryable&lt;T&gt;</code>.</p><p>When you call <code>.Where(p =&gt; p.Price &gt; 100)</code> on an <code>IEnumerable&lt;T&gt;</code>, the lambda compiles to a <code>Func&lt;Product, bool&gt;</code>, a regular delegate that the runtime executes in-process. This is LINQ-to-Objects.</p><p>When you call the same method on an <code>IQueryable&lt;T&gt;</code>, the C# compiler wraps the lambda in an <code>Expression&lt;Func&lt;Product, bool&gt;&gt;</code> instead. This is a data structure that represents the <em>structure</em> of the code rather than its executable form. The expression tree can be inspected, analyzed, and translated into another language at runtime.</p>// IEnumerable: the lambda is a compiled delegate
IEnumerable&lt;Product&gt; local = products.Where(p =&gt; p.Price &gt; 100);

// IQueryable: the lambda is an expression tree, a data structure
IQueryable&lt;Product&gt; remote = queryable.Where(p =&gt; p.Price &gt; 100);<p>The <code>IQueryProvider</code> interface is the extension point. Any provider can implement <code>CreateQuery&lt;T&gt;</code> and <code>Execute&lt;T&gt;</code> to translate these expression trees into a target language. Entity Framework uses this to emit SQL. The LINQ to ES|QL provider uses it to emit ES|QL.</p><p>The expression tree for the query above looks like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt521838e8b9c36649/6a1705b1839dfa5f40dcfdfe/f864cd18a390831f8d28503a29b5835efb1842f7-1000x720.png" alt="Expression tree for the example query." /><p><em>Expression tree for the example query.</em></p><p>The tree is nested inside out: <code>Take</code> wraps <code>OrderByDescending</code>, which wraps <code>Where</code>, which wraps <code>From</code>, which wraps the root <code>EsqlQueryable&lt;Product&gt;</code> constant. The <code>Where</code> predicate is itself a subtree of <code>BinaryExpression</code> nodes for the <code>&amp;&amp;</code>, <code>&gt;=</code>, and <code>==</code> operators, with <code>MemberExpression</code> leaves for property accesses and closure captures for the <code>minPrice</code> and <code>brand</code> variables. This is the data structure that the provider walks to produce the final ES|QL.</p><h2>Under the hood: The translation pipeline</h2><p>The path from a LINQ expression to query results follows a six-stage pipeline:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt930670a505dd61ea/6a1705b3b339d58a54769ecf/2a2c772b63d720f61fc9a28b2f85668fa2db8d38-1999x1036.png" alt="Translation pipeline overview." /><p><em>Translation pipeline overview.</em></p><h3>1. Expression tree capture</h3><p>When you chain <code>.Where()</code>, <code>.OrderBy()</code>, <code>.Take()</code> and other operators on an <code>IQueryable&lt;T&gt;</code>, the standard LINQ infrastructure builds an expression tree. <code>EsqlQueryable&lt;T&gt;</code> implements <code>IQueryable&lt;T&gt;</code> and delegates to <code>EsqlQueryProvider</code>.</p><h3>2. Translation</h3><p>When the query is executed (by enumerating, calling <code>ToList()</code>, or using <code>await foreach)</code>, the <code>EsqlExpressionVisitor</code> walks the expression tree inside out. It dispatches each LINQ method call to a specialized visitor:</p><p>Visitor</p><p>Translates</p><p>Into</p><p>WhereClauseVisitor</p><p>.Where(predicate)</p><p>WHERE condition</p><p>SelectProjectionVisitor</p><p>.Select(selector)</p><p>EVAL + KEEP + RENAME</p><p>GroupByVisitor</p><p>.GroupBy().Select()</p><p>STATS ... BY</p><p>OrderByVisitor</p><p>.OrderBy() / .ThenBy()</p><p>SORT field [ASC\|DESC]</p><p>EsqlFunctionTranslator</p><p>EsqlFunctions.*, Math.*, string methods</p><p>80+ ES|QL functions</p><p>During translation, C# variables referenced in expressions are captured as named parameters.</p><h3>3. Query model</h3><p>The visitors don’t produce strings directly. Instead, they produce <code>QueryCommand</code> objects, an immutable intermediate representation. A <code>FromCommand</code>, a <code>WhereCommand</code>, a <code>SortCommand</code>, and a <code>LimitCommand</code>, each representing one ES|QL processing command. These are collected into an <code>EsqlQuery</code> model.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt788c9936976f2f62/6a1705b50e2e4910da419ff0/2adc349b6cf655b96b7b3e826a134e8a17fe42fd-1999x1036.png" alt="Query model and command pattern." /><p><em>Query model and command pattern.</em></p><p>This intermediate model is decoupled from both the expression tree and the output format. It can be inspected, intercepted (via <code>IEsqlQueryInterceptor</code>), or modified before formatting.</p><h3>4. Formatting</h3><p><code>EsqlFormatter</code> visits each <code>QueryCommand</code> in order and produces the final ES|QL string. Each command becomes one line, separated by the pipe (|) operator that ES|QL uses to chain processing commands. Identifiers containing special characters are automatically escaped with backticks.</p><h3>5. Execution</h3><p>The formatted ES|QL string and captured parameters are sent to Elasticsearch’s <code>/_query</code> endpoint as a JSON payload. The <code>IEsqlQueryExecutor</code> interface abstracts the transport layer, which is where the layered package architecture comes into play.</p><h3>6. Materialization</h3><p><code>EsqlResponseReader</code> streams the JSON response without buffering the entire result set into memory. A <code>ColumnLayout</code> tree, precomputed once per query, maps flat ES|QL column names (like <code>address.street</code>, <code>address.city</code>) to nested POCO properties. Each row is assembled into a <code>T</code> instance and yielded one at a time via <code>IEnumerable&lt;T&gt;</code> or <code>IAsyncEnumerable&lt;T&gt;</code>.</p><h2>The layered architecture</h2><p>The LINQ to ES|QL functionality is split across three packages:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt662bd0dd8861b6b6/6a1705b7a929cf7086ae08a2/41b8aae860ecdc2480edcb1c1d4cc9b03cfb78c9-1999x1036.png" alt="Package architecture." /><p><em>Package architecture.</em>
<a href="https://www.nuget.org/packages/Elastic.Esql"><strong><code>Elastic.Esql</code></strong></a> is the pure translation engine. It has zero HTTP dependencies and contains the expression visitors, query model, formatter, and response reader. You can use it stand alone to build and inspect ES|QL queries without an Elasticsearch connection, which is useful for testing, query logging, or building your own execution layer.</p>// Translation-only: no Elasticsearch connection needed
var provider = new EsqlQueryProvider();
var query = new EsqlQueryable&lt;Product&gt;(provider)
    .From("products")
    .Where(p =&gt; p.InStock)
    .OrderByDescending(p =&gt; p.Price);

Console.WriteLine(query.ToEsqlString());
// FROM products | WHERE in_stock == true | SORT price_usd DESC<p><a href="https://www.nuget.org/packages/Elastic.Clients.Esql"><strong><code>Elastic.Clients.Esql</code></strong></a> is a lightweight stand-alone ES|QL client. It adds HTTP execution on top of <code>Elastic.Esql</code> via <code>Elastic.Transport</code>. If your application only needs ES|QL and none of the other Elasticsearch APIs, this is the minimal dependency option.</p><p><a href="https://www.nuget.org/packages/Elastic.Clients.Elasticsearch"><strong><code>Elastic.Clients.Elasticsearch</code></strong></a> is the full Elasticsearch .NET client. It also builds on <code>Elastic.Esql</code> and exposes the LINQ provider through the <code>client.Esql</code> namespace. This is the recommended entry point for most applications.</p><p>Both execution-layer packages provide their own implementation of <code>IEsqlQueryExecutor</code>, the strategy interface that bridges translation and transport.</p><p>All three packages are compatible with Native AOT when used with a source-generated <code>JsonSerializerContext</code>. For the full client, see the <a href="https://www.elastic.co/docs/reference/elasticsearch/clients/dotnet/source-serialization#native-aot">Native AOT documentation</a>.</p><h2>Beyond the basics</h2><p>The example above covered filtering, sorting, and pagination. The provider supports a broader set of operations.</p><h3>Aggregations</h3><p><code>GroupBy</code>, combined with aggregate functions in <code>Select</code>, translates to ES|QL <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/stats-by"><code>STATS ... BY</code></a>:</p>var stats = client.Esql.Query&lt;Product, object&gt;(q =&gt; q
    .GroupBy(p =&gt; p.Brand)
    .Select(g =&gt; new
    {
        Brand = g.Key,
        Count = g.Count(),
        AvgPrice = g.Average(p =&gt; p.Price),
        MaxPrice = g.Max(p =&gt; p.Price)
    }));

// -&gt; FROM products | STATS COUNT(*), AVG(price_usd), MAX(price_usd) BY brand<h3>Projections</h3><p><code>Select</code>, with anonymous types generates <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/eval"><code>EVAL</code></a>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/keep"><code>KEEP</code></a>, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/rename"><code>RENAME</code></a> commands:</p>var query = client.Esql.CreateQuery&lt;Product&gt;()
    .Select(p =&gt; new { ProductName = p.Name, p.Price, p.InStock });

// -&gt; FROM products | KEEP name, price_usd, in_stock | RENAME name AS ProductName<h3>Rich function library</h3><p>Over 80 ES|QL functions are available through the <code>EsqlFunctions</code> class, covering date/time, string, math, IP, pattern matching, and scoring. Standard <code>Math.*</code> and <code>string.*</code> methods are also translated:</p>.Where(p =&gt; p.Name.Contains("Pro"))       // -&gt; WHERE name LIKE "*Pro*"
.Where(p =&gt; EsqlFunctions.CidrMatch(      // -&gt; WHERE CIDR_MATCH(ip, "10.0.0.0/8")
    p.IpAddress, "10.0.0.0/8"))<h3>LOOKUP JOIN</h3><p>Cross-index lookups translate to ES|QL <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join"><code>LOOKUP JOIN</code></a>:</p>var enriched = client.Esql.Query&lt;Product, object&gt;(q =&gt; q
    .LookupJoin&lt;Product, CategoryLookup, string, object&gt;(
        "category-lookup-index",
        product =&gt; product.Id,
        category =&gt; category.CategoryId,
        (product, category) =&gt; new { product.Name, category!.CategoryLabel }));<h3>Raw ES|QL escape hatch</h3><p>For ES|QL features not yet covered by the LINQ provider, you can append raw fragments:</p>var results = client.Esql.Query&lt;Product&gt;(q =&gt; q
    .Where(p =&gt; p.InStock)
    .RawEsql("| EVAL discounted = price_usd * 0.9"));<h3>Server-side async queries</h3><p>For long-running queries, submit them for background processing on the server:</p>await using var asyncQuery = await client.Esql.SubmitAsyncQueryAsync&lt;Product&gt;(
    q =&gt; q.Where(p =&gt; p.InStock),
    asyncQueryOptions: new EsqlAsyncQueryOptions
    {
        WaitForCompletionTimeout = TimeSpan.FromSeconds(5),
        KeepAlive = TimeSpan.FromMinutes(10)
    });

await asyncQuery.WaitForCompletionAsync();
await foreach (var product in asyncQuery.AsAsyncEnumerable())
    Console.WriteLine(product.Name);<p>Server-side async queries are especially useful for long-running analytical queries / large dataset processing that might exceed typical timeout thresholds, or in timeout-sensitive environments with load balancers, API gateways, or proxies that enforce strict HTTP timeouts. Async queries avoid connection drops by decoupling submission from result retrieval.</p><h2>Getting started</h2><p>LINQ to ES|QL is available starting from:</p><ul><li><p><strong>Elastic.Clients.Elasticsearch v9.3.4</strong> (9.x branch)</p></li><li><p><strong>Elastic.Clients.Elasticsearch v8.19.18</strong> (8.x branch)</p></li></ul><p>Install from NuGet:</p><p><code>dotnet add package Elastic.Clients.Elasticsearch</code></p><p>The entry points are on <code>client.Esql</code>:</p><p>Method</p><p>Returns</p><p>Use case</p><p>Query&lt;T&gt;(...)</p><p>IEnumerable&lt;T&gt;</p><p>Synchronous execution</p><p>QueryAsync&lt;T&gt;(...)</p><p>IAsyncEnumerable&lt;T&gt;</p><p>Async streaming</p><p>CreateQuery&lt;T&gt;()</p><p>IEsqlQueryable&lt;T&gt;</p><p>Advanced composition and inspection</p><p>SubmitAsyncQueryAsync&lt;T&gt;(...)</p><p>EsqlAsyncQuery&lt;T&gt;</p><p>Long-running server-side queries</p><p>For the full feature reference, including query options, multifield access, nested objects, and multivalue field handling, see the <a href="https://www.elastic.co/docs/reference/elasticsearch/clients/dotnet/linq-to-esql">LINQ to ES|QL documentation</a>.</p><h2>Conclusion</h2><p>LINQ to ES|QL brings the full expressiveness of C# LINQ to Elasticsearch's ES|QL query language, letting you write strongly typed, composable queries without handcrafting query strings. With automatic parameter capturing, streaming materialization, and a layered package architecture that scales from stand-alone translation to the full Elasticsearch client, it fits naturally into .NET applications of any size. Install the latest client, point your LINQ expressions at an index, and let the provider handle the rest.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/linq-esql-c-elasticsearch-net-client</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/linq-esql-c-elasticsearch-net-client</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Florian Bernd,Martijn Laarman]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdfa35fbcbbf4959f/6a1705b9dc55de19a4e00d07/e54132e915217063e9ed0ec45059c6cfc38e31dd-1280x720.png" length="0" type="image/png"/>
    <pubDate>Wed, 01 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Faster ES|QL stats with Swiss-style hash tables]]></title>
    <description><![CDATA[How Swiss-inspired hashing and SIMD-friendly design deliver consistent, measurable speedups in Elasticsearch Query Language (ES|QL).]]></description>
    <content:encoded><![CDATA[<p>We recently replaced key parts of Elasticsearch’s hash table implementation with a Swiss-style design and observed up to 2–3x faster build and iteration times on uniform, high-cardinality workloads. The result is lower latency, better throughput, and more predictable performance for Elasticsearch Query Language (ES|QL) stats and analytics operations.</p><h2>Why this matters</h2><p>Most typical analytical workflows eventually boil down to grouping data. Whether it’s computing average bytes per host, counting events per user, or aggregating metrics across dimensions, the core operation is the same — map keys to groups and update running aggregates.</p><p>At a small scale, almost any reasonable hash table works fine. At the large scale (hundreds of millions of documents and millions of distinct groups) details start to matter. Load factors, probing strategy, memory layout, and cache behavior can make the difference between linear performance and a wall of cache misses.</p><p>Elasticsearch has supported these workloads for years, but we’re always looking for opportunities to modernize core algorithms. As such, we evaluated a newer approach inspired by Swiss tables and applied it to how ES|QL computes statistics.</p><h2>What are Swiss tables, really?</h2><p>Swiss tables are a family of modern hash tables popularized by Google’s SwissTable and later adopted in Abseil and other libraries.</p><p>Traditional hash tables spend a lot of time chasing pointers or loading keys just to discover that they don’t match. Swiss tables’ defining feature is the ability to reject most probes using a tiny cache-resident array structure, stored separately from the keys and values, called <em>control bytes</em>, to dramatically reduce memory traffic.</p><p>Each control byte represents a single slot and, in our case, encodes two things: whether the slot is empty, and a short fingerprint derived from the hash. These control bytes are laid out contiguously in memory, typically in groups of 16, making them ideal for <a href="https://en.wikipedia.org/wiki/Single_instruction,_multiple_data">single instruction, multiple data</a> (SIMD) processing.</p><p>Instead of probing one slot at a time, Swiss tables scan an entire control-byte block using vector instructions. In a single operation, the CPU compares the fingerprint of the incoming key against 16 slots and filters out empty entries. Only the few candidates that survive this fast path require loading and comparing the actual keys.</p><p>This design trades a small amount of extra metadata for much better cache locality and far fewer random loads. As the table grows and probe chains lengthen, those properties become increasingly valuable.</p><h2>SIMD at the center</h2><p>The real star of the show is SIMD.</p><p>Control bytes are not just compact, they’re also explicitly designed to be processed with vector instructions. A single SIMD compare can check 16 fingerprints at once, turning what would normally be a loop into a handful of wide operations. For example:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1f710e87dd749ab3/6a170cc46234e052dadb1a49/bd418778f0c6144f8f5f18419f6220ac0c935c7a-903x407.png" alt="SIMD at the center in Elasticsearch" /><p>In practice, this means:</p><ul><li><p>Fewer branches.</p></li><li><p>Shorter probe chains.</p></li><li><p>Fewer loads from key and value memory.</p></li><li><p>Much better utilization of the CPU’s execution units.</p></li></ul><p>Most lookups never make it past the control-byte scan. When they do, the remaining work is focused and predictable. This is exactly the kind of workload that modern CPUs are good at.</p><h2>SIMD under the hood</h2><p>For readers who like to peek under the hood, here’s what happens when inserting a new key into the table. We use the Panama Vector API with 128-bit vectors, thus operating on 16 control bytes in parallel.</p><p>The following snippet shows the code generated on an Intel Rocket Lake with AVX-512. While the instructions reflect that environment, the design does not depend on AVX-512. The same high-level vector operations are emitted on other platforms using equivalent instructions (for example, AVX2, SSE, or NEON).</p>; Load 16 control bytes from the control block
vmovdqu xmm0, XMMWORD PTR [r9+r10*1+0x10]

; Broadcast the 7-bit fingerprint of the new key across the vector
vpbroadcastb xmm1, r11d

; Compare all 16 control bytes to the new fingerprint
vpcmpeqb k7, xmm0, xmm1
kmovq rbx, k7

; Check if any matches were found
test rbx, rbx
jne &lt;handle_match&gt;<p>Each instruction has a clear role in the insertion process:</p><ul><li><p><code>vmovdqu</code>: Loads 16 consecutive control bytes into the 128-bit <code>xmm0</code> register.</p></li><li><p><code>vpbroadcastb</code>: Replicates the 7-bit fingerprint of the new key across all lanes of the <code>xmm1</code> register.</p></li><li><p><code>vpcmpeqb</code>: Compares each control byte against the broadcasted fingerprint, producing a mask of potential matches.</p></li><li><p><code>kmovq</code> + <code>test</code>: Moves the mask to a general-purposes register and quickly checks whether a match exists.</p></li></ul><p>Finally, we settled on probing groups of 16 control bytes at a time, as benchmarking showed that expanding to 32 or 64 bytes with wider registers provided no measurable performance benefit.</p><h2>Integration in ES|QL</h2><p>Adopting Swiss-style hashing in Elasticsearch was not just a drop-in replacement. ES|QL has strong requirements around memory accounting, safety, and integration with the rest of the compute engine.</p><p>We integrated the new hash table tightly with Elasticsearch’s memory management, including the page recycler and circuit breaker accounting, ensuring that allocations remain visible and bounded. Elasticsearch's aggregations are stored densely and indexed by a group ID, keeping the memory layout compact and fast for iteration, as well as enabling certain performance optimizations by allowing random access.</p><p>For variable-length byte keys, we cache the full hash alongside the group ID. This avoids recomputing expensive hash codes during probing and improves cache locality by keeping related metadata close together. During rehashing, we can rely on the cached hash and control bytes without inspecting the values themselves, keeping resizing costs low.</p><p>One important simplification in our implementation is that entries are never deleted. This removes the need for <em>tombstones</em> (markers to identify previously occupied slots) and allows empty slots to remain truly empty, which further improves probe behavior and keeps control-byte scans efficient.</p><p>The result is a design that fits naturally into Elasticsearch’s execution model while preserving the performance characteristics that make Swiss tables attractive.</p><h2>How does it perform?</h2><p>At small cardinalities, Swiss tables perform roughly on par with the existing implementation. This is expected: When tables are small, cache effects dominate less and there is little probing to optimize.</p><p>As cardinality increases, the picture changes quickly.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt09b13af2fe162f59/6a170cc66f7f04485c9148b8/24900afc47ab07b0e9933f6117b99d0f4613f794-962x599.png" alt="ES|QL stats with Swiss-style hash tables" /><p>The heatmap above plots time improvement factors for different key sizes (8, 32, 64, and 128 bytes) across cardinalities from 1,000 up to 10,000,000 groups. As cardinality grows, the improvement factor steadily increases, reaching up to 2–3x for uniform distributions.</p><p>This trend is exactly what the design predicts. Higher cardinality leads to longer probe chains in traditional hash tables, while Swiss-style probing continues to resolve most lookups inside SIMD-friendly control-byte blocks.</p><h2>Cache behavior tells the story</h2><p>To better understand the speedups, we ran the same JMH <a href="https://github.com/elastic/elasticsearch/pull/139343/files#diff-d0e0cc91a7495bf36b2d44eacce95f5185d01879e5f6c38089ac7a89aad17da7"><code>benchmarks</code></a> under Linux <code>perf</code> and captured cache and TLB statistics.</p><p>Compared to the original implementation, the Swiss version performs about 60% fewer cache references overall. Last-level cache loads drop by more than 4x, and LLC load misses fall by over 6x. Since LLC misses often translate directly into main-memory accesses, this reduction alone explains a large portion of the end-to-end improvement.</p><p>Closer to the CPU, we see fewer L1 data cache misses and nearly 6x fewer data TLB misses, pointing to tighter spatial locality and more predictable memory access patterns.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb987a5bd98c0d7eb/6a170cc8a929cf9655ae0a25/6e49b7609fba83e33692cb9834552b6ca7e42a83-998x499.png" alt="Cache behavior: Original vs. ES|QL stats with Swiss-style hash tables" /><p>This is the practical payoff of SIMD-friendly control bytes. Instead of repeatedly loading keys and values from scattered memory locations, most probes are resolved by scanning a compact, cache-resident structure. Less memory touched means fewer misses, and fewer misses mean faster queries.</p><h2>Wrapping up</h2><p>By adopting a Swiss-style hash table design and leaning hard into SIMD-friendly probing, we achieved 2–3x speedups for high-cardinality ES|QL stats workloads, along with more stable and predictable performance.</p><p>This work highlights how modern CPU-aware data structures can unlock substantial gains, even for well-trodded problems, like hash tables. There is more room to explore here, like additional primitive type specializations and use in other high-cardinality paths, like joins, all of which are just part of the broader and ongoing effort to continually modernize Elasticsearch internals.</p><p>If you’re interested in the details or want to follow the work, check out this <a href="https://github.com/elastic/elasticsearch/pull/139343">pull request</a> and <a href="https://github.com/elastic/elasticsearch/issues/138799">meta issue</a> tracking progress on Github.</p><p>Happy hashing!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-swiss-hash-stats</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-swiss-hash-stats</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Chris Hegarty,Matthew Alp,Nik Everett]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf76dd688c5b737e6/6a170cc9839dfa7fc7dcff40/21036e031070f14faccb2b53b22723de2750c391-1280x720.png" length="0" type="image/png"/>
    <pubDate>Mon, 19 Jan 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Hybrid search and multistage retrieval in ES|QL]]></title>
    <description><![CDATA[Explore the multistage retrieval capabilities of ES|QL, using FORK and FUSE commands to integrate hybrid search with semantic reranking and native LLM completions.]]></description>
    <content:encoded><![CDATA[<p>In Elasticsearch 9.2, we’ve introduced the ability to do dense vector search and hybrid search in Elasticsearch Query Language (ES|QL). This continues our investment in making ES|QL the best search language to solve modern search use cases.</p><h2>Multistage retrieval: The challenge of modern search</h2><p>Modern search has evolved beyond simple keyword matching. Today's search applications need to understand intent, handle natural language, and combine multiple ranking signals to deliver the best results.</p><p>Retrieval of the most relevant results happens in multiple stages, with each stage gradually refining the result set. This wasn’t the case in the past, where most use cases would require one or two stages of retrieval: an initial query to get results and a potential rescoring phase.	</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3382265814939417/6a170df3a929cf1246ae0a61/fceada10b0c09d6a4a372f137bb3040e1ff41fbf-1600x895.png" alt="" /><p>We start with an initial retrieval, where we cast a wide net to gather results that are relevant to our query. Since we need to sieve through all the data, we should use techniques that return results fast, even when we index billions of documents.</p><p>We therefore employ trusted techniques, such as lexical search that Elasticsearch has supported and optimized since the beginning, or vector search, where Elasticsearch excels in speed and accuracy.</p><p>Lexical search using BM25 is quite fast and best at exact term matching or phrase matching, and <a href="https://www.elastic.co/docs/solutions/search/vector">vector</a> or <a href="https://www.elastic.co/docs/solutions/search/semantic-search">semantic search</a> is better suited for handling natural language queries. <a href="https://www.elastic.co/what-is/hybrid-search">Hybrid search</a> combines lexical and <a href="https://www.elastic.co/docs/solutions/search/vector">vector search</a> results to bring the best from both. The challenge that hybrid search solves is that vector and lexical search have completely different and incompatible scoring functions which produce values in different intervals, following different distributions. A vector search score close to 1 can mean a very close match, but it doesn’t mean the same for lexical search. Hybrid search methods, such as <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion">reciprocal rank fusion</a> (RRF) and linear combination of scores, assign new scores that blend the original scores from lexical and vector search.</p><p>After hybrid search, we can employ techniques such as <a href="https://www.elastic.co/docs/solutions/search/ranking/semantic-reranking">semantic reranking</a> and <a href="https://www.elastic.co/docs/solutions/search/ranking/learning-to-rank-ltr">Learning To Rank</a> (LTR), which use specialized machine learning models to rerank the result.</p><p>With our most relevant results, we can use large language models (LLMs) to further enrich our response or pass the most relevant results as context to LLMs in agentic workflows in tools such as <a href="https://www.elastic.co/search-labs/blog/elastic-ai-agent-builder-context-engineering-introduction">Elastic Agent Builder</a>.</p><p>ES|QL is able to handle all these stages of retrieval. By design, ES|QL is a piped language, where each command transforms the input and sends the output to the next command. Each stage of retrieval is represented by one or more consecutive ES|QL commands. In this article, we show how each stage is supported in ES|QL.</p><h2>Vector search</h2><p>In Elasticsearch 9.2, we introduced tech preview support for dense vector search in ES|QL. This is as simple as calling the <code>knn</code> function, which only requires a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/dense-vector"><code>dense_vector</code></a> field and a query vector:</p>FROM books METADATA _score
| WHERE KNN(description_vector, ?query_vector)
| SORT _score DESC
| LIMIT 100<p>This query executes an approximate nearest neighbor search, retrieving 100 documents that are the most similar to the <code>query_vector</code>.</p><h2>Hybrid search: Reciprocal rank fusion</h2><p>In Elasticsearch 9.2, we introduced support for hybrid search using RRF and linear combination of results in ES|QL.</p><p>This allows combining vector search and lexical search results into a single result set.</p><p>To achieve this in ES|QL, we need to use the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fork"><code>FORK</code></a> and <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fuse"><code>FUSE</code></a> commands. <code>FORK</code> runs multiple branches of execution, and <code>FUSE</code> merges the results and assigns new relevance scores using RRF or linear combination.</p><p>In the following example, we use <code>FORK</code> to run two separate branches, where one is doing a lexical search using the <code>match</code> function, while the other is doing a vector search using the <code>knn</code> function. We then merge the results together using <code>FUSE</code>:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE KNN(description_vector, ?query_vector) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE // uses RRF by default
| SORT _score DESC<p>Let's decompose the query to better understand the execution model and first look at the output of the <code>FORK</code> command:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE KNN(description_vector, ?query_vector) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)<p>The<code> FORK</code> commands outputs the results from both branches and adds a <code>_fork</code> discriminator column:</p><p>_id</p><p>title</p><p>_score</p><p>_fork</p><p>4001</p><p>The Hobbit</p><p>0.88</p><p>fork1</p><p>3999</p><p>The Fellowship of the Ring</p><p>0.88</p><p>fork1</p><p>4005</p><p>The Two Towers</p><p>0.86</p><p>fork1</p><p>4006</p><p>The Return of the King</p><p>0.84</p><p>fork1</p><p>4123</p><p>The Silmarillion</p><p>0.78</p><p>fork1</p><p>4144</p><p>The Children of Húrin</p><p>0.79</p><p>fork1</p><p>4001</p><p>The Hobbit</p><p>4.55</p><p>fork2</p><p>3999</p><p>The Fellowship of the Ring</p><p>4.25</p><p>fork2</p><p>4123</p><p>The Silmarillion</p><p>4.11</p><p>fork2</p><p>4005</p><p>The Two Towers</p><p>3.8</p><p>fork2</p><p>4006</p><p>The Return of the King</p><p>4.1</p><p>fork2</p><p>As you’ll notice, certain documents appear twice, which is why we then use <code>FUSE</code> to merge rows that represent the same documents and assign new relevance scores. <code>FUSE</code> is executed in two stages:</p><ul><li><p>For each row, <code>FUSE</code> assigns a new relevance score, depending on the hybrid search algorithm that is being used.</p></li><li><p>Rows that represent the same document are merged together, and a new score is computed.</p></li></ul><p>In our example, we’re using RRF. As a first step, <code>FUSE</code> assigns a new score to each row using the RRF formula:</p>score(doc) = 1 / (rank_constant + rank(doc))<p>Where the <code>rank_constant</code> takes a default value of 60 and <code>rank(doc)</code>represents the position of the document in the result set.</p><p>In the first phase, our results become:</p><p>_id</p><p>title</p><p>_score</p><p>_fork</p><p>4001</p><p>The Hobbit</p><p>1 / (60 + 1) = 0.01639</p><p>fork1</p><p>3999</p><p>The Fellowship of the Ring</p><p>1 / (60 + 2) = 0.01613</p><p>fork1</p><p>4005</p><p>The Two Towers</p><p>1 / (60 + 3) = 0.01587</p><p>fork1</p><p>4006</p><p>The Return of the King</p><p>1 / (60 + 4) = 0.01563</p><p>fork1</p><p>4123</p><p> The Silmarillion</p><p>1 / (60 + 5) = 0.01538</p><p>fork1</p><p>4144</p><p>The Children of Húrin</p><p>1 / (60 + 6) = 0.01515</p><p>fork1</p><p>4001</p><p>The Hobbit</p><p>1 / (60 + 1) = 0.01639</p><p>fork2</p><p>3999</p><p>The Fellowship of the Ring</p><p>1 / (60 + 2) = 0.01613</p><p>fork2</p><p>4123</p><p>The Silmarillion</p><p>1 / (60 + 3) = 0.01587</p><p>fork2</p><p>4005</p><p>The Two Towers</p><p>1 / (60 + 4) = 0.01563</p><p>fork2</p><p>4006</p><p>The Return of the King</p><p>1 / (60 + 5) = 0.01538</p><p>fork2</p><p>Then the rows are merged together and a new score is assigned. Since a <code>SORT _score DESC</code> follows the <code>FUSE</code> command, the final results are:</p><p>_id</p><p>title</p><p>_score</p><p>4001</p><p>The Hobbit</p><p>0.01639 + 0.01639 = 0.03279</p><p>3999</p><p>The Fellowship of the Ring</p><p>0.01613 + 0.01613 = 0.03226</p><p>4005</p><p>The Two Towers</p><p>0.01587 + 0.01563 = 0.0315</p><p>4123</p><p>The Silmarillion</p><p>0.01538 + 0.01587 = 0.03125</p><p>4006</p><p>The Return of the King</p><p>0.01563 + 0.01538 = 0.03101</p><p>4144</p><p>The Children of Húrin</p><p>0.01515</p><h2>Hybrid search: Linear combination of scores</h2><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion">Reciprocal rank fusion</a> is the simplest way to do hybrid search, but it isn’t the only hybrid search method that we support in ES|QL.</p><p>In the following example, we use <code>FUSE</code> to combine lexical and <a href="https://www.elastic.co/docs/solutions/search/semantic-search/semantic-search-semantic-text">semantic search</a> results using linear combination of scores:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE MATCH(semantic_description, ?query) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE LINEAR WITH { "weights": { "fork1": 0.7, "fork2": 0.3 } }
| SORT _score DESC<p>Let's first decompose the query and take a look at the input of the <code>FUSE</code> command when we only run the <code>FORK</code> command.</p><p>Notice that we use the <code>match</code> function, which is able to not only query lexical fields, such as <code>text</code> or <code>keyword</code>, but also <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text"><code>semantic_text</code></a> fields.</p><p>The first <code>FORK</code> branch executes a semantic query by querying a <code>semantic_text</code> field, while the second one executes a lexical query:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE MATCH(semantic_description, ?query) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)<p>The output of the <code>FORK</code> command can contain rows with the same <code>_id</code> and <code>_index</code> values representing the same Elasticsearch document:</p><p>_id</p><p>title</p><p>_score</p><p>_fork</p><p>4001</p><p>The Hobbit</p><p>0.88</p><p>fork1</p><p>3999</p><p>The Fellowship of the Ring</p><p>0.88</p><p>fork1</p><p>4005</p><p>The Two Towers</p><p>0.86</p><p>fork1</p><p>4006</p><p>The Return of the King</p><p>0.84</p><p>fork1</p><p>4123</p><p>The Silmarillion</p><p>0.78</p><p>fork1</p><p>4144</p><p>The Children of Húrin</p><p>0.79</p><p>fork1</p><p>4001</p><p>The Hobbit</p><p>4.55</p><p>fork2</p><p>3999</p><p>The Fellowship of the Ring</p><p>4.25</p><p>fork2</p><p>4123</p><p>The Silmarillion</p><p>4.11</p><p>fork2</p><p>4005</p><p>The Two Towers</p><p>3.8</p><p>fork2</p><p>4006</p><p>The Return of the King</p><p>4.1</p><p>fork2</p><p>In the next step, we use <code>FUSE</code> to merge rows that have the same <code>_id</code> and <code>_index</code> values, and assign new relevance scores.</p><p>The new score is a linear combination of the scores the row had in each <code>FORK</code> branch:</p>_score = 0.7 *_score1 + 0.3 * _score2<p>Here, <code>_score1</code> and <code>_score2</code> represent the score a document has in the first <code>FORK</code> branch and the second <code>FORK</code> branch, respectively.</p><p>Notice that we also apply custom weights, giving more weight to the semantic score over the lexical one, resulting in this set of documents:</p><p>_id</p><p>title</p><p>_score</p><p>4001</p><p>The Hobbit</p><p>0.7 * 0.88 + 0.3 * 4.55 = 1.981</p><p>3999</p><p>The Fellowship of the Ring</p><p>0.7 * 0.88 + 0.3 * 4.25 = 1.891</p><p>4006</p><p>The Return of the King</p><p>0.7 * 0.84 + 0.3 * 4.1 = 1.818</p><p>4123</p><p>The Silmarillion</p><p>0.7 * 0.78 + 0.3 * 4.11 = 1.779</p><p>4005</p><p>The Two Towers</p><p>0.7 * 0.86 + 0.3 * 3.8 = 1.742</p><p>4144</p><p>The Children of Húrin</p><p>0.7 * 0.79 + 0.3 * 0 = 0.553</p><p>One challenge is that the semantic and lexical scores can be incompatible to apply the linear combination, since they can follow completely different distributions. To mitigate this, we first need to normalize the scores, employing score normalization methods, such as <code>minmax</code>. This ensures that the scores from each <code>FORK</code> branch are first normalized to take values between 0 and 1, before applying the linear combination formula.</p><p>To achieve this with <code>FUSE</code>, we need to specify the <code>normalizer</code> option:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE MATCH(semantic_description, ?query) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE LINEAR WITH { "weights": { "fork1": 0.7, "fork2": 0.3 }, "normalizer": "minmax" }
| SORT _score DESC<h2>Semantic reranking</h2><p>At this stage, after hybrid search, we should be left with the most relevant documents. We can now use semantic reranking to reorder the results using the <code>RERANK</code> command. By default, <code>RERANK</code> uses the latest Elastic <a href="https://www.elastic.co/docs/solutions/search/ranking/semantic-reranking">semantic reranking</a> machine learning model, so no additional configuration is needed:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE KNN(description_vector, ?query_vector) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE
| SORT _score DESC
| LIMIT 100
| RERANK ?query ON description
| SORT _score DESC<p>We now have our best results, sorted by relevance.</p><p>One key feature that sets the <code>RERANK</code> command apart from other products that offer semantic reranking integrations is that it doesn’t require the input to represent a mapped field from an index. <code>RERANK</code> only expects an expression that evaluates to a string value, making it possible to do semantic reranking using multiple fields:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE KNN(description_vector, ?query_vector) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE
| SORT _score DESC
| LIMIT 100
| RERANK ?query ON CONCAT(title, "\n", description) 
| SORT _score DESC<h2>LLM completions</h2><p>Now we have a set of highly relevant, reranked results.</p><p>At this stage, you might simply decide to return the results back to your application or you might want to further enhance your results using LLM completions.</p><p>If you’re using ES|QL as part of a retrieval-augmented generation (RAG) workflow, you can choose to call your favorite LLM directly from ES|QL.
To achieve this, we’ve added a new <code>COMPLETION</code> command that takes in a prompt, a completion inference ID which designates which LLM to call, and a column identifier to specify where to output the LLM response.</p><p>In the following example, we’re using <code>COMPLETION</code> to add a new <code>_completion</code> column that contains the summary of the <code>content</code> column:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE KNN(description_vector, ?query_vector) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE
| SORT _score DESC
| LIMIT 100
| RERANK ?query ON description
| SORT _score DESC
| LIMIT 10
| COMPLETION CONCAT("Summarize the following:\n", description) WITH { "inference_id" : "my_inference_endpoint" } <p>Each row now contains a summary:</p><p>_id</p><p>title</p><p>_score</p><p>summary</p><p>4001</p><p>The Hobbit</p><p>0.03279</p><p>Bilbo helps dwarves reclaim Erebor from the dragon Smaug.</p><p>3999</p><p>The Fellowship of the Ring</p><p>0.03226</p><p>Frodo begins the quest to destroy the One Ring.</p><p>4005</p><p>The Two Towers</p><p>0.0315</p><p>The Fellowship splits; war comes to Rohan; Frodo nears Mordor.</p><p>4123</p><p>The Silmarillion</p><p>0.03125</p><p>Ancient myths and history of Middle-earth's First Age.</p><p>4006</p><p>The Return of the King</p><p>0.3101</p><p>Sauron is defeated and Aragorn is crowned King.</p><p>4144</p><p>The Children of Húrin</p><p>0.01515</p><p>The tragic tale of Túrin Turambar's cursed life.</p><p>In another use case, you may simply want to answer a question using the proprietary data that you have indexed in Elasticsearch. In this case, the best search results that we’ve computed in the previous stage can be used as context for the prompt:</p>FROM books METADATA _score, _id, _index
| FORK (WHERE KNN(description_vector, ?query_vector) | SORT _score DESC | LIMIT 100)
       (WHERE MATCH(description, ?query) | SORT _score DESC | LIMIT 100)
| FUSE
| SORT _score DESC
| LIMIT 100
| RERANK ?query ON description
| SORT _score DESC
| LIMIT 10
| STATS context = VALUES(CONCAT(title, "\n", description)
| COMPLETION CONCAT("Answer the following question ", ?query, "based on:\n", context) WITH { "inference_id" : "my_inference_endpoint" }<p>Since the <code>COMPLETION</code> command unlocks the ability to send any prompt to an LLM, the possibilities are endless. Although we’re only showing a few examples, the <code>COMPLETION</code> command can be used in a wide range of scenarios, from security analysts using it to assign scores depending on whether a log event can represent a malicious action or data scientists using it to analyze data, to cases where you just need to<a href="https://www.elastic.co/search-labs/blog/esql-completion-command-llm-fact-generator"> generate Chuck Norris facts based on your data</a>.</p><h2>This is only the beginning</h2><p>In the future, we’ll be expanding ES|QL to improve semantic reranking for long documents, better conditional execution of the ES|QL queries using multiple <code>FORK</code> commands, support sparse vector queries, removing close duplicate results to enhance result diversity, allowing full text search on runtime generated columns, and many other scenarios.</p><p>Additional tutorials and guides:</p><ul><li><p><a href="https://www.elastic.co/docs/solutions/search/esql-for-search">ES|QL for search</a></p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-search-tutorial">ES|QL for search tutorial</a></p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text">Semantic_text field type</a></p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fork"><code>FORK</code></a> and <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fuse"><code>FUSE</code></a> documentation</p></li><li><p>ES|QL search functions</p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/hybrid-search-multi-stage-retrieval-esql</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/hybrid-search-multi-stage-retrieval-esql</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Hybrid Search]]></category>
    <category><![CDATA[Relevance]]></category>
    <dc:creator><![CDATA[Ioana Tagirta,Aurélien Foucret,Carlos Delgado]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3382265814939417/6a170df3a929cf1246ae0a61/fceada10b0c09d6a4a372f137bb3040e1ff41fbf-1600x895.png" length="0" type="image/png"/>
    <pubDate>Thu, 08 Jan 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Using ES|QL COMPLETION + an LLM to write a Chuck Norris fact generator in 5 minutes]]></title>
    <description><![CDATA[Discover how to use the ES|QL COMPLETION command to turn your Elasticsearch data into creative output using an LLM in just a few lines of code.]]></description>
    <content:encoded><![CDATA[<p>What if you could turn your Elasticsearch data into creative output using an LLM—in just a few lines of code? With the new <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/completion">COMPLETION command</a> in <strong>ES|QL</strong>, now you can.</p><p>Let’s build something fun to show it off: a Chuck Norris fact generator. We'll combine movie descriptions with a GPT model to generate facts so legendary even Rambo would be impressed.</p><h2>What you'll need</h2><ul><li><p>Access to an LLM (like OpenAI’s GPT-4o in our example below)</p></li><li><p>A dataset of movie descriptions </p></li></ul><p>You can download a <a href="https://www.kaggle.com/datasets/ursmaheshj/top-10000-popular-movies-tmdb-05-2023?resource=download">sample dataset</a> from Kaggle and upload it to your Elasticsearch cluster using the Data Visualizer in Kibana or the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk"><code>_bulk</code></a><a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-bulk"> API</a>.</p><h2>Setting up the inference endpoint</h2><p>Before you can run the <code>COMPLETION</code> command, you need to create an inference endpoint for the model you want to use via the <code>_inference</code>  API.</p><p>Here’s how to set up GPT-4o with OpenAI:</p><p>Once this is in place, you can reference <code>my-gpt-4o-endpoint</code> directly in your query.</p><h2>The query</h2><p>Here’s the magic in action. This single <strong>ES|QL</strong> query handles the entire workflow: it finds a movie based on your input, constructs a prompt from its description, and then calls the LLM to generate a legendary Chuck Norris fact. Below is the full <strong>ES|QL</strong> query that powers our Chuck Norris fact generator. It takes in a movie query, retrieves the most relevant description, turns it into a prompt, and sends it off to the LLM—all in a single, piped query.</p><p>Here’s what comes back:</p><p>Yes, the model really said that. 💪🐐🚁</p><h2>Dissecting the query</h2><p>Let’s dissect the query and break down what’s happening, step by step.</p><h3>Step 1: Retrieve relevant movie data</h3><p>We begin by searching for the most relevant movie for the user query.
We use the <code>MATCH</code> function to search both the title and overview fields for the text provided by the <code>query</code> parameter, keeping only the first result, sorted by relevance using the metadata <code>_score</code> field:</p><p>This narrows down our dataset to the best match, giving us the movie's title and description, which will become the context for the LLM.</p><h3>Step 2: Build the prompt from the context</h3><p>Now we create the input prompt for the LLM by concatenating a static instruction provided as a query parameter, denoted by <code>?instruction</code>, with the movie’s overview:</p><p>This creates a new <code>prompt</code> column combining the provided instruction with the overview field from the returned document, which for our request looks a bit like this:</p>Generate a Chuck Norris Fact from the following description:
Combat has taken its toll on Rambo, but he's finally begun to find inner peace in a monastery. When Rambo's friend and mentor Col. Trautman asks for his help on a top secret mission to Afghanistan, Rambo declines but must reconsider when Trautman is captured.<p>You can easily swap in different instructions to change the tone or style of what the LLM generates by tweaking the instruction parameter. And because the prompt is just another <strong>ES|QL</strong> expression, you can compose it with any string-generating function—whether it’s simple concatenation, conditional logic, or even formatting based on your document content.</p><h3>Step 3: Generate text using the LLM</h3><p>Finally, we pass the prompt to the inference endpoint connected to our model using our new <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/completion"><code>COMPLETION</code></a><a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/completion"> command</a>, and select which fields to return:</p><p>The result? A Chuck Norris fact, rooted in your movie data without any extra tooling required.</p><p>This example also demonstrates the full power of ES|QL's piped structure. Each step flows naturally into the next, letting you express a full retrieval augmented generation (RAG) pipeline in a single, declarative query. It’s clean, composable, and stays entirely inside Elasticsearch.</p><h2>What’s next?</h2><p>While the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/completion">COMPLETION command</a> is still a tech preview, this new feature unlocks a whole new world of possibilities—from summarization and content generation to enrichment and storytelling. Try it yourself! Point it at your favorite movie, tweak the prompt, or go wild and generate haikus from SQL errors. The power is yours.</p><p>Let us know what you build! 💬</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-completion-command-llm-fact-generator</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-completion-command-llm-fact-generator</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Aurélien Foucret]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbca7a25e3c272aa8/6a17ffb1be60868d100049d2/6494d88d51edf6a5b31a92b8439792354eae7190-1536x1024.png" length="0" type="image/png"/>
    <pubDate>Thu, 28 Aug 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[From ES|QL to native Pandas dataframes in Python]]></title>
    <description><![CDATA[Learn how to export ES|QL queries as native Pandas dataframes in Python through practical examples.]]></description>
    <content:encoded><![CDATA[<p>Since Elasticsearch 8.15 or with Elasticsearch Serverless, <a href="https://github.com/elastic/elasticsearch/pull/109873">ES|QL responses support the Apache Arrow streaming format</a>. This blog post will show you how to take advantage of it in Python. In an <a href="https://www.elastic.co/search-labs/blog/esql-pandas-dataframes-python">earlier blog post</a>, I demonstrated how to convert ES|QL queries to Pandas dataframes using CSV as an intermediate representation. Unfortunately, CSV requires explicit type declarations, is slow (especially for larger datasets) and does not handle nested arrays and objects. Apache Arrow lifts all these limitations.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta5d7fdd54f312f06/6a17d7bbfaa913809b93c6db/7decbe330061eae8108f7ad6a32a2df01f55244f-389x144.svg" alt="ES|QL produces tables" /><h2>ES|QL to Pandas dataframes in Python</h2><h3>Importing test data</h3><p>First, let's import some test data. As before, we will be using the <code>employees</code> <a href="https://github.com/elastic/elasticsearch/blob/d46bcc968e6cabca55f1a62b2218e9fc4e84e9d4/x-pack/plugin/esql/qa/testFixtures/src/main/resources/employees.csv">sample data</a> and <a href="https://github.com/elastic/elasticsearch/blob/main/x-pack/plugin/esql/qa/testFixtures/src/main/resources/mapping-default.json">mappings</a>. The easiest way to load this dataset is to <a href="https://gist.github.com/pquentin/7cf29a5932cf52b293699dd994b1a276">run these two Elasticsearch API requests</a> in the <a href="https://www.elastic.co/guide/en/kibana/current/console-kibana.html">Kibana Console</a>.</p><h3>Converting dataset to a Pandas DataFrame object</h3><p>OK, with that out of the way, let's convert the full <code>employees</code> dataset to a Pandas DataFrame object using the ES|QL Arrow export:</p>from elasticsearch import Elasticsearch
import pandas as pd

client = Elasticsearch(
    "https://[host].elastic-cloud.com",
    api_key="...",
)

response = client.esql.query(
    query="""
    FROM employees
    | DROP is_rehired,job_positions,salary_change*
    | LIMIT 500
    """,
    format="arrow",
)
df = response.to_pandas(types_mapper=pd.ArrowDtype)
print(df)
<p>Even though this dataset only contains 100 records, we use a <code>LIMIT</code> command to avoid ES|QL warning us about potentially missing records. This prints the following dataframe:</p>    avg_worked_seconds           birth_date  ...  salary still_hired
0            268728049  1953-09-02 00:00:00  ...   57305        True
1            328922887  1964-06-02 00:00:00  ...   56371        True
2            200296405  1959-12-03 00:00:00  ...   61805       False
3            311267831  1954-05-01 00:00:00  ...   36174        True
4            244294991  1955-01-21 00:00:00  ...   63528        True
..                 ...                  ...  ...     ...         ...
95           204381503  1954-09-16 00:00:00  ...   43889       False
96           206258084  1952-02-27 00:00:00  ...   71165       False
97           272392146  1961-09-23 00:00:00  ...   44817       False
98           377713748  1956-05-25 00:00:00  ...   73578        True
99           223910853  1953-04-21 00:00:00  ...   68431        True

[100 rows x 17 columns]
<p>OK, so what actually happened here?</p><ul><li><p>Given <code>format="arrow"</code>, Elasticsearch returns binary Arrow streaming data</p></li><li><p>The Elasticsearch Python client looks at the Content-Type header and creates a <a href="https://arrow.apache.org/docs/python/index.html">PyArrow object</a></p></li><li><p>Finally, PyArrow's <a href="https://arrow.apache.org/docs/python/pandas.html">Pandas integration</a> converts the PyArrow object to a Pandas dataframe.</p></li></ul><p>Note that the <code>types_mapper=pd.ArrowDtype</code> parameter asks Pandas to use a PyArrow backend instead of a NumPy backend, since the source data is PyArrow. While this backend is not enabled by default for compatibility reasons, it <a href="https://datapythonista.me/blog/pandas-20-and-the-arrow-revolution-part-i">has many advantages</a>: it handles missing values, is faster, more interopable and supports more types. (This is not a <a href="https://arrow.apache.org/docs/python/pandas.html#memory-usage-and-zero-copy">zero copy conversion</a>, however.)</p><p>For this example to work, the Pandas and PyArrow optional dependencies need to be installed. If you want to use another dataframe library such as Polars instead, you don't need Pandas and can directly use <a href="https://docs.pola.rs/api/python/stable/reference/api/polars.from_arrow.html"><code>polars.from_arrow</code></a> to create a Polars DataFrame from the PyArrow table returned by the Elasticsearch client.</p><p>One limitation is that Elasticsearch does not currently handle multi-valued fields, which is why we had to drop the <code>is_rehired</code>, <code>job_positions</code> and <code>salary_change</code> columns. This limitation will be lifted in a future version of Elasticsearch.</p><p>Anyway, you now have a Pandas dataframe that you can use to analyze your data further. But you can also continue massaging the data using ES|QL, which is particularly useful when queries return more than 10,000 rows, the current maximum number of rows that ES|QL queries can return.</p><h3>More complex queries</h3><p>In the next example, we're counting how many employees are speaking a given language by using <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-commands.html#esql-stats-by"><code>STATS ... BY</code></a> (not unlike <code>GROUP BY</code> in SQL). And then we sort the result with the <code>languages</code> column using <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-commands.html#esql-sort"><code>SORT</code></a>:</p>response = client.esql.query(
    query="""
    FROM employees
    | DROP is_rehired,job_positions,salary_change*
    | STATS count = COUNT(emp_no) BY languages
    | SORT languages
    | LIMIT 500
    """,
    format="arrow",
)

df = response.to_pandas(types_mapper=pd.ArrowDtype)
print(df)
<p>Unlike with CSV, we did not have to specify any types, as Arrow data already includes types. Here's the result:</p>   count  languages
0     15          1
1     19          2
2     17          3
3     18          4
4     21          5
5     10       &lt;NA&gt;
<p>21 employees speak 5 languages, wow! And 10 employees did not declare any spoken language. The missing value is denoted by <code>&lt;NA&gt;</code>, which is consistently used for missing data with the PyArrow backend. If we had used the NumPy backend instead, this column would have been converted to floats and the missing value would have been a confusing <code>NaN</code>, as <a href="https://pandas.pydata.org/docs/user_guide/missing_data.html">NumPy integers don't have any sentinel value for missing data</a>.</p><h3>Queries with parameters</h3><p>Finally, suppose that you want to expand the query from the previous section to only consider employees that speak N or more languages, with N being a variable parameter. For this we can use <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-rest.html#esql-rest-params">ES|QL's built-in support for parameters</a>, which eliminates the risk of an injection attack associated with manually assembling queries with variable parts:</p>response = client.esql.query(
    query="""
    FROM employees
    | DROP is_rehired,job_positions,salary_change*
    | STATS count = COUNT(emp_no) BY languages
    | WHERE languages &gt;= (?)
    | SORT languages
    | LIMIT 500
    """,
    format="arrow",
    params=[3],
)

df = response.to_pandas(types_mapper=pd.ArrowDtype)
print(df)
<p>which prints the following:</p>   count  languages
0     17          3
1     18          4
2     21          5
<h2>Conclusion</h2><p>As we saw, ES|QL's native Arrow support makes working with Pandas and other DataFrame libraries even nicer than using CSV and it will continue to improve over time, with the multi-value support coming in a future version of Elasticsearch.</p><h2>Additional resources</h2><p>If you want to learn more about ES|QL, the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">ES|QL documentation</a> is the best place to start. You can also check out <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/Boston-Celtics-Demo/celtics-esql-demo.ipynb">this other Python example using Boston Celtics data</a>. To know more about the Python Elasticsearch client itself, you can <a href="https://www.elastic.co/guide/en/elasticsearch/client/python-api/current/index.html">refer to the documentation</a>, ask a question <a href="https://discuss.elastic.co/tag/language-clients">on Discuss with the language-clients tag</a> or <a href="https://github.com/elastic/elasticsearch-py">open a new issue</a> if you found a bug or have a feature request. Thank you!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-pandas-native-dataframes-python</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-pandas-native-dataframes-python</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Quentin Pradet]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb1808f6b0c1b0ed3/6a17d7bcec0f89c6c35a644e/1b32822c3bf2ad216b21d819c5795f080b6e6cbf-500x500.png" length="0" type="image/png"/>
    <pubDate>Thu, 05 Sep 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[An Elasticsearch Query Language (ES|QL) analysis: Millionaire odds vs. hit by a bus]]></title>
    <description><![CDATA[Use Elasticsearch Query Language (ES|QL) to run statistical analysis on demographic data index in Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch Query Language (ES|QL) is designed for fast, efficient querying of large datasets. It has a straightforward syntax which will allow you to write complex queries easily, with a pipe based language, reducing the learning curve. We're going to use ES|QL to run statistical analysis and compare different odds.</p><p>If you are reading this, you probably want to know how rich you can get before actually reaching the same odds of being hit by a bus. I can't blame you, I want to know too. Let's work out the odds so that we can make sure we win the lottery rather than get in an accident!</p><p>What we are going to see in this blog is figuring out the probability of being hit by a bus and the probability of achieving wealth. We'll then compare both and understand until what point your chances of getting rich are higher, and when you should consider getting life insurance.</p><p>So how are we going to do that? This is going to be a mix of magic numbers pulled from different articles online, some synthetics data and the power of ES|QL, the new Elasticsearch Query Language. Let's get started.</p><h2>Data for the ES|QL analysis</h2><h3>The magic number</h3><p>The challenge starts here as the dataset is going to be somewhat challenging to find. We are then going to assume for the sake of the example that ChatGPT is always right. Let’s see what we get for the following question:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7b0ac87235183705/6a17d75e3e03d768544f2ac1/e53cf27540f30f58af23fc26d3d1b93cc7fd5497-1440x229.png" alt="bus-odds" /><p>Cough Cough… That sounds about right, this is going to be our magic number.</p><h3>Generating the wealth data</h3><h4>Prerequisites</h4><p>Before running any of the scripts below, make sure to install the following packages:</p>
elasticsearch==8.14.0
matplotlib
numpy
panda
scipy

<p>Now, there is one more thing we need, a representative dataset with wealth distribution to compute wealth probability. There is definitely some portion of it here and there, but again, for the example we are going to generate a 500K line dataset with the below python script. I am using python 3.11.5 in this example:</p>
import pandas as pd
import numpy as np
import getpass
from elasticsearch import Elasticsearch, helpers

# Input the Elasticsearch host
hosts = input('Enter your Elasticsearch host address : ')

# Securely input the Elasticsearch API key
api_key = getpass.getpass(prompt='Enter your Elasticsearch API Key: ')

# Initialize Elasticsearch client
client = Elasticsearch(
    hosts=hosts,
    api_key=api_key,
)

# Generate synthetic data with a highly skewed distribution
num_records = 500000
np.random.seed(42)  # Ensure reproducibility

# Generate net worth using a highly skewed distribution
ages = np.random.randint(20, 80, num_records)  # Random ages between 20 and 80
incomes = np.random.exponential(scale=10000, size=num_records)  # Exponential distribution for income
# Use a more skewed distribution for net worth with a much larger range
net_worths = np.random.exponential(scale=100000000, size=num_records)  # Extremely skewed net worth

# Scale up the net worths to reach up to $100 billion
net_worths = np.clip(net_worths, 0, 100000000000)

# Create DataFrame
df = pd.DataFrame({
    'id': range(1, num_records + 1),
    'age': ages,
    'income': incomes,
    'net_worth': net_worths,
    'counter': range(1, num_records + 1)  # Add a counter field for pagination
})

# Index the data into Elasticsearch
index_name = 'raw_wealth_data_large'
try:
    if client.indices.exists(index=index_name):
        client.indices.delete(index=index_name)
except exceptions.NotFoundError:
    pass
client.indices.create(index=index_name)


def generator(df):
    for index, row in df.iterrows():
        yield {
            "_index": index_name,
            "_source": row.to_dict()
        }

helpers.bulk(client, generator(df))

print("Data indexed successfully.")
<p>It should take some time to run depending on your configuration since we are injecting 500K documents here!</p><p>FYI, after playing with a couple of versions of the script above and the ESQL query on the synthetic data, it was obvious that the net worth generated across the population was not really representative of the real world. So I decided to use a log-normal distribution (np.random.lognormal) for income to reflect a more realistic spread where most people have lower incomes, and fewer people have very high incomes.</p><p>Net Worth Calculation: Used a combination of random multipliers (np.random.uniform(0.5, 5)) and additional noise (np.random.normal(0, 10000)) to calculate net worth. Added a check to ensure no negative net worth values by using np.maximum(0, net_worths).</p><p>Not only have we generated 500K documents, but we also used the Elasticsearch python client to bulk ingest all these documents in our deployment. Please note that you will find the endpoint to pass in as hosts Cloud ID in the code above.</p><p>For the deployment API key, open Kibana, and generate the key in Stack Management / API Keys:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb38e3a3a44c66b9e/6a17d7606864a40557b685e2/544c32086c18ca3b5e65fa0e5bbf2d60490a7f66-1440x864.png" alt="api-key" /><p>The good news is that if you have a real data set, all you will need to do is to change the above code to read your dataset and write documents with the same data mapping.</p><p>Ok we're getting there! The next step is pouring our wealth distribution.</p><h2>ES|QL wealth analysis</h2><h3>Introducing ES|QL: A powerful tool for data analysis</h3><p>The arrival of Elasticsearch Query Language (ES|QL) is very exciting news for our users. It largely simplifies querying, analyzing, and visualizing data stored in Elasticsearch, making it a powerful tool for all data-driven use cases.</p><p>ES|QL comes with a variety of functions and operators, to perform aggregations, statistical analyses, and data transformations. We won’t address them all in this blog post, however <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql.html">our documentation</a> is very detailed and will help you familiarize with the language and the possibilities.</p><p>To get started with ES|QL today and run the blog post queries, simply <a href="https://www.elastic.co/getting-started?utm_source=github&amp;utm_content=elasticsearch-labs-notebook">start a trial on Elastic Cloud</a>, load the data and run your first ES|QL query.</p><h3>Understanding the wealth distribution with our first query</h3><p>To get familiar with the dataset, head to Discover in Kibana and switch to ES|QL in the dropdown on the left hand side:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5aba101d232dbb13/6a17d762e8fbce48b53a174c/357703b5c0b543daa61fe98354de29e57182469d-1440x585.png" alt="discover" /><p>Let’s fire our first request:</p>from raw_wealth_data_large | keep age, id, income, net_worth | limit 10
<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd9597115bb01b8c3/6a17d7643e03d7488f4f2ac5/898970fd63e9b2811bf24badd72ef57a45912564-1312x1930.png" alt="result set" /><p>As you could expect from our indexing script earlier, we are finding the documents we bulk ingested, notice the simplicity of pulling data from a given dataset with ES|QL where every query starts with the From clause, then your index.</p><p>In the query above given we have 500K lines, we limited the amount of returned documents to 10. To do this, we are passing the output of the first segment of the query via a pipe to the limit command to only get 10 results. Pretty intuitive, right?</p><p>Alright, what would be more interesting is to understand the wealth distribution in our dataset, for this we will leverage one of the 30 functions ES|QL provides, namely percentile.</p><p>This will allow us to understand the relative position of each data point within the distribution of net worth. By calculating the median percentile (50th percentile), we can gauge where an individual’s net worth stands compared to others.</p>
FROM raw_wealth_data_large
| stats p50 = percentile(net_worth, 50) 

<p>Like our first query, we are passing the output of our index to another function, Stats, which combined with the percentile function will output the median net worth:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt35023abe49d381f3/6a17d7654b055d09c5432048/57f65e58941783994d67dee7ee763a4934bd8ca9-1440x640.png" alt="result set" /><p>The median is about 54K, which unfortunately is probably optimistic compared to the real world, but we are not going to solve this here. If we go a little further, we can look at the distribution in more granularity by computing more percentiles:</p>
FROM raw_wealth_data_large
| STATS  p25 = percentile(net_worth, 25)
       , p50 = percentile(net_worth, 50)
       , p75 = percentile(net_worth, 75)
       , p90 = percentile(net_worth, 90)
       , p95 = percentile(net_worth, 95)
       , p96 = percentile(net_worth, 96)
       , p98 = percentile(net_worth, 98)
       , p97 = percentile(net_worth, 97)
       , p99 = percentile(net_worth, 99)
| keep p25, p25, p50, p75, p90, p95, p96, p97, p98, p99

<p>With the below output:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt48a66c799c539a68/6a17d767414c644eaa944fd6/59f1c99d52887e275422aea7d4e0547c8f3ca886-1440x458.png" alt="Percentile result" /><p>The data reveals a significant disparity in wealth distribution, with the majority of wealth being concentrated among the richest individuals. Specifically, the top 5% (95th percentile) possess a disproportionately large portion of the total wealth, with a net worth starting at $852,988.26 and increasing dramatically in the higher percentiles.</p><p>The 99th percentile individuals hold a net worth exceeding $2 million, highlighting the skewed nature of wealth distribution. This indicates that a substantial portion of the population has modest net worth, which is probably what we want for this example.</p><p>Another way to look at this is to augment the previous query and grouping by age to see if there is, (in our synthetic dataset), a relation between wealth and age:</p>
FROM raw_wealth_data_large
| STATS  p25 = percentile(net_worth, 25)
      , p50 = percentile(net_worth, 50)
      , p75 = percentile(net_worth, 75)
      , p90 = percentile(net_worth, 90)
      , p95 = percentile(net_worth, 95)
      , p96 = percentile(net_worth, 96)
      , p98 = percentile(net_worth, 98)
      , p97 = percentile(net_worth, 97)
      , p99 = percentile(net_worth, 99) by age
| keep p25, p25, p50, p75, p90, p95, p96, p97, p98, p99, age
<p>This could be visualized in a Kibana dashboard. Simply:</p><ul><li><p>Navigate to Dashboard</p></li><li><p>Add a new ES|QL visualization</p></li><li><p>Copy and paste our query</p></li><li><p>Move the age field to the horizontal axis in the visualization configuration</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5e72df525e4fe575/6a17d769b1e113339979f0d3/f47a62215862c5981d2642128c2e41ac707a7476-1066x1864.png" alt="Create ESQL visualization" /><p>Which will output:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2de89c1c00eb23be/6a17d76afbc5f807ff491908/59e17103d0233fa9d79dd65e510b776b96041489-1440x854.png" alt="Visualization output" /><p>The above suggests that the data generator randomized wealth uniformly across the population age, there is no specific trend pattern we can really see.</p><h4>Median Absolute Deviation (MAD)</h4><p>We calculate the median absolute deviation (MAD) to measure the variability of net worth in a robust manner, less influenced by outliers.</p>
FROM raw_wealth_data_large
| stats median_net_worth = MEDIAN(net_worth), mad_net_worth = MEDIAN_ABSOLUTE_DEVIATION(net_worth)
| keep median_net_worth, mad_net_worth

<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc1488254b38e203e/6a17d76c2f4a5c2a81fa8766/535103bff9200ff4c2566418fedf8654de30fc13-1440x335.png" alt="Visualization output" /><p>With a median net worth of 44,205.44, we can infer the typical range of Net Worth: Most individuals’ net worth falls within a range of 9,581.78 to $97,992.66.</p><h3>The statistical showdown between Net Worth and Bus Collision</h3><p>Alright, this is the moment to understand how rich we can get, based on our dataset, before getting hit by a bus. To do that, we are going to leverage ES|QL to pull our entire dataset in chunks and load it into a pandas dataframe to build a net worth probability distribution. Finally, we will determine where the ends meet between the net worth and bus collision probabilities.</p><p>The entire Python <a href="https://github.com/elastic/elasticsearch-labs/blob/main/supporting-blog-content/esql-millionaire/millionaire.ipynb">notebook is available here</a>. I also recommend you read <a href="https://www.elastic.co/search-labs/blog/esql-pandas-dataframes-python">this blog post</a> which walks you through using ES|QL with pandas dataframes.</p><h4>Helper functions</h4><p>As you can see in the previously referred blog post, we introduced support for ES|QL since version 8.12 of the Elasticsearch python client. Thus our notebook first defines the below functions:</p>
from io import StringIO

# Function to execute ESQL query and fetch data in chunks
def execute_esql_query(query):
    response = client.esql.query(query=query, format="csv")
    return pd.read_csv(StringIO(response.body))

# Function to fetch paginated data using the counter field
def fetch_paginated_data(index, num_records, size=10000):
    all_data = pd.DataFrame()
    for start in range(1, num_records + 1, size):
        end = start + size - 1
        query = f"""
        FROM {index}
        | WHERE counter &gt;= {start} AND counter &lt;= {end}
        | limit {size}
        """
        data_chunk = execute_esql_query(query)
        all_data = pd.concat([all_data, data_chunk], ignore_index=True)
    return all_data

<p>The first function is straightforward and executes an ES|QL query, the second is fetching the entire dataset from our index. Notice the trick in there that I am using a counter built-in to a field in my index to paginate through the data. This is workaround I am using while our engineering team is working on <a href="https://github.com/elastic/elasticsearch/issues/100000">the support for pagination in ES|QL</a>.</p><p>Next, knowing that we have 500K documents in our index, we simply call these function to load the data in a data frame:</p>
# Fetch all data using pagination and ES|QL
num_records = 500000
all_data_df = fetch_paginated_data(index_name, num_records)
print(f"Total Data Retrieved: {len(all_data_df)} records")

<h4>Fit Pareto distribution</h4><p>Next, we fit our data to a Pareto distribution, which is often used to model wealth distribution because it reflects the reality that a small percentage of the population controls most of the wealth. By fitting our data to this distribution, we can more accurately represent the probabilities of different net worth levels.</p>from scipy.stats import pareto



# Fit a Pareto distribution to the data
shape, loc, scale = pareto.fit(all_data_df['net_worth'], floc=0)

# Calculate the probability density for each net worth
all_data_df['net_worth_probability'] = pareto.pdf(all_data_df['net_worth'], shape, loc=loc, scale=scale)

# Normalize the probabilities to sum to 1
all_data_df['net_worth_probability'] /= all_data_df['net_worth_probability'].sum()

print("Data with Net Worth Probability:")
print(all_data_df.head())

<p>We can visualize the pareto distribution with the code below: ``</p>
import matplotlib.pyplot as plt
from scipy.stats import pareto

# Assuming all_data_df contains the fetched net worth data from Elasticsearch
# Fit a Pareto distribution to the data
shape, loc, scale = pareto.fit(all_data_df['net_worth'], floc=0)

# Plot the Net Worth Probability Distribution
plt.figure(figsize=(10, 6))

# Plot histogram of empirical net worth data
plt.hist(all_data_df['net_worth'], bins=100, density=True, alpha=0.6, color='g', label='Empirical Data')

# Plot fitted Pareto distribution
xmin, xmax = plt.xlim()
x = np.linspace(xmin, xmax, 100)
p = pareto.pdf(x, shape, loc=loc, scale=scale)
plt.plot(x, p, 'k', linewidth=2, label='Fitted Pareto Distribution')

# Show the plot
plt.xlabel('Net Worth')
plt.y bnblabel('Probability')
plt.title('Net Worth Probability Distribution')
plt.legend()
plt.grid(True)
plt.show()

<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6c446c9de49ec808/6a17d76d2f4a5c686bfa876a/31a715f9d4881c6662c67fe72178635b5836a033-1440x942.png" alt="Pareto" /><h4>Breaking point</h4><p>Finally, with the calculated probability, we determine the target net worth corresponding to the bus hit probability and visualize it. Remember, we use the magic number ChatGPT gave us for the probability of getting hit by a bus:</p>
# Find the Net Worth Corresponding to the Bus Hit Probability
target_probability = 0.0000181
cumulative_probability = all_data_df['net_worth_probability'].cumsum()
target_net_worth_df = all_data_df[cumulative_probability &gt;= target_probability].head(1)
target_net_worth = target_net_worth_df['net_worth'].iloc[0]
print(f"Net Worth with Probability &gt;= {target_probability}: {target_net_worth}")

# Plot the Net Worth Probability Distribution
plt.figure(figsize=(10, 6))
plt.hist(all_data_df['net_worth'], bins=100, density=True, alpha=0.6, color='g', label='Empirical Data')
xmin, xmax = plt.xlim()
x = np.linspace(xmin, xmax, 100)
p = pareto.pdf(x, shape, loc=loc, scale=scale)
plt.plot(x, p, 'k', linewidth=2, label='Fitted Pareto Distribution')
plt.axhline(y=target_probability, color='r', linestyle='--', label='Bus Hit Probability')
plt.axvline(x=target_net_worth, color='g', linestyle='--', label=f'Net Worth = {target_net_worth:.2f}')
plt.xlabel('Net Worth')
plt.ylabel('Probability')
plt.title('Net Worth Probability Distribution')
plt.legend()
plt.grid(True)
plt.show()

<h2>Conclusion</h2><p>Based on our synthetic dataset, this chart vividly illustrates that the probability of amassing a net worth of approximately $12.5 million is as rare as the chance of being hit by a bus. For the fun of it, let’s ask ChatGPT what the probability is:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt224d185dc880bf21/6a17d76f7f6f1581edc09989/077e13f5ebce5019d00b7374ad6fd22dbcf7fe0b-1440x925.png" alt="Probaility Distribution" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7d0ab89f76b8b2e2/6a17d77063baff5bf1741ac3/4aaa32cd60d59770137ae5a3fb582675e606bbd8-1440x210.png" alt="Net worth" /><p>Okay… $439 million? I think ChatGPT might be hallucinating again.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-query-language-esql-statistical-analysis</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-query-language-esql-statistical-analysis</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Python]]></category>
    <dc:creator><![CDATA[Baha Azarmi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1acb1d84387a6310/6a17d7726df7314a250a0d48/274867ef7971390c5d1d4f535c76e50a9f4a8224-1206x1522.png" length="0" type="image/png"/>
    <pubDate>Tue, 20 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch geospatial search with ES|QL]]></title>
    <description><![CDATA[Geospatial search in Elasticsearch Query Language (ES|QL). Elasticsearch has powerful geospatial search features, which are now coming to ES|QL for dramatically improved ease of use and OGC familiarity.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch has had powerful <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/geospatial-analysis.html">geospatial search and analytics capabilities</a> for many years, but the API was quite different from what typical GIS users were used to. In the past year we've <a href="https://www.elastic.co/search-labs/blog/esql-piped-query-language-goes-ga">added the ES|QL query language</a>, a piped query language as easy, or even easier, than SQL. It's particularly suited to the search, security, and observability use cases Elastic excels at. We're also adding support for geospatial search and analytics within ES|QL, making it far easier to use, especially for users coming from SQL or <a href="https://en.wikipedia.org/wiki/Geographic_information_system">GIS</a> communities.</p><p>Elasticsearch 8.12 and 8.13 brought basic support for geospatial types to ES|QL. This was dramatically enhanced with the addition of geospatial search capabilities in 8.14. More importantly, this support was designed to conform closely to the <a href="https://en.wikipedia.org/wiki/Simple_Features">Simple Feature Access</a> standard from the <a href="https://en.wikipedia.org/wiki/Open_Geospatial_Consortium">Open Geospatial Consortium (OGC)</a> used by other spatial databases like PostGIS, making it much easier to use for GIS experts familiar with these standards.</p><p>In this blog, we'll show you how to use ES|QL to perform geospatial searches, and how it compares to the SQL and Query DSL equivalents. We'll also show you how to use ES|QL to perform spatial joins, and how to visualize the results in Kibana Maps. Note that all the features described here are in "technical preview", and we'd love to hear your feedback on how we can improve them.</p><h2>Searching for geospatial data</h2><p>Let's start with an example query:</p>FROM airport_city_boundaries
| WHERE ST_INTERSECTS(
      city_boundary,
      "POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))"::geo_shape
  )
| KEEP abbrev, airport, region, city, city_location
<p>This performs a search for any city boundary polygons that intersect with a rectangular search polygon around the Sanya Phoenix International Airport (SYX).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3897df6bed5d6061/6a17d7c2abe0f29eccdfe861/e48bac8f246c8842f2ea97ddd54910045262aeb1-1440x808.png" alt="ESQL Geospatial Search" /><p>In a sample dataset of airports, cities and city boundaries, this search finds the intersecting polygon and returns the desired fields from the matching document:</p><p>abbrev</p><p>airport</p><p>region</p><p>city</p><p>city_location</p><p>SYX</p><p>Sanya Phoenix Int'l</p><p>天涯区</p><p>Sanya</p><p>POINT(109.5036 18.2533)</p><p>That was easy! Now compare this to the classic Elasticsearch Query DSL for the same query:</p>GET /airport_city_boundaries/_search
{
  "_source": ["abbrev", "airport", "region", "city", "city_location"],
  "query": {
    "geo_shape": {
      "city_boundary": {
        "shape": {
          "type": "polygon",
          "coordinates" : [[
            [109.4, 18.1],
            [109.6, 18.1],
            [109.6, 18.3],
            [109.4, 18.3],
            [109.4, 18.1]
          ]]
        }
      }
    }
  }
}
<p>Both queries are reasonably clear in their intent, but the ES|QL query closely resembles SQL. The same query in PostGIS looks like this:</p>SELECT abbrev, airport, region, city, city_location
FROM airport_city_boundaries
WHERE ST_INTERSECTS(
    city_boundary,
    'SRID=4326;POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))'::geometry
);
<p>Look back at the ES|QL example. So similar, right?</p>FROM airport_city_boundaries
| WHERE ST_INTERSECTS(
      city_boundary,
      "POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))"::geo_shape
  )
| KEEP abbrev, airport, region, city, city_location
<p>We've found that existing users of the Elasticsearch API find ES|QL much easier to use. We now expect that existing SQL users, particularly Spatial SQL users, will find that ES|QL feels very familiar to what they are used to seeing.</p><h4>Why not SQL?</h4><p>What about Elasticsearch SQL? It has been around for a while and has some geospatial features. However, Elasticsearch SQL was written as a wrapper on top of the original Query API, which meant only queries that could be transpiled down to the original API were supported. ES|QL does not have this limitation. Being a completely new stack allows for many optimizations that were not possible in SQL. Our benchmarks show ES|QL is <a href="https://elasticsearch-benchmarks.elastic.co/#tracks/esql/nightly/default/6M">very often faster than the Query API</a>, particularly with aggregations!</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt18bda964c8b24e36/6a17d7c3e3179155d22d568a/b8b6c2b2e45850d832805ed1e71e522f4955f53c-1440x813.png" alt="polygon-intersection-benchmark" /><h2>Differences to SQL</h2><p>Clearly, from the previous example, ES|QL is somewhat similar to SQL, but there are some important differences. For example, ES|QL is a piped query language, starting with a source command like FROM and then chaining all subsequent commands together with the pipe | character. This makes it very easy to understand how each command receives a table of data and performs some action on that table, such as filtering with <code>WHERE</code>, adding columns with <code>EVAL</code>, or performing aggregations with <code>STATS</code>. Rather than starting with <code>SELECT</code> to define the final output columns, there can be one or more <code>KEEP</code> commands, with the last one specifying the final output results. This structure simplifies reasoning about the query.</p><p>Focusing in on the <code>WHERE</code> command in the above example, we can see it looks quite similar to the PostGIS example:</p><p><em>ES|QL</em></p>WHERE ST_INTERSECTS(
    city_boundary,
    "POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))"::geo_shape
)
<p><em>PostGIS</em></p>WHERE ST_INTERSECTS(
    city_boundary,
    'SRID=4326;POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))'::geometry
)
<p>Aside from the difference in string quotation characters, the biggest difference is in how we type-cast the string to a spatial type. In PostGIS, we use the <code>::geometry</code> suffix, while in ES|QL, we use the <code>::geo_shape</code> suffix. This is because ES|QL runs within Elasticsearch, and the type-casting operator <code>::</code> can be used to convert a string to any of the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-limitations.html#_supported_types">supported ES|QL types</a>, in this case, a <code>geo_shape</code>. Additionally, the <code>geo_shape</code> and <code>geo_point</code> types in Elasticsearch imply the spatial coordinate system known as WGS84, more commonly referred to using the SRID number 4326. In PostGIS, this needs to be explicit, hence the use of the <code>SRID=4326;</code> prefix to the WKT string. If that prefix is removed, the SRID will be set to 0, which is more like the Elasticsearch types <code>cartesian_point</code> and <code>cartesian_shape</code>, which are not tied to any specific coordinate system.</p><p>Both ES|QL and PostGIS provide type conversion function syntax as well:</p><p><em>ES|QL</em></p>WHERE ST_INTERSECTS(
    city_boundary,
    TO_GEOSHAPE("POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))")
)
<p><em>PostGIS</em></p>WHERE ST_INTERSECTS(
    city_boundary,
    ST_SetSRID(
      ST_GeomFromText('POLYGON((109.4 18.1, 109.6 18.1, 109.6 18.3, 109.4 18.3, 109.4 18.1))'),
      4326
    )
)
<h2>OGC functions</h2><p>Elasticsearch 8.14 introduces the following four OGC spatial search functions:</p><p>ES|QL</p><p>PostGIS</p><p>Description</p><p>ST_INTERSECTS</p><p>ST_Intersects</p><p>Returns true if two geometries intersect, and false otherwise.</p><p>ST_DISJOINT</p><p>ST_Disjoint</p><p>Returns true if two geometries do not intersect, and false otherwise. The inverse of ST_INTERSECTS.</p><p>ST_CONTAINS</p><p>ST_Contains</p><p>Returns true if one geometry contains another, and false otherwise.</p><p>ST_WITHIN</p><p>ST_Within</p><p>Returns true if one geometry is within another, and false otherwise. The inverse of ST_CONTAINS.</p><p>These function behave similarly to their PostGIS counterparts, and are used in the same way. For example, <code>ST_INTERSECTS</code> returns true if two geometries intersect and false otherwise. If you follow the documentation links in the above table, you might notice that all the ES|QL examples are within a <code>WHERE</code> clause after a <code>FROM</code> clause, while all the PostGIS examples are using literal geometries. In fact, both platforms support using the functions in any part of the query where they make sense.</p><p>The first example in the PostGIS documentation for <code>ST_INTERSECTS</code> is:</p>SELECT ST_Intersects(
    'POINT(0 0)'::geometry,
    'LINESTRING ( 2 0, 0 2 )'::geometry
);
<p>The ES|QL equivalent of this would be:</p>ROW ST_INTERSECTS(
    "POINT(0 0)"::geo_point,
    "LINESTRING ( 2 0, 0 2 )"::geo_shape
)
<p>Note how we did not specify the SRID in the PostGIS example. This is because in PostGIS when using the <code>geometry</code> type, all calculations are done on a planar coordinate system, and so if both geometries have the same SRID, it does not matter what the SRID is. In Elasticsearch, this is also true for most functions, however, there are exceptions where <code>geo_shape</code> and <code>geo_point</code> use spherical calculations, as we'll see in the next blog about spatial distance search.</p><h2>ES|QL versatility</h2><p>So, we've seen examples above for using spatial functions in <code>WHERE</code> clauses, and in <code>ROW</code> commands. Where else would they make sense? One very useful place is in the <code>EVAL</code> command. This command allows you to evaluate an expression and return the result. For example, let's determine if the centroids of all airports grouped by their country names are within a boundary outlining the country:</p>FROM airports
| EVAL in_uk = ST_INTERSECTS(location, TO_GEOSHAPE("POLYGON((1.2305 60.8449, -1.582 61.6899, -10.7227 58.4017, -7.1191 55.3291, -7.9102 54.2139, -5.4492 54.0078, -5.2734 52.3756, -7.8223 49.6676, -5.0977 49.2678, 0.9668 50.5134, 2.5488 52.1065, 2.6367 54.0078, -0.9668 56.4625, 1.2305 60.8449))"))
| EVAL in_iceland = ST_INTERSECTS(location, TO_GEOSHAPE("POLYGON ((-25.4883 65.5312, -23.4668 66.7746, -18.4131 67.4749, -13.0957 66.2669, -12.3926 64.4159, -20.1270 62.7346, -24.7852 63.3718, -25.4883 65.5312))"))
| EVAL within_uk = ST_WITHIN(location, TO_GEOSHAPE("POLYGON((1.2305 60.8449, -1.582 61.6899, -10.7227 58.4017, -7.1191 55.3291, -7.9102 54.2139, -5.4492 54.0078, -5.2734 52.3756, -7.8223 49.6676, -5.0977 49.2678, 0.9668 50.5134, 2.5488 52.1065, 2.6367 54.0078, -0.9668 56.4625, 1.2305 60.8449))"))
| EVAL within_iceland = ST_WITHIN(location, TO_GEOSHAPE("POLYGON ((-25.4883 65.5312, -23.4668 66.7746, -18.4131 67.4749, -13.0957 66.2669, -12.3926 64.4159, -20.1270 62.7346, -24.7852 63.3718, -25.4883 65.5312))"))
| STATS centroid = ST_CENTROID_AGG(location), count=COUNT() BY in_uk, in_iceland, within_uk, within_iceland
| SORT count ASC
<p>The results are expected, the centroid of UK airports are within the UK boundary, and not within the Iceland boundary, and vice versa:</p><p>centroid</p><p>count</p><p>in_uk</p><p>in_iceland</p><p>within_uk</p><p>within_iceland</p><p>POINT (-21.946634463965893 64.13187285885215)</p><p>1</p><p>false</p><p>true</p><p>false</p><p>true</p><p>POINT (-2.597342072712148 54.33551226578214)</p><p>17</p><p>true</p><p>false</p><p>true</p><p>false</p><p>POINT (0.04453958108176276 23.74658354606057)</p><p>873</p><p>false</p><p>false</p><p>false</p><p>false</p><p>In fact, these functions can be used in any part of the query where their signature makes sense. They all take two arguments, which are either a literal spatial object or a field of a spatial type, and they all return a boolean value. One important consideration is that the coordinate reference system (CRS) of the geometries must match, or an error will be returned. This means you cannot mix <code>geo_shape</code> and <code>cartesian_shape</code> types in the same function call. You can, however, mix <code>geo_point</code> and <code>geo_shape</code> types, as the <code>geo_point</code> type is a special case of the <code>geo_shape</code> type, and both share the same coordinate reference system. The documentation for each of the functions defined above lists the supported type combinations.</p><p>Additionally, either argument can be a spatial literal or a field, in either order. You can even specify two fields, two literals, a field and a literal, or a literal and a field. The only requirement is that the types are compatible. For example, this query compares two fields in the same index:</p>FROM airport_city_boundaries
| EVAL in_city = ST_INTERSECTS(city_location, city_boundary)
| STATS count=COUNT(*) BY in_city
| SORT count ASC
| EVAL cardinality = CASE(count &lt; 10, "very few", count &lt; 100, "few", "many")
| KEEP cardinality, count, in_city
<p>The query basically asks if the city location is within the city boundary, which should generally be true, but there are always exceptions:</p><p>cardinality</p><p>count</p><p>in_city</p><p>few</p><p>29</p><p>false</p><p>many</p><p>740</p><p>true</p><p>A far more interesting question would be whether the airport location is within the boundary of the city that the airport serves. However, the airport location resides in a different index than the one containing the city boundaries. This requires a method to effectively query and correlate data from these two separate indexes.</p><h2>Spatial joins</h2><p>ES|QL does not support <code>JOIN</code> commands, but you can achieve a special case of a join using the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-commands.html#esql-enrich"><code>ENRICH</code></a><a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/esql-commands.html#esql-enrich"> command</a>, which behaves similarly to a 'left join' in SQL. This command operates akin to a 'left join' in SQL, allowing you to enrich results from one index with data from another index based on a spatial relationship between the two datasets.</p><p>For example, let's enrich the results from a table of airports with additional information about the city they serve by finding the city boundary that contains the airport location, and then perform some statistics on the results:</p>FROM airports
| ENRICH city_boundaries ON city_location WITH airport, region, city_boundary
| MV_EXPAND city_boundary
| EVAL boundary_wkt_length = LENGTH(TO_STRING(city_boundary))
| STATS centroid = ST_CENTROID_AGG(location), count = COUNT(city_location), min_wkt = MIN(boundary_wkt_length), max_wkt = MAX(boundary_wkt_length) BY region
| SORT count DESC
| LIMIT 5
<p>This returns the top 5 regions with the most airports, along with the centroid of all the airports that have matching regions, and the range in length of the WKT representation of the city boundaries within those regions:</p><p>centroid</p><p>count</p><p>min_wkt</p><p>max_wkt</p><p>region</p><p>POINT (-32.56093470960719 32.598117914802714)</p><p>90</p><p>207</p><p>207</p><p>null</p><p>POINT (-73.94515332765877 40.70366442203522)</p><p>9</p><p>438</p><p>438</p><p>City of New York</p><p>POINT (-83.10398317873478 42.300230911932886)</p><p>9</p><p>473</p><p>473</p><p>Detroit</p><p>POINT (-156.3020245861262 20.176383580081165)</p><p>5</p><p>307</p><p>803</p><p>Hawaii</p><p>POINT (-73.88902732171118 45.57078813901171)</p><p>4</p><p>837</p><p>837</p><p>Montréal</p><p>So, what really happened here? Where did the supposed <code>JOIN</code> occur? The crux of the query lies in the <code>ENRICH</code> command:</p>FROM airports
| ENRICH city_boundaries ON city_location WITH airport, region, city_boundary
<p>This command instructs Elasticsearch to enrich the results retrieved from the <code>airports</code> index, and perform an <code>intersects</code> join between the <code>city_location</code> field of the original index, and the <code>city_boundary</code> field of the <code>airport_city_boundaries</code> index, which we used in a few examples earlier. But some of this information is not clearly visible in this query. What we do see is the name of an enrich policy <code>city_boundaries</code>, and the missing information is encapsulated within that policy definition.</p>{
  "geo_match": {
    "indices": "airport_city_boundaries",
    "match_field": "city_boundary",
    "enrich_fields": ["city", "airport", "region", "city_boundary"]
  }
}
<p>Here we can see that it will perform a <code>geo_match</code> query (<code>intersects</code> is the default), the field to match against is <code>city_boundary</code>, and the <code>enrich_fields</code> are the fields we want to add to the original document. One of those fields, the <code>region</code> was actually used as the grouping key for the <code>STATS</code> command, something we could not have done without this 'left join' capability. For more information on enrich policies, see the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/enrich-setup.html">enrich documentation</a>. While reading those documents, you will notice that they describe using the enrich indexes for enriching data at index time, by configuring ingest pipelines. This is not required for ES|QL, as the <code>ENRICH</code> command works at query time. It is sufficient to prepare the enrich index with the necessary data and enrich policy, and then use the <code>ENRICH</code> command in your ES|QL queries.</p><p>You may also notice that the most commonly found region was <code>null</code>. What could this imply? Recall that I likened this command to a 'left join' in SQL, meaning if no matching city boundary is found for an airport, the airport is still returned but with <code>null</code> values for the fields from the <code>airport_city_boundaries</code> index. It turns out there were 89 airports that found no matching <code>city_boundary</code>, and one airport with a match where the <code>region</code> field was <code>null</code>. This lead to a count of 90 airports with no <code>region</code> in the results. Another interesting detail is the need for the <code>MV_EXPAND</code> command. This is necessary because the <code>ENRICH</code> command may return multiple results for each input row, and <code>MV_EXPAND</code> helps to separate these results into multiple rows, one for each outcome. This also clarifies why "Hawaii" shows different <code>min_wkt</code> and <code>max_wkt</code> results: there were multiple regions with the same name but different boundaries.</p><h2>Kibana Maps</h2><p>Kibana has added support for Spatial ES|QL in the Maps application. This means that you can now use ES|QL to search for geospatial data in Elasticsearch, and visualize the results on a map.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt91922eb4ef43d290/6a16f7161949f737bfe7a7b5/bd78470bd8a4bc60f0db7006bd804b8fe87e2fea-1440x683.png" alt="Kibana Layers ES|QL" /><p>There is a new layer option in the add layers menu, called "ES|QL". Like all of the geospatial features described so far, this is in "technical preview". Selecting this option allows you to add a layer to the map based on the results of an ES|QL query. For example, you could add a layer to the map that shows all the airports in the world.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5185ef7461d7e81d/6a16f718839dfa1559dcfcb1/1dd28d3d0509f92d26b0bb5320a2925f7a54c5d9-1440x736.png" alt="Kibana ES|QL - Airports" /><p>Or you could add a layer that shows the polygons from the <code>airport_city_boundaries</code> index, or even better, how about that complex <code>ENRICH</code> query above that generates statistics for how many airports are in each region?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt51ee3543bb811d98/6a16f71a8b73cbbe63189df1/679a0a401faa613c7cedddd07c64f614ac2b7144-1440x727.png" alt="Kibana ES|QL - Region Statistics" /><h2>What's next</h2><p>You might have noticed in two of the examples above we squeezed in yet another spatial function <code>ST_CENTROID_AGG</code>. This is an aggregating function used in the <code>STATS</code> command, and the first of many spatial analytics features we plan to add to ES|QL. We'll blog about it when we've got more to show!</p><p>Before that, we want to tell you more about a particularly exciting feature we've worked on: the ability to perform spatial distance searches, one of the most used spatial search features of Elasticsearch. Can you imagine what the syntax for distance searches might look like? Perhaps similar to an OGC function? Stay tuned for the next blog in this series to find out!</p><p>Spoiler alert: Elasticsearch 8.15 has just been released, and spatial distance search with ES|QL is included!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-geospatial-search-part-one</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-geospatial-search-part-one</guid>
    <category><![CDATA[Python]]></category>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Craig Taverner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd05627be20e89dfb/6a17d7c6414c640256944fdb/de886289dcb56494920875303b622b030b9b810f-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Mon, 12 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>